Full text
applied
sciences
A icle
Mul i-F ame Labeled Faces Da abase: Towa ds Face
Supe -Resolu ion om Realis ic Video Sequences
Ma in Rajnoha * , Anzhelika Mezina and Radim Bu ge *
Depa men o Telecommunica ions, B no Uni e si y o Technology, 616 00 B no, Czech Republic;
xmezin00@ u b .cz
*Co espondence: ma in. ajnoha@ u b .cz (M.R.); [email p o ec ed].cz (R.B.)
Recei ed: 25 Augus 2020; Accep ed: 10 Oc obe 2020; Published: 16 Oc obe 2020
Fea u ed Applica ion: Du ing police sea ches o pe pe a o s o se ious c imes, o en only a low
de ini ion ideo is a ailable. The p oposed me hodology can use a sequence o aces om a ideo
o econs uc a single high de ini ion ace. Biome ic iden i ica ions a e used wo ldwide, bu he e
is s ill space o imp o emen s in hei accu acy o wo king unde bad en i onmen al condi ions.
Abs ac :
Fo ensically ained acial e iewe s a e s ill conside ed as one o he mos accu a e
app oaches o pe son iden i ica ion om ideo eco ds. The human b ain can u ilize in o ma ion,
no jus om a single image, bu also om a sequence o images (i.e., ideos), and e en in he
case o low-quali y eco ds o a long dis ance om a came a, i can accu a ely iden i y a gi en
pe son. Un o una ely, in many cases, a single s ill image is needed. An example o such a case is a
police sea ch ha is abou o be announced in newspape s. This pape in oduces a ace da abase
ob ained om eal en i onmen coun ing in 17,426 sequences o images. The da ase includes
pe sons o a ious aces and ages and also di e en en i onmen s, di e en ligh ing condi ions o
came a de ice ypes. This pape also in oduces a new mul i- ame ace supe - esolu ion me hod
and compa es his me hod wi h he s a e-o - he-a single- ame and mul i- ame supe - esolu ion
me hods. We p o e ha he p oposed me hod inc eases he quali y o ace images, e en in cases o
low- esolu ion low-quali y inpu images, and p o ides be e esul s han single- ame app oaches
ha a e s ill conside ed he bes in his a ea. Quali y o ace images was e alua ed using se e al
objec i e ma hema ical me hods, and also subjec i e ones, by se e al olun ee s. The sou ce code
and he da ase we e eleased and he expe imen is ully ep oducible.
Keywo ds:
ace ecogni ion; supe esolu ion; mul i ame; image p ocessing; da abase; da ase ;
sequences; deep lea ning
1. In oduc ion
Closed-ci cui ele ision (CCTV) is a widely used echnology o moni o ing public and p i a e
places, which helps o inc ease o e all sa e y, p e en auds, shopli ing, he s, bu gla ies, andalism,
e o ism and o he s. I is o en ins alled in public a eas and businesses h oughou he wo ld o
p e en he abo emen ioned c imes and o inc ease o e all public sa e y. Un o una ely, he e a e
also many abuses and o he ela ed ques ions mainly conce ning p i acy issues. In ela ion o police
sea ches and he in es iga ion o se ious c imes, he e is no doub ha hey ha e a e y posi i e impac
on he cla i ica ion o c iminal o ences and ea ly de en ion o c ime o ende s [1].
In p ac ice, i o en happens ha e en hough he c ime is eco ded using CCTV, he eco ding
canno be used o police in es iga ion o a police sea ch and canno be used as e idence. The cause
o his is o en he poo quali y o he eco ding, he dis ance o he came a om he c ime place,
bad ligh ing, bad wea he condi ions, o o he s ac o s which cause he pe son on he eco ding o be
Appl. Sci. 2020,10, 7213; doi:10.3390/app10207213 www.mdpi.com/jou nal/applsci
Appl. Sci. 2020,10, 7213 2 o 27
unambiguously iden i ied. Such a eco d canno e en be used by he police o ask he public o help
in assis ance wi h acing he pe pe a o , o ins ance in newspape s o online media.
Mos o he compu e ace ecogni ion sys ems usually use a single s ill image and based on i
ecognize a ace. The human expe s— o ensically ained acial e iewe s—a e s ill conside ed one
o he mos accu a e app oaches o pe son iden i ica ion. They a e gene ally mo e esilien o noise
in images, and also mo e e icien in u ilizing he in o ma ion con ained in he ideo, e en when he
quali y o images a e poo . They can ake in o accoun he dynamics o walk, ges u es, i ness s a us o
a pe son, and possibly o he cha ac e is ics. They can also u ilize he in o ma ion om a sequence o
ace images and can be mo e accu a e when compa ed o ecogni ion om each isola ed image.
This pape ocuses on inc easing he quali y o he acial image using a ideo sequence wi h
emphasis on biome ic u iliza ion whe e he p inciple is illus a ed (see Figu e 1). The inpu ace image
is eco ded om a long dis ance (i.e., i s esolu ion is low and hus a single acial image canno be used
o pe son iden i ica ion). This pa icula example demons a es he di icul y o pe son iden i ica ion
om a long dis ance, which is qui e equen o su eillance came as. Inspi ed by human skills,
his pape in es iga es me hods ha can econs uc a high- esolu ion ace image om a se ies o
low- esolu ion and poo quali y images.
Figu e 1.
This pape is inspi ed by he human capaci y o ecognize a ace om a sequence o images
be e han om a s ill pic u e.
The esea ch desc ibed in his pape in oduces a couple o con ibu ions, whose no el y o
imp o emen a e b ie ly desc ibed by he ollowing poin s o be e cla i y:
•
This pape in oduces a new me hodology o supe - esolu ion (SR) di e en om he
gene al-pu pose one, which is limi ed o aces and biome ic applica ions.
•
The sequence o mul iple images is p ocessed (mul i- ame me hods) ins ead o using jus a single
image as mos exis ing da ase s allow (single- ame me hods).
•
We p oposed a me hodology ha was compa ed o o he s a e-o - he-a me hods, cu en ly, he
single- ame supe - esolu ion me hods can be conside ed as one o he bes because o he lack o
he p og ess and non-u iliza ion o deep lea ning in mul i- ame supe - esolu ion me hods [
2
,
3
].
Mo eo e , gene al supe - esolu ion me hods a e o en no applicable o pu poses o biome ics,
e en i hey p o ide good esolu ion.
•
A new unique la ge-scale da ase , con aining 17,426 samples o image sequences aken
om a wide ange o eco ding de ices, is p o ided. They con ain aces o a ious ages,
aces, illumina ion condi ions ha a e aken om eal-wo ld en i onmen s.
•
The pape also de ines me ics o acial quali y measu emen s ha a e ep oducible and based
on open-sou ce solu ions. These objec i e me ics a e widely accessible app oaches and can boos
u he esea ch in his a ea.
•
The whole expe imen desc ibed in his pape is ully ep oducible— he sou ce codes and he
da ase we e eleased online (h p://splab.cz/ml db/#download).
•
The combina ion o he key poin s abo e c ea es a g ound (basis) and i has he po en ial o
c ea ing a s anda d o u u e esea ch in his ield.
The es o he pape is s uc u ed as ollows. Sec ion 2p o ides ela ed wo k ega ding
single- ame and mul i- ame supe - esolu ion me hods. Sec ion 3desc ibes he expe imen which
Appl. Sci. 2020,10, 7213 3 o 27
has been conduc ed. I includes he desc ip ion o he c ea ed da ase and a ious p oposed
me hodologies o ace ecogni ion and e alua ion me ics. Resul s achie ed using ou me hodology
a e shown in Sec ion 4, and Sec ion 5discusses he esul s we achie ed and on how o in e p e hem.
Finally, Sec ion 6concludes he pape .
2. Rela ed Wo k
Face iden i y ecogni ion can usually be done manually by o ensically ained acial e iewe s
o au oma ically, usually by a compu e expe sys em. The au oma ed ace iden i y ecogni ion
is a esea ch a ea which has expe ienced a apid g ow h in e ms o accu acy and eliabili y in
ecen yea s. This success is achie ed mainly hanks o he u iliza ion o he so-called deep neu al
ne wo ks [4]
, highe compu a ional powe , and also he p esence o la ge aining and e alua ion
da ase s such as Labeled Faces in he Wild (LFW) (13,233 images) [
5
], PubFig (58,797 images) [
6
],
FERET (14,126 images) [
7
], CelebA (200 k images, 10,177 iden i ies) [
8
] o You ube aces DB
(1595 iden i ies) [
9
]. These da ase s con ain ace images wi h a ious illumina ion condi ions,
exp essions, aces, and poses. Al hough he e has been g ea p og ess in ace iden i y ecogni ion and
i s accu acy in ecen yea s, he econs uc ion o an image om a low- esolu ion one s ill emains
a challenge [
4
]. The e is cu en ly no la ge-scale ace da ase con aining se e al sequences o images.
Many o he cu en supe - esolu ion me hods ha e a gene a i e na u e and,
he e o e, canno be
used
o iden i ica ion pu poses since hey a e oo c ea i e.
The app oaches, which add esses he p oblem o inc easing esolu ion, can be di ided by
bo h he con en ype and inpu da a ype. The me hods based on he con en ype can be
di ided in o gene al supe - esolu ion echniques and biome y- ocused supe - esolu ion echniques.
F om he poin o iew o he inpu da a ype, he me hods can be di ided in o wo g oups:
wi h a single inpu image ( he so-called single- ame) o wi h many inpu images ( he so-called
mul i- ame,
i.e., ideo). Supe - esolu ion,
mainly hanks o he success o deep lea ning, has made
signi ican p og ess in ecen yea s [
10
]. This is also he eason why he mos success ul me hods o
supe - esolu ion a e mainly based on neu al ne wo ks. In pa icula , hey a e Con olu ional Neu al
Ne wo ks (CNN) o Gene a i e Ad e sa ial Ne wo ks (GANs).
2.1. Gene al Single-F ame Supe -Resolu ion Me hods
Single- ame supe - esolu ion me hods a e me hods ha c ea e a high- esolu ion image om
a single low- esolu ion image. Mos o hese me hods a e gene al pu pose (i.e., hey a e no specialized
o any pa icula applica ion a ea, o example, aces).
In case we do no ake in o accoun he in e pola ion me hods (such as bicubic,
bilinea , nea es neighbou ,
e c.), one o he i s well-known me hods was he Supe -Resolu ion
Con olu ional Neu al Ne wo k (SRCNN) [
11
], which was in oduced in 2015. This neu al ne wo k
(NN) a chi ec u e was ela i ely ligh weigh —i had only h ee laye s, which is a ela i ely low numbe
when conside ing he cu en s a e-o - he-a a chi ec u es. Howe e , i was shown ha e en such
a simple NN ou pe o ms signi ican ly he capabili ies o any o he in e pola ion me hod.
Ano he signi ican p og ess in he a ea o supe - esolu ion was based on he so-called gene a i e
ad e sa ial ne wo ks. Al hough GANs we e in oduced in 2014, hey showed hei applicabili y in
p ac ice o supe - esolu ion only in 2017. One o he well-known p ojec s was Pix2Pix [
12
]. I was
o iginally in oduced as an image- o-image ansla ion model ( he ne wo k has he same dimensions
o he inpu and ou pu laye s). Especially when compa ed o in e pola ion-based me hods, Pix2Pix
has also shown a signi ican imp o emen [
12
]. SRGAN [
13
] is ano he success ul app oach based
on GANs. I s inno a ion was o use he esidual a chi ec u e and imp o emen ega ding he loss
unc ion—in pa icula , i used he pe cep ual loss unc ion du ing aining ins ead o he pixel-wise
loss, such as Mean Squa e E o (MSE). Ano he success ul a chi ec u e was EDSR,
which was
in oduced in [
14
] and i s main idea was o ain high-scale models om p e- ained low-scale models.
An in e es ing inno a ion was o sha e pa ame e s ac oss di e en scales. ESRGAN [
15
] is inspi ed by
Appl. Sci. 2020,10, 7213 4 o 27
SRGAN. The inno a ion o ESRGAN is o emo e ba ch no maliza ion laye s,
he modi ied
pe cep ual
loss unc ion, and ans e lea ning. The disc imina o pa o GAN was eplaced by a ela i is ic
disc imina o [16].
In 2019, mainly due o using huge compu a ional powe and la ge GPU memo y, i was shown
ha he NNs can gene a e e en ela i ely complex images wi h a lo o de ail [
17
]. Al hough hese
ne wo ks ha e achie ed g ea success, hei big disad an age is he ac ha hey do no y o ai h ully
econs uc he in o ma ion ha is on he inpu , bu hey gene a e only pa o he in o ma ion. Fo his
eason, hey a e no sui able o biome ics, ace econs uc ion, and police in es iga ion, as hey end
o gene a e o ally new aces. This is un o una ely a p oblem o mos o he GAN based ne wo ks.
2.2. Single-F ame Supe -Resolu ion Focused on Faces
Supe - esolu ion me hods men ioned in he p e ious sec ion we e designed o gene al use.
This sec ion desc ibes a sub-se o supe - esolu ion me hods ha di ec ly ocus on ace supe - esolu ion
ela ed p oblems, and in connec ion wi h biome ics oo. Al hough many o he me hods men ioned
ea lie o e images in eally high- esolu ion (o en 2
×
, 4
×
, o also 8
×
zoom) and he human eye o en
canno ecognize ha he ace was c ea ed om he o iginal low- esolu ion image, hese images canno
be used o pe son iden i ica ion. The images a e syn hesized om pas expe ience (images used as a
aining da ase o NN) and o en gene a es ake aces. The la es comp ehensi e su ey pape ela ed
o biome ics and supe - esolu ion echniques is p o ided in [
18
]. Among o he s, his pape also
discusses in-dep h di e en equi emen s ega ding gene al supe - esolu ion me hods and biome ics
(o ace iden i y ecogni ion) me hods.
When we jus ocus on me hods ela ed o acial supe - esolu ion, one o he s a e-o - he-a
me hods is he Ul a-Resolu ion by Disc imina i e Gene a i e Ne wo ks (UR-DGN). This a chi ec u e
is based on GAN and allows us o upscale images wi h scale ac o 8 [
19
]. The au ho s used
a decon olu ional ne wo k as he gene a o and a con olu ional ne wo k as he disc imina o .
Ano he app oach in his ield is he end- o-end ainable Face Supe -Resolu ion Ne wo k (FSRNe ),
which is based on CNN, and FSRGAN, which eco e s mo e ealis ic ex u es han FSRNe [
20
].
These a chi ec u es u ilize he es ima ion o landma k hea maps. One o he la es wo ks is also he
P og essi e Face Supe -Resolu ion ia A en ion o Facial Landma k [
21
]. This me hod is also based
on GAN, howe e , i uses he p og essi e me hod o upscaling he image. Mo eo e , his wo k
in oduces a new acial a en ion loss, which allows us o es o e acial landma ks.
A ecen ly in oduced me hod PULSE [
22
], which was p esen ed a he CVPR 2020 h p:
//c p 2020. hec .com/), has e y encou aging esul s. Un o una ely, some c i icism ( o example,
Twi e pos (h ps:// wi e .com/AlexVasilescu/s a us/1277009168319143936): “
PULSE: Medical
and mili a y decision based on hallucina ed pixels?”) o he me hod show he disad an ages o he
PULSE, and in gene al al eady men ioned lack o GANs—i s c ea i eness. The e is a demons a ion
which uses a low- esolu ion image o Ba ack Obama and i s SR al e na i e p ocessed by his me hod.
I c ea ed a high-quali y ace image wi h no ob ious signs ha his ace was econs uc ed by
supe - esolu ion. Un o una ely, i econs uc ed a o ally di e en pe son.
2.3. Mul i-F ame Supe -Resolu ion
Supe - esolu ion has a ac ed g ea a en ion, especially in he a ea o he en e ainmen indus y
he ideo, o in o he wo ds mul i- ame. The e ha e been in oduced plen y o wo ks which in oduce
me hods based no jus on a single inpu image, bu on a sequence o images. These me hods ha e he
po en ial o educe noise and inc ease he quali y o he images, especially in he a ea o con e sion o
old ideos in o HD esolu ion whe e i had g ea success.
One o he la es wo ks ela ed o gene al (non- acial ela ed) mul i- ame supe - esolu ion is
DeepSUM [
2
]. I is based on CNN, which exploi s spa ial and empo al co ela ions. This app oach
includes image egis a ion inside he CNN a chi ec u e. I dynamically compu es and applies
cus om il e s o highe -dimensional image ep esen a ions, ins ead o compensa ing o he mo ion as
Appl. Sci. 2020,10, 7213 5 o 27
a p e-p ocessing s ep as is common in mos mul i- ame supe - esolu ion app oaches. The da ase om
he challenge P oba-V [
23
] was used, which is dealing wi h sa elli e images and alloca es impo an
pa s using segmen a ion based maps.
Ano he me hod o mul i- ame supe - esolu ion is desc ibed in [
24
]. Fi s , all inpu
low- esolu ion ames a e upscaled by he ResNe [
25
] wi h scale ac o 2. Pa allel o ha ,
he low- esolu ion
ames go h ough he image egis a ion block o de e mine he sub-pixel shi s.
The nex s ep is he applica ion o he shi -and-add usion o ob ain he ini ial econs uc ed image.
The inal s ep is he applica ion o he E oIM p ocess, which consis s o i e a i e il e ing o he ini ial
econs uc ed image. This me hod is again gene al pu pose and no acial supe - esolu ion ela ed.
The mos equen u iliza ion o he mul i- ame supe - esolu ion is upscaling he ideo
esolu ion. F ame- ecu en Video Supe - esolu ion [
26
] is an end- o-end ainable ame- ecu en
ideo supe - esolu ion (FRVSR) amewo k. I uses he p e iously es ima ed high- esolu ion ame
as an inpu o i s ollowing s ep. Each ame is p ocessed only once, which allows us o educe
he compu a ional cos . The model consis s o se e al s eps: low es ima ion, upscaling low,
wa ping p e ious ou pu , mapping o low- esolu ion space, supe - esolu ion. I is e y impo an o
keep he p e iously es ima ed high- esolu ion ame in he sys em, o he wise, i is no possible o pass
in o ma ion o u u e es ima ions, which can be c i ical o he econs uc ion o he nex ames.
TecoGAN [
27
] is he a chi ec u e p oposed o sol e he ollowing ideo gene a ion asks:
Video supe - esolu ion (VSR) and Unpai ed Video T ansla ion (UVT). The wo k [
27
] desc ibes he
ad e sa ial lea ning me hod o a ecu en aining app oach, which u ilizes spa ial con en s and
empo al ela ionships. The a chi ec u e uses a ame- ecu en gene a o and a spa io- empo al
disc imina o . This app oach can gene a e e y ealis ic na u al images; howe e , i can lead o
empo ally cohe en ye sub-op imal de ails.
Enhanced De o mable Con olu ional Ne wo ks (EDVR) [
28
] is a amewo k p oposed in
2019, which allows us o do di e en image es o a ion asks, o example, supe - esolu ion
and de-blu ing. I consis s o he ollowing pa s: an alignmen module—Py amid, Cascading
and De o mable con olu ions (PCD), and a usion module—Tempo al and Spa ial A en ion
(TSA).
Mo eo e , he wo-s age
s a egy is used o boos pe o mance and imp o e he quali y o
ou pu ames.
One o he i s a emp s ela ed o he acial supe - esolu ion me hod was in oduced in 2017
in [
29
]. The au ho s p oposed an a chi ec u e ha consis s o h ee modules: ea u e ex ac o ,
ace wa ping, and econs uc ion. This a chi ec u e es o es he cen al ame o each inpu sequence
u ilizing sub-pixel mo emen and aking in o accoun a numbe o adjacen ames. This model was
es ed on he YouTube Faces da ase , which is no dedica ed o supe - esolu ion pu poses. The au ho s
downsampled all images h ough 128
×
128 px, images we e blu ed wi h he Gaussian ke nel (2.4),
and again downsampled o he size 16
×
16 px. The downsampling me hod is in his case no clea and
p obably he same me hod was used o all images, which oge he wi h a cons an blu ,
c ea es a high
bias and does no e lec eal-wo ld scena ios. Fo his pu pose, he e is a isk o o e - i ing.
2.4. Mo i a ion
Mos o he wo ks desc ibed ea lie a e ocused on gene al image supe - esolu ion.
Un o una ely, hose me hods canno o en be used o pe son iden i ica ion since acial
supe - esolu ion has sligh ly di e en equi emen s.
Al hough he gene al supe - esolu ion me hods o en wo ks e y well o gene al images and
some imes hey e en p o ide eally ou s anding esul s, hese me hods s ill c ea e some mis akes in
he images. In he case o gene al images, hese e o s can o en be igno ed. Howe e , a mis ake in
a ace image, whe e an eye o nose is missing o is de o med, e en i i migh be om he poin o pixel
e o a mino e o , om he poin o iew o human pe cep ion, i can be s ongly dis up i e. In his
case, e en a sligh modi ica ion o he image can ha e a signi ican impac on human pe cep ion o he
ace and iden i ica ion o a pe son, as shown in Figu e 2.
Appl. Sci. 2020,10, 7213 6 o 27
Figu e 2.
Di e en ace supe - esolu ion me hods wi h scale ac o 8. F om he human pe cep ion poin
o iew, i is be e o ha e checke boa d a i ac s (U-Ne +GEU) ins ead o a i ac s and de o ma ions
om o he me hods.
Ano he issue may be he c edibili y o he supe - esolu ion me hods om a biome ic pe spec i e.
Fo biome ics, i is absolu ely unaccep able o be c ea i e and gene a e pa s o he image ha do no
o igina e om i s inpu .
The cu en end and di ec ion o esea ch o supe - esolu ion me hods ocus p ima ily
on he gene al single- ame me hods. As he mul i- ame me hods ha e mo e in o ma ion
on i s inpu , i is expec ed hey also p obably ha e a highe po en ial o each be e esul s
ha a e based on he inpu in o ma ion and a e no c ea i ely illing he missing pa s o he
image.
Fu he mo e, conside ing he echnical
capabili ies o mode n CCTV ha dwa e (e.g., ames
independen ly on ps), i is e en possible o go beyond he possibili ies o humans and he e is
an oppo uni y o each e en be e esul s in he u u e. Al hough cogni i e skills o he human
b ain s ill signi ican ly exceed he skills o machines, na owly ocused a i icial in elligence has he
po en ial o b ing new applica ions and, o example, make iden i ica ion mo e objec i e. Examples
o hese applica ions could be an au oma ic sea ch o he bes acial image o a suspec o news
announcemen s, econs uc ion o he ace using a sequence o images, and he supp essing o ad e se
ligh ing condi ions.
One o he obs acles in he de elopmen o he cu en acial mul i- ame supe - esolu ion me hods
is he absence o a publicly a ailable da ase , which would be designed o his pu pose.
Many speci ic equi emen s a e imposed on such a da ase . Da a o such a da ase should no be
biased, i mus be o su icien size, i mus con ain eco ds om eal-wo ld en i onmen s including
di e en ligh ing o wea he condi ions, da a mus be no dependen on a speci ic ideo came a
ha dwa e, and i should e lec also eal-wo ld comp ession a i ac s, he so-called encoding agmen s.
Con ained aces should co e a a ie y o human aces and a ious ages. In each case, he da a mus
con ain a sequence o se e al aces, no jus one ame.
Rega ding a ailable da ase s, he e is cu en ly no da ase ha mee s such equi emen s.
Face images used o be biased because o used celeb i ies, which ends o cause he models o p edic
“p e y” aces supp essing impo an indi idual ace ea u es (LFW, CelebA). The e is a lack o low
quali y image inpu s, especially in he o m o sequences (PubFig, FERET) wi h a high-quali y label
(You ube aces DB). All imp o emen s a e summa ized in Table 1.
Table 1. The imp o emen s o he p oposed da abase in con as o well-known aining da ase s.
Imp o emen LFW FERET PubFig CelebA You ube DB
Numbe o samples × × ×
Numbe o iden i ies × ×
Samples as sequences ×
Supe - esolu ion pu pose
Va iabili y (a oid bias) × ×
3. Ma e ials and Me hods
This sec ion desc ibes he aining and es ing da ase s, examined supe - esolu ion me hods,
and e alua ion me ics. Each pa icula pa is desc ibed in de ail in he ollowing subsec ions.
O e all, he scheme desc ibed in his sec ion and hei mu ual con ex a e depic ed in Figu e 3.
Appl. Sci. 2020,10, 7213 7 o 27
Figu e 3.
Scheme o he expe imen . The da ase consis ing o se e al low- esolu ion images
and one high- esolu ion a ge image was used o aining and e alua ion o a couple o
di e en supe - esolu ion me hods. These me hods we e e alua ed using se e al objec i e and
subjec i e me ics.
3.1. T aining and Tes ing Da a
Fo aining and e alua ion, we in oduce a new da ase called Mul i- ame Labeled Faces
Da abase—MLFDB. The main di e ence o he exis ing da ase s is ha his da ase does no con ain jus
a single acial image, bu i con ains a sequence o consecu i e ames in a ideo eco d. This sequence
o images is in a low- esolu ion (32
×
32 pixels) and se es as an inpu o he supe - esolu ion
me hods. Addi ionally o each o hose sequences, he e was added one ame in a highe esolu ion
(64
×
64 pixels) which se es as a g ound u h, i.e., he op imal ou pu o he supe - esolu ion me hod
(see Figu e 4).
Figu e 4.
An example o sample da a. Se en low- esolu ion images (32
×
32 px) and label (64
×
64 px).
Wha should be emphasized he e is ha he aces in he da ase we e selec ed on he bo de line
whe e iden i ica ion is di icul (i.e., each low- esolu ion ame o 32
×
32 pixels) and when i is easie
o use i o iden i ica ion (i.e., highe esolu ion image wi h 64
×
64 pixels). The main objec i e is no
o ha e he esul ing acial images good looking, bu a he ha hey can be used o iden i ica ion.
The da ase is o a compa able size, as a e he majo ace da ase s (i con ains 17,426 aces and
app oxima ely 6000–7000 unique pe sons). I ies o a oid he p oblem o bias, i.e., he eco dings
a e om a ious eal-wo ld scena ios whe e aces a e nea by o each o he , a ace is pa ially co e ed
including glasses, bea d, sca o o he co e ings (e.g., ace pain ings in some cases). The da ase
also con ains a ious acial images including a ious ligh ing condi ions (i.e., high, low, and poo
ligh ing quali y condi ions), a ious eco ding ha dwa e was used, ideos we e encoded using
a ious algo i hms, aces con ain a wide ange o ages, aces we e eco ded om a ious poses,
and hey con ain a ious emo ional exp essions (including smile, ea , ange o su p ise). I also
Appl. Sci. 2020,10, 7213 8 o 27
co e s all main e hnic g oups including Eu opeans (i.e., Middle Eas e ne s and Medi e aneans),
Eas Indians, Asians, Ame ican Indians, A icans, Melanesians, Mic onesians, Polynesians, Aus alians,
and Abo igines.
The c ea ed da ase was eleased online and can be used o ep oduce he expe imen desc ibed in
his pape . Al hough i was p ima ily designed o mul i- ame p ocessing, i is expec ed i s u iliza ion
will be wide .
You ube Faces DB [
9
] can be conside ed as a simila da ase o in oduced MLFDB.
Howe e , i di e s
in se e al pa ame e s: You ube Faces DB con ains aces ha a e no on i s
bo de line when iden i ica ion is possible and simple downsampling he acial images can b ing
bias and o e - i ing. Mo eo e , i con ains only 1595 di e en people, bu MLFDB con ains
app oxima ely 6000–7000 di e en people (see mo e de ails in Sec ion 5.1), and also i is no designed
o supe - esolu ion pu poses.
3.1.1. Da ase Gene al In o ma ion
The da ase was eleased online and o he objec i i y o he esul s, a pa o he da ase was
kep p i a e oo. Publicly a ailable se s a e TRAIN (12,200 samples) and TEST (2600 samples) all
oge he coun ing 14,800 ace sequences dedica ed o aining and e alua ion pu poses. The es o
he da ase (2626 ace sequences) is ese ed o pe o mance e alua ion in p i a e mode. Each sample
is o ganized in an indi idual olde and in e e y se , hey a e numbe ed om 1—
n
(TRAIN: 1—12200 ,
TEST: 1—2600). Each sample olde con ains:
•
7 images (i.e., ace sequences) in a low- esolu ion (32
×
32 px) named ‚img_1.JPG, . . . , ‚‘img_7.JPG’.
•
The label is de ined om he middle o he sequence ( om img_4.JPG) wi h double esolu ion
64 ×64 px and ile name label.JPG.
•
‘in o. x ’ ile wi h in o ma ion abou he label (bounding box and ame numbe ), ideo sou ce,
and in e pola ion me hod oge he wi h he JPEG comp ession ha we e used o esizing
inpu images.
A inal sample con ains: a sequence o se en images, label, and in o. x ile (see he example in
Figu e 4).
3.1.2. Da ase C ea ion
The da ase o sequences o acial images was c ea ed wi h he help o he YOLO 3 objec de ec o
(h ps://gi hub.com/s hanhng/yolo ace) [
30
]. By including only high-quali y samples, se e al il e s
we e applied be o e he sequence was included in he inal da ase . The i s il e was he condi ion
ha he de ec ed ace had o ha e a esolu ion g ea e han 64
×
64 px. Fo ep oducibili y pu poses
and also o allow e e yone in he u u e o use he same da ase , all he de ails om whe e he aces
we e cap u ed was p o ided. This in o ma ion includes he ace in ull esolu ion (used as a g ound
u h image) oge he wi h i s bounding box coo dina es and ideo ame numbe . To assu e a iabili y
o he da ase only a limi ed numbe o aces we e collec ed om each ideo, and he selec ed aces
we e chosen andomly. Nex , aw ace sequences we e ex ac ed om he ideo using he OpenCV
(h ps://pypi.o g/p ojec /openc -py hon/) lib a y. I sea ched o he co esponding ace ame
(label), i c opped he ace acco ding o he gi en bounding box and inally, i also ook 3 ames be o e
and 3 ames a e he c opped image using he same bounding box. This was an in en ional s ep
on ou pa o make he da ase mo e ealis ic as he e a e many cases when ace de ec o s do no
co ec ly c op he ace (inaccu a e de ec ion, c op scaling ac o ,. ..). Fu he op imiza ions a e le o
la e p ocessing s eps. Un o una ely, no all he cap u ed samples can be used. Fo his eason, a ew
o he s eps we e pe o med (see Figu e 5):
1.
Fo each sequence, i is checked whe he he label con ains a ace o su icien quali y so i
can be used o ace ecogni ion. This was done using an open-sou ce p ojec (Face- ecogni ion
Appl. Sci. 2020,10, 7213 9 o 27
amewo k [
31
]
1.3.0, CNN model
) which is inspi ed by Facene [
32
]. This s ep ac ually alida es
whe he he de ec ed ace is in good quali y so i can be used o he ace ecogni ion simila i y a e
index me ic (see Sec ion 3.4.1). YOLO is a ace de ec o , no a ace ecogni ion sys em, he e o e,
i also de ec s such aces whe e pe son iden i ica ion is impossible (e.g., head om behind).
2.
In some cases, a sequence can con ain mul iple o e lapping aces. I is no a p oblem i mo e
aces o hei pa s occu in one image, bu he p oblem a ises, when—mainly due o a mis ake
o he p ocessing algo i hm— he e is a di e en pe son a he beginning o he sequence and
ano he a he end. F om ime o ime i also happened due o a mo ie spli . The e o e, he label is
compa ed using he Face- ecogni ion amewo k (used h eshold was 0.7) wi h all o he images in
he sequence. Thus a possible p esence o se e al aces in a single image was aken in o accoun .
3.
Due o used andomness in sampling he ideos, i can also happen om ime o ime ha one
pe son appea ed mo e han once in he esul ing da ase . The bigges issue has been how o
dis inguish be ween e y simila images and images wi h di e en ligh ing, ace angle, e c. I has
been shown ha he Face- ecogni ion amewo k has no been an app op ia e solu ion o his
ask due o i s abili y o ace alignmen , i.e., i can pe o m he ecogni ion in di e en condi ions.
We decided o use a s uc u al simila i y me ic—SSIM (MSE, PSNR, e c., a e no e icien me ics
o his case) wi h he h eshold 0.7. G ound u h images a e used o his which a e i s esized
in o a 64 ×64 px esolu ion.
4.
The nex s ep was abou alida ion whe he all he sequence images con ain a ace. This was
done using he Face- ecogni ion amewo k and i s me hod o ace de ec ion (no ecogni ion).
By using his, only hose g ound u h images, which can be used o ecogni ion a e included.
O he images in he sequence mus con ain a ace, bu do no ha e o be ecognizable
(e.g., a di e en angle, pa ially co e ed by ano he pe son, e c.).
5.
The sequences which a e o iginally o a ious esolu ions (highe han 64
×
64 px) a e
downsampled in o esolu ion 32
×
32 px and he g ound u h (label) image in o a 64
×
64 px
esolu ion. When he o iginal esolu ion allowed his, he labels we e also eco ded in esolu ion
128
×
128 px and 256
×
256 px). Each sequence andomly used a chosen me hod o pixel
in e pola ions such as he nea es neighbo , bilinea , bicubic, lanczos (o e 8
×
8 neighbo hood).
The label was sa ed wi h he bes possible JPEG quali y (comp ession 100) and sequence images
we e sa ed by a andomly chosen JPEG comp ession scale om 30–90.
6.
Un o una ely, a e esizing he g ound u h images (labels) in o esolu ion 64
×
64 px some o
he images we e no possible o use o ace ecogni ion. Those sequences o ace images we e
emo ed om he da ase .
Appl. Sci. 2020,10, 7213 16 o 27
ou mul i- ame me hods, and we called hem U-Ne +ResBlock, U-Ne +GEU, U-Ne +GEU2,
and
U-Ne +GEU3.
Some o hem ake inspi a ion om he mos success ul single- ame based
me hods and a e modi ied o wo k wi h mul iple ames. These a chi ec u es we e p e-selec ed a e
expe imen a ion wi h many di e en combina ions o a chi ec u es and hei blocks, whe e hose ou
ha e been shown o be he mos p omising. Fo comple eness, se e al gene ally known in e pola ion
me hods we e also included in he expe imen o he pu pose o compa ison wi h he ea lie
men ioned me hods. Gaussian il e was applied o he ou pu o U-Ne +GEU o supp ess checke boa d
a i ac s [
40
] and due o an in e es ing esul s compa ison, i is conside ed as an indi idual me hod.
Thus, a o al o 12 di e en supe - esolu ion me hods we e e alua ed and compa ed. Al hough GAN
ne wo ks ha e been e y success ul in his a ea in ecen yea s, he p oposed mul i- ame a chi ec u es
we e no included in he expe imen . The eason o his exclusion was ha i was shown hey a e oo
c ea i e as was discussed in Sec ion 2.2. Al hough he aces look igh , hey canno be used o pe son
iden i ica ion since hey gene a e comple ely ake aces, which a e no p ima ily based on hei inpu
in o ma ion. In Sec ion 5) he esul s a e discussed.
4.1. Resul ing Da ase
Based on he p ocess desc ibed in Sec ion 3.1.2, we c ea ed an MLFDB da ase ha con ains a
o al o 17,426 samples. Each sample con ains a sequence o 7 low- esolu ion ace images (32
×
32 px).
Each o hese sequences has i s co esponding label in he 64
×
64 px esolu ion. The label was c ea ed
om he middle image o he sequence ( ame numbe 4). When i was possible, labels in he highe
esolu ion we e also p o ided. The e o e, some o he sequences also ha e labels in 128
×
128 px
(12,165 samples o he whole da ase , i.e., 70%) and 256
×
256 px (4199 samples, i.e., 24%) esolu ion.
An example o da ase samples is shown in Figu e 9.
Figu e 9. An example o sequence samples o he da ase .
This da ase o 17,426 samples was di ided in o a aining se (12,200 samples, 70%), es se
(2600 samples, 15%), a p i a e es se (2500 samples, 15%), and a es se o a ques ionnai e
(126 samples) o subjec i e human e alua ion. The aining da ase is in ended o he aining
and op imiza ion o supe - esolu ion me hods. The es se is only o he e alua ion o esul s. This
da ase should no be used o op imiza ion o sea ch o pa ame e s o a oid o e - i ing. The p i a e
es se is kep p i a e a he B no Uni e si y o Technology and will be used o e alua ion on eques
o ensu e he objec i i y o he esul s (h p://splab.cz/ml db/# esul s). The ques ionnai e da a se is
in ended o e alua ion by human e iewe s. This is mainly because esul s measu ed by objec i e
ma hema ical me hods do no always co espond wi h he quali y pe cei ed by human e alua o s.
Appl. Sci. 2020,10, 7213 17 o 27
4.2. Me hods Compa ison
All objec i e measu es desc ibed in Sec ion 3.4.1 we e used o he pe o mance e alua ion o
he models: SSIM, MSE, PSNR, sha pness as he di e ence be ween he g ound u h image and he
p edic ed image, ace ea u es ex ac ion ail a e (FR ailed) and ace ecogni ion simila i y a e index,
i.e., FR a e. Please no e ha FR a e (all) is compu ed only om he images whe e he ace ea u es
ex ac ion was success ul o he gi en esul ing me hod. FR a e (
∩
) me ic is compu ed om he
in e sec ion o he images wi h success ul ace ea u es ex ac ion by all me hods. The amoun o hese
images is 1518 o all 2500 es images.
The esul s o all 12 me hods a e placed all oge he in o Table 2 o he compendious compa abili y
o esul s be ween he single- ame and mul i- ame app oaches.
Table 2. Single- ame and mul i- ame me hods pe o mance compa ison—mean.
Me hod SSIM MSE PSNR Sha pness FR Ra e (All) FR Ra e (∩) FR Failed
bicubic 0.806 227.196 25.654 0.084 0.476 0.473 674/2500
bilinea 0.809 217.368 25.769 0.158 0.471 0.468 693/2500
lanczos 0.802 238.247 25.504 0.024 0.483 0.480 700/2500
EDSR 0.780 286.999 24.687 −0.025 0.495 0.494 701/2500
SRCNN 0.815 216.940 25.771 0.209 0.461 0.456 643/2500
SRGAN 0.760 320.981 24.234 −0.078 0.511 0.510 771/2500
ESRGAN 0.788 281.192 24.604 −0.016 0.488 0.485 470/2500
U-Ne +GEU 0.760 276.459 24.383 −0.129 0.507 0.503 317/2500
U-Ne +GEU+ il e 0.815 234.552 25.196 0.239 0.464 0.458 394/2500
U-Ne +GEU2 0.799 256.511 24.878 0.031 0.472 0.467 259/2500
U-Ne +GEU3 0.799 243.433 25.135 0.004 0.471 0.466 270/2500
U-Ne +ResBlock 0.812 239.114 25.141 0.130 0.468 0.463 581/2500
As can be seen om he esul s, he SRCNN me hod can be conside ed as he mos sui able
me hod in he way o he highes numbe o bes objec i e me ics ( he bes esul s o 5 om 7 me ics),
while no conside ing subjec i e human e alua ion. Howe e , he sha pness and he FR ailed me ics
o he SRCNN me hod a e signi ican ly wo se when compa ing he esul s o o he me hods.
The bes ace ecogni ion ex ac ion ail a e (FR ailed) has he U-Ne +GEU2 me hod ha ailed
only in 259 cases ou o 2500 compa ed o he SRCNN me hod ha ailed in 643 cases ou o 2500.
Mo eo e , he sha pness di e ence be ween hese wo me hods is no iceable. I should be no ed ha
he esul s desc ibed abo e only p esen s he pe o mance o me hods acco ding o he numbe o he
bes objec i e measu es and i does no ake in o accoun hei signi icance (i.e., SSIM and FR ailed
a e p obably mo e aluable han MSE). The impo an me ic—subjec i e human e alua ion—is no
included in hese esul s, he e o e, hey should no be gene ally aken as he inal ou comes.
U-Ne +GEU2 signi ican ly ou pe o med o he me hods by he subjec i e human e alua ion
(see esul s in Sec ion 4.3), bu among he objec i e me ics i had only one bes esul — ace ea u es
ex ac ion ail a e (FR ailed). The o he alues a e no so good compa ing o he SRCNN me hod,
bu he di e ences be ween hem a e signi ican unlike o he ace ea u es ex ac ion ail a e me ic
whe e he di e ence is e iden . Because o his, and due o he di e en impo ance o each me ic,
hese alues and hei compa ison a e be e eadable by a no malized ep esen a ion in o he ange
(0–1) (see Table 3). Minimum and maximum alues we e de e mined om all me hod’s alues
o he gi en me ic. In he case o he obse ed me ic, whe e he highe alue is be e (SSIM,
PSNR, subjec i e e alua ion), he no malized alue is compu ed as 1
−no m(x)
o ha e all alues
wi h he same logic—lowe alue is be e (i is conside ed as a “penaliza ion”). The inal sco e is
hen simply he sum o all no malized alues, sum1 includes only he objec i e me ics and sum2 also
includes he subjec i e human e alua ion—sub. e .
Appl. Sci. 2020,10, 7213 18 o 27
Table 3. No malized alues compa ison o all me hods o ge he inal sco e.
Me hod SSIM MSE PSNR Sha p. FR Ra e (∩) FR Failed Sub. e . Sum1 Sum2
bicubic 0.164 0.099 0.076 0.340 0.315 0.811 1 1.805 2.805
bilinea 0.109 0.004 0.001 0.655 0.222 0.848 0.989 1.839 2.828
lanczos 0.236 0.205 0.174 0.085 0.444 0.861 0.999 2.005 3.004
EDSR 0.636 0.673 0.705 0.703 0.089 0.863 0.977 3.669 4.646
SRCNN 0 0 0 0.872 0 0.750 0.978 1.622 2.600
SRGAN 1 1 1 1 0.315 1 0.995 5.315 6.310
ESRGAN 0.491 0.618 0.759 0.051 0.537 0.412 0.844 2.868 3.712
U-Ne +GEU 0 0.572 0.903 0.532 0.870 0.113 0.570 2.990 3.560
U-Ne +GEU+ il e 0 0.169 0.374 1 0.037 0.246 0.578 1.844 2.422
U-Ne +GEU2 0.281 0.380 0.581 0.115 0.204 0 0 1.571 1.571
U-Ne +GEU3 0.281 0.255 0.414 0 0.185 0.021 0.204 1.166 1.370
U-Ne +ResBlock 0.055 0.213 0.410 0.536 0.130 0.629 0.834 1.973 2.807
As can be seen om Table 3, a e a desi able ep esen a ion o he esul s and hei o e all sco e
in o m o he sum, U-Ne +GEU2 and U-Ne +GEU3 me hods ha e a be e o e all sco e han he
SRCNN me hod e en wi hou using he subjec i e human e alua ion me ic (sum1). The o e all
sco e including he subjec i e human e alua ion me ic (sum2) enla ges he di e ence be ween he
me hods almos wice as much. The U-Ne +GEU+ il e me hod has a wo se o e all sco e (sum1) han
he SRCNN me hod (1.844 s. 1.622) bu he esul s a e opposi e (2.422 s. 2.600) a e including he
subjec i e human e alua ion (sum2).
An example o esul ing images (p edic ions) on unseen da a om all 12 me hods is shown
in Figu e 10. Each ow ep esen s one inpu sequence and each column shows esul ing images
o each me hod. Finally, he e is a g ound u h image. The p esen ed samples a e no pa
o he MLFDB es se because he labels (g ound u h images) should no be publicly a ailable,
he e o e, hese sequences we e c ea ed o p esen a ion pu poses only.
Figu e 10.
Examples o he esul ing me hods. All se en 32
×
32 px images o he inpu sequence we e
used o he mul i- ame me hods and he 4 h image o single- ame app oaches was used.
Appl. Sci. 2020,10, 7213 19 o 27
As expec ed, he in e pola ion-based me hods canno econs uc he image in a be e quali y,
bu he di e ences be ween hem a e no iceable. Resul s o he EDSR and SRGAN me hods look
simila o he in e pola ion-based me hods. The SRCNN me hod, which has he bes esul s among
objec i e me ics, p edic s he esul ing image a li le be e , compa ed o o he single- ame me hods,
unlike he ESRGAN whose p edic ion quali y is much be e (sha pe ), bu i some imes c ea es
isible de o ma ion a i ac s. All mul i- ame me hods p o ide no ably be e esul s in con as o
single- ame me hods. The U-Ne +ResBlock me hod c ea es a bi smoo h blu y images, which is
annoying o human pe cep ion. The U-Ne +GEU+ il e me hod c ea es simila esul s, bu hey look
no so ema kable. The esul ing images o o he U-Ne +GEU based me hods ac simila ly o each
o he bu a e signi ican ly be e compa ed o all o he me hods.
4.3. Ques ionnai e Resul s
The subjec i e human e alua ion me ic was examined by a ques ionnai e in which 100 people had
o selec he bes image in hei opinion, among 126 compa isons o 12 me hods. In o al 12,600 “ o es”
we e a ailable and hey we e dis ibu ed o all me hods by indi idual p e e ences o a ending people.
The isualiza ion o subjec i e human e alua ion esul s is illus a ed in Figu e 11.
Figu e 11.
Resul s o he ques ionnai e—subjec i e human e alua ion. 100 people a ending illed ou
he ques ionnai e con aining 126 compa isons o 12 me hods.
F om he i s look a he cha , i is clea ha all U-Ne +GEU me hods signi ican ly
o e come o he me hods in human pe cep ion. 3815 o es om 12,600 belong o he U-Ne +GEU2
me hod.
U-Ne +ResBlock
and ESRGAN wi h app oxima ely 700 o es a e ollowing he U-Ne +GEU
base me hods. ESRGAN signi ican ly exceeds all single- ame me hods and as only one single- ame
me hod can be compa ed o esul s achie ed by he mul i- ame me hod—U-Ne +ResBlock, which is
he wo s among mul i- ame me hods, by human pe cep ion. Rema kably, ha SRCNN has he bes
objec i e me ics, bu ESRGAN is much mo e success ul by human pe cep ion. A simila si ua ion is
he same o mul i- ame me hods whe e he U-Ne +GEU+ il e has be e objec i e me ics, bu he
U-Ne +GEU2 has a signi ican ly be e subjec i e human e alua ion.
Appl. Sci. 2020,10, 7213 20 o 27
5. Discussion
5.1. Da ase
Un o una ely, i is qui e challenging o au oma ically coun unique aces in his da ase . Due o he
high numbe o sequences, i is almos impossible o emembe all he aces and he manual app oach is,
he e o e, no possible as well. Fo his eason, we used an au oma ed app oach which is based on he
Face- ecogni ion amewo k and can measu e a simila i y o aces. This amewo k e u ns ze o alue
when wo aces a e o he pe ec simila i y. A highe alue han ze o means less simila i y o hese
aces, which migh be caused no jus by di e en aces, bu also because aces we e aken om di e en
angles, ha e lowe quali y, e c. Acco ding o he p ojec au ho s,
he de aul
h eshold alue by which
he simila i y alue o wo aces should exp ess he same pe son is
≤
0.6. This alue is ecommended
o acial ecogni ion pu poses (compu ed on he LFW da ase ). Howe e , as can be seen in Table 4,
a h eshold
alue o 0.6 is imp ope o his pu pose. The eason o his is ha MLFDB con ains ace
images ha a e on he bo de line when hey s a being use ul o he iden i ica ion (i.e., lowe quali y
han he LFW da ase ). Fo he h eshold alue o 0.6,
only 76 sequences
we e conside ed unique,
which is a e manual e iew and de ini ely no he u h. Based on he empi ical expe imen a ion wi h
di e en alues, we ound he h eshold alue 0.45, which shows good pe o mance. These esul s
we e hen alida ed manually on a selec ed subse .
Table 4. Di e en people in he da ase using he Face- ecogni ion amewo k simila i y h eshold.
Th eshold 0.6 0.5 0.45 0.4 0.35
di e en people 76 2378 6651 10,822 13,369
Sequences om one ideo sou ce we e also il e ed using he SSIM me ic, whe e we used he
h eshold alue 0.7. Only he g ound u h images we e used o his il e ing. This il e ensu es ha
sequences e en o he same pe son a e kep when he e is a di e en scale, ligh ing, angles, e c. (i was
discussed in Sec ion 3.1.2).
The da ase is u he di ided in o public and p i a e pa s. The eason o his is o inc ease he
objec i i y o pe o mance e alua ion. In he case all he pa s would be published, hey can s ill be
used o he aining. Such esul s would wi h high p obabili y ou pe o m he esul s o o he eams,
bu hose esul s would be o e - i ed and would absen gene aliza ion o cou se. Because o ha ,
he es pa
o he da ase was spli in o wo pa s: he es -public and es -p i a e. The es -p i a e
was also eleased bu wi hou g ound u h images. E alua ion o his p i a e da ase is possible on
he MLFDB websi e by submi ing achie ed esul s. The e alua ion eques s a e limi ed by ime o
a oid al eady men ioned o e - i ing.
5.2. Ques ionnai e
The p oblem o ace image supe - esolu ion is signi ican ly mo e challenging han gene al
supe - esolu ion asks. I is clea ha in e e y supe - esolu ion me hod, some pa icula e o s
and mis akes mus be c ea ed since he inpu image con ains less in o ma ion han he e is expec ed
in he esul ing image. Thus, some pe cen ages o he pixels a e es ima ed. In he case o gene al
supe - esolu ion me hods, a mis ake in he image is o en ha d o be no iced and can be igno ed.
Howe e , when he image is checked o some de ails, ypically ex s, b and logos, o o he well-known
complex shapes, some e o s a e appa en a i s glance. This is also he case o acial images.
Fo example, i he esul ing image o a ace is missing an eye, o i has h ee eyes, i migh be o
me ics like PSNR o MSE neglec able e o , bu o human pe cep ion, i can be e y dis up i e.
The same si ua ion is wi h wo noses o no nose, he asymme ical appea ance o he eyes o o he wise
de o med ace.
Appl. Sci. 2020,10, 7213 21 o 27
Ano he opic ha should be aken in o conside a ion is ha he main objec i e is no jus o
gene a e images ha a e good looking. A mo e impo an c i e ion is ha he esul ing acial image
should be o such a quali y ha he iden i ica ion mus be possible. The s udied images a e a he
bo de line whe e images s a being use ul o he ecogni ion. When conside ing each pa icula
low- esolu ion image om a sequence alone, i is no possible o i is di icul o ecognize a pe son.
Ne e heless, he esul ing highe esolu ion image should be o su icien quali y, in which i can be
used o his pu pose. Fo his eason, we add essed se e al olun ee s and asked hem o ill ou
he ques ionnai e. The pu pose o he ques ionnai e was o ob ain eedback ega ding he quali y
o he achie ed ace images om he human pe spec i e. The olun ee s saw all he esul s o he
supe - esolu ion me hods, and hey we e asked o ma k he esul , which hey conside is he bes .
The U-Ne based me hods we e hei main choices, and he in e pola ion-based me hods we e
almos no selec ed. An in e es ing inding was ha he SRGAN me hod was conside ed as he wo s
one. The decisions we e in gene al made based on he essen ial ace ea u es clea ness (same eyes,
nose, mou h)— he en i e y o he ace. Howe e , some cases deg aded ee h and ace bo de s so much
ha he p e e ed image was wi h a lowe quali y image bu wi h an en i e ace. The ESRGAN me hod
was he only me hod ha had signi ican de o ma ions o he ace. Ano he in e es ing in o ma ion
was ha people in ui i ely p e e blu ed edges ins ead o sha p edges and no a i ac s, e en when
he image looks good. These eedbacks co ela e wi h he objec i e esul s, excep o he ESRGAN
me hod. I is he second sha pes me hod and i has go he mos o es among all single- ame me hods
om he ques ionnai e.
5.3. Resul s Compa ison
Resul s we e e alua ed by wo app oaches, using objec i e ma hema ical me ics (see Table 2) and
subjec i e e alua ion based on he human olun ee s’ o es (see Figu e 11). Fu he mo e, since e e y
objec i e me ic e alua es he images using a di e en poin o iew, we ha e me ged all he esul s
(using hei no malized ep esen a ion) and show he o e all sco e which is deno ed as sum1 and
sum2 (see Table 3).
Acco ding o he subjec i e me ics, he mul i- ame me hods achie ed he bes esul s.
Fu he mo e, a e he no maliza ion o he objec i e me ics alues and imp o ed ep esen a ion
(ins ead o absolu e alues) hey ha e had e en be e han he SRCNN me hod, which was e alua ed
as he bes as poin ed in Table 2. The p obable eason is ha he esul s o indi idual me ics o
mul i- ame me hods a e e y close o he bes alues among he es ed me hods (basically he SRCNN
me hod) and p ima ily hey do no ha e any signi ican ou lie s among all me ics compa ed o
single- ame me hods, especially he SRCNN me hod.
No GAN-based mul i- ame supe - esolu ion me hods we e used in his wo k. The eason o
his is hei gene a i e na u e, which was discussed in mo e de ail in Sec ion 2.2. In he scope o his
wo k, he e we e conduc ed expe imen s wi h GANs, bu hese expe imen s concluded, he cu en
me hods should no be used o pe son iden i ica ion pu poses. I should be emphasized ha gene al
supe - esolu ion and supe - esolu ion o iden i ica ion pu poses ha e di e en objec i es.
Nowadays, one o he obs acles o he as e de elopmen o hese me hods is he lack o objec i e
me ics which will be e co espond o human pe cep ion. In his pape , we in oduce an ex ended se
o me ics like SSIM, PSNR, e c. Un o una ely, i was no a a e case when he bes esul s selec ed by
he human olun ee s did no ma ch he op esul s by he objec i e me ics. Because o ha , he deep
lea ning models c ea ed images wi h small checke boa d a i ac s such as small di e en colou in s
o small ace e-posi ioning, which a ec hese me ics. Those issues become e en mo e impo an
i he e a e no g ound- u h images a ailable, i.e., he model is used in eal si ua ions. On he o he
hand, i is p obably be e o ma k se e al suspicious ace images wi h lowe accu acy, a he han no
selec ing any o he aces. Fo example, i is be e o p eemp i ely check en people ins ead o no-one
when looking o dange ous suspec s.
Appl. Sci. 2020,10, 7213 22 o 27
5.4. Face Recogni ion Based Me ics
Fo all 2500 es images, he ace ecogni ion simila i y a e index (FR a e (all)) was compu ed.
The same me ic was also compu ed o he subse o he es da ase whe e ace ecogni ion was
success ul (1518 samples). Only hose samples a e conside ed o be success ul, whe e all he
models c ea ed such a esul ing image, which succeeded wi h ace ecogni ion. The e was an
assump ion ha he me hods wi h be e ace ea u es ex ac ion ail a e (lowe FR ailed) will
ha e a be e ace ecogni ion simila i y a e index on hose subse images ins ead o he whole es se .
Howe e , acco ding o achie ed esul s, i seems ha his has only a minimal e ec .
A mo e in e es ing me ic is he ace ea u es ex ac ion ail a e (FR ailed), i.e., in how many cases
i was no possible o ex ac success ully ace ea u es. Acco ding o he esul s, he
U-Ne +GEU2
me hod has he lowes numbe o cases whe e ace ecogni ion ailed. I was also examined in how
many cases i ailed on he same images. The numbe o di e en images ha a e no in he se o ailed
images o he U-Ne +GEU2 me hod is used o make he ou pu mo e in e p e able. These alues a e
shown as he i s alue in Table 5. The second alue a e he slash ep esen s he o e all di e ence
om he U-Ne +GEU2 FR ailed me ic.
Table 5. Face ecogni ion (FR) ailed di e ences compa ed o he U-Ne +GEU2 me hod.
Me hod Di e ences
bicubic 12/415
bilinea 9/434
lanczos 16/441
EDSR 18/442
SRCNN 11/384
SRGAN 15/512
ESRGAN 35/211
U-Ne +GEU 38/58
U-Ne +GEU+ il e 16/135
U-Ne +GEU3 55/11
U-Ne +ResBlock 10/322
The bes ma ch o ailed images is pa adoxically o he bilinea in e pola ion. On he o he
hand, he bilinea in e pola ion ailed in 693 cases ins ead o he U-Ne +GEU2 me hod ha ailed
in 259 cases ( he di e ence is 434) and he bilinea in e pola ion se o ailed images is big enough
o ma ch he majo i y o he U-Ne +GEU2 se o images. An in e es ing inding is an opposi e issue.
The
U-Ne +GEU3
me hod has almos he same numbe o FR ailed me ic ( he di e ence is only
11 cases), bu i s se o images di e s in 55 cases compa ed o he U-Ne +GEU2 me hod.
A e u he examina ion o he me hod, he e we e no pa e ns ound ( he same in e pola ions,
comp ession, images backg ound, e c.), which would explain why he e a e so di e en esul s.
Howe e , he e a e some di e si ies be ween U-Ne +GEU2 and U-Ne +GEU3 me hods in some cases.
Figu e 12a shows a de ailed iew o he esul ing quali y c ea ed using U-Ne +GEU2 and U-Ne +GEU3.
Su p isingly, he e a e also con a y cases, bu hey a e no so signi ican as illus a ed in Figu e 12b.
U-Ne +GEU3 me hod ailed in ace ea u es ex ac ion in bo h cases. The e a e p obably wo essen ial
explana ions. Fi s , human pe cep ion di e s om he machine lea ning image p ocessing and how he
quali y o he inpu image is assessed. Second, he well-known and used Face- ecogni ion amewo k
does no wo k as well as p esen ed, a leas o his case. Among he images which ailed in ace
ea u es ex ac ion a e a couple o p o ile images (see examples o labels in Figu e 12c). The ace
ea u es ex ac ion wo ks in his case, bu wi h wha accu acy? How do ace landma k es ima ions
and ace ans o ma ions wo k in hese cases?
Appl. Sci. 2020,10, 7213 23 o 27
Figu e 12.
Examples o low ace ea u e ex ac ion pe o mance. (
a
,
b
) U-Ne + GEU2 (passed) and
U-Ne + GEU3 ( ailed) me hods compa ison. (
c
) Examples o p o ile ace images o da abase samples
ha passed ace ea u es ex ac ion as well du ing he da abase c ea ion il a ion s ep.
5.5. Benchma k and Leade boa d
MLFDB da ase is he i s mul i- ame acial da ase o compa able size, which has he po en ial
o suppo he de elopmen o his ield. Ano he impo an s ep is o se a baseline in he o m o
esul s om a ious me hods, so anyone can compa e he achie ed esul s wi h new me hods.
The gene al p oblem o many da ase s is ha hey can be easily o e - i ed. This o en happens
when pa ame e s a e op imized no jus using he aining se , bu also he es se . In he si ua ion
when e e y hing is publicly a ailable, i is no possible o ha e con ol o e his. Millions o ails can
lead in some cases o esul s ha do no e lec he eal pe o mance o he me hod. Some imes he e is
no clea ly de ined how he compa ison me ics should be compu ed, so he p og ess is di icul o be
objec i ely e alua ed.
To deal wi h his p oblem, we decided o c ea e an au oma ic benchma k (inspi ed by sys ems like
Kaggle (h ps://www.kaggle.com/)) o an objec i e esul s compa ison on his da ase . G ound u h
images o he p i a e pa o he es da ase a e no publicly a ailable, and hus me ics a e compu ed
in he same manne as hey we e compu ed in his pape . Fo his pu pose, he MLFDB se e is
a ailable online, whe e anyone can submi hei esul s (h p://splab.cz/ml db/). The inal sco e will
be compu ed as sum1 desc ibed in Sec ion 4.2. The numbe o esul s o compu a ion que ies will be
limi ed acco ding o he ules o he benchma k.
5.6. Fu u e Resea ch
The e a e plen y o ways how o ollow up on his wo k. Especially, i could be in e es ing o
expe imen wi h models wi h a highe scale ac o , o example, 4
×
o 8
×
, as was poin ed ou in he
example in Figu e 2. The MLFDB da ase has all g ound u h images in esolu ion 64
×
64 px,
bu he e
a e also many images in e en highe esolu ion—128 ×128 px (12,165 samples) o 256 ×256 px (4199
samples). Ano he way wo h ying could be changing he numbe o images in sequences o a mo e
complex and be e u iliza ion in eal si ua ions would be an applica ion o some ea u e selec ion [
47
].
The ea u e selec ion in gene al leads o a be e pe o mance because o ocusing on he impo an
pa s o inpu da a and emo ing ou lie s o noisy da a ha cause model inaccu acy.
6. Conclusions
Fo ensically ained acial e iewe s a e s ill conside ed o be one o he mos accu a e app oaches
o pe son iden i ica ion, especially in he case o low- esolu ion o low-quali y ideos. The human b ain
can u ilize in o ma ion no jus om a single image bu also om a sequence o aces
(i.e., ideos) and,
e en in he case o low-quali y eco ds o a long dis ance om a came a. They can accu a ely iden i y
a pe son. Fo compu e me hods, his emains a challenge. Howe e , on he o he hand, hey ha e he
po en ial o suppo human e iewe s and help o p e-p ocess he da a and ex ac as much in o ma ion
om he da a as possible. One o such a use-case would be, o example, o econs uc he acial image
in highe quali y om a ideo, which can police use o sea ch and announce i in newspape s.
This pape in oduced a la ge-scale ace da ase con aining 17,426 sequences o ace images.
The da ase co e s di e en aces, ages, ligh ing condi ions, and many ypes o di e en came a de ices.
Reco ds we e ob ained om he eal en i onmen , including common de ec s and impe ec ions ha
occu in he ideos (e.g., comp ession a i ac s, blu , e c.) Acco ding o ou knowledge, i is he i s
Appl. Sci. 2020,10, 7213 24 o 27
da ase o a compa able size con aining sequences o ace images. The pape also in oduces a new
mul i- ame ace supe - esolu ion me hod. This me hod has been p o en o p oduce be e esul s han
single- ame me hods which a e conside ed as s a e-o - he-a , and hey a e one o he mos e ec i e
and o en compa ed me hods in his esea ch a ea. The e o e, he hypo hesis ha he mul i- ame
me hod p oduces be e esul s han he deep lea ning-based single- ame me hods was p o en.
The esul s we e also e alua ed using se e al objec i e me ics and also subjec i ely by olun ee s
whe e he U-Ne +GEU2 me hod has he bes esul s ega ding human pe cep ion, and he U-Ne +GEU3
me hod achie ed he bes o e all sco e o objec i e me ics. The p oposed me hod is no gene a i e
and i was p o en o imp o e he quali y o he ace image. The sou ce code and he da ase we e
eleased, and he expe imen is ully ep oducible.
The new and unique MLFDB da ase has been published o s udying ace supe - esolu ion
om sequences o images (i.e., mul i- ame p oblem). The benchma k wi h gi en ules and me ics
was c ea ed, and esul s achie ed in his pape we e used as an ini ial s a ing poin and as a
challenge o o he esea che s. They a e p esen ed in he leade boa d able as a pa o he benchma k.
Rega ding esul s hemsel es, i was ound ha human pe cep ion plays a signi ican ole du ing
model e alua ion. The e o e, i is essen ial o use some subjec i e human e alua ion, o example,
in he o m o ques ionnai es o su eys.
Au ho Con ibu ions:
Concep ualiza ion, M.R. and R.B.; me hodology, A.M.; so wa e, M.R. and A.M.;
alida ion, M.R. and R.B.; o mal analysis, A.M. and M.R.; in es iga ion, A.M.; esou ces, R.B.; da a cu a ion, M.R.
and A.M.; w i ing—o iginal d a p epa a ion, M.R.; w i ing— e iew and edi ing, R.B.; isualiza ion, M.R. and
A.M.; supe ision, R.B. All au ho s ha e ead and ag eed o he published e sion o he manusc ip .
Funding: This esea ch was unded by he In e eg Cen al Eu ope niCE-li e by he g an CE1581.
Acknowledgmen s:
Resea ch desc ibed in his pape was inanced by he In e eg Cen al Eu ope niCE-li e by
he g an CE1581. We would like o hank all o he olun ee s ha a ended he ques ionnai e, especially he ones
who p o ided he eedback.
Con lic s o In e es : The au ho s decla e no con lic o in e es .
Abb e ia ions
The ollowing abb e ia ions a e used in his manusc ip :
CelebA La ge-scale CelebFaces A ibu es
CCTV Closed ci cui ele ision
CNN Con olu ional Neu al Ne wo k
CPBD Cumula i e P obabili y o Blu De ec ion
CVPR Con e ence on Compu e Vision and Pa e n Recogni ion
DB Da aBase
DeepSUM Deep neu al ne wo k o Supe - esolu ion o Un egis e ed Mul i empo al images
EDSR Enhanced Deep Supe -Resolu ion Ne wo k
ESRGAN Enhanced Supe -Resolu ion Gene a i e Ad e sa ial Ne wo k
EDVR Enhanced De o mable Con olu ional Ne wo k
FERET Face Recogni ion Technology
FR Face Recogni ion
FRVSR F ame- ecu en ideo supe - esolu ion
FSRGAN Face Supe -Resolu ion Gene a i e Ad e sa ial Ne wo k
FSRNe Face Supe -Resolu ion Ne wo k
GAN Gene a i e Ad e sa ial Ne wo k
GPU G aphics P ocessing Uni
JNB Jus No iceable Blu
LFW Labeled Faces in he Wild
MLFDB Mul i- ame Labeled Faces Da abase
MSE Mean Squa e E o
NN Neu al Ne wo k
PCD Py amid, Cascading and De o mable con olu ions
Appl. Sci. 2020,10, 7213 25 o 27
PSNR Peak signal- o-noise a io
PubFig Public Figu es Face Da abase
PULSE Sel -Supe ised Pho o Upsampling ia La en Space Explo a ion o Gene a i e Models
ReLU Rec i ied Linea Uni
ResNe Residual Ne wo k
ResBlock Residual Block
SR Supe -Resolu ion
SRCNN Supe -Resolu ion Con olu ional Neu al Ne wo k
SRGAN Supe -Resolu ion Gene a i e Ad e sa ial Ne wo k
SSIM S uc u al simila i y
TSA Tempo al and Spa ial A en ion
UR-DGN Ul a- esolu ion by disc imina i e gene a i e ne wo k
UVT Unpai ed Video T ansla ion
VSR Video Supe -Resolu ion
YOLO You Only Look Once
Re e ences
1.
Hollis, M.E. Secu i y o su eillance? Examina ion o CCTV came a usage in he 21s cen u y.
C iminol. Public Policy 2019,18, 131–134. [C ossRe ]
2.
Molini, A.B.; Valsesia, D.; F acas o o, G.; Magli, E. DeepSUM: Deep neu al ne wo k o Supe - esolu ion o
Un egis e ed Mul i empo al images. IEEE T ans. Geosci. Remo e Sens. 2019,58, 3644–3656. [C ossRe ]
3.
Sal e i, F.; Mazzia, V.; Khaliq, A.; Chiabe ge, M. Mul i-Image Supe Resolu ion o Remo ely Sensed Images
Using Residual A en ion Deep Neu al Ne wo ks. Remo e Sens. 2020,12, 2207. [C ossRe ]
4.
Zangeneh, E.; Rahma i, M.; Mohsenzadeh, Y. Low esolu ion ace ecogni ion using a wo-b anch deep
con olu ional neu al ne wo k a chi ec u e. Expe Sys . Appl. 2020,139, 112854. [C ossRe ]
5.
Huang, G.B.; Ma a , M.; Be g, T.; Lea ned-Mille , E. Labeled Faces in he Wild: A Da abase o S udying Face
Recogni ion in Uncons ained En i onmen s; Technical Repo 07-49; Uni e si y o Massachuse s o Amhe s :
Amhe s , MA, USA, Oc obe 2007.
6.
Kuma , N.; Be g, A.C.; Belhumeu , P.N.; Naya , S.K. A ibu e and simile classi ie s o ace e i ica ion.
In P oceedings o he 2009 IEEE 12 h In e na ional Con e ence on Compu e Vision, Kyo o, Japan,
27 Sep embe –4 Oc obe 2009; pp. 365–372.
7.
F eeman, W.T.; Pasz o , E.C.; Ca michael, O.T. Lea ning low-le el ision. In . J. Compu . Vis.
2000
,40, 25–47.
[C ossRe ]
8.
Liu, Z.; Luo, P.; Wang, X.; Tang, X. La ge-scale celeb aces a ibu es (celeba) da ase . Re ie ed Augus
2018,15, 2018.
9.
Wol , L.; Hassne , T.; Maoz, I. Face ecogni ion in uncons ained ideos wi h ma ched backg ound simila i y.
In P oceedings o he CVPR 2011, Colo ado Sp ings, CO, USA, 20–25 June 2011; pp. 529–534.
10.
O’Mahony, N.; Campbell, S.; Ca alho, A.; Ha apanahalli, S.; He nandez, G.V.; K palko a, L.; Rio dan, D.;
Walsh, J. Deep lea ning s. adi ional compu e ision. In P oceedings o he Science and In o ma ion
Con e ence, Leipzig, Ge many, 1–4 Sep embe 2019; pp. 128–144.
11.
Dong, C.; Loy, C.C.; He, K.; Tang, X. Image supe - esolu ion using deep con olu ional ne wo ks. IEEE T ans.
Pa e n Anal. Mach. In ell. 2015,38, 295–307. [C ossRe ]
12.
Isola, P.; Zhu, J.Y.; Zhou, T.; E os, A.A. Image- o-image ansla ion wi h condi ional ad e sa ial ne wo ks.
In P oceedings o he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, Honolulu, HI, USA,
21–26 July 2017; pp. 1125–1134.
13.
Ledig, C.; Theis, L.; Huszá , F.; Caballe o, J.; Cunningham, A.; Acos a, A.; Ai ken, A.; Tejani, A.; To z, J.;
Wang, Z.; e al. Pho o- ealis ic single image supe - esolu ion using a gene a i e ad e sa ial ne wo k.
In P oceedings o he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, Honolulu, HI, USA,
21–26 July 2017; pp. 4681–4690.
14.
Lim, B.; Son, S.; Kim, H.; Nah, S.; Mu Lee, K. Enhanced deep esidual ne wo ks o single image
supe - esolu ion. In P oceedings o he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion,
Honolulu, HI, USA, 21–26 July 2017; pp. 136–144.