scieee Open visual document viewer

Multi-Frame Labeled Faces Database: Towards Face Super-Resolution from Realistic Video Sequences

Rajnoha, Martin; Mezina, Anzhelika; Burget, Radim

Abstract

Forensically trained facial reviewers are still considered as one of the most accurate approaches for person identification from video records. The human brain can utilize information, not just from a single image, but also from a sequence of images (i.e., videos), and even in the case of low-quality records or a long distance from a camera, it can accurately identify a given person. Unfortunately, in many cases, a single still image is needed. An example of such a case is a police search that is about to be announced in newspapers. This paper introduces a face database obtained from real environment counting in 17,426 sequences of images. The dataset includes persons of various races and ages and also different environments, different lighting conditions or camera device types. This paper also introduces a new multi-frame face super-resolution method and compares this method with the state-of-the-art single-frame and multi-frame super-resolution methods. We prove that the proposed method increases the quality of face images, even in cases of low-resolution low-quality input images, and provides better results than single-frame approaches that are still considered the best in this area. Quality of face images was evaluated using several objective mathematical methods, and also subjective ones, by several volunteers. The source code and the dataset were released and the experiment is fully reproducible.

Full text

applied sciences A icle Mul i-F ame Labeled Faces Da abase: Towa ds Face Supe -Resolu ion om Realis ic Video Sequences Ma in Rajnoha * , Anzhelika Mezina and Radim Bu ge * Depa men o Telecommunica ions, B no Uni e si y o Technology, 616 00 B no, Czech Republic; xmezin00@ u b .cz *Co espondence: ma in. ajnoha@ u b .cz (M.R.); [email p o ec ed].cz (R.B.) Recei ed: 25 Augus 2020; Accep ed: 10 Oc obe 2020; Published: 16 Oc obe 2020   Fea u ed Applica ion: Du ing police sea ches o pe pe a o s o se ious c imes, o en only a low de ini ion ideo is a ailable. The p oposed me hodology can use a sequence o aces om a ideo o econs uc a single high de ini ion ace. Biome ic iden i ica ions a e used wo ldwide, bu he e is s ill space o imp o emen s in hei accu acy o wo king unde bad en i onmen al condi ions. Abs ac : Fo ensically ained acial e iewe s a e s ill conside ed as one o he mos accu a e app oaches o pe son iden i ica ion om ideo eco ds. The human b ain can u ilize in o ma ion, no jus om a single image, bu also om a sequence o images (i.e., ideos), and e en in he case o low-quali y eco ds o a long dis ance om a came a, i can accu a ely iden i y a gi en pe son. Un o una ely, in many cases, a single s ill image is needed. An example o such a case is a police sea ch ha is abou o be announced in newspape s. This pape in oduces a ace da abase ob ained om eal en i onmen coun ing in 17,426 sequences o images. The da ase includes pe sons o a ious aces and ages and also di e en en i onmen s, di e en ligh ing condi ions o came a de ice ypes. This pape also in oduces a new mul i- ame ace supe - esolu ion me hod and compa es his me hod wi h he s a e-o - he-a single- ame and mul i- ame supe - esolu ion me hods. We p o e ha he p oposed me hod inc eases he quali y o ace images, e en in cases o low- esolu ion low-quali y inpu images, and p o ides be e esul s han single- ame app oaches ha a e s ill conside ed he bes in his a ea. Quali y o ace images was e alua ed using se e al objec i e ma hema ical me hods, and also subjec i e ones, by se e al olun ee s. The sou ce code and he da ase we e eleased and he expe imen is ully ep oducible. Keywo ds: ace ecogni ion; supe esolu ion; mul i ame; image p ocessing; da abase; da ase ; sequences; deep lea ning 1. In oduc ion Closed-ci cui ele ision (CCTV) is a widely used echnology o moni o ing public and p i a e places, which helps o inc ease o e all sa e y, p e en auds, shopli ing, he s, bu gla ies, andalism, e o ism and o he s. I is o en ins alled in public a eas and businesses h oughou he wo ld o p e en he abo emen ioned c imes and o inc ease o e all public sa e y. Un o una ely, he e a e also many abuses and o he ela ed ques ions mainly conce ning p i acy issues. In ela ion o police sea ches and he in es iga ion o se ious c imes, he e is no doub ha hey ha e a e y posi i e impac on he cla i ica ion o c iminal o ences and ea ly de en ion o c ime o ende s [1]. In p ac ice, i o en happens ha e en hough he c ime is eco ded using CCTV, he eco ding canno be used o police in es iga ion o a police sea ch and canno be used as e idence. The cause o his is o en he poo quali y o he eco ding, he dis ance o he came a om he c ime place, bad ligh ing, bad wea he condi ions, o o he s ac o s which cause he pe son on he eco ding o be Appl. Sci. 2020,10, 7213; doi:10.3390/app10207213 www.mdpi.com/jou nal/applsci Appl. Sci. 2020,10, 7213 2 o 27 unambiguously iden i ied. Such a eco d canno e en be used by he police o ask he public o help in assis ance wi h acing he pe pe a o , o ins ance in newspape s o online media. Mos o he compu e ace ecogni ion sys ems usually use a single s ill image and based on i ecognize a ace. The human expe s— o ensically ained acial e iewe s—a e s ill conside ed one o he mos accu a e app oaches o pe son iden i ica ion. They a e gene ally mo e esilien o noise in images, and also mo e e icien in u ilizing he in o ma ion con ained in he ideo, e en when he quali y o images a e poo . They can ake in o accoun he dynamics o walk, ges u es, i ness s a us o a pe son, and possibly o he cha ac e is ics. They can also u ilize he in o ma ion om a sequence o ace images and can be mo e accu a e when compa ed o ecogni ion om each isola ed image. This pape ocuses on inc easing he quali y o he acial image using a ideo sequence wi h emphasis on biome ic u iliza ion whe e he p inciple is illus a ed (see Figu e 1). The inpu ace image is eco ded om a long dis ance (i.e., i s esolu ion is low and hus a single acial image canno be used o pe son iden i ica ion). This pa icula example demons a es he di icul y o pe son iden i ica ion om a long dis ance, which is qui e equen o su eillance came as. Inspi ed by human skills, his pape in es iga es me hods ha can econs uc a high- esolu ion ace image om a se ies o low- esolu ion and poo quali y images. Figu e 1. This pape is inspi ed by he human capaci y o ecognize a ace om a sequence o images be e han om a s ill pic u e. The esea ch desc ibed in his pape in oduces a couple o con ibu ions, whose no el y o imp o emen a e b ie ly desc ibed by he ollowing poin s o be e cla i y: • This pape in oduces a new me hodology o supe - esolu ion (SR) di e en om he gene al-pu pose one, which is limi ed o aces and biome ic applica ions. • The sequence o mul iple images is p ocessed (mul i- ame me hods) ins ead o using jus a single image as mos exis ing da ase s allow (single- ame me hods). • We p oposed a me hodology ha was compa ed o o he s a e-o - he-a me hods, cu en ly, he single- ame supe - esolu ion me hods can be conside ed as one o he bes because o he lack o he p og ess and non-u iliza ion o deep lea ning in mul i- ame supe - esolu ion me hods [ 2 , 3 ]. Mo eo e , gene al supe - esolu ion me hods a e o en no applicable o pu poses o biome ics, e en i hey p o ide good esolu ion. • A new unique la ge-scale da ase , con aining 17,426 samples o image sequences aken om a wide ange o eco ding de ices, is p o ided. They con ain aces o a ious ages, aces, illumina ion condi ions ha a e aken om eal-wo ld en i onmen s. • The pape also de ines me ics o acial quali y measu emen s ha a e ep oducible and based on open-sou ce solu ions. These objec i e me ics a e widely accessible app oaches and can boos u he esea ch in his a ea. • The whole expe imen desc ibed in his pape is ully ep oducible— he sou ce codes and he da ase we e eleased online (h p://splab.cz/ml db/#download). • The combina ion o he key poin s abo e c ea es a g ound (basis) and i has he po en ial o c ea ing a s anda d o u u e esea ch in his ield. The es o he pape is s uc u ed as ollows. Sec ion 2p o ides ela ed wo k ega ding single- ame and mul i- ame supe - esolu ion me hods. Sec ion 3desc ibes he expe imen which Appl. Sci. 2020,10, 7213 3 o 27 has been conduc ed. I includes he desc ip ion o he c ea ed da ase and a ious p oposed me hodologies o ace ecogni ion and e alua ion me ics. Resul s achie ed using ou me hodology a e shown in Sec ion 4, and Sec ion 5discusses he esul s we achie ed and on how o in e p e hem. Finally, Sec ion 6concludes he pape . 2. Rela ed Wo k Face iden i y ecogni ion can usually be done manually by o ensically ained acial e iewe s o au oma ically, usually by a compu e expe sys em. The au oma ed ace iden i y ecogni ion is a esea ch a ea which has expe ienced a apid g ow h in e ms o accu acy and eliabili y in ecen yea s. This success is achie ed mainly hanks o he u iliza ion o he so-called deep neu al ne wo ks [4] , highe compu a ional powe , and also he p esence o la ge aining and e alua ion da ase s such as Labeled Faces in he Wild (LFW) (13,233 images) [ 5 ], PubFig (58,797 images) [ 6 ], FERET (14,126 images) [ 7 ], CelebA (200 k images, 10,177 iden i ies) [ 8 ] o You ube aces DB (1595 iden i ies) [ 9 ]. These da ase s con ain ace images wi h a ious illumina ion condi ions, exp essions, aces, and poses. Al hough he e has been g ea p og ess in ace iden i y ecogni ion and i s accu acy in ecen yea s, he econs uc ion o an image om a low- esolu ion one s ill emains a challenge [ 4 ]. The e is cu en ly no la ge-scale ace da ase con aining se e al sequences o images. Many o he cu en supe - esolu ion me hods ha e a gene a i e na u e and, he e o e, canno be used o iden i ica ion pu poses since hey a e oo c ea i e. The app oaches, which add esses he p oblem o inc easing esolu ion, can be di ided by bo h he con en ype and inpu da a ype. The me hods based on he con en ype can be di ided in o gene al supe - esolu ion echniques and biome y- ocused supe - esolu ion echniques. F om he poin o iew o he inpu da a ype, he me hods can be di ided in o wo g oups: wi h a single inpu image ( he so-called single- ame) o wi h many inpu images ( he so-called mul i- ame, i.e., ideo). Supe - esolu ion, mainly hanks o he success o deep lea ning, has made signi ican p og ess in ecen yea s [ 10 ]. This is also he eason why he mos success ul me hods o supe - esolu ion a e mainly based on neu al ne wo ks. In pa icula , hey a e Con olu ional Neu al Ne wo ks (CNN) o Gene a i e Ad e sa ial Ne wo ks (GANs). 2.1. Gene al Single-F ame Supe -Resolu ion Me hods Single- ame supe - esolu ion me hods a e me hods ha c ea e a high- esolu ion image om a single low- esolu ion image. Mos o hese me hods a e gene al pu pose (i.e., hey a e no specialized o any pa icula applica ion a ea, o example, aces). In case we do no ake in o accoun he in e pola ion me hods (such as bicubic, bilinea , nea es neighbou , e c.), one o he i s well-known me hods was he Supe -Resolu ion Con olu ional Neu al Ne wo k (SRCNN) [ 11 ], which was in oduced in 2015. This neu al ne wo k (NN) a chi ec u e was ela i ely ligh weigh —i had only h ee laye s, which is a ela i ely low numbe when conside ing he cu en s a e-o - he-a a chi ec u es. Howe e , i was shown ha e en such a simple NN ou pe o ms signi ican ly he capabili ies o any o he in e pola ion me hod. Ano he signi ican p og ess in he a ea o supe - esolu ion was based on he so-called gene a i e ad e sa ial ne wo ks. Al hough GANs we e in oduced in 2014, hey showed hei applicabili y in p ac ice o supe - esolu ion only in 2017. One o he well-known p ojec s was Pix2Pix [ 12 ]. I was o iginally in oduced as an image- o-image ansla ion model ( he ne wo k has he same dimensions o he inpu and ou pu laye s). Especially when compa ed o in e pola ion-based me hods, Pix2Pix has also shown a signi ican imp o emen [ 12 ]. SRGAN [ 13 ] is ano he success ul app oach based on GANs. I s inno a ion was o use he esidual a chi ec u e and imp o emen ega ding he loss unc ion—in pa icula , i used he pe cep ual loss unc ion du ing aining ins ead o he pixel-wise loss, such as Mean Squa e E o (MSE). Ano he success ul a chi ec u e was EDSR, which was in oduced in [ 14 ] and i s main idea was o ain high-scale models om p e- ained low-scale models. An in e es ing inno a ion was o sha e pa ame e s ac oss di e en scales. ESRGAN [ 15 ] is inspi ed by Appl. Sci. 2020,10, 7213 4 o 27 SRGAN. The inno a ion o ESRGAN is o emo e ba ch no maliza ion laye s, he modi ied pe cep ual loss unc ion, and ans e lea ning. The disc imina o pa o GAN was eplaced by a ela i is ic disc imina o [16]. In 2019, mainly due o using huge compu a ional powe and la ge GPU memo y, i was shown ha he NNs can gene a e e en ela i ely complex images wi h a lo o de ail [ 17 ]. Al hough hese ne wo ks ha e achie ed g ea success, hei big disad an age is he ac ha hey do no y o ai h ully econs uc he in o ma ion ha is on he inpu , bu hey gene a e only pa o he in o ma ion. Fo his eason, hey a e no sui able o biome ics, ace econs uc ion, and police in es iga ion, as hey end o gene a e o ally new aces. This is un o una ely a p oblem o mos o he GAN based ne wo ks. 2.2. Single-F ame Supe -Resolu ion Focused on Faces Supe - esolu ion me hods men ioned in he p e ious sec ion we e designed o gene al use. This sec ion desc ibes a sub-se o supe - esolu ion me hods ha di ec ly ocus on ace supe - esolu ion ela ed p oblems, and in connec ion wi h biome ics oo. Al hough many o he me hods men ioned ea lie o e images in eally high- esolu ion (o en 2 × , 4 × , o also 8 × zoom) and he human eye o en canno ecognize ha he ace was c ea ed om he o iginal low- esolu ion image, hese images canno be used o pe son iden i ica ion. The images a e syn hesized om pas expe ience (images used as a aining da ase o NN) and o en gene a es ake aces. The la es comp ehensi e su ey pape ela ed o biome ics and supe - esolu ion echniques is p o ided in [ 18 ]. Among o he s, his pape also discusses in-dep h di e en equi emen s ega ding gene al supe - esolu ion me hods and biome ics (o ace iden i y ecogni ion) me hods. When we jus ocus on me hods ela ed o acial supe - esolu ion, one o he s a e-o - he-a me hods is he Ul a-Resolu ion by Disc imina i e Gene a i e Ne wo ks (UR-DGN). This a chi ec u e is based on GAN and allows us o upscale images wi h scale ac o 8 [ 19 ]. The au ho s used a decon olu ional ne wo k as he gene a o and a con olu ional ne wo k as he disc imina o . Ano he app oach in his ield is he end- o-end ainable Face Supe -Resolu ion Ne wo k (FSRNe ), which is based on CNN, and FSRGAN, which eco e s mo e ealis ic ex u es han FSRNe [ 20 ]. These a chi ec u es u ilize he es ima ion o landma k hea maps. One o he la es wo ks is also he P og essi e Face Supe -Resolu ion ia A en ion o Facial Landma k [ 21 ]. This me hod is also based on GAN, howe e , i uses he p og essi e me hod o upscaling he image. Mo eo e , his wo k in oduces a new acial a en ion loss, which allows us o es o e acial landma ks. A ecen ly in oduced me hod PULSE [ 22 ], which was p esen ed a he CVPR 2020 h p: //c p 2020. hec .com/), has e y encou aging esul s. Un o una ely, some c i icism ( o example, Twi e pos (h ps:// wi e .com/AlexVasilescu/s a us/1277009168319143936): “ PULSE: Medical and mili a y decision based on hallucina ed pixels?”) o he me hod show he disad an ages o he PULSE, and in gene al al eady men ioned lack o GANs—i s c ea i eness. The e is a demons a ion which uses a low- esolu ion image o Ba ack Obama and i s SR al e na i e p ocessed by his me hod. I c ea ed a high-quali y ace image wi h no ob ious signs ha his ace was econs uc ed by supe - esolu ion. Un o una ely, i econs uc ed a o ally di e en pe son. 2.3. Mul i-F ame Supe -Resolu ion Supe - esolu ion has a ac ed g ea a en ion, especially in he a ea o he en e ainmen indus y he ideo, o in o he wo ds mul i- ame. The e ha e been in oduced plen y o wo ks which in oduce me hods based no jus on a single inpu image, bu on a sequence o images. These me hods ha e he po en ial o educe noise and inc ease he quali y o he images, especially in he a ea o con e sion o old ideos in o HD esolu ion whe e i had g ea success. One o he la es wo ks ela ed o gene al (non- acial ela ed) mul i- ame supe - esolu ion is DeepSUM [ 2 ]. I is based on CNN, which exploi s spa ial and empo al co ela ions. This app oach includes image egis a ion inside he CNN a chi ec u e. I dynamically compu es and applies cus om il e s o highe -dimensional image ep esen a ions, ins ead o compensa ing o he mo ion as Appl. Sci. 2020,10, 7213 5 o 27 a p e-p ocessing s ep as is common in mos mul i- ame supe - esolu ion app oaches. The da ase om he challenge P oba-V [ 23 ] was used, which is dealing wi h sa elli e images and alloca es impo an pa s using segmen a ion based maps. Ano he me hod o mul i- ame supe - esolu ion is desc ibed in [ 24 ]. Fi s , all inpu low- esolu ion ames a e upscaled by he ResNe [ 25 ] wi h scale ac o 2. Pa allel o ha , he low- esolu ion ames go h ough he image egis a ion block o de e mine he sub-pixel shi s. The nex s ep is he applica ion o he shi -and-add usion o ob ain he ini ial econs uc ed image. The inal s ep is he applica ion o he E oIM p ocess, which consis s o i e a i e il e ing o he ini ial econs uc ed image. This me hod is again gene al pu pose and no acial supe - esolu ion ela ed. The mos equen u iliza ion o he mul i- ame supe - esolu ion is upscaling he ideo esolu ion. F ame- ecu en Video Supe - esolu ion [ 26 ] is an end- o-end ainable ame- ecu en ideo supe - esolu ion (FRVSR) amewo k. I uses he p e iously es ima ed high- esolu ion ame as an inpu o i s ollowing s ep. Each ame is p ocessed only once, which allows us o educe he compu a ional cos . The model consis s o se e al s eps: low es ima ion, upscaling low, wa ping p e ious ou pu , mapping o low- esolu ion space, supe - esolu ion. I is e y impo an o keep he p e iously es ima ed high- esolu ion ame in he sys em, o he wise, i is no possible o pass in o ma ion o u u e es ima ions, which can be c i ical o he econs uc ion o he nex ames. TecoGAN [ 27 ] is he a chi ec u e p oposed o sol e he ollowing ideo gene a ion asks: Video supe - esolu ion (VSR) and Unpai ed Video T ansla ion (UVT). The wo k [ 27 ] desc ibes he ad e sa ial lea ning me hod o a ecu en aining app oach, which u ilizes spa ial con en s and empo al ela ionships. The a chi ec u e uses a ame- ecu en gene a o and a spa io- empo al disc imina o . This app oach can gene a e e y ealis ic na u al images; howe e , i can lead o empo ally cohe en ye sub-op imal de ails. Enhanced De o mable Con olu ional Ne wo ks (EDVR) [ 28 ] is a amewo k p oposed in 2019, which allows us o do di e en image es o a ion asks, o example, supe - esolu ion and de-blu ing. I consis s o he ollowing pa s: an alignmen module—Py amid, Cascading and De o mable con olu ions (PCD), and a usion module—Tempo al and Spa ial A en ion (TSA). Mo eo e , he wo-s age s a egy is used o boos pe o mance and imp o e he quali y o ou pu ames. One o he i s a emp s ela ed o he acial supe - esolu ion me hod was in oduced in 2017 in [ 29 ]. The au ho s p oposed an a chi ec u e ha consis s o h ee modules: ea u e ex ac o , ace wa ping, and econs uc ion. This a chi ec u e es o es he cen al ame o each inpu sequence u ilizing sub-pixel mo emen and aking in o accoun a numbe o adjacen ames. This model was es ed on he YouTube Faces da ase , which is no dedica ed o supe - esolu ion pu poses. The au ho s downsampled all images h ough 128 × 128 px, images we e blu ed wi h he Gaussian ke nel (2.4), and again downsampled o he size 16 × 16 px. The downsampling me hod is in his case no clea and p obably he same me hod was used o all images, which oge he wi h a cons an blu , c ea es a high bias and does no e lec eal-wo ld scena ios. Fo his pu pose, he e is a isk o o e - i ing. 2.4. Mo i a ion Mos o he wo ks desc ibed ea lie a e ocused on gene al image supe - esolu ion. Un o una ely, hose me hods canno o en be used o pe son iden i ica ion since acial supe - esolu ion has sligh ly di e en equi emen s. Al hough he gene al supe - esolu ion me hods o en wo ks e y well o gene al images and some imes hey e en p o ide eally ou s anding esul s, hese me hods s ill c ea e some mis akes in he images. In he case o gene al images, hese e o s can o en be igno ed. Howe e , a mis ake in a ace image, whe e an eye o nose is missing o is de o med, e en i i migh be om he poin o pixel e o a mino e o , om he poin o iew o human pe cep ion, i can be s ongly dis up i e. In his case, e en a sligh modi ica ion o he image can ha e a signi ican impac on human pe cep ion o he ace and iden i ica ion o a pe son, as shown in Figu e 2. Appl. Sci. 2020,10, 7213 6 o 27 Figu e 2. Di e en ace supe - esolu ion me hods wi h scale ac o 8. F om he human pe cep ion poin o iew, i is be e o ha e checke boa d a i ac s (U-Ne +GEU) ins ead o a i ac s and de o ma ions om o he me hods. Ano he issue may be he c edibili y o he supe - esolu ion me hods om a biome ic pe spec i e. Fo biome ics, i is absolu ely unaccep able o be c ea i e and gene a e pa s o he image ha do no o igina e om i s inpu . The cu en end and di ec ion o esea ch o supe - esolu ion me hods ocus p ima ily on he gene al single- ame me hods. As he mul i- ame me hods ha e mo e in o ma ion on i s inpu , i is expec ed hey also p obably ha e a highe po en ial o each be e esul s ha a e based on he inpu in o ma ion and a e no c ea i ely illing he missing pa s o he image. Fu he mo e, conside ing he echnical capabili ies o mode n CCTV ha dwa e (e.g., ames independen ly on ps), i is e en possible o go beyond he possibili ies o humans and he e is an oppo uni y o each e en be e esul s in he u u e. Al hough cogni i e skills o he human b ain s ill signi ican ly exceed he skills o machines, na owly ocused a i icial in elligence has he po en ial o b ing new applica ions and, o example, make iden i ica ion mo e objec i e. Examples o hese applica ions could be an au oma ic sea ch o he bes acial image o a suspec o news announcemen s, econs uc ion o he ace using a sequence o images, and he supp essing o ad e se ligh ing condi ions. One o he obs acles in he de elopmen o he cu en acial mul i- ame supe - esolu ion me hods is he absence o a publicly a ailable da ase , which would be designed o his pu pose. Many speci ic equi emen s a e imposed on such a da ase . Da a o such a da ase should no be biased, i mus be o su icien size, i mus con ain eco ds om eal-wo ld en i onmen s including di e en ligh ing o wea he condi ions, da a mus be no dependen on a speci ic ideo came a ha dwa e, and i should e lec also eal-wo ld comp ession a i ac s, he so-called encoding agmen s. Con ained aces should co e a a ie y o human aces and a ious ages. In each case, he da a mus con ain a sequence o se e al aces, no jus one ame. Rega ding a ailable da ase s, he e is cu en ly no da ase ha mee s such equi emen s. Face images used o be biased because o used celeb i ies, which ends o cause he models o p edic “p e y” aces supp essing impo an indi idual ace ea u es (LFW, CelebA). The e is a lack o low quali y image inpu s, especially in he o m o sequences (PubFig, FERET) wi h a high-quali y label (You ube aces DB). All imp o emen s a e summa ized in Table 1. Table 1. The imp o emen s o he p oposed da abase in con as o well-known aining da ase s. Imp o emen LFW FERET PubFig CelebA You ube DB Numbe o samples × × × Numbe o iden i ies × × Samples as sequences × Supe - esolu ion pu pose Va iabili y (a oid bias) × × 3. Ma e ials and Me hods This sec ion desc ibes he aining and es ing da ase s, examined supe - esolu ion me hods, and e alua ion me ics. Each pa icula pa is desc ibed in de ail in he ollowing subsec ions. O e all, he scheme desc ibed in his sec ion and hei mu ual con ex a e depic ed in Figu e 3. Appl. Sci. 2020,10, 7213 7 o 27 Figu e 3. Scheme o he expe imen . The da ase consis ing o se e al low- esolu ion images and one high- esolu ion a ge image was used o aining and e alua ion o a couple o di e en supe - esolu ion me hods. These me hods we e e alua ed using se e al objec i e and subjec i e me ics. 3.1. T aining and Tes ing Da a Fo aining and e alua ion, we in oduce a new da ase called Mul i- ame Labeled Faces Da abase—MLFDB. The main di e ence o he exis ing da ase s is ha his da ase does no con ain jus a single acial image, bu i con ains a sequence o consecu i e ames in a ideo eco d. This sequence o images is in a low- esolu ion (32 × 32 pixels) and se es as an inpu o he supe - esolu ion me hods. Addi ionally o each o hose sequences, he e was added one ame in a highe esolu ion (64 × 64 pixels) which se es as a g ound u h, i.e., he op imal ou pu o he supe - esolu ion me hod (see Figu e 4). Figu e 4. An example o sample da a. Se en low- esolu ion images (32 × 32 px) and label (64 × 64 px). Wha should be emphasized he e is ha he aces in he da ase we e selec ed on he bo de line whe e iden i ica ion is di icul (i.e., each low- esolu ion ame o 32 × 32 pixels) and when i is easie o use i o iden i ica ion (i.e., highe esolu ion image wi h 64 × 64 pixels). The main objec i e is no o ha e he esul ing acial images good looking, bu a he ha hey can be used o iden i ica ion. The da ase is o a compa able size, as a e he majo ace da ase s (i con ains 17,426 aces and app oxima ely 6000–7000 unique pe sons). I ies o a oid he p oblem o bias, i.e., he eco dings a e om a ious eal-wo ld scena ios whe e aces a e nea by o each o he , a ace is pa ially co e ed including glasses, bea d, sca o o he co e ings (e.g., ace pain ings in some cases). The da ase also con ains a ious acial images including a ious ligh ing condi ions (i.e., high, low, and poo ligh ing quali y condi ions), a ious eco ding ha dwa e was used, ideos we e encoded using a ious algo i hms, aces con ain a wide ange o ages, aces we e eco ded om a ious poses, and hey con ain a ious emo ional exp essions (including smile, ea , ange o su p ise). I also Appl. Sci. 2020,10, 7213 8 o 27 co e s all main e hnic g oups including Eu opeans (i.e., Middle Eas e ne s and Medi e aneans), Eas Indians, Asians, Ame ican Indians, A icans, Melanesians, Mic onesians, Polynesians, Aus alians, and Abo igines. The c ea ed da ase was eleased online and can be used o ep oduce he expe imen desc ibed in his pape . Al hough i was p ima ily designed o mul i- ame p ocessing, i is expec ed i s u iliza ion will be wide . You ube Faces DB [ 9 ] can be conside ed as a simila da ase o in oduced MLFDB. Howe e , i di e s in se e al pa ame e s: You ube Faces DB con ains aces ha a e no on i s bo de line when iden i ica ion is possible and simple downsampling he acial images can b ing bias and o e - i ing. Mo eo e , i con ains only 1595 di e en people, bu MLFDB con ains app oxima ely 6000–7000 di e en people (see mo e de ails in Sec ion 5.1), and also i is no designed o supe - esolu ion pu poses. 3.1.1. Da ase Gene al In o ma ion The da ase was eleased online and o he objec i i y o he esul s, a pa o he da ase was kep p i a e oo. Publicly a ailable se s a e TRAIN (12,200 samples) and TEST (2600 samples) all oge he coun ing 14,800 ace sequences dedica ed o aining and e alua ion pu poses. The es o he da ase (2626 ace sequences) is ese ed o pe o mance e alua ion in p i a e mode. Each sample is o ganized in an indi idual olde and in e e y se , hey a e numbe ed om 1— n (TRAIN: 1—12200 , TEST: 1—2600). Each sample olde con ains: • 7 images (i.e., ace sequences) in a low- esolu ion (32 × 32 px) named ‚img_1.JPG, . . . , ‚‘img_7.JPG’. • The label is de ined om he middle o he sequence ( om img_4.JPG) wi h double esolu ion 64 ×64 px and ile name label.JPG. • ‘in o. x ’ ile wi h in o ma ion abou he label (bounding box and ame numbe ), ideo sou ce, and in e pola ion me hod oge he wi h he JPEG comp ession ha we e used o esizing inpu images. A inal sample con ains: a sequence o se en images, label, and in o. x ile (see he example in Figu e 4). 3.1.2. Da ase C ea ion The da ase o sequences o acial images was c ea ed wi h he help o he YOLO 3 objec de ec o (h ps://gi hub.com/s hanhng/yolo ace) [ 30 ]. By including only high-quali y samples, se e al il e s we e applied be o e he sequence was included in he inal da ase . The i s il e was he condi ion ha he de ec ed ace had o ha e a esolu ion g ea e han 64 × 64 px. Fo ep oducibili y pu poses and also o allow e e yone in he u u e o use he same da ase , all he de ails om whe e he aces we e cap u ed was p o ided. This in o ma ion includes he ace in ull esolu ion (used as a g ound u h image) oge he wi h i s bounding box coo dina es and ideo ame numbe . To assu e a iabili y o he da ase only a limi ed numbe o aces we e collec ed om each ideo, and he selec ed aces we e chosen andomly. Nex , aw ace sequences we e ex ac ed om he ideo using he OpenCV (h ps://pypi.o g/p ojec /openc -py hon/) lib a y. I sea ched o he co esponding ace ame (label), i c opped he ace acco ding o he gi en bounding box and inally, i also ook 3 ames be o e and 3 ames a e he c opped image using he same bounding box. This was an in en ional s ep on ou pa o make he da ase mo e ealis ic as he e a e many cases when ace de ec o s do no co ec ly c op he ace (inaccu a e de ec ion, c op scaling ac o ,. ..). Fu he op imiza ions a e le o la e p ocessing s eps. Un o una ely, no all he cap u ed samples can be used. Fo his eason, a ew o he s eps we e pe o med (see Figu e 5): 1. Fo each sequence, i is checked whe he he label con ains a ace o su icien quali y so i can be used o ace ecogni ion. This was done using an open-sou ce p ojec (Face- ecogni ion Appl. Sci. 2020,10, 7213 9 o 27 amewo k [ 31 ] 1.3.0, CNN model ) which is inspi ed by Facene [ 32 ]. This s ep ac ually alida es whe he he de ec ed ace is in good quali y so i can be used o he ace ecogni ion simila i y a e index me ic (see Sec ion 3.4.1). YOLO is a ace de ec o , no a ace ecogni ion sys em, he e o e, i also de ec s such aces whe e pe son iden i ica ion is impossible (e.g., head om behind). 2. In some cases, a sequence can con ain mul iple o e lapping aces. I is no a p oblem i mo e aces o hei pa s occu in one image, bu he p oblem a ises, when—mainly due o a mis ake o he p ocessing algo i hm— he e is a di e en pe son a he beginning o he sequence and ano he a he end. F om ime o ime i also happened due o a mo ie spli . The e o e, he label is compa ed using he Face- ecogni ion amewo k (used h eshold was 0.7) wi h all o he images in he sequence. Thus a possible p esence o se e al aces in a single image was aken in o accoun . 3. Due o used andomness in sampling he ideos, i can also happen om ime o ime ha one pe son appea ed mo e han once in he esul ing da ase . The bigges issue has been how o dis inguish be ween e y simila images and images wi h di e en ligh ing, ace angle, e c. I has been shown ha he Face- ecogni ion amewo k has no been an app op ia e solu ion o his ask due o i s abili y o ace alignmen , i.e., i can pe o m he ecogni ion in di e en condi ions. We decided o use a s uc u al simila i y me ic—SSIM (MSE, PSNR, e c., a e no e icien me ics o his case) wi h he h eshold 0.7. G ound u h images a e used o his which a e i s esized in o a 64 ×64 px esolu ion. 4. The nex s ep was abou alida ion whe he all he sequence images con ain a ace. This was done using he Face- ecogni ion amewo k and i s me hod o ace de ec ion (no ecogni ion). By using his, only hose g ound u h images, which can be used o ecogni ion a e included. O he images in he sequence mus con ain a ace, bu do no ha e o be ecognizable (e.g., a di e en angle, pa ially co e ed by ano he pe son, e c.). 5. The sequences which a e o iginally o a ious esolu ions (highe han 64 × 64 px) a e downsampled in o esolu ion 32 × 32 px and he g ound u h (label) image in o a 64 × 64 px esolu ion. When he o iginal esolu ion allowed his, he labels we e also eco ded in esolu ion 128 × 128 px and 256 × 256 px). Each sequence andomly used a chosen me hod o pixel in e pola ions such as he nea es neighbo , bilinea , bicubic, lanczos (o e 8 × 8 neighbo hood). The label was sa ed wi h he bes possible JPEG quali y (comp ession 100) and sequence images we e sa ed by a andomly chosen JPEG comp ession scale om 30–90. 6. Un o una ely, a e esizing he g ound u h images (labels) in o esolu ion 64 × 64 px some o he images we e no possible o use o ace ecogni ion. Those sequences o ace images we e emo ed om he da ase . Appl. Sci. 2020,10, 7213 16 o 27 ou mul i- ame me hods, and we called hem U-Ne +ResBlock, U-Ne +GEU, U-Ne +GEU2, and U-Ne +GEU3. Some o hem ake inspi a ion om he mos success ul single- ame based me hods and a e modi ied o wo k wi h mul iple ames. These a chi ec u es we e p e-selec ed a e expe imen a ion wi h many di e en combina ions o a chi ec u es and hei blocks, whe e hose ou ha e been shown o be he mos p omising. Fo comple eness, se e al gene ally known in e pola ion me hods we e also included in he expe imen o he pu pose o compa ison wi h he ea lie men ioned me hods. Gaussian il e was applied o he ou pu o U-Ne +GEU o supp ess checke boa d a i ac s [ 40 ] and due o an in e es ing esul s compa ison, i is conside ed as an indi idual me hod. Thus, a o al o 12 di e en supe - esolu ion me hods we e e alua ed and compa ed. Al hough GAN ne wo ks ha e been e y success ul in his a ea in ecen yea s, he p oposed mul i- ame a chi ec u es we e no included in he expe imen . The eason o his exclusion was ha i was shown hey a e oo c ea i e as was discussed in Sec ion 2.2. Al hough he aces look igh , hey canno be used o pe son iden i ica ion since hey gene a e comple ely ake aces, which a e no p ima ily based on hei inpu in o ma ion. In Sec ion 5) he esul s a e discussed. 4.1. Resul ing Da ase Based on he p ocess desc ibed in Sec ion 3.1.2, we c ea ed an MLFDB da ase ha con ains a o al o 17,426 samples. Each sample con ains a sequence o 7 low- esolu ion ace images (32 × 32 px). Each o hese sequences has i s co esponding label in he 64 × 64 px esolu ion. The label was c ea ed om he middle image o he sequence ( ame numbe 4). When i was possible, labels in he highe esolu ion we e also p o ided. The e o e, some o he sequences also ha e labels in 128 × 128 px (12,165 samples o he whole da ase , i.e., 70%) and 256 × 256 px (4199 samples, i.e., 24%) esolu ion. An example o da ase samples is shown in Figu e 9. Figu e 9. An example o sequence samples o he da ase . This da ase o 17,426 samples was di ided in o a aining se (12,200 samples, 70%), es se (2600 samples, 15%), a p i a e es se (2500 samples, 15%), and a es se o a ques ionnai e (126 samples) o subjec i e human e alua ion. The aining da ase is in ended o he aining and op imiza ion o supe - esolu ion me hods. The es se is only o he e alua ion o esul s. This da ase should no be used o op imiza ion o sea ch o pa ame e s o a oid o e - i ing. The p i a e es se is kep p i a e a he B no Uni e si y o Technology and will be used o e alua ion on eques o ensu e he objec i i y o he esul s (h p://splab.cz/ml db/# esul s). The ques ionnai e da a se is in ended o e alua ion by human e iewe s. This is mainly because esul s measu ed by objec i e ma hema ical me hods do no always co espond wi h he quali y pe cei ed by human e alua o s. Appl. Sci. 2020,10, 7213 17 o 27 4.2. Me hods Compa ison All objec i e measu es desc ibed in Sec ion 3.4.1 we e used o he pe o mance e alua ion o he models: SSIM, MSE, PSNR, sha pness as he di e ence be ween he g ound u h image and he p edic ed image, ace ea u es ex ac ion ail a e (FR ailed) and ace ecogni ion simila i y a e index, i.e., FR a e. Please no e ha FR a e (all) is compu ed only om he images whe e he ace ea u es ex ac ion was success ul o he gi en esul ing me hod. FR a e ( ∩ ) me ic is compu ed om he in e sec ion o he images wi h success ul ace ea u es ex ac ion by all me hods. The amoun o hese images is 1518 o all 2500 es images. The esul s o all 12 me hods a e placed all oge he in o Table 2 o he compendious compa abili y o esul s be ween he single- ame and mul i- ame app oaches. Table 2. Single- ame and mul i- ame me hods pe o mance compa ison—mean. Me hod SSIM MSE PSNR Sha pness FR Ra e (All) FR Ra e (∩) FR Failed bicubic 0.806 227.196 25.654 0.084 0.476 0.473 674/2500 bilinea 0.809 217.368 25.769 0.158 0.471 0.468 693/2500 lanczos 0.802 238.247 25.504 0.024 0.483 0.480 700/2500 EDSR 0.780 286.999 24.687 −0.025 0.495 0.494 701/2500 SRCNN 0.815 216.940 25.771 0.209 0.461 0.456 643/2500 SRGAN 0.760 320.981 24.234 −0.078 0.511 0.510 771/2500 ESRGAN 0.788 281.192 24.604 −0.016 0.488 0.485 470/2500 U-Ne +GEU 0.760 276.459 24.383 −0.129 0.507 0.503 317/2500 U-Ne +GEU+ il e 0.815 234.552 25.196 0.239 0.464 0.458 394/2500 U-Ne +GEU2 0.799 256.511 24.878 0.031 0.472 0.467 259/2500 U-Ne +GEU3 0.799 243.433 25.135 0.004 0.471 0.466 270/2500 U-Ne +ResBlock 0.812 239.114 25.141 0.130 0.468 0.463 581/2500 As can be seen om he esul s, he SRCNN me hod can be conside ed as he mos sui able me hod in he way o he highes numbe o bes objec i e me ics ( he bes esul s o 5 om 7 me ics), while no conside ing subjec i e human e alua ion. Howe e , he sha pness and he FR ailed me ics o he SRCNN me hod a e signi ican ly wo se when compa ing he esul s o o he me hods. The bes ace ecogni ion ex ac ion ail a e (FR ailed) has he U-Ne +GEU2 me hod ha ailed only in 259 cases ou o 2500 compa ed o he SRCNN me hod ha ailed in 643 cases ou o 2500. Mo eo e , he sha pness di e ence be ween hese wo me hods is no iceable. I should be no ed ha he esul s desc ibed abo e only p esen s he pe o mance o me hods acco ding o he numbe o he bes objec i e measu es and i does no ake in o accoun hei signi icance (i.e., SSIM and FR ailed a e p obably mo e aluable han MSE). The impo an me ic—subjec i e human e alua ion—is no included in hese esul s, he e o e, hey should no be gene ally aken as he inal ou comes. U-Ne +GEU2 signi ican ly ou pe o med o he me hods by he subjec i e human e alua ion (see esul s in Sec ion 4.3), bu among he objec i e me ics i had only one bes esul — ace ea u es ex ac ion ail a e (FR ailed). The o he alues a e no so good compa ing o he SRCNN me hod, bu he di e ences be ween hem a e signi ican unlike o he ace ea u es ex ac ion ail a e me ic whe e he di e ence is e iden . Because o his, and due o he di e en impo ance o each me ic, hese alues and hei compa ison a e be e eadable by a no malized ep esen a ion in o he ange (0–1) (see Table 3). Minimum and maximum alues we e de e mined om all me hod’s alues o he gi en me ic. In he case o he obse ed me ic, whe e he highe alue is be e (SSIM, PSNR, subjec i e e alua ion), he no malized alue is compu ed as 1 −no m(x) o ha e all alues wi h he same logic—lowe alue is be e (i is conside ed as a “penaliza ion”). The inal sco e is hen simply he sum o all no malized alues, sum1 includes only he objec i e me ics and sum2 also includes he subjec i e human e alua ion—sub. e . Appl. Sci. 2020,10, 7213 18 o 27 Table 3. No malized alues compa ison o all me hods o ge he inal sco e. Me hod SSIM MSE PSNR Sha p. FR Ra e (∩) FR Failed Sub. e . Sum1 Sum2 bicubic 0.164 0.099 0.076 0.340 0.315 0.811 1 1.805 2.805 bilinea 0.109 0.004 0.001 0.655 0.222 0.848 0.989 1.839 2.828 lanczos 0.236 0.205 0.174 0.085 0.444 0.861 0.999 2.005 3.004 EDSR 0.636 0.673 0.705 0.703 0.089 0.863 0.977 3.669 4.646 SRCNN 0 0 0 0.872 0 0.750 0.978 1.622 2.600 SRGAN 1 1 1 1 0.315 1 0.995 5.315 6.310 ESRGAN 0.491 0.618 0.759 0.051 0.537 0.412 0.844 2.868 3.712 U-Ne +GEU 0 0.572 0.903 0.532 0.870 0.113 0.570 2.990 3.560 U-Ne +GEU+ il e 0 0.169 0.374 1 0.037 0.246 0.578 1.844 2.422 U-Ne +GEU2 0.281 0.380 0.581 0.115 0.204 0 0 1.571 1.571 U-Ne +GEU3 0.281 0.255 0.414 0 0.185 0.021 0.204 1.166 1.370 U-Ne +ResBlock 0.055 0.213 0.410 0.536 0.130 0.629 0.834 1.973 2.807 As can be seen om Table 3, a e a desi able ep esen a ion o he esul s and hei o e all sco e in o m o he sum, U-Ne +GEU2 and U-Ne +GEU3 me hods ha e a be e o e all sco e han he SRCNN me hod e en wi hou using he subjec i e human e alua ion me ic (sum1). The o e all sco e including he subjec i e human e alua ion me ic (sum2) enla ges he di e ence be ween he me hods almos wice as much. The U-Ne +GEU+ il e me hod has a wo se o e all sco e (sum1) han he SRCNN me hod (1.844 s. 1.622) bu he esul s a e opposi e (2.422 s. 2.600) a e including he subjec i e human e alua ion (sum2). An example o esul ing images (p edic ions) on unseen da a om all 12 me hods is shown in Figu e 10. Each ow ep esen s one inpu sequence and each column shows esul ing images o each me hod. Finally, he e is a g ound u h image. The p esen ed samples a e no pa o he MLFDB es se because he labels (g ound u h images) should no be publicly a ailable, he e o e, hese sequences we e c ea ed o p esen a ion pu poses only. Figu e 10. Examples o he esul ing me hods. All se en 32 × 32 px images o he inpu sequence we e used o he mul i- ame me hods and he 4 h image o single- ame app oaches was used. Appl. Sci. 2020,10, 7213 19 o 27 As expec ed, he in e pola ion-based me hods canno econs uc he image in a be e quali y, bu he di e ences be ween hem a e no iceable. Resul s o he EDSR and SRGAN me hods look simila o he in e pola ion-based me hods. The SRCNN me hod, which has he bes esul s among objec i e me ics, p edic s he esul ing image a li le be e , compa ed o o he single- ame me hods, unlike he ESRGAN whose p edic ion quali y is much be e (sha pe ), bu i some imes c ea es isible de o ma ion a i ac s. All mul i- ame me hods p o ide no ably be e esul s in con as o single- ame me hods. The U-Ne +ResBlock me hod c ea es a bi smoo h blu y images, which is annoying o human pe cep ion. The U-Ne +GEU+ il e me hod c ea es simila esul s, bu hey look no so ema kable. The esul ing images o o he U-Ne +GEU based me hods ac simila ly o each o he bu a e signi ican ly be e compa ed o all o he me hods. 4.3. Ques ionnai e Resul s The subjec i e human e alua ion me ic was examined by a ques ionnai e in which 100 people had o selec he bes image in hei opinion, among 126 compa isons o 12 me hods. In o al 12,600 “ o es” we e a ailable and hey we e dis ibu ed o all me hods by indi idual p e e ences o a ending people. The isualiza ion o subjec i e human e alua ion esul s is illus a ed in Figu e 11. Figu e 11. Resul s o he ques ionnai e—subjec i e human e alua ion. 100 people a ending illed ou he ques ionnai e con aining 126 compa isons o 12 me hods. F om he i s look a he cha , i is clea ha all U-Ne +GEU me hods signi ican ly o e come o he me hods in human pe cep ion. 3815 o es om 12,600 belong o he U-Ne +GEU2 me hod. U-Ne +ResBlock and ESRGAN wi h app oxima ely 700 o es a e ollowing he U-Ne +GEU base me hods. ESRGAN signi ican ly exceeds all single- ame me hods and as only one single- ame me hod can be compa ed o esul s achie ed by he mul i- ame me hod—U-Ne +ResBlock, which is he wo s among mul i- ame me hods, by human pe cep ion. Rema kably, ha SRCNN has he bes objec i e me ics, bu ESRGAN is much mo e success ul by human pe cep ion. A simila si ua ion is he same o mul i- ame me hods whe e he U-Ne +GEU+ il e has be e objec i e me ics, bu he U-Ne +GEU2 has a signi ican ly be e subjec i e human e alua ion. Appl. Sci. 2020,10, 7213 20 o 27 5. Discussion 5.1. Da ase Un o una ely, i is qui e challenging o au oma ically coun unique aces in his da ase . Due o he high numbe o sequences, i is almos impossible o emembe all he aces and he manual app oach is, he e o e, no possible as well. Fo his eason, we used an au oma ed app oach which is based on he Face- ecogni ion amewo k and can measu e a simila i y o aces. This amewo k e u ns ze o alue when wo aces a e o he pe ec simila i y. A highe alue han ze o means less simila i y o hese aces, which migh be caused no jus by di e en aces, bu also because aces we e aken om di e en angles, ha e lowe quali y, e c. Acco ding o he p ojec au ho s, he de aul h eshold alue by which he simila i y alue o wo aces should exp ess he same pe son is ≤ 0.6. This alue is ecommended o acial ecogni ion pu poses (compu ed on he LFW da ase ). Howe e , as can be seen in Table 4, a h eshold alue o 0.6 is imp ope o his pu pose. The eason o his is ha MLFDB con ains ace images ha a e on he bo de line when hey s a being use ul o he iden i ica ion (i.e., lowe quali y han he LFW da ase ). Fo he h eshold alue o 0.6, only 76 sequences we e conside ed unique, which is a e manual e iew and de ini ely no he u h. Based on he empi ical expe imen a ion wi h di e en alues, we ound he h eshold alue 0.45, which shows good pe o mance. These esul s we e hen alida ed manually on a selec ed subse . Table 4. Di e en people in he da ase using he Face- ecogni ion amewo k simila i y h eshold. Th eshold 0.6 0.5 0.45 0.4 0.35 di e en people 76 2378 6651 10,822 13,369 Sequences om one ideo sou ce we e also il e ed using he SSIM me ic, whe e we used he h eshold alue 0.7. Only he g ound u h images we e used o his il e ing. This il e ensu es ha sequences e en o he same pe son a e kep when he e is a di e en scale, ligh ing, angles, e c. (i was discussed in Sec ion 3.1.2). The da ase is u he di ided in o public and p i a e pa s. The eason o his is o inc ease he objec i i y o pe o mance e alua ion. In he case all he pa s would be published, hey can s ill be used o he aining. Such esul s would wi h high p obabili y ou pe o m he esul s o o he eams, bu hose esul s would be o e - i ed and would absen gene aliza ion o cou se. Because o ha , he es pa o he da ase was spli in o wo pa s: he es -public and es -p i a e. The es -p i a e was also eleased bu wi hou g ound u h images. E alua ion o his p i a e da ase is possible on he MLFDB websi e by submi ing achie ed esul s. The e alua ion eques s a e limi ed by ime o a oid al eady men ioned o e - i ing. 5.2. Ques ionnai e The p oblem o ace image supe - esolu ion is signi ican ly mo e challenging han gene al supe - esolu ion asks. I is clea ha in e e y supe - esolu ion me hod, some pa icula e o s and mis akes mus be c ea ed since he inpu image con ains less in o ma ion han he e is expec ed in he esul ing image. Thus, some pe cen ages o he pixels a e es ima ed. In he case o gene al supe - esolu ion me hods, a mis ake in he image is o en ha d o be no iced and can be igno ed. Howe e , when he image is checked o some de ails, ypically ex s, b and logos, o o he well-known complex shapes, some e o s a e appa en a i s glance. This is also he case o acial images. Fo example, i he esul ing image o a ace is missing an eye, o i has h ee eyes, i migh be o me ics like PSNR o MSE neglec able e o , bu o human pe cep ion, i can be e y dis up i e. The same si ua ion is wi h wo noses o no nose, he asymme ical appea ance o he eyes o o he wise de o med ace. Appl. Sci. 2020,10, 7213 21 o 27 Ano he opic ha should be aken in o conside a ion is ha he main objec i e is no jus o gene a e images ha a e good looking. A mo e impo an c i e ion is ha he esul ing acial image should be o such a quali y ha he iden i ica ion mus be possible. The s udied images a e a he bo de line whe e images s a being use ul o he ecogni ion. When conside ing each pa icula low- esolu ion image om a sequence alone, i is no possible o i is di icul o ecognize a pe son. Ne e heless, he esul ing highe esolu ion image should be o su icien quali y, in which i can be used o his pu pose. Fo his eason, we add essed se e al olun ee s and asked hem o ill ou he ques ionnai e. The pu pose o he ques ionnai e was o ob ain eedback ega ding he quali y o he achie ed ace images om he human pe spec i e. The olun ee s saw all he esul s o he supe - esolu ion me hods, and hey we e asked o ma k he esul , which hey conside is he bes . The U-Ne based me hods we e hei main choices, and he in e pola ion-based me hods we e almos no selec ed. An in e es ing inding was ha he SRGAN me hod was conside ed as he wo s one. The decisions we e in gene al made based on he essen ial ace ea u es clea ness (same eyes, nose, mou h)— he en i e y o he ace. Howe e , some cases deg aded ee h and ace bo de s so much ha he p e e ed image was wi h a lowe quali y image bu wi h an en i e ace. The ESRGAN me hod was he only me hod ha had signi ican de o ma ions o he ace. Ano he in e es ing in o ma ion was ha people in ui i ely p e e blu ed edges ins ead o sha p edges and no a i ac s, e en when he image looks good. These eedbacks co ela e wi h he objec i e esul s, excep o he ESRGAN me hod. I is he second sha pes me hod and i has go he mos o es among all single- ame me hods om he ques ionnai e. 5.3. Resul s Compa ison Resul s we e e alua ed by wo app oaches, using objec i e ma hema ical me ics (see Table 2) and subjec i e e alua ion based on he human olun ee s’ o es (see Figu e 11). Fu he mo e, since e e y objec i e me ic e alua es he images using a di e en poin o iew, we ha e me ged all he esul s (using hei no malized ep esen a ion) and show he o e all sco e which is deno ed as sum1 and sum2 (see Table 3). Acco ding o he subjec i e me ics, he mul i- ame me hods achie ed he bes esul s. Fu he mo e, a e he no maliza ion o he objec i e me ics alues and imp o ed ep esen a ion (ins ead o absolu e alues) hey ha e had e en be e han he SRCNN me hod, which was e alua ed as he bes as poin ed in Table 2. The p obable eason is ha he esul s o indi idual me ics o mul i- ame me hods a e e y close o he bes alues among he es ed me hods (basically he SRCNN me hod) and p ima ily hey do no ha e any signi ican ou lie s among all me ics compa ed o single- ame me hods, especially he SRCNN me hod. No GAN-based mul i- ame supe - esolu ion me hods we e used in his wo k. The eason o his is hei gene a i e na u e, which was discussed in mo e de ail in Sec ion 2.2. In he scope o his wo k, he e we e conduc ed expe imen s wi h GANs, bu hese expe imen s concluded, he cu en me hods should no be used o pe son iden i ica ion pu poses. I should be emphasized ha gene al supe - esolu ion and supe - esolu ion o iden i ica ion pu poses ha e di e en objec i es. Nowadays, one o he obs acles o he as e de elopmen o hese me hods is he lack o objec i e me ics which will be e co espond o human pe cep ion. In his pape , we in oduce an ex ended se o me ics like SSIM, PSNR, e c. Un o una ely, i was no a a e case when he bes esul s selec ed by he human olun ee s did no ma ch he op esul s by he objec i e me ics. Because o ha , he deep lea ning models c ea ed images wi h small checke boa d a i ac s such as small di e en colou in s o small ace e-posi ioning, which a ec hese me ics. Those issues become e en mo e impo an i he e a e no g ound- u h images a ailable, i.e., he model is used in eal si ua ions. On he o he hand, i is p obably be e o ma k se e al suspicious ace images wi h lowe accu acy, a he han no selec ing any o he aces. Fo example, i is be e o p eemp i ely check en people ins ead o no-one when looking o dange ous suspec s. Appl. Sci. 2020,10, 7213 22 o 27 5.4. Face Recogni ion Based Me ics Fo all 2500 es images, he ace ecogni ion simila i y a e index (FR a e (all)) was compu ed. The same me ic was also compu ed o he subse o he es da ase whe e ace ecogni ion was success ul (1518 samples). Only hose samples a e conside ed o be success ul, whe e all he models c ea ed such a esul ing image, which succeeded wi h ace ecogni ion. The e was an assump ion ha he me hods wi h be e ace ea u es ex ac ion ail a e (lowe FR ailed) will ha e a be e ace ecogni ion simila i y a e index on hose subse images ins ead o he whole es se . Howe e , acco ding o achie ed esul s, i seems ha his has only a minimal e ec . A mo e in e es ing me ic is he ace ea u es ex ac ion ail a e (FR ailed), i.e., in how many cases i was no possible o ex ac success ully ace ea u es. Acco ding o he esul s, he U-Ne +GEU2 me hod has he lowes numbe o cases whe e ace ecogni ion ailed. I was also examined in how many cases i ailed on he same images. The numbe o di e en images ha a e no in he se o ailed images o he U-Ne +GEU2 me hod is used o make he ou pu mo e in e p e able. These alues a e shown as he i s alue in Table 5. The second alue a e he slash ep esen s he o e all di e ence om he U-Ne +GEU2 FR ailed me ic. Table 5. Face ecogni ion (FR) ailed di e ences compa ed o he U-Ne +GEU2 me hod. Me hod Di e ences bicubic 12/415 bilinea 9/434 lanczos 16/441 EDSR 18/442 SRCNN 11/384 SRGAN 15/512 ESRGAN 35/211 U-Ne +GEU 38/58 U-Ne +GEU+ il e 16/135 U-Ne +GEU3 55/11 U-Ne +ResBlock 10/322 The bes ma ch o ailed images is pa adoxically o he bilinea in e pola ion. On he o he hand, he bilinea in e pola ion ailed in 693 cases ins ead o he U-Ne +GEU2 me hod ha ailed in 259 cases ( he di e ence is 434) and he bilinea in e pola ion se o ailed images is big enough o ma ch he majo i y o he U-Ne +GEU2 se o images. An in e es ing inding is an opposi e issue. The U-Ne +GEU3 me hod has almos he same numbe o FR ailed me ic ( he di e ence is only 11 cases), bu i s se o images di e s in 55 cases compa ed o he U-Ne +GEU2 me hod. A e u he examina ion o he me hod, he e we e no pa e ns ound ( he same in e pola ions, comp ession, images backg ound, e c.), which would explain why he e a e so di e en esul s. Howe e , he e a e some di e si ies be ween U-Ne +GEU2 and U-Ne +GEU3 me hods in some cases. Figu e 12a shows a de ailed iew o he esul ing quali y c ea ed using U-Ne +GEU2 and U-Ne +GEU3. Su p isingly, he e a e also con a y cases, bu hey a e no so signi ican as illus a ed in Figu e 12b. U-Ne +GEU3 me hod ailed in ace ea u es ex ac ion in bo h cases. The e a e p obably wo essen ial explana ions. Fi s , human pe cep ion di e s om he machine lea ning image p ocessing and how he quali y o he inpu image is assessed. Second, he well-known and used Face- ecogni ion amewo k does no wo k as well as p esen ed, a leas o his case. Among he images which ailed in ace ea u es ex ac ion a e a couple o p o ile images (see examples o labels in Figu e 12c). The ace ea u es ex ac ion wo ks in his case, bu wi h wha accu acy? How do ace landma k es ima ions and ace ans o ma ions wo k in hese cases? Appl. Sci. 2020,10, 7213 23 o 27 Figu e 12. Examples o low ace ea u e ex ac ion pe o mance. ( a , b ) U-Ne + GEU2 (passed) and U-Ne + GEU3 ( ailed) me hods compa ison. ( c ) Examples o p o ile ace images o da abase samples ha passed ace ea u es ex ac ion as well du ing he da abase c ea ion il a ion s ep. 5.5. Benchma k and Leade boa d MLFDB da ase is he i s mul i- ame acial da ase o compa able size, which has he po en ial o suppo he de elopmen o his ield. Ano he impo an s ep is o se a baseline in he o m o esul s om a ious me hods, so anyone can compa e he achie ed esul s wi h new me hods. The gene al p oblem o many da ase s is ha hey can be easily o e - i ed. This o en happens when pa ame e s a e op imized no jus using he aining se , bu also he es se . In he si ua ion when e e y hing is publicly a ailable, i is no possible o ha e con ol o e his. Millions o ails can lead in some cases o esul s ha do no e lec he eal pe o mance o he me hod. Some imes he e is no clea ly de ined how he compa ison me ics should be compu ed, so he p og ess is di icul o be objec i ely e alua ed. To deal wi h his p oblem, we decided o c ea e an au oma ic benchma k (inspi ed by sys ems like Kaggle (h ps://www.kaggle.com/)) o an objec i e esul s compa ison on his da ase . G ound u h images o he p i a e pa o he es da ase a e no publicly a ailable, and hus me ics a e compu ed in he same manne as hey we e compu ed in his pape . Fo his pu pose, he MLFDB se e is a ailable online, whe e anyone can submi hei esul s (h p://splab.cz/ml db/). The inal sco e will be compu ed as sum1 desc ibed in Sec ion 4.2. The numbe o esul s o compu a ion que ies will be limi ed acco ding o he ules o he benchma k. 5.6. Fu u e Resea ch The e a e plen y o ways how o ollow up on his wo k. Especially, i could be in e es ing o expe imen wi h models wi h a highe scale ac o , o example, 4 × o 8 × , as was poin ed ou in he example in Figu e 2. The MLFDB da ase has all g ound u h images in esolu ion 64 × 64 px, bu he e a e also many images in e en highe esolu ion—128 ×128 px (12,165 samples) o 256 ×256 px (4199 samples). Ano he way wo h ying could be changing he numbe o images in sequences o a mo e complex and be e u iliza ion in eal si ua ions would be an applica ion o some ea u e selec ion [ 47 ]. The ea u e selec ion in gene al leads o a be e pe o mance because o ocusing on he impo an pa s o inpu da a and emo ing ou lie s o noisy da a ha cause model inaccu acy. 6. Conclusions Fo ensically ained acial e iewe s a e s ill conside ed o be one o he mos accu a e app oaches o pe son iden i ica ion, especially in he case o low- esolu ion o low-quali y ideos. The human b ain can u ilize in o ma ion no jus om a single image bu also om a sequence o aces (i.e., ideos) and, e en in he case o low-quali y eco ds o a long dis ance om a came a. They can accu a ely iden i y a pe son. Fo compu e me hods, his emains a challenge. Howe e , on he o he hand, hey ha e he po en ial o suppo human e iewe s and help o p e-p ocess he da a and ex ac as much in o ma ion om he da a as possible. One o such a use-case would be, o example, o econs uc he acial image in highe quali y om a ideo, which can police use o sea ch and announce i in newspape s. This pape in oduced a la ge-scale ace da ase con aining 17,426 sequences o ace images. The da ase co e s di e en aces, ages, ligh ing condi ions, and many ypes o di e en came a de ices. Reco ds we e ob ained om he eal en i onmen , including common de ec s and impe ec ions ha occu in he ideos (e.g., comp ession a i ac s, blu , e c.) Acco ding o ou knowledge, i is he i s Appl. Sci. 2020,10, 7213 24 o 27 da ase o a compa able size con aining sequences o ace images. The pape also in oduces a new mul i- ame ace supe - esolu ion me hod. This me hod has been p o en o p oduce be e esul s han single- ame me hods which a e conside ed as s a e-o - he-a , and hey a e one o he mos e ec i e and o en compa ed me hods in his esea ch a ea. The e o e, he hypo hesis ha he mul i- ame me hod p oduces be e esul s han he deep lea ning-based single- ame me hods was p o en. The esul s we e also e alua ed using se e al objec i e me ics and also subjec i ely by olun ee s whe e he U-Ne +GEU2 me hod has he bes esul s ega ding human pe cep ion, and he U-Ne +GEU3 me hod achie ed he bes o e all sco e o objec i e me ics. The p oposed me hod is no gene a i e and i was p o en o imp o e he quali y o he ace image. The sou ce code and he da ase we e eleased, and he expe imen is ully ep oducible. The new and unique MLFDB da ase has been published o s udying ace supe - esolu ion om sequences o images (i.e., mul i- ame p oblem). The benchma k wi h gi en ules and me ics was c ea ed, and esul s achie ed in his pape we e used as an ini ial s a ing poin and as a challenge o o he esea che s. They a e p esen ed in he leade boa d able as a pa o he benchma k. Rega ding esul s hemsel es, i was ound ha human pe cep ion plays a signi ican ole du ing model e alua ion. The e o e, i is essen ial o use some subjec i e human e alua ion, o example, in he o m o ques ionnai es o su eys. Au ho Con ibu ions: Concep ualiza ion, M.R. and R.B.; me hodology, A.M.; so wa e, M.R. and A.M.; alida ion, M.R. and R.B.; o mal analysis, A.M. and M.R.; in es iga ion, A.M.; esou ces, R.B.; da a cu a ion, M.R. and A.M.; w i ing—o iginal d a p epa a ion, M.R.; w i ing— e iew and edi ing, R.B.; isualiza ion, M.R. and A.M.; supe ision, R.B. All au ho s ha e ead and ag eed o he published e sion o he manusc ip . Funding: This esea ch was unded by he In e eg Cen al Eu ope niCE-li e by he g an CE1581. Acknowledgmen s: Resea ch desc ibed in his pape was inanced by he In e eg Cen al Eu ope niCE-li e by he g an CE1581. We would like o hank all o he olun ee s ha a ended he ques ionnai e, especially he ones who p o ided he eedback. Con lic s o In e es : The au ho s decla e no con lic o in e es . Abb e ia ions The ollowing abb e ia ions a e used in his manusc ip : CelebA La ge-scale CelebFaces A ibu es CCTV Closed ci cui ele ision CNN Con olu ional Neu al Ne wo k CPBD Cumula i e P obabili y o Blu De ec ion CVPR Con e ence on Compu e Vision and Pa e n Recogni ion DB Da aBase DeepSUM Deep neu al ne wo k o Supe - esolu ion o Un egis e ed Mul i empo al images EDSR Enhanced Deep Supe -Resolu ion Ne wo k ESRGAN Enhanced Supe -Resolu ion Gene a i e Ad e sa ial Ne wo k EDVR Enhanced De o mable Con olu ional Ne wo k FERET Face Recogni ion Technology FR Face Recogni ion FRVSR F ame- ecu en ideo supe - esolu ion FSRGAN Face Supe -Resolu ion Gene a i e Ad e sa ial Ne wo k FSRNe Face Supe -Resolu ion Ne wo k GAN Gene a i e Ad e sa ial Ne wo k GPU G aphics P ocessing Uni JNB Jus No iceable Blu LFW Labeled Faces in he Wild MLFDB Mul i- ame Labeled Faces Da abase MSE Mean Squa e E o NN Neu al Ne wo k PCD Py amid, Cascading and De o mable con olu ions Appl. Sci. 2020,10, 7213 25 o 27 PSNR Peak signal- o-noise a io PubFig Public Figu es Face Da abase PULSE Sel -Supe ised Pho o Upsampling ia La en Space Explo a ion o Gene a i e Models ReLU Rec i ied Linea Uni ResNe Residual Ne wo k ResBlock Residual Block SR Supe -Resolu ion SRCNN Supe -Resolu ion Con olu ional Neu al Ne wo k SRGAN Supe -Resolu ion Gene a i e Ad e sa ial Ne wo k SSIM S uc u al simila i y TSA Tempo al and Spa ial A en ion UR-DGN Ul a- esolu ion by disc imina i e gene a i e ne wo k UVT Unpai ed Video T ansla ion VSR Video Supe -Resolu ion YOLO You Only Look Once Re e ences 1. Hollis, M.E. Secu i y o su eillance? Examina ion o CCTV came a usage in he 21s cen u y. C iminol. Public Policy 2019,18, 131–134. [C ossRe ] 2. Molini, A.B.; Valsesia, D.; F acas o o, G.; Magli, E. DeepSUM: Deep neu al ne wo k o Supe - esolu ion o Un egis e ed Mul i empo al images. IEEE T ans. Geosci. Remo e Sens. 2019,58, 3644–3656. [C ossRe ] 3. Sal e i, F.; Mazzia, V.; Khaliq, A.; Chiabe ge, M. Mul i-Image Supe Resolu ion o Remo ely Sensed Images Using Residual A en ion Deep Neu al Ne wo ks. Remo e Sens. 2020,12, 2207. [C ossRe ] 4. Zangeneh, E.; Rahma i, M.; Mohsenzadeh, Y. Low esolu ion ace ecogni ion using a wo-b anch deep con olu ional neu al ne wo k a chi ec u e. Expe Sys . Appl. 2020,139, 112854. [C ossRe ] 5. Huang, G.B.; Ma a , M.; Be g, T.; Lea ned-Mille , E. Labeled Faces in he Wild: A Da abase o S udying Face Recogni ion in Uncons ained En i onmen s; Technical Repo 07-49; Uni e si y o Massachuse s o Amhe s : Amhe s , MA, USA, Oc obe 2007. 6. Kuma , N.; Be g, A.C.; Belhumeu , P.N.; Naya , S.K. A ibu e and simile classi ie s o ace e i ica ion. In P oceedings o he 2009 IEEE 12 h In e na ional Con e ence on Compu e Vision, Kyo o, Japan, 27 Sep embe –4 Oc obe 2009; pp. 365–372. 7. F eeman, W.T.; Pasz o , E.C.; Ca michael, O.T. Lea ning low-le el ision. In . J. Compu . Vis. 2000 ,40, 25–47. [C ossRe ] 8. Liu, Z.; Luo, P.; Wang, X.; Tang, X. La ge-scale celeb aces a ibu es (celeba) da ase . Re ie ed Augus 2018,15, 2018. 9. Wol , L.; Hassne , T.; Maoz, I. Face ecogni ion in uncons ained ideos wi h ma ched backg ound simila i y. In P oceedings o he CVPR 2011, Colo ado Sp ings, CO, USA, 20–25 June 2011; pp. 529–534. 10. O’Mahony, N.; Campbell, S.; Ca alho, A.; Ha apanahalli, S.; He nandez, G.V.; K palko a, L.; Rio dan, D.; Walsh, J. Deep lea ning s. adi ional compu e ision. In P oceedings o he Science and In o ma ion Con e ence, Leipzig, Ge many, 1–4 Sep embe 2019; pp. 128–144. 11. Dong, C.; Loy, C.C.; He, K.; Tang, X. Image supe - esolu ion using deep con olu ional ne wo ks. IEEE T ans. Pa e n Anal. Mach. In ell. 2015,38, 295–307. [C ossRe ] 12. Isola, P.; Zhu, J.Y.; Zhou, T.; E os, A.A. Image- o-image ansla ion wi h condi ional ad e sa ial ne wo ks. In P oceedings o he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, Honolulu, HI, USA, 21–26 July 2017; pp. 1125–1134. 13. Ledig, C.; Theis, L.; Huszá , F.; Caballe o, J.; Cunningham, A.; Acos a, A.; Ai ken, A.; Tejani, A.; To z, J.; Wang, Z.; e al. Pho o- ealis ic single image supe - esolu ion using a gene a i e ad e sa ial ne wo k. In P oceedings o he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, Honolulu, HI, USA, 21–26 July 2017; pp. 4681–4690. 14. Lim, B.; Son, S.; Kim, H.; Nah, S.; Mu Lee, K. Enhanced deep esidual ne wo ks o single image supe - esolu ion. In P oceedings o he IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, Honolulu, HI, USA, 21–26 July 2017; pp. 136–144.