Dynamic Facial Emotion Recognition Oriented to HCI Applications
Abstract
Producción Científica
Full text
1
Dynamic Facial Emo ion Recogni ion O ien ed o HMI Applica ions
Abs ac
As pa o a mul imodal anima ed in e ace p e iously p esen ed in [38], in his pape we desc ibe a
me hod o dynamic ecogni ion o displayed acial emo ions on low esolu ion s eaming images. Fi s , we
add ess he de ec ion o Ac ion Uni s o he Facial Ac ion Coding Sys em upon Ac i e Shape Models and
Gabo il e s. No malized ou pu s o he Ac ion Uni ecogni ion s ep a e hen used as inpu s o a neu al
ne wo k which is based on eal cogni i e sys ems a chi ec u e, and consis s on a habi ua ion ne wo k plus a
compe i i e ne wo k. Bo h he compe i i e and he habi ua ion laye use di e en ial equa ions hus aking
in o accoun he dynamic in o ma ion o acial exp essions h ough ime. Expe imen al esul s ca ied ou
on li e ideo sequences and on he Cohn-Kanade ace da abase show ha he p oposed me hod p o ides
high ecogni ion hi a es.
Keywo ds: Ac i e Shape Models,Cogni i e Sys ems, Facial Emo ion Recogni ion,Gabo il e s, Human-
Machine In e aces
2
I. In oduc ion
O e he las ew yea s, a g owing in e es has been obse ed in he de elopmen o new in e aces in
Human-Machine In e ac ion (HMI) ha allow humans o in e ac wi h machines in a na u al way. Ideally, a
man-machine in e ace should be anspa en o he use , i.e. i should allow he use o in e ac wi h he
machine wi hou equi ing any cogni i e e o . Wha is sough is ha anyone can use machines and
compu e s in hei daily li es, e en people no e y amilia wi h he echnology, making i easie bo h he
way people communica e wi h de ices and he way de ices p esen hei da a o he use .
Conside ing human o human in e ac ion as a e e ence, one o he main sou ces o in o ma ion
in e change esides in he ace ges u e capabili ies. People a e ex emely skilled in in e p e ing he
beha iou o o he pe sons acco ding o changes in hei acial cha ac e is ics. Fo ha eason, ace- o- ace
communica ion is one o he mos in ui i e, obus and e ec i e ways o communica ion be ween humans.
Using acial exp essions as a communica ion channel in HMI is one way o making use o he inhe en
human abili y o exp ess hemsel es ia changes in hei acial appea ance. Thus, he obus ness and
na u alness o human-machine in e ac ion is g ea ly imp o ed [1].
In he pas , a lo o e o was dedica ed o ecognizing acial exp essions and emo ions in s ill images.
Using s ill images has he limi a ion ha hey a e usually cap u ed in he apex o he exp ession, whe e he
pe son is pe o ming he indica o s o emo ion mo e ma kedly. In day- o-day communica ion, howe e , he
apex o acial exp essions a e a ely shown, unless o e y speci ic cases. Also, esea ch has demons a ed
ha people pe cei e o he s’ emo ions in a con inuous ashion, and ha me ely ep esen ing emo ions in a
bina y (on/o ) manne is no an accu a e e lec ion o he eal human pe cep ion o acial emo ions [37].
Mo e ecen ly, esea ches ha e changed hei e o s owa ds modelling dynamic aspec s o acial
exp ession. This is because di e ences be ween emo ions a e mo e powe ully modelled by s udying he
dynamic a ia ions o acial appea ance, a he han jus as disc e e s a es. This is especially ue o na u al
exp essions wi hou any delibe a e exagge a ed posing.
I is also necessa y o dis inguish be ween wo concep s which, al hough in e ela ed, should be
independen ly conside ed in he ield o acial ecogni ion: acial exp ession ecogni ion and acial emo ion
ecogni ion. Exp ession ecogni ion e e s o hose sys ems whose pu pose is he de ec ion and analysis o
acial mo emen s, and changes in acial ea u es om isual in o ma ion. On he o he hand, o pe o m
emo ional exp ession ecogni ion, i is also necessa y o associa e hese acial displacemen s o hei
emo ional signi icance, gi en ha he e will be ce ain mo emen s associa ed wi h mo e han one emo ion.
This pape ocuses on he la e : he au oma ic isual ecogni ion o emo ional exp essions in dynamic
sequences. Howe e , he p oposed me hod is di ided in o wo di e en s ages: he i s one ocuses on he
de ec ion and ecogni ion o small g oups o acial muscle ac ions and, based on he ou pu , he second
de e mines he associa ed emo ional exp ession.
The p oposed sys em is pa o a mul imodal anima ed head which was p e iously de eloped by ou
esea ch g oup. The de eloped anima ed head has p o en o be a pe ec complemen o a humanoid
obo ic cons uc ion o e en a subs i u e o a ha dwa e equi alen solu ion [38].
The es o he pape is o ganized as ollows. Rela ed wo k is e iewed in sec ion II. The p oposed
app oach o acial ac i i y de ec ion h ough compu e ision echniques is p esen ed in sec ion III. The
me hod o emo ion ecogni ion is desc ibed in sec ion IV. Expe imen al esul s a e p esen ed in V.
Conclusions and u u e wo k a e desc ibed in sec ion VI.
II. Rela ed wo k
The common p ocedu e o he ecogni ion o acial exp essions can be di ided in o h ee s eps: image
segmen a ion, ea u e ex ac ion and ea u e classi ica ion. Image segmen a ion is he p ocess ha allows an
3
image o be di ided in o g oups o pixels ha sha e common cha ac e is ics. The aim o segmen a ion is o
simpli y image con en in o de o analyse i easily. The e a e many di e en echniques o image
segmen a ion ha use ei he colou [18] o g ey scale [19] in o ma ion. As a as acial image segmen a ion
is conce ned, one o he mos widesp ead echniques is he Viola and Jones algo i hm [19], due o i s
obus ness agains ace a ia ions and i s compu a ional speed. Image segmen a ion o acial ecogni ion
can also include a no maliza ion s ep, which elimina es o a ions, ansla ions and scaling o cap u ed aces.
The goal o he ea u e ex ac ion s ep is o ob ain a se o acial ea u es ha a e able o desc ibe he ace
in an e ec i e way. An op imum se o acial ea u es should be able o minimize a ia ions be ween
cha ac e is ics ha belong o he same class and maximize he di e ences be ween classes. The e exis
di e en app oaches o ea u e ex ac ion. Fo example, in [20] op ical low is used o modelling
muscula ac i i y h ough he es ima ion o he displacemen s o a g oup o cha ac e is ic poin s. The main
disad an age o op ical low is ha i is highly sensi i e o ligh ing a ia ions and no e y accu a e i
applied o low esolu ion images. Fo ha eason, i is usually used along wi h o he me hods, such as
olume local bina y pa e ns VLBPN [21]. Analysis o acial geome y is ano he me hod used o ea u e
ex ac ion, whe e he geome ic posi ions o ce ain acial ea u es a e used o ep esen he ace. In [22] he
posi ion o manually ma ked acial ea u es a e used o ep esen he ace, and hen pa icle il e s a e used
o s udy a ia ions in hose cha ac e is ics.
Changes in acial appea ance a e also used o acial ea u e ex ac ion. Mo e common echniques a e
P incipal Componen Analysis (PCA) [23], Linea Disc iminan Analysis (LDA) [24] and many o he s,
such as Gabo il e s [25]. Gabo il e s a e widely used o ea u e ex ac ion as hey allow acial
cha ac e is ics o be ex ac ed p ecisely and in a ian ly agains changes in luminosi y o pose [26].
Howe e , Gabo il e s demand ela i ely high compu a ional cos s, which ge highe as he image
esolu ion inc eases. Fo ha eason, hey a e usually used in small egions o he ace a e a p e ious
segmen a ion s ep. Recen ly, Local Bina y Pa e ns ha e begun o be used o image segmen a ion as hey
equi e lowe compu a ional cos s han Gabo il e s [27], al hough hey a e less p ecise han he la e
when he numbe o ex ac ed ea u es inc eases [28].
Bo h geome y based and appea ance based me hods ha e hei ad an ages and disad an ages. Fo ha
eason, ecen p oposals y o combine bo h echniques o ea u e ex ac ion in wha a e called de o mable
models. De o mable models a e buil om a aining da abase and a e able o ex ac bo h geome ic and
appea ance da a om a acial image. The e a e many di e en app oaches o building de o mable models,
bu hose ha a e mo e commonly used a e Ac i e Shape Models (ASMs) [31] and Ac i e Appea ance
Models (AAMs) [29]. Ac i e shape and appea ance models we e de eloped by Coo es and Taylo [30].
They a e based on he s a is ical in o ma ion o a poin dis ibu ion model (shape) which is i ed o ack
aces in images using he s a is ical in o ma ion o he image ex u e. AAMs use all ex u e in o ma ion
p esen in he image, whe eas ASMs only use ex u e in o ma ion a ound model poin s. As o ea u e
acking, ASMs ha e shown o be mo e obus han AAMs i used in eal ime ea u e ex ac ion [6]. This is
because ASM acking is done in p ecise image egions which ha e p e iously been modelled du ing a
aining s age (a ound model poin s), whe eas AAM acking is pe o med using he in o ma ion p esen in
he whole image.
Once acial ea u es had been ex ac ed, he e a e many di e en app oaches o hei classi ica ion in
o de o ca ego ize and ecognize acial exp essions, such as Suppo Vec o Machines [32] o ule based
classi ie s [33]. Fo eal ime applica ions, he analysis should be done o e a se o acial ea u es ob ained
o e ime om ideo sequences, using neu al ne wo ks [34], Hidden Ma ko Models [35] o Bayesian
ne wo ks [36]. Among classi ica ion echniques, wo big g oups can also be dis inguished: hose ha y o
classi y emo ional exp essions by analysing he ace as a whole [13] [15], and hose which y o de ec
small acial exp essions sepa a ely, and hen combine hem o complex exp ession classi ica ion [32] [34].
4
III. Ac ion Uni de ec ion
The e a e se e al me hods o measu ing and desc ibing acial muscula ac i i y, and he changes which
ha ac i i y p oduces on acial appea ance. F om hese, he Facial Ac ion Coding Sys em (FACS) [9] is he
mos widely used in psychological esea ch, and has been used in many app oaches o au oma ic
exp ession ecogni ion [29] [34]. As a inal goal, FACS seeks o ecognize and desc ibe he so-called
Ac ion Uni s (AUs), which ep esen he smalles g oup o muscles ha can be con ac ed independen ly
om he es and p oduce momen a y changes in acial appea ance. These Ac ion Uni s can, in
combina ion, desc ibe any global beha io o he ace.
The de eloped acial exp ession ecogni ion sys em i s looks o he au oma ic ecogni ion o FACS
Ac ion Uni s by using a combina ion o Ac i e shape models (ASMs) [4] and Gabo il e s [5]. ASMs a e
s a is ical models o he shape o objec s, which a e cons ained o a y only in ways lea ned du ing a
aining s age o labelled samples. Once ained, he ASM algo i hm aims o ma ch he shape model o a
new objec in a new image. ASMs allow accu a e ace segmen a ion in images, in o de o pe o m a mo e
de ailed analysis o he di e en acial egions. Based on he wo k de eloped by [6], a shape model o 75
poin s has been used o ack and segmen human aces om images cap u ed wi h a webcam as seen in
Fig.1.
The human ace is ex emely dynamic, and p esen s a ia ions in i s shape due o he con ac ion o acial
muscles when acial exp essions a e pe o med. ASMs a e able o accu a ely ack aces, e en in he
p esence o acial shape a ia ions (mou h opening, eyeb ows ising, e c.). Howe e , as ASM aining and
acking is based on a local p o ile sea ch, i is di icul o ain he model o ack ansien acial ea u es
such as w inkles, o o ake in o accoun he p esence o elemen s such as glasses, bea d o sca es. The
amoun o a ia ions allowed o a shape du ing he aining s age d ama ically a ec s acking speed and
accu acy [4]. Fo ha eason, acial a ia ions pe mi ed in he images used du ing he aining s age ha e
been cons ained o pe manen ea u es (wid h, heigh , eyeb ows o nose and mou h posi ion). Howe e ,
many acial exp essions in ol e he appa i ion o ansien acial ea u es ha may only be isible while he
exp ession is being pe o med. (E.g. he e ical w inkle ha appea s be ween he eyeb ows when
owning). Since he ASMs a e no able o de ec ansien cha ac e is ics o acial exp essions, ASM
acking has been used o ack he ace and de ec changes in pe manen acial ea u es. I has also been
used o ace segmen a ion in smalle egions, which a e p ocessed using a bank o Gabo il e s ha le us
look o ansien acial ea u es.
III.A. AU – ela ed ea u e ex ac ion
The ollowing sec ions show he p oposed me hod used o he de ec ion o acial ea u es when
pe o ming acial exp essions. Fou di e en egions o he ace ha e been conside ed o ea u e de ec ion:
o ehead and eyeb ows, eyes, nose, mou h and lowe and la e al sides o he mou h. Fo each egion, he
mos dis inc i e AUs associa ed o each o he six uni e sal emo ional exp essions, as desc ibed in FACs
(joy, ange , ea , disgus , su p ise and sadness) ha e been conside ed.
Fo ehead and eyeb ow egion
The ac ion uni s conside ed o his egion a e AU1, AU2 and AU4 (see Table I). The con ac ion o he
on alis pa s medialis makes he eyeb ows ise and be sligh ly pulled oge he , wi h he o ma ion o
ho izon al w inkles in he cen al egion o he o ehead (coded in FACS as AU1), while he con ac ion o
he on alis pa s la e alis muscle also o igina es ho izon al w inkles and eyeb ow ising bu in he ou e
egions o he o ehead (AU2). The con ac ion o he co uga o supe cilii muscle pulls he inne egion o
5
he eyeb ows oge he and downwa ds. Also, he eyelids a e sligh ly pushed oge he . As a esul , e ical
w inkles appea in he space be ween he eyeb ows (AU4).
Va ia ions in eyeb ow posi ion ha e been de e mined by compu ing he dis ance be ween he eyeb ows
and he dis ances be ween he inne and ou e co ne s o he eyeb ows and he line connec ing he co ne s
o he eyes. These dis ances a e compu ed using shape model poin s’ posi ion as shown in
Fig.2. Thus, i AU1 o AU2 a e ac i e, he e will be an inc ease in he h1 and h2 dis ances espec i ely,
while, i AU4 is ac i e, he dis ance d1 will dec ease.
Howe e , as compu ed ASM poin s’ posi ion may su e om noise, he p oposed me hod complemen s
hose measu es wi h he de ec ion o o he cha ac e is ics, such as he appea ance o w inkles in he uppe
egion o he ace. The de ec ion o ansien ea u es, such as w inkles, has a numbe o p oblems esul ing
om e lec i e and e lexi e p ope ies o he skin, he p esence o acial hai , skin one, e c. These
p oblems can cause he condi ions in which ce ain acial ea u es appea o a y e en o he same
indi idual. To de ec hese ea u es in as in a ian and gene ic a way as possible, Gabo il e s ha e been
used. Gabo il e s can exploi salien isual p ope ies, such as spa ial localiza ion, o ien a ion selec i i y,
and spa ial equency cha ac e is ics, and a e qui e obus agains illumina ion changes [7]. A 2D Gabo
il e is a Gaussian ke nel unc ion modula ed by a sinusoidal plane wa e. The conside ed il e equa ions
in spa ial (𝑔𝑔𝜎𝜎,𝐹𝐹,𝜃𝜃) and equency (𝐺𝐺𝜎𝜎,𝐹𝐹,𝜃𝜃) domains a e shown in equa ions (1) and (4):
𝑔𝑔𝜎𝜎,𝐹𝐹,𝜃𝜃(𝑥𝑥,𝑦𝑦)=𝑔𝑔′𝜎𝜎(𝑥𝑥,𝑦𝑦)𝑒𝑒𝑗𝑗2𝜋𝜋𝐹𝐹𝑥𝑥′
(1)
whe e:
𝑔𝑔′𝜎𝜎(𝑥𝑥,𝑦𝑦)=1
2𝜋𝜋𝜎𝜎𝑥𝑥𝜎𝜎𝑦𝑦𝑒𝑒𝑥𝑥𝑒𝑒�−1
2��𝑥𝑥′
𝜎𝜎𝑥𝑥�2+�𝑦𝑦′
𝜎𝜎𝑦𝑦�2��
(2)
𝑥𝑥′=𝑥𝑥𝑐𝑐𝑐𝑐𝑐𝑐𝜃𝜃+𝑦𝑦𝑐𝑐𝑠𝑠𝑠𝑠𝜃𝜃𝑦𝑦′=𝑥𝑥𝑐𝑐𝑠𝑠𝑠𝑠𝜃𝜃+𝑦𝑦𝑐𝑐𝑐𝑐𝑐𝑐𝜃𝜃
(3)
and
𝐺𝐺𝜎𝜎,𝐹𝐹,𝜃𝜃(𝑥𝑥,𝑦𝑦)=𝑒𝑒𝑥𝑥𝑒𝑒�−1
2�(𝑢𝑢′−𝐹𝐹)2
𝜎𝜎𝑢𝑢2+𝑣𝑣′2
𝜎𝜎𝑣𝑣2��
(4)
whe e:
𝜎𝜎𝑢𝑢=1
2𝜋𝜋𝜎𝜎𝑥𝑥 , 𝜎𝜎𝑣𝑣=1
2𝜋𝜋𝜎𝜎𝑦𝑦
(5)
𝑢𝑢′=𝑢𝑢𝑐𝑐𝑐𝑐𝑐𝑐𝜃𝜃+𝑣𝑣𝑐𝑐𝑠𝑠𝑠𝑠𝜃𝜃 , 𝑣𝑣′=−𝑢𝑢𝑐𝑐𝑠𝑠𝑠𝑠𝜃𝜃+𝑣𝑣𝑐𝑐𝑐𝑐𝑐𝑐𝜃𝜃
(6)
F is he modula ed equency, which de ines he equency o he complex sinusoid, θ is he o ien a ion
o he Gaussian unc ion and σx and σy de ine he Gaussian unc ion scale on bo h axes.
Gabo il e s a e loca ed bo h in he spa ial and equency domains, which makes hem sui able o
ela ing equency cha ac e is ics o hei posi ion in he image space. One ad an age o using Gabo il e s
esides in he possibili y o o ien ing hem so ha acial cha ac e is ics wi h a known o ien a ion can be
highligh ed. Fo he de ec ion o w inkles in he o ehead egion, shape poin s loca ion is used o ex ac he
ace egion be ween he eyeb ows and he base o he hai on he o ehead. The ex ac ed egion is esized
6
o 32x32 pixels size and g ay-scaled. Tha egion size coincides wi h Gabo il e sizes, which a e
compu ed o -line in o de o minimize il e ing p ocessing ime. The ex ac ed egion is hen con ol ed
wi h he co esponding il e s in equency domain. To achie e mo e obus ness agains ligh ing condi ions,
all il e s a e u ned o ze o DC:
𝑓𝑓(𝑥𝑥,𝑦𝑦)⨂𝑔𝑔(𝑥𝑥,𝑦𝑦)=𝑇𝑇−1�𝐹𝐹(𝑢𝑢,𝑣𝑣)𝐺𝐺(𝑢𝑢,𝑣𝑣)�
𝑤𝑤𝑠𝑠𝑤𝑤ℎ 𝐺𝐺(0,0)= 0
(7)
A e applying he il e , he esul ing phase componen o he con olu ion is h esholded. Whene e
possible, we ha e wo ked wi h he phase componen , since i is he one ha su e s om less a ia ion due
o di e en ligh ing condi ions o he p esence o acial hai o shiny skin. Since Gabo il e ing causes he
appea ance o a e ac s a he co ne s o he images, a egion la ge han he in e es a ea has been
conside ed. The co ne pixels o he il e ed image pixels a e hen masked ou , gi en ha hey do no
p o ide in o ma ion. Masking also allows image egions o be easily segmen ed wi h i egula shapes.
In he case o he o ehead, h ee egions a e conside ed in he esul ing il e ed and masked image: one
cen al and wo la e al. I he a e age image in ensi y in he cen al egion exceeds a h eshold, i is
encoded as he p esence o w inkles, con ibu ing o he ac i a ion o AU1. Simila ly, he mean in ensi y in
he la e al egions allows w inkles ha a e conside ed o con ibu e o he ac i a ion o AU2, o igh and
le eyeb ows espec i ely, o be de ec ed. Fig.3 shows he de ec ion o w inkles o AU1, AU2 and he
combina ion AU1+AU2. I has o be no ed ha il e ing and h esholding calcula ions a e done wi h
loa ing poin p ecision, al hough he esul ing image is shown as an 8 bi g ey image o isualiza ion
pu poses.
The same p ocedu e has been applied o he AU4, whe e he egion o in e es is loca ed be ween he
wo eyeb ows and he op o he nose (Fig.3). In his case a Gabo il e wi h ho izon al o ien a ion has been
used, as he w inkles gene a ed by AU4 ac i a ion ha e a e ical o ien a ion ( il e pa ame e s: F = 0.07
pixels-1, θ = 0 ad, σx= 7.9, σy = 4 and size 32x32).
Eyes egion
Ac ion Uni s conside ed in his egion a e AU5, AU6, AU7, AU43, AU45 and AU46 (Table II).
In he p oposed app oach, AU de ec ion in he egion o he eyes is done acco ding o h ee
cha ac e is ics: wo ansi ional, he posi ion o he eyelids and he p esence o absence o w inkles on he
ou e edges o he eyes, and one pe manen , he amoun o isible scle a.
The opening and closing o he eyelids is de e mined h ough he dis ance be ween eyelids compu ed
om he shape model poin s. The amoun o isible scle a is also used o eye ape u e de ec ion. To
compu e his, a il e o pa ame e s F = 0.12 pixels-1, θ = 0 ad, σx = 3.2 pixels and σy = 2 pixels and size
32x32 pixels has been applied o each ex ac ed eye egion. The esul an phase image is masked using an
ellipsoidal mask, and h esholded o ex ac he wo bigges con ou s o isible scle a. The size o he
bigges con ou con ibu es o de e mining whe he he eyes a e open, pa ially open o comple ely closed,
as shown in Fig.4.
A pa ially isible scle a (Fig.4, middle column) could indica e ha ei he AU6 o AU7 a e ac i e. To
dis inguish hose AUs, an a ea nea he ou e edges o he eyes is ex ac ed and il e ed wi h a Gabo il e
o pa ame e s F = 0.12pixels-1, θ = 1.5 ad, σx = 4.2 pixels and σy= 1pixels. I he esul ing image in ensi y
is high enough o conside he p esence o w inkles, AU6 is conside ed ac i e, o he wise AU7 is ac i e.
Nose egion
7
The egion o he nose has dis inc i e ansien ea u es ela ed o he ac i a ion o AU9 and AU10 (Table
III).
W inkle de ec ion in he nose egion ollows a simila schema o he ones desc ibed in p e ious sec ions.
In his case, h ee egions a e conside ed: he egion o he nose be ween he eyes (Fig.5 egion A), and on
bo h sides o he nose base (Fig.5 egions B and C). W inkles in bo h egions indica e he ac i a ion o
AU9, while, i hey only appea nea he base o he nose, hey indica e he ac i a ion o AU10.
Mou h egion
The main ac ion uni s whose ac i a ion p oduces isual changes in he egion o he mou h a e AU10,
AU12, AU15, AU16, AU17, AU25 and AU26 (Table VI). The main acial cha ac e is ics in he mou h
egion a e he posi ion o he lips and he ee h.
Mou h opening and closing is de e mined based on he a ea o he ellipse de ined by he dis ances
be ween shape poin s in he uppe and lowe lips and on he mou h co ne s. These poin s a e also used o
de ec he mou h co ne s being pulled down. The ising and alling o he lip co ne s is also de e mined by
shape model poin s loca ed in he mou h co ne s. Along wi h ASM poin s’ posi ion, wo Gabo il e s a e
used in he mou h egion o complemen AU de ec ion wi h pa ame e s F = 0.9 pixels-1, θ = 1.57 ad, σx = 4
pixels,σy = 1.3 pixels and F = 0.21 pixels-1, θ = 2.74, σx = 1.9 pixels, σy = 0.7 pixels, bo h 32x32 pixels size.
Nine di e en egions a e conside ed in he esul ing image, as shown in Fig.6. The combina ion o he
h esholded in ensi ies on hose sub- egions con ibu es o he ac i a ion o he di e en AUs. Fo example,
i he in ensi y in sub- egion 1 in Fig.6 exceeds a h eshold, hen AU12 (le ) is conside ed o be ac i e,
whe eas, i he in ensi y in sub- egions 7, 8 and 9 exceed a h eshold, hen AU17 is conside ed ac i e. The
esul o il e ing in he mou h egion when pe o ming o he Ac ion Uni s a e also shown in Fig.6.
AU16 and AU10 make he uppe and lowe ee h, espec i ely, become isible when ac i a ed in
conjunc ion wi h AU25; so only he combina ion o AU25 wi h hose ac ion uni s has been conside ed. Fo
he de ec ion o he ee h, a il e simila o ha used o he scle a has been used, wi h pa ame e s F = 0.07
pixels-1, θ = 0 ad, σx = 7.9 pixels, σy = 4 and 32x32 pixels size.
Lowe and la e al sides o he mou h egion
In his egion, dis inc i e w inkles appea as he esul o he ac i a ion o AU12, AU15 and AU17 (Table
V). Ac i a ion o AU12 causes a w inkle o appea om he nos ils o he co ne s o he mou h. Bo h i s
angle wi h espec o he sagi al plane o he ace and i s dep h de e mine he in ensi y o AU12. To de ec
his ea u e, a ace egion bounded by he cheeks, nos ils and he angle o he mou h is ex ac ed and
il e ed wi h he pa ame e s F = 0.21 pixel-1, θ = 2.3 ad, σx = 1.9 and σy = 0.7 and size 32x32 pixels o he
igh w inkle, as shown in Fig.7, and θ = 0.84 ad o he le one.
As desc ibed be o e, along wi h making he lips adop a con ex a c o m, bo h AU 15 and 17 p oduce
w inkles in he chin egion when ac i e. To complemen he de ec ion o bo h AUs, he chin egion has
been il e ed wi h pa ame e s F = 0.52 pixels-1, θ = 0 ad, σx= 0.9 and σy = 1.5 and size 32x32, ollowing a
simila p ocess o he one desc ibed o he es o he w inkle de ec ion.
III.B. Su ey o conside ed pa ame e s o AU de ec ion
Table I o Table V show he conside ed pa ame e s o he de ec ion o each o he Ac ion Uni s in he
i e acial egions.
8
III.C. Compu a ion o AU ac i a ion
The ac i a ion alue o each Ac ion Uni is hen compu ed as he weigh ed sum o all no malized
ea u es i depends on (see Table I o Table V).
Equa ion (8) shows he calcula ion o AU ac i a ion whe e cni a e he no malized ea u es alue and Pni
a e he weigh s o each ea u e. Weigh s ha e been manually adjus ed so hey gi e mo e impo ance o
hose ea u es ob ained om he il e ing p ocess, a he han om he ASM poin s’ posi ion, as he la e
a e mo e noise sensi i e.
𝐴𝐴𝐴𝐴𝑖𝑖= 𝑃𝑃1𝑖𝑖𝑐𝑐1𝑖𝑖+𝑃𝑃2𝑖𝑖𝑐𝑐2𝑖𝑖+. . +𝑃𝑃𝑛𝑛𝑖𝑖𝑐𝑐𝑛𝑛𝑖𝑖
(8)
whe e
�𝑃𝑃𝑗𝑗𝑖𝑖 = 1
𝑛𝑛
𝑗𝑗=1
(9)
Las ly, ou pu s ob ained o each Ac ion Uni a e low-pass il e ed in o de o a oid luc ua ions in hei
alues due o noise, as shown in
𝑐𝑐𝑗𝑗𝑖𝑖[𝑠𝑠]=𝑐𝑐𝑗𝑗𝑖𝑖[𝑠𝑠−2]+𝐴𝐴(𝑐𝑐𝑗𝑗𝑖𝑖[𝑠𝑠−1]−𝑐𝑐𝑗𝑗𝑖𝑖[𝑠𝑠−2])
(10)
IV. Emo ional exp ession ecogni ion
The p e iously desc ibed Ac ion Uni ecogni ion me hod allows di e en acial mo emen s ha happen
when pe o ming acial exp essions o be de e mined and ca ego ized. In human-compu e and human-
obo in e ac ion, ges u e ecogni ion could be use ul pe se, as i p o ides humans na u al ways o
communica e wi h de ices in e ms o simple non- e bal communica ion. Howe e , when de eloping mo e
complex cogni i e sys ems in ended o na u al in e ac ion wi h humans, acial exp essions should be
endowed wi h a meaning.
Conside ing Ekman’s wo k on acial emo ions [9], emo ions a e displayed in he ace as mo emen s
which could be isually ca ego ized in e ms o Ac ion Uni s. Fo each uni e sal emo ional exp ession, he
Facial Ac ion Coding Sys em is able o pa ame e ize i wi h he co esponding se o AUs ha a e ac i e.
Based on Ekman’s wo k, no malized ou pu s o he Ac ion Uni ecogni ion s age a e used as inpu s o a
neu al ne wo k which pe o ms emo ion ecogni ion. The a chi ec u e is based on eal cogni i e sys ems
and consis s o a habi ua ion based ne wo k plus a compe i i e based ne wo k. Fig.8 shows an o e iew o
he ne wo k a chi ec u e, while he ollowing subsec ions desc ibe he unc ionali y and pu pose o each
laye in de ail.
IV.A. Inpu il e
Ac i a ion le els o he de ec ed Ac ion Uni s se e as s imuli o he inpu bu e in he F0 le el (Fig.8).
Each s imulus is no malized in o he [0,1] ange and i s salience is modula ed wi h a sensi i i y gain Ki.
Those gains allow ce ain inpu s o be inhibi ed in case he e is a acial ea u e ha causes e o s du ing
ecogni ion (e.g. sunglasses, sca s). In hose cases, inpu s ha a e known o be e oneous can be
deac i a ed, and he ne wo k could s ill wo k using he es o he de ec ed acial ea u es. Timing is one o
he main aspec s o be aken in o accoun when pe o ming acial exp essions, gi en ha no all exp essions
9
a e pe o med a he same speed. Some acial exp essions could be pe o med quicke han he ne wo k is
able o p ocess hem, hus a il e ing s age is also applied o he inpu s. Thus, he inpu ’s e ec i e salience
will decay mo e slowly o e ime.
Equa ion (11) shows how sensi i i y gain and il e ing a e applied o he inpu s, whe e Ii a e he inpu
s imuli, Si is he il e ed inpu s imuli, Ki he sensi i i y gain and a he a enua ion a e. Thus, he nex
le el’s neu on ac i i y decays wi h a a e a when he e is no inpu .
𝑑𝑑𝑆𝑆𝑖𝑖
𝑑𝑑𝑤𝑤 =−𝑎𝑎𝑆𝑆𝑖𝑖+𝑎𝑎𝐾𝐾𝑖𝑖𝐼𝐼𝑖𝑖
(11)
IV.B. Habi ua ion laye
The p oposed ne wo k has habi ua ion capabili ies, ha is, i loses in e es in pe manen s imuli o e
ime. S imuli decay due o habi ua ion allows he ne wo k o dynamically adap i sel agains pe manen
inpu s no caused by acial exp essions, bu by acial ea u es p esen in some use s e en in es posi ion. I
could be possible, o example, ha a use ’s nasolabial w inkle is p onounced e en in es posi ion because
o his/he physiognomy, o ha pe manen aging w inkles a e p esen in he use ’s ace. By using a
habi ua ion laye , hose con inuous s imuli will lose p eponde ance o e ime, allowing he es o he acial
ea u es o acqui e impo ance. Habi ua ion is pe o med by mul iplying he inpu s imuli wi h a gain ha is
ac ualized o e ime. The gain compu a ion is based on G ossbe g’s Slow T ansmi e Habi ua ion and
Reco e y Model [10]:
𝑑𝑑𝑔𝑔𝑖𝑖
𝑑𝑑𝑤𝑤 =𝐸𝐸(1−𝑔𝑔𝑖𝑖)−𝐹𝐹𝑆𝑆𝑖𝑖𝑔𝑔𝑖𝑖
(12)
whe e Si is he il e ed s imulus and gi is he habi ua ion gain o ha s imulus. When a s imulus is ac i e,
habi ua ion gain dec eases om he maximum alue o 1 o a minimum alue gi en by 𝐸𝐸/(𝐸𝐸+𝐹𝐹𝑆𝑆𝑖𝑖),
which is p opo ional o he s imulus alue Si. This gain is echa ged o i s ini ial uni y alue when he
s imulus ends. Cha ge and discha ge a es a e de e mined by he pa ame e s E and F.
IV.C. AUs compe i i e laye
Habi ua ed s imuli a e he inpu s o he F1 laye (Fig.8), ha is, a compe i i e neu on laye . The
compe i i e model used is on-cen e o -su ound [11], so each neu on is ein o ced wi h i s own ac i i y,
bu is a enua ed by he ac i i y o he neu ons i is connec ed o, which is known as la e al inhibi ion.
La e al inhibi ion is used o hose AU which a e complemen a y, such as AU5 (eyes wide open) and AU43
(eyes closed).
Neu on ac i i y is compu ed using equa ions (13) and (14), whe e xi is he ac i i y o neu on i and Ai is
he decay a e. The second e m is he au o- ein o cemen (on-cen e), which makes neu on ac i i y end o
i s sa u a ion alue B. The las e m in he equa ion ep esen s la e al inhibi ion (o -su ound). This
equa ion is based on he Hodking-Huxley model [12]:
𝑑𝑑𝑥𝑥𝑖𝑖
𝑑𝑑𝑤𝑤 =−𝐴𝐴𝑥𝑥𝑖𝑖+(𝐵𝐵−𝑥𝑥𝑖𝑖)[𝑆𝑆𝑖𝑖𝑔𝑔𝑖𝑖+𝑓𝑓(𝑥𝑥𝑖𝑖)]−𝑥𝑥𝑖𝑖�𝑓𝑓(𝑥𝑥𝑗𝑗)
𝑖𝑖≠𝑗𝑗
(13)
whe e
16
Re e ences
[1] Edlund, J. &Beskow, J. 2007, ‘Pushy e sus meek - using a a a s o in luence u n- aking beha iou ’,
INTERSPEECH-2007, pp. 682-685.
[2] Ma cos-Pablos S., Gómez-Ga cía-Be mejo, J., Zalama, E. 2008. ‘A ealis ic acial anima ion sui able
o human- obo in e acing’.P oceedings o he 2008 IEEE/RSJ In e na ional Con e ence on
In elligen Robo s and Sys ems, IROS2008, pp. 3810-3815, ISBN 978-1-4244-2058-2.
[3] Ma cos-Pablos S., Gómez-Ga cía-Be mejo, J., Zalama, E. 2010. ‘A ealis ic, i ual head o human-
compu e in e ac ion’.In e ac ing wi h Compu e s, Vol.22 (3), pp. 176-192, May 2010, ISSN 0953-
5438.
[4] Coo es,T.F., Taylo ,C.J. and Coope D.H. , G aham, J., 1995. Ac i e shape models - hei aining and
applica ion. Compu e Vision and Image Unde s anding (61): pp. 38-59.
[5] Daugman, J.G., 1985. Unce ain y ela ions o esolu ion in space, spa ial equency, and o ien a ion
op imized by wo-dimensional isual co ical il e s. Jou nal o he Op ical Socie y o Ame ica A, ol.
2, pp. 1160-1169.
[6] Wei, Y., 2009. Resea ch on Facial Exp ession Recogni ion and Syn hesis.Mas e Thesis, Depa men o
Compu e Science and Technology, Nanjing Uni e si y.
[7] Daugman, J.G., 1985. Unce ain y ela ions o esolu ion in space, spa ial equency, and o ien a ion
op imized by wo-dimensional isual co ical il e s. Jou nal o he Op ical Socie y o Ame ica A, ol.
2, pp. 1160-1169.
[8] Kanade, T., Cohn, J.F., Tian. Y., 2000.Comp ehensi e da abase o acial exp ession analysis.
P oceedings o he Fou h IEEE In e na ional Con e ence on Au oma ic Face and Ges u e Recogni ion
(FG’00). G enoble, F ance, pp. 46-53.
[9] Ekman, P., F iesen, W.V., Hage , J.C., 2002. The Facial Ac ion Coding Sys em - Second
edi ion.Weiden eld& Nicolson, London, UK.
[10] G ossbe g, S, 1968. Some Nonlinea Ne wo ks capable o Lea ning a Spa ial Pa e n o A bi a y
Complexi y. P oceedings o he Na ional Academy o Sciences, USA, ol. 59, pp. 368 – 372.
[11] G ossbe g, S., 1973. Con ou enhancemen , sho - e m memo y and cons ancies in e e be a ing
neu al ne wo ks. S udies in Applied Ma hema ics, ol. 52, pp. 217-257.
[12] Hodgkin, A. L., Huxley, A. F., 1952. A quan i a i e desc ip ion o ion cu en s and i s applica ions o
conduc ion and exci a ion in ne e memb anes.J. Physiol. (Lond.), ol. 117, pp. 500-544.
[13] Buenaposada, J. M., Muñoz, E., Baumela, L., 2008. Recognising acial exp essions in ideo
sequences. Pa e n Analysis and Applica ions, ol. 11, pp. 101 – 116.
[14] McAnd ew, F. T., 1986. A C oss-Cul u al S udy o Recogni ion Th esholds o Facial Exp essions o
Emo ion. Jou nal o C ossCul u al Psychology, ol. 17(2), pp. 211-224.
[15] Xiao, R., Zhao, Q., Zhang, D., Shi, P., 2011. Facial exp ession ecogni ion on mul iple
mani olds.Pa e n Recogni ion, ol.11, pp. 107-116.
[16] V e os, N., Nikolaidis, N., Pi as, I., 2009. A model – based acial exp ession ecogni ion algo i hm
using P incipal Componen s Analysis. IEEE In e na ional Con e ence on Image P ocessing, pp. 3301 –
3304.
[17] Wimme , M., MacDonald, B. A., Jayamuni, D., Yaday, A, 2008. Facial Exp ession Recogni ion o
Human – Robo In e ac ion – A p o o ype. Lec u e No es in Compu e Science, ol. 4931, pp. 139-152.
[18] Kakumanu, P., Mak ogiannis , S., Bou bakis, N., 2007. A su ey o skin-colo modeling and de ec ion
me hods. Pa e n Recogni ion, ol. 40, pp. 1106-1122.
[19] Viola, P., Jones, M. J., 2004. Robus Real-Time Face De ec ion. In e na ional Jou nal on Compu e
Vision, ol. 57, pp. 137-154.
[20] Yeasin, M., Bullo , B., Sha ma, R., 2004. F om acial exp ession o le el o in e es s: a spa io-
empo al app oach. En IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR).
[21] Kong, J., Zhan, Y., Chen, Y., 2009. Exp ession Recogni ion Based on VLBP and Op ical Flow Mixed
Fea u es. Fi h In e na ional Con e ence on Image and G aphics, pp.933-937.
17
[22] T ipa hi, R., A a ind, R., 2007. Recognizing acial exp ession using pa icle il e based ea u e poin s
acke . P oceedings o he 2nd in e na ional con e ence on Pa e n ecogni ion and machine
in elligence, pp. 584-591.
[23] Wang, F., Huang, C., Liu, X., 2009. A Fusion o Face Symme y o Two-Dimensional P incipal
Componen Analysis and Face Recogni ion. In e na ional Con e ence on Compu a ional In elligence
and Secu i y, ol. 1, pp.368-371.
[24] Zhao, T., Liang, Z., Zhang, D., Zou, Q., 2008. In e es il e s. in e es ope a o : Face ecogni ion
using Fishe linea disc iminan based on in e es il e ep esen a ion. Pa e n Recogni ion Le e s, ol.
29 (1), pp. 1849-1857.
[25] Li, J.B., Pan, J. S., Lu, Z. M., 2009. Face ecogni ion using Gabo -based comple e Ke nel Fishe
Disc iminan analysis wi h ac ional powe polynomial models. Neu al Compu ing & Applica ions,
Vol. 8 (6), pp. 613-621.
[26] Ou, J., Bai, X. B., Pei, Y., Ma, L., Liu, W., 2010. Au oma ic Facial Exp ession Recogni ion Using
Gabo Fil e and Exp ession Analysis. Second In e na ional Con e ence on Compu e Modeling and
Simula ion, ol. 2, pp.215-218.
[27] Laje a di, S. M., Hussain, Z. M., 2009. Facial exp ession ecogni ion using log-Gabo il e s and local
bina y pa e n ope a o s.In e na ional Con e ence on Communica ion, Compu e and Powe
(ICCCP’08), pp. 349-353, Oman, 2009.
[28] Shan, C., Gong, S., McOwan., P. W., 2009. Facial exp ession ecogni ion based on Local Bina y
Pa e ns: A comp ehensi e s udy. Image and Vision Compu ing, ol. 27 (6), pp. 803-816.
[29] Lucey, P., Cohn, J., Lucey, S., S idha an, S., P kachin, K.M., 2009. Au oma ically de ec ing ac ion
uni s om aces o pain: Compa ing shape and appea ance ea u es, 2009 IEEE Compu e Socie y
Con e ence on Compu e Vision and Pa e n Recogni ion Wo kshops, pp.12-18.
[30] Coo es, T. F., Coope , D. H., Taylo , C. J., G aham, J., 1991. A T ainable Me hod o Pa ame ic Shape
Desc ip ion. 2nd B i ish Machine Vision Con e ence pp. 54-61.
[31] Milbo ow, S., Nicolls, F., 2008. Loca ing Facial Fea u es wi h an Ex ended Ac i e Shape Model.
P oceedings o he 10 h Eu opean Con e ence on Compu e Vision: Pa IV, pp. 504 – 513.
[32] Su end an, N., Xie, S., 2009. Au oma ed acial exp ession ecogni ion – an in eg a ed app oach wi h
op ical low analysis and Suppo Vec o Machines.In e na ional Jou nal o In elligen Sys ems
Technologies and Applica ions, ol 7 (3), pp. 316 – 346.
[33] Pan ic, M., Pa as, I., 2006. Dynamics o acial exp ession: ecogni ion o acial ac ions and hei
empo al segmen s om ace p o ile image sequences. IEEE T ansac ions on Sys ems, Man, and
Cybe ne ics ol. 36 (2) pp. 433–449.
[34] Tian, Y., Kanade, T., Cohn, J. Recognizing ac ion uni s o acial exp ession analysis. IEEE
T ansac ions on Pa e n Analysis and Machine In elligence ol. 23 (2) , pp. 97–115.
[35] Schmid , M., Schels, M., Schwenke , F., 2010. A Hidden Ma ko Model Based App oach o Facial
Exp ession Recogni ion in Image Sequences. A i icial Neu al Ne wo ks in Pa e n Recogni ion pp.
149-160.
[36] Valen i, R., Sebe, N., Ge e s, T., 2007. Facial Exp ession Recogni ion: A Fully In eg a ed App oach.
14 h In e na ional Con e ence o Image Analysis and P ocessing - Wo kshops (ICIAPW 2007), pp.125-
130.
[37] Schiano, D.J., Eh lich, S.M., She idan, K., 2004. Ca ego ical Impe a i e NOT: Facial A ec is
Pe cei ed Con inuously. CHI '04 P oceedings o he SIGCHI con e ence on Human ac o s in
compu ing sys ems, pp. 49–56.
[38] Samuel, M., Jaime, G., Edua do, Z. “A ealis ic, i ual head o human-compu e in e ac ion”.
In e ac ing wi h Compu e s, Vol.22 (3),pp. 176-192, May 2010, ISSN 0953-5438.
18
Figu es and Tables
Fig.1.T iangula ed 75 poin s shape used o ea u e ex ac ion and segmen a ion, and esul o image
no maliza ion.
Fig.2. Dis ances compu ed om he shape model o eyeb ow displacemen . h2: ou e eyeb ow co ne
dis ance, h1: inne eyeb ow co ne dis ance. d1: dis ance be ween eyeb ows.
19
Fig.3.De ec ion o o ehead w inkles. Rows om op o bo om: AU1, AU1+AU2 and AU4. Columns om
le o igh : g ey-scaled image, il e ed and masked image and h esholded image.
Fig.4. F om op o bo om: o iginal image, ex ac ed egion, il e ed+masked egion and h esholded egion
esul o he p ocess o isible scle a ex ac ion in he eye egion. Le column images co espond o AU5,
middle column o AU6+12 and igh column o AU43. The a ea o bigges con ou (hsc * dsc) is compu ed
o de e mine he amoun o isible scle a.
20
Fig.5. W inkle de ec ion o AU9 and AU10 in he nose egion. Gabo il e pa ame e s: F = 0.12 pixels-1,
θ = 1.5 ad, σx = 4.2 pixels and σy = 1 pixels and size 32x32 pixels o egion A. Gabo il e pa ame e s: F
= 0.21 pixels-1, θ = 1.9 ad, σx = 1.9 pixels, σy = 0.7 pixels and size 32x32 pixels o egion C.
Fig.6. Fil e ing esul in he mou h egion when pe o ming di e en AUs. Top igh image displays he
labelling o he nine sub egions conside ed. F om op o bo om: es posi ion, AU12 le and igh ac i e,
AU15 ac i e, AU17 ac i e.
Fig.7. O iginal, il e ed and h esholded image in he p occess o ex ac ion o he nasolabial w inkle.
Angle and in ensi y a e compu ed om he h esholded image by pixel pa e n sea ching.
21
Fig.8. Compe i i e + Habi ua ion based ne wo k schema used o acial emo ion ecogni ion.
22
Fig.9. Image sequences co esponding o su p ise, sadness and ea along wi h hei co esponding sys em
ou pu s. Images co espond o one o each 4 ames o he o iginal sequence. Neu al Ne wo k pa ame e s: h
= 0.02, A = 0.3, B = 1, D = 5, E = 0.04, F = 0.08.
23
Fig.10. Image sequences co esponding o pe son 113, sequence 8 o he Cohn-Kanade da abase [8] wi h i s
co esponding sys em ou pu s. Neu al Ne wo k pa ame e s: h = 0.03, A = 0.3, B = 1, D = 5, E = 0.04, F =
0.08.
24
Fig.11: Image sequences co esponding o pe son 111, sequence 7 o he Cohn-Kanade da abase [8] wi h i s
co esponding sys em ou pu s o he o iginal sequence (bo om le ) and a longe sequence buil om
epea ing he las ame (bo om igh ). Neu al Ne wo k pa ame e s: h = 0.03, A = 0.3, B = 1, D = 5, E =
0.04, F = 0.08.
25
Fig.12. Sys em ou pu s o a 40 second sequence whe e he six uni e sal emo ion exp essions a e
pe o med. F om le o igh pe o med exp essions we e: joy, ea , ange , disgus , sadness, joy, alking,
su p ise. Neu al Ne wo k pa ame e s: h = 0.02, A = 0.3, B = 1, D = 5, E = 0.04, F = 0.30.