Saliency om hie a chical adap a ion h ough deco ela ion and a iance
no maliza ion
An ´
on Ga cia-Diaz, Xos´
e R. Fdez Vidal, Xos´
e M. Pa do, Raquel Dosil
Compu e Vision G oup, Dep . o Elec onics and Compu e Science, Uni e si y o San iago de Compos ela
Abs ac
This pape p esen s a no el app oach o isual saliency ha elies on a con ex ually adap ed ep esen a ion p oduced
h ough adap i e whi ening o colo and scale ea u es. Unlike p e ious models, he p oposal is g ounded on he
speci ic adap a ion o he basis o low le el ea u es o he s a is ical s uc u e o he image. Adap a ion is achie ed
h ough deco ela ion and con as no maliza ion in se e al s eps in a hie a chical app oach, in compliance wi h coa se
ea u es desc ibed in biological isual sys ems. Saliency is simply compu ed as he squa e o he ec o no m in he
esul ing ep esen a ion. The pe o mance o he model is compa ed wi h se e al s a e-o - he-a models, in p edic ing
human ixa ions using h ee di e en eye- acking da ase s. Re e ing his measu e o he pe o mance o human
p io i y maps, he model is p o ed o be he only one able o keep he same beha io h ough di e en da ase s,
showing ee o biases. Mo eo e , i is able o p edic a wide se o ele an psychophysical obse a ions, o ou
knowledge, no ep oduced oge he by any o he model be o e.
Keywo ds: saliency, bo om-up, eye ixa ions, deco ela ion, whi ening, isual a en ion
1. In oduc ion
Resea ch on he es ima ion o isual saliency has ex-
pe ienced an inc easing ac i i y in he las yea s om
bo h compu e ision and neu oscience pe spec i es,
gi ing ise o a numbe o imp o ed app oaches. Fu -
he mo e, a wide di e si y o applica ions based on
saliency a e being p oposed ha ange om image e-
a ge ing [1] o human-like obo su eillance [2], objec
lea ning and ecogni ion [3, 4, 5], objec ness de ini ion
[6], image p ocessing o e inal implan s [7], and many
o he s.
Exis ing app oaches o isual saliency ha e adop ed
a numbe o qui e di e en s a egies. A i s g oup, in-
cluding many ea ly models, is e y in luenced by psy-
chophysical heo ies suppo ing a pa allel p ocessing
o se e al ea u e dimensions. Models in his g oup
a e pa icula ly conce ned wi h biological plausibili y
in hei o mula ion, and hey eso o he modeling o
isual unc ions. Ou s anding examples can be ound in
[8] o in [9]. Mos ecen models a e in a second g oup
Email add esses: [email p o ec ed] (An ´
on
Ga cia-Diaz), [email p o ec ed] (Xos´
e R. Fdez Vidal),
[email p o ec ed] (Xos´
e M. Pa do), [email p o ec ed]
(Raquel Dosil)
ha b oadly aims o es ima e he in e se o he p oba-
bili y densi y o a se o low le el ea u es by di e en
p ocedu es. In his kind o models, low le el ea u es
a e usually ob ained by an o -line p ocess o s a is ical
analysis o a la ge se o images, aiming o ep esen he
se o na u al images. Saliency is compu ed on hese
ea u es h ough a pa icula es ima ion o imp obabil-
i y. Ou s anding examples o hese models a e he ap-
p oaches o [10, 11, 12, 13]. O he models, al hough
wi hou an explici g ound, can also be in e p e ed om
an in o ma ion heo e ic pe spec i e in e ms o es ima-
ions o he in e se o he p obabili y densi y. Fo in-
s ance, hose models ha seek dis inc i e spec al ea-
u es in he domain o he spa ial equencies, like he
models p oposed in [14, 15], bu also a mo e ecen
model ha compu es dis ances in a colo space o di -
e en spa ial equency bands [16].
App oaches ha a e no s ic ly da a-d i en include
he combina ion o saliency wi h seman ic maps, ying
o ca ch a ac i eness o aces, pe sons, and o he ob-
jec s [17, 18], o he ad-hoc adap a ion o saliency mod-
els o di e en da ase s by lea ning weigh s ha op i-
mize he p edic ion o human ixa ions in hose da ase s
[19]. Also, he adap a ion o he spa io-ch oma ic ep-
esen a ion o he speci ic da ase om lea ning speci i-
P ep in submi ed o Image and Vision Compu ing Augus 30, 2011
cally deco ela ed coo dina es [20] o independen com-
ponen s [21] has been p oposed. These wo las ap-
p oaches al eady poin o bene i s om he adap a ion o
he ea u e basis. Howe e hese app oaches a e mos ly
ad-hoc, since hey ely on an o -line compu a ion ap-
plied o a speci ic da ase . They do no p oduce a ep e-
sen a ion adap ed o each speci ic image.
1.1. Ou app oach
A na u al app oxima ion o sample dis inc i eness
can be done by compu ing he s a is ical dis ance in a
ep esen a i e coo dina e sys em. This can be simply
done h ough a ec o no m compu a ion i such sys-
em is s a is ically whi ened. The eby, i makes sense
o hink in he adap a ion – h ough whi ening– o he
ea u e basis o he speci ic s a is ical s uc u e o a pa -
icula image, conside ing pixels as samples. The e-
sul ing ep esen a ion would yield a simple and s aigh
measu e o poin saliency h ough a ec o no m com-
pu a ion.
Howe e , ypical schemes o s a is ical whi ening
ha e a cubic o highe complexi y on he numbe o co-
o dina es, while linea on he numbe o samples. These
ac s p e en hei use wi h ep esen a ions o images in-
ol ing h ee colo componen s, se e al scales and se -
e al o ien a ions. In his pape we gene alize a p elim-
ina y app oach [22] o o e come his p oblem by im-
possing a whi ening ans o ma ion independen ly on
educed g oups o ea u e componen s.
The p oposed app oach g ounds on a classical hie a -
chical decomposi ion o images ha i s sepa a es ch o-
ma ic componen s, and nex pe o ms on each o hem
a mul iscale and mul io ien ed decomposi ion. Such ap-
p oach is coa sely inspi ed in he image ep esen a ion
desc ibed in ea ly s ages o he isual pa hway. Be-
sides, di e en implemen a ions and simpli ica ions o
he same can be ound in a a ie y o ea ly and e-
cen models o compu e ision wi h di e en pu poses.
The e o e, we p opose o apply on-line whi ening on
ch oma ic componen s in a i s s age. This ope a ion is
ollowed by a mul io ien a ion and mul iscale decompo-
si ion o he esul ing whi ened ch oma ic componen s.
Nex , u he whi ening is impossed o g oups o o i-
en ed scales o each whi ened ch oma ic componen .
This s a egy allows o keep he numbe o componen s
in ol ed in whi ening limi ed, o e coming p oblems o
compu a ional complexi y.
As a esul , an speci ically adap ed ep esen a ion o
he image a ises. The esul ing image componen s ha e
ze o mean and uni s o a iance. As well, hey a e pa ly
deco ela ed. To ob ain a saliency map, we simply com-
pu e poin dis inc i eness by aking, o each pixel, he
squa ed ec o no m in his ep esen a ion di ided by
he sum o he same ac oss all he pixels.
The p oposed model is alida ed and compa ed wi h
s a e-o - he-a app oaches by measu ing he p edic-
i e capabili y o human ixa ions in h ee open access
da ase s h ough s a e-o he-a p ocedu es based on
Recei e Ope a ing Cha ac e is ic (ROC) analysis and
Kullback-Leible di e gences (KLD). Addi ionally, he
model will be shown o ep oduce a wide se o ele an
psychophysical esul s in which o he models show ail-
u es.
The pape is o ganized as ollows. Sec ion 2 p o ides
a de ailed desc ip ion o he AWS model o saliency
compu a ion. Sec ion 3 e alua es he capabili y o he
model in p edic ing eye- ixa ions. Sec ion 4 shows he
abili y o ep oduce a selec ion o psychophysical and
pe cep ual obse a ions. Finally, in Sec ion 5 he main
conclusions o he wo k a e p esen ed.
2. Model
The key poin o he model o saliency p oposed e-
lies on he on-line adap a ion o he basis used o ep-
esen a ion o he speci ic s a is ical s uc u e o he im-
age. This implies a s ep beyond he adap a ion o a gi en
se –like he se o na u al images– ha is unde he
decomposi ion me hods o mos exis ing app oaches o
saliency. This adap a ion uses pixels as s a is ical sam-
ples and seeks o a se o deco ela ed and whi ened
coo dina es, able o deal wi h he in o ma ion p esen in
he image and o also p o ide a eliable es ima ion o he
s a is ical dis ance o each sample –pixel– o he cen e
o he dis ibu ion. The e o e, he p oposed model will
be e e ed o as he adap i e whi ening saliency (AWS)
model.
2.1. Ch oma ic decomposi ion and whi ening
Ch oma ic componen s unde go he i s adap a i e
s age. Each pixel in he image has an associa ed ec o
o ed ( ), g een (g) and blue (b) componen s. In gen-
e al, he ( ,g,b) coo dina es a e highly co ela ed in he
ensemble o samples. P o ided he co a iance ma ix in
hese coo dina es is:
C gb =
σ2
σ g σ b
σ g σ2
gσgb
σ b σgb σ2
b
(1)
we ypically ha e ha all he elemen s a e non-ze o and
non-negligible. Some colo spaces (e.g. he Lab model)
educe his co ela ion be ween componen s by p oduc-
ing a ep esen a ion ha is deco ela ed in he se o na -
u al images, bu no necessa ily in speci ic images.
2
To deco ela e colo in o ma ion, he whi ening p o-
cedu e consis ing in deco ela ion and a iance no mal-
iza ion (as desc ibed in he appendix) is simply applied
o he ,g,bcomponen s o he image. Being x1= ,
x2=gand x3=b he RGB coo dina es o any pixel in
he image, hey a e in ol ed in he ans o ma ion
(x1,x2,x3)→(z1,z2,z3) (2)
The eby, we ge a z=(zch
1,zch
2,zch
3) whi ened ep e-
sen a ion wi h a new ec o associa ed o each pixel. In
his ep esen a ion he co a iance ma ix is he inden i y
ma ix, and hus each coo dina e has uni s o a iance.
Indeed, he ec o no m gi es a measu e o ch oma ic
dis inc i eness as he s a is ical dis ance o each colo
poin o he a e age colo . Such a simple measu e is
equi alen o he explana ion p oposed by [23] o colo
sea ch asymme y phenomena epo ed o humans in
a se o simple syn he ic images. Howe e , o compu e
poin saliency in clu e ed na u al scenes, spa ial dis inc-
i eness also mus be aken in o accoun .
Al e na i ely o a RGB colo space, we ha e also
es ed he use o o he colo spaces like he Lab model.
The con e sion om RGB o a Lab model in ol es a
non-linea ans o ma ion. Besides, he Lab model p o-
duces a ep esen a ion ha p ese es, on a e age, pe -
cep ual dis ances. The e o e, di e ences in he esul -
ing deco ela ed componen s and e en an ad an age o
he Lab model may be expec ed. Howe e , we did no
ind a signi ican ad an age o any o he al e na i e
colo spaces o e RGB in ou expe imen al e alua ion.
The eby, he e was no appa en eason o ecode he
RGB images o o he colo space be o e whi ening.
2.2. O ien ed mul iscale decomposi ion and whi ening
We ep esen he spa ial s uc u e by decomposing
each o he whi ened ch oma ic componen s (i.e. zch
1,
zch
2, and zch
3) h ough a measu e o local ene gy a di -
e en spa ial equency bands cen e ed a di e en e-
quency modulus alues (scales) and di e en o ien a-
ions.
To ob ain local ene gy, we use a bank o log-Gabo
il e s, since hei eal and imagina y pa s in he spa ial
domain o m a pai o il e s in phase quad a u e. These
il e s p esen se e al ad an ages o e he Gabo il e s.
Namely, hey ha e a ze o DC componen and a long ail
owa ds high equencies, app oaching be e he ecep-
i e ields o co ical cells [24]. The exp ession o hese
il e s in he equency domain is gi en by:
log Gabo so (ρ, α)=exp
−log (ρ/ρs)2
2log σρs/ρs2
·
exp −(α−αo)2
2(σαo)2!
(3)
being (ρ, α) he spa ial equency in pola coo dina es,
(ρs, αo) he cen al equency o he il e , s he scale
index, and o he o ien a ion index.
In he implemen a ion employed in his pape , ou
o ien a ions (0◦,45◦,90◦,135◦) a e used, se en scales
o he i s z-sco e ( oughly equi alen o luminance),
and only 5 scales o he emaining wo componen s.
This di e ence is jus i ied by he obse a ion ha he
ines and coa ses scales o hese componen s ba ely
showed any ele an in o ma ion. Acco dingly, while
he minimum wa eleng h o he i s z-sco e is 3 pix-
els, 6 pixels o colo ha e been used ins ead. The
use o o ien a ions in colo componen s has been ob-
se ed o imp o e pe o mance, compa ed o he use o
iso opic esponses. Besides, o ien a ion selec i i y o
ch oma ic mul iscale ecep i e ields has been shown o
ake place in V1 and is hough o in luence saliency
[25]. I has been also ied o include iso opic esponses
o luminance in addi ion o he o ien ed esponses, bu
he esul s we e p ac ically he same. Consequen ly,
hey we e conside ed edundan in he compu a ion o
saliency, and disca ded o he sake o e iciency.
The bank o il e s is applied on each o he whi ened
ch oma ic componen s p e iously ob ained. F om he
complex esponse o he il e in a gi en equency band
we compu e local ene gy as he modulus o he esponse
[26][27]. Tha is:
ecos =q(zc∗ os)2+(zc∗hos)2(4)
whe e index cdeno es a whi ened ch oma ic compo-
nen , zcis a e ino opic ep esen a ion o such compo-
nen , and and hdeno e espec i ely he e en symme -
ic log-Gabo gi ing he eal pa o he esponse, and
he odd symme ic log-Gabo gi ing he imagina y pa
o he esponse. They o m indeed a pai o il e s in
phase quad a u e.
The e o e, we ob ain a ep esen a ion o he image
in e ms o he local ene gy co esponding o di e en
whi ened colo componen s, di e en scales and o ien-
a ions.
The nex s ep deals wi h he adap a ion o his ep-
esen a ion ha al eady codes he spa ial s uc u e. To
do so, we ha e chosen o deco ela e and whi en, in-
dependen ly and in pa allel, each se o o ien ed scales
3
o each o he ch oma ic componen s. Fo a gi en
whi ened ch oma ic componen and o ien a ion, each
pixel has a local ene gy alue o each scale. The e o e,
each scale sican be iewed as an o iginal coo dina e
axis. The ensemble o scales de e mines a se o o ig-
inal axis in which each pixel is ep esen ed by a poin
wi h i s coo dina es de e mined by he local ene gy al-
ues o he co esponding scales. F om he ensemble o
samples (all he pixels), we can compu e he co a iance
ma ix in such scale coo dina es. I he numbe o scales
is Ms, hen he co a iance ma ix is a Ms×Msma ix.
Cco =
σ2
co;s1. . . σco;s1sMs
.
.
.....
.
.
σco;s1sMs . . . σ2
co;sMs
(5)
As well known, in na u al images di e en scales a e
highly co ela ed, which in gene al makes all he ma ix
elemen s non-ze o. The e o e, he whi ening p ocedu e
based on deco ela ion and a iance no maliza ion is ap-
plied again o achie e a new se o whi ened scale coo -
dina es zsc
i. The esul ing co a iance ma ix becomes
he iden i y. This new whi ened ea u e basis is com-
posed o axis ha a e shi ed, o a ed and escaled om
he o iginal scale axis. In sum, o each se o scales we
ha e ans o med he o iginal scale coo dina es o pix-
els o new coo dina es ha a e deco ela ed and wi h he
a iance as he no m.
2.3. Saliency
To compu e saliency, we simply use he sum o he
squa ed no m o he ec o s in he ob ained ep esen a-
ion as an es ima ion o pixel (i.e. sample) dis inc i e-
ness and we no malize i o he sum ac oss all he pixels.
Tha is, o each pixel i
kzicok2=zT
icozico (6)
whe e zico is he ec o associa ed o he pixel o a colo
componen cand an o ien a ion owi h as many compo-
nen s as whi ened scales (Ms).
This p o ides a e ino opic measu e o he local ea-
u e con as . In his way, a measu e o conspicui y is
ob ained o each o ien a ion o each o he colo com-
ponen s. The nex s eps in ol e a Gaussian smoo hing
and he addi ion o he maps co esponding o all o
he o ien a ions. Tha is, o a gi en colo componen
c=1...Mcand pixel i, he co esponding saliency (Sic)
is calcula ed:
Sic =
Mo
X
o=1
kzicok2(7)
Colo componen s unde go he same summa ion s ep
o ge a inal map o saliency. Addi ionally, o ease in e -
p e a ion o his map as p obabili y o ecei e a en ion,
i is no malized by he in eg al o he saliency in he im-
age domain (i.e. he o al popula ion ac i i y). Hence,
saliency o a pixel i(Si) is gi en by:
Si=PMc
c=1Sic
PN
i=1PMc
c=1Sic
(8)
The alues ob ained o he ensemble o poin s, a -
anged in a 2D ma ix deli e a map o saliency So he
same dimensions han he inpu image.
The igu e 1 shows a g aphic ou line o he model.
I mus be no iced ha o he app oaches o in eg a-
ion di e en om he squa ed no m ha e been explo ed
like he aw ec o no m, highe powe exponen s o he
ec o no m, and e en an exponen ial ans o ma ion o
he di e en componen s ollowed by a summa ion.The
use o he aw ec o no m achie ed close -bu in e io -
pe o mance in p edic ing ixa ions, while all he o he
app oaches beha ed much wo se. All he al e na i es
ailed in se e al o he psychophysical expe imen s de-
sc ibed in his pape . This obse a ions ag ee wi h he
iew ha saliency is ela ed o he classical s a is ical
dis ance o he ea u e ec o associa ed o a poin om
he cen e o he dis ibu ion o ea u es p esen in he
image. The use o he squa ed ec o no m in ou hie a -
chical whi ening app oach can be iewed as an e icien
es ima ion o an o e all T2o Ho elling in an o iginal
high-dimensional ep esen a ion.
Rega ding he compu a ional complexi y o his im-
plemen a ion, PCA implies a load ha linea ly g ows
wi h he numbe o pixels (N), and in a cubic man-
ne wi h he numbe o componen s (M), speci ically
O(M3+M2N). Since we ha e kep he numbe o com-
ponen s (colo componen s o scales) ixed and small,
he asymp o ic complexi y depends on he numbe o
pixels. This is de e mined by he use o he FFT in he
il e ing p ocess, which is O(Nlog(N)). Mos saliency
models ha e a complexi y which is O(N2) o highe .
3. Compa ison wi h human ixa ions
In he las yea s, he mos ex ended alida ion p oce-
du e o no el models o saliency has been he abili y o
p edic ixa ions eco ded om humans du ing he ee-
iewing o na u al images, wi hou any speci ic goal o
ask [10, 11, 13]. The e a e some al e na i e p ocedu es,
mo e di icul o in e p e s ic ly in e ms o saliency.
Fo example, objec segmen a ion o ask-d i en isual
4
Figu e 1: Adap i e whi ening saliency model.
sea ch. A e y ecen wo k analyses a numbe o mod-
els o saliency h ough he compa ison wi h human a -
ings in a ask o isual sea ch o mili a y ehicles in
pho og aphs, inding s a is ically signi ican co ela ion
o mos models [28]. Howe e , he da ase employed is
op-down biased and also has impo an ea u e biases
(mos ly open g een landscapes and sky). The a ge is
in mos cases non salien due o i s camou lage design,
o akes up a la ge po ion o he image –holding a num-
be o salien and non salien pa s. The eby, he impli-
ca ions o such signi icance in co ela ion esul s aise
impo an di icul ies o in e p e a ion. Is such co ela-
ion ela ed o a gene al es ima ion o saliency o a he
o an e icien de ec ion –o e en segmen a ion– o mil-
i a y ehicles in coun yside scenes?. These conce ns
ha e no been se ou o he es s based on he p edic-
ion o human ixa ions. This p ocedu e does no ely
on e lec i e decisions abou wha is conpicuous, bu e-
lies on a as ac ion when aced o an image. Mo eo e ,
i is p ecisely ela ed o posi ions in he space, no o
an objec o unde e mined and changeable a ea on he
image.
The e o e, a majo goal in he modeling o saliency
is pushing his benchma k u he on.
3.1. Da ase s and models
Th ee open-access eye- acking da ase s o na u al
images ha e been used. In he h ee da ase s he sub-
jec s did no ecei e any speci ic ins uc ion. Tha is,
hey mee he equi emen o ee- iewing. The images
ha e been shown in a amdom o de o each subjec .
The igu e 2 shows h ee example images om each o
he da ase s.
The i s da ase has been published by B uce and
Tso sos and has 120 images and ixa ions om 20 sub-
jec s [29]. Each image has been iewed du ing 4 sec-
onds. I has al eady been used o alida e many s a e-
o - he-a models o bo om-up saliency using di e en
p ocedu es [10, 11, 13]. The e o e, i p o ides a sui -
able e e ence o a ai assessmen o a no el model in
ela ion o exis ing app oaches.
The second da ase has been published by Koo s a
e al. and consis s o 99 images and he co esponding
ixa ions o 31 subjec s [30]. The iewing ime was 5
seconds o each image. One in e es ing p ope y o his
da ase is ha i is o ganized in i e di e en g oups o
images (12 images o animals, 12 o s ee s, 16 o build-
ings, 40 o na u e, and 19 o lowe s o na u al symme-
ies). This ea u e may be expec ed o e eal possible
biases in he models. Besides, i may help o analyze
he causes o a iabili y unde he same expe imen al
condi ions.
Finally, he NUSEF da ase has been chosen because
i is supposed o ha e a s ong emo ional bu den [31].
This a ec i e con en may be expec ed o p oduce an
inc eased human consis ency no ela ed o low le el
ea u es, bu o emo ions ela ed o abs ac concep s
s ongly sugges ed by he images. The e o e, models
o saliency may be expec ed o explain less amoun o
in e subjec consis ency han in he o he da ase s. I is
composed o 758 images obse ed on a e age by 25.3
subjec s. The iewing ime was 5 seconds o each im-
5
Figu e 2: Examples o images om he h ee da ase s used. B uce and Tso sos (le ); Ko s a e al. (cen e ); NUSEF ( igh )
age.
O he wise, we compa e he esul s o he p oposed
model wi h o he 4 models. Namely: The model o Seo
and Milan a based on sel esemblance [13]; he SUN
model [11] ha adop s a bayessian app oach based on
p e iously lea ned image s a is ics; he AIM model ha
decomposes he image wi h independen componen s o
na u al images and uses sel -in o ma ion as a measu e
o dis inc i eness [10]; and inally he classic model o
saliency p oposed by I i e al. [8].
3.2. ROC analysis and KL di e gence
To assess he use ulness o he saliency maps o dis-
c imina e be ween ixa ed and non ixa ed poin s, we
ha e used he a ea unde he cu e (AUC), ob ained
om a ecei e ope a ing cha ac e is ic (ROC) analysis,
and a Kullback-Leible di e gence (KLD) compa ison.
Bo h me hods a ge he capabili y o he saliency maps
o p edic he spa ial dis ibu ion o ixa ions h ough
he compa ison o he dis ibu ions o saliency in ixa ed
e sus non- ixa ed poin s.
To a oid cen e -bias, in each image, only poin s ix-
a ed in ano he image om he same da ase a e used
as non ixa ed poin s. As sugges ed in [32], s anda d e -
o is compu ed h ough a boo s ap echnique, shu ling
he o he images used o ake he non ixa ed poin s,
exac ly like in [11] and in [13]. This las s ep should
no be adop ed i he goal is o assess a combina ion
o saliency and a cen e -bias model. Howe e , since
saliency is da a-d i en and he cen e -bias is a spa ial
bias (wo king ega dless o he speci ic da a), hey a e
di e en mechanisms. Thus, i makes sense o use di -
e en e alua ions ocusing on each o he componen s.
Fu he mo e, he use o he boo s apping me hod
yields a high sensi i i y. As ecen ly shown in [19], a
ROC analysis as used by many au ho s (wi hou a boo -
s apping o p e en he in luence o cen e -bias) aises
p oblems o sensi i i y. Howe e , he use o he boo -
6
s apping p ocedu e yields a s anda d e o ha is ipi-
cally below he 3% o he dynamic ange spanned by
he ob ained alues o he di e en models o bo h he
ROC analysis and he KLD compa ison.
Ne e heless, a double assessmen h ough a ROC
analysis and a KLD compa ison is p o ided o ensu e
he eliabili y o he e alua ion.
3.3. Resul s
The able 1 ga he s he esul s ob ained on he h ee
da ase s. The igu es 3 o 5 p o ide examples o saliency
maps o he bes pe o ming models ha ocus on spe-
ci ic ea u es o suppo he discussion.
The alues shown o he model o I i e al. [8] on he
da ase o B uce and Tso sos a e highe han epo ed
in p e ious wo ks [11, 13] because, ins ead o using
hei saliency oolbox, he o iginal implemen a ion has
been used, as made a ailable o Ma lab (h p://www.
klab.cal ech.edu/~ha el/sha e/gb s.php). Fo
he o he models in his da ase he alues ha we ha e
ob ained a e compa ible wi h hose published in [11]
and [13], hus we ha e espec ed he epo ed alues.
3.3.1. Discussion
Fi s ly, i is wo h no ing ha he esul s wi h bo h
measu es, ROC analysis and KLD, yield an equi alen
anking and equi alen dis ances be ween models on
each o he da ase s. The e a e mino di e ences in
wo g oups o he da ase o Koo s a bu hey do no
gi e ise o any ema kable di e ence in he e alua ion.
The e o e, he commen s ha ollow hold o bo h. Re-
ga ding he sensi i i y o he measu es, he s anda d e -
o emains below he 3% o he spanned ange o al-
ues.
Conside ing he h ee da ase s, he anking o mod-
els yields only a single change o posi ions be ween he
model o Seo and Milan a and he AIM model in he
NUSEF da ase . Al hough wi h a iable dis ances, he
es o models ank he same posi ion. The AWS model
holds clea ly he i s posi ion in he h ee da ase s.
Howe e , looking a he i e g oups o he da ase o
Koo s a e al., se e al changes o posi ion in ol ing di -
e en models occu . As a esul , he e is no mo e a
clea cohe en anking. E en hough, he AWS main-
ains he bes pe o mance wi h a dis ance on he nex
a beyond he s anda d e o , excep o he buildings
g oup in which he model o Seo and Milan a achie es
a sligh ly be e esul , bu wi hin he unce ain y limi s
es ablished.
O he wise, he a ia ions in pe o mance ac oss he
da ase s may be used o look o biases in he models.
The s ong ad an age o AWS in he g oup o lowe s
and na u al symme ies inds an explana ion in he ex-
amples shown in he igu e 3. The model o Seo and Mi-
lan a and he AIM model miss comple ely he saliency
o na u al symme ies ha ca ch a conside able amoun
o ixa ions in hese images. In con as , he AWS model
manages o cap u e he saliency o symme ies in na u-
al scenes.
Besides, o he ac o s appea o con ibu e o he ad-
an age o he AWS model. I shows a sensi i i y o
salien high equency pa e ns like he s iped pa e n
on he small head o a bu e ly ha is shown in he
igu e 4. Fo he model o Seo and Milan a , he head
seems o be jus ano he pa o an edge a ound he bu -
e ly. Addi ionally, he beha io o ou model when
aced o colo single ons appea s o be mo e obus .
The igu e 5 shows a e ealing example. The yellow
and ed peppe s a e among he mos salien objec s o
bo h he AWS model and humans. In con as , he AIM
model and pa icula ly he model o Seo and Milan a
ind mo e salien he objec s in he uppe pa o he im-
age, hus showing a lack o sensi i i y o colo pop-ou
in his na u al con ex .
3.4. Compa ison wi h human p io i y
Some ques ions in ela ion o he assessmen p oce-
du e a ise. Is i sui able he s a is ical signi icance o
compa e he models o is i oo igh in p ac ice?. I
may occu ha di e ences in model pe o mance a e
simila o he a iabili y shown by humans hemsel es,
while being s a is ically signi ican . O he wise, wha a e
he easons o he high a ia ion in he absolu e al-
ues ac oss he da ase s? Is he e any means o c ea e
a da ase ha p o ides eliable and de ini i e esul s in
anking he models?.
I is clea ha he explana ion is no a di e en expe -
imen al se up since he la ges a ia ion is ound ac oss
he g oups o he Koo s a da ase , all unde he same
se up. Any explana ion should be ela ed o di e ences
in he image con en , di e ences in he associa ed hu-
man beha io , and di e en biases in he models o
saliency.
In o de o explo e answe s o he aised ques ions,
we p opose o compa e he pe o mance o models wi h
he pe o mance o single subjec s. To assess he p edic-
i e pe o mance o a single subjec we eso o p io i y
maps de i ed om he ixa ions o he subjec on each
image. F om his measu e, we may compu e an es ima-
ion o he a e age subjec pe o mance and an es ima-
ion o human pe o mance a iabili y.
7
Table 1: AUC alues ob ained wi h di e en models o saliency o bo h o he da ase s o B uce and Tso sos and Koo s a e al. S anda d e o s,
ob ained like in [11], ange 0.0004-0.0008. Fo he g oups o he Koo s a e al. da ase , s anda d e o s ange 0.0010-0.0018. (* Resul s epo ed
by [11]; ** Resul s epo ed by he au ho s).
Model B uce and
Tso sos
da ase
NUSEF
da ase
Koo s a e al. da ase
Whole
da ase
Buildings Na u e Animals Flowe s S ee
AWS 0.7106 0.6035 0.6205 0.6105 0.5815 0.6565 0.6374 0.7020
Seo and Mil. 0.6896** 0.5802 0.5933 0.6136 0.5530 0.6445 0.5602 0.6907
AIM 0.6727* 0.5902 0.5842 0.5766 0.5628 0.5953 0.5881 0.6393
SUN 0.6682* 0.5782 0.5705 0.5514 0.5484 0.5401 0.6100 0.6458
I i e al. 0.6456 0.5655 0.5702 0.5814 0.5478 0.6200 0.5217 0.6509
Gao e al. 0.6395* – – – – – –
Table 2: KL di e genge alues ob ained wi h di e en models o saliency o bo h o he da ase s o B uce and Tso sos and Koo s a e al. S anda d
e o s, ob ained like in [11], ange 0.001-0.002 o bo h o he da ase s and he g oups. (* Resul s epo ed by [11]; ** Resul s epo ed by he
au ho s).
Model B uce and
Tso sos
da ase
NUSEF
da ase
Koo s a e al. da ase
Whole
da ase
Buildings Na u e Animals Flowe s S ee
AWS 0.321 0.071 0.099 0.109 0.058 0.188 0.142 0.307
Seo and Mil. 0.278** 0.048 0.071 0.110 0.049 0.175 0.057 0.281
AIM 0.203* 0.055 0.055 0.070 0.045 0.105 0.085 0.197
SUN 0.210* 0.043 0.039 0.046 0.033 0.049 0.097 0.173
I i e al. 0.175 0.033 0.038 0.069 0.032 0.109 0.025 0.200
3.4.1. Human p io i y pe o mance
To implemen his measu e, p io i y maps de i ed
om ixa ions ha e been used, ollowing he me hod
desc ibed by [30]. This me hod lies in he sub ac ion
o he dis ance be ween each poin and i s nea es ixa-
ion om he maximum possible dis ance in he image.
As a esul , ixa ed poin s ha e he maximum alue and
non ixa ed poin s ha e a alue ha dec eases linea ly
wi h he dis ance o he nea es ixa ion. The esul ing
maps can be used as p obabili y dis ibu ions o sub-
jec s ixa ions (p io i y maps), and can be conside ed as
subjec i e measu es o saliency. A leas wi h ew ix-
a ions pe subjec , as i is he case, his me hod yields
be e p edic i e esul s han he app oach o compu e
p io i y maps based on il e ing o ixa ions wi h Gaus-
sians ke nels [29]. This las app oach ipically assigns
ze o o decimal p io i y o poin s beyond 2.3◦o isual
angle, since usually he wid h o he Gaussian is ixed
o 1◦and he ampli ude is ixed o 255, which is he
maximum o he dynamic ange used in he ROC anal-
ysis. Consequen ly, wi h ew ixa ions by subjec , he
p io i y maps p esen alues below 1 o posi ions ha
can be close o ixa ions. The e o e, in a ROC analysis
all hese loca ions a e equally conside ed ze o p io i y
poin s. Exac ly he same as poin s much u he om
any ixa ion.
Fu he mo e, he linea dis ance-based me hod is pa-
ame e ee. The eby, we do no need o make assump-
ions on he ange o poin s ha may ha e in luenced
a gi en ixa ion. O cou se, i can be a gued ha i is
no jus i ied o assume ha p io i y d ops linea ly wi h
dis ance o ixa ions. Ne e heless, i seems ac ually
easonable o assume ha p io i y d ops mono onically
wi h dis ance o he nea es ixa ion. I he me hod o
compa e and e alua e maps is in a ian o mono onic
ans o ma ions, as ROC analysis is, hen he e is no
issue wi h using linea , o any o he mono onic maps.
Hence, h ough a ROC analysis, he same one employed
o e alua e models o saliency, he capabili y o hese
maps o p edic he ixa ions o he se o subjec s can
be assessed, wi hou conce ns on he dynamic ange o
he ROC analysis. I mus be no iced ha we ha e no
used he a e aged p io i y maps shown in he igu es
3 o 5, bu maps compu ed speci ically o each o he
subjec s using he p ocedu e desc ibed abo e.
The p e ious e alua ion o each subjec has been
done, only o hose wi h ixa ions o all o he im-
ages. One indi idual has been excluded o he da ase
o B uce and Tso sos, whose de ia ion om he a e -
age o humans was la ge han wice he s anda d de ia-
ion, and who also had jus one ixa ion in many images.
This yields p io i y maps om 9 subjec s o he da ase
o B uce and Tso sos, and 25 subjec s o he da ase o
Koo s a. On he NUSEF da ase ou app oach o de i e
8
Figu e 3: Examples o esul s wi h 4 images wi h a dominan symme ic poin . Fo compa ison, human p io i y maps p o ided by he au ho s a e
shown. B uce and Tso sos de i ed p io i y om ixa ions using Gaussian ke nels [29], while Koo s a e al. used a dis ance- o- ixa ion ans o m
[30]. This explains he no iceable di e ences in he dynamic ange o he p io i y maps.
subjec p io i y maps inds a p oblem: no subjec has
obse ed all he images. Ne e heless, all he images
ha e been obse ed by a leas 13 subjec s. The e o e,
we ha e buil 13 pseudo-subjec s ga he ing he p io i y
maps o he i s 13 obse e s ha iewed each o he
images, ollowing he o de o subjec s p o ided by he
au ho s.
The pe o mance o he p io i y maps associa ed o a
gi en subjec was ob ained h ough he assessmen wi h
ixa ions o o he subjec s in he da ase . Compu ing he
a e age, we ha e he a e age pe o mance o p io i y
maps associa ed o di e en subjec s. Besides, he dou-
ble o he s anda d de ia ion p o ides an es ima ion o
he ange o p edic i e pe o mance o he 95% o hu-
mans, unde he assump ion o a no mal dis ibu ion o
AUC p io i y alues. This was ue o he da ase s and
g oups s udied, wi h a ku osis alue e y close o 3.
Mo eo e , his in e al o a iabili y be ween subjec s
can be also used as a measu e o he minimum ele an
dis ance be ween wo models. Di e ences lowe han
such a iabili y may be ega ded as p oducing no p ac-
ical e ec on pe o mance. The esul s a e gi en in he
able 3.
3.4.2. Saliency e sus p io i y
A a i s look, we can see how he pe o mance
o p io i y a ies ac oss da ase s simila ly o saliency
maps. The a iabili y o subjec pe o mance in a gi en
da ase is well an o de o magni ude highe han he
s a is ical signi icance o he measu e o pe o mance,
excep o he NUSEF ha shows much less a iabili y.
Rema kably, he pe o mance o he AWS model is
compa ible wi h he es ima ed human p io i y pe o -
mance o he h ee da ase s and all he g oups. The
model o Seo and Milan a is also compa ible wi h he
a e age human o he da ase o B uce and Tso sos, and
o wo o he i e g oups o Koo s a e al.. Howe e
his compa ibili y does no hold o ei he he whole
da ase o Koo s a e al. o he NUSEF da ase . The
model by B uce and Tso sos is only ma ginally com-
pa ible wi h he a e age human wi h hei own da ase .
9
Figu e 12: Examples o saliency-based segmen a ion: o iginal image (le s), saliency maps (cen e ), and p o o-objec s ( ig h) a aganged in wo
e ical blocks. Six o he images ha e been ob ained om [12], he es a e ou s.
Tha is,
x=xj→y=yj→z=zj(A.1)
wi h j=1...M, whe e Mis he numbe o componen s.
The whi ening p ocedu e can be summa ized in wo
s eps. Fi s , as well known, p incipal componen s esul
om diagonaliza ion o he co a iance ma ix, o de ing
eigen alues (lj) om highe o lowe . To compu e he
co a iance ma ix he e a e Nsamples, as many as he
numbe o pixels in he inpu image. The whi ened z
ep esen a ion is hen ob ained h ough no maliza ion
by a iance, gi en by he eigen alues. This means ha
o each p incipal componen :
zj=yj
plj
;j∈[1,M] (A.2)
These z-sco es yield a whi ened ep esen a ion, wi h
he co a iance ma ix being he uni y ma ix. The
squa ed no m o a ec o in hese coo dina es is in ac
he s a is ical dis ance in he o iginal xcoo dina es.
Appendix B. Rep oducibili y
A Ma lab p-code ile o ep oduce he expe imen-
al esul s epo ed in his pape as well as all he
saliency maps compu ed wi h he AWS model o
he h ee eye- acking da ase s a e a ailable on he
web page h p://www-g a.dec.usc.es/pe soal/
xose. idal/ esea ch/aws/AWSmodel.h ml.
Re e ences
[1] Z. Liu, H. Yan, L. Shen, K. N. Ngan, Z. Zhang, Adap i e im-
age e a ge ing using saliency-based con inuous seam ca ing,
Op ical Enginee ing 49 (2010) 1–10.
[2] J. Ruesch, M. Lopes, A. Be na dino, J. Ho ns ein, J. San os-
Vic o , R. P ei e , Mul imodal saliency-based bo om-up a en-
ion a amewo k o he humanoid obo icub, in: In . Con . on
Robo ics and Au oma ion (ICRA), pp. 962–967.
16
[3] C. Kanan, G. Co ell, Robus classi ica ion o objec s, aces,
and lowe s using na u al image s a is ics, in: IEEE in . Con .
on Compu e Vision and Pa e n Recogni ion (CVPR).
[4] J. Ha el, C. Koch, On he op imali y o spa ial a en ion o
objec de ec ion, in: A en ion in Cogni i e Sys ems 2009, pp.
1–14.
[5] D. Gao, S. Han, N. Vasconcelos, Disc iminan saliency, he de-
ec ion o suspicious coincidences, and applica ions o isual
ecogni ion, IEEE T ansac ions on Pa e n Analysis and Ma-
chine In elligence 31 (2009) 989.
[6] B. Alexe, T. Deselae s, V. Fe a i, Wha is an objec ?, in: IEEE
Con . on Compu e Vision and Pa e n Recogni ion (CVPR), pp.
73–80.
[7] N. Pa ikh, L. I i, J. Weiland, Saliency-based image p ocessing
o e inal p os heses, Jou nal o Neu al Enginee ing 7 (2010)
016006.
[8] L. I i, C. Koch, E. Niebu , A model o saliency-based isual
a en ion o apid scene analysis, IEEE T ansac ions on Pa e n
Analysis and Machine In elligence 20 (1998) 1254–1259.
[9] O. Le Meu , P. Le Calle , D. Ba ba, D. Tho eau, A cohe -
en compu a ional app oach o model bo om-up isual a en-
ion, IEEE T ansac ions on Pa e n Analysis and Machine In el-
ligence 28 (2006) 802–817.
[10] N. D. B uce, J. K. Tso sos, Saliency, a en ion, and isual
sea ch: An in o ma ion heo e ic app oach, Jou nal o Vision
9 (2009) 5.
[11] L. Zhang, M. H. Tong, T. K. Ma ks, H. Shan, G. W. Co ell,
SUN: a bayesian amewo k o saliency using na u al s a is ics,
Jou nal o Vision 8 (2008) 32.
[12] X. Hou, L. Zhang, Dynamic isual a en ion: Sea ching o
coding leng h inc emen s, in: Ad ances in Neu al In o ma ion
P ocessing Sys ems (NIPS), olume 21, pp. 681–688.
[13] H. J. Seo, P. Milan a , S a ic and space- ime isual saliency
de ec ion by sel - esemblance, Jou nal o Vision 9 (2009) 12–
15.
[14] X. Hou, L. Zhang, Thumbnail gene a ion based on global
saliency, in: Ad ances in Cogni i e Neu odynamics (ICCN),
pp. 999–1003.
[15] C. Guo, Q. Ma, L. Zhang, Spa io- empo al saliency de ec ion
using phase spec um o qua e nion ou ie ans o m, in: IEEE
Con on Compu e Vision and Pa e n Recogni ion (CVPR).
[16] R. Achan a, S. Hemami, F. Es ada, S. Sss unk, F equency-
uned salien egion de ec ion, in: IEEE Con . on Compu e
Vision and Pa e n Recogni ion (CVPR).
[17] M. Ce , E. P. F ady, C. Koch, Faces and ex a ac gaze in-
dependen o he ask: Expe imen al da a and compu e model,
Jou nal o ision 9 (2009).
[18] T. Judd, K. Ehinge , F. Du and, A. To alba, Lea ning o p edic
whe e humans look, in: IEEE 12 h In .l Con . on Compu e
Vision, IEEE, pp. 2106–2113.
[19] Q. Zhao, C. Koch, Lea ning a saliency map using ixa ed loca-
ions in na u al scenes, Jou nal o ision 11 (2011).
[20] J. an de Weije , T. Ge e s, A. D. Bagdano , Boos ing colo
saliency in image ea u e de ec ion, IEEE T ansac ions on Pa -
e n Analysis and Machine In elligence 28 (2006) 150–156.
[21] N. D. B uce, P. Ko np obs , On he ole o con ex in p obabilis-
ic models o isual saliency, in: IEEE In e na ional con e ence
on image p ocessing (ICIP), p. 30893092.
[22] A. Ga cia-Diaz, X. Fdez-Vidal, X. Pa do, R. Dosil, Deco ela-
ion and dis inc i eness p o ide wi h human-like saliency, in:
Ad anced Concep s o In elligen Vision Sys ems, pp. 343–
354.
[23] R. Rosenhol z, A. L. Nagy, N. R. Bell, The e ec o backg ound
colo on asymme ies in colo sea ch, Jou nal o Vision 4 (2004)
224–240.
[24] D. J. Field, Rela ions be ween he s a is ics o na u al images
and he esponse p ope ies o co ical cells, Jou nal o he Op-
ical Socie y o Ame ica A 4 (1987) 2379–2394.
[25] L. Zhaoping, R. J. Snowden, A heo y o a saliency map in
p ima y isual co ex (V1) es ed by psychophysics o colou o -
ien a ion in e e ence in ex u e segmen a ion, Visual Cogni ion
14 (2006) 911–933.
[26] P. Ko esi, In a ian measu es o image ea u es om phase in-
o ma ion, Ph.D. hesis, Depa men o Psychology, Uni e si y
o Wes e n Aus alia, 1996.
[27] M. C. Mo one, D. C. Bu , Fea u e de ec ion in human ision:
A phase-dependen ene gy model 1998, in: P oc. o he Royal
Socie y o London. Se ies B, Biological Sciences, pp. 221–245.
[28] A. Toe , Compu a ional e sus psychophysical image saliency:
A compa a i e e alua ion s udy, IEEE T ansac ions on Pa e n
Analysis and Machine In elligence (p ep in ) (2011).
[29] N. B uce, J. Tso sos, Saliency based on in o ma ion maximiza-
ion, in: Ad ances in Neu al In o ma ion P ocessing Sys ems
(NIPS), olume 18, p. 155.
[30] G. Koo s a, A. Nede een, B. de Boe , Paying a en ion o
symme y, in: P oc. o he B i ish Machine Vision Con e ence
(BMVC), pp. 1115–1125.
[31] S. Ramana han, H. Ka i, N. Sebe, M. Kankanhalli, T. S. Chua,
An eye ixa ion da abase o saliency de ec ion in images, in:
Eu opean Con . on Compu e Vision (ECCV), pp. 30–43.
[32] B. W. Ta le , R. J. Baddeley, I. D. Gilch is , Visual co ela es o
ixa ion selec ion: E ec s o scale and ime, Vision Resea ch 45
(2005) 643–659.
[33] W. Einhuse , M. Spain, P. Pe ona, Objec s p edic ixa ions be -
e han ea ly saliency, Jou nal o Vision 8 (2008) 18.
[34] E. Bi mingham, W. F. Bischo , A. Kings one, Saliency does
no accoun o ixa ions o eyes wi hin social scenes, Vision
Resea ch 49 (2009) 2992–3000.
[35] H. C. No hdu , The conspicuousness o o ien a ion and mo ion
con as , Spa ial Vision 7 (1993) 341–363.
[36] D. Gao, V. Mahade an, N. Vasconcelos, On he plausibili y o
he disc iminan cen e -su ound hypo hesis o isual saliency,
Jou nal o Vision 8 (2008) 13.
[37] A. T eisman, S. Go mican, Fea u e analysis in ea ly ision:
E idence om sea ch asymme ies, Psychological Re iew 95
(1988) 15–48.
[38] J. M. Wol e, T. S. Ho owi z, Wha a ibu es guide he deploy-
men o isual a en ion and how do hey do i ?, Na u e Re iews
Neu oscience 5 (2004) 495–501.
[39] A. Hy inen, E. Oja, A as ixed-poin algo i hm o indepen-
den componen analysis, Neu al Compu a ion 9 (1997) 1483–
1492.
[40] J. F. Ca doso, A. Souloumiac, T. Pa is, Blind beam o ming o
non-Gaussian signals, in: IEE P oceedings on Rada and Signal
P ocessing 1993, olume 140, pp. 362–370.
17