scieee Science in your language
[en] (orig)

Saliency from hierarchical adaptation through decorrelation and variance normalization

Abstract

This paper presents a novel approach to visual saliency that relies on a contextually adapted representation produced through adaptive whitening of color and scale features. Unlike previous models, the proposal is grounded on the specific adaptation of the basis of low level features to the statistical structure of the image. Adaptation is achieved through decorrelation and contrast normalization in several steps in a hierarchical approach, in compliance with coarse features described in biological visual systems. Saliency is simply computed as the square of the vector norm in the resulting representation. The performance of the model is compared with several state-of-the-art approaches, in predicting human fixations using three different eye-tracking datasets. Referring this measure to the performance of human priority maps, the model proves to be the only one able to keep the same behavior through different datasets, showing free of biases. Moreover, it is able to predict a wide set of relevant psychophysical observations, to our knowledge, not reproduced together by any other model before.

Read accessible full text

Saliency from hierarchical adaptation through decorrelation and variance normalization

Author: García Díaz, Antón; Fernández Vidal, Xosé Ramón; Pardo López, Xosé Manuel; Dosil Lago, Raquel
Publisher: Elsevier
Year: 2012
DOI: 10.1016/j.imavis.2011.11.007
Source: https://minerva.usc.es/bitstreams/13b13c65-fd34-4939-88f5-6bf576c6810f/download
Saliency om hie a chical adap a ion h ough deco ela ion and a iance
no maliza ion
An ´
on Ga cia-Diaz, Xos´
e R. Fdez Vidal, Xos´
e M. Pa do, Raquel Dosil
Compu e Vision G oup, Dep . o Elec onics and Compu e Science, Uni e si y o San iago de Compos ela
Abs ac
This pape p esen s a no el app oach o isual saliency ha elies on a con ex ually adap ed ep esen a ion p oduced
h ough adap i e whi ening o colo and scale ea u es. Unlike p e ious models, he p oposal is g ounded on he
speci ic adap a ion o he basis o low le el ea u es o he s a is ical s uc u e o he image. Adap a ion is achie ed
h ough deco ela ion and con as no maliza ion in se e al s eps in a hie a chical app oach, in compliance wi h coa se
ea u es desc ibed in biological isual sys ems. Saliency is simply compu ed as he squa e o he ec o no m in he
esul ing ep esen a ion. The pe o mance o he model is compa ed wi h se e al s a e-o - he-a models, in p edic ing
human ixa ions using h ee di e en eye- acking da ase s. Re e ing his measu e o he pe o mance o human
p io i y maps, he model is p o ed o be he only one able o keep he same beha io h ough di e en da ase s,
showing ee o biases. Mo eo e , i is able o p edic a wide se o ele an psychophysical obse a ions, o ou
knowledge, no ep oduced oge he by any o he model be o e.
Keywo ds: saliency, bo om-up, eye ixa ions, deco ela ion, whi ening, isual a en ion
1. In oduc ion
Resea ch on he es ima ion o isual saliency has ex-
pe ienced an inc easing ac i i y in he las yea s om
bo h compu e ision and neu oscience pe spec i es,
gi ing ise o a numbe o imp o ed app oaches. Fu -
he mo e, a wide di e si y o applica ions based on
saliency a e being p oposed ha ange om image e-
a ge ing [1] o human-like obo su eillance [2], objec
lea ning and ecogni ion [3, 4, 5], objec ness de ini ion
[6], image p ocessing o e inal implan s [7], and many
o he s.
Exis ing app oaches o isual saliency ha e adop ed
a numbe o qui e di e en s a egies. A i s g oup, in-
cluding many ea ly models, is e y in luenced by psy-
chophysical heo ies suppo ing a pa allel p ocessing
o se e al ea u e dimensions. Models in his g oup
a e pa icula ly conce ned wi h biological plausibili y
in hei o mula ion, and hey eso o he modeling o
isual unc ions. Ou s anding examples can be ound in
[8] o in [9]. Mos ecen models a e in a second g oup
Email add esses: [email p o ec ed] (An ´
on
Ga cia-Diaz), [email p o ec ed] (Xos´
e R. Fdez Vidal),
[email p o ec ed] (Xos´
e M. Pa do), [email p o ec ed]
(Raquel Dosil)
ha b oadly aims o es ima e he in e se o he p oba-
bili y densi y o a se o low le el ea u es by di e en
p ocedu es. In his kind o models, low le el ea u es
a e usually ob ained by an o -line p ocess o s a is ical
analysis o a la ge se o images, aiming o ep esen he
se o na u al images. Saliency is compu ed on hese
ea u es h ough a pa icula es ima ion o imp obabil-
i y. Ou s anding examples o hese models a e he ap-
p oaches o [10, 11, 12, 13]. O he models, al hough
wi hou an explici g ound, can also be in e p e ed om
an in o ma ion heo e ic pe spec i e in e ms o es ima-
ions o he in e se o he p obabili y densi y. Fo in-
s ance, hose models ha seek dis inc i e spec al ea-
u es in he domain o he spa ial equencies, like he
models p oposed in [14, 15], bu also a mo e ecen
model ha compu es dis ances in a colo space o di -
e en spa ial equency bands [16].
App oaches ha a e no s ic ly da a-d i en include
he combina ion o saliency wi h seman ic maps, ying
o ca ch a ac i eness o aces, pe sons, and o he ob-
jec s [17, 18], o he ad-hoc adap a ion o saliency mod-
els o di e en da ase s by lea ning weigh s ha op i-
mize he p edic ion o human ixa ions in hose da ase s
[19]. Also, he adap a ion o he spa io-ch oma ic ep-
esen a ion o he speci ic da ase om lea ning speci i-
P ep in submi ed o Image and Vision Compu ing Augus 30, 2011
cally deco ela ed coo dina es [20] o independen com-
ponen s [21] has been p oposed. These wo las ap-
p oaches al eady poin o bene i s om he adap a ion o
he ea u e basis. Howe e hese app oaches a e mos ly
ad-hoc, since hey ely on an o -line compu a ion ap-
plied o a speci ic da ase . They do no p oduce a ep e-
sen a ion adap ed o each speci ic image.
1.1. Ou app oach
A na u al app oxima ion o sample dis inc i eness
can be done by compu ing he s a is ical dis ance in a
ep esen a i e coo dina e sys em. This can be simply
done h ough a ec o no m compu a ion i such sys-
em is s a is ically whi ened. The eby, i makes sense
o hink in he adap a ion – h ough whi ening– o he
ea u e basis o he speci ic s a is ical s uc u e o a pa -
icula image, conside ing pixels as samples. The e-
sul ing ep esen a ion would yield a simple and s aigh
measu e o poin saliency h ough a ec o no m com-
pu a ion.
Howe e , ypical schemes o s a is ical whi ening
ha e a cubic o highe complexi y on he numbe o co-
o dina es, while linea on he numbe o samples. These
ac s p e en hei use wi h ep esen a ions o images in-
ol ing h ee colo componen s, se e al scales and se -
e al o ien a ions. In his pape we gene alize a p elim-
ina y app oach [22] o o e come his p oblem by im-
possing a whi ening ans o ma ion independen ly on
educed g oups o ea u e componen s.
The p oposed app oach g ounds on a classical hie a -
chical decomposi ion o images ha i s sepa a es ch o-
ma ic componen s, and nex pe o ms on each o hem
a mul iscale and mul io ien ed decomposi ion. Such ap-
p oach is coa sely inspi ed in he image ep esen a ion
desc ibed in ea ly s ages o he isual pa hway. Be-
sides, di e en implemen a ions and simpli ica ions o
he same can be ound in a a ie y o ea ly and e-
cen models o compu e ision wi h di e en pu poses.
The e o e, we p opose o apply on-line whi ening on
ch oma ic componen s in a i s s age. This ope a ion is
ollowed by a mul io ien a ion and mul iscale decompo-
si ion o he esul ing whi ened ch oma ic componen s.
Nex , u he whi ening is impossed o g oups o o i-
en ed scales o each whi ened ch oma ic componen .
This s a egy allows o keep he numbe o componen s
in ol ed in whi ening limi ed, o e coming p oblems o
compu a ional complexi y.
As a esul , an speci ically adap ed ep esen a ion o
he image a ises. The esul ing image componen s ha e
ze o mean and uni s o a iance. As well, hey a e pa ly
deco ela ed. To ob ain a saliency map, we simply com-
pu e poin dis inc i eness by aking, o each pixel, he
squa ed ec o no m in his ep esen a ion di ided by
he sum o he same ac oss all he pixels.
The p oposed model is alida ed and compa ed wi h
s a e-o - he-a app oaches by measu ing he p edic-
i e capabili y o human ixa ions in h ee open access
da ase s h ough s a e-o he-a p ocedu es based on
Recei e Ope a ing Cha ac e is ic (ROC) analysis and
Kullback-Leible di e gences (KLD). Addi ionally, he
model will be shown o ep oduce a wide se o ele an
psychophysical esul s in which o he models show ail-
u es.
The pape is o ganized as ollows. Sec ion 2 p o ides
a de ailed desc ip ion o he AWS model o saliency
compu a ion. Sec ion 3 e alua es he capabili y o he
model in p edic ing eye- ixa ions. Sec ion 4 shows he
abili y o ep oduce a selec ion o psychophysical and
pe cep ual obse a ions. Finally, in Sec ion 5 he main
conclusions o he wo k a e p esen ed.
2. Model
The key poin o he model o saliency p oposed e-
lies on he on-line adap a ion o he basis used o ep-
esen a ion o he speci ic s a is ical s uc u e o he im-
age. This implies a s ep beyond he adap a ion o a gi en
se –like he se o na u al images– ha is unde he
decomposi ion me hods o mos exis ing app oaches o
saliency. This adap a ion uses pixels as s a is ical sam-
ples and seeks o a se o deco ela ed and whi ened
coo dina es, able o deal wi h he in o ma ion p esen in
he image and o also p o ide a eliable es ima ion o he
s a is ical dis ance o each sample –pixel– o he cen e
o he dis ibu ion. The e o e, he p oposed model will
be e e ed o as he adap i e whi ening saliency (AWS)
model.
2.1. Ch oma ic decomposi ion and whi ening
Ch oma ic componen s unde go he i s adap a i e
s age. Each pixel in he image has an associa ed ec o
o ed ( ), g een (g) and blue (b) componen s. In gen-
e al, he ( ,g,b) coo dina es a e highly co ela ed in he
ensemble o samples. P o ided he co a iance ma ix in
hese coo dina es is:
C gb =









σ2
σ g σ b
σ g σ2
gσgb
σ b σgb σ2
b










(1)
we ypically ha e ha all he elemen s a e non-ze o and
non-negligible. Some colo spaces (e.g. he Lab model)
educe his co ela ion be ween componen s by p oduc-
ing a ep esen a ion ha is deco ela ed in he se o na -
u al images, bu no necessa ily in speci ic images.
2
To deco ela e colo in o ma ion, he whi ening p o-
cedu e consis ing in deco ela ion and a iance no mal-
iza ion (as desc ibed in he appendix) is simply applied
o he ,g,bcomponen s o he image. Being x1= ,
x2=gand x3=b he RGB coo dina es o any pixel in
he image, hey a e in ol ed in he ans o ma ion
(x1,x2,x3)→(z1,z2,z3) (2)
The eby, we ge a z=(zch
1,zch
2,zch
3) whi ened ep e-
sen a ion wi h a new ec o associa ed o each pixel. In
his ep esen a ion he co a iance ma ix is he inden i y
ma ix, and hus each coo dina e has uni s o a iance.
Indeed, he ec o no m gi es a measu e o ch oma ic
dis inc i eness as he s a is ical dis ance o each colo
poin o he a e age colo . Such a simple measu e is
equi alen o he explana ion p oposed by [23] o colo
sea ch asymme y phenomena epo ed o humans in
a se o simple syn he ic images. Howe e , o compu e
poin saliency in clu e ed na u al scenes, spa ial dis inc-
i eness also mus be aken in o accoun .
Al e na i ely o a RGB colo space, we ha e also
es ed he use o o he colo spaces like he Lab model.
The con e sion om RGB o a Lab model in ol es a
non-linea ans o ma ion. Besides, he Lab model p o-
duces a ep esen a ion ha p ese es, on a e age, pe -
cep ual dis ances. The e o e, di e ences in he esul -
ing deco ela ed componen s and e en an ad an age o
he Lab model may be expec ed. Howe e , we did no
ind a signi ican ad an age o any o he al e na i e
colo spaces o e RGB in ou expe imen al e alua ion.
The eby, he e was no appa en eason o ecode he
RGB images o o he colo space be o e whi ening.
2.2. O ien ed mul iscale decomposi ion and whi ening
We ep esen he spa ial s uc u e by decomposing
each o he whi ened ch oma ic componen s (i.e. zch
1,
zch
2, and zch
3) h ough a measu e o local ene gy a di -
e en spa ial equency bands cen e ed a di e en e-
quency modulus alues (scales) and di e en o ien a-
ions.
To ob ain local ene gy, we use a bank o log-Gabo
il e s, since hei eal and imagina y pa s in he spa ial
domain o m a pai o il e s in phase quad a u e. These
il e s p esen se e al ad an ages o e he Gabo il e s.
Namely, hey ha e a ze o DC componen and a long ail
owa ds high equencies, app oaching be e he ecep-
i e ields o co ical cells [24]. The exp ession o hese
il e s in he equency domain is gi en by:
log Gabo so (ρ, α)=exp 









−log (ρ/ρs)2
2log σρs/ρs2









·
exp −(α−αo)2
2(σαo)2!
(3)
being (ρ, α) he spa ial equency in pola coo dina es,
(ρs, αo) he cen al equency o he il e , s he scale
index, and o he o ien a ion index.
In he implemen a ion employed in his pape , ou
o ien a ions (0◦,45◦,90◦,135◦) a e used, se en scales
o he i s z-sco e ( oughly equi alen o luminance),
and only 5 scales o he emaining wo componen s.
This di e ence is jus i ied by he obse a ion ha he
ines and coa ses scales o hese componen s ba ely
showed any ele an in o ma ion. Acco dingly, while
he minimum wa eleng h o he i s z-sco e is 3 pix-
els, 6 pixels o colo ha e been used ins ead. The
use o o ien a ions in colo componen s has been ob-
se ed o imp o e pe o mance, compa ed o he use o
iso opic esponses. Besides, o ien a ion selec i i y o
ch oma ic mul iscale ecep i e ields has been shown o
ake place in V1 and is hough o in luence saliency
[25]. I has been also ied o include iso opic esponses
o luminance in addi ion o he o ien ed esponses, bu
he esul s we e p ac ically he same. Consequen ly,
hey we e conside ed edundan in he compu a ion o
saliency, and disca ded o he sake o e iciency.
The bank o il e s is applied on each o he whi ened
ch oma ic componen s p e iously ob ained. F om he
complex esponse o he il e in a gi en equency band
we compu e local ene gy as he modulus o he esponse
[26][27]. Tha is:
ecos =q(zc∗ os)2+(zc∗hos)2(4)
whe e index cdeno es a whi ened ch oma ic compo-
nen , zcis a e ino opic ep esen a ion o such compo-
nen , and and hdeno e espec i ely he e en symme -
ic log-Gabo gi ing he eal pa o he esponse, and
he odd symme ic log-Gabo gi ing he imagina y pa
o he esponse. They o m indeed a pai o il e s in
phase quad a u e.
The e o e, we ob ain a ep esen a ion o he image
in e ms o he local ene gy co esponding o di e en
whi ened colo componen s, di e en scales and o ien-
a ions.
The nex s ep deals wi h he adap a ion o his ep-
esen a ion ha al eady codes he spa ial s uc u e. To
do so, we ha e chosen o deco ela e and whi en, in-
dependen ly and in pa allel, each se o o ien ed scales
3
o each o he ch oma ic componen s. Fo a gi en
whi ened ch oma ic componen and o ien a ion, each
pixel has a local ene gy alue o each scale. The e o e,
each scale sican be iewed as an o iginal coo dina e
axis. The ensemble o scales de e mines a se o o ig-
inal axis in which each pixel is ep esen ed by a poin
wi h i s coo dina es de e mined by he local ene gy al-
ues o he co esponding scales. F om he ensemble o
samples (all he pixels), we can compu e he co a iance
ma ix in such scale coo dina es. I he numbe o scales
is Ms, hen he co a iance ma ix is a Ms×Msma ix.
Cco =











σ2
co;s1. . . σco;s1sMs
.
.
.....
.
.
σco;s1sMs . . . σ2
co;sMs












(5)
As well known, in na u al images di e en scales a e
highly co ela ed, which in gene al makes all he ma ix
elemen s non-ze o. The e o e, he whi ening p ocedu e
based on deco ela ion and a iance no maliza ion is ap-
plied again o achie e a new se o whi ened scale coo -
dina es zsc
i. The esul ing co a iance ma ix becomes
he iden i y. This new whi ened ea u e basis is com-
posed o axis ha a e shi ed, o a ed and escaled om
he o iginal scale axis. In sum, o each se o scales we
ha e ans o med he o iginal scale coo dina es o pix-
els o new coo dina es ha a e deco ela ed and wi h he
a iance as he no m.
2.3. Saliency
To compu e saliency, we simply use he sum o he
squa ed no m o he ec o s in he ob ained ep esen a-
ion as an es ima ion o pixel (i.e. sample) dis inc i e-
ness and we no malize i o he sum ac oss all he pixels.
Tha is, o each pixel i
kzicok2=zT
icozico (6)
whe e zico is he ec o associa ed o he pixel o a colo
componen cand an o ien a ion owi h as many compo-
nen s as whi ened scales (Ms).
This p o ides a e ino opic measu e o he local ea-
u e con as . In his way, a measu e o conspicui y is
ob ained o each o ien a ion o each o he colo com-
ponen s. The nex s eps in ol e a Gaussian smoo hing
and he addi ion o he maps co esponding o all o
he o ien a ions. Tha is, o a gi en colo componen
c=1...Mcand pixel i, he co esponding saliency (Sic)
is calcula ed:
Sic =
Mo
X
o=1
kzicok2(7)
Colo componen s unde go he same summa ion s ep
o ge a inal map o saliency. Addi ionally, o ease in e -
p e a ion o his map as p obabili y o ecei e a en ion,
i is no malized by he in eg al o he saliency in he im-
age domain (i.e. he o al popula ion ac i i y). Hence,
saliency o a pixel i(Si) is gi en by:
Si=PMc
c=1Sic
PN
i=1PMc
c=1Sic
(8)
The alues ob ained o he ensemble o poin s, a -
anged in a 2D ma ix deli e a map o saliency So he
same dimensions han he inpu image.
The igu e 1 shows a g aphic ou line o he model.
I mus be no iced ha o he app oaches o in eg a-
ion di e en om he squa ed no m ha e been explo ed
like he aw ec o no m, highe powe exponen s o he
ec o no m, and e en an exponen ial ans o ma ion o
he di e en componen s ollowed by a summa ion.The
use o he aw ec o no m achie ed close -bu in e io -
pe o mance in p edic ing ixa ions, while all he o he
app oaches beha ed much wo se. All he al e na i es
ailed in se e al o he psychophysical expe imen s de-
sc ibed in his pape . This obse a ions ag ee wi h he
iew ha saliency is ela ed o he classical s a is ical
dis ance o he ea u e ec o associa ed o a poin om
he cen e o he dis ibu ion o ea u es p esen in he
image. The use o he squa ed ec o no m in ou hie a -
chical whi ening app oach can be iewed as an e icien
es ima ion o an o e all T2o Ho elling in an o iginal
high-dimensional ep esen a ion.
Rega ding he compu a ional complexi y o his im-
plemen a ion, PCA implies a load ha linea ly g ows
wi h he numbe o pixels (N), and in a cubic man-
ne wi h he numbe o componen s (M), speci ically
O(M3+M2N). Since we ha e kep he numbe o com-
ponen s (colo componen s o scales) ixed and small,
he asymp o ic complexi y depends on he numbe o
pixels. This is de e mined by he use o he FFT in he
il e ing p ocess, which is O(Nlog(N)). Mos saliency
models ha e a complexi y which is O(N2) o highe .
3. Compa ison wi h human ixa ions
In he las yea s, he mos ex ended alida ion p oce-
du e o no el models o saliency has been he abili y o
p edic ixa ions eco ded om humans du ing he ee-
iewing o na u al images, wi hou any speci ic goal o
ask [10, 11, 13]. The e a e some al e na i e p ocedu es,
mo e di icul o in e p e s ic ly in e ms o saliency.
Fo example, objec segmen a ion o ask-d i en isual
4
Figu e 1: Adap i e whi ening saliency model.
sea ch. A e y ecen wo k analyses a numbe o mod-
els o saliency h ough he compa ison wi h human a -
ings in a ask o isual sea ch o mili a y ehicles in
pho og aphs, inding s a is ically signi ican co ela ion
o mos models [28]. Howe e , he da ase employed is
op-down biased and also has impo an ea u e biases
(mos ly open g een landscapes and sky). The a ge is
in mos cases non salien due o i s camou lage design,
o akes up a la ge po ion o he image –holding a num-
be o salien and non salien pa s. The eby, he impli-
ca ions o such signi icance in co ela ion esul s aise
impo an di icul ies o in e p e a ion. Is such co ela-
ion ela ed o a gene al es ima ion o saliency o a he
o an e icien de ec ion –o e en segmen a ion– o mil-
i a y ehicles in coun yside scenes?. These conce ns
ha e no been se ou o he es s based on he p edic-
ion o human ixa ions. This p ocedu e does no ely
on e lec i e decisions abou wha is conpicuous, bu e-
lies on a as ac ion when aced o an image. Mo eo e ,
i is p ecisely ela ed o posi ions in he space, no o
an objec o unde e mined and changeable a ea on he
image.
The e o e, a majo goal in he modeling o saliency
is pushing his benchma k u he on.
3.1. Da ase s and models
Th ee open-access eye- acking da ase s o na u al
images ha e been used. In he h ee da ase s he sub-
jec s did no ecei e any speci ic ins uc ion. Tha is,
hey mee he equi emen o ee- iewing. The images
ha e been shown in a amdom o de o each subjec .
The igu e 2 shows h ee example images om each o
he da ase s.
The i s da ase has been published by B uce and
Tso sos and has 120 images and ixa ions om 20 sub-
jec s [29]. Each image has been iewed du ing 4 sec-
onds. I has al eady been used o alida e many s a e-
o - he-a models o bo om-up saliency using di e en
p ocedu es [10, 11, 13]. The e o e, i p o ides a sui -
able e e ence o a ai assessmen o a no el model in
ela ion o exis ing app oaches.
The second da ase has been published by Koo s a
e al. and consis s o 99 images and he co esponding
ixa ions o 31 subjec s [30]. The iewing ime was 5
seconds o each image. One in e es ing p ope y o his
da ase is ha i is o ganized in i e di e en g oups o
images (12 images o animals, 12 o s ee s, 16 o build-
ings, 40 o na u e, and 19 o lowe s o na u al symme-
ies). This ea u e may be expec ed o e eal possible
biases in he models. Besides, i may help o analyze
he causes o a iabili y unde he same expe imen al
condi ions.
Finally, he NUSEF da ase has been chosen because
i is supposed o ha e a s ong emo ional bu den [31].
This a ec i e con en may be expec ed o p oduce an
inc eased human consis ency no ela ed o low le el
ea u es, bu o emo ions ela ed o abs ac concep s
s ongly sugges ed by he images. The e o e, models
o saliency may be expec ed o explain less amoun o
in e subjec consis ency han in he o he da ase s. I is
composed o 758 images obse ed on a e age by 25.3
subjec s. The iewing ime was 5 seconds o each im-
5

Figu e 2: Examples o images om he h ee da ase s used. B uce and Tso sos (le ); Ko s a e al. (cen e ); NUSEF ( igh )
age.
O he wise, we compa e he esul s o he p oposed
model wi h o he 4 models. Namely: The model o Seo
and Milan a based on sel esemblance [13]; he SUN
model [11] ha adop s a bayessian app oach based on
p e iously lea ned image s a is ics; he AIM model ha
decomposes he image wi h independen componen s o
na u al images and uses sel -in o ma ion as a measu e
o dis inc i eness [10]; and inally he classic model o
saliency p oposed by I i e al. [8].
3.2. ROC analysis and KL di e gence
To assess he use ulness o he saliency maps o dis-
c imina e be ween ixa ed and non ixa ed poin s, we
ha e used he a ea unde he cu e (AUC), ob ained
om a ecei e ope a ing cha ac e is ic (ROC) analysis,
and a Kullback-Leible di e gence (KLD) compa ison.
Bo h me hods a ge he capabili y o he saliency maps
o p edic he spa ial dis ibu ion o ixa ions h ough
he compa ison o he dis ibu ions o saliency in ixa ed
e sus non- ixa ed poin s.
To a oid cen e -bias, in each image, only poin s ix-
a ed in ano he image om he same da ase a e used
as non ixa ed poin s. As sugges ed in [32], s anda d e -
o is compu ed h ough a boo s ap echnique, shu ling
he o he images used o ake he non ixa ed poin s,
exac ly like in [11] and in [13]. This las s ep should
no be adop ed i he goal is o assess a combina ion
o saliency and a cen e -bias model. Howe e , since
saliency is da a-d i en and he cen e -bias is a spa ial
bias (wo king ega dless o he speci ic da a), hey a e
di e en mechanisms. Thus, i makes sense o use di -
e en e alua ions ocusing on each o he componen s.
Fu he mo e, he use o he boo s apping me hod
yields a high sensi i i y. As ecen ly shown in [19], a
ROC analysis as used by many au ho s (wi hou a boo -
s apping o p e en he in luence o cen e -bias) aises
p oblems o sensi i i y. Howe e , he use o he boo -
6
s apping p ocedu e yields a s anda d e o ha is ipi-
cally below he 3% o he dynamic ange spanned by
he ob ained alues o he di e en models o bo h he
ROC analysis and he KLD compa ison.
Ne e heless, a double assessmen h ough a ROC
analysis and a KLD compa ison is p o ided o ensu e
he eliabili y o he e alua ion.
3.3. Resul s
The able 1 ga he s he esul s ob ained on he h ee
da ase s. The igu es 3 o 5 p o ide examples o saliency
maps o he bes pe o ming models ha ocus on spe-
ci ic ea u es o suppo he discussion.
The alues shown o he model o I i e al. [8] on he
da ase o B uce and Tso sos a e highe han epo ed
in p e ious wo ks [11, 13] because, ins ead o using
hei saliency oolbox, he o iginal implemen a ion has
been used, as made a ailable o Ma lab (h p://www.
klab.cal ech.edu/~ha el/sha e/gb s.php). Fo
he o he models in his da ase he alues ha we ha e
ob ained a e compa ible wi h hose published in [11]
and [13], hus we ha e espec ed he epo ed alues.
3.3.1. Discussion
Fi s ly, i is wo h no ing ha he esul s wi h bo h
measu es, ROC analysis and KLD, yield an equi alen
anking and equi alen dis ances be ween models on
each o he da ase s. The e a e mino di e ences in
wo g oups o he da ase o Koo s a bu hey do no
gi e ise o any ema kable di e ence in he e alua ion.
The e o e, he commen s ha ollow hold o bo h. Re-
ga ding he sensi i i y o he measu es, he s anda d e -
o emains below he 3% o he spanned ange o al-
ues.
Conside ing he h ee da ase s, he anking o mod-
els yields only a single change o posi ions be ween he
model o Seo and Milan a and he AIM model in he
NUSEF da ase . Al hough wi h a iable dis ances, he
es o models ank he same posi ion. The AWS model
holds clea ly he i s posi ion in he h ee da ase s.
Howe e , looking a he i e g oups o he da ase o
Koo s a e al., se e al changes o posi ion in ol ing di -
e en models occu . As a esul , he e is no mo e a
clea cohe en anking. E en hough, he AWS main-
ains he bes pe o mance wi h a dis ance on he nex
a beyond he s anda d e o , excep o he buildings
g oup in which he model o Seo and Milan a achie es
a sligh ly be e esul , bu wi hin he unce ain y limi s
es ablished.
O he wise, he a ia ions in pe o mance ac oss he
da ase s may be used o look o biases in he models.
The s ong ad an age o AWS in he g oup o lowe s
and na u al symme ies inds an explana ion in he ex-
amples shown in he igu e 3. The model o Seo and Mi-
lan a and he AIM model miss comple ely he saliency
o na u al symme ies ha ca ch a conside able amoun
o ixa ions in hese images. In con as , he AWS model
manages o cap u e he saliency o symme ies in na u-
al scenes.
Besides, o he ac o s appea o con ibu e o he ad-
an age o he AWS model. I shows a sensi i i y o
salien high equency pa e ns like he s iped pa e n
on he small head o a bu e ly ha is shown in he
igu e 4. Fo he model o Seo and Milan a , he head
seems o be jus ano he pa o an edge a ound he bu -
e ly. Addi ionally, he beha io o ou model when
aced o colo single ons appea s o be mo e obus .
The igu e 5 shows a e ealing example. The yellow
and ed peppe s a e among he mos salien objec s o
bo h he AWS model and humans. In con as , he AIM
model and pa icula ly he model o Seo and Milan a
ind mo e salien he objec s in he uppe pa o he im-
age, hus showing a lack o sensi i i y o colo pop-ou
in his na u al con ex .
3.4. Compa ison wi h human p io i y
Some ques ions in ela ion o he assessmen p oce-
du e a ise. Is i sui able he s a is ical signi icance o
compa e he models o is i oo igh in p ac ice?. I
may occu ha di e ences in model pe o mance a e
simila o he a iabili y shown by humans hemsel es,
while being s a is ically signi ican . O he wise, wha a e
he easons o he high a ia ion in he absolu e al-
ues ac oss he da ase s? Is he e any means o c ea e
a da ase ha p o ides eliable and de ini i e esul s in
anking he models?.
I is clea ha he explana ion is no a di e en expe -
imen al se up since he la ges a ia ion is ound ac oss
he g oups o he Koo s a da ase , all unde he same
se up. Any explana ion should be ela ed o di e ences
in he image con en , di e ences in he associa ed hu-
man beha io , and di e en biases in he models o
saliency.
In o de o explo e answe s o he aised ques ions,
we p opose o compa e he pe o mance o models wi h
he pe o mance o single subjec s. To assess he p edic-
i e pe o mance o a single subjec we eso o p io i y
maps de i ed om he ixa ions o he subjec on each
image. F om his measu e, we may compu e an es ima-
ion o he a e age subjec pe o mance and an es ima-
ion o human pe o mance a iabili y.
7
Table 1: AUC alues ob ained wi h di e en models o saliency o bo h o he da ase s o B uce and Tso sos and Koo s a e al. S anda d e o s,
ob ained like in [11], ange 0.0004-0.0008. Fo he g oups o he Koo s a e al. da ase , s anda d e o s ange 0.0010-0.0018. (* Resul s epo ed
by [11]; ** Resul s epo ed by he au ho s).
Model B uce and
Tso sos
da ase
NUSEF
da ase
Koo s a e al. da ase
Whole
da ase
Buildings Na u e Animals Flowe s S ee
AWS 0.7106 0.6035 0.6205 0.6105 0.5815 0.6565 0.6374 0.7020
Seo and Mil. 0.6896** 0.5802 0.5933 0.6136 0.5530 0.6445 0.5602 0.6907
AIM 0.6727* 0.5902 0.5842 0.5766 0.5628 0.5953 0.5881 0.6393
SUN 0.6682* 0.5782 0.5705 0.5514 0.5484 0.5401 0.6100 0.6458
I i e al. 0.6456 0.5655 0.5702 0.5814 0.5478 0.6200 0.5217 0.6509
Gao e al. 0.6395* – – – – – –
Table 2: KL di e genge alues ob ained wi h di e en models o saliency o bo h o he da ase s o B uce and Tso sos and Koo s a e al. S anda d
e o s, ob ained like in [11], ange 0.001-0.002 o bo h o he da ase s and he g oups. (* Resul s epo ed by [11]; ** Resul s epo ed by he
au ho s).
Model B uce and
Tso sos
da ase
NUSEF
da ase
Koo s a e al. da ase
Whole
da ase
Buildings Na u e Animals Flowe s S ee
AWS 0.321 0.071 0.099 0.109 0.058 0.188 0.142 0.307
Seo and Mil. 0.278** 0.048 0.071 0.110 0.049 0.175 0.057 0.281
AIM 0.203* 0.055 0.055 0.070 0.045 0.105 0.085 0.197
SUN 0.210* 0.043 0.039 0.046 0.033 0.049 0.097 0.173
I i e al. 0.175 0.033 0.038 0.069 0.032 0.109 0.025 0.200
3.4.1. Human p io i y pe o mance
To implemen his measu e, p io i y maps de i ed
om ixa ions ha e been used, ollowing he me hod
desc ibed by [30]. This me hod lies in he sub ac ion
o he dis ance be ween each poin and i s nea es ixa-
ion om he maximum possible dis ance in he image.
As a esul , ixa ed poin s ha e he maximum alue and
non ixa ed poin s ha e a alue ha dec eases linea ly
wi h he dis ance o he nea es ixa ion. The esul ing
maps can be used as p obabili y dis ibu ions o sub-
jec s ixa ions (p io i y maps), and can be conside ed as
subjec i e measu es o saliency. A leas wi h ew ix-
a ions pe subjec , as i is he case, his me hod yields
be e p edic i e esul s han he app oach o compu e
p io i y maps based on il e ing o ixa ions wi h Gaus-
sians ke nels [29]. This las app oach ipically assigns
ze o o decimal p io i y o poin s beyond 2.3◦o isual
angle, since usually he wid h o he Gaussian is ixed
o 1◦and he ampli ude is ixed o 255, which is he
maximum o he dynamic ange used in he ROC anal-
ysis. Consequen ly, wi h ew ixa ions by subjec , he
p io i y maps p esen alues below 1 o posi ions ha
can be close o ixa ions. The e o e, in a ROC analysis
all hese loca ions a e equally conside ed ze o p io i y
poin s. Exac ly he same as poin s much u he om
any ixa ion.
Fu he mo e, he linea dis ance-based me hod is pa-
ame e ee. The eby, we do no need o make assump-
ions on he ange o poin s ha may ha e in luenced
a gi en ixa ion. O cou se, i can be a gued ha i is
no jus i ied o assume ha p io i y d ops linea ly wi h
dis ance o ixa ions. Ne e heless, i seems ac ually
easonable o assume ha p io i y d ops mono onically
wi h dis ance o he nea es ixa ion. I he me hod o
compa e and e alua e maps is in a ian o mono onic
ans o ma ions, as ROC analysis is, hen he e is no
issue wi h using linea , o any o he mono onic maps.
Hence, h ough a ROC analysis, he same one employed
o e alua e models o saliency, he capabili y o hese
maps o p edic he ixa ions o he se o subjec s can
be assessed, wi hou conce ns on he dynamic ange o
he ROC analysis. I mus be no iced ha we ha e no
used he a e aged p io i y maps shown in he igu es
3 o 5, bu maps compu ed speci ically o each o he
subjec s using he p ocedu e desc ibed abo e.
The p e ious e alua ion o each subjec has been
done, only o hose wi h ixa ions o all o he im-
ages. One indi idual has been excluded o he da ase
o B uce and Tso sos, whose de ia ion om he a e -
age o humans was la ge han wice he s anda d de ia-
ion, and who also had jus one ixa ion in many images.
This yields p io i y maps om 9 subjec s o he da ase
o B uce and Tso sos, and 25 subjec s o he da ase o
Koo s a. On he NUSEF da ase ou app oach o de i e
8
Figu e 3: Examples o esul s wi h 4 images wi h a dominan symme ic poin . Fo compa ison, human p io i y maps p o ided by he au ho s a e
shown. B uce and Tso sos de i ed p io i y om ixa ions using Gaussian ke nels [29], while Koo s a e al. used a dis ance- o- ixa ion ans o m
[30]. This explains he no iceable di e ences in he dynamic ange o he p io i y maps.
subjec p io i y maps inds a p oblem: no subjec has
obse ed all he images. Ne e heless, all he images
ha e been obse ed by a leas 13 subjec s. The e o e,
we ha e buil 13 pseudo-subjec s ga he ing he p io i y
maps o he i s 13 obse e s ha iewed each o he
images, ollowing he o de o subjec s p o ided by he
au ho s.
The pe o mance o he p io i y maps associa ed o a
gi en subjec was ob ained h ough he assessmen wi h
ixa ions o o he subjec s in he da ase . Compu ing he
a e age, we ha e he a e age pe o mance o p io i y
maps associa ed o di e en subjec s. Besides, he dou-
ble o he s anda d de ia ion p o ides an es ima ion o
he ange o p edic i e pe o mance o he 95% o hu-
mans, unde he assump ion o a no mal dis ibu ion o
AUC p io i y alues. This was ue o he da ase s and
g oups s udied, wi h a ku osis alue e y close o 3.
Mo eo e , his in e al o a iabili y be ween subjec s
can be also used as a measu e o he minimum ele an
dis ance be ween wo models. Di e ences lowe han
such a iabili y may be ega ded as p oducing no p ac-
ical e ec on pe o mance. The esul s a e gi en in he
able 3.
3.4.2. Saliency e sus p io i y
A a i s look, we can see how he pe o mance
o p io i y a ies ac oss da ase s simila ly o saliency
maps. The a iabili y o subjec pe o mance in a gi en
da ase is well an o de o magni ude highe han he
s a is ical signi icance o he measu e o pe o mance,
excep o he NUSEF ha shows much less a iabili y.
Rema kably, he pe o mance o he AWS model is
compa ible wi h he es ima ed human p io i y pe o -
mance o he h ee da ase s and all he g oups. The
model o Seo and Milan a is also compa ible wi h he
a e age human o he da ase o B uce and Tso sos, and
o wo o he i e g oups o Koo s a e al.. Howe e
his compa ibili y does no hold o ei he he whole
da ase o Koo s a e al. o he NUSEF da ase . The
model by B uce and Tso sos is only ma ginally com-
pa ible wi h he a e age human wi h hei own da ase .
9
Figu e 12: Examples o saliency-based segmen a ion: o iginal image (le s), saliency maps (cen e ), and p o o-objec s ( ig h) a aganged in wo
e ical blocks. Six o he images ha e been ob ained om [12], he es a e ou s.
Tha is,
x=xj→y=yj→z=zj(A.1)
wi h j=1...M, whe e Mis he numbe o componen s.
The whi ening p ocedu e can be summa ized in wo
s eps. Fi s , as well known, p incipal componen s esul
om diagonaliza ion o he co a iance ma ix, o de ing
eigen alues (lj) om highe o lowe . To compu e he
co a iance ma ix he e a e Nsamples, as many as he
numbe o pixels in he inpu image. The whi ened z
ep esen a ion is hen ob ained h ough no maliza ion
by a iance, gi en by he eigen alues. This means ha
o each p incipal componen :
zj=yj
plj
;j∈[1,M] (A.2)
These z-sco es yield a whi ened ep esen a ion, wi h
he co a iance ma ix being he uni y ma ix. The
squa ed no m o a ec o in hese coo dina es is in ac
he s a is ical dis ance in he o iginal xcoo dina es.
Appendix B. Rep oducibili y
A Ma lab p-code ile o ep oduce he expe imen-
al esul s epo ed in his pape as well as all he
saliency maps compu ed wi h he AWS model o
he h ee eye- acking da ase s a e a ailable on he
web page h p://www-g a.dec.usc.es/pe soal/
xose. idal/ esea ch/aws/AWSmodel.h ml.
Re e ences
[1] Z. Liu, H. Yan, L. Shen, K. N. Ngan, Z. Zhang, Adap i e im-
age e a ge ing using saliency-based con inuous seam ca ing,
Op ical Enginee ing 49 (2010) 1–10.
[2] J. Ruesch, M. Lopes, A. Be na dino, J. Ho ns ein, J. San os-
Vic o , R. P ei e , Mul imodal saliency-based bo om-up a en-
ion a amewo k o he humanoid obo icub, in: In . Con . on
Robo ics and Au oma ion (ICRA), pp. 962–967.
16

[3] C. Kanan, G. Co ell, Robus classi ica ion o objec s, aces,
and lowe s using na u al image s a is ics, in: IEEE in . Con .
on Compu e Vision and Pa e n Recogni ion (CVPR).
[4] J. Ha el, C. Koch, On he op imali y o spa ial a en ion o
objec de ec ion, in: A en ion in Cogni i e Sys ems 2009, pp.
1–14.
[5] D. Gao, S. Han, N. Vasconcelos, Disc iminan saliency, he de-
ec ion o suspicious coincidences, and applica ions o isual
ecogni ion, IEEE T ansac ions on Pa e n Analysis and Ma-
chine In elligence 31 (2009) 989.
[6] B. Alexe, T. Deselae s, V. Fe a i, Wha is an objec ?, in: IEEE
Con . on Compu e Vision and Pa e n Recogni ion (CVPR), pp.
73–80.
[7] N. Pa ikh, L. I i, J. Weiland, Saliency-based image p ocessing
o e inal p os heses, Jou nal o Neu al Enginee ing 7 (2010)
016006.
[8] L. I i, C. Koch, E. Niebu , A model o saliency-based isual
a en ion o apid scene analysis, IEEE T ansac ions on Pa e n
Analysis and Machine In elligence 20 (1998) 1254–1259.
[9] O. Le Meu , P. Le Calle , D. Ba ba, D. Tho eau, A cohe -
en compu a ional app oach o model bo om-up isual a en-
ion, IEEE T ansac ions on Pa e n Analysis and Machine In el-
ligence 28 (2006) 802–817.
[10] N. D. B uce, J. K. Tso sos, Saliency, a en ion, and isual
sea ch: An in o ma ion heo e ic app oach, Jou nal o Vision
9 (2009) 5.
[11] L. Zhang, M. H. Tong, T. K. Ma ks, H. Shan, G. W. Co ell,
SUN: a bayesian amewo k o saliency using na u al s a is ics,
Jou nal o Vision 8 (2008) 32.
[12] X. Hou, L. Zhang, Dynamic isual a en ion: Sea ching o
coding leng h inc emen s, in: Ad ances in Neu al In o ma ion
P ocessing Sys ems (NIPS), olume 21, pp. 681–688.
[13] H. J. Seo, P. Milan a , S a ic and space- ime isual saliency
de ec ion by sel - esemblance, Jou nal o Vision 9 (2009) 12–
15.
[14] X. Hou, L. Zhang, Thumbnail gene a ion based on global
saliency, in: Ad ances in Cogni i e Neu odynamics (ICCN),
pp. 999–1003.
[15] C. Guo, Q. Ma, L. Zhang, Spa io- empo al saliency de ec ion
using phase spec um o qua e nion ou ie ans o m, in: IEEE
Con on Compu e Vision and Pa e n Recogni ion (CVPR).
[16] R. Achan a, S. Hemami, F. Es ada, S. Sss unk, F equency-
uned salien egion de ec ion, in: IEEE Con . on Compu e
Vision and Pa e n Recogni ion (CVPR).
[17] M. Ce , E. P. F ady, C. Koch, Faces and ex a ac gaze in-
dependen o he ask: Expe imen al da a and compu e model,
Jou nal o ision 9 (2009).
[18] T. Judd, K. Ehinge , F. Du and, A. To alba, Lea ning o p edic
whe e humans look, in: IEEE 12 h In .l Con . on Compu e
Vision, IEEE, pp. 2106–2113.
[19] Q. Zhao, C. Koch, Lea ning a saliency map using ixa ed loca-
ions in na u al scenes, Jou nal o ision 11 (2011).
[20] J. an de Weije , T. Ge e s, A. D. Bagdano , Boos ing colo
saliency in image ea u e de ec ion, IEEE T ansac ions on Pa -
e n Analysis and Machine In elligence 28 (2006) 150–156.
[21] N. D. B uce, P. Ko np obs , On he ole o con ex in p obabilis-
ic models o isual saliency, in: IEEE In e na ional con e ence
on image p ocessing (ICIP), p. 30893092.
[22] A. Ga cia-Diaz, X. Fdez-Vidal, X. Pa do, R. Dosil, Deco ela-
ion and dis inc i eness p o ide wi h human-like saliency, in:
Ad anced Concep s o In elligen Vision Sys ems, pp. 343–
354.
[23] R. Rosenhol z, A. L. Nagy, N. R. Bell, The e ec o backg ound
colo on asymme ies in colo sea ch, Jou nal o Vision 4 (2004)
224–240.
[24] D. J. Field, Rela ions be ween he s a is ics o na u al images
and he esponse p ope ies o co ical cells, Jou nal o he Op-
ical Socie y o Ame ica A 4 (1987) 2379–2394.
[25] L. Zhaoping, R. J. Snowden, A heo y o a saliency map in
p ima y isual co ex (V1) es ed by psychophysics o colou o -
ien a ion in e e ence in ex u e segmen a ion, Visual Cogni ion
14 (2006) 911–933.
[26] P. Ko esi, In a ian measu es o image ea u es om phase in-
o ma ion, Ph.D. hesis, Depa men o Psychology, Uni e si y
o Wes e n Aus alia, 1996.
[27] M. C. Mo one, D. C. Bu , Fea u e de ec ion in human ision:
A phase-dependen ene gy model 1998, in: P oc. o he Royal
Socie y o London. Se ies B, Biological Sciences, pp. 221–245.
[28] A. Toe , Compu a ional e sus psychophysical image saliency:
A compa a i e e alua ion s udy, IEEE T ansac ions on Pa e n
Analysis and Machine In elligence (p ep in ) (2011).
[29] N. B uce, J. Tso sos, Saliency based on in o ma ion maximiza-
ion, in: Ad ances in Neu al In o ma ion P ocessing Sys ems
(NIPS), olume 18, p. 155.
[30] G. Koo s a, A. Nede een, B. de Boe , Paying a en ion o
symme y, in: P oc. o he B i ish Machine Vision Con e ence
(BMVC), pp. 1115–1125.
[31] S. Ramana han, H. Ka i, N. Sebe, M. Kankanhalli, T. S. Chua,
An eye ixa ion da abase o saliency de ec ion in images, in:
Eu opean Con . on Compu e Vision (ECCV), pp. 30–43.
[32] B. W. Ta le , R. J. Baddeley, I. D. Gilch is , Visual co ela es o
ixa ion selec ion: E ec s o scale and ime, Vision Resea ch 45
(2005) 643–659.
[33] W. Einhuse , M. Spain, P. Pe ona, Objec s p edic ixa ions be -
e han ea ly saliency, Jou nal o Vision 8 (2008) 18.
[34] E. Bi mingham, W. F. Bischo , A. Kings one, Saliency does
no accoun o ixa ions o eyes wi hin social scenes, Vision
Resea ch 49 (2009) 2992–3000.
[35] H. C. No hdu , The conspicuousness o o ien a ion and mo ion
con as , Spa ial Vision 7 (1993) 341–363.
[36] D. Gao, V. Mahade an, N. Vasconcelos, On he plausibili y o
he disc iminan cen e -su ound hypo hesis o isual saliency,
Jou nal o Vision 8 (2008) 13.
[37] A. T eisman, S. Go mican, Fea u e analysis in ea ly ision:
E idence om sea ch asymme ies, Psychological Re iew 95
(1988) 15–48.
[38] J. M. Wol e, T. S. Ho owi z, Wha a ibu es guide he deploy-
men o isual a en ion and how do hey do i ?, Na u e Re iews
Neu oscience 5 (2004) 495–501.
[39] A. Hy inen, E. Oja, A as ixed-poin algo i hm o indepen-
den componen analysis, Neu al Compu a ion 9 (1997) 1483–
1492.
[40] J. F. Ca doso, A. Souloumiac, T. Pa is, Blind beam o ming o
non-Gaussian signals, in: IEE P oceedings on Rada and Signal
P ocessing 1993, olume 140, pp. 362–370.
17