Decision-Theo e ic Planning wi h Pe son
T ajec o y P edic ion o Social Na iga ion
Ignacio Pé ez-Hu ado, Jesús Capi án, Fe nando Caballe o and Luis Me ino
Abs ac Robo s na iga ing in a social way should eason abou people in en ions
when ac ing. Fo ins ance, in applica ions like obo guidance o mee ing wi h a
pe son, he obo has o conside he goals o he people. In en ions a e inhe en ly non-
obse able, and hus we p opose Pa ially Obse able Ma ko Decision P ocesses
(POMDPs) as a decision-making ool o hese applica ions. One o he issues wi h
POMDPs is ha he p edic ion models a e usually handc a ed. In his pape , we use
machine lea ning echniques o build p edic ion models om obse a ions. A no el
echnique is employed o disco e poin s o in e es (goals) in he en i onmen , and a
a ian o G owing Hidden Ma ko Models (GHMMs) is used o lea n he ansi ion
p obabili ies o he POMDP. The app oach is applied o an au onomous elep esence
obo .
Keywo ds Ma ko decision p ocesses ·Social obo na iga ion ·GHMM
1 In oduc ion
Social obo s a e becoming a s ong end in he las yea s. I is clea ha u u e
obo ic applica ions will equi e obo s o coexis wi h human beings. In scena ios
popula ed wi h people, obo s should beha e in a social manne [2]. This includes
no only conside ing humans in a di e en way as o he "obs acles", bu also eason-
ing abou people in en ions and eac ing o hem. Fo ins ance, in applica ions like
obo guidance o mee ing wi h a pe son, obo s ha e o conside people’s goals and
commi men in o de o ac ua e in ad ance [6]. Mo eo e , obo s may wan o a oid
ce ain places when humans in end o go no o dis u b hem [14].
I. Pé ez-Hu ado(B)·L. Me ino
Pablo de Ola ide Uni e si y Se ille, Se ille, Spain
e-mail: [email p o ec ed]
J. Capi án ·F. Caballe o
Uni e si y o Se ille, Se ille, Spain
In his pape , we conside he applica ion o elep esence obo s in scena ios like
mee ing a pe son a a pa icula place. The objec i e is o inc emen he social in-
elligence and au onomy o he elep esence obo , so ha he obo can execu e he
low-le el na iga ion asks and he use can concen a e in he in e ac ion wi h his/he
pee s. This is ele an , as i has been obse ed ha one o he p oblems o elep esence
sys ems is he cogni i e o e load ha a ises by ha ing o ake low-le el (na iga ion
commands) and high-le el decisions (in e ac ion) a he same ime. This may lead o
mis akes a low le el and o gi e less a en ion o he high-le el asks [15].
These scena ios a e unce ain by na u e: obo ac ions a e no always de e minis ic;
he en i onmen may be dynamic; and senso s a e noisy. Fu he mo e, he main
sou ce o unce ain y comes om people in en ions, which a e non-obse able and
should be modeled p obabilis ically. Fo his eason, Pa ially Obse able Ma ko
Decision P ocesses (POMDPs) a e p oposed in his pape o decision making in hese
se ups. POMDPs p o ide a sound ma hema ical amewo k o decision- heo e ic
p oblems in unce ain domains, and ha e al eady been used o social applica ions
[2,3,6]. E en hough hey ha e adi ionally aced scalabili y issues and can be
compu a ionally cos ly, ecen ad ances in online [12,13] and o line [8] sol e s a e
making POMDPs inc easingly p ac ical o obo planning in la ge domains.
POMDPs use p edic ion models o in e he s a e. Fo ins ance, people mo ion
models a e equi ed in mos social asks. Howe e , in mos wo ks, basic o hand-
c a ed models o people mo emen a e used [2,6]. Howe e , mo ion pa e ns and
people in en ions a e a ec ed by poin s o in e es in he en i onmen , and ollow
epe i i e pa e ns: people mo e be ween doo s and co ido s ollowing common
ajec o ies; places o in e es a e common goals o a ec people beha io , a ac ing
hem (e.g., ending machines) o epelling hem (e.g., g ass lawns); e c. The e o e,
machine lea ning echniques a e conside ed in he li e a u e o es ima e such in e -
es ing poin s and lea n models o people in en ions [3,17].
In his pape , we con ibu e by using a POMDP model o a social ask o a obo
mee ing wi h a pe son, whe e human mo ion and in en ions a e modeled au oma -
ically by an ex ension o G owing Hidden Ma ko Models (GHMMs) [16]. This
GHMM model is i s lea n om da a obse ed by he obo and hen in eg a ed
wi hin he POMDP in o de o p edic he pe son in en ions o goals. La e , an o line
sol e is used o compu e an app oxima e op imal policy o he obo . Du ing he
execu ion phase, as he obo in e ac wi h he pe son, a p obabili y dis ibu ion o e
he pe son posi ion and in en ion is es ima ed using he lea n GHMM. Tha belie
is also used o eed he POMDP policy ha selec s bes ac ions o he obo in o de
o mee he pe son o a oid him/he depending on his/he goal.
We show esul s o p o e he easibili y o he me hod in an indoo scena io whe e
a elep esence obo [11] has o na iga e au onomously a ound mee ing people a
ce ain poin s o in e es (e.g., co ee machine) and no bo he ing hem a o he s (e.g.,
oile s).
The emainde o he pape is as ollows: Sec ion 2de ines he p oblem as a
social ask o obo na iga ion; Sec ion 3desc ibes ou app oach and models o
decision making; Sec ion 4p o ides expe imen al esul s; and Sec ion 5discusses
conclusions and u u e wo k.
Fig. 1 The elep esence obo is loca ed in a mee ing a ea, and he ask consis s o app oaching
people ha go o ypical in e ac ion poin s (like co ee machines). These poin s o in e es a e lea n
p e iously om da a.
2 P oblem De ini ion
In his wo k, we ocus on a social ask whe e a elep esence obo needs o in e ac
wi h people in he en i onmen . The main objec i e o a elep esence obo is o
ac as an a a a o a emo e use , ca ying a ideo-con e ence sys em on boa d
and allowing ha emo e use o sense and in e ac wi h he en i onmen om he
dis ance.
As commen ed, he objec i e is ha he obo ca ies ou he low-le el na iga ion
asks, allowing he use o concen a e on he in e ac ion h ough he obo . In pa ic-
ula , he ask we conside he e allows he use o connec o he obo , which ope a es
au onomously in a ce ain a ea ( o ins ance, a mee ing oom o co ee a ea), wai ing
o people o appea (see Fig. 1). Then, he obo au oma ically should go and ca ch
he pe son a hei des ina ion so ha he emo e use can es ablish a con e sa ion.
Fo ha , he obo mus eason abou he possible in en ions o he pe son and dis in-
guish be ween wo ypes o des ina ions: adequa e and inadequa e spo s o ha ing
a con e sa ion. Fo example, he e a e some places whe e people go o in e ac wi h
o he s, like a co ee machine o a es a ea. Howe e , he obo should no dis u b
people when hey in end o go o he oile o exi he a ea.
The obo can de ec and ack people nea by, bu i s only a ailable in o ma ion
o he ask is a map o he scena io. The in en ion o each pe son en e ing in he
ope a ional a ea is no obse able, so he obo needs o plan whe e o go aking
in o accoun unce ain ies in people’s posi ions and in en ions. Mo eo e , we wan
he obo o disco e au oma ically which a e he loca ions o he ho spo s o he
scena io, whe e people in end o go, ei he o in e ac wi h o he s o o lea e he
scena io.
We p opose a POMDP o model and sol e his social ask, since i allows he
obo o deal wi h he unce ain ies associa ed wi h i s senso s, ac ions and he people
a ound in a compac manne . POMDPs a e also adequa e o di e en mul i-objec i e
p oblems i hei ewa d and cos unc ions a e designed p ope ly. Fu he mo e,
we aim o lea n he spa ial s uc u e o human ajec o ies in a speci ic en i onmen
by building an Ins an aneous Topological Map (ITM) ha can be iewed as a dynamic
occupancy g id map. A Hidden Ma ko Model (HMM) is hen buil o e his ITM
and used as he ansi ion model o he o me POMDP. The de ails a e desc ibed
in he nex sec ion.
3 A POMDP o Social Na iga ion
The p oblem in Sec ion 2can be modeled as a POMDP whe e he obo main ains
a belie o e a pe son de ec ed and i s in en ions. He e, we assume ha he obo
can only eason abou a pe son a once, so when he e a e mul iple people a ound, i
only ocuses on a pa icula one. We use a GHMM o model pe son mo ion pa e ns
and disco e au oma ically he poin s o in e es in he scena io by obse ing people
a ound.
3.1 POMDP P elimina ies
Fo mally, a disc e e POMDP is de ined by he uple S,A,Z,T,O,R,h,γ[5].
–Thes a e space is he ini e se o possible s a es s∈S, o ins ance obo and
people poses.
–Theac ion space is de ined as he ini e se o possible ac ions ha he obo can
ake, a∈A.
–Theobse a ion space consis s o he ini e se o possible obse a ions z∈Z
om he onboa d senso s.
– A e pe o ming an ac ion a, he s a e ansi ion is modeled by he condi ional
p obabili y unc ion T(s,a,s)=p(s|a,s), which indica es he p obabili y o
eaching s a e si ac ion ais pe o med a s a e s.
– The obse a ions a e modeled by he condi ional p obabili y unc ion O(z,a,s)=
p(z|a,s), which gi es he p obabili y o ge ing obse a ion zgi en ha he s a e
is sand ac ion ais pe o med.
– The ewa d ob ained o pe o ming ac ion aa s a e sis R(s,a).
The s a e is no ully obse able; a e e y ime ins an he agen has only access
o obse a ions zwhich gi e incomple e in o ma ion abou he s a e. Thus, a belie
unc ion bis main ained by using he Bayes ule. I ac ion ais applied a belie b
and obse a ion zis ob ained, a new belie bis gi en by:
b(s)=ηO(z,a,s)
s∈S
T(s,a,s)b(s), (1)
whe e he no maliza ion cons an :
Fig. 2 The space is disc e ized by lea ning a opological map om people acks ob ained wi h he
obo senso s (in blue). Some o he nodes a e iden i ied du ing he lea ning phase as goals (in ed).
η=p(z|b,a)=
s∈S
O(z,a,s)
s∈S
T(s,a,s)b(s)(2)
gi es he p obabili y o ob aining a ce ain obse a ion za e execu ing ac ion a o
a belie b.
The objec i e o a POMDP is o ind a policy ha maps belie s in o ac ions in
he o m π(b)→a, so ha he alue is maximized. This alue unc ion ep esen s
he expec ed o al ewa d ea ned by ollowing πdu ing h ime s eps s a ing a he
cu en belie b:Vπ(b)=Eh
=0γ (b ,π(b ))|b0=b, whe e (b ,π(b )) =
s∈SR(s,π(b ))b (s). Rewa ds a e weigh ed by a discoun ac o γ∈[0,1) o
ensu e ha he sum is ini e when h→∞. The e o e, he op imal policy π∗is he
one ha maximizes ha alue unc ion: π∗(b)=a g max
π
Vπ(b).
3.2 S a es
The s a e so ou POMDP consis s o h ee ac o s: he obo posi ion, he pe son
posi ion and he pe son goal. As we employ a disc e e POMDP, he scena io is di ided
in o non-o e lapping egions in o de o disc e ize he obo and pe son posi ions (see
Fig. 2). Each egion has a cen oid and all egions a e combined in o a opological
map ha is disco e ed au oma ically, as i will be desc ibed in he nex sec ion. Also,
he e is a ini e se o goals (each goal co esponds o a egion) whe e he pe son can
go, which a e disco e ed au oma ically oo, as explained la e .
We assume ha he localiza ion sys em o he obo is good enough o be able
o de e mine i s egion wi h high ce ain y. The e o e, he obo posi ion is assumed
obse able, being jus necessa y o keep a belie o e he pe son posi ion and in en-
ion. No e ha he mos ele an unce ain y o he p oblem comes om he pe son
in en ions which a e non-obse able by na u e.
3.3 S a e T ansi ions: A G owing Hidden Ma ko Model
As desc ibed in Sec ion 3.1, he POMDP planne needs a p obabilis ic ansi ion
unc ion o he s a e T(s,a,s), which models he dynamics o people loca ions and
mo ion in en ions. Ins ead o handc a ing his ansi ion model (a Ma ko model),
we ha e de eloped an ex ension o GHMMs [16] o lea n his ansi ion unc ion
om da a.
In a GHMM, he e is a disc e e ep esen a ion o he space, which is di ided in o
egions. T ansi ions a e only allowed be ween neighbo ing egions. The lea ning p o-
cess consis s o es ima ing he bes space disc e iza ion, and iden i ying neighbo ing
egions and ansi ion p obabili ies om obse ed da a. Thus, i s a opological map
is buil wi h he ITM algo i hm [4]; hen, an HMM is buil om he ITM, and i s
ansi ion (and p io ) p obabili ies a e ained wi h he inc emen al Baum-Welch
echnique [7]. Once he GHMM has been ained, i is in eg a ed wi h he POMDP
be o e s a ing he ask execu ion.
Lea ning Phase. The ITM algo i hm s uc u es he space in a g aph whose nodes
ep esen he cen oid o Vo onoi egions whe e people ha e been obse ed and
edges ep esen connec ions be ween adjacen /neighbo ing egions (in he ma he-
ma ical sense). In o de o en o ce an a e age geome ical dis ance be ween nodes, a
h eshold τ o inse new nodes is de ined. Each node main ains and upda es a Gaus-
sian dis ibu ion in ela ion wi h he obse a ions (2D posi ions o people) wi hin i s
co esponding egion. We p opose a a ian o he o iginal ITM algo i hm whe e a bi-
a ia e Gaussian dis ibu ion N(μn,
n)is upda ed o each egion na e each new
obse a ion, whe e μnand na e espec i ely he mean and he co a iance ma ix
o all he obse a ions (x,y) ela ed o n. Thus, ins ead o using a ixed co a iance
ma ix o all he nodes, each node s o es and upda es a speci ic co a iance ma ix,
so he obse a ion model can adap o he cha ac e is ics o he di e en pa s o he
scena io. The cen oids a e also upda ed in a di e en manne as in he o iginal ITM
algo i hm, and each cen oid is compu ed as he mean o i s associa ed obse a ions.
People goals a e au oma ically disco e ed by applying hypo hesis es ing ( - es ),
he e a e wo ypes o goals: en y/exi poin s whe e people appea o disappea and
s anding poin s whe e people s op longe han usual in he scene. The algo i hm is
adap i e, i.e, nodes, edges and goals a e c ea ed, e ased and upda ed dynamically as
mo e people a e obse ed. A p− alue h eshold is de ined in o de o accep o
e use new goals.
Finally, a HMM is used o model all he ansi ion s a e p obabili ies, being he
s a e o a pe son i s node o he g aph and i s goal o in en ion. In pa icula , p io
and ansi ion p obabili ies a e compu ed by applying he Baum-Welch algo i hm,
including people posi ions and eloci ies as obse a ions. A sampling a io pa ame e
Tsis used o sample he obse ed ajec o ies a a cons an a e and eed he Baum-
Welch algo i hm (hence, each s a e ansi ion co esponds wi h a ime Ts). In he
GHMM amewo k, he HMM can be ained se e al imes du ing he lea ning phase.
Fo his pape , he HMM has been gene a ed and ained once a e he c ea ion o
he opological map and goal disco e y, since we a e using an o line app oach. Fo
mo e de ails abou he lea ning phase, please e e o [9].
Belie Es ima ion. Once he ansi ion p obabili ies o he GHMM ha e been lea n
( o he pe son posi ion and goal), he belie o e he pe son posi ion and in en ion can
be upda ed each Tswi h Equa ion 1. No e ha he obo ac ions do no a ec people
posi ions no in en ions in ou model. Mo eo e , he obo posi ions a e conside ed
obse able and i s ansi ion p obabili ies o each ac ion a e hand-coded.
3.4 Obse a ions, Ac ions and Rewa ds
The obo has senso s onboa d o measu e i s own pose and es ima e he pe son
posi ion. A each momen , i can ei he de e mine he egion whe e he pe son is
o no de ec any hing. Senso s a e noisy, and he p obabili y o non-de ec ing he
pe son ( alse nega i e) is p . Mo eo e , i he pe son is in a ce ain egion, i could
be de ec ed in he adjacen egions.
This p obabili y o e oneous de ec ion depends on he dis ance be ween he cen-
oids o he ac ual egion and he obse ed egion. The p obabili y o de ec ing a
pe son o each egion is modeled by a Gaussian cen e ed in he cen oid o he e-
gion. Those Gaussian dis ibu ions a y o each egion and a e lea n oge he wi h
he GHMM [9].
In addi ion, he obo can ake mo emen ac ions a each i e a ion o he planne .
In pa icula , he obo can decide ei he o s ay whe e i is o o mo e o an adjacen
egion. Those ansi ions a e no modeled as de e minis ic and he e is ce ain p ob-
abili y ha he obo may end up in an e oneous egion. Mo eo e , he e is a cos
associa ed wi h mo ing o an adjacen egion, whe eas he e is no cos associa ed
wi h s aying in he same egion.
The objec i e o he obo is o come ac oss he pe son in o de o ha e a con-
e sa ion, he e o e he ewa d unc ion is designed wi h his pu pose. Fo ha , wo
di e en ypes o goals a e conside ed: adequa e o inadequa e. Adequa e goals a e
hose whe e he obo can go and ha e an in e ac ion wi h he pe son. Inadequa e
goals a e hose whe e he obo should no bo he he pe son and go o i s home
posi ion, de ined be o ehand. Thus:
– I he pe son in ends o go o an adequa e goal and he obo is he e, i ge s a
posi i e ewa d Rpos.
– I he pe son in ends o go o an inadequa e goal and he obo is a home posi ion,
i also ge s a posi i e ewa d Rpos.
– I he pe son in ends o go o an adequa e goal and he obo is no he e when he
pe son a i es, i ge s a nega i e penal y Rneg.
– I he pe son in ends o go o an inadequa e goal and he obo is he e, i ge s a
nega i e penal y Rneg.
(a) (b)
Fig. 3 (a) Expe imen al a ea a uni e si y. (b) Schema ic iew wi h he main poin s o in e es and
he home posi ion o he obo .
The home posi ion is a egion mo e o less cen e ed in he scena io whe e he
obo can wai o people o a i e. The ewa d unc ion ies o encou age he obo
o ca ch he people who go o adequa e goals and o a i e he e be o e hem. I also
o ces he obo o go back o home i he pe son in ends o go o an inadequa e goal.
The ac ha he pe son comes ac oss he obo in an inadequa e place is penalized
because i may be conside ed dis u bing.
4 Expe imen s
In his sec ion, we p esen some expe imen al esul s o show he easibili y o ou
app oach. We implemen ed ou decision-making algo i hm in a eal elep esence
obo .
4.1 Expe imen al Se up
The scena io used o ou social ask is one o he es a eas in Pablo de Ola ide
Uni e si y (Fig. 3), which is a space o 4.30 ×11.80 me e s wi h a single en y/exi
poin a one side and se e al poin s o in e es : a spo wi h a co ee and a snack
machine, a doo o he oile s, a wa e on and a es a ea wi h magazines.
We implemen ed ou me hods o people mo ion model and decision making in
C++ unde he Robo Ope a ing Sys em (ROS) amewo k. We used o he expe i-
men s he TERESA obo [11], which is equipped wi h wo lase -scanne s ( on and
back) and a ideo-con e ence sys em. The obo had a map o he scena io and was
able o localize i sel and na iga e be ween waypoin s hanks o he ROS na iga ion
s ack (amcl and mo e_base packages).
Fig. 4 Topological map (blue) and disco e ed goals ( ed) o he scena io. Node 6 o he opological
map is used as home posi ion o he obo .
Fi s , we placed he obo in he scena io wi hou mo ing om he home posi ion,
jus obse ing ajec o ies o people passing by. Wi h ha in o ma ion, we an he
algo i hm desc ibed in Sec ion 3.3 o lea n a opological map o he scena io wi h
he cen oids o he egions, he possible goals o people and a GHMM wi h he
ansi ion p obabili ies.
The obo used he wo lase -scanne s o pe son de ec ion and acking, applying
he algo i hm in [1] and a Kalman Fil e o empo al acking and eloci y es ima ion.
Mo e han 200 people ajec o ies we e eco ded in a da ase and used o ain he
models1, gene a ing a opological map wi h 21 nodes, 32 edges and 5 disco e ed goals
(see Fig. 4). The algo i hm was able o disco e as goals all he poin s o in e es in he
scene: (1) en y/exi doo ; (2) wa e on , (3) oile s, (4) co ee/snack machines, (5)
es a ea. The co esponding GHMM was ained by sampling he people ajec o ies
a Ts=1 Hz, esul ing in 105 s a es and 1,705 ansi ion p obabili ies.
Once he GHMM was lea n , we implemen ed ou POMDP model2 o ob ain
a policy o he obo . In his case, we used an o line POMDP sol e , Symbolic
Pe seus [10]. Du ing he expe imen s, we an wo di e en modules: a module using
he GHMM o es ima e he belie o he pe son posi ion and in en ion, and a module
o de e mine he bes ac ion o he obo a each ime gi en he cu en belie .
The es ima o module is execu ed a 1 Hz whe eas he decision-make a 0.33 Hz.
Mo eo e , he decision-make commands he obo o s ay a he same egion o
o go o adjacen ones, which means sending o he mo e_base na iga o he
co esponding waypoin (cen oid o he des ina ion egion).
4.2 Resul s
In o de o e alua e he beha io o he obo wi h he compu ed policy, we an
di e en ials whe e people we e appea ing a he scena io and going o di e en
places. In gene al, we obse ed a common beha io : he obo wai s be o e mo ing
1The ITM algo i hm was execu ed wi h τ=1 me e o node inse ion and p− alue =10−4
o hypo hesis es ing.
2The pa ame e s we e se as p =0.1, Rpos =10 and Rneg =−10.