E icien Monocula Pose Es ima ion o Complex 3D Models
A. Rubio, M. Villamiza , L. Fe az, A. Pena e-Sanchez,
A. Ramisa, E. Simo-Se a, A. San eliu and F. Mo eno-Nogue
Ins i u de Robò ica i In o mà ica Indus ial, CSIC-UPC
Llo ens A igas 4-6, 08028 Ba celona, Spain
Abs ac — We p opose a obus and e icien me hod o
es ima e he pose o a came a wi h espec o complex 3D
ex u ed models o he en i onmen ha can po en ially con ain
mo e han 100,000 poin s. To ackle his p oblem we ollow a
op down app oach whe e we combine high-le el deep ne wo k
classi ie s wi h low le el geome ic app oaches o come up wi h
a solu ion ha is as , obus and accu a e. Gi en an inpu
image, we ini ially use a p e- ained deep ne wo k o compu e
a ough es ima ion o he came a pose. This ini ial es ima e
cons ains he numbe o 3D model poin s ha can be seen
om he came a iewpoin . We hen es ablish 3D- o-2D co es-
pondences be ween hese po en ially isible poin s o he model
and he 2D de ec ed image ea u es. Accu a e pose es ima ion
is inally ob ained om he 2D- o-3D co espondences using a
no el PnP algo i hm ha ejec s ou lie s wi hou he need o use
a RANSAC s a egy, and which is be ween 10 and 100 imes
as e han o he me hods ha use i . Two eal expe imen s
dealing wi h e y la ge and complex 3D models demons a e
he e ec i eness o he app oach.
I. INTRODUCTION
Robus came a localiza ion is a undamen al p oblem in
a wide ange o obo ics applica ions, going om p ecise
objec manipula ion o au onomous ehicle na iga ion. De-
spi e being a opic esea ched o decades i is s ill an open
challenge. The e exis app oaches based on in a ed came as
and high- equency sys ems such as Vicon [1], which ha e
shown excellen esul s o localiza ion and na iga ion o
obo s. Howe e , hese sys ems a e limi ed o indoo en i-
onmen s whe e ligh ing condi ions a e con olled.
Ano he al e na i e o obus ly localize he obo is o
equip i wi h mul iple senso s, such as lase s o s e eo
came as, and hen using he da a om each o hem.
Al hough hese sys ems inc ease he eliabili y o he obo
o sel -localiza ion, hey ha e a nega i e impac on he
compu a ional cos and payload. This is specially c i ical in
some obo ics asks whe e small and low-cos obo s (e.g.
ae ial obo s) a e equen ly used.
In con as o hese mul i-senso app oaches, in his wo k
we p opose an e icien and obus sys em based uniquely on
a monocula came a and a known p e-compu ed 3D ex u ed
model o he en i onmen , as shown in Fig. 1. Indeed, he
p oposed me hod is able o e icien ly es ima e he ull pose
( o a ion and ansla ion) o he came a wi hin he 3D map,
gi en solely one single inpu image. No empo al in o ma ion
is used abou he p e ious poses ha can cons ain he egion
o he 3D model whe e he came a is poin ing o. While his
makes ou p oblem signi ican ly mo e complex, i makes he
esul ing pose es ima ion obus o issues such as d i ing,
Fig. 1: P oblem de ini ion: Gi en an inpu image (a) and
a known ex u ed 3D model o he en i onmen (b), he
p oblem is o es ima e he pose o he came a ha cap u ed
he image wi h espec o he model (c). The main challenge
add essed in his pape is o pe o m he co espondence o
poin s e icien ly and eliably o complex 3D ha con ain
a la ge numbe o poin s. In his example, he model o he
Sag ada Familia has o e 100,000 poin s.
occlusions o he model du ing sho pe iods o ime, o
sudden came a mo ions.
Mo e p ecisely, le us assume ou 3D model is made
o n3D poin s, each associa ed o a isual desc ip o
ep esen ing i s appea ance. Fig. 2 shows he o e all scheme
o he p oposed me hod. In ou case hese isual desc ip o s
co espond o SIFT ea u es [12] ob ained om a se o
aining images p e iously used, in an o -line s ep, o build
he model (Fig. 2(g)). A un ime, he desc ip o s o he 3D
model a e compa ed agains he mdesc ip o s ex ac ed om
he inpu image (Fig. 2(d)) in o de o de e mine a i s se
o 3D- o-2D ma ch candida es, o hen es ima e he pose
(Fig 1(c)). Howe e , sol ing his co espondence p oblem
has an O(n·m)complexi y, which is ex emely cos ly i we
conside ha ncan be e y la ge (e.g. 100,000 poin s).
In o de o alle ia e he compu a ional load, we p opose
including a p elimina y s ep based on a deep-lea ning ne -
wo k ha yields an app oxima e ini ial pose, wi hou he
need o explici ly compu e co espondences. This me hod
p o ides he kmos simila images (among he aining
images used o build he 3D model) o he inpu image
(Fig. 2(e)). Then, he desc ip o s o hese images a e used
o pe o m he 3D- o-2D ma ching (Fig. 2( )). The e o e, he
co espondence is done wi h a small pa o he model bu no
wi h he comple e model. And mos impo an ly, his ini ial
Image ep esen a ion k-nea es images
R,
Deep Ne wo k
3D model o he objec
(Images and SIFT ea u es)
REPPnP
Ma ching
Inpu image
(a)
(d)
Fea u e de ec ion
and ex ac ion (SIFT)
(b) (e)
(g)
( ) (h)
(c)
Fig. 2: O e all scheme o he p oposed me hod o e icien came a pose es ima ion using highly complex models.
p e-selec ion o mos simila images is done in a ma e o
milliseconds. Since hese co espondences may s ill con ain
alse ma ches, we ge id o hem by using REPPnP [4],
a no el RANSAC-less algo i hm o ejec ing ou lie s. The
expe imen s ha e shown ha he p oposed me hod no only
educes he compu a ion cos bu ha i also achie es high
accu acy a es in highly complex models.
The es o he pape desc ibes each componen s o he
p oposed me hod. Sec. III explains he cons uc ion o he
3D model. Sec. IV desc ibes he ini ial pose es ima ion
using he deep-lea ning ne wo k and he REPPnP algo i hm
used o he subsequen e inemen . In Sec. V he me hod
is ex ensi ely e alua ed o e wo di e en models. Finally,
Sec. VI summa izes he main con ibu ions and u u e wo k.
II. RELATED WORK
Me hods o 3D pose es ima ion om monocula images
can be oughly spli in o wo ca ego ies: Geome ic ap-
p oaches ha ely on local image ea u es (e.g., poin s)
and use geome ic ela ions o compu e he pose; and
Appea ance-based me hods ha compu e global desc ip o s
o he image, and hen use machine lea ning app oaches o
es ima e which image wi hin he aining se is he closes
one o a gi en inpu image.
Geome ic app oaches use local desc ip o s o es ima e
2D- o-3D ma ches be ween one inpu image and one o
se e al e e ence images egis e ed o a 3D model. PnP
algo i hms such as EPnP [10], [13] a e hen used o en o ce
geome ic cons ain s and sol e o he pose pa ame e s. On
op o ha , obus RANSAC-based s a egies [2], [14] can
be used bo h o speed up he ma ching p ocess and o il e
ou lie co espondences. Ye , while hese me hods p o ide
e y accu a e esul s, hey equi e bo h he e e ence and
inpu images o be o high quali y, such ha local ea u es
can be eliably and epe i i ely ex ac ed. Addi ionally, i he
numbe o poin s is e y la ge, he ou lie ejec ion scheme
can become ex emely slow. Recen ly, [23] has shown ha
p io s abou he o ien a ion o he came a ela i e o he
g ound plane can speed up his p ocess. We do no conside
hese kind o p io s o ou me hod, hough.
On he o he hand, app oaches elying on global desc ip-
ions o he image a e less sensi i e o a p ecise localiza ion
o indi idual ea u es. These me hods ypically use a se o
aining images acqui ed om di e en iewpoin s o s a-
is ically model he spa ial ela ionship o he local ea u es,
ei he using one single de ec o o all poses [7], [11], [22] o
a combina ion o a ious pose-speci ic de ec o s [15], [16],
[24], [26], [27]. Ano he al e na i e is o bind image ea u es
wi h poses du ing aining and ha e hem o e in he pose
space [6]. The limi a ion o hese app oaches is ha he
es ima ed pose ends o be innacu a e, and highly depends
on he spa ial esolu ion a which he aining images ha e
been acqui ed. The highes g anula i y o he aining se ,
he mo e p ecise can be he es ima ed pose, al hough he
esul ing de ec o is mo e p one o gi e alse posi i es.
In his pape we combine he bes o bo h wo lds. On
he one hand, we will use a global desc ip o based on a
deep con olu ional ne wo k o ge a i s es ima ion o he
pose. Deep ne wo ks ha e ecen ly shown imp essi e esul s
in image classi ica ion asks [9]. This ini ial es ima e will in
i s u n educe he numbe o po en ial 2D- o-3D ma ches,
and make geome ic app oaches applicable. On he geome ic
side, we will make use o a e y ecen PnP app oach,
which inhe en ly inco po a es an ou lie ejec ion scheme
wi hou he need o un RANSAC [5]. The combina ions
o bo h ing edien s will esul in a powe ul pose es ima ion
algo i hm capable o dealing wi h e y la ge models.
III. BUILDING THE 3D MODEL
Ou app oach assumes a 3D ex u ed model o he scene
o be a ailable. To build hese models we used Bundle [20],
a s uc u e- om-mo ion sys em o uno de ed image collec-
ions. Gi en a se o Nimages o he scene we seek o model,
his package ini ially ex ac s SIFT desc ip o s o all images,
and hen simul aneously es ima es he Ncame a poses and
3D s uc u e using bundle adjus men . This is usually a ime
consuming p ocess ha can ake a ew hou s.
Fo each 3D poin o he model, Bundle p o ides i s SIFT
desc ip o , i s 3D posi ion, i s RGB colo and he numbe
o images whe e he poin appea s. Also, o each image i
gi es he calib a ion ma ix, he o a ion ma ix, R, and he
ansla ion ec o , . This da a will be used as “g ound u h”
when e alua ing he p ecision o he algo i hm.
Fo his pape , we used wo 3D models: The Sag ada
Familia da ase ( om [17]), composed o 478 images o
his chu ch in Ba celona. The esul ing 3D model con ains
Fig. 3: Cou ya d 3D model buil and wo sample images.
100,532 3D poin s. Fig. 1(b) shows he model, and plo s
in yellow he e ie ed came a pose o each one o he N
aining se images. The Cou ya d o he ma hema ics school
da ase is composed o 265 images, which esul ed in a 3D
model wi h 30,196 poin s. The model and a wo sample
images a e shown in Fig. 3. No e ha while he amoun o
poin s in he i s model is la ge , i s images p esen less
di e ence in e ms o isual appea ance.
IV. MERGING APPEARANCE AND GEOMETRIC METHODS
We nex desc ibe how appea ance and geome ic me h-
ods a e combined in a op-down coa se- o- ine app oach,
o ackle he p oblem o pose es ima ion om e y la ge
models. Le us i s o mula e ou p oblem:
Assume we a e gi en a 3D model wi h npoin s
{p1,...,pn}and N aining images {T1, . . . , TN}which
ha e been used o compu e he 3D model. Each o hese
images has an associa ed pose {Ri, i}. Each poin pihas
an associa ed SIFT desc ip o and a isibili y lis iwi h he
indexes o he aining images om whe e i is seen. Gi en
an inpu image I, ou goal is o accu a ely compu e he pose
om whe e i was acqui ed.
One s aigh o wa d solu ion would be o ex ac m2D
ea u es {u1,...,um} om he inpu image Iand compu e
2D- o-3D co espondences {ui↔pi}by jus compa ing
SIFT desc ip o s. This se o co espondences could hen be
il e ed using a RANSAC+PnP scheme o ge id o ou lie s
and accu a ely es ima e he pose. None heless, since we a e
conside ing cases whe e n >> m hose co espondences a e
p one o con ain a e y la ge pe cen age o ou lie s, which
migh d ama ically slow down he p ocess.
In his wo k we p opose a wo s age s a egy ha combines
he so-called appea ance and geome ic me hods. The o me
will compu e he subse o he aining images which is
mo e simila o ou inpu image. This will cons ain he
se o candida e poses. The la e will use his subse o
poses o limi he numbe o 3D poin s o he model ha
a e po en ially isible, and e ine he pose using a geome ic
app oach. These wo s eps a e nex discussed.
A. Coa se Pose Es ima ion
Gi en ou inpu image Iand aining images
{T1, . . . , TN}we seek o design a as and obus s a egy
o ge he k aining images which a e mo e simila o I.
Fo ob aining his subse we could ep esen he images
using any global image desc ip o , e.g, bag o wo ds, GIST,
and simply compa e such desc ip o s.
In his pape , hough, we ha e used he ecen gene ic
image-le el ea u es ob ained om a deep ne wo k o image
ep esen a ion. In pa icula , we use he 4,096-dimensional
second o las laye o a Con olu ional Neu al Ne wo k
(CNN) as ou high-le el image ep esen a ion. The ull
ne wo k has 5 con olu ional laye s ollowed by 3 ully
connec ed laye s, and ob ained he bes pe o mance in he
ILSVRC-2012 challenge. The ne wo k is ained on a subse
o ImageNe [3] o classi y 1,000 di e en classes. We use
he publicly a ailable implemen a ion and p e- ained model
p o ided by [8]. The ea u es ob ained wi h his p ocedu e
ha e been shown o gene alize well and ou pe o m adi-
ional hand-c a ed ea u es, hus hey a e al eady being used
in a wide di e si y o asks [21].
Compa ing he esul ing ec o s o his ep esen a ion we
ob ain, o each inpu image I, a subse o he kmos
“simila ” aining images {T1, . . . , Tk}. These images may
o may no be clus e ed in a simila egion o he space.
Indeed, he alue o kis chosen su icien ly la ge o ensu e
ha he subse o “simila ” images con ains a leas one
image which is ac ually close o I, i.e, he se o poses
associa ed o hese images ep esen s jus a e y ough
es ima ion o he g ound u h pose. We nex explo e hem
and e ine he pose using a geome ic app oach.
B. Fine Pose Es ima ion
The kcloses aining images p o ide a coa se es ima e
o he came a pose. To inc ease he accu acy we adap ed he
REPPnP me hod [4], which simul aneously allows o disca d
ou lie co espondences while es ima es he came a pose.
S anda d PnP app oaches assume he 2D- o-3D co espon-
dences o be ee o ou lie s. The e o e, when dealing wi h
eal images hese me hods need an ou lie ejec ion p ep o-
cessing s ep (e.g. RANSAC + P3P), which may signi ican ly
educe he o e all compu a ional e iciency.
Gi en c2D- o-3D co espondences {ui↔pi}and he
came a in e nal calib a ion ma ix A, PnP me hods build
upon he pe spec i e cons ain o each 2D ea u e poin i,
diui
1=A[R| ]pi
1,(1)
whe e diis he dep h o he ea u e poin .
REPPnP e o mula es he p e ious equa ions as a low-
ank homogeneous se o equa ions Mx = 0, whe e Mis
a2c×12 ma ix ep esen ing he pe spec i e cons ain s o
all cco espondences. In [4] is shown ha he solu ions x
can be es ima ed assuming ha he ank o he null-space
o Mis equal o 1. Once xis es ima ed, he came a pose
[R| ]is sol ed using a gene aliza ion o he O hogonal
P oc us es p oblem [18], combined wi h a p ojec ed g adi-
en s op imiza ion. In REPPnP, he ou lie ejec ion is done
by i e a ing he es ima ion o x. A each i e a ion, hose
equa ions in M ha a e being p ojec ed on o xp o ide he
lowes algeb aic e o s a e chosen. This p ocedu e shows
kTime (s) # 3D poin s # Ma ches e o e ans
Sag ada Familia model
All 45.6694 100532 470 0.0124 0.0231
612.3344 24855 424 0.0178 0.0305
814.1634 29046 436 0.0176 0.0306
12 16.9470 35422 447 0.0177 0.0303
16 20.2644 43005 452 0.0178 0.0306
Cou ya d model
All 50.2912 30196 379 0.0279 0.0449
315.5296 6189 296 0.0255 0.0441
621.2358 10038 322 0.0267 0.0476
824.3889 12134 328 0.0212 0.0343
12 28.9528 15152 334 0.0206 0.0360
16 33.5211 18145 337 0.0303 0.0538
TABLE I: Pe o mance o he p oposed app oach in he wo
models o di e en alues o he pa ame e k. We also
conside he case when all he model poin s a e conside ed.
con e gence wi h up o 50% o ou lie s and equi ing up
o wo o de s o magni ude less compu a ional ime han
s anda d P3P + RANSAC + PnP algo i hms.
In o de o adap REPPnP o he case o his pape wi h
la ge scena ios, we p opose building he Mma ix as k
di e en pa s, each coming om one o he ksimila images
e ie ed in he p e ious s age. By doing his we can educe
d as ically he numbe o ou lie s, o a es below he 50%,
o which he REPPnP is shown o succeed. Conc e ely, we
p opose o sol e,
hM1>M2>... Mk>i>
x= 0 (2)
whe e Mj o j= [1,2, ..., k] ep esen he homogeneous
equa ions ob ained as in [4] by ma ching he 2D poin s
uiwi h he 3D poin s pj
iwhich a e isible (acco ding o
,. . . , k) in each o he knea es images.
Wi h he p oposed me hod, aking k= 6, we dec ease
he numbe o 3D poin s o be ma ched o one qua e
o he o al in he Sag ada Familia scena io and o one
hi d in he Cou ya d scena io, inc easing he numbe o
inlie co espondences. Table I shows he pe o mance o
ou app oach o bo h models and di e en alues o k.
V. RESULTS
In his sec ion, he e iciency and accu acy o he p oposed
me hod will be e alua ed using he wo models p esen ed in
Sec.III. P e iously, we will show how ou image e ie al
algo i hm pe o ms on he es da ase s.
Fo he expe imen s, bo h da ase s ha e been e enly spli
in o wo disjoin se s. One o hem is used as aining se o
he cons uc ion o he model, and he o he one as es ing
se o he e alua ion o he me hod.
A. Image e ie al o Coa se Pose Es ima ion
In o de o e alua e he quali y o he image e ie al using
he deep ne wo k desc ip o s, we ha e plo ed in Fig. 4 he
Euclidean dis ance be ween he desc ip o s o each pai
o es / ain images in he Sag ada Familia da ase . Fo
each es image ( ow) he aining images a e o de ed by
angula dis ance o he o a ion ma ix. The accumula ion o
bluish egions on he le hand side o he g aph indica es
ha he dis ance be ween desc ip o s inc eases as we mo e
Fig. 4: Compa ison o he L2 dis ances be ween dis ance
deep ne wo k ea u es and he known angula dis ance
be ween he es images and aining images. Each ow is
so ed acco ding o he angula dis ance. We can see a s ong
co ela ion be ween images ha a e close o each o he and
hei desc ip o s ob ained om he deep ne wo k.
away om he o iginal iewpoin , and hus con i ms a high
con idence in he image e ie al p ocess we use.
As illus a i e examples, in Fig. 5 we show one sample
que y o each da ase . We plo bo h he inpu image and he
closes images in he aining se acco ding o he dis ances
be ween he deep ne wo k desc ip o s. In bo h cases nea ly
all selec ed images p esen simila poses o he es image.
B. E iciency assessmen
In o de o ha e a clea pic u e o he cos o ou me hod,
we ha e measu ed he compu a ional ime o each o he
di e en s ages: SIFT ea u e ex ac ion on he inpu image,
image e ie al h ough ou deep lea ning app oach, selec ion
o he co esponding SIFT desc ip o s om he model,
desc ip o ma ching and ou lie emo al along wi h pose
es ima ion done by he REPPnP algo i hm. To compa e he
compu a ion ime o he app oach agains a ypical baseline,
we ha e eplaced REPPnP in ou pipeline, by RANSAC
combined wi h OPnP [28], one o he as es and mos
accu a e app oaches in he s a e-o - he-a .
A e age imes o e he comple e es se a e gi en in
Fig. 6. As can be obse ed, he coa se il e ing o model
images educes signi ican ly he compu a ion ime o he
geome ical es ima ion o he came a pose.
I can also be seen in he igu e ha he mos ime
consuming s ep o he whole pipeline is by a he SIFT
desc ip o ma ching. This s ep is pe o med using he MAT-
LAB implemen a ion o he VLFea open sou ce lib a y [25].
I is, he e o e, c i ical o ob ain a good se o neighbo ing
images: he be e i is, he less model desc ip o s will be
equi ed in he geome ical es ima ion s ep. In ou expe i-
men s, his coa se o ine app oach led o 75% educ ion o
he compu a ional ime, bu i could be e en mo e signi ican
on la ge models.
Finally, in Table II we show he a e age ime equi ed
o compu e he deep ne wo k desc ip o s and he bag o
isual wo ds om a new image (le ), and he REPPnP
Fig. 5: Inpu image 422 o he Sag ada Familia model ( op) and 18 o he Cou ya d model (bo om), wi h simila images
selec ed by ou image e ie al algo i hm (k= 4 and k= 5) and a plo showing he dis ance o he aining images.
Compu a ion Time Compu a ion Time (Zoom)
All 6 8 12 16
0
25
50
k
Time (s)
COMPUTATION TIME
Ma ching
Inpu image SIFT
REPPNP
Nea SIFT selec ion
Deep lea ning
All 6 8 12 16
0
0.05
0.1
0.15
0.2
k
Time (s)
COMPUTATION TIME (ZOOM)
Fig. 6: Analysis o compu a ion ime. Le : Compu a ion
ime o he app oaches wi h di e en alues o k. The
case All is equi alen o no doing any appea ance based
p e-compu a ion o e he 3D model. Righ : Zoom o he
compu a ion imes.
and RANSAC ( igh ). While a deep ne wo k desc ip o can
be ex ac ed di ec ly om an inpu image in a ma e o
milliseconds, he bag o isual wo ds equi es i s compu ing
SIFT desc ip o s o ha image (on he o de o seconds,
depending on he size o he image), and hen inding he
co esponding isual wo d o each desc ip o wi h a p e-
compu ed dic iona y. Rega ding he geome ical es ima ion
pa , by using REPPnP we a e able o educe he ime
equi ed by a ac o o 5.
C. Accu acy
The accu acy is compu ed as he o a ion and he ans-
la ion e o s in he calcula ed pose. The o a ion e o
is es ima ed using qua e nions as e o =kqua (R)−
qua (R ue)k/kqua (R ue)k, and he ansla ion e o as
e ans =k − uek/k uek. The es ima ed pose is {R, }, and
{R ue, ue}, co esponds o he g ound u h gi en by he
Bundle algo i hm, as men ioned in Sec. III. As obse ed in
Fig. 7, he e o s ound using he coa se o ine app oach a e
compa able o he ones ob ained wi h he comple e model.
In all expe imen s, we ob ain a signi ican educ ion o he
e o when compa ed o RANSAC+OPnP.
Fig. 8 shows a quali a i e compa ison o he ep ojec ion
o he SIFT desc ip o s in he inpu image ob ained wi h he
Rand gi en by REPPnP and RANSAC+OPnP, along wi h
he g ound u h ep ojec ion (gi en by Bundle ). By using
REPPnP we ob ain good esul s o e all he da ase s while in
Model Bag o Wo ds CNN Fea u es Gain (%)
Sag. Familia 1.63 0.0208 98.72
Cou ya d 5.48 0.0207 99.62
kRANSAC+OPnP REPPnP Gain (%)
60.0217 0.0047 78.34
80.0193 0.0030 84.46
12 0.0188 0.0030 84.04
16 0.0187 0.0030 83.96
TABLE II: Time alues in seconds o he di e en me hods
e alua ed. The uppe pa o he able compa es he wo
me hods e alua ed o he coa se es ima ion o he pose (Bag
o wo ds s. Deep ne wo k ea u es). The lowe pa epo s
he compu a ion ime o he me hods used o ine pose
es ima ion (RANSAC + OPnP s. REPPnP).
con as , RANSAC+OPnP, does no always p o ide a good
es ima ion in all images.
VI. CONCLUSIONS
The ime equi ed o es ima e pose on a la ge scale 3D
models has been signi ican ly educed using a me hod ha
combines a pu ely appea ance based echnique wi h a geo-
me ical app oach. By using i s global image appea ance
we educe he numbe o ma ches o es bu a he same
ime by pe o ming a PnP ma ch o e he candida e images
he e o in ou pose es ima ion becomes nea ly negligible.
In his wo k we ha e shown how i is possible o le e age
deep lea ning echniques o imp o e appea ance based ap-
p oaches o he obo ic communi y. I is ema kable ha we
a e able o pe o m accu a e pose es ima ion o e hund eds
o housands o poin s in a ew seconds pe image, especially
conside ing ha we a e using a MATLAB implemen a ion.
As u u e wo k, we will conside he possibili y o u he
exploi ing deep ne wo ks by eplacing he ubiqui ous SIFT
desc ip o by desc ip o s lea n wi h CNNs [19].
VII. ACKNOWLEDGMENTS
This wo k has been pa ially unded by he Spanish Min-
is y o Economy and Compe i i eness unde p ojec s ERA-
Ne Chis e a p ojec ViSen PCIN-2013-047, PAU+ DPI2011-
27510 and ROBOT-INT-COOP DPI2013-42458-P, and by
he EU p ojec ARCAS FP7-ICT-2011-28761.
Fig. 7: Ro a ion and ansla ion e o s o di e en alues o knea es images o bo h models. All cha s show, as a
e e ence, he co esponding e o when es ima ing he pose wi h he whole model.
(a)!(b)!(d)!
(c)!
Ro a ion e o s
REPPnP RANSAC+OPnP
(a) 0.0068 0.0007
(b) 0.0042 0.0812
(c) 0.0033 0.0024
(d) 0.0012 0.0079
T ansla ion e o s
REPPnP RANSAC+OPnP
(a) 0.0306 0.0383
(b) 0.0109 0.1497
(c) 0.0091 0.0091
(d) 0.0109 0.0267
Fig. 8: Examples o pose es ima ion esul s o he Sag ada Familia (le ) and Cou ya d ( igh ) models. Fo each image
we show he ep ojec ed 3D coo dina es o he SIFT poin s using he g ound u h Rand ( om Bundle ) and he ones
es ima ed by bo h REPPnP and RANSAC+OPnP wi h k= 6.
REFERENCES
[1] Vicon. www. icon.com.
[2] O. Chum and J. Ma as. Ma ching wi h PROSAC-p og essi e sample
consensus. In CVPR, 2005.
[3] J. Deng, W. Dong, R Soche , L.-J. Li, K. Li, and L Fei-Fei. Imagene :
A la ge-scale hie a chical image da abase. In CVPR, 2009.
[4] L. Fe az, X. Bine a, and F. Mo eno-Nogue . Ve y as solu ion o he
PnP p oblem wi h algeb aic ou lie ejec ion. In CVPR, 2014.
[5] M. A. Fischle and R. C. Bolles. Random sample consensus: a
pa adigm o model i ing wi h applica ions o image analysis and
au oma ed ca og aphy. In Communica ions ACM, 1981.
[6] D. Glasne , M. Galun, S. Alpe , R. Bas i, and G. Shakhna o ich.
Viewpoin -awa e objec de ec ion and pose es ima ion. In ICCV, 2011.
[7] W. Hu and S.-C. Zhu. Lea ning a p obabilis ic model mixing 3D and
2D p imi i es o iew in a ian objec ecogni ion. In CVPR, 2010.
[8] Y. Jia. Ca e: An open sou ce con olu ional a chi ec u e o as ea u e
embedding. In h p://ca e.be keley ision.o g/, 2013.
[9] A. K izhe sky, I. Su ske e , and G. Hin on. Imagene classi ica ion
wi h deep con olu ional neu al ne wo ks. In NIPS, 2012.
[10] V. Lepe i , F. Mo eno-Nogue , and P. Fua. EPnP: An accu a e O(n)
solu ion o he PnP p oblem. IJCV, 81(2):155–166, 2009.
[11] J. Liebel and C. Schmid. Mul i- iew objec class de ec ion wi h a 3d
geome ic model. In CVPR, 2010.
[12] D. Lowe. Dis inc i e image ea u es om scale-in a ian keypoin s.
IJCV, 60(2):91–110, 2004.
[13] F. Mo eno-Nogue , V. Lepe i , and P. Fua. Accu a e noni e a i e O(n)
solu ion o he PnP p oblem. In ICCV, 2007.
[14] F. Mo eno-Nogue , V. Lepe i , and P. Fua. Pose p io s o simul ane-
ously sol ing alignmen and co espondence. In ECCV, 2008.
[15] M. Ozuysal, V. Lepe i , and P. Fua. Pose es ima ion o ca ego y
speci ic mul i iew objec localiza ion. In CVPR, 2009.
[16] N. Paye and S. Todo o ic. F om con ou s o 3D objec de ec ion and
pose es ima ion. In ICCV, 2011.
[17] A. Pena e-Sanchez, F. Mo eno-Nogue , J. And ade-Ce o, and
F. Fleu e . LETHA: Lea ning om high quali y inpu s o 3D pose
es ima ion in low quali y images. In 3DV, 2014.
[18] P.H. Schönemann and R.M. Ca oll. Fi ing one ma ix o ano he
unde choice o a cen al dila ion and a igid mo ion. Psychome ika,
35(2):245–255, 1970.
[19] Edga Simo-Se a, Edua d T ulls, Luis Fe az, Iasonas Kokkinos,
and F ancesc Mo eno Nogue . F acking Deep Con olu ional Image
Desc ip o s. CoRR, abs/1412.6537, 2014.
[20] N. Sna ely, S. M. Sei z, and R. Szeliski. Pho o ou ism: Explo ing
image collec ions in 3d. In SIGGRAPH, 2006.
[21] R. Soche , A. Ka pa hy, Q.V. Le, C.D. Manning, and A.Y. Ng.
G ounded composi ional seman ics o inding and desc ibing images
wi h sen ences. In TACL, 2014.
[22] H. Su, M. Sun, L. Fei-Fei, and S. Sa a ese. Lea ning a dense
mul i- iew ep esen a ion o de ec ion, iewpoin classi ica ion and
syn hesis o objec ca ego ies. In ICCV, 2009.
[23] L. S a m, O. Enq is , M. Oska sson, and F. Kahl. Accu a e localiza-
ion and pose es ima ion o la ge 3d models. In CVPR, June 2014.
[24] A. Thomas, V. Fe a i, B. Leibe, T. Tuy elaa s, B. Schiel, and L. Van
Gool. Towa ds mul i- iew objec class de ec ion. In CVPR, 2006.
[25] A. Vedaldi and B. Fulke son. VLFea : An open and po able lib a y
o compu e ision algo i hms. 2008.
[26] M. Villamiza , A. Ga ell, A. San eliu, and F. Mo eno-Nogue . Online
human-assis ed lea ning using andom e ns. In ICPR, 2012.
[27] M. Villamiza , A. San eliu, and F. Mo eno-Nogue . Fas online
lea ning and de ec ion o na u al landma ks o au onomous ae ial
obo s. In ICRA, 2014.
[28] Y. Zheng, Y. Kuang, S. Sugimo o, K. As öm, and M. Oku omi.
Re isi ing he pnp p oblem: A as , gene al and op imal solu ion. In
ICCV, 2013.