scieee Science in your language
[en] (orig)

Repositorio Institucional de Documentos

Abstract

Los seres humanos comprendemos los entornos que nos rodean sin esfuerzo y bajo una amplia variedad de condiciones, lo cual es debido principalmente a nuestra percepción visual. Desarrollar algoritmos de Computer Vision que logren una comprensión visual similar es muy deseable, para permitir que las máquinas puedan realizar tareas complejas e interactuar con el mundo real, con el principal objectivo de ayudar y entretener a los seres humanos. <br />En esta tesis, estamos especialmente interesados en los problemas que surgen durante la búsqueda de la comprensión visual de espacios interiores, ya que es dónde los seres humanos pasamos la mayor parte de nuestro tiempo, así como en la búsqueda del sensor más adecuado para logar dicha comprensión. Con respecto a los sensores, en este trabajo proponemos utilizar cámaras no convencionales, en concreto imágenes panorámicas y sensores 3D. Con respecto a la comprensión de interiores, nos centramos en tres aspectos clave: estimación del diseño 3D de la escena (distribución de paredes, techo y suelo); detección, localización y segmentación de objetos; y modelado de objetos por categoría, para los que se proporcionan soluciones novedosas y eficientes. El enfoque de la tesis se centra en los siguientes desafíos subyacentes. <br />En primer lugar, investigamos métodos de reconstrucción 3D de habitaciones a partir de una única imagen de 360, utilizado para lograr el nivel más alto de modelado y comprensión de la escena. Para ello combinamos ideas tradicionales, como la asunción del mundo Manhattan por la cual la escena se puede definir en base a tres direcciones principales ortogonales entre si, con técnicas de aprendizaje profundo, que nos permiten estimar probabilidades en la imagen a nivel de pixel para detectar los elementos estructurales de la habitación. Los modelos propuestos nos permiten estimar correctamente incluso partes de la habitación no visibles en la imágen, logrando reconstrucciones fieles a la realidad y generalizando por tanto a modelos de escena más complejos. Al mismo tiempo, se proponen nuevos métodos para trabajar con imágenes panorámicas, destacando la propuesta de una convolución especial que deforma el kernel para compensar las distorsiones de la proyección equirrectangular propia de dichas imágenes.<br />En segundo lugar, considerando la importancia del contexto para la comprensión de la escena, estudiamos el problema de la localización y segmentación de objetos, adaptando el problema para aprovechar todo el potencial de las imágenes de $360^\circ$. También aprovechamos la interacción escena-objetos para elevar las detecciones 2D en la imagen de los objetos al modelo 3D de la habitación.<br />La última línea de trabajo de esta tesis se centra en el análisis de la forma de los objetos directamente en 3D, trabajando con nubes de puntos. Para ello proponemos utilizar un modelado explícito de la deformación de los objetos e incluir una noción de la simetría de estos para aprender, de manera no supervisada, puntos clave de la geometría de los objetos que sean representativos de los mismos. Dichos puntos estan en correspondencia, tanto geométrica como semántica, entre todos los objetos de una misma categoría.<br />Nuestros modelos avanzan el estado del arte en las tareas antes mencionadas, siendo evaluados cada uno de ellos en varios datasets y en los benchmarks correspondientes.<br /> <br /> Fernández Labrador, Clara; Guerrero Campo, José Jesús; Demonceaux, Cédric

Read accessible full text

Repositorio Institucional de Documentos

Publisher: Universidad de Zaragoza, Prensas de la Universidad
Year: 2020
Source: https://zaguan.unizar.es/record/100733/files/TESIS-2021-090.pdf
2021
90
Cla a Fe nández Lab ado
Indoo Scene
Unde s anding using Non-
Con en ional Came as
Di ec o /es
Gue e o Campo, José Jesús
Demonceaux, Céd ic
© Uni e sidad de Za agoza
Se icio de Publicaciones
ISSN 2254-7606
Cla a Fe nández Lab ado
INDOOR SCENE UNDERSTANDING USING NON-
CONVENTIONAL CAMERAS
Di ec o /es
Gue e o Campo, José Jesús
Demonceaux, Céd ic
Tesis Doc o al
Au o
2020
UNIVERSIDAD DE ZARAGOZA
Escuela de Doc o ado
P og ama de Doc o ado en Ingenie ía de Sis emas e In o má ica
Reposi o io de la Uni e sidad de Za agoza – Zaguan h p://zaguan.uniza .es
Indoo Scene Unde s anding using
Non-Con en ional Came as
Cla a Fe nández Lab ado
Supe iso : José J. Gue e o
Céd ic Demonceaux
I3A, Uni e sidad de Za agoza
VIBOT ERL CNRS 6000, ImViA, Uni e si é de Bo gogne
F anche-Com é
This disse a ion is submi ed o he deg ee o
Doc o o Philosophy
Dual P og am
Oc obe 2020

To my amily.
Acknowledgemen s
And I would like o acknowledge ...

Lis o igu es
1.1 Examples o algo i hms in his hesis. . . . . . . . . . . . . . . . . . . . 6
2.1 Lea ning s a egies wi h equi ec angula images . . . . . . . . . . . . . 25
2.2 Edgemapscompa ison........................... 26
2.3 S uc u allines ............................... 27
2.4 Roomsolu ion................................ 28
2.5 Layou hypo hesis gene a ion . . . . . . . . . . . . . . . . . . . . . . . 29
2.6 Re e encemaps............................... 31
2.7 Lines and VPs compa ison . . . . . . . . . . . . . . . . . . . . . . . . . 33
2.8 Combining geome y and deep lea ning . . . . . . . . . . . . . . . . . . 34
2.9 Quan i a i e compa ison wi h PanoCon ex . . . . . . . . . . . . . . . 36
2.10 Quali a i e compa ison wi h PanoCon ex . . . . . . . . . . . . . . . . 36
2.11Challenging esul .............................. 37
2.12Quali a i e esul s ............................. 37
2.13PanoRooma chi ec u e........................... 39
2.14 S uc u al lines and co ne s . . . . . . . . . . . . . . . . . . . . . . . . 41
2.15P edic ededgemaps ............................ 43
2.16Quali a i e esul s ............................. 44
2.17 EquiCon s pa ame iza ion . . . . . . . . . . . . . . . . . . . . . . . . 45
2.18 Modi ying he ield o iew and esolu ion in EquiCon s. . . . . . . . . 46
2.19Ke nelo se s ................................ 47
2.20EquiCon s.................................. 48
2.21CFLEquiCon s............................... 50
2.22 Layou om co ne p edic ions . . . . . . . . . . . . . . . . . . . . . . 50
2.23EquiCon spadding............................. 54
2.24Handlingocclusions............................. 55
2.25 Syn he ic images o obus ness analysis . . . . . . . . . . . . . . . . . 57
2.26Syn he ic es images............................ 58
xi Lis o igu es
2.27Quali a i e esul s ............................. 60
2.28 Quali a i e esul s on SUN360 . . . . . . . . . . . . . . . . . . . . . . . 62
2.29 Quali a i e esul so on SUN360 (non-cuboid) . . . . . . . . . . . . . . . 63
2.30 Quali a i e esul s on S an o d 2D-3D . . . . . . . . . . . . . . . . . . 64
3.1 Da ase c ea ion............................... 69
3.2 F ommask o3D.............................. 72
3.3 Quali a i e esul s o 3D objec s . . . . . . . . . . . . . . . . . . . . . . 73
3.4 Ins ance segmen a ion masks . . . . . . . . . . . . . . . . . . . . . . . . 77
3.5 E ec s o objec -layou combina ion . . . . . . . . . . . . . . . . . . . . 77
3.6 Quali a i e esul s o objec de ec ion and seman ic segmen a ion . . . 79
4.1 Coe icien s dis ibu ion . . . . . . . . . . . . . . . . . . . . . . . . . . 90
4.2 Ne wo ka chi ec u e............................ 90
4.3 Keypoin s co espondence/ epea abili y ac oss ins ances . . . . . . . . 95
4.4 Keypoin s co espondence/ epea abili y ac oss ins ances - all ca ego ies 96
4.5 Seman ic pa co espondence . . . . . . . . . . . . . . . . . . . . . . . 97
4.6 In a-ca ego y egis a ion . . . . . . . . . . . . . . . . . . . . . . . . . 99
4.7 Label ans e ................................ 99
4.8 Resul sin ealda a.............................100
4.9 Quali a i e esul s in ModelNe 10 da ase . . . . . . . . . . . . . . . . 101
4.10 Quali a i e esul s in ShapeNe pa s da ase . . . . . . . . . . . . . . 102
4.11 Quali a i e esul s in Dynamic FAUST da ase . . . . . . . . . . . . . . 102
4.12 Quali a i e esul s in Basel Face Model 2017 da ase . . . . . . . . . . 103
Lis o ables
2.1 E ec o di e en e e ence maps . . . . . . . . . . . . . . . . . . . . . 35
2.2 Quan i a i e esul s pe a ea and da ase . . . . . . . . . . . . . . . . . 35
2.3 FCNe alua ion............................... 43
2.4 Benchma king o SUN360 and S an o d 2D-3D . . . . . . . . . . . . . 44
2.5 Abla ions udy ............................... 54
2.6 Quan i a i e obus ness analysis . . . . . . . . . . . . . . . . . . . . . . 56
2.7 Layou esul s................................ 59
2.8 A e age compu ing ime pe image . . . . . . . . . . . . . . . . . . . . 60
3.1 E ec o ini ializa ion............................ 75
3.2 E ec o adap ing he CNN o pano amas . . . . . . . . . . . . . . . . 75
3.3 Ins ance segmen a ion esul s . . . . . . . . . . . . . . . . . . . . . . . 76
3.4 Objec de ec ion quan i a i e esul s . . . . . . . . . . . . . . . . . . . 76
3.5 Seman ic segmen a ion quan i a i e esul s . . . . . . . . . . . . . . . . 78
4.1 P ope iesAnalysis............................. 94
Chap e 1
In oduc ion
“Some imes science is mo e a han science. Lo o people don’ ge ha .”
— Rick Sanchez
1.1 Mo i a ion
Can we c ea e au onomous algo i hms ha unde s and he scenes as humans do? The
wo ld is made o objec s wi h a wild a ie y o shapes, appea ances and s uc u es. We
humans see he wo ld h ough he images o med by he ligh e lec ed om he objec s
in ou en i onmen . These images allow ou b ain o unde s and he shape and ex u e
o he objec s, c ucial o highe le el unde s andings. Mo eo e , we unde s and he
isual scene e o lessly, unde a wide a ie y o condi ions, and his is because human
pe cep ion eme ges om he gene ic code ueled by millions o yea s o e olu ion and,
a he same ime, om a li e ime o expe ience. This unde s anding is achie ed om
an speci ic iewpoin , om which he scene is obse ed. In o de o au oma ically
ep oduce his unde s anding wi h algo i hms, we need a came a o cap u e he ligh
o he scene a some loca ion and addi ionally, in elligen models ha eason abou
he isual inpu .
The human isual sys em pe cei es an ho izon al angle o iew o abou 140
◦
,
wi hou conside ing he eyes mo emen . Fo compa ison, he ho izon al ield o iew
(FoV) o con en ional came as anges om 40
◦
o 60
◦
. Mo eo e , human s e eopsis
allows a 3D pe cep ion ha may no be di ec ly achie ed om single 2D con en ional
images. Such educed ield o iew o he lack o dep h pe cep ion, c ucially limi s he
goal o de eloping in elligen sys ems o ma ch he pe o mance o human ision.

2In oduc ion
An inc easing demand o ex end he came as FoV, led o he appea ance o isheye
and ca adiop ic lenses, achie ing up o 360
◦
o ho izon al FoV. Today, 360
◦
images
can be easily ob ained wi h special lenses, bu also wi h came a a ays o au oma ic
image s i ching algo i hms [
18
]. Ul a wide-angle lenses ha e demons a ed o be
bene icial, pa icula ly in indoo scena ios, o many di e en asks including indoo
scene unde s anding [
172
] o isual odome y [
174
]. This is no su p ising, since a FoV
o 360
◦
allows o see he whole scene a once, gi ing a s ong con ex abou he space,
allows o ack ea u es longe , g ea ly inc easing he obus ness o isual localiza ion,
and allows ea u es o be mo e e enly dis ibu ed in space, which s abilizes pose
es ima ion. Addi ionally, wi h he apid de elopmen o 3D acquisi ion echnologies,
3D senso s a e becoming inc easingly a ailable and a o dable, including 3D scanne s,
LIDAR senso s, and RGB-D came as. Dep h pe cep ion signi ican ly con ibu es o
sol e se e al challenging asks ela ed o 3D objec shape analysis [
2
,
95
]. As an
example, when we look a an image o a 3D objec , we see only i s p ojec ion om a
speci ic iewpoin . The e o e, di e en iewpoin s may c ea e en i ely di e en ende s,
limi ing he use o 2D images in applica ions whe e shape in o ma ion is c i ical, o
when shape abs ac ion i sel is he scope o he s udy. The e o e, he choice o he
senso has a emendous impac on he obus ness and accu acy o he de eloped
models. We a e pa icula ly mo i a ed o see i scenes can be unde s ood be e beyond
he adi ional senso s, be ing o 360 images and 3D poin clouds. And, o be mo e
speci ic, we a e in e es ed on explo ing his unde s anding on indoo scena ios.
Why a e indoo scenes “special”? A ound 10.000 yea s ago, he e was a ime
o ansi ion om a hun e -ga he e mode o subsis ence o an ag icul u al way o
li e. This enabled humans o li e in mo e pe manen se lemen s, o he poin ha
nowadays, humans spend app oxima ely 90 pe cen o hei ime in in e io spaces
[
74
], which means mo e han six days pe week. This shocking ac ansla es in o an
u gen need o unde s and well indoo scena ios o imp o e li e quali y indoo s. We
he e o e a e conce ned abou gi ing machines he equi ed isual sensing mechanism
o unde s and he indoo en i onmen s whe e hey o en ope a e, o assis o en e ain.
This unde s anding is no i ial, as he e a e hund eds o di e en man-made sce-
na ios. To pu an example, he Places Da abase[
176
] is a eposi o y ha con ains 10
million scene pic u es, comp ising a la ge and di e se lis o he ypes o en i onmen s
encoun e ed in he wo ld, consis en wi h eal-wo ld equencies o occu ence. They
di ide he da ase in o 365 scene ca ego ies and classi y hem in o indoo s, ou doo s
and ou doo s man-made classes. As a esul , 159 ca ego ies ou o he 365 belong o
indoo scenes, 80 o ou doo s and 159 o ou doo s man-made (some o he ca ego ies
1.1 Mo i a ion 3
a e classi ied as ou doo s and ou doo s man-made a he same ime). Acco ding o
hese numbe s, a ound he 80%o he images aken a e om man-made scena ios.
Man-made scena ios can di e o many easons, mainly due o he use o which hey
a e designed o due o he ime o cul u al en i onmen in which hey a e buil . These
di e ences a e usually in e ms o he scenes layou o he ype o objec s we can ind
inside hem e.g. supe ma ke s and hea e s. E en be ween di e en scenes ha belong
o he same ca ego y we can see hese di e ences, as one does no ac he same way in
a home bed oom, a ho el bed oom o a nu se y. Achie ing a comple e unde s anding
o indoo scenes, would equip he discipline o compu e ision wi h many exci ing
ools, hus making i mo e powe ul and ubiqui ous. Bu he challenges ha need o
be sol ed a e nume ous. To s a , he layou can ange om e y simple i.e. 4 walls,
o highly complex layou s as in museums. To con inue, he di e si y o objec s o
in e es is e y high, many o which appea in equen ly. Indoo spaces usually con ain
many ins ances, gene a ing in he isual scenes clu e and high deg ee o occlusions,
which makes e y ha d o know he sepa abili y o objec s and su aces. No o say
ha objec s come in a ious shapes, sizes and in di e en poses. Addi ionally, while
some spaces can be ecognized by global spa ial p ope ies e.g. co ido s, o he s a e
be e cha ac e ized by he objec s inside e.g. books o es [
109
]. In ac , p oposed
solu ions o p oblems such as scene ecogni ion o seman ic segmen a ion [
109
,
7
] ha e
demons a ed a high pe o mance on ou doo scena ios, while pe o ming poo ly in
he indoo domain. This sugges s ha indoo scenes equi e special and dedica ed
algo i hms o hei unde s anding.
Sol ing he a o emen ioned challenges is no only s imula ing, bu also necessa y
o many exci ing applica ions. One clea example, ha is e olu ionizing he eal
es a e indus y, a e he i ual ou s o homes
1
, which a e helping selle s sell as e
while eeling mo e con iden , and buye s unde s and he home layou and imagine
wha i would be like o call i home. Such i ual ou s a e also becoming e y
popula o isi museums o a galle ies, in pa as a consequence o he COVID-19
lockdown ac oss he wo ld. In ac , he A s &Cul u e ini ia i e by Google cu en ly
o e s i ual isi s o abou 500 museums h oughou he wo ld, including he MoMA,
Ams e dam’s Rijksmuseum, he Na ional Galle y and he Palace o Ve sailles, o name
a ew. Indoo scene unde s anding is also i al o au onomous mobile obo s such us
acuum cleane s
2
, su eillance d ones, o assis i e obo s, ha need o mo e eely
in he same space as humans do, o e en ge o mo e in complex spaces whe e less
1h ps://www.zillow.com/ma ke ing/3d-home/
2h ps://www.i obo .es/ oomba/
4In oduc ion
exposu e o humans is desi ed, like buildings unde cons uc ion, mines o hospi al
a eas wi h con agious people. Au onomous mobile obo s can be also use ul o isual
da a collec ion o 3D modeling and domes ic obo s wi h cogni i e abili ies can be
specially help ul o ake ca e o isually impai ed people, which ep esen he 17% o
he wo ld’s popula ion [
99
], and elde ly people, whose numbe is expec ed o each
1.5 billion by he yea 2050 [
148
]. Ano he example, which is ge ing mo e and mo e
demanded, is i ual and augmen ed eali y o educa ion, games o in e io design
3
.
This echnology equi es a de ailed le el o unde s anding o he scene o a ious
easons, such as delimi ing a sa e a ea o he use o easoning abou he eal- i ual
objec s in e ac ion.
We s ongly belie e ha b inging oge he indoo scene unde s anding and non-
con en ional came as, such as 360 and 3D senso s, has many possibili ies. Some o
hem a e scene ecogni ion, s uc u e analysis such as oom layou econs uc ion o
loo plan es ima ion, objec de ec ion, objec -layou in e ac ion, shape analysis, objec
pose es ima ion, saliency p edic ion, e c. While all o hese p oblems a e exci ing, we
ocus in his hesis on some o hem, selec ed o hei ele ance, no only o he ask
i sel , bu also o hei po en ial use o bene i o o he ision asks. Mo e speci ically,
he con ibu ions o his hesis a e ocused on he hie a chical unde s anding o indoo
en i onmen s, ha we summa ize in h ee di e en le els o de ails:
•Layou le el: unde s anding o he main s uc u e o an indoo scena io.
•Scene le el: localiza ion o objec s and hei dis ibu ion in he 3D scene.
•Objec le el: geome y and shape modelling o la ge collec ion o objec s.
1.2 Con ex and Challenges
Compu e Vision is he science ha seeks o gi ing compu e s a ull h ee-dimensional
scene unde s anding om images, by emula ing he b ain’s abili y o make sense o
wha he eyes see. Is his idea ha challenging? In he ea ly 1960s, Seymou Pape ,
one o he pionee s o a i icial in elligence, did no hink so, and p oposed o a couple
o his s uden s o sol e ’Compu e Vision’ as a summe p ojec [
101
]. Howe e , oday
we a e s ill a away om seeing me hods ha achie e human-le el obus ness and
gene aliza ion. So, wha makes Compu e Vision so demanding and why do we ca e
abou i ?
3h ps://www.ikea.com/au/en/cus ome -se ice/mobile-apps/say-hej- o-ikea-place-pub1 8a 050
1.2 Con ex and Challenges 5
Visual pe cep ion is a e y complex piece o ou o ganic echnology. I no only
in ol es ou eyes and isual co ex, bu also akes in o accoun ou uncoun able pe sonal
expe iences and in e ac ions wi h he wo ld, as well as ou abs ac unde s anding
o concep s and men al models o objec s. In o de o unde s and how ou lea ning
p ocess wo ks, ea ly s udies analyze he p inciples o objec pe cep ion wi h human
in an s [
134
]. Unde s anding how we lea n in ou ea lies s age o li e, gi es he hin s as
o how we ha e o design lea ning algo i hms. Mode n Compu e Vision models aim a
ep oducing how ou b ain shapes all he inpu s we ecei e, om he simples ea u es
o he mos de ailed unde s anding, la gely by obse a ion o he geome y o he wo ld.
I is equally impo an o unde s and how posi i e expe iences and ailu es help o
he lea ning p ocess, as human in an s lea n h ough a mix u e o semi-supe ised and
unsupe ised lea ning.
The ea ly op imism o Seymou Pape led o huge imp o emen s in compu e ision
models, and also in he capabili ies o he compu e s ha un hem. The 1980’s saw
he backp opaga ion algo i hm o neu al ne wo ks being laid ou by Geo ey Hin on
[
117
] and en yea s la e , Yann LeCun and Yoshua Bengio among o he s, p oposed he
i s con olu ional neu al ne wo k (CNN) a chi ec u e [
79
]. The apid ad ancemen s o
Machine Lea ning and Deep Lea ning echniques [
77
] b ough u he li e o he ield
2012 onwa ds, whe e backp opaga ion and CNNs became ubiqui ous in AI. And now,
we li e in an exci ing ime o compu e ision, since we a e p oducing mo e isual da a
han e e be o e, and we coun on powe ul algo i hms o p ocess i . E en i he ield has
been able o ake g ea leaps in ecen yea s, and e en o su pass humans in some asks
[
128
,
58
], signi ican ly mo e e o s a e equi ed o achie e uly au onomous sys ems
ha enable complex asks, like home obo ics o au onomous d i ing. These complex
asks equi e a deep unde s anding o 3D scenes, ac oss mul iple le els, connec ing
ision, g aphics and obo ics esea ch.
This hesis shows how o combine geome y and deep lea ning echniques o a
hie a chical unde s anding o indoo en i onmen s. The p oposed me hods ad ance
s a e-o - he-a a h ee di e en le els ha a e de ailed below. An o e iew o ou
esul s is gi en in Figu e 1.1.
Layou le el.
Wha is he con igu a ion o his oom? how much space do I ha e?
These a e he i s ques ions ha need o be answe ed when we a i e o a new space.
In o de o answe hem, we need o know he scene s uc u e. The scene s uc u e
can be simply de ined by a se o geome ic p imi i es. They can be a se o planes,
co esponding o he walls, ceiling and loo . They can be de ined by lines, ep esen ing
12 In oduc ion
Associa ed publica ion:
•
“Unsupe ised Lea ning o Ca ego y-Speci ic Symme ic 3D Keypoin s om
Poin Se s”
Cla a Fe nández Lab ado
, Ajad Chha kuli, Danda Pani Paudel, José J.
Gue e o, Céd ic Demonceaux, Luc Van Gool.
ECCV, 2020. Glasgow, Sco land.
1.3.2 Open-Sou ce So wa e/ Da ase s
We ha e eleased he ollowing open-sou ce so wa e and da ase s:
•
Sou ce code and da ase o he wo k “Co ne s o Layou : End- o-End Layou
Reco e y om 360 Images” can be downloaded om he p ojec websi e: h ps:
//c e nandezlab.gi hub.io/CFL/
•
360 Scene Unde s anding. I c ea ed he Gi Hub eposi o y 360 Scene Unde -
s anding, whe e di e en ools o wo k wi h 360 images, de eloped du ing my
PhD, a e a ailable: h ps://gi hub.com/c e nandezlab/360-Scene-Unde s anding.
•
The ex ended da ase o he wo k “Wha ’s in my Room? Objec Recogni ion
on Indoo Pano amic Images” can be downloaded om he p ojec websi e:
h ps://webdiis.uniza .es/~jgue e / oom_OR/.
•
Sou ce code o “Unsupe ised Lea ning o Ca ego y-Speci ic Symme ic 3D
Keypoin s om Poin Se s” can be downloaded om gi hub: h ps://gi hub.com/
c e nandezlab/Ca ego y-Speci ic-Keypoin s.
1.3.3 Resea ch S ays
Du ing my PhD I did he ollowing esea ch s ays ab oad:
Compu e Vision Labo a o y, ETH Zu ich.
F om Augus 2019, I spen 7 won-
de ul mon hs as a isi ing esea che in he Compu e Vision Labo a o y a ETH
Zu ich lead by P o esso Luc Van Gool, whe e I enjoyed wo king oge he wi h D .
Ajad Chha kuli and D . Danda Pani Paudel.
Addi ionally o ou ECCV submission [
40
], we wo ked on he unde s anding o
he Indus y Founda ion Classes (IFC) da a model and on he de elopmen o ans-
la o ools be ween IFC and poin clouds. The IFC schema is a s anda dized (ISO

1.4 Ou line 13
16739-1:2018) da a model in ended o desc ibe building indus y da a. The schema
inco po a es no only 3D geome y bu also all he ele an da a ela ing o he building,
i s componen s and he p ojec schedules. We also e alua ed objec de ec ion pipelines
on 3D poin clouds gene a ed om IFC models. This pa o he wo k is no p esen ed
in he hesis.
Disney Resea ch S udios.
Du ing he summe o 2020 I did a 3 mon hs in e nship
a Disney Resea ch S udios in Zu ich, Swi ze land. Du ing his s ay, I had he
oppo uni y o wo k unde he supe ision o D . Hayko Riemenschneide on 3D shape
co espondences, a he in e sec ion be ween compu e ision and compu e g aphics.
1.3.4 Supe ision o S uden s
Du ing my PhD I also had se e al en iching and ewa ding ad ising expe iences.
•
Juan Ca los Medina (Bachelo Thesis in Indus ial Enginee ing a he Uni e si y
o Za agoza – 2018): “Single View Layou Recons uc ion”.
•
Julia Gue e o Campo (Bachelo Thesis in Compu e Science a he Uni e si y
o Za agoza – 2019): “Objec Recogni ion in 360 Images”.
1.4 Ou line
This disse a ion is di ided in o he ollowing chap e s:
Chap e 2
answe s many ques ions ega ding he 3D layou es ima ion om single
iew p oblem and how o le e age 360 images. This chap e guides h ough h ee
app oaches showing an e olu ion o ou esea ch in his ield.
Chap e 3
in oduces and p oposes solu ions o he p oblem o objec de ec ion using
360 images. Addi ionally, we explo e how o le e age objec -layou in e ac ion o
place he de ec ed objec s inside he 3D oom model.
Chap e 4
explo es how o au oma ically disco e 3D keypoin s om a collec ion o
objec s o he same ca ego y, so ha hey a e co esponden . Fo he i s ime,
we p opose o do so using misaligned 3D poin clouds, including he no ion o
symme y and in an unsupe ised manne .
Chap e 5
gi es he inal discussion o his hesis and ideas o u u e wo k on he
p esen ed p oblems.
Chap e 2
Room Layou Es ima ion
“The e is geome y in he humming o he s ings, he e is music in he spacing o he
sphe es.”
— Py hago as
This chap e guides h ough h ee app oaches showing an e olu ion o ou esea ch
in he ask o oom layou es ima ion. The p oblem o 3D layou eco e y in indoo
scenes has been a co e esea ch opic o o e a decade. Howe e , he e a e s ill se e al
majo challenges ha emain unsol ed. Among he mos ele an ones, a majo pa o
he s a e-o - he-a me hods make implici o explici assump ions on he scenes –e.g.
box-shaped o Manha an layou s. Also, cu en me hods a e compu a ionally expensi e
and no sui able o eal- ime applica ions like obo na iga ion and AR/VR. A he end
o his chap e , we will end up wi h a as geome ic deep lea ning model, lexible and
obus o came a pose, ha gene alizes o cuboid and non-cuboid layou s om single
360 images.
Inpu Image Room Layou Room Layou
co ne /edge ep esen a ion segmen a ion ep esen a ion
3D Model
16 Room Layou Es ima ion
2.1 In oduc ion
Room layou es ima ion aims a inding he 3D box ha bes i s an indoo scene
ega dless o unca ions o occlusions, whe e he bounda ies o he 3D box a e he
in e sec ions be ween walls, ceiling and loo .
We a e pa icula ly exci ed o sol e he layou es ima ion p oblem due o i s u ili y
o many eal applica ions and o o he compu e ision asks. 3D scene modeling
is a key echnology in se e al eme ging applica ion ma ke s, such as augmen ed and
i ual eali y, in elligen obo na iga ion o na iga ional aid o isually impai ed
people [
121
]. And also o mo e adi ional ones, like eal es a e [
85
]. Fo he o me
applica ions, i is highly use ul o ha e a p ecise easoning abou he a ailable space,
e.g. whe e he subjec can mo e. Fo he la e , e e y ime mo e and mo e companies
op o democ a izing his echnology o deli e alue o hei use s. Knowing he
oom layou , also p o ides a s ong p io o o he isual asks like dep h eco e y
[
36
,
23
], ealis ic inse ions o i ual objec s in o indoo images [
72
], indoo objec
ecogni ion [
8
,
132
], indoo place ecogni ion [
65
], human pose es ima ion [
46
] o scene
econs uc ion/ ende ing [66].
A la ge a ie y o me hods ha e been de eloped o es ima e oom models using
mul iple inpu images [
146
,
44
] o dep h senso s [
170
], which deli e high-quali y
econs uc ion esul s. Fo he common case when a single RGB image is a ailable, he
p oblem becomes conside ably mo e challenging. In ac , in e ing 3D in o ma ion om
a single monocula image is one o he holy g ails o compu e ision. The p oblem
i sel is ill-posed: dep h is i eco e ably los . Howe e , p io knowledge abou he scene
geome y and seman ics can help esol e some o he ambigui ies. The Manha an
wo ld assump ion, p oposed by Coughlan and Yuille [
24
], has led he as majo i y o
wo k on monocula layou es ima ion, whe eby indoo scenes a e h ee-dimensional
g ids. Consequen ly, all walls a e a igh angles o each o he and pe pendicula o he
loo and ceiling planes. This assump ion allows o eco e bo h ex insic and in insic
came a pa ame e s, as well as o ex ac he 3D s uc u e o he oom om a single
image.
A c ucial limi a ion o p e ious wo ks [
31
,
81
,
60
,
61
,
125
,
80
,
92
,
175
] lies in he use
o con en ional images wi h limi ed ield o iew (FoV). On he one hand, his p e en s
he econs uc ion o he eal closed geome y o he whole oom. This limi a ion leads
o he o e -simpli ica ion o he oom ypes, e.g. 4-wall layou s, o en unde i ing
he ichness o eal indoo spaces. On he o he hand, he ceiling does no usually
appea in con en ional images, being ne e heless an impo an pa o de ec he
main s uc u e o he oom, as i usually has much less occlusions. In his ega d,
2.1 In oduc ion 17
he app oxima e ho izon al FoV o he human ision sys em is almos 180
◦
, wi hou
conside ing eyes mo emen . Howe e , he FoV o a s anda d came a is much smalle ,
only a ound 50
◦
. E en he new sma phones only ha e an ho izon al FoV o a ound
70
◦
. This c ucially limi s he use o con ex cues o unde s and he su ounding scene.
Fo example, we expec a bed oom o ha e a leas one bed o a ki chen o ha e a
idge. Howe e , i we use a came a wi h a small FoV, depending on he di ec ion he
came a looks a , he e migh no be a bed o a idge, while he e migh be a able
ha does no gi e us much in o ma ion abou he scene. Since he goal o compu e
ision is o mimic he human isual pe cep ion, i is no ai o ask compu e ision
algo i hms o ma ch he pe o mance o human ision wi h such educed FoV.
The e o e, a mo e ecen esea ch di ec ion looks o ex end he FoV. Lopez-Nicolas
e al. in [
88
] pe o m he layou eco e y using a ca adiop ic sys em. In [
105
], layou
hypo heses a e made combining isheye images wi h dep h in o ma ion ha p o ides
scale. Bu he eal impac comes wi h he 360
◦
images, which nowadays can be easily
ob ained wi h came a a ays, special lenses o au oma ic image s i ching algo i hms [
18
].
Pano amic images ha e b oken he ba ie s o pe o mance on his ask. These images
allows o acqui e he whole scene a once and hence, i is possible o exploi hei wide
FoV o gene a e closed oom solu ions based on he bes consensus dis ibu ed a ound
he scene. [
69
] shows he ad an ages o ha ing a comple e scene iew o e pa ial
iews o he same scene [81]. Howe e , he me hods o con en ional came as a e no
sui able o wide FoV images, due o he image dis o ions. This limi a ion becomes a
majo bo leneck in some ecen wo ks ha use 360 images [
172
,
159
,
161
], as ex a
wo k is needed o le e age con en ional algo i hms.
To summa ize, he p oblem o es ima ing he 3D oom layou om a single iew is
no i ial, as i is an ill-posed p oblem. Addi ionally, he e a e se e al majo challenges
ha s ill emain unsol ed. Mos exis ing me hods use con en ional images wi h educed
FoV, ha p e en s gene a ing closed oom solu ions. Such educed FoV also leads o
oom simpli ica ions, e.g. simple 4-wall cuboids. Mo eo e , s a e o he a me hods
a e s ill a om being sol ed in eal- ime, as expensi e p e- and/o pos -p ocessing
s eps a e needed. Addi ionally, i we wan o le e age wide FoV images, adi ional
me hods a e no sui able and new ones ha e o be de eloped.
In his chap e we p o ide a 3D unde s anding o he oom layou beyond he
ield o iew, p oposing se e al app oaches ha a e p esen ed in Sec ions 2.4,2.5 and
2.6. Ou goal is h ee old: i) p o ide ai h ul scene geome y p edic ions, wi h he
mo i a ion o lea e behind he 4-wall oom simpli ica ion, ii) c ea e as e me hods,

18 Room Layou Es ima ion
a oiding expensi e p e-pos p ocessing s ages and iii) p opose e ec i e ways o le e age
he ad an ages o e ed by 360 images.
In Sec ion 2.4, we p esen ou i s wo k in his di ec ion. We p opose a model o
eco e ing oom models, gene alizing o cuboid and non-cuboid layou s. The me hod
combines adi ional geome ic easoning and deep lea ning echniques o ge al eady
po en ial s uc u al lines and co ne s, om which he layou hypo heses a e gene a ed.
Wo king di ec ly wi h po en ial s uc u al p imi i es, leads o a meaning ul educ ion o
he numbe o hypo heses needed and consequen ly, o a educ ion in he pos -p ocesing
compu a ion ime. An impo an con ibu ion o gene alizing o non-cuboid layou s is
p esen ed in he hypo heses gene a ion p ocess. We obse ed ha bo h con en ional
and deep lea ning algo i hms, s uggle o de ec s uc u al co ne s ha a e non- isible
in he image due o occlusions. This is in ac a e y common scena io whe e specially
loo co ne s a e occluded by he objec s in he scene. Addi ionally, depending on he
came a iewpoin and he complexi y o he scene, some walls may occlude en i e pa s
o he oom. To sol e his p oblem, we allow he model o add ex a co ne s du ing
he hypo heses gene a ion, so ha ooms meaning ully sa is y he Manha an Wo ld
assump ion.
Sec ion 2.5 p esen s ou second wo k. In his sec ion, we p esen a deep lea ning
model o es ima e s uc u al lines and co ne s di ec ly on pano amic images. The
ad an age is wo old. Fi s , i allows a oiding expensi e p e-p ocessing algo i hms
o le e age me hods designed o con en ional images. Second, wo king di ec ly on
he pano ama allows o ake ad an age o all he con ex . The e o e, he p esen ed
amewo k p o ides as e and mo e accu a e oom solu ions.
The inal me hod is p esen ed in Sec ion 2.6. We make he obse a ion ha he use
o s anda d con olu ions in equi ec angula images can lead o a loss o pe o mance. We
p opose a no el con olu ion ha adap s he ke nel size and shape o he equi ec angula
image dis o ions, b inging nume ous ad an ages. We demons a e how he ob ained
p edic ions a e no only mo e accu a e, bu also mo e obus o came a pose a ia ions.
The pe o mance imp o emen allows us o p edic layou s in an end- o-end manne ,
and elax he scene assump ions. P edic ing he co ne s in an end- o-end ashion,
makes ou me hod up o 100 imes as e han p e ious app oaches.
2.2 Rela ed Wo k
The seminal monocula app oach o au oma ically eco e 3D econs uc ions was [
31
],
which shows how p io knowledge abou indoo scenes, i.e. loo -wall bounda ies, can
2.2 Rela ed Wo k 19
be lea ned using a dynamic Bayesian ne wo k. In pa allel, Lee e al. [
81
] gene a e layou
hypo heses om de ec ed line segmen s, and selec he bes - i ing one e alua ing wi h
an O ien a ion map. While e ec i e, O ien a ion maps ge limi ed wi h he p esence
o clu e , since no easoning abou he lines is made. Mo i a ed by he p oblem o he
p esence o clu e , [
60
] models he layou o he oom wi h an aligned 3D box while
localizing isible objec s. This inspi ing idea was ollowed by [
61
,
125
]. Howe e , he
3D box simpli ica ion does no ma ch eali y in many cases, being hence cons ained o
his pa icula oom geome y and unable o gene alize o o he oom con igu a ions.
Typically, hese me hods ollow a p oposing- anking scheme and ely on Geome ic
Con ex [
63
] o e alua e. Geome ic Con ex imp o es clu e de ec ion compa ed wi h
he O ien a ion maps, bu p o ides wo se esul s a he highe pa s o he scenes.
Mo e ecen ly, [
126
] in oduces he concep o in eg al geome y and pai wise po en ials
decomposi ion which esul s in an e icien s uc u ed p edic ion amewo k.
Since 2012, CNNs achie ed b eak h ough pe o mance in a wide ange o applica ions
such as image classi ica ion [
77
], segmen a ion [
7
], de ec ion [
110
] op ical low [
138
]
and keypoin de ec ion [
122
]. This unp eceden ed le el o da a abs ac ion inspi ed by
neu onal p ocesses, became popula in all a eas o Compu e Vision, including ha o
es ima ing he layou o ooms. Mallya e al. [
92
] ain a Fully con olu ional Ne wo k
(FCN) o join ly p edic in o ma i e edges and geome ic con ex om con en ional
images. The pixel-wise edge labeling dis inguishes be ween backg ound, wall- loo edge,
wall-wall edge and wall-ceiling edge. Mo e ecen ly, o he wo ks ocus on pixel-wise
edge labeling. [
114
,
171
] add ess he p oblem in a coa se- o- ine manne . Ins ead, he
p oposal o [
175
] is inspi ed by mechanics concep s. Al e na i ely, Dasgup a e al. [
30
]
p opose a FCN o p edic seman ic su ace labels o he ooms, p o iding sepa a e
belie maps o he walls, ceiling and loo o he scene. All hese me hods equi e ex a
compu a ion added o he o wa d p opaga ion o he ne wo k o e ie e he ac ual
layou . In [
80
], an end- o-end ne wo k p edic s he layou co ne s in a pe spec i e
image, as well as a label ha indica es which co ne s a e isible. A e wa ds, he oom
ype is in e ed wi hin a limi ed se o manually chosen con igu a ions. O he deep
lea ning wo ks ex ac an es ima ion o he dep h o /and su ace no mals om simple
RGB images, which also p oduces an in e es ing ou come o he p oblem o oom
layou es ima ion [
36
,
78
]. The main d awback o hese CNNs is ha hey a e designed
o wo k on con en ional images wi h limi ed FoV, wi h he a o emen ioned consequen
limi a ions.
While layou eco e y om con en ional images has p og essed apidly wi h bo h
con en ional me hods and deep lea ning, he wo ks ha add ess hese challenges using
20 Room Layou Es ima ion
omnidi ec ional images a e s ill e y ew in compa ison. Omnidi ec ional came as ha e
he po en ial o imp o e he pe o mance o he ask: hei 360
◦
ield o iew cap u es
he en i e iewing sphe e su ounding i s op ical cen e , allowing o acqui e he whole
oom a once and hence o p edic layou s wi h mo e isual in o ma ion. PanoCon ex
[
172
] was he i s wo k ha ex ended he amewo ks designed o pe spec i e images
o pano amas. I eco e s he oom layou , which is also assumed as a simple 3D box,
and bounding boxes o he mos salien objec s inside he oom. Pano2CAD [
159
]
ex ends he me hod o non-cuboid ooms, bu i is limi ed by i s dependence on he
ou pu o objec de ec o s. In [
161
] hey ea he p oblem as a g aph wi h lines and
supe pixels as nodes, sol ing i wi h complex geome ic cons ain s ins ead. The mos
ecen wo ks along his line a e con empo a y o he las app oach o his chap e .
Layou Ne [
177
] ains a FCN om pano amas and anishing lines, gene a ing he
layou models om edge and co ne maps, and DuLa-Ne [
163
] p edic s Manha an-
wo ld layou s le e aging a pe spec i e ceiling- iew o he oom. All o hese app oaches
equi e p e- o pos -p ocessing s eps like line and anishing poin ex ac ion o oom
model i ing, ha inc ease hei cos .
In he las wo yea s, he main imp o emen s in layou eco e y om pano amas
ha e come om he applica ion o deep lea ning. The high-le el ea u es lea ned by deep
ne wo ks ha e p o en o be as use ul o his p oblem as o many o he s. Ne e heless,
hese echniques en ail o he p oblems such as he lack o da a o o e i ing. In his
ega d, s a e-o - he-a me hods equi e addi ional p e- and/o pos -p ocessing. As a
consequence hey a e slow, and his is a majo d awback conside ing he a o emen ioned
applica ions o eal- ime layou eco e y.
In addi ion o all he challenges men ioned abo e, we also no ice ha he e is
an incong uence be ween pano amic images and con en ional CNNs. The space-
a ying dis o ions caused by he equi ec angula ep esen a ion makes he ansla ional
weigh sha ing ine ec i e. Ve y ecen ly, Cohen e .al. [
22
] did a ele an heo e ical
con ibu ion by s udying con olu ions on he sphe e using spec al analysis. Howe e ,
i is no clea ly demons a ed whe he Sphe ical CNNs can each he same accu acy
and e iciency on equi ec angula images. A ela ed wo k [
141
] p oposes dis o ion-
awa e con olu ional il e s o sol e dense p edic ion asks such as dep h p edic ion and
seman ic segmen a ion by le e aging commonly used da ase s wi h anno a ions o
pe spec i e images du ing aining.
In he ollowing sec ions we p esen he h ee di e en app oaches p oposed o he
oom layou es ima ion p oblem.
2.3 Backg ound and Theo y 21
2.3 Backg ound and Theo y
Pano ama Geome y.
A big pa o he solu ions p oposed in his hesis o he
p oblem o 3D indoo scene unde s anding use pano amic images. The e o e, we s a
explaining he basics o he sphe ical came a geome y, which will help us p og ess
smoo hly o he ac ual solu ions. Fo con enience, we will use he e ms equi ec angula
image, 360 image, sphe ical image and pano amic image in e changeably.
We can de ine a cen al came a as a collec ion o ays passing h ough a single poin
in a space, which is he came a cen e . Fo he pa icula case o a sphe ical came a
model, i consis s o a came a cen e ed inside a su ace o a uni sphe e.
How do we i he su ace o he uni sphe e on o a single image? Acco ding o
he Gauss’s Theo ema Eg egium, he Gaussian cu a u e o an embedded smoo h
su ace in
R3
is in a ian unde he local isome ies. Since he sphe e o adius
has
cons an posi i e cu a u e 1
/ 2
and a la plane has ze o cons an cu a u e, hese
wo su aces a e no isome ic. This means ha a piece o pape canno be ben on o a
sphe e wi hou c umpling and con e sely, he sphe e su ace canno be un olded on o
a plane wi hou dis o ing. Thus, all plana p ojec ions o a sphe e ha e dis o ions.
Among all he possible plana p ojec ions o he sphe e, he Equi ec angula p ojec ion
is usually p e e ed in compu e ision as i p ese es dis ances be ween poin s i.e. i
is equidis an , meaning ha he image g id can be indexed di ec ly wi h sphe ical
coo dina es. This is because Equi ec angula p ojec ion maps me idians and pa allels
o he sphe e o e ical and ho izon al s aigh lines o cons an spacing espec i ely.
The p ojec ion howe e , is nei he equia eal no con o mal. This ine i ably gene a es
some dis o ions ha a e mo e p onounced nea he poles, whe e he a eas ge s e ched
ho izon ally o he en i e wid h o he image, i.e. he en i e op edge co esponds o a
single poin , as does he lowe edge. Fu he , he le and he igh edges o he image,
a e he same spo in eali y, loosing he con inui y o he scene in he image.
Le ’s deno e he esolu ion o he equi ec angula image o be
W×H
pixels.
Because he sphe ical images co e s 360
◦
ield o iew ho izon ally and 180
◦
ield o
iew e ically, we know ha
W
= 2
H
, and he ocal leng h is
W
2π
pixels. To ake hese
images, he came a is ypically posi ioned so ha he op o he p ojec ion sphe e is
poin ing o he sky. The e o e, we can sa ely assume ha he ho izon al anishing
line o he g ound plane is a 0 heigh o he image coo dina e [
−W
2,W
2
]
×
[
−H
2,H
2
].
O he wise, we can pe o m an up igh -alignmen .
We o mula e he ela ion be ween a poin in a space, a poin on he uni sphe e
su ace and a poin in he equi alen image plane.
The i s s ep consis s o p ojec ing he scene poin Xon o he uni sphe e; he e o e:
28 Room Layou Es ima ion
de elop a me hod o gene a e layou hypo heses om co ne s, using he p edic ed
subse o lines, ha a e al eady po en ial s uc u al lines.
c1c2c3
c4
c1
c2c3
h
y
x
1
2 3
4
y
x
z
c1
c2c3
c4
?
c4
?
z
1..4 quad an di ision
ho izon al VP
candida e co ne s
bad loo plane es ima ions
ceiling e e ence plane
bes i ing loo plane
~90
p'
y
p'
x
p
y
p
x
ho izon line
Fig. 2.4
Room solu ion
. Ou algo i hm e u ns a solu ion such ha he walls con-
nec ed by he co ne s a e as pe pendicula as possible (g een) ollowing he Manha an
wo ld assump ion. See how non-Manha an solu ions (pink and pu ple) esul in loo
planes ha a e no pa allel o he ceiling plane (yellow). The solu ion also p o ide us
an es ima ed oom heigh o he layou hypo hesis.
Candida e co ne s ex ac ion
Ou layou gene a ion p ocess is based on co ne s, i.e. s uc u al in e sec ions be ween
wo walls and ceiling o loo . In a Manha an Wo ld, wo line segmen s a e enough
o de ine a co ne . We in e sec he p edic ed lines among hemsel es in pai s, as
long as hey do no c oss each o he and hey ha e di e en di ec ions (
x, y, z
). The
di ec ion ec o o he co ne is compu ed using he lines in e sec ing a ha co ne ,
cab
= (
na×nb
). Thanks o he p e ious line il e ing s ep, he ex ac ed co ne s a e
al eady good candida es o gene a e oom layou hypo heses. Figu e 2.3 shows he
la ge di e ence be ween ob aining co ne s wi h he o iginal se o lines (le ) and
wi h he subse o s uc u al lines ( igh ). By emo ing non-s uc u al lines, he
numbe o co ne s ge s as ly educed, ye he impo an ones emain de ec ed. This
educ ion makes u he s ages o he me hod as e and mo e e icien while imp o ing
he eliabili y o he esul s, since mos co ne candida es coming om clu e and
i ele an s uc u es a e no conside ed.
Pano amic images ha e he ad an age o p o iding a ull iew o he oom, allowing
us o look a ound, up and down in he scene. This is unlike he con en ional images,
whe e he ceiling and some walls use o be ou o he FoV. Taking his in o accoun ,
we ca y ou a classi ica ion o he de ec ed co ne s ollowing wo c i e ia:

2.4 Layou s wi h Geome y and Deep Lea ning 29
a) Ve ical di ec ion.
Co ne s de ec ed below he ho izon line (
lH
), which in
cen al pano amas is a he middle ow, a e conside ed as loo co ne candida es
and hose de ec ed abo e, a e conside ed as ceiling co ne candida es.
b) XY -plane.
We di ide he scene in o ou quad an s a ound he came a cen e
using he ho izon al VPs as quad an di ide s,
Q
=
{q1, q2, q3, q4}
. Hence, e.g.
o de e mine when a co ne belongs o he ou h quad an :
c∈q4⇐⇒ cx∈
R+∧cy∈R−.
See Figu es 2.4 and 2.5 o mo e de ails abou he ep esen a ion.
c1
c2c3
1
2 3
4
c4
c1
c2c3
y
x
1
2 3
4
c1
c2
c3c4
c4
c5
c5
c6
c6
1 2 3 4
~90
1 2 3 4
c1
c2
c3c4
y
x
Fig. 2.5
Layou hypo hesis gene a ion
: We show wo examples o layou hypo heses
gene a ion. The i s example co esponds o a alid hypo hesis whe eas he second
one p esen s a non- alid disca ded hypo hesis.
Layou hypo heses gene a ion
Many wo ks simpli y he oom layou o ha e only ou walls. Usually, his simpli ica ion
comes om he lack o con ex ual in o ma ion when con en ional images a e used
[
60
,
61
,
125
]. Howe e , mo e ecen wo ks using 360 images also adop his simpli ica ion
[
172
]. He e, we handle mo e complex designs which will be ai h ul o he ac ual shapes
30 Room Layou Es ima ion
o he ooms, in oducing he possibili y o es ima ing in-be ween hidden co ne s when
equi ed, i.e. when hey a e occluded by clu e o due o scene non con exi y. We
gene a e layou hypo heses by means o an i e a i e me hod ha a emp s o join
consecu i e co ne s wi h al e na i ely o ien ed walls. We assume he ollowing:
a) Manha an wo ld.
The e a e h ee main o hogonal di ec ions o each o he
ha de ine he indoo scene.
b) Ceiling- loo pa allelism.
Floo co ne s a e on he same loo plane and ceiling
co ne s a e di ec ly abo e he loo ones. The no mal di ec ion o bo h planes is
he e ical anishing di ec ion pz.
c) Came a heigh .
Since no dep h in o ma ion is a ailable, we need o assume he
dis ance om he came a o he loo o ceiling planes. This is i ial as esul s
a e up o scale, bu needed o p edic he o al heigh o he oom. Gene ally, he
dis ance o he loo is assumed. We obse ed ha he p edic ed ceiling co ne s
a e mo e eliable, being in a less clu e ed a ea, and we assume he dis ance o
he ceiling plane ins ead.
The p oposed algo i hm andomly samples a g oup o co ne s among he p edic ed
ones,
Gc
, which a e o de ed clockwise in he
XY
-plane. The numbe o sampled
co ne s
NGc
may a y a each i e a ion and can be di ec ly ela ed o he maximum
numbe o walls ha ou algo i hm can handle,
Nmax
W
= 2(
NGc−
1). Fo example,
we can d aw oom layou s wi h six walls om a minimum numbe o ou co ne s,
allowing he algo i hm o in oduce wo new co ne s ha may be occluded in he
image. Addi ionally, we obse e ha Manha an Wo ld ooms always ha e an e en
numbe o walls and an odd numbe o co ne s a each quad an . As an example,
a simple layou has only one co ne pe quad an , while mo e complex layou s may
ha e h ee o e en i e co ne s a some o hei quad an s. The e o e, he p oposed
quad an di ision p o ides a con enien way o sample co ne s. The co ne sampling
mus include co ne s om a leas h ee quad an s, so ha he co ne in he emaining
quad an can be es ima ed assuming closed Manha an layou s, and he e mus be a
leas one co ne o each hemisphe e, so ha he o al oom heigh can be p edic ed.
We p oceed wi h he geome ic easoning in 2D wi h a op iew o he 3D scene,
see igh side o Figu e 2.4 o 2.5. No e ha we do no ha e he eal 3D coo dina es
o he co ne s, bu only he 3D ay ha goes om he cen e o he sphe ical image
h ough he co ne posi ion on he su ace o he sphe e.
2.4 Layou s wi h Geome y and Deep Lea ning 31
We use Figu e 2.4 o desc ibe how ou hypo heses gene a ion algo i hm wo ks.
Fi s , he 3D ays o he sampled ceiling co ne s a e in e sec ed in o a e e ence ceiling
plane a an assumed dis ance, ob aining he po en ial 3D ceiling co ne s c1, c2and c3
(yellow). We keep he 3D ay o he sampled loo co ne (cyan), as he dis ance o he
loo is ye unknown. Then, we use he Manha an wo ld equi emen o es ima e he
co ec loo co ne
c4
posi ion along i s 3D ay, so ha he walls connec ed by he
co ne s a e as pe pendicula as possible (90
◦±
5
◦
). This p ocess e u ns a Manha an
Wo ld solu ion o he oom ha also allows us o compu e he complemen a y dis ance
o he loo plane ha e i ies he ay equa ion.
In Figu e 2.5 we show a complex layou and wo possible layou hypo heses de-
pending on he ini ial sampled co ne s. We i s show a alid hypo hesis ( op), wi h
sampled co ne s
Gc
=
{c1, c2, c3, c5}
. This means ha he algo i hm will be able o
sol e a layou hypo hesis wi h
Nmax
W
= 6. A e p ojec ing he ceiling co ne s in o a
e e ence plane, a joining co ne p ocess s a s om
c1
. As be o e, we ind he op imal
loo co ne posi ion along i s ay using i s nea es co ne s and d aw an in e media e
solu ion,
c2
. In he hi d quad an , aking in o accoun he di ec ion (
x−y
) om
p e ious unions, ou algo i hm selec s he bes solu ion o
c4
by choosing he one which
p oduces al e na i ely o ien ed consecu i e walls. In he emp y quad an , Manha an
walls om nea es co ne s gi e
c6
. We also show a non- alid hypo hesis (bo om),
wi h sampled co ne s
Gc
=
{c1, c2, c3, c4}
. Following he same p ocess, he co ne s a e
o de ly joined esul ing in his case a non-Manha an layou .
(a) Layou hypo he-
sis (IHi)
(b) No mals Map
(INM )
(c) O ien a ion Map
(IOM )
(d) Geome ic Con-
ex (IGC )
(e) Me ge Map
(IMM )
Fig. 2.6 (a) Example o labeled image gene a ed om layou hypo heses. (b)-(e) Visual
ep esen a ion o how each o he e e ence maps IR, looks like.
Layou hypo heses e alua ion
We e alua e a numbe o layou hypo heses
Nh
o ge he bes and inal oom layou
solu ion. Fo each hypo hesis
Hi
, we gene a e a segmen ed image
IHi
, encoding he
o ien a ion o he p edic ed su aces, i.e. walls in
x
, walls in
y
and loo /ceiling in
z
.
32 Room Layou Es ima ion
In Figu e 2.6 (a) he e is an example o a segmen ed image, whe e each o ien a ion is
encoded wi h a di e en colo .
We e alua e he segmen ed image
IHi
by compa ing i o a e e ence map
IR
ha oughly encodes he o ien a ion o he pixels, and can be ob ained om se e al
me hods. We compu e he a io o pixels ha a e equally o ien ed in
IHi
and
IR
o e
he o al size o he image (
H, W
), ha we name Equally O ien ed Pixel a io (
EOP
):
EOP IHi,IR=1
H·W
C
X
x,y,z
H,W
X
i,j
IHi&IR,
being C he numbe o channels co esponding o he labels i.e. o ien a ions x,y,z.
In his wo k, we es ou me hods o compu e he e e ence map
IR
. The ou
me hods a e designed o con en ional images so we epea he same p ocess as in
Sec ion 2.4.1 o compu e hem. The O ien a ion Map [
81
],
IOM
(Figu e 2.6 (c)), and
Geome ic Con ex [
60
],
IGC
(Figu e 2.6 (d)) a e wo me hods widely used in he
li e a u e o e alua e oom models [
63
,
81
]. Recen ly, [
172
,
68
] combine he s eng hs
o bo h o hem in one single map, ha we name Me ge Map, IMM (Figu e 2.6 (e)).
We addi ionally p opose o use a No mal Map (
INM
). We choose he wo k om
Eigen and Fe gus [
36
], which p oposes a mul iscale CNN ha e u ns dep h p edic ion,
su ace no mal es ima ion and seman ic labeling o indoo images. He e, we ake
ad an age o he su ace no mal es ima ion o c ea e ou e e ence map. In his case,
in o de o s i ch he local esul s back o he pano ama, we need o o a e he no mals
o se hem in a common e e ence ame. O e lapping a eas a e ackled by doing he
pe -pixel a e age o achie e a con inui y o he o e all image. The esul ing No mal
Map is shown in Figu e 2.6 (b). We also de e mine whe he o no he no mals om
each pixel belong o a main di ec ion (VPs) and label hem acco dingly. We se o
black he pixels ha do no belong o any main di ec ion. I can be no iced in Figu e
2.6 ha he ceiling is he wo s es ima ed pa by all he me hods. This happens
because he ceiling does no usually appea in con en ional images.
2.4.3 Expe imen al Resul s
We e alua e ou p oposal using 360 images o indoo scena ios om wo public da ase s.
In pa icula , mos o ou quan i a i e esul s ha e been ob ained om a subse o 85
pano amas o bed ooms and li ing ooms o he SUN360 da ase [
158
]. Addi ionally,
we also show esul s on he S an o d (2D-3D-S) da ase [6].
2.4 Layou s wi h Geome y and Deep Lea ning 33
(a) Ou s (b) Bazin e .al. [11] (c) PanoCon ex [172]
Fig. 2.7 Lines and anishing poin s de ec ion using h ee di e en me hods.
Fo each pano ama we c ea e he g ound u h as a segmen ed image
IGT
, like
hose in Figu e 2.6, whe e each pixel encodes he di ec ion o he su ace i belongs
o. A p e ious g ound u h was p o ided by [
172
], bu images we e labeled assuming
ooms ha e only 4 walls.
The accu acy o ou esul s is e alua ed by compu ing
EOP IHb,IGT 
, measu ing
he a io o equally-o ien ed pixels be ween he bes layou hypo hesis and he g ound
u h. Each EOP alue shown is a median o 10 imes pe o ming he expe imen . The
numbe o hypo heses d awn (
Nh
) is speci ied in each expe imen . Fo he expe imen s
we allow he algo i hm o ini ially selec om h ee o i e co ne s, i.e. o sol e layou s
wi h ou o eigh walls.
Lines and anishing poin s.
The p oposed algo i hm in Sec ion 2.4.1 wo ks di ec ly
on he equi ec angula image, allowing us o ob ain unique line segmen s, a oiding
hus duplica e lines coming om di e en spli s.
In [
172
] hey spli he pano ama in o de o un a speci ic algo i hm ha only
wo ks wi h pe spec i e images, wa ping hen all de ec ed line segmen s back o he
pano ama, whe eas in [
11
], hey sol e he p oblem by a b anch-and-bound amewo k
associa ed wi h a o a ion space sea ch, wo king di ec ly on pano amas. Fo [
172
] we
un di ec ly he code p o ided by he au ho s. [
11
] does no p o ide any code and
he e a e no public expe imen al esul s wi h omnidi ec ional images. Howe e , same
au ho s p o ide code o p e ious wo k [12,10] ha is used o his e alua ion.
Ou RANSAC-based algo i hm achie es eally simila esul s o [
172
,
11
] being also
much as e ,
∼
3.8s pe image in ou p oposal, compa ed o
∼
67s pe image wi h [
11
]
and
∼
42s pe image using [
172
]. Visual esul s om each wo k a e shown in Figu e 2.7.
Ad an ages o combining geome ic easoning wi h deep lea ning.
A com-
pa a i e s udy showing he e ec s o selec ing s uc u al lines (Sec ion 2.4.1) can be

34 Room Layou Es ima ion
G G+DL
0.7
0.8
0.9
1
EOP
Fig. 2.8
Ad an ages o combining.
He e we highligh he ad an ages o using
s uc u al lines om Geome y and Deep Lea ning combina ion [
92
] o e lines ob ained
only wi h Geome y. The mean is ep esen ed in solid black and he median in do ed
black. Also he s anda d de ia ion is shown in ligh colo and ji e ed aw da a a e
plo ed o each g oup.
ound in Figu e 2.8. Fo his expe imen we choose
Nh
=100 and he
INM
as e e ence
map. Each poin ep esen s an image. We show in ed he EOP when he comple e
se o lines p edic ed by ou geome ic app oach (G) is used o ob ain he candida e
co ne s. We show in g een he esul s when only he subse o s uc u al lines ob ained
combining geome y and deep lea ning (G+DL) is used. Mean and especially median
alues highligh he imp o emen when combining adi ional app oaches wi h deep
lea ning: 0.889 s. 0.925. The de ec ion o s uc u al lines allows o emo e clu e
e ec i ely, which ansla es in o be e accu acy.
Re e ence maps.
We compa e he pe o mance o ou model using he ou al e -
na i e e e ence maps in he e alua ion s ep. Fo his expe imen we also conside
Nh
=100. Table 2.1 shows he median EOP alue and he compu ing ime o c ea ing
each map. In e ms o accu acy,
INM
and
IMM
pe o m simila ly in median, al hough
he smalle s anda d de ia ion o he
INM
indica es mo e consis en esul s. Bo h
a e conside ably be e han
IOM
and
IGC
. Howe e , he
IOM
is abou en imes
as e o compu e han he
INM
and, he e o e, i s usage would be ecommendable i
he p io i y lies in ge ing as esul s in spi e o losing some accu acy. The smalle
s anda d de ia ion on he compu ing ime o he
INM
shows ha i does no a y
h ough images, unlike he o he s whose ime depends on scene-speci ic ea u es such
as he numbe o lines.
2.4 Layou s wi h Geome y and Deep Lea ning 35
EOP Compu ing Time
No mal Map (INM )0.925±0.061 243.36±1.42
O ien a ion Map (IOM ) 0.906±0.133 23.54±4.16
Geome ic Con ex (IGC) 0.883±0.114 174.07±13.28
Me ge Map (IMM ) 0.923±0.147 197.61±17.44
Table 2.1 Ra io o equally-o ien ed pixels when compa ing he bes inal hypo heses,
IHb
, wi h he g ound u h
IGT
, e alua ing in each case wi h a e e ence map. Also
he compu ing ime in seconds o gene a ing each map is shown.
Da ase Ca ego y EOP (Nh=100)
LSUN360 bed oom 0.921
li ing oom 0.933
S an o d (2D-3D-S) a ea1 0.873
a ea3 0.885
Table 2.2 Ra io o equally-o ien ed pixels e alua ed in di e en scena ios om wo
public da ase s.
Compa ison wi h he s a e o he a .
We pe o m a compa ison wi h PanoCon-
ex [
172
] since i is, o ou knowledge, he only di ec ly ela ed me hod wi h a ailable
code. In Figu e 2.9 we show he EOP a io and he compu ing ime necessa y o
gene a e he hypo heses o each me hod, a ying he numbe o hypo heses
Nh
. Ou
me hod clea ly ou pe o ms [
172
], being he di e ence la ge when only a ew hypo he-
ses a e conside ed. Al hough he di e ence dec eases as he amoun o hypo heses
ises, when bo h me hods each a s able EOP alue, ou p oposal achie es be e
esul s. Mo eo e , ou me hod wi h only 10 hypo heses (91,26%) bea s [
172
] wi h
100 hypo heses (89.66%). This shows he good pe o mance o ou s uc u al lines
selec ion which inc eases he likelihood o ge ing good hypo heses wi h only a ew
a emp s. Compu ing imes show again bigge di e ence when ewe hypo hesis a e
e alua ed. Only ooms up o 4 walls a e conside ed in his e alua ion o ge a ai
nume ical compa ison, bu ou me hod is also able o deal wi h mo e complex ooms,
see Figu e 2.10 o a quali a i e compa ison.
E alua ion in SUN360 and S an o d 2D-3D-S da ase s.
Besides he 85 images
om he SUN360 da ase , we addi ionally es ed ou me hod wi h 25 pano amas om
he S an o d (2D-3D-S) da ase . In Table 2.2 we show he EOP a io eached in bo h
36 Room Layou Es ima ion
10 20 30 40 50 60 70 80 90 100
0.8
0.85
0.9
0.95
Numbe o hypo heses e alua ed
EOP
Ou EOP
PanoCon ex EOP
10 20 30 40 50 60 70 80 90 100
0
2
4
6
8
10
Numbe o hypo heses e alua ed
Compu ing Time (s)
Ou CT
PanoCon ex CT
Fig. 2.9
Compa ison wi h PanoCon ex
[
172
] (wi h only ou -wall ooms). We
show he a io o equally-o ien ed pixels and compu ing ime agains he numbe o
hypo heses. Ou me hod ou pe o ms PanoCon ex and is able o p o ide much be e
esul s and much as e wi h ewe hypo heses.
Fig. 2.10
Compa ison wi h PanoCon ex
[
172
] in complex geome ies. Ou me hod
(cyan) is able o ind 6 walls whe eas [172] (da k blue) always inds jus 4 walls.
2.4 Layou s wi h Geome y and Deep Lea ning 37
Fig. 2.11
Top
: challenging co ido well es ima ed by ou app oach on S an o d
(2D-3D-S) da ase . Bo om: a clea case o ailu e.
Fig. 2.12 Layou p edic ions (cian) and g ound u h ( ed) on he SUN360 da ase .
Le : cuboid layou s. Righ : non-cuboid layou s.
44 Room Layou Es ima ion
Da ase Me hod 3DIoU(%) CE(%) PESS(%) PECS(%)
SUN360 PanoCon ex [172] 67.22 1.60 4.55 10.34
F-L C. e al. [43] - - - 7.26
Layou Ne [177] 74.48 1.06 3.34 -
Ou s 76.82 0.79 2.59 3.13
S n d.2D-3D F-L C. e al. [43] - - - 12.1
Ou s 70.64 1.15 3.95 4.98
Table 2.4 Pe o mance benchma king o SUN360 and S an o d 2D-3D da ase s aining
on SUN360 da a. SS: Simple Segmen a ion (3 ca ego ies): ceiling, loo and walls [
177
].
CS: Comple e Segmen a ion: ceiling, loo , wall1,..., walln[43].
Fig. 2.16 Layou p edic ions (yellow) and g ound u h (g een) on bo h da ase s.
he FCN has been ained on he same da ase , howe e esul s on S and o d 2D-3D
da ase a e also e y compe i i e. See Figu e 2.16 o some quali a i e esul s on he
SUN360 da ase .
2.6 CFL: Co ne s o Layou
We make he obse a ion ha he use o s anda d con olu ions in equi ec angula
images can lead o a loss o pe o mance o se e al easons. Fi s , p e- aining on
con en ional images is key due o he lack o aining 360 da a. Howe e , he p e-
ained ea u e space is non-dis o ed, which makes he p e- aining less e ec i e on
he dis o ed space o he equi ec angula images. Mo eo e , he igh and he le
side o he equi ec angula images a e he same spo in eali y. I we apply s anda d
con olu ions on hese images, he ne wo k simply do no unde s and he con inui y o
he scene, as he ke nel will die on he image bo de s. We p esen CFL in his sec ion.
We p opose an special ype o con olu ion, named EquiCon , ha adap s he size and

2.6 CFL: Co ne s o Layou 45
Fig. 2.17
Sphe ical pa ame iza ion o EquiCon s
. The sphe ical ke nel, de ined
by i s angula size (
αw×αh
) and esolu ion (
w× h
), is con ol ed a ound he sphe e
wi h angles ϕand θ.
shape o he ke nel o he equi ec angula image dis o ions. EquiCon s can di ec ly
subs i u e he s anda d con olu ions, and demons a ed o ha e se e al ad an ages o
wo k wi h 360 images.
2.6.1 Equi ec angula Con olu ions
Sphe ical images a e ecei ing an inc easing a en ion due o he g owing numbe o
omnidi ec ional senso s in d ones, obo s and au onomous ca s. We obse e ha a
naï e applica ion o con olu ional ne wo ks o an equi ec angula p ojec ion, is no ,
in p inciple, a good choice due o he space- a ying dis o ions in oduced by such
p ojec ion.
In his sec ion we p esen a con olu ion ha we name EquiCon , which is de ined
in he sphe ical domain ins ead o he image domain and i is implici ly in a ian
o equi ec angula ep esen a ion dis o ions. The ke nel in EquiCon s is de ined
as a sphe ical su ace pa ch –see Figu e 2.17. We pa ame ize i s ecep i e ield by
he angles
αw
and
αh
. Thus, we di ec ly de ine a con olu ion o e he ield o iew.
The ke nel is o a ed and applied along he sphe e and i s posi ion is de ined by he
sphe ical coo dina es (
ϕ
and
θ
in he igu e) o i s cen e . Unlike s anda d ke nels, ha
a e pa ame e ized by hei size
kw×kh
, wi h EquiCon s we de ine he angula size
(
αw×αh
) and esolu ion (
w× h
). In p ac ice, we keep he aspec a io,
αw
w
=
αh
h
,
and we use squa e ke nels, so we will e e he ield o iew as
α
(
αw
=
αh
) and he
46 Room Layou Es ima ion
+
-
-+
Fig. 2.18
E ec o changing ield o iew α( ad) and esolu ion in
EquiCon s
. 1 column shows a na ow ield o iew
α
= 0
.
2. 2 column shows a
wide ke nel keeping i s esolu ion (a ous-like),
α
= 0
.
5. 3 column shows an e en
la ge ield o iew o he ke nel,
α
= 0
.
8. No ice how he ke nel adap s o he
equi ec angula dis o ion. Rows a e esolu ions = 3 and = 5.
esolu ion as
(
w
=
h
) espec i ely om now on. In his wo k, we choose alues o
esolu ion and ield o iew o be he same as he image.
Al hough we use by de aul he same esolu ion and ield o iew om he image
in ou model, i can be di e en . As we inc ease he esolu ion o he ke nel, he
angula dis ance be ween he elemen s dec eases, wi h he in ui i e uppe limi o no
gi ing mo e esolu ion o he ke nel han he image i sel . In o he wo ds, he ke nel is
de ined in a sphe e, being i s adius less o equal o he image sphe e adius. The e o e,
EquiCon s can also be seen as a gene al model o sphe ical A ous Con olu ions
[
20
,
21
] whe e he ke nel size is wha we call esolu ion, and he a e is he ield o iew
o he ke nel di ided by he esolu ion. An example o he di e ences o EquiCon s by
modi ying αand can be seen in Figu e 2.18.
EquiCon s De ails
In [
27
], hey in oduce de o mable con olu ions by lea ning addi ional o se s om
he p eceding ea u e maps. O se s a e added o he egula ke nel loca ions in he
S anda d Con olu ion enabling ee o m de o ma ion o he ke nel.
2.6 CFL: Co ne s o Layou 47
S anda d De o mable Equi ec angula
Fig. 2.19
E ec o o se s on a
3
×
3
ke nel
. Le : Regula ke nel in S anda d
Con olu ion. Cen e : De o mable ke nel in [
27
]. Righ : Sphe ical su ace pa ch in
EquiCon s.
Inspi ed by his wo k, we de o m he shape o he ke nels acco ding o he geome -
ical p io s o he equi ec angula image p ojec ion. To do ha , we gene a e o se s
ha a e no lea ned bu ixed gi en he sphe ical dis o ion model and cons an o e
he same ho izon al loca ions. He e, we desc ibe how o ob ain he dis o ed pixel
loca ions om he o iginal ones.
Le us de ine (
u0,0, 0,0
)as he pixel loca ion on he equi ec angula image whe e
we apply he con olu ion ope a ion (i.e. he image coo dina e whe e he cen e o he
ke nel is loca ed). Fi s , we de ine he coo dina es o e e y elemen in he ke nel and
a e wa ds we o a e hem o he poin o he sphe e whe e he ke nel is being applied.
We de ine each poin o he ke nel as ollows,
ˆpij =



ˆxij
ˆyij
ˆzij



=



i
j
d



,(2.8)
whe e
i
and
j
a e in ege s in he ange [
− −1
2, −1
2
]and
d
is he dis ance om he cen e
o he sphe e o he ke nel g id. In o de o co e he ield o iew α,
d=
2 an(α
2).(2.9)
48 Room Layou Es ima ion
Fig. 2.20
EquiCon s on sphe ical images.
We show h ee ke nel posi ions o
highligh he di e ences be ween he o se s. As we app oach o he poles (la ge
θ
angles) he de o ma ion o he ke nel on he equi ec angula image is bigge , in o de
o ep oduce a egula ke nel on he sphe e su ace. Addi ionally, wi h EquiCon s, we
do no use padding when he ke nel is on he bo de o he image since o se s ake he
poin s o hei co ec posi ion on he o he side o he 360◦image.
2.6 CFL: Co ne s o Layou 49
We p ojec each poin in o he sphe e su ace by no malizing he ec o s, and o a e
hem o align he ke nel cen e o he poin whe e he ke nel is applied.
pij =



xij
yij
zij



=Ry(ϕ0,0)Rx(θ0,0)ˆpij
|ˆpij|,(2.10)
whe e
Ra
(
β
)s ands o a o a ion ma ix o an angle
β
a ound he
a
axis.
ϕ0,0
and
θ0,0
a e he sphe ical angles o he cen e o he ke nel –see Figu e 2.17, and a e de ined as
ϕ0,0= (u0,0−W
2)2π
W;θ0,0=−( 0,0−H
2)π
H,(2.11)
whe e
W
and
H
a e, espec i ely, he wid h and heigh o he equi ec angula image
in pixels. Finally, he es o elemen s a e back-p ojec ed o he equi ec angula image
domain. Fi s , we con e he uni sphe e coo dina es o la i ude and longi ude angles:
ϕij = a c an(xij
zij
) ; θij = a csin(yij).(2.12)
And hen, o he o iginal 2D equi ec angula image domain:
uij = (ϕij
2π+1
2)W; ij = (−θij
π+1
2)H. (2.13)
In Figu e 2.19 we show how hese o se s a e applied o a egula ke nel; and in
Figu e 2.20 h ee ke nel samples on he sphe ical and on he equi ec angula images.
2.6.2 Lea ning co ne s o layou
He e we desc ibe ou end- o-end app oach o eco e ing he oom co ne s ha allow us
o es ima e he layou , i.e. he main s uc u e o he oom, om a single 360◦image.
We use he ne wo k a chi ec u e p esen ed in Sec ion 2.5, which con ol es he
ea u e maps wi h s anda d con olu ions and use up-con olu ions o decode he ou pu .
We name i he e CFL S dCon s. In his sec ion, we p opose a ne wo k a ia ion, see
Figu e 2.21, ha we name CFL EquiCon s, which uses Equi ec angula Con olu ions
in he encode and he decode , using unpooling o upsample he ou pu . We use he
loss unc ion Lmaps p esen ed in Sec ion 2.5.
F om Co ne Maps o 3D Layou .
Cu en me hods [
177
,
43
,
172
] use p e-
compu ed anishing poin s and pos e io op imiza ions, being cons ained o p oduce

50 Room Layou Es ima ion
3
Skip-connec ions
P elimina y
p edic ions
ResNe -50
4
63
256x128
Inpu Pano ama
128x64
Ou pu Co ne /Edge Maps
Equi ec angula con olu ion Equi ec angula Con olu ion
+ Unpooling
Fig. 2.21
CFL a chi ec u e
. Ou ne wo k is buil upon ResNe -50, adding a single
decode ha join ly p edic s edge and co ne maps. The e a e wo ne wo k a ia ions:
he o iginal one, p esen ed in Sec ion 2.5, applies s anda d con olu ions and upcon-
olu ions on he equi ec angula pano ama, whe eas his one applies Equi ec angula
Con olu ions and Equi ec angula Con olu ions + unpooling di ec ly on he sphe e.
1
1'
23
2' 3'
4
4' 4
4'
1
1'
2
2'
2
2'
3
3'
(a) The 2D co ne s coo dina es a e he maximum ac i a ions in
he p obabili y map. F om he 2D co ne s, we can di ec ly eco e
he 3D layou by doing a couple o assump ions.
h
g
(b) Assump ions. (i) ceiling
and loo planes a e pa allel
and o ien ed wi h he g a i y
di ec ion, (ii) he came a is lo-
ca ed a a ce ain heigh .
Fig. 2.22
Layou om co ne p edic ions
. F om he co ne p obabili y map, he
coo dina es wi h maximum alues a e di ec ly selec ed o gene a e he layou .
2.6 CFL: Co ne s o Layou 51
s ic Manha an 3D layou s. Aiming o a as end- o-end simple model, CFL a oids
ex a compu a ion and adop a ep esen a ion usually e e ed as So /Weak Manha an
[
47
] o A lan a Wo ld [
71
]. Following his, ho izon al di ec ions a e no necessa ily
o hogonal o each o he , hus elaxing he model assump ions. To his end, we
simply ollow a na u al ans o ma ion om co ne s coo dina es o 2D and 3D layou .
The 2D co ne s coo dina es a e he maximum ac i a ions in he p obabili y map.
Assuming ha he co ne se is consis en , hey a e di ec ly joined, om le o igh ,
in he uni sphe e space and e-p ojec ed o he equi ec angula image plane. The 3D
layou is in e ed by only assuming ceiling- loo pa allelism, lea ing he wall s uc u e
uncons ained i.e. we assume ha he loo co ne s a e on he same plane and he op
co ne s a e di ec ly abo e he loo ones, bu we do no o ce he usual Manha an
pe pendicula i y be ween walls. Co ne s a e p ojec ed o loo and ceiling planes gi en
a uni a y came a heigh ( i ial as esul s a e up o scale). See Figu e 2.22.
We di ec ly join co ne s om le o igh , meaning ha ou model would no
wo k i any wall is occluded because o he con exi y o he scene. In hose pa icula
cases, he joining p ocess should ollow a di e en o de . In Sec ion 2.4.2 we p opose a
geome y-based pos -p ocessing ha could alle ia e his p oblem, bu i s cos is high
and i equi es he Manha an Wo ld assump ion.
We p o ide now a de ailed explana ion o how, om he p edic ed 2D co ne
posi ions, we can di ec ly eco e he 3D layou .
We can de ine a plane as he se o all poin s
P
= (
x, y, z
)such ha
P·N
+
d
= 0,
whe e he no mal N= (nx, ny, nz)is a no malized ec o pe pendicula o i s su ace
and
d
is he dis ance ha sepa a es i om he o igin o coo dina es in he di ec ion o
he no mal. Since we assume ceiling- loo pa allelism and a came a heigh ,
N
o bo h
he loo and ceiling planes is equal and co esponds o he e ical di ec ion, and he
dis ance
d
om he loo o he came a is known. The dis ance o he ceiling is ye
unknown.
Addi ionally, hanks o he na u e o sphe ical images, we can easily ob ain he
3D ay
R
(
) =
O
+

V·
(pa ame ic ep esen a ion) going om he cen e o he
sphe e
O
= (
ox, oy, oz
) h ough he co ne posi ion, wi h no malized di ec ion ec o

V
= (
x, y, z
). To ob ain he no malized di ec ion ec o

V
, we need he co ne
posi ion in he sphe e, hus we ans o m he image coo dina es o he co ne s (
u,
)
in o sphe ical coo dina es and hen o he Euclidean 3D space. Equa ions o his a e
p esen ed in Sec ion 2.3.
52 Room Layou Es ima ion
In he i s place, Eq (2.14) gi e us he angles ha de ine he poin (
u,
)in he
sphe e.
ϕ= (u−W
2)2π
W;θ=−( −H
2)π
H(2.14)
Whe e W and H a e he wid h and heigh o he equi ec angula image. Second, once
hese o a ions a e known we can compu e he di ec ion o he ay. The e o e, using
Eq (2.15) we can calcula e 
V.

V=



−cos(θ)sin(ϕ)
sin(θ)
cos(θ)cos(ϕ)



(2.15)
The in e sec ion be ween he co ne ay and he co esponding loo o ceiling plane
will gi e us he ac ual 3D co ne poin
P
= (
x, y, z
)(up o scale), i.e. he in e sec ion
ep esen s ha poin
P
on he su ace o he plane ha e i ies he ay equa ion:
(
ox
+
x·
)
nx
+(
oy
+
y·
)
ny
+(
oz
+
z·
)
nz
+
d
= 0. The poin
P
o in e sec ion would
simply be he esul o e alua ing he calcula ed
, Eq (2.16), in he ay equa ion
R
(
).
=−oxnx+oyny+oznz+d
xnx+ yny+ znz
(2.16)
Le ’s conside we ha e pe o med he ope a ions o compu e one co ne poin on
he loo plane,
PF
= (
xF, yF, zF
). The co esponding poin on he ceiling plane (
PC
)
will be on op o i (ie.
xF
=
xC
and
yF
=
yC
). The e o e, we can use his o compu e
C, Eq (2.17), and hus he ceiling poin :
C=(xF−ox)
C
x
(2.17)
whe e

VC
= (
C
x, C
y, C
z
)is compu ed as in (2.15) wi h he co esponding ceiling poin
in he image. No ice ha wi h
PC
we ha e he in o ma ion we we e missing o eco e
he ceiling plane.
2.6.3 Expe imen al Resul s
We p esen a se o expe imen s o e alua e CFL using bo h S anda d Con olu ions
(S dCon s) and he p oposed Equi ec angula Con olu ions (EquiCon s). We do no
only analyze he co ne maps p edic ed by ou model, bu also he impac o each
algo i hmic componen h ough abla ion s udies. We epo he pe o mance o ou
2.6 CFL: Co ne s o Layou 53
p oposal in wo di e en da ase s, and show quali a i e 2D and 3D models o di e en
indoo scenes.
Da ase s.
We use wo public da ase s ha comp ise se e al indoo scenes, SUN360
[
158
] and S an o d (2D-3D-S) [
6
] in equi ec angula p ojec ion (360
◦
). The o me is
used o abla ion s udies, and bo h a e used o compa ison agains se e al s a e-o -
he-a baselines.
SUN360 [
158
]: We use
∼
500 bed oom and li ing oom pano amas om his da ase
labeled by Zhang e al. [
172
]. We use hese labels bu , since all pano amas we e labeled
as box- ype ooms, we hand-label and subs i u e 35 pano amas ep esen ing mo e
ai h ully he ac ual shapes o he ooms. We spli he aw da ase in 85% aining
scenes and 15% es scenes andomly by making su e ha he e we e ooms o mo e
han 4 walls in bo h pa i ions.
S an o d 2D-3D-S [
6
]: This da ase con ains mo e challenging scena ios like clu e ed
labo a o ies o co ido s. In [
177
], hey use a eas 1, 2, 4, 6 o aining, and a ea 5 o
es ing. Fo ou expe imen s we use same pa i ions and he g ound u h p o ided by
hem.
Implemen a ion de ails.
The inpu o he ne wo k is a single pano amic RGB
image o esolu ion 256
×
128. The ou pu s a e, on he one hand, he oom layou edge
map and on he o he hand, he co ne map, bo h o hem a esolu ion 128
×
64. A
widely used s a egy o imp o e gene aliza ion o neu al ne wo ks is da a augmen a ion.
We apply andom e asing, ho izon al mi o ing as well as ho izon al o a ion om 0
◦
o 360
◦
o inpu images du ing aining. The weigh s a e all ini ialized using ResNe -50
[
59
] ained on ImageNe [
119
]. Fo CFL EquiCon s we use he same ke nel esolu ions
and ield o iews as in ResNe -50. This means ha o a s anda d 3
×
3 ke nel applied
o a W
×
H ea u e map,
= 3 and
α
=
o
W
, whe e
o
= 360
◦
o pano amas. We
minimize he c oss-en opy loss using Adam [
73
], egula ized by penalizing he loss wi h
he sum o he L2 o all weigh s. The ini ial lea ning a e is 2
.
5
e−4
and is exponen ially
decayed by a a e o 0.995 e e y epoch. We apply a d opou a e o 0.3.
The ne wo k is implemen ed using Tenso Flow [
1
] and ained and es ed in a
NVIDIA Ti an X. The aining ime o S dCon s is a ound 1hou and he es ime
is 0
.
31 seconds pe image. Fo EquiCon s, aining akes 3hou s and es a ound 3
.
32
seconds pe image.
Ne wo k’s ou pu e alua ion.
We measu e he quali y o ou p edic ed p obabili y
co ne maps using i e s anda d me ics: in e sec ion o e union IoU, p ecision P,
60 Room Layou Es ima ion
Me hod Compu a ion Time (s)
PanoCon ex [172]>300
Layou Ne [177]44.73
DuLa-Ne [163]13.43
CFL EquiCon s 3.47
CFL S dCon s 0.46
Table 2.8
A e age compu ing ime pe image.
E e y app oach is e alua ed using
NVIDIA Ti an X and In el Xeon 3.5 GHz (6 co es) excep DuLa-Ne , e alua ed using
NVIDIA 1080Ti GPU. Ou end- o-end me hod is mo e han 100 imes as e han
o he me hods.
Fig. 2.27 Layou p edic ions (ligh magen a) and g ound u h (da k magen a) on bo h
da ase s.
2.7 Quali a i e Resul s
He e we show addi ional quali a i e esul s o ou eco e ed layou s in SUN360 [
158
]
and S an o d 2D-3D [
6
] da ase s. Figu es 2.28 and 2.29 collec examples in SUN360
da ase and show indoo scenes wi h di e en geome ies, no only cuboid shapes.
Figu e 2.30 shows examples in S an o d 2D-3D da ase . Pano amas in his da ase do
no co e ull iew e ically and he indoo scenes ep esen mo e challenging scena ios
like clu e ed labo a o ies o co ido s.
2.8 Conclusion
In his chap e we p esen h ee di e en app oaches ha show an e olu ion o ou
esea ch on he 3D oom layou es ima ion p oblem om single 360 images.
We i s p opose a no el pipeline ha combines geome y and deep lea ning o
ob ain s uc u al lines and co ne s, om which he layou hypo heses a e gene a ed.

2.8 Conclusion 61
We also demons a e how o deal wi h non- isible s uc u al co ne s by au oma ically
p edic ing new co ne s du ing he hypo heses gene a ion p ocess, so ha he gene a ed
oom layou s sa is y he Manha an wo ld assump ion. This idea allows us o gene alize
o cuboid and non-cuboid layou s, lea ing behind he simpli ica ion o 4 wall ooms.
We addi ionally p esen a new deep lea ning model o p edic s uc u al lines
and co ne s di ec ly on pano amic images. The CNN allows us o a oid expensi e
p e-p ocessing s ages imp o ing he o e all e iciency o he me hod. Addi ionally,
wo king di ec ly on pano amic images ensu es a ull le e age o he oom con ex ,
gi ing be e p edic ions.
In he las app oach, we p esen CFL, he i s end- o-end algo i hm o layou
eco e y in 360
◦
images. Ou expe imen al esul s demons a e ha ou p edic ed
layou s a e mo e accu a e han he s a e o he a . Addi ionally, he emo al o
ex a p e- and pos -p ocessing s ages makes ou me hod much as e han o he
wo ks. Finally, being en i ely da a-d i en elaxes he geome ic assump ions ha a e
commonly used in he s a e o he a and limi s hei usabili y in complex geome ies.
We p esen wo di e en a ian s o CFL. The i s one, implemen ed using S anda d
Con olu ions, educes he compu a ion in 100 imes and i is e y sui able o images
aken wi h a ipod ( ecommended i he ime is a c i ical issue). The second one uses
ou p oposed implemen a ion o Equi ec angula Con olu ions ha adap hei shape
o he equi ec angula p ojec ion o he sphe ical image ( ecommended i looking o
obus ness and be e gene aliza ion). This p o es o be mo e obus o ansla ions
and o a ions o he came a making i ideal o pano amas aken by a hand-held
came a.
62 Room Layou Es ima ion
Fig. 2.28 Layou p edic ions (ligh magen a) and g ound u h (da k magen a) on he
SUN360 anno a ion da ase [158]. Bes iewed in colo .
2.8 Conclusion 63
Fig. 2.29 Layou p edic ions (ligh magen a) and g ound u h (da k magen a) o
complex oom geome ies
on he SUN360 anno a ion da ase [
158
]. Bes iewed
in colo .
64 Room Layou Es ima ion
Fig. 2.30 Layou p edic ions (ligh magen a) and g ound u h (da k magen a) on he
S an o d 2D-3D anno a ion da ase [6]. Bes iewed in colo .
Chap e 3
Objec Recogni ion
“The Th ee R’s o Compu e Vision: Recogni ion, Recons uc ion &Reo ganiza ion.”
— Ji end a Malik
In he las ew yea s, he e has been a g owing in e es in pano amic images. While
se e al asks ha e been imp o ed hanks o he con ex ual in o ma ion hese images
o e , objec ecogni ion in indoo scenes s ill emains a challenging p oblem ha has no
been deeply in es iga ed. We p o ide an objec ecogni ion sys em ha pe o ms objec
de ec ion and seman ic segmen a ion asks by using a deep lea ning model adap ed o
ma ch he na u e o equi ec angula images. F om hese esul s, ins ance segmen a ion
masks a e eco e ed, e ined and ans o med in o 3D bounding boxes ha a e placed
in o he 3D model o he oom. The p oposed me hod ou pe o ms he s a e o he a
by a la ge ma gin and shows a comple e unde s anding o he main objec s in indoo
scenes.
- pain ing
- bed
- bed
- able
- able
- mi o
- window
- window
- window
- chai
- so a
- doo
- doo
- doo
- doo
- cabine
- cabine
- bedside

66 Objec Recogni ion
3.1 In oduc ion
The inc easing in e es in au onomous mobile sys ems, like d ones, obo ic acuum
cleane s o assis an obo s, makes de ec ion and ecogni ion o objec s in indoo
en i onmen s a e y impo an and demanded ask.
Since ecognizing a isual concep is ela i ely i ial o a human, i is wo h
conside ing he ha d challenges inhe en ly in ol ed. Objec s in images can be o ien ed
in many di e en ways, a y hei size, be occluded, blended in o he en i onmen
because o hei colo o appea ance, o a ec ed by di e en illumina ion condi ions,
which changes d as ically hei aspec on he pixel le el. Mo eo e , he concep behind
an objec ’s name is some imes b oad, including non-clea on ie s o o he concep s.
Fo example, whe e do you conside he limi s be ween a so a and an a mchai ?
Con olu ional Neu al Ne wo ks (CNNs) ha e al eady demons a ed o be he bes
known models o pe o m objec ecogni ion, as hey a e capable o dealing wi h hose
challenges by au oma ically lea ning objec s’ inhe en ea u es and co ec ly iden i y
hei in insic concep s.
Howe e , images om con en ional came as ha e a small ield o iew, much smalle
han human ision, which implies ha con ex ual in o ma ion canno be as use ul as i
should. To o e come his limi a ion, a eal impac came wi h he a i al o he 360
◦
ull- iew pano amic images, which a e ecen ly a ising mo e and mo e in e es in he
obo ics and compu e ision communi y, as hey allow us o isualize, in a single
image, he whole scene a he same ime. Toge he wi h all o hei po en ial we ha e
o deal wi h challenges p oduced by hei own sphe ical p ojec ion, such as dis o ion,
o he lack o comple e, labeled and massi e da ase s. This equi es he de elopmen
o speci ic echniques ha ake ad an age o hei s eng hs and allow wo king wi h
pano amic images in an e icien and e ec i e way.
In his Chap e , we p opose an objec ecogni ion sys em ha p o ides a comple e
unde s anding o he main objec s in an indoo scene om a single 360
◦
image in
equi ec angula p ojec ion. Ou me hod ex ends he Bli zNe model [
35
] o pe o m
bo h objec de ec ion and seman ic segmen a ion asks bu adap ed o ma ch he
na u e o he equi ec angula image inpu . We ain he ne wo k o p edic 14 di e en
classes o main indoo scenes ela ed objec s. Resul s o he CNN a e pos -p ocessed o
ob ain ins ance segmen a ion masks, which a e success ully e ined by aking ad an age
o he spa ial con ex ual clues ha he oom layou p o ides. In his wo k, we no
only show he po en ial o exploi ing he 2D oom layou o imp o e he ins ance
segmen a ion mask, bu also he possibili y o le e aging he 3D layou o gene a e 3D
objec bounding boxes di ec ly om he imp o ed masks.
3.2 Rela ed Wo k 67
3.2 Rela ed Wo k
Objec de ec ion ield has been mainly domina ed by wo di e en app oaches: one-
s age and wo-s age de ec o s. Two-s age de ec o s, as he i s R-CNN[
51
] a chi ec u e
ollowed by i s a ian s Fas R-CNN[
50
], Fas e R-CNN[
113
] and Mask R-CNN[
57
]
achie e g ea accu acy bu lowe speed. They equi e i s ly o e ine p oposals o
ob ain he ea u es needed o classi y he objec s. On he o he hand, one-s age
de ec o s, ollowing YOLO[
111
] and SSD[
86
] simul aneous bounding box e inemen
and classi ica ion, signi ican ly educe compu a ional cos . They achie e eal- ime
pe o ming main aining high accu acy, which is needed o mos applica ions in au-
onomous mobile sys ems. SSD mul i-scale py amid idea p o es o help in conduc ing
mo e accu a e de ec ions and manage widely a ious objec sizes, app oach ollowed in
mos s a e-o - he-a objec de ec o s.
While all hose models op imize bounding box de ec ion, no so many in eg a e
in hei pipeline he pixel-wise ecogni ion needed o many applica ions. In his way,
Bli zNe [
35
] is a one-s age mul i-scale model ha adds seman ic segmen a ion and
he e o e ecognizes objec s a pixel le el. I also p o es he ad an ages o join ly
lea ning wo scene unde s anding asks: objec de ec ion and seman ic segmen a ion,
which bene i om each o he by sha ing almos he comple e ne wo k a chi ec u e.
Howe e , s a e-o - he-a esea ch mainly ocuses on using con en ional images.
Thei limi ed ield o iew p e en s con ex ual in o ma ion om being as c ucial as
i is in scene unde s anding o humans. Di e en ly om ou doo objec ecogni ion,
whe e hanks o he inc easing esea ch on au onomous d i ing, he e a e ecen wo ks
using pano amic images [
93
] [
164
], he e is no wide esea ch on objec ecogni ion om
indoo pano amas. A ecen wo k ha add esses his p oblem is [
32
], whe e Deng e .
al use a R-CNN app oach, and also e alua e hei own implemen a ion o DPM [
45
]
on pano amas. The mos ele an wo k on indoo pano amic objec ecogni ion is
PanoCon ex [
172
]. I includes 2D objec de ec ion and seman ic segmen a ion among
o he 3D scene unde s anding asks, p o ing he po en ial o ha ing a la ge ield o iew
o ecogni ion p oblems. Thei me hod, ne e heless, is based on geome ical easoning
and adi ional compu e ision ea u e ex ac o s and can be s ill conside ed as s a e-
o - he-a in indoo objec ecogni ion on pano amic images. Recen esea ch on his
kind o images includes 3D layou eco e y [
98
] [
69
] [
177
] [
41
] and scene modeling [
165
],
which p o ides global con ex and gi es a 3D in e p e a ion o he scene om a single
iew. In [
90
], hey show ha his asks can also bene i and augmen an omnidi ec ional
SLAM. Combining objec ecogni ion and 3D layou eco e y mo i a es ou p oposal
o ob ain he 3D ecogni ion and loca ion o main objec s in ou oom.
68 Objec Recogni ion
3.3 Da ase ex ension
Pano amic images da ase s wi h objec ecogni ion labels a e no as s anda d o
comple e as con en ional images ones [
67
] [
133
] [
6
]. The e o e, in his wo k we decide
o ex end he SUN360 da abase [
67
] wi h segmen a ion labels. Fo e e y pano ama,
we gene a e indi idual masks encoding each objec ’s spa ial layou . Addi ionally,
we combine all he masks ob aining a seman ic segmen a ion pano amic image wi h
pe -pixel classi ica ion. Bed oom and li ing oom se s, o med by 418 and 248 images
espec i ely, a e used and 14 di e en objec classes a e conside ed. The da ase is
di ided in o 85% o ain and alida ion and 15% o es .
We gene a e segmen a ion masks based on 2D bounding poin s o he objec s, aken
om PanoCon ex [
172
] wo k. We p ojec hem on he sphe ical domain o ollow
dis o ion pa e ns in con ou s and o co ec ly manage objec s ha appea c opped
on he ho izon al image limi s (see Sec ion 2.3 o mo e de ails abou he sphe ical
geome y). To combine he bina y masks and c ea e he seman ic segmen a ion
pano ama, wi h he lack o dep h o o he 3D in o ma ion, an hypo hesis o occlusion
among objec s is needed. We conside he assump ion ha objec s a e no in gene al
comple ely occluded, and he e o e o each pai o objec s in con lic hei a ea o
o e lap and size a e compu ed. I a ea o o e lap is bigge han a h eshold, he
smalles objec is conside ed close and comple ely isible and o he wise he bigges
one is selec ed. Wi h i s e iden limi a ions, his hypo hesis expe imen ally p o es o
wo k well in mos o he cases, allowing o co ec ly segmen mos o he isible and
cuboid-shaped objec s in images as shown in Figu e 3.1. The comple e da ase used
in his wo k is eleased o public access and can be ound in he p ojec webpage1.
3.4 Model
In his sec ion we p esen ou objec ecogni ion model, called Pano amic Bli zNe ,
ha is based on he o iginal CNN Bli zNe [
35
] bu adap ed o wo k speci ically wi h
comple e equi ec angula images. I add esses bo h objec de ec ion and seman ic
segmen a ion asks, ollowing Bli zNe a chi ec u e: a Fully Con olu ional model
ha ollows he encode -decode app oach wi h skip connec ions. I pe o ms mul i-
scale ecogni ion and akes ad an age o join lea ning. Main changes o hei base
implemen a ion include he use o he comple e ec angula pano ama, modi ying he
inpu aspec a io. We also change he ancho boxes p oposals, as he new inpu shape
1A ailable a h ps://webdiis.uniza .es/∼jgue e / oom_OR/
3.4 Model 69
Fig. 3.1 Resul o ou me hod o
c ea e seman ic segmen a ion masks
, assuming
hypo hesis o occlusion. No ice on he le he di e ences be ween c ea ing s aigh
con ou s on image domain ( op) s. sphe ical domain (bo om).
needs o be conside ed because hey a e cen e ed on pixels g id. Ou bounding boxes
p oposals a e done by i s ly con e ing image o a egula g id, co e ing he whole
ec angula -shaped image. G id has di e en dimensions in each laye , om 128x256
o 1x2, because o he i e a i ely lowe scale o he ea u e maps. In each g id cell
5 di e en p oposals a e c ea ed wi h 5 di e en aspec a ios: 1, 2, 1/2, 3 and 1/3,
allowing he ne wo k o manage di e en objec shapes.
Special men ion dese es da a augmen a ion as an impo an echnique o a oid
o e i ing, pa icula ly on non-massi e da ase s like in ou case. He e, we modi y
he o iginal da a augmen a ion by emo ing andom c ops on images (con ex ual
in o ma ion is impo an ) and adding ho izon al o a ion om 0
°
o 360
°
o co e all
di e en posi ions on he sphe e.
How can we deal wi h
360
◦images dis o ion?
We exploi he po en ial o om-
nidi ec ional images co e ing 360
◦
ho izon al and 180
◦
e ical ield o iew ep esen ed
in equi ec angula p ojec ion. While hese images allow us o analyse he whole scene
a once aking ad an age o all he con ex , hey p esen g ea dis o ions due o hei
p ojec ion o he sphe e. He e, we eplace all s anda d con olu ions o ou Pano amic
Bli zNe by equi ec angula con olu ions (EquiCon s [
41
]), o s udy hei impac on he
ask o ecognizing objec s. Wi h his kind o con olu ions, he ke nel adap s i s shape
and size acco dingly o he dis o ions p oduced by he equi ec angula p ojec ion.
As men ioned in [
41
], he dis o ion p esen ed is loca ion dependen , speci ically, i
depends on he pola angle. They demons a e how EquiCon s can be eally con enien
o gene alize o di e en came a posi ions since he layou shape can su e om many
76 Objec Recogni ion
inpu
mIoU
backg bed pic u e able mi o window cu ain chai ligh so a doo cabine bedside shel
Ou s 53.0 90.7 61.7 32.1 75.2 42.3 55.8 54.0 55.1 31.4 34.6 63.6 48.5 40.7 52.2 57.4
Ou s+I
53.1 90.7 63.3 30.1 75.0 41.5 56.7 55.3 55.5 34.0 31.8 62.9 48.8 40.2 53.5 57.5
Table 3.3 Seman ic segmen a ion esul s be o e and a e applying he ins ance seg-
men a ion pos -p ocessing. Ini ial seman ic segmen a ion is aken om ou CNN
ou pu .
model mAP bed pic u e able mi o window cu ain chai ligh so a doo cabine bedside shel
∗[39] 29.4 35.2 56.0 21.6 19.2 21.8 29.5 26.0 — 22.2 31.9 — — 31.0 —
∗[32] 68.7 76.3 68.0 73.6 58.7 62.6 69.5 68.0 — 72.5 67.3 — — 70.0 —
Ou s
SC 76.8 94.9 85.0 83.3 71.9 72.2 72.2 71.9 35.0 89.3 75.5 57.9 87.9 91.1 30.5
Ou s
EC 77.8 95.3 83.9 82.1 76.2 70.9 75.9 80.9 41.0 85.4 72.5 55.6 91.4 93.3 40.2
Table 3.4
Objec de ec ion
esul s on SUN360 es se wi h ou me hod Pano amic
Bli zNe using s anda d con olu ions (SC) e sus equi ec angula con olu ions (EC),
compa ed wi h p e ious me hods.
∗
Resul s ained and e alua ed on a combina ion
o da ase s (including SUN360) by [32]
a y and scenes a e ela i ely simila ) does no d as ically damage es esul s, bu will
p obably be c ucial when wo king on di e en da ase s. Finally, EquiCon s also p o e
o make de ec ions wi h highe con idence as, when aising he con idence h eshold
o 0
.
95, hei ecall is main ained o e 40% compa ed o 28% achie ed wi h s anda d
con olu ions.
3.5.4 Ins ance segmen a ion
In his expe imen we compa e he seman ic segmen a ion ou pu o he ne wo k wi h
he esul o applying ou ins ance segmen a ion me hod o c ea e imp o ed seman ic
segmen a ion maps. Al hough his is no he objec i e o he pos -p ocessing ( he
goal is o di e en ia e among di e en ins ances), i is e alua ed o p o e he in luence
ha i can ha e in segmen a ion pe o mance. Ou in ui ion was ha he ins ance
segmen a ion me hod would imply an imp o emen o he ini ial segmen a ion maps
because i gi es highe con idence o de ec ions, whose pe o mance is clea ly highe
han segmen a ion’s one in ou model. Resul s, shown in Table 3.3, a e e y simila
and lead us o conclude ha he pos -p ocessing does no p o e o be in luen ial in his
way. Howe e , quali a i e esul s suppo ou in ui i e idea by showing some clea ly
imp o ing cases ha a e ema ked in Figu e 3.4.

3.5 Expe imen al Resul s 77
Fig. 3.4
Ins ance segmen a ion pos -p ocessing
esul s. Top is ini ial seman ic
segmen a ion (ou pu o CNN) and bo om is esul o pos -p ocessing. No ice ha
apa om co ec ly di e en ia e among ins ances (highligh ed in blue) i imp o es
o iginal segmen a ion (highligh ed in ed and g een o ailed and imp o ed segmen a ion
espec i ely).
Be o e co ec ion A e co ec ion
Fig. 3.5 A e combining he oom layou wi h ou segmen a ion masks, he model
expe iences a clea imp o emen as a whole. Howe e , he e we wan o show a
ailu e
case
whe e, when assuming ha doo s mus each he loo , we may ha e o e lapping
wi h o he occluding objec s in he image, damaging segmen a ion esul s bu imp o ing
he doo 3D localiza ion.
78 Objec Recogni ion
model
mIoU
backg bed pic u e able mi o window cu ain chai ligh so a doo cabine bedside shel
[172] 37.5 86.9 78.6 38.7 29.6 38.2 35.6 — 09.6 — 11.1 19.4 27.4 39.7 34.8 —
Ou s
SC 53.0 90.7 61.7 32.1 75.2 42.3 55.8 54.0 55.1 31.4 34.6 63.6 48.5 40.7 52.2 57.4
Ou s
EC 54.4 91.3 62.1 61.2 72.3 41.1 53.4 53.7 55.2 26.5 32.9 63.8 51.1 36.6 52.3 61.9
Table 3.5
Seman ic segmen a ion
esul s on SUN360 ex ended es se wi h ou
p oposed model Pano amic Bli zNe using s anda d con olu ions (SC) e sus equi ec -
angula con olu ions (EC) and compa ison wi h PanoCon ex [172].
Ou app oach p o es o wo k well on se e al di e en scenes by co ec ly sepa a ing
same ca ego y objec s, ha ini ially o e lapped in seman ic maps, in o di e en
ins ances. Limi a ions o he me hod can be seen when he ne wo k ails de ec ing
an objec , which is he e o e no di e en ia ed as an ins ance on he inal map and
when managing objec s wi h complex shapes ha can no be modelled wi h a gaussian
dis ibu ion.
He e, we inally analyze he imp o emen o e ou segmen a ion masks by le e aging
he con ex ual in o ma ion o he oom layou . In ou expe imen , logical assump ions
used o his e inemen en ail a signi ican imp o emen o up o 7
.
2% mIoU wi h
espec o he segmen a ion ou pu o Pano amic Bli zNe wi h S dCon s, achie ing
a inal
mIoU = 60.3%
. I should be no ed ha he classes ha con ibu e mos
o his imp o emen a e mi o , window and pic u e. Howe e , while one would also
expec a clea imp o emen in he doo ca ego y, we ha e seen a d op in pe o mance
in some cases such as he one shown in Figu e 3.5, al hough i de ini ely has a posi i e
e ec on i s loca ion in he 3D oom space. As al eady suppo ed by his p elimina y
expe imen , we p opose a p omising me hod o no iceably bene i 2D and 3D objec
ecogni ion asks om oom layou knowledge, and encou age he idea ha i is wo h
con inuing o wo k in his di ec ion.
3.5.5 Compa ison wi h he S a e o he A
De ec ion.
In Table 3.4 we show ou de ec ion esul s on he SUN360 ex ended
da ase . Ou Pano amic Bliz Ne wi h EquiCon s achie es e y sa is ac o y esul s,
wi h a global
mAP = 77.8%
. Fo comple eness, we include he e he esul s o [
32
],
ecen wo k on indoo pano amic objec ecogni ion wi h deep lea ning, oge he wi h
hei e alua ion o he De o mable Pa s Model (DPM) [
39
] on pano amas. Ou me hod
achie es he bes esul s in de ec ion o all 10 common classes compa ed o hem.
I is wo h no ing ha ou app oach achie es hese esul s jus aining wi h
∼
400
3.5 Expe imen al Resul s 79
Pano amicBli zNe S dCon s Pano amicBli zNe EquiCon s
Fig. 3.6
Quali a i e e alua ion o objec de ec ion and seman ic segmen a-
ion:
Examples o esul s ob ained wi h ou Pano amic Bli zNe using bo h s anda d
con olu ions and EquiCon s [41].
80 Objec Recogni ion
pano amas om he SUN360 da ase while hey use addi ional pano amas o ain hei
model. Since hei da ase is no public and no code is a ailable, we epo di ec ly
he esul s collec ed in [32].
Segmen a ion.
Table 3.5 summa izes he seman ic segmen a ion esul s on he
SUN360 ex ended da ase . A di ec compa ison is possible wi h he wo k o PanoCon-
ex [
172
]. The esul s clea ly show ha ou me hod signi ican ly imp o es o e he s a e
o he a . In pa icula , we add h ee new objec classes and boos
mIoU = 54.4%
,
which ep esen s an imp o emen o 16.9% o e PanoCon ex ’s me hod.
3.6 Conclusion
F om a single pano amic image, we p opose a me hod ha p o ides a comple e
unde s anding o he main objec s in an indoo scene. By managing he inhe en
cha ac e is ics and challenges ha equi ec angula pano amas in ol e, we ou pe o m
s a e o he a in addi ion o c ea ing a mo e comple e sys em, which no only ob ains
2D de ec ion and pixel-wise segmen a ion o objec s bu also places hem in o a 3D
econs uc ion o he oom. Exploi ing he ad an ages o ha ing a wide ield o iew
in indoo en i onmen s, his isual sys em becomes a p omising key elemen o u u e
au onomous mobile obo s. Fu u e wo k includes he inclusion o ins ance segmen a ion
p edic ions in o he deep lea ning pipeline and a u he s udy o he po en ial in
combining layou eco e y and objec ecogni ion asks.
Chap e 4
Objec Ca ego y Shape Modelling
“- Wha a e he h ee mos impo an p oblems in compu e ision? - Co espondence,
co espondence, co espondence!”
— Takeo Kanade
Au oma ic disco e y o ca ego y-speci ic 3D keypoin s om a collec ion o objec s o a
ca ego y is a challenging p oblem. The di icul y is added when objec s a e ep esen ed by
3D poin clouds, wi h a ia ions in shape and seman ic pa s and unknown coo dina e
ames. We de ine keypoin s o be ca ego y-speci ic, i hey meaning ully ep esen
objec s’ shape and hei co espondences can be simply es ablished o de -wise ac oss
all objec s in he ca ego y. We aim a lea ning such 3D keypoin s, in an unsupe ised
manne , using a collec ion o misaligned 3D poin clouds o objec s om an unknown
ca ego y. We model shapes de ined by he keypoin s using symme ic linea basis shapes
wi hou assuming he plane o symme y o be known. The usage o symme y p io
leads us o lea n s able keypoin s sui able o highe misalignmen s. To he bes o ou
knowledge, his is he i s wo k on lea ning such keypoin s di ec ly om 3D poin clouds
o a gene al ca ego y. Using objec s om ou benchma k da ase s, we demons a e
he quali y o ou lea ned keypoin s by quan i a i e and quali a i e e alua ions.

82 Objec Ca ego y Shape Modelling
4.1 In oduc ion
A se o keypoin s ep esen ing any objec is his o ically o la ge in e es o geome ic
easoning, due o hei simplici y and ease o handling. Keypoin s-based me hods [
89
,
143
,
9
] ha e been c ucial o he success o many ision applica ions. A ew examples
include; 3D econs uc ion [
97
,
28
,
130
], egis a ion [
166
,
75
,
91
,
87
], human body
pose [
127
,
96
,
19
,
14
], ecogni ion [
57
,
123
], and gene a ion [
140
,
169
]. Tha being said,
many keypoin s a e de ined manually, while conside ing hei seman ic loca ions such
as acial landma ks and human body join s, o add ess he p oblem a hand. To u he
bene i om hei widesp ead u ili y, se e al a emp s ha e been made on lea ning o
de ec keypoin s [
64
,
104
,
173
,
33
,
168
], as well as on au oma ically disco e ing hem
[
4
,
84
,
83
,
139
]. In his ega d, he ask o lea ning o de ec keypoin s om se e al
supe ision examples, has achie ed many successes [
155
,
104
]. Howe e , disco e ing
hem au oma ically om unlabeled 3D da a –such ha hey meaning ully ep esen
shapes and seman ics– so as o ha e a simila u ili y as hose o manually de ined, has
ecei ed only limi ed a en ion due o i s di icul y.
As objec s o in e es eside in he 3D space, i is no su p ising ha 3D keypoin s
a e p e e ed o geome ic easoning. Fo he gi en 3D keypoin s, hei coun e pa s in
2D images can be associa ed by me ely using came a p ojec ion models [
160
,
62
,
153
].
Howe e , being able o di ec ly p edic keypoin s on p o ided 3D da a (poin clouds)
has he ad an age ha he ask can be achie ed when mul iple came a iews o images
a e no a ailable. In his wo k, we a e in e es ed on lea ning keypoin s using only
3D s uc u es. In ac , 3D s uc u es wi h keypoin s su ice o se e al applica ions
including, egis a ion [
106
], shape comple ion [
94
], and shape modeling [
112
]; wi hou
equi ing hei 2D coun e pa s.
When 3D objec s go h ough shape a ia ions, due o de o ma ion o when wo
di e en objec s o a ca ego y a e compa ed, consis en keypoin s a e desi ed o
meaning ul geome ic easoning. Recall he examples o seman ic keypoin s such as
acial landma ks and body join s. To se e a simila pu pose, can we au oma ically
ind keypoin s ha a e consis en o e in e -subjec shape a ia ions and in a-subjec
de o ma ions in a ca ego y? This is he p ima y ques ion ha we a e in e es ed o
answe in his chap e . Fu he mo e, we wish o disco e such keypoin s di ec ly om
3D poin se s, in an unsupe ised manne . We call hese keypoin s “ca ego y-speci ic",
which a e expec ed o meaning ully ep esen objec s’ shape and o e hei co espon-
dence o de -wise ac oss all objec s. Mo e o mally, we de ine he desi ed p ope ies
o ca ego y-speci ic keypoin s as: i) gene alizabili y o e di e en shape ins ances
and alignmen s in a ca ego y, ii) one- o-one o de ed co espondences and seman ic
4.1 In oduc ion 83
consis ency, iii) ep esen a i e o he shape as well as he ca ego y while p ese ing
shape symme y. These p ope ies no only make he ep esen a ion meaning ul, bu
also end o enhance he use ulness o keypoin s. Lea ning ca ego y-speci ic keypoin s
on poin clouds, howe e , is a challenging p oblem because no all he objec pa s
a e always p esen in a ca ego y. The challenges a e exace ba ed when he p ac ical
cases o misaligned da a and unsupe ised lea ning a e conside ed. Rela ed wo ks
do no add ess all hese p oblems, bu ins ead op o ; d opping ca ego y-speci ici y
and using aligned da a [
83
], employing manual supe ision on 2D images [
104
], o
using aligned 3D and mul iple 2D images wi h known pose [
139
]. The la e me hod
achie es ca ego y-speci ici y wi hou explici ly easoning on he shapes. Ye ano he
wo k le e ages p ede ined local shape desc ip o s and a empla e model [
26
] speci ically
on aces.
In his chap e , we show ha he ca ego y-speci ic keypoin s wi h he lis ed p op-
e ies can be lea ned unsupe ised by modeling hem wi h non- igidi y, based on
unknown linea basis shapes. We u he impose an unknown e lec i e symme y
on he de o ma ion model, when conside ing ca ego ies wi h ins ance-wise symme y.
Fo ca ego ies whe e ins ance-wise symme y is no applicable, we p opose he use o
symme ic linea basis shapes in o de o be e model, wha we de ine as symme ic
de o ma ion spaces, e.g., human body de o ma ions. This allows us o be e cons ain
he pose and he shape coe icien s p edic ion. Ou p oposed lea ning me hod does
no assume aligned shapes [
139
], p e-compu ed basis shapes [
104
] o known planes
o symme y [
135
] and all quan i ies a e lea ned in an end- o-end manne . Ou sym-
me y modeling is powe ul and mo e lexible compa ed o ha o p e ious NRS M
me hods [
48
,
135
]. We achie e his by conside ing he shape basis o a ca ego y and
he e lec i e plane o symme y as he neu al ne wo k weigh a iables, op imized
du ing he aining p ocess. The aining is done on a single inpu , ci cum en ing he
Siamese-like a chi ec u e used in [
83
,
166
]. A in e ence ime, he ne wo k p edic s
he basis coe icien s and he pose in o de o es ima e he ins ance-speci ic keypoin s.
Using mul iple ca ego ies om ou benchma k da ase s, we e alua e he quali y o
ou lea ned keypoin s bo h quan i a i ely and wi h quali a i e isualiza ion. Ou
expe imen s show ha he keypoin s disco e ed by ou me hod a e geome ically and
seman ically consis en , which a e measu ed espec i ely by in a-ca ego y egis a ion
and seman ic pa -wise assignmen s. We u he show ha symme ic basis shapes can
be used o model symme ic de o ma ion space o ca ego ies such as he human body.
84 Objec Ca ego y Shape Modelling
4.2 Rela ed Wo k
Ca ego y-speci ic keypoin s on objec s ha e been ex ensi ely used in NRS M me hods,
howe e , only ew me hods ha e ackled he p oblem o es ima ing hem. In e ms o
he ou come, ou wo k is closes o [
139
], which lea ns ca ego y-speci ic 3D keypoin s
by sol ing an auxilia y ask o igid egis a ion be ween mul iple ende s o he same
shape and by conside ing he ca ego y ins ances o be p e-aligned. Al hough he
me hod shows p omising esul s on 2D and 3D, i does so wi hou explici ly modeling
he shapes. Consequen ly, i equi es ende s o di e en ins ances o be p e-aligned o
eason on keypoin co espondences be ween ins ances. A simila ask is also sol ed
in [
104
] o 6-deg ees o eedom (DoF) es ima ion which uses low- ank shape p io
o condi ion keypoin s in 3D. Al hough, he low- ank shape modeling is a powe ul
ool, [
104
] equi es supe ision o hea map p edic ion and elies on aligned shapes and
p e-compu ed shape basis. [
155
] also p edic s keypoin s o ca ego ies wi h low- ank
shape p io bu he me hod is again ained on ully supe ised manne . Mo eo e ,
all o he men ioned me hods lea n keypoin s on images as hea maps and he ea e
li hem o 3D. Di e en om he o he wo ks, [
26
] exploi s de o ma ion model and
symme y o di ec ly p edic keypoin s on 3D bu equi es a ace empla e, aligned
shapes and known basis. Shape modeling o ca ego y shape ins ances has been widely
explo ed in NRS M wo ks. Linea low- ank shape basis[
16
,
145
,
28
], low- ank ajec o y
basis [
3
], isome y o piece-wise igidi y [
142
,
102
] a e some o he di e en me hods
used o NRS M. Recen ly, a ew numbe o wo ks ha e used low- ank shape basis in
o de o de ise lea ned me hods [
97
,
76
,
155
,
135
]. Ano he use ul ool in modeling
shape ca ego y is he e lec i e symme y, which is also di ec ly ela ed o he objec
pose. Al hough [
48
] showed ha he low- ank shape basis can be o mula ed wi h
unknown e lec i e symme y, i s adap a ion o lea ned NRS M me hods is no i ial.
Recen me hods, in ac , assume ha he plane o symme y is one among a ew
known planes [
154
]. Mo eo e , none o he me hods o mula e symme y applicable o
non- igidly de o ming objec s such as he human body. A pa allel wo k [
156
] on his
ega d models symme y p obabilis ically in a wa ped canonical space o econs uc
3D o di e en objec s.
While shape modeling is a key aspec o ou wo k, ano he challenge is o in e
o de ed keypoin s by lea ning on uno de ed poin se s. Despi e se e al ad ances on
deep neu al ne wo ks o poin se s [
107
,
108
,
149
], cu en achie emen s o lea ning
on images dwa hose o lea ning on poin se s. A ela ed wo k lea ns o p edic
3D keypoin s unsupe ised by again sol ing he auxilia y ask o co ec ly es ima ing
o a ions in a Siamese a chi ec u e [
17
]. The keypoin p edic ion is done wi hou
4.3 Backg ound and Theo y 85
o de by pooling ea u es o ce ain poin neighbo hoods. Ano he p e ious wo k [
166
]
p oposes lea ning poin ea u es o ma ching, again using alignmen as he auxilia y
ask. Ma ching such keypoin s ac oss shapes is no an easy ask as he keypoin s a e
no p edic ed in any o de . In he ollowing sec ions we show how one can model shape
ins ances using he low- ank symme ic shape basis and use he shape modeling o
p edic o de ed ca ego y-speci ic keypoin s.
4.3 Backg ound and Theo y
4.3.1 Ca ego y-speci ic Shape and Keypoin s
We ep esen shapes as poin clouds, de ined as an uno de ed se o poin s S=
{
s
1,
s
2,...,
s
M},
s
j∈R3
,
j∈ {
1
,
2
, . . . , M}
. The se o all such shapes in a ca ego y
de ines he ca ego y shape space
C
. We w i e a pa icula
i
- h ca ego y-speci ic shape
ins ance in
C
as S
i
. Fo con enience, we will use he e ms ca ego y-speci ic shape
and shape in e changeably. The ca ego y shape space
C
can be any hing om a
se o disc e e shapes o a smoo h mani old o ca ego y-speci ic shapes spanned by
a de o ma ion unc ion Ψ
C
. The ocus o he wo k is on lea ning meaning ul 3D
keypoin s om he poin se ep esen a ion o S
i
. To ha end, his sec ion de ines
ca ego y-speci ic keypoin s and de elops hei modeling.
Ca ego y-speci ic keypoin s.
We ep esen ca ego y-speci ic keypoin s o a shape
S
i
as a spa se uple o poin s, P
i
= (p
i1,
p
i2,...,
p
iN
)
,
p
ij ∈R3
,
j∈ {
1
,
2
, . . . , N}
.
Unlike he shape, i s keypoin s a e ep esen ed as o de ed poin s. Ou objec i e is
o lea n a mapping Π
C
:S
i→
P
i
in o de o ob ain he ca ego y-speci ic keypoin s
om an inpu shape S
i
in
C
. Al hough no comple ely unambiguous, we can de ine
he ca ego y-speci ic keypoin s using he p ope ies lis ed in Sec. 4.1. In ma hema ical
no a ions hey a e:
(i) Gene aliza ion: ΠC(Si) = Pi,∀Si∈ C.
(ii)
Co esponding poin s and seman ic consis ency: Gi en S
a,
S
b∈ C
, we wan
paj ⇔pbj. Simila ly, paj and pbj should ha e he same seman ics.
(iii)
Rep esen a i e-ness:
ol
(S
i
) =
ol
(P
i
)and p
ij ∈
S
i
, whe e
ol(.)
is he Volume
ope a o o a shape. I S
i∈ C
has a e lec i e symme y, P
i
should ha e he
same symme y.
92 Objec Ca ego y Shape Modelling
he de o ma ion unc ion P
i
= Φ
C
(R
i,
c
i
;
BC,
n
C
), hus ob aining P
i
=X
i
. Howe e , as
con i med by ou e alua ions as well as in [
83
], he
ℓ2
loss does no con e ge as he
ne wo k is unable o p edic he poin o de . Al e na i ely, he Cham e loss [
37
] does
con e ge, minimizing he dis ance be ween each poin x
ik
in he i s se X
i
and i s
nea es neighbo pij in he second se Piand ice e sa.
Lch =
N
X
k=1
min
pij ∈Pi
∥xik −pij∥2
2+
N
X
j=1
min
xik∈Xi
∥xik −pij∥2
2,(4.6)
The Cham e loss in Eq.
(4.6)
ensu es ha he lea ned keypoin s ollow a gene aliz-
able ca ego y-speci ic p ope y – ha hey a e a linea combina ion o common basis
lea ned speci ically o he ca ego y. To addi ionally model symme y, Eq.
(4.3)
o
(4.4)
is di ec ly used in Eq.
(4.6)
. The e o e, wo di e en Cham e losses a e possible
modeling wo di e en ypes o symme ies.
Co e age and inclusi i y loss.
The Cham e loss in Eq.
(4.6)
does no ensu e ha
he keypoin s ollow he objec shape. Howe e , one can add he ollowing condi ions:
a) he keypoin s co e he whole ca ego y shape (co e age loss), b) he keypoin s a e
no a om he poin cloud (inclusi i y loss). The co e age loss can be de ined as a
Hube loss be ween he olume o he nodes X
i
and ha o he inpu shape S
i
, using
he p oduc o he singula alues. Howe e , we ins ead app oxima e he olume using
he 3D bounding box de ined by he poin s. This imp o es he aining speed and,
based on ou ini ial e alua ions, also does no ha m pe o mance. The co e age loss is
hus gi en by:
Lco =∥ ol(Xi)− ol(Si)∥(4.7)
The inclusi i y loss is o mula ed as a single side Cham e loss [
13
] which penalizes
nodes in Xi ha a e a om he o iginal shape Si, simila ly o Eq. (4.6):
Linc =
N
X
k=1
min
sij ∈Si
∥xik −sij∥2
2.(4.8)
4.5 Expe imen al Resul s
We conduc expe imen s o e alua e he desi ed p ope ies o he p oposed ca ego y-
speci ic keypoin s and show hei gene aliza ion o e indoo /ou doo objec s and

4.5 Expe imen al Resul s 93
igid/non- igid objec s wi h ou di e en da ase s in o al (Sec ions 4.5.1 and 4.5.2).
All hese p ope ies a e also compa ed wi h a p oposed baseline. We hen e alua e
he p ac ical use o ou keypoin s o in a-ca ego y shapes egis a ion (Sec ion 4.5.3),
analyzing he in luence o symme y, and o segmen a ion label ans e (Sec ion 4.5.4).
Fu he mo e, an expe imen showing he gene aliza ion o ou me hod on eal da a is
included in Sec ion 4.5.5. Addi ional quali a i e esul s a e shown in Sec ion 4.5.6.
Da ase s.
We use ou main da ase s. They a e ModelNe 10 [
157
], ShapeNe
pa s [167], Dynamic FAUST [15] and Basel Face Model 2017 [49]. Since ou me hod
is ca ego y-speci ic, we equi e sepa a e aining da a o each class in he da ase s.
Fo indoo igid objec s, we choose h ee ca ego ies om ModelNe 10 [
157
]; chai ,
able and bed. Th ee ou doo igid objec ca ego ies: ai plane, ca and mo o bike, a e
e alua ed om ShapeNe pa s [
167
]. Fo non- igid objec s, we andomly choose a
sequence o he Dynamic Faus [
15
], ha p o ides high- esolu ion 4D scans o human
subjec s in mo ion. Finally, we gene a e shape models o aces using he Basel Face
Model 2017 [
49
] combining 50 di e en shapes and 20 di e en exp essions. All models
a e no malized in he ange
−
1 o 1and a e andomly misaligned wi hin
±
45 deg ees.
Baseline.
Since his is he i s wo k compu ing ca ego y-speci ic keypoin s om
poin se s, we cons uc ou own baseline based on he ecen wo k USIP [
83
]. The
me hod de ec s s able in e es poin s in 3D poin clouds unde a bi a y ans o ma ions
and is also unsupe ised, which makes i he closes me hod o compa ison. The
USIP de ec o is no ca ego y-based, so we ain he ne wo k pe ca ego y o c ea e
he baseline. Addi ionally, we adap he numbe o p edic ed keypoin s so ha he
esul s a e di ec ly compa able o ou s. While aining wi h some o he ca ego ies,
speci ically ca and bed, we obse e ha p edic ing lowe numbe o keypoin s can
lead o some degene acies [83].
Implemen a ion de ails.
Inpu poin clouds o dimension 3
×
2000 a e used. We
implemen he ne wo k in Py o ch [
103
] and ain i end- o-end om sc a ch using he
Adam op imize [
73
]. The ini ial lea ning a e is 10
−3
, which is exponen ially decayed
by a a e o 0
.
5e e y 40 epochs. We use a ba ch size o 32 and ain each model un il
con e gence, o 200 epochs. The inal loss unc ion combines he h ee aining losses,
Eqs.
(4.6)
,
(4.7)
and
(4.8)
, and a e weigh ed as ollows:
wch
=
wco
= 1 and
winc
= 2.
Fo ModelNe 10 and ShapeNe pa s, we use he aining and es ing spli p o ided
by he au ho s. Fo he Basel Face Model 2017, we ollow he common p ac ice and
94 Objec Ca ego y Shape Modelling
spli he 1000 gene a ed aces in 85% aining and 15% es . We use he same spli
s a egy o he sequence ‘50009_jiggle_on_ oes’ o Dynamic Fuaus , which con ains
244 examples.
Ca ego y Co e age Model E Co espondence Inclusi i y Sym E De ini ion
% % % % ◦
chai 88.83 0.72 100 90.46 0.40 10
able 93.33 0.99 100 93.38 2.86 6
bed 80.31 0.94 100 95.33 0.13 6
ai plane 89.15 0.64 100 96.35 0.20 8
ca 92.39 0.72 100 97.77 2.21 8
mo o bike 96.13 0.79 100 90.53 1.42 8
human body 85.59 0.72 100 97.73 33.30 11
aces 97.93 0.41 100 100 0.15 9
chai 79.73 −55.698.50 −10
able 79.72 −34.599.83 −6
bed 42.18 −49.33 70.00 −6
ai plane 69.24 −47.5 87.13 −8
ca 26.87 −32.18 74.0−8
mo o bike 75.29 −48.14 84.57 −8
human body 72.66 −50.45 100 −11
aces 42.98 −30.11 100 −9
Table 4.1 P ope ies Analysis: Top (ou s) and bo om (baseline [
83
]). Fo co e age,
co espondence and inclusi i y highe is be e , and o model and symme y e o
lowe is be e . We empi ically show he desi ed p ope ies o ou keypoin s, as well
as he gene aliza ion o ou me hod o e indoo /ou doo and igid/non- igid objec s.
Bes esul s a e in bold.
4.5.1 Desi ed P ope ies Analysis
As desc ibed in Sec. 4.1 and 4.3, he ca ego y-speci ic keypoin s sa is y ce ain desi ed
p ope ies. We p opose six di e en me ics o e alua e he p ope ies which a e also
used o compa ison agains he baseline. All he esul s a e p esen ed in Table 4.1,
and a e a e aged ac oss he es samples.
4.5 Expe imen al Resul s 95
Fig. 4.3
Keypoin s co espondence/ epea abili y ac oss ins ances
. We clus e
he p edic ed keypoin s o all he ins ances in he ca ego y o show hei geome ic
consis ency. No e how ou keypoin s a e nea ly clus e ed as hey a e consis en ly
p edic ed in he co esponding geome ic loca ions, unlike he baseline keypoin s.
(No e: clus e colo s do no co espond o keypoin colo s.)
Co e age: Acco ding o p ope y iii), we seek keypoin s ha a e ep esen a i e o
each ins ance shape as well as o he ca ego y i sel . To measu e i , we calcula e he
pe cen age o he inpu shape co e ed by he keypoin s’ 3D bounding box. On a e age,
we achie e a 29.4% mo e co e age han he baseline.
Model E o : This me ic e e s o he Cham e dis ance be ween he es ima ed nodes
and he lea ned ca ego y-speci ic keypoin s, no malized by he model’s scale. We ob ain
less han 1% o e o in all he ca ego ies, meaning ha he ne wo k sa is ac o ily
manages o gene alize, desc ibing he nodes wi h he symm e ic non- igidi y modeling
(P ope ies i) and iii)).
Co espondence/ Repea abili y: We measu e he abili y o he model o ind he same
se o keypoin s on di e en ins ances o a gi en ca ego y (P ope y ii)). Fo ou
me hod, we clus e he keypoin s using hei inhe en o de whe eas o he baseline,
we use K-means clus e ing o e alua e and compa e his p ope y. We show a de ailed
e alua ion o he chai ca ego y in Fig. 4.4, he es o he ca ego ies a e p o ided
in Fig. 4.3. One can see a a glance how ou keypoin s a e well clus e ed, unlike he
baseline keypoin s. Nume ically, we show he % occu ence o each speci ic keypoin
belonging o he same clus e ac oss ins ances. Ou keypoin s sa is y 100% he co e-
spondence/ epea abili y es hanks o ou geome ic non- igidi y modelling.
96 Objec Ca ego y Shape Modelling
Chai
Ou s USIP USIP USIP
Ou s
Table
Ou s
Bed
USIP USIP USIP
Ou s
Ca Mo o bike
USIP
Faces
Ou s
USIP
Human body
Ai plane
Ou s
Ou s
Ou s
Fig. 4.4
Keypoin s co espondence ac oss ins ances
. We clus e he keypoin s
p edic ed o all he ins ances o a ca ego y o show hei geome ic consis ency. No e
how ou keypoin s ge nea ly clus e ed c ea ing a gene al 3D shape empla e.
Inclusi i y: We measu e he pe cen age o keypoin s ha lie inside he poin cloud
(o scale 2) wi hin a chosen h eshold o 0.015, which also p o es p ope y iii). This
is he only me ic in which ou me hod doesn’ ou pe o m he baseline in all cases.
On a e age, ou me hod achie es
∼
95% inclusi i y compa ed o
∼
89% o he baseline.
Symme y: The me ic shows he angle e o o he p edic ed e lec i e plane o symme-
y. We ob ain highly accu a e p edic ion o igid ca ego ies. In he non- igid human
body shape howe e , he ambigui ies a e se e e. Despi e ha , he lea ned keypoin s
sa is y he o he p ope ies, pa icula ly ha o seman ic co espondence.
De ini ion: inal numbe o keypoin s
N′
p edic ed pe ca ego y a e he Non-Maximal
Supp ession.
4.5.2 Seman ic Consis ency
We use he ShapeNe pa da ase [
167
] o show he seman ic consis ency o he
p oposed keypoin s. Following he low- ank non- igidi y modelling, he keypoin s lie
4.5 Expe imen al Resul s 97
kp1
kp2
kp3
kp4
kp5
kp6
kp7
kp8
body
wings
ail
0.10 0.96 0.02 0.95 0.97 0.08 0.02 0.07
0.01 0.03 0.97 0.02 0.01 0.90 0.97 0.87
0.89 0.00 0.01 0.03 0.01 0.00 0.01 0.02
Ai plane
oo
body
wheels
0.00 0.80 0.00 0.00 0.00 0.00 0.00 0.00
1.00 0.20 1.00 1.00 0.00 1.00 1.00 0.10
0.00 0.00 0.00 0.00 0.90 0.00 0.00 0.90
0.00 0.00 0.00 0.00 0.10 0.00 0.00 0.00
Ca
wheels
handle
body
gas ank
sea
ligh
1.00 1.00 0.50 1.00 0.25 0.00 0.00 0.00
0.00 0.00 0.00 0.00 0.00 0.00 1.00 1.00
0.00 0.00 0.50 0.00 0.75 1.00 0.00 0.00
0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
Mo o bike
0.2 0.4 0.6 0.8
0.0 1.0
body wings/ wheels ail/ oo / handle o he s keypoin s
kp1
kp2
kp3
kp4
kp5
kp6
kp7
kp8
kp1
kp2
kp3
kp4
kp5
kp6
kp7
kp8
oo
body
wheels
0.79 0.21 0.76 0.73 0.83 0.66 0.25 0.14
0.06 0.75 0.23 0.27 0.16 0.20 0.37 0.42
0.14 0.03 0.00 0.00 0.01 0.13 0.38 0.44
wheels
handle
body
body
wings
ail
0.05 0.05 0.03 0.00 0.00 0.00 0.00 0.01
0.94 0.83 0.80 0.84 0.91 0.95 0.93 0.93
0.00 0.01 0.01 0.03 0.03 0.04 0.07 0.06
kp1
kp2
kp3
kp4
kp5
kp6
kp7
kp8
kp1
kp2
kp3
kp4
kp5
kp6
kp7
kp8
kp1
kp2
kp3
kp4
kp5
kp6
kp7
kp8
0.23 0.04 0.26 0.28 0.09 0.26 0.28 0.13
0.00 0.00 0.00 0.09 0.00 0.00 0.00 0.00
0.77 0.94 0.72 0.51 0.74 0.72 0.72 0.85
Fig. 4.5
Seman ic pa co espondence.
Top o bo om: he seman ic co e-
spondence o he p oposed keypoin s, quali a i e esul s and he baseline seman ic
co espondence. Ou p edic ed keypoin s show he co ec seman ic co espondence
ac oss he ca ego y.
on geome ically co esponding loca ions. The idea o he expe imen is o measu e
keypoin -seman ics ela ionship o e e y keypoin ac oss ins ances o he ca ego y.
The esul s a e p esen ed in Fig. 4.5 as co a iance ma ices, along wi h keypoin
isualiza ions pe ca ego y o ou me hod. On a e age, he p oposed keypoin s ha e
a high seman ic consis ency o 93% ac oss ins ances, despi e he la ge in a-ca ego y
a iabili y. The same expe imen is pe o med o he baseline and p esen ed in bo om
o Fig. 4.5. He e, he degene acy causes all he keypoin s o app oach he objec
cen oid o ‘Ca ’. None heless, we obse e no seman ic consis ency e en o ‘Ai plane’
wi hou degene acies. Ou model, aiming o a common ep esen a ion o all he
ins ances o he ca ego y, a oids placing keypoin s in less ep esen a i e pa s o
unique pa s, e.g., a m es s in chai s (in Fig. 4.9), engines in ai planes o gas ank in
mo o bikes. This highligh s signi ican obus ness achie ed in modelling and lea ning
he keypoin s.
4.5.3 Objec s Pose and In a-ca ego y Regis a ion
P e ious me hods do no handle misaligned da a due o he ob ious di icul y i poses
o unsupe ised lea ning. This dese es special a en ion since eal da a is ne e
aligned. In his sec ion we e alua e he in a-ca ego y egis a ion pe o mance o ou

98 Objec Ca ego y Shape Modelling
model and show he impac o he di e en symme y models p oposed. These esul s
implici ly measu e he objec poses es ima ed as well.
Ro a ion Ambigui ies.
Recen unsupe ised app oaches o keypoin de ec ion
ac ually sel -supe ise o a ion du ing aining, e.g., [
139
,
83
], and highligh ha i is
c ucial o achie ing a good pe o mance. In ou case, we do no di ec ly supe ise he
o a ions. The e o e, he di e en combina ion o basis shapes can esul in di e en
alignmen s. This implies ha compu ing P
i
wi h he de o ma ion unc ion Φ
C
will
gi e he co ec se o keypoin s along wi h he co ec plane o symme y, bu he
p edic ed o a ion alone is no meaning ul o egis a ion. As we show in Fig. 5 in
he ex , p edic ing he symme y plane o he objec ca ego y allows o ha e mo e
con ol o e he p edic ed ins ance poses. We came up wi h he idea o lea ning an
addi ional common pa ame e , R
C
, which is di ec ly ela ed o he symme y plane. By
adding his ca ego y-speci ic pa ame e , he ne wo k lea ns a common o a ion o all
he objec s in he ca ego y. As a consequence, he ins ance-wise o a ion, R
i
, can be
hough like an o se om he e e ence basis alignmen . Se e al e alua ions con i med
ha his s a egy helps he lea ning p ocess, educing he o a ion ambigui ies.
Expe imen al se up.
Despi e he abo e ambigui y, an impo an cha ac e is ic
o he p oposed keypoin s is ha hey a e o de ed, which empowe s di ec in e -
ins ances egis a ion since no ex a desc ip o s a e needed o ma ching. We pe o m
expe imen s o he chai ca ego y, using 10 keypoin s (Table 4.1) and a misalignmen
o
±
45 deg ees. Th ee di e en models a e compa ed. The i s one is ained wi hou
symme y awa eness ollowing Eq.
(4.2)
. A second one uses shape symme y du ing
aining as shown in Eq.
(4.3)
. The las model is ained wi h basis symme y as in
Eq.
(4.4)
. We a emp o egis e keypoin s in each ins ance o hose o andomly
chosen h ee aligned empla es by compu ing a simila i y ans o ma ion and obse e
he mean e o . Fig. 4.6 shows ha symme y helps o ha e mo e con ol o e he
o a ions and ackle highe misalignmen .
4.5.4 Segmen a ion Label T ans e
Ou p edic ed keypoin s co espond o seman ically meaning ul loca ions. The e o e,
he e we explo e he u ili y o he p oposed ca ego y-speci ic keypoin s o he segmen a-
ion label ans e ask. In his expe imen , o e e y poin in he o iginal shape s
ij ∈
S
i
,
we ind i s closes ca ego y-speci ic keypoin p
ik ∈
P
i
, and ans e he co esponding
4.5 Expe imen al Resul s 99
Inpu Ro a ion (°)
RRE (°)
15 20 25 30 35 40 45
9
8
7
6
5
4
3
2
1
No Symme y
Basis Symme y
Shape Symme y
In a-ca ego y egis a ion
Fig. 4.6 Le : Rela i e o a ion e o o di e en symme y modelings. Righ : 3
examples o egis a ion be ween di e en ins ances o he same ca ego y.
Fig. 4.7
Fi s ow:
esul s o pe o ming seman ic label ans e wi h ou keypoin s.
Second ow:
g ound u h. This is e alua ed in ShapeNe pa da ase [
167
] using
eigh keypoin s o he label ans e .
seman ic label o i . We assume he keypoin s labels a e known and co espond o
hose in Figu e 4.5.
Some quali a i e esul s a e shown in Fig. 4.7. Ou me hod achie es ull co espon-
dence be ween ins ances, he e o e a oiding placing keypoin s in less ep esen a i e
pa s. An example is he engine, in g ey, in he case o ai planes. This is e lec ed
in he label ans e since he e is no dis inc ion o hese pa s. Besides ha , only
wi h eigh keypoin s in he example, we achie e easonable esul s, close o he g ound
u h da a.
4.5.5 Real Da a
In his sec ion, we show he pe o mance o ou me hod o eal da a in Fig. 4.8. Fo
his expe imen , he ne wo k is ained on he chai ca ego y om he ModelNe 10
da ase [
157
] and es ed on eal chai s om he SUNRGBD da ase [
131
]. To gene a e
100 Objec Ca ego y Shape Modelling
Fig. 4.8 Resul s in eal chai s om SUNRGBD da ase [
131
] aining wi h CAD chai s
om ModelNe 10 da ase [157].
he eal da a da ase om [
131
], we c op he poin s inside he g ound u h 3D bounding
boxes p o ided by he au ho s. Real da a en ail addi ional challenges. This is no
only because shapes appea incomple e and noisy, bu also because o he objec s may
cause occlusions, e.g. pa o a able occluding a chai . As illus a ed in Fig. 4.8, e en
hough eal da a is ai ly challenging, ou ne wo k can s ill p oduce co esponding
meaning ul keypoin s.
Being able o gene alize o p e iously unseen eal objec s as demons a ed in Fig.
4.8 is c ucial and eally use ul o many asks such as guide o shape comple ion o
shape gene a ion.
4.5.6 Quali a i e esul s
In his sec ion, we p o ide addi ional quali a i e esul s on a ious objec ca ego ies
om he da ase s e alua ed; ModelNe 10 [
157
] in Fig. 4.9, ShapeNe pa s [
167
] in Fig.
4.10, Dynamic FAUST [15] in Fig. 4.11 and Basel Face Model 2017 [49] in Fig. 4.12.
Again, we no e ha ou ne wo k p edic s co esponding keypoin s be ween ins ances
o he same ca ego y and consis en ly associa es he same keypoin wi h he same
seman ic pa . Fo ins ance, o he chai objec ca ego y, he keypoin colo ed in pink
4.6 Conclusions 101
Fig. 4.9 Quali a i e esul s in able, chai and bed ca ego ies om ModelNe 10 da ase
[157].
is always associa ed wi h he chai back, he keypoin colo ed in cyan is associa ed
wi h he on le leg, e c.
4.6 Conclusions
This wo k in es iga es au oma ic disco e y o kepoin s in 3D misaligned poin clouds
ha a e consis en o e in e -subjec shape a ia ions and in a-subjec de o ma ions
in a ca ego y. We ind ha his can be sol ed, wi h unsupe ised lea ning, by modeling
keypoin s wi h non- igidi y, based on symme ic linea basis shapes. Addi ionally,
he p oposed ca ego y-speci ic keypoin s ha e one- o-one o de ed co espondences
and seman ic consis ency. Applica ions o he lea ned keypoin s include egis a ion,
ecogni ion, gene a ion, shape comple ion and many mo e. Ou expe imen s showed
ha high quali y keypoin s can be ob ained using he p oposed me hods and ha he
me hod can be ex ended o complex non- igid de o ma ions. Fu u e wo k could ocus
on be e modeling complex de o ma ions wi h non-linea app oaches.
108 Discussion and Conclusions
p oposed in Sec ion 2.4. This pos -p ocessing will lead o an inc ease o he compu ing
ime, bu also will help ill he non- isible co ne s. A second op ion o a oid inc easing
he compu ing ime, could be o p edic di ec ly he o de o he oom co ne s inside
he ne wo k. Howe e , he me hod could s ill ail i any co ne is no p edic ed.
We ha e come a long way in he ask o layou eco e y, bu we a e s ill a om
achie ing human obus ness and gene aliza ion on his p oblem. The nex s ep should
ocus on sol ing inc easingly complex oom geome ies and aim o eal- ime algo i hms.
The e a e de ini ely many exci ing ideas o be ied like; exploi ing symme y o p edic
mo e meaning ul oom co ne s, elaxing he Manha an Wo ld assump ion o adding
mo e geome ic cons ain s inside he deep lea ning algo i hms. All hese ideas would
no only help o achie e mo e accu a e esul s, bu also o head owa ds unsupe ised
lea ning, p o iding mo e scalable amewo ks.
5.2 Objec Recogni ion
In Chap e 3we p esen , up o ou knowledge, he i s objec de ec ion sys em wo king
di ec ly on 360 images, ocused on indoo scene unde s anding. We s udy how o adap
exis ing CNNs, in his case designed o he ask o objec de ec ion, o ma ch he
na u e o he equi ec angula image inpu . We adap he ancho box p oposals and
subs i u e s anda d con olu ions by EquiCon s. We al eady demons a ed in Sec ion
2.6, how EquiCon s can help o gene alize o di e en came a pose a ia ions. Fo he
ask o objec ecogni ion, we obse ed ha he use o EquiCon s is con enien e en i
he came a is always a he same place, since objec s can be a many di e en loca ions
inside he scene, e.g., objec s close o he came a will ha e g ea e dis o ions han
objec s a ound he ho izon line. Addi ionally, since he dis o ion depends only on he
pola angle, objec s ha appea o hogonal o he iewpoin will be symme ically
dis o ed, whe eas objec s a di e en poses will no . The e o e, EquiCon s he e
play an impo an ole since hey can lea n objec appea ance by igno ing sphe ical
dis o ion pa e ns. Addi ionally, we show he po en ial o exploi ing he 2D oom
layou o imp o e he ins ance segmen a ion masks, and how o le e age he 3D layou
o gene a e 3D objec bounding boxes, di ec ly om he imp o ed segmen a ion masks.
A limi a ion is ha he e is no lea ning on he 3D s uc u e o objec de ec ion.
We s ongly belie e ha including he layou p io and he 2D-3D li ing inside he
ne wo k, would imp o e ou esul s.

5.3 Objec Ca ego y Shape Modelling 109
5.3 Objec Ca ego y Shape Modelling
In Chap e 4we p opose o lea n 3D keypoin s om a collec ion o objec s o some
ca ego y, so ha hey meaning ully ep esen objec s’ shape and hei co espon-
dences can be simply es ablished o de -wise ac oss all objec s. Ou mo i a ion is ha
keypoin s-based me hods a e c ucial o he success o many ision applica ions like
3D econs uc ion, egis a ion, human body pose, ecogni ion, o gene a ion. The
challenges we conside a e he ollowing: inpu shapes a e misaligned 3D poin clouds,
3D objec s go h ough shape a ia ions and a gi en seman ic pa may no be p esen
in all objec s in a ca ego y. We demons a e ha his p oblem can be sol ed, in an
unsupe ised manne , by modeling keypoin s wi h non- igidi y, based on symme ic
linea basis shapes. We do no assume he plane o symme y o be known and conside
wo di e en p io s: ins ance-wise symme y ( igid objec s) and symme ic de o ma ion
space (non- igid objec s). We show ha he keypoin s disco e ed by ou me hod ha e
one- o-one o de ed co espondences and a e geome ically and seman ically consis-
en . A limi a ion o his wo k is ela ed o he model pe o mance on o ganic shapes
like he human body bu also p obably on animals, o gans and non- igid shapes in
gene al. Co espondences be ween pai s o such de o mable shapes has been ackled
using shape simila i ies, unc ional maps [
100
], e c. These me hods usually ely on
g ound u h co espondences, p e-compu ed desc ip o s, shape empla es o aligned
da a. Howe e , ackling co espondences be ween a collec ions o non- igid objec s
is conside ably mo e challenging. P e ious me hods ha e explo ed echniques such
us pa h in a iance o cycle consis ency. The challenges add u he when no labels
a e a ailable. We belie e ha i would be in e es ing o combine low ank cons ain s
wi h unc ional maps and cycle consis ency echniques. Ano he op ion ha seems
in e es ing, would be o explo e mul ilinea (bilinea ) models o analyze shape and
pose a ia ions independen ly [56].
5.4 Fu u e Wo k
Imp o emen s o he indi idual p oposed wo ks ha e been discussed abo e. To conclude
his hesis, we would like o highligh some aspec s o u u e esea ch, which a e
pa icula ly exci ing o us, wi h he goal o de eloping in elligen sys ems ha ma ch
he pe o mance o human ision.
In his hesis we ad ance s a e o he a in se e al opics ela ed o indoo scene
unde s anding bu o cou se, he e is much mo e ou he e. Aiming o a comple e
110 Discussion and Conclusions
scene unde s anding, complemen a y asks as well as ways o connec ing se e al asks
oge he a e desi ed. In his ega d, once we ge o au oma ically unde s and he
geome y o an indi idual oom, we migh be in e es ed in es ima ing how he oom
is connec ed o o he ooms, o how is he building dis ibu ed. Simila ly, once we
iden i y he objec s inside a oom and know hei loca ion inside he oom, we ca e
abou he ela ionships be ween he en i ies o abou he ac ions we can pe o m wi h
hem, e.g. chai s a e usually a ound a able, we can si on a chai , lay on a bed, e c.
In he same way, once we a e awa e o he shape model and geome y o a pa icula
objec , i is in e es ing o know i s colo , ma e ial and physical p ope ies in gene al. A
ecen wo k [
5
] demons a es how o hos hese di e se ypes o seman ics in a uni ied
s uc u e. They gene a e a 3D scene g aph using a 3D mesh and egis e ed pano amic
images o he building, and combine exis ing de ec ion me hods in o de o collec
all he in o ma ion. The p oposed me hod is s ill semi-au oma ic and lea es oom o
many exci ing imp o emen s and no el ideas. A simila ecen wo k [116] models he
scene dynamics as well, e.g. a e sabili y be ween places o ooms: “agen A is in
oom B a ime ”. S ill, i we wan o mimic he human isual pe cep ion, we should
aspi e o es ima e such scene g aphs inc emen ally and in eal- ime.
The e a e many ob ious applica ions o ge ing such uni ied unde s anding which
ha e been al eady discussed in his hesis, such as indoo na iga ion o i ual o
augmen ed eali y. One less ob ious bu e y exci ing applica ion is o ans e he
indoo space in o ma ion o he Building In o ma ion Modelling (BIM) me hodology,
o model he exis ing building s ock, ei he o acili y managemen pu poses, he i age
conse a ion, building esea ch p ojec s o s uc u al s abili y analyses. This echnol-
ogy is impo an as i inc eases he in e ope abili y be ween mul iple he e ogeneous
disciplines such as a chi ec u e, cons uc ion, plumbing, ligh ing/elec ical, mechanical
o enginee ing. Howe e , his line o wo k s ill needs a lo o e o , as we need o
consis en ly model he ou doo and indoo pa s o he building, ge as much de ails
as possible om he s uc u al componen s, and ind a new da a ep esen a ion ha
acili a es he image (o poin cloud) o BIM model con e sion. Compu e Vision
applica ions in gene al, ela ed o indoo scene unde s anding, a e al eady e y p esen
in ou socie y, and al hough some sec o s emain skep ical, many ha e al eady emb aced
his echnology. We ha e jus expe ienced an unp eceden ed e en in he las 100
yea s, he co ona i us. This disease, known as COVID-19, has hi he en i e wo ld
popula ion, making us wonde how we can change he u u e, and mo e speci ically,
how we can c ea e a socie y ha su e s less exposu e o his ype o diseases. F om he
Compu e Vision side, i is mo e u gen han e e be o e ha we specially con ibu e
5.4 Fu u e Wo k 111
c ea ing in elligen sys ems ha a e able o assis and in e ac wi h humans, making
he discipline mo e powe ul and ubiqui ous. This will no only ha e a di ec impac
o help on hese c i ical si ua ions, whe e ace- o- ace in e ac ions should be educed,
bu also in he daily li e o humans, seeking an imp o emen in hei quali y o li e.
Re e ences
[1]
Abadi, M., Ba ham, P., Chen, J., Chen, Z., Da is, A., Dean, J., De in, M.,
Ghemawa , S., I ing, G., Isa d, M., e al. (2016). Tenso low: A sys em o
la ge-scale machine lea ning. In OSDI, olume 16, pages 265–283.
[2]
Achliop as, P., Diaman i, O., Mi liagkas, I., and Guibas, L. (2018). Lea ning ep e-
sen a ions and gene a i e models o 3d poin clouds. In In e na ional Con e ence
on Machine Lea ning, pages 40–49.
[3] Akh e , I., Sheikh, Y., Khan, S., and Kanade, T. (2008). Non igid s uc u e om
mo ion in ajec o y space. In NIPS.
[4]
Alahi, A., O iz, R., and Vande gheyns , P. (2012). F eak: Fas e ina keypoin . In
2012 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, pages 510–517.
Ieee.
[5]
A meni, I., He, Z.-Y., Gwak, J., Zami , A. R., Fische , M., Malik, J., and Sa a ese,
S. (2019). 3d scene g aph: A s uc u e o uni ied seman ics, 3d space, and came a.
In P oceedings o he IEEE In e na ional Con e ence on Compu e Vision, pages
5664–5673.
[6]
A meni, I., Sax, A., Zami , A. R., and Sa a ese, S. (2017). Join 2D-3D-Seman ic
Da a o Indoo Scene Unde s anding. a Xi :1702.01105.
[7]
Bad ina ayanan, V., Kendall, A., and Cipolla, R. (2017). Segne : A deep con olu-
ional encode -decode a chi ec u e o image segmen a ion. IEEE ansac ions on
pa e n analysis and machine in elligence, 39(12):2481–2495.
[8]
Bao, S. Y., Sun, M., and Sa a ese, S. (2011). Towa d cohe en objec de ec ion
and scene layou unde s anding. Image and Vision Compu ing, 29(9):569–579.
[9]
Bay, H., Ess, A., Tuy elaa s, T., and Van Gool, L. (2008). Speeded-up obus
ea u es (su ). Compu e Vision and Image Unde s anding, 110(3):346 – 359.
[10]
Bazin, J.-C. and Polle eys, M. (2012). 3-line ansac o o hogonal anishing poin
de ec ion. In 2012 IEEE/RSJ In e na ional Con e ence on In elligen Robo s and
Sys ems, pages 4282–4287. IEEE.
[11]
Bazin, J.-C., Seo, Y., Demonceaux, C., Vasseu , P., Ikeuchi, K., Kweon, I., and
Polle eys, M. (2012a). Globally op imal line clus e ing and anishing poin es ima ion
in manha an wo ld. In IEE CVPR, pages 638–645.

114 Re e ences
[12]
Bazin, J.-C., Seo, Y., and Polle eys, M. (2012b). Globally op imal consensus se
maximiza ion h ough o a ion sea ch. In Asian Con e ence on Compu e Vision,
pages 539–551. Sp inge .
[13]
Besl, P. J. and McKay, N. D. (1992). Me hod o egis a ion o 3-d shapes. In
Senso usion IV: con ol pa adigms and da a s uc u es, olume 1611, pages 586–606.
In e na ional Socie y o Op ics and Pho onics.
[14]
Bogo, F., Kanazawa, A., Lassne , C., Gehle , P., Rome o, J., and Black, M. J.
(2016). Keep i smpl: Au oma ic es ima ion o 3d human pose and shape om a
single image. In Eu opean Con e ence on Compu e Vision, pages 561–578. Sp inge .
[15]
Bogo, F., Rome o, J., Pons-Moll, G., and Black, M. J. (2017). Dynamic aus :
Regis e ing human bodies in mo ion. In CVPR, pages 6233–6242.
[16]
B egle , C., He zmann, A., and Bie mann, H. (2000). Reco e ing non- igid 3D
shape om image s eams. In CVPR.
[17]
B omley, J., Guyon, I., LeCun, Y., Säckinge , E., and Shah, R. (1994). Signa u e
e i ica ion using a" siamese" ime delay neu al ne wo k. In Ad ances in neu al
in o ma ion p ocessing sys ems, pages 737–744.
[18]
B own, M. and Lowe, D. G. (2007). Au oma ic pano amic image s i ching using
in a ian ea u es. In e na ional jou nal o compu e ision, 74(1):59–73.
[19]
Cao, Z., Simon, T., Wei, S.-E., and Sheikh, Y. (2017). Real ime mul i-pe son 2d
pose es ima ion using pa a ini y ields. In P oceedings o he IEEE Con e ence on
Compu e Vision and Pa e n Recogni ion, pages 7291–7299.
[20]
Chen, L.-C., Papand eou, G., Kokkinos, I., Mu phy, K., and Yuille, A. L. (2018).
Deeplab: Seman ic image segmen a ion wi h deep con olu ional ne s, a ous con-
olu ion, and ully connec ed c s. ansac ions on pa e n analysis and machine
in elligence, 40(4):834–848.
[21]
Chen, L.-C., Papand eou, G., Sch o , F., and Adam, H. (2017). Re hinking a ous
con olu ion o seman ic image segmen a ion. a Xi :1706.05587.
[22]
Cohen, T. S., Geige , M., Köhle , J., and Welling, M. (2018). Sphe ical cnns.
a Xi :1801.10130.
[23]
Concha, A., Hussain, M. W., Mon ano, L., and Ci e a, J. (2014). Manha an and
Piecewise-Plana Cons ain s o Dense Monocula Mapping. In Robo ics: Science
and sys ems.
[24]
Coughlan, J. M. and Yuille, A. L. (1999). Manha an wo ld: Compass di ec ion
om a single image by bayesian in e ence. In P oceedings o he Se en h IEEE
In e na ional Con e ence on Compu e Vision, olume 2, pages 941–947. IEEE.
[25]
Coughlan, J. M. and Yuille, A. L. (2003). Manha an wo ld: O ien a ion and
ou lie de ec ion by bayesian in e ence. Neu al compu a ion, 15(5):1063–1088.
Re e ences 115
[26]
C euso , C., Pea s, N., and Aus in, J. (2012). 3d landma k model disco e y om
a egis e ed se o o ganic shapes. In 2012 IEEE Compu e Socie y Con e ence on
Compu e Vision and Pa e n Recogni ion Wo kshops, pages 57–64. IEEE.
[27]
Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., and Wei, Y. (2017).
De o mable con olu ional ne wo ks. CoRR, abs/1703.06211, 1(2):3.
[28]
Dai, Y., Li, H., and He, M. (2012). A simple p io - ee me hod o non- igid
s uc u e- om-mo ion ac o iza ion. In CVPR.
[29]
Dalal, N. and T iggs, B. (2005). His og ams o o ien ed g adien s o human
de ec ion. In 2005 IEEE compu e socie y con e ence on compu e ision and pa e n
ecogni ion (CVPR’05), olume 1, pages 886–893. IEEE.
[30]
Dasgup a, S., Fang, K., Chen, K., and Sa a ese, S. (2016). Delay: Robus spa ial
layou es ima ion o clu e ed indoo scenes. In IEEE CVPR, pages 616–624.
[31]
Delage, E., Lee, H., and Ng, A. Y. (2006). A dynamic bayesian ne wo k model
o au onomous 3D econs uc ion om a single indoo image. In IEEE Compu e
Socie y Con e ence on Compu e Vision and Pa e n Recogni ion, olume 2, pages
2418–2428. IEEE.
[32]
Deng, F., Zhu, X., and Ren, J. (2017). Objec de ec ion on pano amic images based
on deep lea ning. In 2017 3 d In e na ional Con e ence on Con ol, Au oma ion and
Robo ics (ICCAR), pages 375–380. IEEE.
[33]
Dong, X., Yan, Y., Ouyang, W., and Yang, Y. (2018). S yle agg ega ed ne wo k
o acial landma k de ec ion. In P oceedings o he IEEE Con e ence on Compu e
Vision and Pa e n Recogni ion, pages 379–388.
[34]
Doso i skiy, A., Fische , P., Ilg, E., Hausse , P., Hazi bas, C., Golko , V., an de
Smag , P., C eme s, D., and B ox, T. (2015). Flowne : Lea ning op ical low wi h
con olu ional ne wo ks. In IEEE ICCV, pages 2758–2766.
[35]
D o nik, N., Shmelko , K., Mai al, J., and Schmid, C. (2017). Bli zne : A
eal- ime deep ne wo k o scene unde s anding. In ICCV, pages 4154–4162.
[36]
Eigen, D. and Fe gus, R. (2015). P edic ing dep h, su ace no mals and seman ic
labels wi h a common mul i-scale con olu ional a chi ec u e. In IEEE In e na ional
Con e ence on Compu e Vision, pages 2650–2658.
[37]
Fan, H., Su, H., and Guibas, L. J. (2017). A poin se gene a ion ne wo k o 3d
objec econs uc ion om a single image. In P oceedings o he IEEE con e ence on
compu e ision and pa e n ecogni ion, pages 605–613.
[38]
Felzenszwalb, P., McAlles e , D., and Ramanan, D. (2008). A disc imina i ely
ained, mul iscale, de o mable pa model. In 2008 IEEE Con e ence on Compu e
Vision and Pa e n Recogni ion, pages 1–8. IEEE.
[39]
Felzenszwalb, P. F., Gi shick, R. B., McAlles e , D., and Ramanan, D. (2009).
Objec de ec ion wi h disc imina i ely ained pa -based models. IEEE ansac ions
on pa e n analysis and machine in elligence, 32(9):1627–1645.
116 Re e ences
[40]
Fe nandez-Lab ado , C., Chha kuli, A., Paudel, D. P., Gue e o, J. J., Demon-
ceaux, C., and Van Gool, L. (2020a). Unsupe ised lea ning o ca ego y-speci ic
symme ic 3d keypoin s om poin se s.
[41]
Fe nandez-Lab ado , C., Facil, J. M., Pe ez-Yus, A., Demonceaux, C., Ci e a, J.,
and Gue e o, J. J. (Ap il 2020b). Co ne s o layou : End- o-end layou eco e y
om 360 images. IEEE Robo ics and Au oma ion Le e s, 5 (2), pp: 1255-1262.
[42]
Fe nandez-Lab ado , C., Facil, J. M., Pe ez-Yus, A., Demonceaux, C., and Gue -
e o, J. J. (2018a). Pano oom: F om he sphe e o he 3d layou . a Xi :1808.09879.
[43]
Fe nandez-Lab ado , C., Pe ez-Yus, A., Lopez-Nicolas, G., and Gue e o, J. J.
(2018b). Layou s om pano amic images wi h geome y and deep lea ning.
a Xi :1806.08294.
[44]
Flin , A., Mu ay, D., and Reid, I. (2011). Manha an scene unde s anding using
monocula , s e eo, and 3d ea u es. In Compu e Vision (ICCV), 2011 In e na ional
Con e ence on, pages 2228–2235. IEEE.
[45]
Fo sy h, D. A. (2014). Objec de ec ion wi h disc imina i ely ained pa -based
models. IEEE Compu e , 47:6–7.
[46]
Fouhey, D. F., Delai e, V., Gup a, A., E os, A. A., Lap e , I., and Si ic, J. (2014).
People wa ching: Human ac ions as a cue o single iew geome y. In e na ional
jou nal o compu e ision, 110(3):259–274.
[47]
Fu lan e al (2013). F ee you came a: 3d indoo scene unde s anding om
a bi a y came a mo ion. BMVC.
[48]
Gao, Y. and Yuille, A. L. (2016). Symme ic non- igid s uc u e om mo ion o
ca ego y-speci ic objec s uc u e es ima ion. In Eu opean Con e ence on Compu e
Vision, pages 408–424. Sp inge .
[49]
Ge ig, T., Mo el-Fo s e , A., Blume , C., Egge , B., Lu hi, M., Schönbo n, S.,
and Ve e , T. (2018). Mo phable ace models-an open amewo k. In 2018 13 h
IEEE In e na ional Con e ence on Au oma ic Face & Ges u e Recogni ion (FG 2018),
pages 75–82. IEEE.
[50]
Gi shick, R. (2015). Fas -cnn. In P oceedings o he IEEE in e na ional con e ence
on compu e ision, pages 1440–1448.
[51]
Gi shick, R., Donahue, J., Da ell, T., and Malik, J. (2014). Rich ea u e hie a chies
o accu a e objec de ec ion and seman ic segmen a ion. In P oceedings o he IEEE
con e ence on compu e ision and pa e n ecogni ion, pages 580–587.
[52]
González, Á. (2010). Measu emen o a eas on a sphe e using ibonacci and
la i ude–longi ude la ices. Ma hema ical Geosciences, 42(1):49.
[53]
Gue e o-Viu, J., Fe nandez-Lab ado , C., Demonceaux, C., and Gue e o, J. J.
(2019). Wha ’s in my oom? objec ecogni ion on indoo pano amic images. a Xi
p ep in a Xi :1910.06138.
Re e ences 117
[54]
Gup a, S., A beláez, P., Gi shick, R., and Malik, J. (2015). Aligning 3d models o
gb-d images o clu e ed scenes. In P oceedings o he IEEE con e ence on compu e
ision and pa e n ecogni ion, pages 4731–4740.
[55]
Gu ié ez-Gómez, D., Mayol-Cue as, W., and Gue e o, J. J. (2015). Wha should
i landma k? en opy o no mals in dep h ju s o place ecogni ion in changing
en i onmen s using gb-d da a. In 2015 IEEE In e na ional Con e ence on Robo ics
and Au oma ion (ICRA), pages 5468–5474. IEEE.
[56]
Hasle , N., Acke mann, H., Rosenhahn, B., Tho mählen, T., and Seidel, H.-P.
(2010). Mul ilinea pose and body shape es ima ion o d essed subjec s om image
se s. In 2010 IEEE Compu e Socie y Con e ence on Compu e Vision and Pa e n
Recogni ion, pages 1823–1830. IEEE.
[57]
He, K., Gkioxa i, G., Dollá , P., and Gi shick, R. (2017). Mask -cnn. In P oceedings
o he IEEE in e na ional con e ence on compu e ision, pages 2961–2969.
[58]
He, K., Zhang, X., Ren, S., and Sun, J. (2015). Del ing deep in o ec i ie s:
Su passing human-le el pe o mance on imagene classi ica ion. In P oceedings o
he IEEE in e na ional con e ence on compu e ision, pages 1026–1034.
[59]
He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep esidual lea ning o image
ecogni ion. In IEEE CVPR, pages 770–778.
[60]
Hedau, V., Hoiem, D., and Fo sy h, D. (2009). Reco e ing he spa ial layou
o clu e ed ooms. In IEEE In e na ional Con e ence on Compu e Vision, pages
1849–1856.
[61]
Hedau, V., Hoiem, D., and Fo sy h, D. (2010). Thinking inside he box: Using
appea ance models and con ex based on oom geome y. Eu opean Con e ence on
Compu e Vision, pages 224–237.
[62]
Hej a i, M. and Ramanan, D. (2012). Analyzing 3d objec s in clu e ed images.
In Ad ances in Neu al In o ma ion P ocessing Sys ems, pages 593–601.
[63]
Hoiem, D., E os, A. A., and Hebe , M. (2005). Geome ic con ex om a single
image. In Compu e Vision, 2005. ICCV 2005. Ten h IEEE In e na ional Con e ence
on, olume 1, pages 654–661. IEEE.
[64]
Huang, S., Gong, M., and Tao, D. (2017). A coa se- ine ne wo k o keypoin
localiza ion. In P oceedings o he IEEE In e na ional Con e ence on Compu e
Vision, pages 3028–3037.
[65]
Hussain, W., Ci e a, J., Mon ano, L., and Hebe , M. (2016). Dealing wi h small
da a and aining blind spo s in he Manha an wo ld. In Win e Con e ence on
Applica ions o Compu e Vision (WACV), pages 1–9. IEEE.
[66]
Izadinia, H., Shan, Q., and Sei z, S. M. (2017). Im2cad. In P oceedings o he
IEEE Con e ence on Compu e Vision and Pa e n Recogni ion, pages 5134–5143.
124 Re e ences
[149]
Ve ma, N., Boye , E., and Ve beek, J. (2018). Feas ne : Fea u e-s ee ed g aph
con olu ions o 3d shape analysis. In CVPR.
[150]
Viola, P. and Jones, M. (2001). Rapid objec de ec ion using a boos ed cascade
o simple ea u es. In P oceedings o he 2001 IEEE compu e socie y con e ence on
compu e ision and pa e n ecogni ion. CVPR 2001, olume 1, pages I–I. IEEE.
[151]
Viola, P. and Jones, M. J. (2004). Robus eal- ime ace de ec ion. In e na ional
jou nal o compu e ision, 57(2):137–154.
[152]
on Gioi, R. G., Jakubowicz, J., Mo el, J.-M., and Randall, G. (2012). Lsd:
a line segmen de ec o , image p ocessing on line,(2012). URL: h p://dx. doi.
o g/10.5201/ipol.
[153]
Wang, C., Wang, Y., Lin, Z., Yuille, A. L., and Gao, W. (2014). Robus es ima ion
o 3d human poses om a single image. In P oceedings o he IEEE Con e ence on
Compu e Vision and Pa e n Recogni ion, pages 2361–2368.
[154]
Wang, H., S idha , S., Huang, J., Valen in, J., Song, S., and Guibas, L. J.
(2019). No malized objec coo dina e space o ca ego y-le el 6d objec pose and
size es ima ion. In P oceedings o he IEEE Con e ence on Compu e Vision and
Pa e n Recogni ion, pages 2642–2651.
[155]
Wu, J., Xue, T., Lim, J. J., Tian, Y., Tenenbaum, J. B., To alba, A., and
F eeman, W. T. (2016). Single image 3d in e p e e ne wo k. In Eu opean Con e ence
on Compu e Vision, pages 365–382. Sp inge .
[156]
Wu, S., Rupp ech , C., and Vedaldi, A. (2020). Unsupe ised lea ning o p obably
symme ic de o mable 3d objec s om images in he wild. In CVPR.
[157]
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. (2015). 3d
shapene s: A deep ep esen a ion o olume ic shapes. In CVPR, pages 1912–1920.
[158]
Xiao, J., Ehinge , K., Oli a, A., and To alba, A. (2012). Recognizing scene
iewpoin using pano amic place ep esen a ion. In IEEE CVPR, pages 2695–2702.
[159]
Xu, J., S enge , B., Ke ola, T., and Tung, T. (2017). Pano2CAD: Room layou
om a single pano ama image. In IEEE WACV, pages 354–362.
[160]
Yang, H. and Ca lone, L. (2019). In pe ec shape: Ce i iably op imal 3d shape
econs uc ion om 2d landma ks. a Xi p ep in a Xi :1911.11924.
[161]
Yang, H. and Zhang, H. (2016a). E icien 3D oom shape eco e y om a single
pano ama. In IEEE CVPR, pages 5422–5430.
[162]
Yang, H. and Zhang, H. (2016b). E icien 3d oom shape eco e y om a single
pano ama. In IEEE CVPR, pages 5422–5430.
[163]
Yang, S.-T., Wang, F.-E., Peng, C.-H., Wonka, P., Sun, M., and Chu, H.-K.
(2018a). Dula-ne : A dual-p ojec ion ne wo k o es ima ing oom layou s om a
single gb pano ama. a Xi :1811.11977.

Re e ences 125
[164]
Yang, W., Qian, Y., Kämä äinen, J.-K., C ic i, F., and Fan, L. (2018b). Objec
de ec ion in equi ec angula pano ama. In 2018 24 h In e na ional Con e ence on
Pa e n Recogni ion (ICPR), pages 2190–2195. IEEE.
[165]
Yang, Y., Jin, S., Liu, R., Bing Kang, S., and Yu, J. (2018c). Au oma ic 3d
indoo scene modeling om single pano ama. In The IEEE Con e ence on Compu e
Vision and Pa e n Recogni ion (CVPR).
[166]
Yew, Z. J. and Lee, G. H. (2018). 3d ea -ne : Weakly supe ised local 3d ea u es
o poin cloud egis a ion. In Eu opean Con e ence on Compu e Vision, pages
630–646. Sp inge .
[167]
Yi, L., Kim, V. G., Ceylan, D., Shen, I.-C., Yan, M., Su, H., Lu, C., Huang, Q.,
She e , A., and Guibas, L. (2016). A scalable ac i e amewo k o egion anno a ion
in 3d shape collec ions. ACM T ansac ions on G aphics (TOG), 35(6):1–12.
[168]
Yu, X., Zhou, F., and Chand ake , M. (2016). Deep de o ma ion ne wo k o
objec landma k localiza ion. In Eu opean Con e ence on Compu e Vision, pages
52–70. Sp inge .
[169]
Za ei iou, S., Ch ysos, G. G., Roussos, A., Ve e as, E., Deng, J., and T igeo gis,
G. (2017). The 3d menpo acial landma k acking challenge. In P oceedings o he
IEEE In e na ional Con e ence on Compu e Vision Wo kshops, pages 2503–2511.
[170]
Zhang, J., Kan, C., Schwing, A. G., and U asun, R. (2013). Es ima ing he 3d
layou o indoo scenes and i s clu e om dep h senso s. In 2013 In e na ional
Con e ence on Compu e Vision, pages 1273–1280. IEEE.
[171]
Zhang, W., Zhang, W., Liu, K., and Gu, J. (2017). Lea ning o p edic high-
quali y edge maps o oom layou es ima ion. T ansac ions on Mul imedia, 19(5):935–
943.
[172]
Zhang, Y., Song, S., Tan, P., and Xiao, J. (2014a). PanoCon ex : A whole-
oom 3D con ex model o pano amic scene unde s anding. In IEEE ECCV, pages
668–686.
[173]
Zhang, Z., Luo, P., Loy, C. C., and Tang, X. (2014b). Facial landma k de ec ion
by deep mul i- ask lea ning. In Eu opean con e ence on compu e ision, pages
94–108. Sp inge .
[174]
Zhang, Z., Rebecq, H., Fo s e , C., and Sca amuzza, D. (2016). Bene i o la ge
ield-o - iew came as o isual odome y. In 2016 IEEE In e na ional Con e ence
on Robo ics and Au oma ion (ICRA), pages 801–808. IEEE.
[175]
Zhao, H., Lu, M., Yao, A., Guo, Y., Chen, Y., and Zhang, L. (2017). Physics
inspi ed op imiza ion on seman ic ans e ea u es: An al e na i e me hod o oom
layou es ima ion. a Xi :1707.00383.
[176]
Zhou, B., Laped iza, A., Khosla, A., Oli a, A., and To alba, A. (2017). Places:
A 10 million image da abase o scene ecogni ion. IEEE ansac ions on pa e n
analysis and machine in elligence, 40(6):1452–1464.
126 Re e ences
[177]
Zou, C., Colbu n, A., Shan, Q., and Hoiem, D. (2018). Layou ne : Recons uc ing
he 3d oom layou om a single gb image. In P oceedings IEEE Con e ence on
Compu e Vision and Pa e n Recogni ion, pages 2051–2059.