UNIVERSITÀ DEGLI STUDI DI PARMA
DOTTORATO DI RICERCA IN
“TECNOLOGIE DELL’INFORMAZIONE”
CICLO XXXIII
Deep Lea ning-based Objec
De ec ion o Au onomous D i ing
Coo dina o e:
Chia .mo P o . Ma co Loca elli
Tu o e:
Chia .mo P o . Massimo Be ozzi
Do o ando: And ea Zinelli
Anni 2018/2021
Con en s
In oduc ion 1
1 P io A 7
1.1 Wha a e Deep Neu al Ne wo ks? . . . . . . . . . . . . . . . . . 7
1.1.1 Linea Reg ession . . . . . . . . . . . . . . . . . . . . . . 7
1.1.2 Linea Basis Func ion Reg ession . . . . . . . . . . . . . 9
1.1.3 Pa ame ic Basis Func ions: The mos basic Neu al Ne -
wo k............................. 10
1.1.4 F om Shallow o Deep Models . . . . . . . . . . . . . . . 12
1.1.5 T aining Neu al Ne wo ks . . . . . . . . . . . . . . . . . 13
1.2 Con olu ional Neu al Ne wo ks . . . . . . . . . . . . . . . . . . 15
1.2.1 The con olu ional ope a o . . . . . . . . . . . . . . . . 16
1.2.2 Con olu ional Neu al Ne wo ks . . . . . . . . . . . . . . 17
1.3 The Fi s La ge Scale Models o Image Classi ica ion . . . . . 20
1.4 Seman ic Segmen a ion . . . . . . . . . . . . . . . . . . . . . . . 24
1.5 Objec de ec ion .......................... 26
1.5.1 2D Objec De ec ion on Images . . . . . . . . . . . . . . 27
1.5.2 T adeo be ween One-s age and Two-s age echniques . 32
1.6 3D Objec De ec ion . . . . . . . . . . . . . . . . . . . . . . . . 33
1.6.1 LiDAR-based 3D Objec De ec ion . . . . . . . . . . . . 34
1.6.2 Image-based 3D Objec De ec ion . . . . . . . . . . . . 41
ii Con en s
2 Objec De ec ion o Pa king Slo De ec ion 47
2.1 P io A and Mo i a ion . . . . . . . . . . . . . . . . . . . . . 47
2.2 Fas e R-CNN............................ 49
2.2.1 T aining P ocedu e . . . . . . . . . . . . . . . . . . . . . 53
2.3 Pa king Slo De ec o . . . . . . . . . . . . . . . . . . . . . . . 54
2.3.1 T aining P ocedu e . . . . . . . . . . . . . . . . . . . . . 57
2.3.2 Da ase Cons uc ion and Da a P epa a ion . . . . . . . 60
2.4 Expe imen al Resul s . . . . . . . . . . . . . . . . . . . . . . . . 61
2.4.1 Seman ic Shi P oblem . . . . . . . . . . . . . . . . . . 63
2.4.2 Abla ion S udy . . . . . . . . . . . . . . . . . . . . . . . 65
2.4.3 Model Simpli ica ion and Spa si ica ion . . . . . . . . . 68
2.4.4 P edic ion Di ec ly on Sphe ical Images . . . . . . . . . 73
2.5 Discussion.............................. 77
3 Monocula 3D Objec De ec ion ia Gene alized In e sec ion-
o e -Union Minimiza ion 79
3.1 P io A and Mo i a ion . . . . . . . . . . . . . . . . . . . . . 79
3.2 BaselineModel ........................... 82
3.3 3D De ec ion Module . . . . . . . . . . . . . . . . . . . . . . . . 83
3.4 Model Op imiza ion . . . . . . . . . . . . . . . . . . . . . . . . 87
3.4.1 Gene alized In e sec ion-o e -Union . . . . . . . . . . . 87
3.4.2 T aining P ocedu e . . . . . . . . . . . . . . . . . . . . . 89
3.5 Expe imen al Resul s . . . . . . . . . . . . . . . . . . . . . . . . 92
3.5.1 The KITTI Da ase . . . . . . . . . . . . . . . . . . . . 92
3.5.2 Compa ison wi h he S a e o he A . . . . . . . . . . 93
3.5.3 Compa ison wi h o he Loss Fo mula ions . . . . . . . . 98
3.5.4 Quali a i e esul s . . . . . . . . . . . . . . . . . . . . . 102
3.6 A case s udy:
3D GIoU applied o F us um-Poin Ne s . . . . . . . . . . . . . 105
3.6.1 3D de ec o : s anda d op imiza ion . . . . . . . . . . . . 106
3.6.2 3D de ec o : p oposed op imiza ion . . . . . . . . . . . . 107
Con en s iii
3.6.3 Expe imen al Resul s . . . . . . . . . . . . . . . . . . . 108
3.7 Discussion.............................. 110
4 3D Objec De ec ion on LiDAR scans ia Vo ing and Sel -
A en ion Mechanisms 113
4.1 P io A and Mo i a ion . . . . . . . . . . . . . . . . . . . . . 113
4.2 Vo ene o 3D Objec De ec ion on D i ing Scena ios . . . . . 115
4.2.1 Poin Ne s o poin cloud p ocessing . . . . . . . . . . . 115
4.2.2 Baseline Vo ene Model o 3D Objec De ec ion . . . . 119
4.2.3 Modi ica ions o he baseline o Au onomous D i ing
Scena ios .......................... 122
4.3 Enhancing SA laye s ia Sel -A en ion . . . . . . . . . . . . . . 127
4.3.1 The a en ion mechanism . . . . . . . . . . . . . . . . . 128
4.3.2 Sel -a en ion applied o SA laye s . . . . . . . . . . . . 131
4.4 T ainingSe up ........................... 135
4.4.1 KITTI LiDAR Da ase . . . . . . . . . . . . . . . . . . . 135
4.4.2 Model a chi ec u e . . . . . . . . . . . . . . . . . . . . . 135
4.4.3 T aining Rou ine . . . . . . . . . . . . . . . . . . . . . . 138
4.5 Expe imen al Resul s . . . . . . . . . . . . . . . . . . . . . . . . 139
4.5.1 Abla ion s udy on he design choice . . . . . . . . . . . 140
4.5.2 Abla ion s udy on a en ion ype . . . . . . . . . . . . . 142
4.5.3 Compa ison wi h S a e-o - he-A Sys ems . . . . . . . . 144
4.5.4 Quali a i e esul s . . . . . . . . . . . . . . . . . . . . . 147
4.6 Discussion.............................. 150
Conclusions 153
Bibliog aphy 157
In oduc ion
Recen ad ancemen s in pa allel compu ing ha dwa e and amewo ks, cou-
pled wi h he eno mous inc ease in publicly a ailable anno a ed da ase s, has
caused a esu gence in in e es om he esea ch communi y on he opic o
machine lea ning and, in pa icula , neu al ne wo ks. In he las decade alone,
an ex emely wide a ie y o di e en neu al a chi ec u es and models ha e
been p oposed o ackle he mos dispa a e p oblems, including bu no lim-
i ed o sen ence ansla ion, image gene a ion, image classi ica ion, de ec ion,
localiza ion and planning. In pa icula , he elease o inc easingly powe ul
G aphical P ocessing Uni s (GPUs) allowed o he de elopmen o mo e com-
plex neu al ne wo k designs, gi ing ise o he esea ch ield ha is cu en ly
known as deep lea ning. Amid hese designs, pe haps he mos impo an and
no o ious one is ep esen ed by Con olu ional Neu al Ne wo ks (CNNs), a
ca ego y o neu al ne wo k ha u ilizes mul iple s acked laye s o lea nable
con olu ional ke nels o build powe ul hie a chical ep esen a ions o inpu
images and ca y ou complex isual asks [1, 2, 3, 4, 5, 6, 7, 8]. All hese
solu ions ha e shown ema kable pe o mance, o en su passing " adi ional"
algo i hms by wide ma gins while being mo e e icien due o hei high pa -
allelism. As a esul , many "classical" algo i hms a e g adually being eplaced
by deep lea ning-based solu ions, b inging abou he so-called age o So wa e
2.0: he so wa e enginee no longe hand-designs he logic o he algo i hm
di ec ly, bu a he designs he neu al ne wo k model ha lea ns he mos
sui able logic om he da a.
2 In oduc ion
Pe haps one o he indus ies mos a ec ed om his shi in pa adigm is
au omo i e: powe ed by his eme gen echnology, ex ensi ely esea ched and
applied o he ield o compu e ision, Ad anced D i e Assis ance Sys ems
(ADAS) and Au onomous D i ing solu ions ha e bene i ed om emendous
leaps in pe o mance. Among all he asks ha a e equi ed o success ul au-
onomous na iga ion, one o he mos c i ical is pe cep ion, which consis s o
in e p e ing he signals cap u ed by he senso sui e displaced on he ehicle
such ha hey can success ully be used o he subsequen asks o acking,
planning and con ol. Mos o he esea ch on deep lea ning-based pe cep ion
in he ield o au onomous d i ing concen a es on wo speci ic kinds o sen-
so s, RGB came as and LiDARs, which a e cha ac e ized by complemen a y
s eng hs and weaknesses. Came as p o ide a much dense and iche ype o
signal, while also being conside ably cheape han he LiDAR solu ions cu -
en ly a ailable on he ma ke ; on he o he hand, hey do no p o ide any
in o ma ion abou he dis ance o he pe cei ed objec s om he senso , and
hus equi e s e eoscopic se ups o be able o in e he scene dep h. Con e sely,
LiDARs p o ide an accu a e econs uc ion o he su ounding en i onmen in
he o m o a poin cloud, bu such signal is conside ably spa se and does no
con ain addi ional seman ic in o ma ion such as colo .
Gi en his da a, many deep lea ning models and echniques ha e been
de eloped ha ca y ou di e en ypes o pe cep ion asks. A pa icula ly
impo an one owa ds ully au oma ed na iga ion is objec de ec ion, which
can be de ined as ollows: gi en an inpu , which is he signal e u ned by one
o mo e senso s, he objec i e is o iden i y all en i ies o in e es inside such
inpu and de e mine hei s a e. Such s a e can ep esen any p ope ies o
he de ec ed objec : image-based 2D de ec o s [9, 10, 11], o ins ance, usually
es ima e an image-aligned bounding box ha con ains said objec , as well as
he class i belongs o. 3D de ec o s [12, 13, 14, 15, 16, 17] es ima e 3D bound-
ing boxes ha minimally enclose each objec , which a e ep esen ed by hei
posi ion in he wo ld, hei dimensions and hei o ien a ion. By in eg a ing
addi ional in o ma ion, like ada da a o pas measu emen s [18], o he quan-
In oduc ion 3
i ies can be es ima ed, such as objec eloci y o ajec o y. By summa izing
he inpu signal ia a ini e se o elemen s, each one ep esen ing he s a e
o a pa icula de ec ion (e.g. an obs acle on he oad), objec de ec o s p o-
ide an ou pu ha is e y simple and compac , and he e o e immedia ely
use ul o he subsequen planning phase, as i equi es li le o no addi ional
pos -p ocessing. Mo eo e , deep lea ning-based objec de ec o s a e ex emely
lexible, allowing o de ec almos any ca ego y o in e es , as long as he e
exis sui able aining da a o model op imiza ion.
Gi en i s cen al ole in con empo a y au onomous d i ing pipelines, his
hesis ocuses on he opic o deep objec de ec ion. Mo e speci ically, I p opose
h ee di e en sys ems, each one wo king on di e en inpu da a and ackling
a dis inc de ec ion p oblem.
The i s sys em consis s o an deep lea ning model o pa king space de-
ec ion and acancy classi ica ion on su ound- iew images [19]. This app oach
builds upon he exis ing wo-s age objec de ec o Fas e R-CNN [9], modi y-
ing i s s uc u e and logic o adap i o he peculia na u e o he ask o
pa king slo de ec ion. Pa king slo s can be o di e en ypes (e.g. ec angula
o slan ed) and can be obse ed om di e en angles, ende ing he o iginal
implemen a ion, which p edic s axis-aligned bounding boxes, ine ec i e o
accu a e slo p edic ion. As such, I p opose a a ia ion o he o iginal model,
emo ing he ancho -based egion p oposal and allowing gene ic quad ila e als
o be es ima ed ins ead o axis-aligned boxes. To ain and e alua e he model,
I collec ed and anno a ed a small da ase depic ing di e en oad scenes and
pa king lo s, s i ching oge he he bi d’s eye iew p ojec ions o ou isheye
came as o cons uc su ound- iew images. Se e al expe imen s show he e -
ec i eness o he p oposed o mula ion in unobse ed si ua ions, as well as
unde noise and di e en obse a ion condi ions.
The second wo k consis s in an ex en ion o Fas e R-CNN o he ask o
monocula 3D ca de ec ion. This is accomplished by in oducing a simple ad-
di ional neu al ne wo k module in o he o iginal a chi ec u e which is asked
o pe o m image-based 3D de ec ion, essen ially by lea ning he mapping be-
10 Chap e 1. P io A
linea basis unc ion eg ession. In his case, each inpu is ans o med in o a
ea u e ec o by a se o p ede e mined unc ions, known as basis unc ions
Φ(x)=[φ1(x), ..., φM(x)], be o e being ed o he linea model:
(x)W,b=Φ(x)·W+b.(1.10)
The use o his in e media e ep esen a ion allows o model mo e complex ela-
ionships be ween inpu and ou pu . Fo ins ance, i we know be o ehand ha
xand ya e bo h scala s and he p ocess ggene a ing y om xis a polynomial
o deg ee n, we could choose Φ(x)=[x,x2, ..., xn]as basis unc ions, allowing
he model o ep esen polynomial unc ions up o deg ee n. Op imizing ia he
minimiza ion o Eq.1.3 would yield he polynomial ha bes i s he da ase
D.
Despi e no longe being linea wi h espec o he inpu , his new model is
s ill linea wi h espec o he weigh s, which means ha a closed o m solu ion
can be ob ained using he no mal equa ions i he quad a ic loss (Eq. 1.4) is
adop ed: he only di e ence is ha now he design ma ix Xdoes no con ain
he inpu alues, bu a he hei ea u e ep esen a ions ob ained h ough he
basis unc ions. No e ha simple linea eg ession can be ob ained by se ing
Φ(x) = x.
1.1.3 Pa ame ic Basis Func ions: The mos basic Neu al Ne -
wo k
Fo complex da a i is o en un easible o de e mine he na u e o he unde -
lying unc ion g, and he e o e a sui able se o in e media e basis unc ions.
These cases can be app oached by u ilizing pa ame ic basis unc ions [28]:
Φ(x) = [φW1,b1
1(x), ..., φWM,bM
M(x)].(1.11)
He e, each φWi,bi
i(x)is commonly a scala unc ion composed o a pa ame ic
linea mapping ollowed by a non-linea unc ion:
φWi,bi
i(x) = σ(x·Wi+bi).(1.12)
1.1. Wha a e Deep Neu al Ne wo ks? 11
The non-linea unc ion σ(·), commonly e e ed o as ac i a ion unc ion, is
necessa y o allow he se o basis unc ions o model non-linea ela ionships
be ween inpu and ou pu .
In his o mula ion, he se o pa ame e s {(Wi,bi)}M
iis also op imized,
allowing o he mos sui able in e media e ep esen a ion o be de e mined
di ec ly om he aining da a. Con a ily o he p e ious wo app oaches,
howe e , his model is no longe linea wi h espec o he pa ame e s, due o
he p esence o he ac i a ion unc ion σ(·): as a esul , he co esponding mini-
miza ion p oblem (Eqs. 1.3, 1.4) is no longe con ex. This has wo implica ions:
on he one hand, he e is no closed o m solu ion o he se o pa ame e s;
on he o he hand, he e is no gua an ee ha g adien descen echniques will
each he global minimum, as he loss unc ion migh now con ain local minima
and saddle poin s.
This o mula ion ep esen s he mos basic o m o neu al ne wo k:
• he se basis unc ions Φpa ame ized by {(Wi,bi)}M
i, is commonly
e e ed o as hidden laye . This name s ems om he ac ha he in-
e media e ep esen a ion p oduced by Φis gene ally abs ac and no
immedia ely in e p e able by an ex e nal obse e ;
•each Wi∈RI ep esen s a weigh e m, and each bi∈R ep esen s a
bias e m;
•each elemen φWi,bi
i(x)o he ea u e ec o is commonly e e ed o as
neu on, and he numbe o neu ons in he hidden laye is called wid h;
• he hidden laye Φas de ined in Equa ion 1.11 ep esen s a ully-connec ed
laye , since each ou pu neu on is a unc ion o he en i e inpu ;
• he ou wa d linea mapping W,b, pa ame ized by {W,b}, ep esen s
he ou pu laye o he ne wo k.
The Uni e sal App oxima ion Theo em [29, 30] s a es ha his neu al ne wo k
is heo e ically capable o app oxima ing any con inuous unc ion, unde he
12 Chap e 1. P io A
assump ion ha he ac i a ion unc ion σ(·)is non-cons an and con inuous
and ha he wid h o he ne wo k (i.e. he numbe o basis unc ions) is su i-
cien ly high. Gi en ha any algo i hm can be ep esen ed by a unc ion (e.g.
classi ying an image can be seen as a unc ion ha maps an inpu ma ix o
pixels o a p obabili y dis ibu ion), his model could be used o sol e any
ask. Fo many p ac ical applica ions, howe e , ob aining a su icien ly accu-
a e app oxima ion o he desi ed unc ion would equi e an ex emely high
wid h, which leads o p ohibi i e compu a ional and memo y cos s. Mo eo e ,
his kind o model exhibi s a endency o o e i he da a, expecially when he
da ase is limi ed. O e i ing is a phenomenon whe e he model, ins ead o
lea ning an app oxima ion o he unknown unde lying unc ion gene a ing he
da a, i memo izes he aining se ins ead. As a esul , he model pe o ms
well on samples belonging o he aining se , bu exhibi s poo esul s on new,
unobse ed da a poin s.
1.1.4 F om Shallow o Deep Models
The basic idea in Deep Lea ning and Deep Neu al Ne wo ks consis s o ex end-
ing basic neu al ne wo ks by applying mul iple hidden laye s in a cascaded way,
leading o he ollowing o mula ion:
(x)W,b=WT·ΦL◦ΦL−1... ◦Φ1(x) + b,(1.13)
whe e Φl(x)=[φWl,1,bl,1
1(x), ..., φWl,Ml,bl,Ml
M(x)] ep esen s he l- h hidden
laye (i.e.: he l- h se o basis unc ions) o he model. The numbe o hid-
den laye s ha comp ise his model is commonly e e ed o as he dep h o
he neu al ne wo k. A as amoun o esea ch on a wide a ie y o di e -
en asks highligh ed how deepe models a e capable o a aining compa able
pe o mance o shallow ne wo ks (i.e. ne wo ks wi h one o only ew hidden
laye s) wi h conside ably less pa ame e s, while also being less p one o o e -
i ing. The common consensus on why his is he case is a ibu ed o how
deep models ex ac in o ma ion om he inpu compa ed o shallow ones.
Despi e he in e media e ep esen a ion p oduced by a ne wo k wi h a single
1.1. Wha a e Deep Neu al Ne wo ks? 13
hidden laye is enough o ensu e he model is a uni e sal app oxima o , i is
belie ed o be highly ine icien [31]. Con e sely, s acking a sequence o hidden
laye s one a e he o he allows a deepe ne wo k o lea n a ep esen a ion o
he inpu a mul iple scales and le els o abs ac ions, which in u n allows o
mo e complex unc ion amilies o be ep esen ed mo e e icien ly. Mo eo e ,
his hie a chical s uc u e o in e media e ep esen a ions ac i ely con ibu es
o p e en ing o e i ing, as i is mo e likely o cap u e meaning ul pa e ns in
he inpu a he han jus memo izing he da a poin s. Inc easing he dep h,
howe e , is ine i ably me wi h an highe di icul y in op imizing he model.
Inc easing he numbe o hidden laye s leads o mo e cascaded non-linea ac-
i a ions, which in u n ende s he cos unc ion highly non-con ex.
Ea ly on, common choices o he ac i a ion unc ion used o be he logis ic
unc ion, o sigmoid σ(x) = (1 + e−x)−1, and he hype bolic angen unc ion
anh(x) = (1 −e−2x)/(1 + e−2x). When inc easing he dep h o he model,
howe e , hese unc ions o en cause op imiza ion issues: his is due o he
p ope ies o he g adien s o hese unc ions, which is always ≤1(in case
o sigmoid ≤1/4) and quickly sa u a es o ze o when mo ing away om he
o igin. As a esul , chaining mul iple hidden laye s ha use his kind o non-
linea i y leads o g adien s ha ge p og essi ely smalle owa ds he ea ly
laye s. This phenomenon is known as he anishing g adien p oblem, and i
hinde s op imiza ion as he pa ame e s o hese laye s canno upda e p ope ly.
As a esul , mos mode n deep lea ning pipelines ha e g a i a ed owa ds non-
sa u a ing unc ions. The mos widely adop ed choice is he Rec i ied Linea
Uni ReLU (x) = max (0, x)[32], whose g adien is cons an o 1 in he ac i e
pa o he unc ion, hus a oiding he anishing o he g adien . In o de o
ha e use ul aining signal o e e y inpu alue, a ia ions o he ReLU ha e
been p oposed, such as he Leaky ReLU, he PReLU [33] and he SELU [34].
1.1.5 T aining Neu al Ne wo ks
Despi e he non-con exi y o he loss unc ion, g adien descen echniques can
s ill be u ilized o op imize Neu al Ne wo ks: while he e is no gua an ee ha
14 Chap e 1. P io A
he op imiza ion will con e ge owa ds he global minimum o he unc ion,
he hope is ha he p ocess will esul o a local minimum ha is app op ia e
enough o he ask o pe o m. In o de o compu e he g adien s o use o
g adien descen , backp opaga ion [35] is o en employed: his me hod elies
on he abili y o de e mine he g adien o simple unc ions, as well as on he
chain ule o di e en iabili y, o p og essi ely de e mine he nume ical alue
o he g adien o each weigh wi h espec o he loss unc ion. Fo ins ance,
gi en he loss unc ion alue L, ∂L/∂ can be compu ed di ec ly, which can
hen be used o eco e he de i a i es o he weigh s using he chain ule:
∂L
∂W=∂L
∂ ·∂
∂W,∂L
∂b=∂L
∂ ·∂
∂b.(1.14)
He e, ∂ /∂Wand ∂ /∂ba e i ial o compu e gi en he linea ela ionship
be ween he weigh s and biases and he unc ion . Gi en hese quan i ies,
he chain ule can hen be en o ced i e a i ely o compu e he g adien s wi h
espec o he hidden laye pa ame e s:
∂L
∂WL
=∂L
∂ ·∂
∂ΦL·∂ΦL
∂WL
,(1.15)
∂L
∂Wl
=∂L
∂ ·∂
∂ΦL·∂ΦL
∂ΦL−1·. . . ∂Φl
∂Wl
l= 1, . . . , L −1(1.16)
.
Again, all he pa ial de i a i es in he abo e equa ions can be compu ed
analy ically, as he co esponding unc ions a e ei he linea o depend on he
non-linea i y σ(·)whose de i a i e should be known.
Mode n Neu al Ne wo ks a e o en compu a ionally expensi e and hey a e
usually op imized o e e y la ge da ase s. As a esul , applying g adien de-
scen di ec ly o hei op imiza ion is o en in ac able. This is due o he ac
ha g adien descen equi es he model o be e alua ed on all aining sam-
ples (i.e. Eq. 1.4) e e y ime a pa ame e upda e (Eq. 1.9) is o be pe o med,
which leads o a p ohibi i e amoun o ime necessa y o each he local mini-
mum. As a esul , mos cu en models a e ained using S ochas ic G adien
1.2. Con olu ional Neu al Ne wo ks 15
Descen (SGD) ins ead: in his a ian , each op imiza ion s ep is pe o med
using only BN andomly sampled da a poin s a a ime, whe e each g oup
o Belemen s is o en e e ed o as ba ch and B ep esen s he ba ch size.
Consequen ly, pe o ming each pa ame e upda e is signi ican ly as e , as i
equi es only a ac ion o he da a each ime. Once he en i e aining se is
used o pe o m upda es, an epoch is comple ed, he aining se is e-shu led
and new ba ches a e c ea ed o u he aining. The heo e ical d awback o
his me hod is ha i does no gua an ee ha e en a local minimum will be
eached du ing op imiza ion. In p a ice, howe e , SGD ends o each a sa is-
ac o y alue o he cos unc ion a he quickly and, by in oducing noise in o
he aining p ocess by andomly sampling ba ches o da a each ime, o en
ac s as a egula ize , educing o e i ing and leading o models ha pe o m
be e on new da a samples.
Many a ia ions o SGD ha e been p oposed, wi h he objec i e o im-
p o ing and speeding up aining. Adding a momen um e m [36] allows o a
as e con e gence owa ds a minimum and helps a oid oscilla ions in badly-
condi ioned egions o he loss unc ion. Adag ad [37], ins ead o upda ing each
pa ame e using he same lea ning a e, adap s he lea ning a e indi idually
o each pa ame e . ADAM [38], one o he mos used op imize s cu en ly o-
ge he wi h SGD wi h momen um, imp o es upon he idea o Adag ad while
also keeping ack o pas g adien upda es like in momen um SGD.
1.2 Con olu ional Neu al Ne wo ks
Con olu ional Neu al Ne wo ks (CNNs) ep esen pe haps he mos well-known
and widely-used ca ego y o deep lea ning models. These neu al ne wo ks a e
designed o ha ness he p ope ies o ce ain classes o signals o being o ga-
nized in hie a chies and p esen ing local pa e ns. Examples o hese kinds o
signals include:
•na u al images, in which neighbo ing pixels a e co ela ed ( he alue o
each pixel usually depends on i s icini y) and a e o ganized o o m
16 Chap e 1. P io A
spacial pa e ns and s uc u es;
• ideos, whe e he co ela ion akes place no only in space bu also
h ough ime;
•na u al ex , whe e neighbo ing cha ac e s o m wo ds and neighbo ing
wo ds o m sen ences.
1.2.1 The con olu ional ope a o
Con olu ion is a undamen al ma hema ical ope a o which is widely used in
image p ocessing. I can be in e p e ed as a way o pe o m mul iplica ion
be ween wo signals ha ing di e en numbe s o elemen s. Fo ins ance, gi en
an inpu image x∈RH×W×C, whe e H×W ep esen i s spa ial dimensions
and Ci s numbe o channels, he con olu ion o his image wi h a ke nel
W∈RK1×K2×Ccan be de ined as ollows:
y(h, w)=(x∗W)(h, w) = X
(k1,k2)∈S
x(h+k1, w +k2)·W(k1, k2),(1.17)
whe e
S=−K1
2,...,K1
2×−K2
2,...,K2
2 (1.18)
The ou pu y∈RH0×W0p ese es he numbe o spa ial dimensions o he
inpu , and each elemen (h, w)is he esul o a scala p oduc be ween he
ke nel and he neighbo hood o (h, w)in he o iginal image. In o he wo ds,
con olu ion is a unc ion ha ope a es locally and ha e u ns a map o local
esponses o he ke nel o he inpu . Due o he ac ha con olu ion canno
be compu ed o hose posi ions in which he ke nel is no comple ely inside
he inpu , he spa ial dimensions H0×W0o ydo no necessa ily coincide
wi h hose o he inpu . A commonly adop ed echnique, especially in deep
con olu ional models, o p ese e he spa ial dimensions consis s o padding
he inpu , commonly wi h ze os, such ha e e y loca ion can be p ocessed.
No e ha , al hough Eqs. 1.17 and 1.18 desc ibe a bidimensional con olu ion,
1.2. Con olu ional Neu al Ne wo ks 17
his ope a o can be applied o signals ha ing any a bi a y numbe o spa ial
dimensions by choosing he ke nels app op ia ely.
E en be o e CNNs, con olu ion has been widely used in compu e ision
o many di e en asks. Fo ins ance, con olu ion is cen al in he well-known
Canny edge de ec o algo i hm [39] as i is used bo h o image smoo hing
(using gaussian ke nels) and o g adien compu a ion ( ypically done using
Sobel ke nels). Ano he di ec applica ion o con olu ion in compu e ision is
image sha pening, which makes use o high-pass ke nels.
1.2.2 Con olu ional Neu al Ne wo ks
A Con olu ional Neu al Ne wo k (CNN) can be de ined as a Neu al Ne wo k
in which a leas one o i s hidden laye s is con olu ional ins ead o ully-
connec ed. A con olu ional laye is a laye in which he basis unc ions in
Φ={φWi,bi
i}M
i=1 ake he ollowing o m:
φWi,bi
i(x) = σ(x∗Wi+bi).(1.19)
The di e ences wi h espec o a ully-connec ed laye s a e as ollows:
•each weigh e m Wi∈RK1×K2×Cis now a ke nel, applied o he inpu
x∈RH×W×C(in case o bidimensional con olu ion);
•each bias e m bi∈Ris a scala alue which is added independen ly o
each ou pu o he con olu ion;
• he ac i a ion unc ion σ(·)is applied independen ly o each elemen
esul ing om he abo e ans o ma ion.
As a esul , he ou pu o each unc ion φiis no longe a scala , bu a he
a ma ix in RH×W(assuming p ope padding is used), and he ou pu o he
con olu ional laye Φis in RH×W×M. A bidimensional con olu ional laye can
he e o e be in e p e ed as a laye ha akes a signal ha ing spa ial dimensions
H×Wand Cchannels as inpu , and p oduces a new signal ha ing he same
spa ial size and Mchannels in which each elemen is a ea u e ep esen a ion
18 Chap e 1. P io A
o he neighbo hood o he same elemen in he inpu . This ou pu signal is
commonly e e ed o as ea u e map, while each o i s elemen s ep esen s a
neu on. The con olu ional ke nels used o compu e he ea u e maps, ins ead
o being hand designed o pe o m speci ic ope a ions (e.g. smoo hing o g a-
dien compu a ion) a e lea ned du ing op imiza ion, yielding he se o local
ans o ma ions ha a e mos app op ia e o he ask a hand.
Con olu ional laye s b ing se e al ad an ages o e hei ully-connec ed
coun e pa :
• he esul ing model con ains conside ably less pa ame e s. Fo example,
i he inpu o a ully-connec ed laye con ains Ielemen s, he weigh s o
each basis unc ion mus con ain Ielemen s. Fo high dimensional inpu s,
such as images, his would quickly lead o p ohibi i ely la ge models. The
numbe o weigh s in con olu ional laye s, on he o he hand, depends
only on he amoun o channels o he inpu , which a e gene ally limi ed
compa ed o hei spa ial dimensions;
• he con olu ional ope a ion is equi a ian o ansla ions in he inpu .
In o he wo ds, i he inpu signal is shi ed along one o mo e o i s
spa ial dimensions, he co esponding ou pu alue does no change bu
i is subjec o he same shi . This p ope y is e y powe ul, as i im-
plies ha con olu ion is able o cap u e local pa e ns wi hin an inpu
independen ly o whe e hose pa e ns a e. This is in con as o ully-
connec ed laye s, in which a shi o he inpu signal leads o an en i ely
di e en ou pu .
The dep h o he model plays an especially impo an ole o CNNs, as s acking
mul iple con olu ional laye s one a e he o he allows he neu ons o he
model o p og essi ely expand hei ecep i e ield o e he inpu . The ecep i e
ield o a neu on can be de ined as he po ion o he inpu ha neu on is a
unc ion o . Suppose ha he i s hidden laye o a CNN is comp ised o
ke nels ha ing spa ial size 3×3. The ecep i e ield o he neu ons p oduced
1.2. Con olu ional Neu al Ne wo ks 19
by his laye would be 3×3, as each neu on is a unc ion o a 3×3 egion
o he inpu . I we now apply a second con olu ional hidden laye , again wi h
ke nel size 3×3, o he ea u e map p oduced by he i s laye , he esul ing
neu ons would ha e a ecep i e ield o 5×5o e he inpu (IMAGE). By
s acking mul iple con olu ional laye s, i is possible o p og essi ely inc ease
he ecep i e ields o he neu ons a bi a ily, po en ially allowing each one o
"see" he en i e inpu image. This has p o en o be e y powe ul in p ac ice
as i allows o e y "s ong" hie a chical ep esen a ions o be lea ned by
he model: he neu ons o he ea ly hidden laye s lea n o ecognize low-le el
gene al pa e ns, such as edges, do s o o he basic shapes. Neu ons om la e
laye s become p og essi ely mo e specialized and abs ac , as hey ha e access
o mo e con ex ual in o ma ion, and lea n o ecognize high-le el s uc u es
ha a e use ul o sol ing he speci ic ask [40, 41]. This phenomenon is one
o he mos impo an con ibu o s o he obus ness o his class o models
o o e i ing, as well as hei abili y o pe o m well e en when ained on
limi ed da a.
In p ac ice, mode n deep con olu ional models do no keep he spa ial di-
mensions cons an h oughou he en i e model, bu a he hey p og essi ely
educe hem, gene a ing inc easingly smalle ea u e maps. This is achie ed
ei he by using s ided con olu ional laye s o pooling ope a o s. S ided con-
olu ions ope a e exac ly like no mal con olu ions, wi h he di e ence ha
some inpu loca ions a e skipped. In pa icula , he s ide o his ope a o de-
e mines how many posi ions he ke nel "mo es" when compu ing he nex
alue: o ins ance, is he s ide is equal o 2 his means ha e e y o he el-
emen in he inpu will be igno ed, leading o an ou pu whose spa ial size is
hal o ha o he inpu . Pooling ope a o s can be seen as a special case o
s ided con olu ions in which he ke nels a e no op imized, bu a he pe -
o m a speci ic kind o ope a ion o e he inpu . The mos widely used a e Max
Pooling and A e age Pooling, in which he ke nels pe o m a max and mean
ope a ion espec i ely. No mal con olu ions can also be seen as a special case
o s ided con olu ion whe e s ide is equal o 1. The ad an age o p og es-
26 Chap e 1. P io A
The cu en s a e-o - he-a o segmen a ion models is ep esen ed by he
DeepLab amily o a chi ec u es [4, 57, 58]. In DeepLab [4], he au ho s adop
an ImageNe -p e ained VGG16 model as base o he segmen a ion. To deal
wi h he limi a ion ha VGG16 p oduces a ea u e map a 1/32 he o iginal
esolu ion (which in u n would p oduce e y coa se segmen a ion maps), he
au ho s emo e he las wo pooling laye s o he model, and employ dila ed
con olu ions in he ollowing laye s o make up o he loss o ecep i e ield.
The esul ing segmen a ion map is hen upsampled om 1/8 he esolu ion o
he o iginal inpu size using bilinea in e pola ion, and Condi ional Random
Fields (CRF) a e used o u he imp o e he esul and be e segmen he
ine de ails. DeepLab 2 [57], ex ends DeepLab by in eg a ing ResNe as addi-
ional backbone o be e pe o mance and by in oducing he A ous Spa ial
Py amid Pooling (ASPP) module, whe e mul iple ke nels ha ing mul iple dila-
ion a es a e used o cap u e objec s a di e en sizes and scales. DeepLab 3
[58] augmen s he ASPP module wi h image-le el ea u es and emo es he
cos ly CRF pos -p ocessing, while s ill pe o ming conside ably be e han i s
p edecesso s.
1.5 Objec de ec ion
Seman ic segmen a ion p o ides e y powe ul in o ma ion o au onomous
d i ing sys ems. Fo example, i can be applied o de e mine he ee d i ing
space, o o loca e a ic lanes o o he kinds o ho izon al a ic signs. I
ep esen s, howe e , edundan in o ma ion when he a ge s o de ec ion a e
obs acles such as ca s o pedes ians o mo e complex en i ies such as e ical
a ic signs. In hese cases, since seman ic segmen a ion does no p o ide an
ins ance-le el dis inc ion bu only e u ns he class each pixel belongs o, addi-
ional pos -p ocessing s eps would be equi ed o de e mine he exac loca ion
o each single objec .
A class o p oblems ha can be conside ed complemen a y o seman ic
segmen a ion is objec de ec ion. Objec de ec ion can be de ined as ollows:
1.5. Objec de ec ion 27
gi en an inpu x(e.g. an image, a LiDAR scan e c...), he goal is o loca e
wi hin his inpu all he en i ies o in e es as well as o desc ibe hei s a es
{Si}N
i=1. This s a e migh include he posi ion o he en i y, i s class and o he
p ope ies such as size o eloci y. Wha makes objec de ec ion pa icula ly
in e es ing and challenging is ha he numbe o de ec ions, and he e o e
he numbe o ou pu s o be e u ned by sys em, g ea ly a ies depending
on he inpu . Fo ins ance, one image migh display se e al ens o ca s o
be de ec ed o none a all. This is in con as wi h he p e iously p esen ed
p oblems, whe e he dimension o he ou pu was independen o he con en
o he inpu : in classi ica ion he ou pu is always a single p obabili y ec o ;
in image segmen a ion he ou pu is always a p obabili y ec o pe image
pixel. To deal wi h his complica ion, he gene al app oach adop ed in deep
lea ning consis s on es ima ing a e y high numbe o possible objec s and
p og essi ely supp ess edundan and nega i e in o ma ion o ob ain he inal
se o high-quali y es ima ions.
1.5.1 2D Objec De ec ion on Images
P obably he mos esea ched a ian o objec de ec ion in deep lea ning is 2D
objec de ec ion on images. In his case, he s a e o each objec is ep esen ed
by i s class and by he smalles axis-aligned 2D bounding box con aining i ,
ha is Si= (ci, xi, yi, wi, hi), whe e ci ep esen s he class o he i- h objec , xi
and yi he cen e o i s bounding box in he image and hiand wii s dimensions.
Pionee ing esea ch in he di ec ion o ully deep lea ning-based objec de ec-
ion is ep esen ed by R-CNN (Region-based CNN) [5]. In his wo k, objec
de ec ion is achie ed ia a combina ion o Selec i e Sea ch [59] and ImageNe -
p e ained AlexNe . Selec i e Sea ch is an algo i hm ha , gi en an image,
e u ns a se o bounding boxes ha a e likely o be loca ed in in e es ing
egions o he image. The gene a ed se o boxes has high ecall (i is likely
o con ain all in e es ing objec s), low p ecision (mos o he boxes a e w ong,
inaccu a e o duplica es) and is gene ally e y la ge (a ound ∼2000 boxes a e
gene a ed o a 600x600 image). Mo eo e , he algo i hm is class-agnos ic, ha
28 Chap e 1. P io A
is i does no speci y wha kind o objec is wi hin each o he boxes. RCNN
uses Selec i e Sea ch o de e mine an ini ial se o de ec ions, called egion p o-
posals; each p oposal is hen used o gene a e a new image, o size 227 ×227,
by c opping and eshaping i s con en , ha is subsequen ly gi en as inpu o
a p e ained AlexNe model, which is asked o classi y he p oposal con en
as well as o compu e a co ec ion o he box loca ion and size. In pa icu-
la , he las ully-connec ed laye o AlexNe is eplaced wi h wo new pa allel
laye s: a laye ha ing N+ 1 ou pu s, which ep esen s he p obabili y ec o
o he desi ed Nclasses plus he backg ound class o nega i e boxes, and
a laye o bounding box eg ession, which e u ns ou alues co esponding
o he box co ec ions in posi ion and dimensions. While su passing all o he
me hods by a la ge ma gin, howe e , RCNN is ex emely slow, equi ing o e
40 seconds pe image. This is due o he ac ha AlexNe mus be applied o
each o he 2000+ egion p oposals, esul ing in a sys em ha is imp ac ical
o au onomous d i ing si ua ions, whe e high ame a es a e essen ial.
A s ep owa ds imp o ing upon his limi a ion is he ollow-up wo k Fas
R-CNN [60]. The majo con ibu ion o his wo k consis s in ea u e sha -
ing among di e en p oposals: he e he con olu ional backbone, which is now
VGG16, is applied o he en i e inpu image, ob aining a ea u e map ha
is sha ed o all objec s. Then, o each egion p oposal e u ned by Selec i e
Sea ch, he co esponding a ea in he ea u e map is de e mined by p ojec ing
he box and wa ped o a s anda d 7x7 size using he so called RoI Pooling
ope a ion. This pooled se o ea u es is inally used o classi ica ion and box
eg ession. This o mula ion has wo majo ad an ages o e he o iginal design:
on he one hand, applying he con olu ional backbone o he en i e image in-
s ead o once pe p oposal esul s in a conside ably as e compu a ion, o e
200 imes as e han RCNN. On he o he hand, sha ing he same se o ea-
u es o all p oposals esul s in a s onge in e media e ep esen a ion ha is
mo e awa e o he con ex , which leads o imp o ed pe o mance.
Despi e being conside ably as e han RCNN, Fas R-CNN is s ill bo -
lenecked by Selec i e Sea ch, which equi es a ound 2 seconds pe image o
1.5. Objec de ec ion 29
gene a e he p oposal boxes, slowing down he en i e pipeline conside ably. A
solu ion o his p oblem is in oduced by Fas e R-CNN [9]. In his wo k, a
speci ic subne wo k, called Region P oposal Ne wo k (RPN), is used ins ead
o Selec i e Sea ch o es ima e he se o p oposal boxes. RPN gene a es he
p oposal boxes by assigning o each pixel o he ea u e map gene a ed by he
backbone a se o p ede ined boxes, called ancho s. Fo each ancho , i hen
es ima es whe he i is an in e es ing o a backg ound box, as well as compu -
ing a co ec ion o i . The p oposal boxes a e inally ob ained by emo ing
he backg ound ancho s and by il e ing duplica es ia Non-Maximum Sup-
p ession. These p oposals a e hen p ocessed as in Fas R-CNN o ob ain he
inal de ec ions. Fu he de ails abou his app oach will be e iewed in he
nex chap e , as Fas e R-CNN is cen al o bo h he p oposed Pa king Slo
De ec ion ne wo k as well as he 3D Objec De ec ion ne wo k. By a oiding
Selec i e Sea ch, and pe o ming bo h egion p oposal and bounding box es i-
ma ion in a uni ied o wa d pass o he model, Fas e R-CNN is conside ably
as e han i s p edecesso s, allowing o mo e han 15 images o be p ocessed
pe second. Mo eo e , i achie es op pe o mance, making i one o he cu en
s a e-o - he-a app oaches. This me hod is u he imp o ed in a subsequen
wo k [11], whe e a Fea u e Py amid Ne wo k is used as backbone in o de o
ex ac ea u e maps a mul iple scales. Ancho s and p oposal boxes a e hen
assigned o he p ope scale depending on hei size.
The R-CNN app oaches a e widely ega ded as wo-s age app oaches. This
is due o he ac ha he inal de ec ions a e ob ained h ough wo sepa a e
phases: i s , a se o p oposal boxes is gene a ed. Then, each p oposal is ana-
lyzed indi idually (i.e. h ough he pooling o i s ea u es) in o de o compu e
i s class and e ine i s s a e. Ano he class o objec de ec ion echniques is
ep esen ed by single-s age me hods: he e, he inal de ec ions a e ob ained
di ec ly, wi hou exploi ing an in e media e se o objec p oposals.
One o he i s wo ks in his di ec ion is ep esen ed by YOLO (You Only
Look Once) [6]. He e, objec de ec ion is cas as a single con olu ional (nd :
he e a e 2 FC laye s ho) neu al ne wo k ha es ima es he bounding boxes
30 Chap e 1. P io A
di ec ly om he inpu image. To achie e his, he inpu image is di ided in o
an S×Sg id o cells, and he ne wo k is asked o es ima e Bpo en ial boxes
pe cell. To do his, he inpu is p ocessed by a s ack o con olu ional and
pooling ope a ions such ha an S×S×(B∗5 + C) ea u e map is ob ained.
Each pixel o his map is esponsible o e u ning he boxes o he objec s
whose cen e s lie in he co esponding g id cell o he inpu : in pa icula ,
he ne wo k e u ns Bpossible boxes pe cell wi h he espec i e con idences,
cen e coo dina es and dimensions, as well as a C-class p obabili y ec o ha
indica es he ca ego y o he de ec ed objec . The uni ied design adop ed by
his app oach allows o ex emely as p edic ions: he base model, which is
a modi ied GoogLeNe a chi ec u e, is able o pe o m de ec ion a 45 ames
pe second, while he ligh model, which is simila o he base model bu wi h
less laye s, eaches 155 ames pe second.
Simila ly o YOLO, SSD (Single-Sho Mul ibox De ec o ) [7] app oaches
objec de ec ion using a single, uni ied, con olu ional neu al ne wo k. Ins ead
o p edic ing he boxes di ec ly, howe e , in SSD he de ec ions a e ob ained
as e inemen s o a se o p ede e mined boxes (simila ly o he ancho s in
Fas e R-CNN). Mo eo e , ins ead o using a single ea u e map o pe o m
de ec ion, mul iple ea u e maps a di e en scales a e used o p edic ion:
he low- esolu ion maps a e adop ed o bigge p ede e mined boxes, while he
high- esolu ion ones a e used o smalle p ede e mined boxes. Ha nessing ea-
u es a mul iple scales o di e en ly sized boxes allows he esul ing de ec ions
o be mo e accu a e, conside ably imp o ing pe o mance o e YOLO.
YOLO was subsequen ly imp o ed, leading o wo new models: YOLO 2
[10] and YOLO 3 [61]. YOLO 2 is simila o i s p edecesso wi h a ew weaks
and changes aimed a imp o ing ecall and localiza ion accu acy o he boxes.
Fi s ly, hey imp o e upon he ne wo k a chi ec u e: hey adop a new model
simila o VGG16 as base, called Da kne -19, and hey in eg a e Ba ch No mal-
iza ion, which speeds up aining and imp o es pe o mance. Mo eo e , hey
s eng hen he p e aining on ImageNe . Secondly, hey adop an ancho -based
design, whe e he p edic ions a e ob ained om e ining p ede e mined boxes
1.5. Objec de ec ion 31
ins ead o being compu ed di ec ly. In pa icula , ins ead o hand-picking he
ancho dimensions like SSD o Fas e R-CNN, hey ob ain hem by unning k-
means clus e ing on he da ase s’ g ound u h boxes; his way he dis ibu ion
o ancho s is close o ha o he g ound u h boxes, which simpi ies op imiza-
ion and leads o imp o ed pe o mance. Finally, hey inc ease he obus ness
o he sys em o di e en objec scales by adop ing mul i-scale aining: e e y
10 epochs, a new andom scale is selec ed and he model is ained on images
esized o ha scale. YOLO 3 u he boos s pe o mance by adop ing a much
deepe model wi h esidual connec ions and by pe o ming p edic ion using
ea u e maps a mul iple scales.
All p e ious app oaches sha e a common p oblem: on a e age, he numbe
o o e all backg ound g id cells o ancho s is a g ea e han hose ac ually
con aining objec s. This disc epancy o en leads o subop imal op imiza ion o
he classi ie , since he g adien alue is o e whelmed by backg ound examples.
Some app oaches, such as he RCNN amily and SSD, deal wi h his p oblem
by o cing he loss unc ion o be compu ed on a balanced se o posi i e and
nega i e samples. E en in his case, howe e , mos o he aining signal is
domina ed by easily classi ied backg ound examples.
A echnique o en adop ed o deal wi h his limia a ion is OHEM [62]
(Online Ha d Example Mining): he e, he loss unc ion is applied o all he
samples, bu he g adien is compu ed only o he nsamples o which he
ne wo k pe o ms he wo s , ha is he nsamples o which he loss alue is
highes . This has he e ec o o cing he aining o ocus only on he mo e
di icul cases, igno ing he easie ones.
Ano he solu ion is p oposed in [63]. In his wo k, he au ho s in oduce
Re inaNe , a single-s age objec de ec o . This model adop s ResNe wi h FPN
as backbone, and pe o ms objec de ec ion a mul iple scales by classi ying
and e ining ancho s, simila ly o SSD. To deal wi h he inbalance be ween
posi i e and nega i e classes, hey in oduce he ocal loss: in his o mula ion,
ins ead o compu ing he loss alue as he a e age o c oss-en opy losses o
each sample, he loss alue is compu ed as he sum o he c oss-en opy losses
32 Chap e 1. P io A
o each elemen , bu each one is downweigh ed depending on how well he
model pe o ms on i . Ha d examples will ecei e high weigh s, whe eas simple
examples will be assigned p og essi ely lowe weigh s. As a esul , du ing op-
imiza ion, aining au oma ically concen a es on he ha d examples; his is
con a y o he o iginal c oss-en opy o mula ion, in which he con ibu ion o
he aining o he ew ha d examples is decima ed by he a e aging ope a ion
o e all samples. The ocal loss, mo eo e , ca ies an ad an age o e OHEM:
in OHEM easy samples a e igno ed; ins ead, when using ocal loss all samples
con ibu e o he op imiza ion, albei wi h educed impac i easy.
1.5.2 T adeo be ween One-s age and Two-s age echniques
The p e ious pa ag aph in oduced some o he main s a e-o - he-a objec
de ec o s, ca ego izing hem in wo-s age and single-s age pipelines.
Two-s age de ec o s pe o m de ec ion in wo dis inc phases: i s , hey
gene a e a se o p oposals, ha is a se o in e media e de ec ions ha ing
high- ecall (mos equi ed en i ies a e de ec ed) bu low p ecision (mos p o-
posals a e w ong o inaccu a e). Then, a second pa o he model analyzes
each p oposal by pooling i s ea u es and ou pu s a classi ica ion decision as
well as a co ec ion o i s s a e. P oposal gene a ion can ei he be pe o med
using exis ing algo i hms, such as Selec i e Sea ch, o by employing a Region
P oposal Ne wo k. On he o he hand, single-s age me hods pe o m es ima-
ion di ec ly, wi hou making use o an in e media e se o p oposals.
The e is no clea winne be ween he wo classes o me hods; which pa adigm
o use depends on he speci ic use-case and si ua ion. Single-s age me hods end
o be as e han hei wo-s age coun e pa while also being simple om a
logical s andpoin , equi ing less complex code and being mo e s aigh o wa d
o implemen and debug. Two-s age me hods, on he o he hand, ha e an edge
pe o mance-wise. Whe e he g ea es ad an age o wo-s age me hods lies,
howe e , is hei a g ea e lexibili y: by adop ing egion p oposals and by
pooling hei ea u es, i is conside ably easie o in eg a e addi ional asks
in o he o iginal pipeline.
1.6. 3D Objec De ec ion 33
The mos no o ious wo k in his di ec ion is ep esen ed by Mask R-CNN
[64]. Mask R-CNN expands Fas e R-CNN o include a segmen a ion mask p e-
dic ion o each p oposal, on op o he classi ica ion and posi ion e inemen .
This is achie ed by adding an ex a con olu ional module which p ocesses he
pooled ea u e map o each objec and e u ns an upsampled bina y mask
o ha objec . To u he imp o e pe o mance and be e align he pooled
ea u es o he seman ic ask, he au ho s p opose RoIAlign, an upg ade o
he RoIPooling ope a o , which a oids he quan iza ion e ec s o he la e by
adop ing adop ing bilinea in e pola ion.
Ano he in e es ing line o esea ch ha di ec ly akes ad an age o he
wo-s age s uc u e is T acking wi hou Bells and Whis les [65]. In his wo k,
he au ho s p opose a small modi ica ion o Fas e R-CNN o di ec ly pe o m
acking, wi hou he need o addi ional aining o acking speci ic da a. In
pa icula , he de ec ions o each ame a e added o he p oposal pool o he
successi e ame and co ec ed by he e inemen ne wo k o di ec ly c ea e
ajec o ies, while he o iginal se o p oposals is used o iden i y new objec s
and ini ialize new acks.
Gi en his ex a lexibili y and pe o mance, wo-s ages me hods, and es-
pecially Fas e R-CNN, cons i u e he ounda ion o wo o he solu ions p o-
posed in his hesis: he pa king slo de ec ion ne wo k and he monocula 3D
objec de ec ion ne wo k.
1.6 3D Objec De ec ion
While being use ul o applica ions such as moni o ing sys ems o secu i y
came as, in au onomous d i ing de ec ing objec s a image le el is o en in-
su icien , as i gi es no in o ma ion on whe e each ins ance ac ually is in he
wo ld. As a esul , new sys ems and algo i hms a e con inuously being de el-
oped ha pe o m de ec ion in 3D. Fo mally, 3D objec de ec ion consis s o
loca ing in he inpu all ins ances o objec s o in e es and e u n, o each
one:
34 Chap e 1. P io A
•i s class c;
•i s 3D bounding box, exp essed wi h espec o some ame o e e ence.
This box is iden i ied by i s cen e poin (x, y, z), i s dimensions (h, w, l)
and i s o ien a ion exp essed as oll, pi ch and yaw angles (φ, ψ, θ).
In o he wo ds, he s a e o he objec iis now ep esen ed by he ec o
Si= (ci, xi, yi, zi, hi, wi, li, φi, ψi, θi).(1.20)
When ope a ing in d i ing scena ios, a common assump ion is ha all objec s
o in e es lie on a plane, and can he e o e be subjec only o o a ions a ound
one axis, simpli ying he s a e o Si= (ci, xi, yi, zi, hi, wi, li, θi).
Inpu da a used o 3D objec de ec ion sys ems o en include images, ei he
coming om monocula o s e eo came a se ups, and poin clouds, ob ained
by dispa i y map iangula ion o LiDAR senso s.
1.6.1 LiDAR-based 3D Objec De ec ion
LiDAR (Ligh De ec ion And Ranging) senso s a e commonly adop ed choices
o au onomous d i ing as hey a e able o p o ide highly accu a e dep h in-
o ma ion abou he en i onmen . They ope a e on he ime-o - ligh p inciple:
dis ance om an objec is calcula ed om he ime equi ed by a ligh impulse,
emi ed by he senso i sel , o each he objec and be e lec ed back o he
senso ecei e . The esul ing aw da a e u ned by his class o senso s is a
poin cloud, ha is a se o 3D poin s in space. Some senso s migh also e-
u n addi ional in o ma ion o each o hese poin s, such as he e lec i i y
o he co esponding objec s. The mos ypically adop ed choice in au omo-
i e is ep esen ed by Velodyne senso s, due o hei abili y o apidly pe o m
360◦scans while also ha ing easonable e ical Field-o -View and poin den-
si y h ough he use o mul iple emi e - ecei e pai s ( e i y). The esul ing
scans usually con ain om se e al ens o housands o a ew hund ed housand
poin s, depending on he senso model and he numbe o scan planes.
1.6. 3D Objec De ec ion 35
Poin clouds a e conside ably di e en om images: he la e a e dense
g ids o elemen s, o ganized in a speci ic s uc u e whe e posi ion ma e s. On
he o he hand, poin clouds a e spa se and a e no cha ac e ized by a speci ic
o de ing: a pe mu a ion o a poin cloud is equi alen o he o iginal poin
cloud. As a esul , con olu ional models, which a e based on he assump ion
ha he inpu is dense and egula , canno be di ec ly applied o his da a
ype.
An ea ly wo k in deep lea ning-based 3D objec de ec ion [66] a emp s a
b idging he gap be ween image and poin cloud da a modali ies by p ojec ing
he LiDAR poin cloud on a Bi d’s Eye View (BEV) plane, which is hen used
as inpu o a con olu ional model ha di ec ly pe o ms 3D bounding box
p edic ion and classi ica ion. Ano he simila wo k is ep esen ed by MV3D
[67]: he e, bo h LiDAR and image da a a e used oge he o pe o m de ec-
ion. In pa icula , h ee sepa a e con olu ional backbones a e used o ex ac
ea u e maps: one om he image, one om a Bi d’s Eye View p ojec ion o
he LiDAR poin cloud and one om a on al iew p ojec ion o he same
cloud. Then, he Bi d’s Eye iew ea u es a e used o gene a e a se o ini ial
3D p oposal boxes ia an RPN. Finally, hese p oposals a e p ojec ed back on
all h ee ea u e maps, whose ea u es a e pooled acco dingly, used and used
o es ima e he inal boxes. AVOD [12] adop s a simila s a egy o using
image and LiDAR in o ma ion, bu only uses he Bi d’s Eye View p ojec ion
o he poin cloud, gene a ed simila ly o MV3D. In his me hod, howe e , he
p oposal boxes a e gene a ed using bo h image and Bi d’s eye iew LiDAR
ea u es, esul ing in imp o ed ecall o small ins ances.
Despi e enabling all p e ious me hods o ha ness he ep esen a ional powe
o con olu ional models and p e-exis ing a chi ec u es, poin cloud p ojec ion
ine i ably leads o in o ma ion loss, limi ing he abili y o such models o ea-
son in 3D and ul ima ely comp omising pe o mance. Cu en ly, he s a e o
he a app oaches o 3D Objec De ec ion can be classi ied in wo mac o
ca ego ies: oxel-based me hods and poin -based me hods.
In oxel-based me hods, he inpu poin cloud is i s con e ed in o a oxel
42 Chap e 1. P io A
p oaches o 3D de ec ion: s e eo-based me hods, which use pai s o ec i ied
images as inpu , and monocula -based me hods, which aim a de ec ing 3D
objec s om a single image.
S e eo-based De ec ion
A ecen wo k on deep lea ning-based s e eo 3D de ec ion can be iden i ied in
T iangula ion Lea ning Ne wo k [76]. This app oach i s p oposes a monocula
baseline ha pe o ms 3D de ec ion om a single ame, and hen ex ends i
o he s e eoscopic case. The baseline ope a es simila ly o Fas e R-CNN,
wi h he di e ence ha all es ima ions e ol e a ound 3D boxes: in pa icula ,
3D ancho s a e displaced in he 3D space and hen p ojec ed on he image o
de e mine hei 2D coun e pa s. Fea u es om he 2D ancho s a e hen pooled
and p ocessed by he RPN, which de e mines 3D p oposals. Finally, he same
p ocess is epea ed using he 3D p oposals, whose ea u es a e passed o he
de ec ion po ion o he model o de e mine he inal p edic ions. To ex end
his amewo k o s e eoscopic da a, bo h le and igh ames a e p ocessed
in pa allel by he ne wo k, he 3D ancho s and p oposals a e p ojec ed on
bo h ames and he espec i e ea u es a e used using an ad-hoc scheme ha
accoun s o po en ial misma ches due o di e en dep hs. The addi ion o
he second image and he usion scheme leads o mode a e imp o emen s in
pe o mance; despi e his, howe e , his me hod pe o ms conside ably wo se
han o he con empo a y s e eo app oaches and is e en ou pe o med by some
monocula pipelines.
Ano he simila app oach is S e eo R-CNN [17], which also builds upon
Fas e R-CNN. Ins ead o elying on 3D ancho s o p oposals, howe e , his
me hod uses he le and igh ea u es compu ed by he same backbone o
p edic ma ching le - igh 2D p oposals om 2D ancho s. These p oposals
a e hen used o pool he co esponding le - igh ea u es, which a e inally
employed o es ima e he le - igh 2D boxes, a se o image keypoin s, as
well as he co esponding 3D objec dimensions and o ien a ion. Using his
in o ma ion, he inal 3D box is ob ained ia iangula ion. To u he imp o e
1.6. 3D Objec De ec ion 43
pe o mance, a inal alignmen phase is pe o med, in which he dep h o he
de ec ions is co ec ed by minimizing he pho ome ic e o o objec s in he
le and igh images. This co ec ion, in pa icula , p o es o be cen al o he
me hod, accoun ing o mos o he pe o mance gain.
A seminal wo k owa ds accu a e s e eoscopic (and monocula ) 3D objec
de ec ion is ep esen ed by Pseudo-LiDAR [77]. The idea behind his esea ch
is su p isingly simple: he pe o mance gap ha exis s be ween LiDAR-based
and S e eo-based de ec o s is no o be a ibu ed solely o he echnological
di e ences be ween he wo ypes o senso s, bu also o he da a ep esen-
a ion ha is used o ain he models. In ac , wha he au ho s obse ed
is ha he poin clouds esul ing om dispa i y iangula ion a e indeed no
inaccu a e enough o jus i y such a wide di e ence in esul quali y. To ali-
da e his hypo hesis, hey in oduced a wo-s ep pipeline: i s , s a e-o - he-a
app oaches a e used o ex ac a dispa i y map om he s e eo pai , which is
hen con e ed in o a poin cloud; second, s a e-o - he-a LiDAR-based 3D de-
ec o s, such as AVOD and F us um-Poin ne s, a e applied o he poin cloud
o pe o m de ec ion. The esul ing sys em decisi ely ou pe o ms all pu ely
image-based s e eo sys ems, alida ing he claim. One o he main easons as
o why a poin cloud ep esen a ion allows o inc eased pe o mance com-
pa ed o an image-based one is as ollows: p ocessing poin -cloud da a, ei he
by using 3D con olu ions, 2D con olu ions on BEV ep esen a ions o Poin -
Ne s, ensu es ha he elemen s ha a e ope a ed upon oge he a e physically
close in space. Con olu ions on images, on he o he hand, ope a e iden ically
on pa ches co esponding o objec s a di e en scales ( a away objec s a e
smalle on images compa ed o nea by objec s) o pa ches in-be ween objec s
and backg ound (and he e o e in ol ing en i ies e y a away in physical
space), making hem less sui able o easoning in 3D.
Monocula De ec ion
Con a y o LiDAR-based o S e eo-based 3D de ec ion, monocula de ec ion is
an ill-posed p oblem, as a single image does no p o ide enough in o ma ion o
44 Chap e 1. P io A
eco e he scale o he po ayed scene, and he e o e he dep h o he objec s.
One way o eco e a good enough app oxima ion o he posi ion o in e es ing
objec s is o use a p io i knowledge abou hem, such as hei dimension.
Fo ins ance, i he heigh o a speci ic a ic sign is known hen i would
be possible, gi en he came a in insic pa ame e s, o in e he app oxima e
dis ance om he senso . Ano he possible way is by lea ning he ela ionship
be ween he way objec s appea in he image plane and hei co esponding
s a e in he wo ld using accu a e g ound u h and deep lea ning models. This
would be simila o how humans wi h one eye co e ed would s ill be able
o es ima e he 3D s uc u e o he wo ld, despi e no ha ing, heo e ically,
enough in o ma ion o do so.
Due o he challenging na u e o he p oblem, o ease lea ning and imp o e
pe o mance mos monocula 3D de ec ion pipelines embed some o m o a
p io i knowledge o 3D easoning mechanism di ec ly wi hin hei models. In
Mono3D [15], he au ho s le e age he assump ion ha all objec s should lie
on he g ound plane in o de o gene a e objec p oposals. In pa icula , hey
use came a calib a ion in o ma ion o de e mine a ixed plane, hey gene a e
3D candida e p oposals on his plane and hey p ojec hem on he image in
o de o sco e hem and keep only he mos p omising ones. The o e all sco e
o each candida e is de e mined using seman ic, ins ance, shape, loca ion and
con ex cues, which a e compu ed using ex e nal me hods. The bes candida e
p oposals a e hen p ocessed u he by a second s age simila ly o Fas -RCNN:
a VGG16 model is used o ex ac ea u es om he inpu image, RoIPooling
is used o ob ain he ea u es o each p oposal and ully-connec ed laye s a e
used o es ima e he class, he posi ion o he objec as well as i s o ien a ion.
OFTNe [78] adop s a ResNe backbone o ex ac mul i-scale ea u e maps
om he inpu image. Then, an o hog aphic ea u e ans o m is in oduced
o map he image-le el ea u es o a BEV ep esen a ion, in o de o allow
u he compu a ion o eason in 3D wi hou pe spec i e e ec s. To ob ain
his ep esen a ion, a oxel g id ixed o he g ound plane is gene a ed, hen
each oxel o he g id is p ojec ed on o he ea u e maps and all ea u es
1.6. 3D Objec De ec ion 45
wi hin he p ojec ion a e accumula ed in o he oxel. Finally, he oxel g id is
collapsed in o a 2D ep esen a ion by accumula ing ea u es along he heigh
di ec ion. As las s ep, each loca ion o he BEV ea u e map is p ocessed in
o de o classi y whe he he e is an objec as well as o de e mine he cen e ,
dimensions and o ien a ion o he objec .
In MonoGRNe [79] he 3D bounding box es ima ion p ocess is spli in o
ou sequen ial sub asks pe o med by a single, uni ied neu al ne wo k. In
he i s s ep, image ea u es a e compu ed using a VGG16 backbone and
2D bounding boxes a e ex ac ed by adop ing a wo-s age objec de ec ion
pipeline. Gi en he 2D boxes, RoIAlign is used o pool he espec i e se o
ea u es which a e used o he emaining h ee s eps. Fi s , he dep h o each
ins ance is es ima ed. Then, gi en he dep h, he ue coo dina es o each in-
s ance cen e a e eg essed. Finally, he posi ions o he eigh e ices o each
box a e compu ed wi h espec o each box local ame o e e ence placed in
each es ima ed cen e .
Mul iFusion [80] le e ages a p e ained model o monocula dep h p edic-
ion o compu e a dep h map o each image. This map is hen conca ena ed o
he image i sel and ed o a VGG16 backbone ollowed by an RPN o ex ac
egion p oposals. Gi en he 2D p oposals loca ions, RoI Max Pooling is used o
ex ac he co esponding ea u es om he ea u e map and RoI Mean Pooling
is used o ex ac a ea u e ep esen a ion om he poin cloud gene a ed om
he dep h map. These wo se s o ea u es a e hen conca ena ed and used o
classi y each p oposal and de e mine hei co esponding 3D box.
Deep MANTA[81] de ec s ehicles ia pa es ima ion ollowed by em-
pla e ma ching. Fi s , a s anda d VGG16 backbone ollowed by an RPN a e
applied o he inpu image o ob ain egion p oposals. Then, hese p oposals
go h ough wo cascaded e inemen s ages, in ol ing RoIAlign on he p oposal
coo dina es ollowed by box co ec ion. Besides he co ec ion, he second e-
inemen s age also ou pu s, o each box, i s classi ica ion sco e, a ec o o
2D image coo dina es co esponding o he objec pa s (i.e. y es, headligh s
e c...) as well as an es ima ion o he simila i y o i s 3D dimensions wi h e-
46 Chap e 1. P io A
spec o a se o ixed empla es. Follows a inal 2D/3D ma ching phase in
which he p edic ed simila i ies as well as he 2D objec pa s loca ions a e
used o eco e he 3D pose o he objec as well as he 3D coo dina es o
i s pa s by ma ching hem agains a da abase o empla es using he PnP
algo i hm [82].
MonoPSR [16], builds upon p e-exis ing high pe o mance 2D de ec o s and
uses LiDAR da a as addi ional in o ma ion du ing he aining o he model o
imp o e pe o mance. Fi s , a p e ained 2D objec de ec o is applied o he
inpu image in o de o compu e he image-le el boxes. Fo each de ec ion, a se
o ea u es is hen ex ac ed by using oge he he ull-image ea u es pooled
a he box loca ion and a second se o ea u es ob ained by applying a ResNe
backbone o an image-le el c op o he objec . These ea u es a e hen ed o a
p oposal gene a ion module, which es ima es he o ien a ion, dimensions and
posi ion o he objec . Follows a p oposal e inemen module, which u he
co ec s he loca ion o he de ec ion. Finally, his in o ma ion as well as he
a ailable LiDAR da a du ing aining a e used o guide an addi ional ins ance
econs uc ion module o es ima e a poin cloud ep esen a ion o each objec ,
which is hen used o se up addi ional auxilia y loss unc ions.
Chap e 2
Objec De ec ion o Pa king
Slo De ec ion
In his chap e I p esen he p oposed deep lea ning me hod o pa king slo de-
ec ion om su ound iew images. I i s p o ide mo i a ion o he esea ch,
highligh ing he impo ance o he p oblem as well as de ailing o he me hod-
ologies adop ed in li e a u e and hei weaknesses. I hen b ie ly e iew Fas e
R-CNN, as i cons i u es he baseline objec de ec ion model adop ed as s a -
ing poin o his wo k. I ollow up by illus a ing he app oach and he da ase
used o op imiza ion, including he da a p epa a ion p ocess used o c ea e
i . Finally, I desc ibe he expe imen s pe o med o alida e he e ec i eness
and obus ness o he sys em and I discuss he ob ained esul s.
2.1 P io A and Mo i a ion
Ad anced D i e -Assis ance Sys ems (ADAS) a e expe iencing a spike in in-
e es om he esea ch communi y and a e cu en ly one he mos esea ched
echnologies. Among hese sys ems a e, o ins ance, Adap i e C uise Con ol,
which allows he ehicle o au oma ically main ain a speci ic dis ance om
he ehicle in on , o Lane Keeping, which au oma ically keeps he ehicle
48 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
cen e ed in i s lane. Ano he echnology equen ly a ailable on mode n ca s
is he Pa king Assis an which moni o s he su ounding space and, once i lo-
ca es an a ailable pa king spo , i assis s he d i e h oughou he maneu e .
The localiza ion o ee pa king space is o en pe o med using sona senso s
by de ec ing enough unoccupied space o allow he maneu e o be comple ed
success ully. While such a sys em migh wo k well unde he supe ision o a
human d i e , howe e , i is unsui able in he con ex o a ully au oma ed
ehicle: sona s a e only capable o de ec ing acan space, no pa king slo s
di ec ly, and he e o e a e only use ul unde he assump ion ha he ehicle is
in p oximi y o a pa king lo . On he o he hand, came as a e able o pe cie e
ho izon al a ic signs and he e o e can enable ully au oma ed pa king slo
de ec ion.
A lo o esea ch has ocused on ision o pa king slo localiza ion and
occupancy classi ica ion. Many algo i hms ha e been de eloped ha exploi
s a ic came as o moni o occupancy in o de o manage pa king lo s [83,
84, 85]. F om an au onomous d i ing sys em poin o iew, howe e , hese
app oaches a e unsui able as hey all ely on he a p io i knowledge abou he
loca ion o each slo , in o ma ion ha is una ailable when he came as a e
dispaced on a mo ing ehicle.
Resea ch in he di ec ion o au oma ic pa king de ec ion du ing na iga ion
s a ed wi h [86], whe e colo was used as cue o segmen pa king slo ma kings
di ec ly in he image. Mo e ecen app oaches pe o m de ec ion on a Bi d’s Eye
View ep esen a ion o he image ins ead, in which ho izon al oad ma kings a e
mos ly ee o pe spec i e e ec s: o ins ance, in BEV ec angula slo s always
appea ec angula , and line hickness does no depend on he icini y o he
slo o he senso . Mo eo e , in o de o ob ain a comple e 360◦pe cep ion o
he su ounding en i onmen , mos se ups in ol e mul iple calib a ed came as,
whose BEVs a e hen s i ched oge he o ob ain su ound iew images. These
ep esen a ions a e used in app oaches such as [87, 88], in which pa king slo s
a e iden i ied using low le el isual ea u es, such as co ne s and lines. [89] uses
boos ing [90] in o de o classi y c oss-poin s be ween pa king-line segmen s,
2.2. Fas e R-CNN 49
and hen de e mines he en y poin o each slo om hose. The classi ie is
ained on a da ase comp ised o 8600 images, in which he posi ion o each
c oss-poin and he o ien a ion o each slo is anno a ed. In [91], sobel il e s
ollowed by a p obabilis ic Hough ans o m a e used o ex ac lines om
he su ound iew images. Then, he a ailable pa king slo s a e de ec ed by
exploi ing ela ions be ween pa allel lines.
All he abo e app oaches, howe e , su e om a common limi a ion: hey
all ely on hand-designed isual ea u es, and he e o e hey end o pe o m
well only o he speci ic and con olled en i onmen s o which hey a e de-
signed. [89], o example, is only able o de ec ho izon al and e ical slo s,
so i ends o ail in p esence o slan ed pa king slo s on du ing he execu ion
o he pa king maneu e . [91] is ela i ely obus o di e en obse a ion con-
di ions, bu is unable o de ec slan ed slo s and equi es a compu a ionally
expensi e pos p ocessing phase. Colo -based app oaches such as [86] migh ail
in p esence o occlusions, noise o a ia ions in illumina ion condi ions. To deal
wi h hese weaknesses, he me hod p oposed in his hesis pe o ms pa king
slo de ec ion and occupancy classi ica ion om su ound- iew images di ec ly
using a deep con olu ional neu al ne wo k. The co e idea is ha , by allowing
he model o au oma ically lea n om da a wha ea u es a e use ul o de-
ec ing pa king slo s, he esul ing sys em should, gi en enough he e ogeneous
aining da a, show highe obus ness and adap abili y o di e en obse a-
ion condi ions and slo ypes. Be o e going in o de ail abou he app oach, I
b ie ly e iew he 2D objec de ec o Fas e R-CNN [9], as i cons i u es he
ounda ion o he p oposed pipeline.
2.2 Fas e R-CNN
As al eady s a ed in Sec ion 1.5.1, Fas e R-CNN is a wo-s age neu al ne wo k
o 2D objec de ec ion om images. In pa icula , o each de ec ed objec , i
e u ns:
•i s class ci;
50 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
•i s axis-aligned 2D bounding box, desc ibed by i s cen e coo dina es
(ui, i)and i s dimensions (hi, wi).
Mo e speci ically, his sys ems is composed o h ee main modules:
•aBackbone Ne wo k, which is used o ex ac gene ic ea u e ep esen a-
ions om he inpu ;
•aRegion P oposal Ne wo k (RPN), which es ima es a high- ecall ini ial
se o po en ial bounding boxes, called p oposals;
•aDe ec ion Head, which analyzes each p oposal in o de o de e mine
whe he i con ains an objec , classi y he objec and compu e a co ec-
ion o he p oposal box o make i be e i he objec .
Backbone Ne wo k As i s s ep, he image is p ocessed by he backbone,
which is esponsible o ex ac ing a high-le el ea u e ep esen a ion o he
inpu . The mos commonly adop ed choice o his ne wo k consis s o a ResNe
model whose weigh s a e ini ialized ia a aining on he ImageNe da ase o
he classi ica ion ask. Recen ly, he ResNe model is o en enhanced wi h a
Fea u e Py amid ollowing FPN [11], which gene a es a py amid o ea u e
maps a di e en esolu ions by p og essi ely upsampling he las ea u e map
p oduced by ResNe and by using i wi h ea ly laye s maps.
Region P oposal Ne wo k Gi en he gene a ed ea u e map, he RPN
es ima es an ini ial se o bounding boxes po en ially enclosing egions o in-
e es . To do so, i exploi s a se o p ede ined boxes, called ancho s, as well
as he ela ionship be ween each pixel o he ea u e map and he cen e o i s
ecep i e ield in he inpu image. The chosen ancho s usually span mul iple
sizes and aspec a io, as o be able o co e he as majo i y o po en ial
objec s. Mo e speci ically, he RPN ope a es as ollows:
•each pixel o he ea u e map is assigned he same se o kancho s,
2.2. Fas e R-CNN 51
excep ha he ancho s a e shi ed a he cen e o he co esponding
pixel ecep i e ield in he image;
• he RPN maps he ea u e map, using a 3×3con olu ional laye ollowed
by wo pa allel 1×1con olu ional laye s, in o wo ou pu s: a classi ica ion
map and a eg ession map. Bo h maps ha e he same spa ial size as he
ea u e map and 2·kand 4·kchannels espec i ely.
• he classi ica ion map con ains he classi ica ion decision o each se
o ancho s a each loca ion, exp essed as a disc e e p obabili y ec o
p= (p0, p1), whe e p0 ep esen s he p obabili y ha he ancho con ains
backg ound and p1 he p obabili y ha i con ains an objec ; likewise, he
eg ession map con ains he shi and scale co ec ions o apply o each
ancho (ua, a, ha, wa)in o de o gene a e he co esponding p oposal
box (u, , h, w), encoded as ollows:
u=u−ua
wa
= − a
ha
w= log w
wa
h= log h
ha
(2.1)
In case a ea u e py amid is used ins ead o a single ea u e map, ancho s a
di e en scales a e assigned o di e en py amid le els: bigge ancho s a e as-
signed o low- esolu ion ea u e maps, as hese maps end o ocus mo e on
he global con ex o he inpu ; con e sely, smalle ancho s a e assigned o
high- esolu ion ea u e maps, as hese embed mo e local in o ma ion. Using
bigge ea u e maps o smalle ancho s also means ha mo e o hese ancho s
a e p esen , which imp o es ecall o smalle objec s. This a chi ec u al choice
ensu es ha objec s can be e ec i ely de ec ed e en i hey appea a consid-
e ably di e en scales in he inpu . No e ha he same RPN model is used o
p ocess all ea u e maps in he py amid.
De ec ion Head Gi en he gene a ed p oposals and he ea u e maps om
he backbone, he de ec ion head is asked o p edic he inal se o objec s.
In pa icula , his second s age ope a es he ollowing way:
58 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
(a) (b)
(c) (d)
(e)
Figu e 2.1: Visualiza ion o he in e media e s eps o he ne wo k. (a): Lo-
ca ions o he e e ence poin s on he image. (b) Region p oposals es ima ed
om he e e ence poin s. The g een colo indica es p oposals ha ing objec ness
sco e abo e 0.5. (c): P uning o he nega i e egion p oposals. (d) Remaining
posi i e egion p oposals a e NMS. (e): esul o he de ec ion head on he
emaining egion p oposals.
2.3. Pa king Slo De ec o 59
enough posi i e samples a e p esen , enough nega i e samples a e selec ed o
each 256 elemen s o al. Since no ancho s a e used a his s age, posi i i y is
de e mined simply by whe he each e e ence poin is con ained in a g ound
u h pa king slo o no : i i is, i is labelled as posi i e and associa ed o
ha slo ; o he wise, i is labelled as nega i e. The loss unc ion o aining
he RPN is gi en by:
LRP N =1
N·Lcls (pin, c∗) + λ
Npos
[c∗= 1] L eg ( , ∗),(2.7)
whe e Lcls is a s anda d bina y c oss-en opy loss:
Lcls(pin, c∗) = −c∗·log pin −(1 −c∗)·log (1 −pin)(2.8)
and L eg ollows Eqs. 2.3 and 2.4. pin and ep esen , espec i ely, he p e-
dic ed con idence p obabili es and e ices esiduals (pa ame ized as in Eq.
2.6), while ∗con ains he co esponding g ound u h esiduals, and c∗is se
equal o 1 o posi i e e e ence poin s and 0 o he wise. The balancing hype -
pa ame e λis empi ically se o 3.
The aining o he de ec ion head ollows Fas e -RCNN, wi h he only
di e ence ha , in his case, a se o 128 p oposals pe image a e used, wi h
a a io o 1:1 be ween posi i e and nega i e p oposals. In his s age posi i i y
is de e mined by measu ing he in e sec ion o e union be ween he minimum
enclosing ec angle o each p oposal and he minimum enclosing ec angle o
each g ound u h quad ila e al. In pa icula , o each p oposal, he g ound
u h box ha ing maximum IoU wi h i is de e mined. Then, i he IoU is
g ea e han 0.5, i is labelled as posi i e and associa ed wi h ha g ound
u h. O he wise i is ma ked as nega i e. The loss unc ion LDET o he
de ec ion head is equal o Eq. 2.2 wi h λ= 1, and he o al loss o he sys em
is gi en by he sum o he RPN loss unc ion and he de ec ion head loss
unc ion: L=LRP N +LDET .
The model is ained join ly using S ochas ic G adien Descen wi h mo-
men um 0.9, ba ch size 16 and lea ning a e 10−3 o 10.000 i e a ions. Du ing
aining, da a augmen a ion is employed in o de o en ich he da a samples
60 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
and educe o e i ing. In pa icula , each inpu image is subjec o bo h ho -
izon al and e ical lipping, applied independen ly each wi h p obabili y 0.5,
and is pe u bed in b igh ness, sa u a ion and con as in he ange ±20%, in
o de o simula e addi ional obse a ion condi ions.
2.3.2 Da ase Cons uc ion and Da a P epa a ion
In o de o ain he model, a small aining da ase composed o 467 su ound-
iew images was c ea ed and manually anno a ed wi h he image coo dina es
o he co ne s o each pa king slo , as well as he in o ma ion ega ding he
occupancy o he slo s. The e ices o each slo ha e been p ep ocessed in o de
o ensu e a consis en o de ing among di e en examples. Mo e speci ically, he
e ices a e so ed clockwise wi h espec o he pa king slo cen oid, iden i ied
as he mean o i s ou co ne s. Ensu ing he consis ency o he e ices o de
is c ucial o success ul op imiza ion, as he model mus be able o a ibu e a
seman ic meaning o each one. I g ound u h e ices we e no o de ed, e y
simila ins ances would co espond o di e en aining objec i es, which would
lead o ins abili y and inabili y o each con e gence. Examples o anno a ion
a e isible in Fig. 2.2.
(a) (b)
Figu e 2.2: Examples o anno a ed images. Colo s a e used o highligh bo h
e ex o de ing and occupancy.
2.4. Expe imen al Resul s 61
To ob ain he aining images, a ehicle se up wi h ou calib a ed and
synch onized ish-eye came as was used o collec se e al sequences, bo h o
oad scenes and pa king lo a eas. The images om each came a we e hen
cas o a Bi d’s Eye View ep esen a ion using he came as p ojec ion model
and s i ched oge he using he calib a ion o c ea e he su ound- iew images.
Meaning ul ames depic ing di e en scena ios, pa king slo ypes and obse -
a ion condi ions we e hen manually selec ed o anno a ion. To a oid biasing
he ne wo k owa ds si ua ions in which pa king slo s a e always p esen , 167
o he 467 chosen ames con ain no pa king slo s. The gene a ed images ha e
shape 1100 ×900, which is downsampled o 544 ×384 be o e being gi en as
inpu o he model. The eason o downsampling he inpu is wo old: on he
one hand, educing he spa ial size educes he compu a ional and memo y
cos s o he sys em conside ably. On he o he hand, he chosen esolu ion is
di isible by 32 on bo h dimensions, which simpli ies he compu a ion o he
e e ence poin s posi ions.
2.4 Expe imen al Resul s
To alida e he e ec i eness o he p esen ed app oach, he ained model was
es ed on new sequences acqui ed unde di e en obse a ion condi ions and
con aining pa king lo s unobse ed du ing aining. Some quali a i e esul s
a e isible in Fig. 2.3. As can be obse ed, despi e he ex emely limi ed ain-
ing se , he model exhibi s a ema kable capabili y o gene alize o new, unseen
scenes and i is able o co ec ly iden i y pa king slo s ha a e simila o hose
obse ed du ing aining. In pa icula , he ne wo k is cu en ly capable o
handling pa king slo s displaying di e en pa e ns (Fig. 2.3b) as well as o-
a ed slo s (Figs. 2.3c, 2.3d). Mo eo e , despi e he ac ha he aining se
con ains only as ew as 20 examples o slan ed pa king slo s, he ne wo k is
able o de ec hem wi h accep able accu acy, as shown in Fig. 2.3e. The sys-
em is also able o wi hs and noise in he da a: in Figs. 2.3c, 2.3d and 2.3e
some di is isible in he lenses. Also, he came as a e no pe ec ly calib a ed,
62 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
(a) (b)
(c) (d)
(e) ( )
Figu e 2.3: Resul s o he ne wo k on di e en ypes o pa king slo s unde
di e en obse a ion condi ions: (a) he mos common scena io. (b) Di e en
pa e n. (c-d) Di e en o a ions. (e) Slan ed pa king slo s. ( ) Failu e case.
2.4. Expe imen al Resul s 63
which leads o misalignmen s be ween he ou bi d’s eye iew images. Ne -
e heless, he ne wo k is able o p edic he obse able slo s well enough. The
sca si y o aining da a, howe e , migh lead o inco ec p edic ions, whe e
unobse ed ho izon al a ic signs and pa e ns a e e oneously in e p e ed as
pa king slo s, such as in Fig. 2.3 .
The p oposed sys em is able o un a o e 13 ames pe second ( ps) on
a NVIDIA Ge o ce GTX 1080 GPU.
2.4.1 Seman ic Shi P oblem
In he da a p epa a ion sec ion (2.3.2) I men ioned he impo ance o ensu ing
a consis en o de ing among he ou e ices o each g ound u h bounding
box, as he ne wo k mus be able o associa e a seman ic meaning o each poin
(e.g. he i s p edic ion is he op-le poin , he second is he op igh and
so on) in o de o be able o con e ge. In o de o gua an ee his consis ency,
he e ices we e so ed in a clockwise o de wi h espec o he cen oid o
each bounding box. The e a e, howe e , some speci ic obse a ion angles o
which he esul o he so ing ule changes ab up ly, causing wha I call a
seman ic shi (see Fig. 2.4a o a schema ic ep esen a ion o he p oblem).
As a consequence, he ne wo k exhibi s e a ic beha io when asked o pe o m
p edic ions o ins ances ha a e e y close o his c i ical angle (see Figs. 2.4b,
2.4c, 2.4d). This is due o he ac ha he model is unable o iden i y which
is he co ec o de o he poin s and, as a esul , ends o p edic coo dina es
ha a e in be ween he igh ones. No e ha his p oblem is no due o he
speci ic o de ing ule chosen, bu a he o any o de ing ule. Di e en o de ing
s a egies migh co espond o di e en con igu a ions a which he seman ic
shi occu s, bu he o e all p oblem emains. The hope is ha , by inc easing
he numbe o aining examples close o he seman ic shi , he ne wo k can
lea n o handle hese cases mo e e ec i ely and na ow down he ange o
angles o which he con usion happens. O cou se, he ideal way o handle
his p oblem would be o adop a ep esen a ion o he ou pu ha does
no depend on any speci ic o de ing, bypassing he seman ic shi al oge he .
64 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
1 2
34
1
2
3
4
2
3
1
4
(a)
(b) (c) (d)
Figu e 2.4: Seman ic shi p oblem. (a) Schema ic illus a ion o he p oblem.
(b) Ou pu be o e he shi angle. (c) Ou pu a ound he shi angle. (d) Ou pu
a e he shi angle.
2.4. Expe imen al Resul s 65
One such ep esen a ion migh ake ispi a ion om he e y ecen wo k Poly-
YOLO [94], in which he au ho s ex end he la es YOLO 3 model o also
p edic he seman ic mask o each de ec ed objec by in e p e ing each mask
as a se o e ices ha a e es ima ed using a pola g id.
2.4.2 Abla ion S udy
To alida e he choice o adop ing a wo-s age de ec ion app oach o added
localiza ion accu acy, I ca ied ou an abla ion s udy ha is made up o wo
dis inc expe imen s:
• i s , I emo ed he second s age, en us ing he RPN o pe o m de-
ec ion di ec ly. To achie e his, he classi ica ion b anch o he RPN
was modi ied so ha i p edic s he class o each e e ence poin (e.g.
backg ound, occupied, acan ) ins ead o jus de e mining whe he each
e e ence poin is con ained in a pa king slo . The cos unc ion o he
classi ica ion was changed acco dingly o a s anda d mul i-class c oss-
en opy o mula ion;
•second, he sampling p ocedu e o picking a balanced se o o eg ound
and backg ound examples a each i e a ion was emo ed. Ins ead, all
e e ence poin s a e used o compu e he loss unc ion LRP N .
To e alua e each expe imen , a small es se o 107 unobse ed pa king
spaces was anno a ed.
As e alua ion me ic I chose he A e age P ecision (AP), u ilized o e alu-
a e Objec De ec ion pe o mance in he PascalVOC benchma k [95]. To com-
pu e his me ic, he ou pu s o he sys em a e anked acco ding o hei con i-
dence and a e hen used o de e mine he p ecision/ ecall cu e. The AP alue
is hen ob ained by compu ing he mean p ecision alue a a se o 11 equally
spaced ecall in e als. To de e mine whe he i is a ue o alse posi i e, each
p edic ion is checked agains he g ound u h acco ding o a ule. Commonly,
in adi ional Objec De ec ion his ule consis s in he In e sec ion-o e -Union
66 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
be ween he de ec ion and he g ound u h boxes. Mo e speci ically, in o de
o be conside ed a co ec de ec ion, an ou pu o he model mus ha e an IoU
abo e a ce ain h eshold wi h a leas one g ound u h elemen o he same
class as he p edic ion. Duplica e de ec ions o he same g ound u h ins ance
a e handled as alse posi i es (e.g. i he same g ound u h objec is de ec ed
3 imes, one is conside ed as ue posi i e and he es a e alse posi i es).
Di e en ly om he aining p ocedu e, whe e he IoU compu a ion has been
app oxima ed using he minimum enclosing ec angle o e iciency, he e he
compu a ion is exac , as pe o mance is no a conce n.
The esul s o his s udy a e shown in Tab. 2.1. The displayed mAP (mean
AP) sco es a e ob ained by a e aging he AP alues o he acan and occu-
pied classes. The subsc ip indica es he IoU h eshold used o de e mine he
posi i i y o nega i i y o each de ec ion, while no subsc ip means ha he
co espoding alues a e ob ained by a e aging he mAP alues o e mul iple
h esholds, om 50% o 95% in in e als o 5%.
Poin sampling Second s age mAP mAP50 mAP70
X X 44.9 60.2 54.8
X20.1 36.5 25.4
14.1 29.6 17.4
Table 2.1: Resul s o e he es se . mAP50 and mAP70 ep esen he mean
A e age P ecision using IoU h eshold o 50% and 70%, espec i ely. mAP
ep esen s he mean A e age P ecision a e aged o e mul iple h esholds ( om
50% o 95%, in in e als o 5%).
As can be obse ed, he in oduc ion o he second s age o he pipeline
accoun s o mos o he pe o mance gain o he sys em. This di e ence in
pe o mance can be jus i ied by he ac ha , when he second s age is no
used, he p edic ions a e no longe based on ad-hoc se o ea u es ex ac ed
by RoIAlign, bu a he on gene ic pa ches o he ea u e map. These pa ches
2.4. Expe imen al Resul s 67
a e no as accu a ely localized as ea u e c ops, a e less in o ma i e and ha e
a smalle ecep i e ield, which hinde s he abili y o he model o p oduce
accu a e e ex p edic ions.
Remo ing he sampling s a egy o he e e ence poin s du ing aining
u he deg ades he quali y o he esul s. The eason o his is wo old: on
he one hand, using all e e ence poin s a each aining i e a ion in la es he
classi ica ion loss wi h backg ound samples, slowing down p og ess o posi i e
e e ence poin s; on he o he hand, he sampling p ocess p o ides s ochas ici y
du ing op imiza ion, meaning ha he same inpu p o ides di e en eedback
e e y ime i is p esen ed o he model. This has a egula izing e ec , a o -
ing gene aliza ion especially when he aining da a is limi ed. Mo eo e , he
imp o emen in pe o mance induced by he sampling s a egy as well as he
second s age is mo e p onounced a highe IoU h esholds (i.e. mAP70), which
is ep esen a i e o a be e localiza ion accu acy o he ull model. An exam-
ple o he di e en es - ime beha io o he model in he h ee cases can be
obse ed in Fig. 2.5.
(a) (b) (c)
Figu e 2.5: Resul s o he abla ion s udy. (a): Ne wo k wi hou he de ec ion
head and using all e e ence poin s. (b): Ne wo k wi hou he de ec ion head
using he usual sampling s a egy o he RPN. (c): Full pipeline
I is wo h no ing ha he blu ing a he edges o he images caused by
he bi d’s eye iew ans o ma ion, migh lead o de ec ion ins abili y in hose
egions, esul ing in mAP sco es o he ull model ha a e no ep esen a i e
o i s ue pe o mance. Mo eo e , some alse posi i es migh s ill be de ec ed
74 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
da a acqui ed by ish-eye lenses on a hemisphe e using a sphe ical p ojec ion
model. To ob ain he g ound u h pa king slo s in his new ep esen a ion,
I di ec ly eused he p e iously gene a ed anno a ions on he bi d’s eye iew
images, in e ing he ans o ma ion o ob ain hei co esponding coo dina es
in he sphe ical images. The gene a ed bi d’s eye iew images, howe e , depic
only a limi ed po ion o he obse ed space, as a away egions a e c opped
ou due o being oo noisy. As a esul , since he anno a ions a e eused om
he bi d’s eye iew case, he e exis pa king slo s ha a e clea ly isible in he
sphe ical images which a e no labelled; also, pa king slo s ha a e cu o in
he bi d’s eye iew bu a e ully isible in he sphe ical images would ha e hei
anno a ions also cu o (see Fig. 2.7 o an example). I such da a we e o be
used di ec ly o ain he sys em, i would lead o op imiza ion ins abili y and
subop imal con e gence, as he model would ecei e con addic o y aining
signal.
(a) (b)
Figu e 2.7: P ojec ions o he anno a ions gene a ed on su ound iew images
on o he co esponding sphe ical images. The yellow polygon is used o high-
ligh he po ion o image isible om i s bi d’s eye iew.
To o e come his limi a ion, each inpu image is appended a ou h channel
which is a bina y mask highligh ing he po ion o he image ha is isible
in i s co esponding bi d’s eye iew ep esen a ion. This solu ion allows he
2.4. Expe imen al Resul s 75
model o ha e knowledge abou egions ha a e alid o p edic ion while being
conside ably cheape o implemen han simply anno a ing all he missing slo s.
An al e na i e solu ion o he bina y mask could consis in simply ze oing
ou all pixels ha a e no isible in bi d’s eye iew. This app oach, howe e ,
is subop imal as i would dep i e he model om being awa e o con ex ual
in o ma ion ha migh p o e use ul o he inal ask.
O e all, he esul ing aining se consis s o 1868 images o al, as each
o iginal su ound- iew image is gene a ed om a o al o 4 sphe ical images.
Each sphe ical image has an o iginal esolu ion o 1024×992, which is c opped
o 1024 ×704 by emo ing he ows co esponding o he sky, be o e using i
as inpu . The adop ed model is iden ical o he one used on su ound iew,
wi h he excep ion ha he i s laye o he backbone is modi ied o accep
4-channel inpu s.
Quali a i e esul s on unobse ed pa kings scenes a e isible in Fig. 2.8. I
can be seen ha he model is capable o handling ela i ely well pa king slo s
obse ed om e y di e en poin o iews, and he e o e cha ac e ized by
e y di e en shapes, sizes and appea ances due o he pe spec i e p ojec ion.
Mo eo e , he ne wo k appea s o ha e p ope ly lea ned he seman ic meaning
o he inpu bina y mask, as i only e u ns de ec ions wi hin i s a ea.
No e ha he expe imen s on sphe ical images we e conduc ed exclusi ely
o e alua e he pe o mance and obus ness o he p oposed sys em o di e en
condi ions and poin s o obse a ions. Indeed, ope a ing on sphe ical images
(o e en pinhole ones) ins ead o su ound iew ep esen a ions is subop imal
o he ask o pa king slo de ec ion, o se e al easons. Pa king slo ma kings,
and oad ma kings as a whole, a e gene ally well-beha ed in bi d’s eye iew
as hey a e mos ly ee o pe spec i e e ec s and p ese e hei shape and size
independen ly om hei dis ance om he obse a ion poin . Also, he bi d’s
eye iew p ojec ion na u ally elimina es all in o ma ion ha lies abo e he
chosen plane, which is mos ly useless o he ask. Fo bo h o hese easons,
su ound iew images ep esen a much mo e sui able domain o he ask,
which leads o be e model op imiza ion and inc eased pe o mance. Ano he
76 Chap e 2. Objec De ec ion o Pa king Slo De ec ion
(a) (b)
(c) (d)
(e) ( )
Figu e 2.8: Resul s o he model on sphe ical images. I can be no ed ha he
ne wo k has lea ned o u ilize co ec ly he in o ma ion abou he ield o iew.
2.5. Discussion 77
non negligible ad an age o su ound iew is ha i allows o co e a ield o
iew o 360◦wi h a single inpu , while a leas ou inpu s would be equi ed
in cases sphe ical images a e used. This ansla es in o highe compu a ional
equi emen s, as e e y image would need o be p ocessed by he model o
ob ain a 360◦awa e de ec ion.
2.5 Discussion
In his chap e I p esen ed an end- o-end deep lea ning-based app oach o pa k-
ing slo de ec ion and occupancy classi ica ion on su ound- iew images. Mo e
speci ically, I buil upon he exis ing 2D objec de ec ion amewo k Fas e
R-CNN, edesigning i o allow o gene ic quad ila e al p edic ion ins ead o
axis-aligned bounding box es ima ion.
To ain and e alua e he sys em, wo small da ase s, con aining 467 and
107 su ound- iew images espec i ely, we e collec ed and manually anno a ed
wi h he loca ion and he occupancy o isible pa king spaces. The sys em dis-
played p omising esul s, exhibi ing a ema kable capabili y o adap o unseen
scena ios con aining pa king slo s o he same ype as hose obse ed in he
aining se while being obus o noise and misalignmen s be ween he s i ched
images ha cons i u e he su ound- iew ep esen a ions. I also showed he
abili y o unc ion p ope ly on an en i ely di e en and mo e di icul domain,
as p o en by i s e ec i eness when applied on he na i e sphe ical images.
Model simpli ica ion and spa si ica ion expe imen s highligh ed he ac ha
he p oposed model is capable o p ese ing compa able pe o mance by ac-
i ely using only 1/14 o i s o al numbe o ainable pa ame e s, which lea es
plen y o oom o e iciency imp o emen s.
Chap e 3
Monocula 3D Objec De ec ion
ia Gene alized
In e sec ion-o e -Union
Minimiza ion
3.1 P io A and Mo i a ion
A key challenge in ADAS and au onomous d i ing sys ems consis s o pe -
cei ing he su ounding en i onmen and loca ing obs acles in i , such ha
planning and con ol can be pe o med accu a ely. Pa icula ly c i ical is he
de ec ion o mo ing en i ies such as ca s, pedes ians and cyclis s, as hey pose
a majo challenge o sa e na iga ion.
Amongs he mos esea ched solu ions o 3D objec de ec ion a e he
LiDAR-based ones, as hese kinds o senso s a e capable o p o iding a e y
accu a e, albei spa se, econs uc ion o he su ounding en i onmen . These
senso s, howe e , a e gene ally qui e expensi e, o en making up a big ac ion
o he o al cos o he ehicle. As a esul , came as a e usually adop ed as a
cheape , mo e consume - iendly, al e na i e o pe cep ion. Came as ha e he
80
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
ad an age o pe cei ing iche in o ma ion compa ed o LiDAR, as i is dense
and con ains colo , and can be used o econs uc he geome y o he en i-
onmen by adop ing s e eo se ups in conjunc ion wi h dispa i y compu a ion
me hods [97, 98, 99], albei no as accu a ely as LiDARs, especially a long
anges. This in o ma ion can hen be used ei he implici ly [100, 76, 17] o
explici ely [77] o es ima e he 3D loca ions o he objec s o in e es .
Pe haps he mos in e es ing a ian o 3D objec de ec ion is, howe e ,
he monocula one. He e, he ask is o de e mine a 3D bounding box o e e y
objec o in e es using as inpu only a single image. This p oblem is e iden ly
unde cons ained, as a single image alone does no p o ide enough in o ma ion
o de e mine he o e all scale o he scene and, he e o e, he dis ance o he
objec s om he senso . To de e mine he dep h, addi ional in o ma ion abou
he scene is equi ed, such as he a p io i knowledge abou he dimension o
he obse ed objec s. Mos con empo a y s a e o he a monocula sys ems
le e age he a ailabili y o his knowledge, which o en comes in he o m o
g ound u h 3D bounding boxes, o ain deep models, wi h he objec i e o
implici ly encoding he ela ionship be ween objec appea ance on he image
plane and he co esponding posi ion in he wo ld wi hin i s pa ame e s, such
ha he model can be used o pe o m de ec ions in new scena ios. Ob iously,
o such a sys em o be able o unc ion accu a ely, i equi es a conside able
amoun o aining da a, and he new en i onmen s ha i is exposed o mus
belong o a domain ha is simila o he one i is ained on. Fo ins ance,
i he came a in insic pa ame e s change om he aining se , he sys em is
unlikely o p oduce accu a e localiza ions, as he lea ned unde lying mapping
be ween appea ance and posi ion is no longe alid o he new came a model.
Simila p oblems a ise i he came a is posi ioned di e en ly o i he po ayed
objec s a e isually e y di e en .
Due o he di icul y o his ask, mos cu en 3D objec de ec ion ame-
wo ks op o dedica ed models ha in eg a e 3D easoning mechanisms di-
ec ly in o hei a chi ec u es [15, 78, 79, 81, 16], in he hope ha he esul ing
sys em lea ns ea u es ha a e mo e sui able o 3D asks and gene alize be -
3.1. P io A and Mo i a ion 81
e o new si ua ions. Please e e o sec ion 1.6.2 o a mo e in-dep h e iew
o such app oaches. Con e sely, in he p oposed me hod, I a gue ha explici
3D easoning di ec ly encoded in o he ne wo k s uc u e is no manda o y o
good monocula 3D de ec ion pe o mance, as long as he unde lying model
has su icien capaci y. To his end, I p opose an ex en ion o he 2D de ec o
Fas e R-CNN in which a small subne wo k is added o he de ec ion head o
pe o m 3D box es ima ion. This subne wo k is simple, does no con ain any
kind o explici 3D easoning in i s s uc u e and is ained join ly wi h he
es o he model. See Fig. 3.1 o a schema ic ep esen a ion o he p oposed
sys em.
Fas e R-CNN 3D Head
Figu e 3.1: O e iew o he p oposed 3D de ec ion pipeline: I ex end
Fas e R-CNN wi h an addi ional module esponsible o es ima ing 3D bound-
ing boxes gi en he 2D de ec ions. This ex a module is ained end- o-end wi h
he es o he ne wo k using a no el loss based on he Gene alized In e sec ion-
o e -Union.
Commonly, 3D es ima o s a e ained ia a loss unc ion ha minimizes
he e o be ween he p edic ed box pa ame e s (i.e. cen e , dimensions, o i-
en a ion) and hei co esponding g ound u hs di ec ly. Ins ead, I p opose a
no el objec i e unc ion ha allows o eason in e ms o boxes as a whole ia
he minimiza ion o an app oxima ion o hei Gene alized In e sec ion-o e -
82
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
Union [20]. The expe imen s show ha his o mula ion leads o conside ably
be e esul s, likely due o be e ea u e ep esen a ions induced by a mo e
sui able choice o he loss unc ion.
3.2 Baseline Model
As al eady s a ed in Sec. 3.1, he p oposed me hod consis s o an ex en ion
o he anilla Fas e R-CNN model o 2D de ec ion ( e e o Sec. 2.2 o an
illus a ion o his model).
Mo e speci ically, as backbone o choice I adop he s anda d FPN buil
upon an ImageNe -p e ained ResNe -50 model, and I ex end i , ollowing [63],
wi h an addi ional downsampling s age consis ing o a 3×3con olu ion ha ing
s ide 2 applied o he las ea u e map e u ned by ResNe . Fo mally, he
esul ing ea u e ex ac o gene a es a ea u e py amid comp ised o ea u e
maps a 5 di e en esolu ions. These maps a e usually labelled as P2 o P6,
whe e Pliden i ies he ea u e map ha ing esolu ion 1/2lo ha o he inpu .
The Region P oposal Ne wo k ollows he o iginal implemen a ion. To han-
dle objec s ha ing di e en sizes, he RPN comp ises 5 di e en ancho scales
ha ing a eas 162,322,642,1282,2562which a e assigned o he le els P2 o
P6o he ea u e py amid. Each scale is made up o h ee di e en ancho s
ha ing aspec a ios {0.5,1,2}, o a o al o 15 ancho s o e he en i e py a-
mid.
Likewise, he de ec ion head ollows s anda d p ocedu e. As pooling me hod
I adop RoIAlign, which ex ac s a 7×7 ixed-size ea u e map om he py a-
mid o each o he op 300 sco ing p oposals (pos NMS) e u ned by he
RPN. These ea u es a e hen p opaga ed o he de ec ion head o objec
classi ica ion and 2D box e inemen . Again, in o de o handle objec s ha -
ing di e en sizes, each p oposal is assigned o he p ope le el o he ea u e
py amid be o e pe o ming RoIAlign, acco ding o he ollowing ule:
l=$l0+ log2 √w·h
224 !%.(3.1)
3.3. 3D De ec ion Module 83
He e, wand h ep esen he wid h and heigh o he egion p oposal and
l0 ep esen s he le el in he ea u e py amid ha a p oposal ha ing a ea
w·h= 2242should be mapped o. Following he o iginal implemen a ion o
FPN, I se l0= 4.
3.3 3D De ec ion Module
To allow he sys em o pe o m de ec ion o 3D bounding boxes, I p opose
o ex end Fas e R-CNN wi h an addi ional 3D module. In pa icula , gi en
he inal 2D de ec ions p oduced by he de ec ion head, a second RoIAlign
s ep is pe o med o ex ac hei speci ic se s o ea u es, ollowing he same
assignmen ule illus a ed in Eq. 3.1. Then, gi en each 2D de ec ion and i s
co esponding se o ea u es, he 3D module is esponsible o es ima ing he
3D bounding box B= (x, y, z, h, w, l, θ)co esponding o ha objec , whe e
(x, y, z)a e i s cen e coo dina es wi h espec o he came a ame o e e ence,
(h, w, l)a e i s heigh , wid h and leng h espec i ely and θis i s o ien a ion,
exp essed as a o a ion angle a ound he came a y-axis. See Fig. 3.2 o a bi d’s
eye iew illus a ion o he a ge s o be es ima ed.
In o de o simpli y and s abilize aining, hese alues a e no es ima ed
di ec ly by he 3D module, bu a e a he encoded as ollows.
Objec Dimensions To es ima e he objec dimensions, he de ec ion head
ou pu s he ollowing quan i ies:
log h
¯
h,log w
¯w,log l
¯
l,(3.2)
whe e ¯
h,¯w,¯
l ep esen class-speci ic p io alues ob ained by a e aging he
dimensions o each g ound u h objec ac oss he en i e aining se . This
o mula ion allows o ame he es ima ion o he dimensions in e ms o a
ela i e co ec ion, whe e nega i e alues co espond o a educ ion in size
wi h espec o he p io and posi i e alues o an inc ease in size.
90
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
he 3D module ope a es on he same se o posi i e p oposals also used by he
de ec ion head.
Gene ally, he commonly adop ed app oach o aining 3D box p edic-
o s consis s o minimizing di ec ly he adop ed pa ame iza ion (in his case
Eqs. 3.2, 3.3, 3.4, 3.7) agains he g ound u h using some so o dis ance
unc ion (e.g. Eq. 2.4). Ins ead, I p opose o ex end he GIoU o mula ion
o he 3D case. As al eady s a ed in Sec. 3.4.1, he IoU be ween wo axis-
aligned bounding boxes has closed o m solu ion, and his is due o he ac
ha hei in e sec ion is s ill an axis-aligned box. Mo eo e , by app oxima ing
he minimum enclosing box as ano he axis-aligned box, he GIoU can also be
calcula ed analy ically and has a well-beha ed g adien . 3D boxes, howe e ,
a e subjec o o a ions and hei in e sec ion is a cuboid only i hey ha e
a ela i e o ien a ion which is a mul iple o π/2. In all o he cases, hei in-
e sec ion would be a gene ic, i egula , polihed on. To ensu e he exis ance
o an analy ical solu ion, I disen angle he es ima ion o he angle om he
es o he dimensions, which a e op imized ia he GIoU by conside ing hei
espec i e boxes a a canonical o ien a ion. The loss unc ion o his module
is he e o e comp ised o wo dis inc componen s:
L3D=Lang +L3IoU ,(3.11)
whe e Lang is esponsible o op imizing he o ien a ion and L3IoU is asked o
op imize posi ion and dimen ions h ough he GIoU.
Fo mally, le B= (x, y, z, h, w, l, θ)be he 3D box p edic ed by he mod-
ule and ˆ
B=ˆx, ˆy, ˆz, ˆ
h, ˆw, ˆ
l, ˆ
θi s assigned g ound u h box. Le ˆα=ˆ
θ−
a an2 (−ˆx, ˆz)be he g ound u h obse a ion angle. Lang is de ined as he
smoo h-L1 loss be ween he es ima ed and he a ge obse a ion angles:
Lang =smoo hL1(sin ˆα−sin α) + smoo hL1(cos ˆα−cos α).(3.12)
In o de o op imize he posi ion and dimensions o he boxes h ough he
GIoU, I i s p e o a e hem such ha hei o ien a ion angle is equal o 0,
yielding:
3.4. Model Op imiza ion 91
B0= (x, y, z, h, w, l, 0),(3.13)
ˆ
B0= (ˆx, ˆy, ˆz, ˆ
h, ˆw, ˆ
l, 0).(3.14)
Unde his assump ion, he boxes can be di ec ly de ined in e ms o hei
opposing co ne s:
x1,2=x±l/2y1,2=y±h/2z1,2=z±w/2,(3.15)
ˆx1,2= ˆx±ˆ
l/2 ˆy1,2= ˆy±ˆ
h/2 ˆz1,2= ˆz±ˆw/2.(3.16)
Gi en hese alues, and by app oxima ing he minimum enclosing box as an-
o he cuboid ha ing θ= 0, compu ing he in e sec ion a ea I, he union a ea
Uand he minimum enclosing a ea Acis a i ial ex ension o he 2D case (see
Alg. ?? o he comple e o mula ion). Finally, gi en he GIoU alue, he loss
unc ion is ob ained using Eq. 3.10.
Mo e speci ically, ins ead o minimizing he GIoU loss unc ion be ween
B0and ˆ
B0di ec ly, inspi ed by he disen angling ans o ma ion in oduced
in [102] I spli he op imiza ion in o six sepa a e con ibu ions, each esponsible
o a single deg ee o eedom:
L3IoU =1
6X
i∈{x,y,z,h,w,l}1−GIoU ˆ
B0,Bi
0.(3.17)
He e, Bi
0is used o ep esen he box ob ained om B0by eplacing all alues
excep o iwi h he g ound u h (e.g. Bz
0= (ˆx, ˆy, z, ˆ
h, ˆw, ˆ
l, 0))). This o -
mula ion leads o a conside able op imiza ion speedup, especially ea ly in he
aining whe e mos p edic ions a e disjoin om hei co esponding g ound
u h boxes. Fu he analysis will be p esen ed in Sec. 3.5.3.
T aining De ails The model is ained end- o-end on ull esolu ion images
o 90k i e a ions, using S ochas ic G adien Descen wi h ba ch size 4, weigh
92
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
decay 5e-4 and momen um 0.9. The lea ning a e is ini ially se o 10−2and
is educed by a ac o o 10 e e y 30k i e a ions. The ResNe backbone is
ini ialized wi h ImageNe p e aining alues and i s Ba ch No maliza ion laye s
as well as i s i s wo con olu ional blocks a e kep ixed du ing aining. To
en ich he aining da a, each sample is independen ly augmen ed by andom
ho izon al lipping wi h p obabili y 0.5 as well as by ji e ing i s sa u a ion,
b igh ness and con as by ±30%.
3.5 Expe imen al Resul s
In his sec ion I in oduce KITTI [22], he au onomous d i ing da ase used
o ain and e alua e he p oposed app oach. Then, I pe o m a quan i a i e
compa ison agains cu en s a e o he a monocula 3D objec de ec o s.
Finally, I analyze he loss unc ion used o op imizing he 3D module and
compa e i s e ec i eness agains o he al e na i es.
3.5.1 The KITTI Da ase
The p oposed sys em is ained and e alua ed on he KITTI [22] da ase , which
cu en ly cons i u es he de ac o choice o he au onomous d i ing communi y
o esea ch.
This da ase p o ides bo h image, LiDAR and odome y da a, as well as
g ound u h anno a ions o a wide a ie y o asks including seman ic seg-
men a ion, ins ance segmen a ion, isual odome y/SLAM, 2D objec de ec-
ion and acking, 3D objec de ec ion, dep h es ima ion and op ical low.
Image da a is acqui ed using wo s e eo came a se ups, one o g eyscale and
one o colo , bo h displaced a he on o he ehicle and p o iding images
a a esolu ion o 1382 ×512. Due o ec i ica ion, he images p o ided o
aining a e smalle and ha e an app oxima e esolu ion o 1240×375. LiDAR
scans a e eco ded using a 64-planes Velodyne spinning a 10 ames pe sec-
ond and cap u ing app oxima ely 100k poin s pe e olu ion. The came as a e
3.5. Expe imen al Resul s 93
synch onized wi h he Velodyne and cap u e images a he beginning o each
e olu ion, also a 10Hz.
The 2D/3D Objec De ec ion da ase is ga he ed by anno a ing dissimi-
la ames om se e al eco ded sequences wi h he co esponding obse able
2D and 3D bounding boxes, o a o al o 7481 aining samples and 7518
es samples. The es se anno a ions a e no publicly a ailable and a e used
exclusi ely by he online e alua ion se e o pe o mance e alua ion. Mo e
speci ically, KITTI p o ides box anno a ions o 7 di e en classes, ha is Ca ,
Pedes ian, Cyclis , Van, T uck, Si ing Pe son and T am, bu only he i s
h ee a e conside ed o e alua ion by he o icial benchma k, as he o he s
a e oo sca ce in numbe o p ope model aining. S ill, like many cu en
me hods I only conside he Ca class o p edic ion, as i is conside ably mo e
equen and e enly dis ibu ed wi hin he da ase compa ed o Pedes ians
and Cyclis s. Also, each anno a ed box is a ibu ed one o h ee ca ego ies,
easy,mode a e o ha d, depending on i s size on he image plane and on how
much i is occluded and unca ed.
Following p e ious wo k [67, 17, 76], I spli he a ailable 7481 anno a ed
images in o a aining and a alida ion se , comp ised o 3712 ad 3769 samples
espec i ely. I is impo an o no e ha , in o de o ensu e p ope pe o mance
e alua ion, hese wo spli s a e o igina ed om wo disjoin se s o sequences,
such ha no simila scenes a e sha ed be ween aining and alida ion.
3.5.2 Compa ison wi h he S a e o he A
I e alua e he 3D localiza ion and de ec ion pe o mance o he sys em using
he KITTI A e age P ecision me ic o bi d’s eye iew (APBEV) and 3D de-
ec ion (AP3D). Fo a exhaus i e compa ison, I conside bo h he o icial 0.7
IoU h eshold and he mo e pe missi e 0.5 IoU h eshold. The esul s o he
wo asks a e shown in Tab. 3.1 and Tab. 3.2 espec i ely.
The p oposed me hod exhibi s s a e-o - he-a pe o mance on he alida-
ion se , su passing all o he monocula app oaches by a good ma gin on he
o icial 0.7 IoU h eshold, while also being compe i i e wi h he s e eo-based
94
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
(a) (b)
(c) (d)
(e) ( )
Figu e 3.4: Resul s o he p oposed me hod on he alida ion se . Red bounding
boxes co espond o he g ound u h, while g een boxes ep esen he de ec-
ions. The LiDAR poin clouds a e used exclusi ely o isualiza ion. Bes
iewed in colo .
3.5. Expe imen al Resul s 95
(a) (b)
(c) (d)
(e) ( )
Figu e 3.5: (con .) Resul s o ou me hod on he alida ion se . The ed bound-
ing boxes co espond o he g ound u h, while he g een boxes a e ou de-
ec ions. The LiDAR poin clouds a e used exclusi ely o isualiza ion. Bes
iewed in colo .
96
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
Me hod Da a APBEV @ 0.5 IoU APBEV @ 0.7 IoU
Easy Mode a e Ha d Easy Mode a e Ha d
TLNe (S) [76] S e eo 62.46 45.99 41.92 29.22 21.88 18.83
Mono3D [15] Mono 30.50 22.39 19.16 5.22 5.19 4.13
OFNe [78] Mono - - - 11.06 8.79 8.91
Mul iFusion [80] Mono 55.02 36.73 31.27 22.03 13.63 11.60
TLNe (M) [76] Mono 52.72 37.22 32.16 21.91 15.72 14.32
MonoGRNe [79] Mono 54.21 39.69 33.06 24.97 19.44 16.30
MonoPSR [16] Mono 56.97 43.39 36.00 20.63 18.67 14.45
MonoDIS [102] Mono - - - 24.26 18.43 16.95
Ou s Mono 60.17 43.45 36.53 29.70 21.86 18.13
Table 3.1: Resul s o localiza ion on he KITTI alida ion se o he Ca class.
3.5. Expe imen al Resul s 97
Me hod Da a AP3D @ 0.5 IoU AP3D @ 0.7 IoU
Easy Mode a e Ha d Easy Mode a e Ha d
TLNe (S) [76] S e eo 59.51 43.71 37.99 18.15 14.26 13.72
Mono3D [15] Mono 25.19 18.20 15.22 2.53 2.31 2.31
OFNe [78] Mono - - - 4.07 3.27 3.29
Mul iFusion [80] Mono 47.88 29.48 26.44 10.53 5.69 5.39
TLNe (M) [76] Mono 48.34 33.98 28.67 13.77 9.72 9.29
MonoGRNe [79] Mono 50.51 36.97 30.82 13.88 10.19 7.62
MonoPSR [16] Mono 49.65 41.71 29.95 12.75 11.48 8.59
MonoDIS [102] Mono - - - 18.05 14.98 13.42
Ou s Mono 55.70 40.34 34.40 22.48 16.67 15.08
Table 3.2: Resul s o de ec ion on he KITTI alida ion se o he Ca class.
98
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
app oach TLNe [76]. Likewise, a he 0.5 IoU h eshold, he p esen ed ap-
p oach ou pe o ms all o he me hods on bo h asks excep o he mode a e
samples on he de ec ion ask, whe e he esul s a e sligh ly lowe han hose
o MonoPSR [16]. The ela i ely la ge imp o emen in pe o mance o he
0.7 IoU case compa ed o he 0.5 IoU one sugges s ha he p oposed app oach
e u ns, on a e age, de ec ions ha ing highe localiza ion accu acy.
On a NVIDIA Tesla V100 GPU, he in e ence ime o he 3D de ec ion
module is app oxima ely 5 ms, wi h sligh a ia ions depending on he numbe
o 2D de ec ions ha mus be p ocessed by he 3D module. The o al in e ence
ime o he en i e pipeline is abou 50ms on ull esolu ion KITTI images.
3.5.3 Compa ison wi h o he Loss Fo mula ions
To alida e he choice o u ilizing an app oxima ion o he GIoU loss o mula-
ion, I conduc ed se e al expe imen s in which di e en loss o mula ions a e
used.
Is lea ning each dimension disjoin ly impo an ? As i s expe imen ,
I in es iga ed whe he u ilizing a sepa a e loss componen o each deg ee o
eedom o posi ion and dimensions is bene icial o he pe o mance o he
esul ing sys em. In pa icula , I ca ied ou an expe imen whe e all six pa-
ame e s a e op imized oge he in a single loss unc ion. Wha I no iced was a
endency o he sys em o educe he GIoU loss alue in case o non-o e lapping
boxes by inc easing hei size a he han by ying o ma ch hei posi ions in
space. This beha io led o a signi ican numbe o spu ious de ec ions ea ly
in he aining, which in u n caused a pla eau o he loss unc ion a ound he
alue o 1, ul ima ely slowed down he lea ning p ocess and leading o wo se
accu acy. A isualiza ion o his phenomenon on alida ion images is isible
in Fig. 3.6: as i can be seen, anomalous de ec ions esul ing om his beha -
io a e e y p ominen a he ea ly s ages o op imiza ion, and s ill pe sis ,
al hough less ex emely, once he sys em is ully ained.
3.5. Expe imen al Resul s 99
(a) (b)
Figu e 3.6: Anomalous beha io o he 3D de ec ion module when ained wi h
he join GIoU. Fig. 3.6a shows a sample esul ea ly in he aining, while
3.6b shows he same sample a he end o op imiza ion esul s o he ained
sys em.
I specula e ha his phenomenon is induced by how he Gene alized In-
e sec ion o e -Union is o mula ed: wo e y simila boxes, close oge he bu
disjoin , always leads o a highe loss alue han wo e y di e en boxes such
ha one is comple ely inside he o he . The eason o his beha io is illus-
a ed in Fig. 3.7. This also holds ue o e y di e en boxes ha sha e a
low enough o e lap such ha E > IoU (see Eq. 3.8). As a esul , he endency
o he sys em, especially du ing he i s hal o he aining p ocedu e whe e
mos p edic ed 3D boxes end no o o e lap hei associa ed g ound u h, is
o educe he loss alue by inc easing he size, such ha Eis d i en o ze o.
This beha io was p obably no no iced when he GIoU was applied o 2D de-
ec ion, as mos 2D de ec o s op imize he e inemen b anch o he de ec ion
head only on ancho s o p oposals ha sha e a signi ican o e lap wi h he
g ound u h boxes o begin wi h. As a esul , since he inal p edic ions a e
ob ained as a co ec ion o hese p io boxes, i is unlikely ha a la ge po ion
o hem a e disjoin om hei co esponding g ound u h. Con e sely, in he
106
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
3.6.1 3D de ec o : s anda d op imiza ion
All h ee modules in ol ed in he es ima ion o he 3D box om he sampled
poin cloud a e op imized join ly, using a mul i- ask loss unc ion:
L=Lseg +λ(Lc1- eg +Lc2- eg +Lh-cls +Lh- eg +Ls-cls +Ls- eg +γLco ne ).
(3.19)
The segmen a ion submodule e u ns a scala alue o each inpu poin ,
encoding he p obabili y ha said poin belongs o he objec o in e es . As
such, he loss Lseg used o ain his submodule akes he o m o a s anda d
bina y c oss-en opy loss (see Eq. 2.8), a e aged o e all inpu poin s.
The T-Ne o cen e es ima ion is ained o di ec ly es ima e he posi-
ion o he objec cen e gi en he se o o eg ound poin s shi ed a ound hei
cen oid. I is ained ia he loss Lc1- eg, which akes he o m o a smoo h-L1
dis ance (Eq. 2.4) be ween he es ima ed and he g ound u h cen e .
The second Poin Ne o 3D box es ima ion is asked o e u n he 3D
box gi en he se o o eg ound poin s shi ed a he es ima ed cen e coo di-
na es. Mo e speci ically, his submodule e u ns a o al o 3 + 4 ×NS + 2 ×NH
ou pu s, whe e:
• he i s 3 ou pu s ep esen s he es ima ed posi ion o he objec cen e
wi h espec o he inpu o eg ound poin cloud. These a e ained ia
he loss Lc2- eg which is a smoo h-L1 dis ance be ween he p edic ed and
he g ound u h cen e ;
• he subsequen 4×NS ou pu s encode he dimensions es ima ion, which
is pe o med ela i e o a se o NS p ede e mined dimension empla es.
Mo e speci ically, he i s NS ou pu s ep esen a disc e e p obabili y
dis ibu ion ha de e mines which empla e is used, and he ollowing
3×NS alues encode heigh , wid h and leng h co ec ions wi h espec o
each empla e. The classi ica ion is ained ia Ls-cls, which is a s anda d
mul i-class c oss-en opy loss (see Sec. 2.2.1). The co ec ion is ained
3.6. A case s udy:
3D GIoU applied o F us um-Poin Ne s 107
ia Ls- eg, which is a smoo h-L1 dis ance be ween he es ima ed co ec-
ion and he g ound u h one: no e ha in his case he loss unc ion is
only compu ed o he igh empla e, de e mined by g ound u h;
• he las 2×NH ou pu s encode he o ien a ion es ima ion which, like o
he dimensions, is spli in o a classi ica ion componen and a co ec ion
componen . In pa icula , he 360◦a e di ided in o NH bins: he i s NH
ou pu s ep esen a disc e e p obabili y dis ibu ion ha he o ien a ion
angle lies in ha bin and he emaining NH alues encode co ec ions
wi h espec o each bin cen e . T aining ollows he one used o he
dimensions: Lh-cls is a mul i-class c oss-en opy loss o bin classi ica ion
and Lh- eg is a smoo h-L1 dis ance compu ed on he co ec ed bin, which
is de e mined by g ound u h.
Addi ionally, o egula ize he es ima ion o he 3D box and imp o e he ac-
cu acy, an addi ional loss e m, he co ne loss Lco ne is used, which consis s
o minimizing he smoo h-L1 dis ance be ween he es ima ed and he g ound
u h box co ne s. In o de o compu e his loss e m, he 3D box o igina ed
using he es ima ed co ec ions on he g ound u h bins is conside ed.
3.6.2 3D de ec o : p oposed op imiza ion
I p opose o modi y he mul i- ask loss (Eq. 3.19) such ha he es ima ion
o he 3D box cen e and dimensions is op imized h ough he p oposed dis-
join 3D GIoU loss. Mo e speci ically, I emo e he losses p e iously asked
o es ima e cen e (Lc2- eg) and dimensions (Ls- eg) as well as he co ne loss
(Lco ne ), and I sub i u e hem wi h he o mula ion in oduced in Eq. 3.17:
L=Lseg +λ(Lc1- eg +Lh-cls +Lh- eg +Ls-cls +γL3IoU)(3.20)
As o iginally done o he co ne loss compu a ion, in o de o compu e he
GIoU loss alue I conside he box ob ained om he p edic ed co ec ion
on he g ound u h dimension empla e. No o he modi ica ion is done o
108
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
APBEV / AP3D @ 0.7 IoU
Ca Easy Mode a e Ha d
O iginal 1 87.82 / 83.26 82.44 / 69.28 74.77 / 62.56
Baseline 1 87.64 / 82.97 83.09 / 71.11 75.75 / 63.43
GIoU 1 87.59 / 85.36 83.89 /73.47 76.16 /65.29
O iginal 2 88.16 / 83.76 84.02 / 70.92 76.44 / 63.65
Baseline 2 88.19 / 83.33 85.38 / 71.74 77.15 / 64.01
GIoU 2 88.41 /86.08 85.99 /74.47 77.64 /66.14
Table 3.4: Resul s o he s udy on applying he GIoU as cos unc ion o op i-
mizing F us um-Poin Ne s. The Ca class is conside ed o e alua ion. O igi-
nal and O iginal 2 indica e he esul s epo ed by he o iginal pape [23],
using he Poin Ne and Poin Ne ++ based a chi ec u es espec i ely. Base-
line and Baseline 2 ep esen he esul s ob ained by my eimplemen a ion
o he sys em. GIoU indica es ha he eimplemen ed sys em makes use o
he GIoU loss o mula ion in Eq. 3.17 o op imize he es ima ion o cen e and
dimensions.
he sys em: he 3D box o ien a ion and dimensions a e s ill es ima ed using a
hyb id classi ica ion-co ec ion app oach, and he model a chi ec u e is exac ly
he same. Also, I adop he exac same aining ou ine as he o iginal wo k.
3.6.3 Expe imen al Resul s
The esul s o he expe imen s o he Ca class a e displayed in Table 3.4.
Mo e speci ically, I epo bo h he o iginal esul s shown in he pape [23], as
well as hose ob ained by my eimplemen a ion o he sys em in PyTo ch. I
expe imen ed wo h bo h e sions o F us um-Poin Ne s:
• e sion 1 adop s simple Poin Ne ne wo ks o all h ee modules;
• e sion 2 u ilizes Poin Ne ++ models bo h o he segmen a ion module
3.6. A case s udy:
3D GIoU applied o F us um-Poin Ne s 109
and he box es ima ion module, esul ing in a mo e powe ul and con ex -
awa e, albei qui e slowe , ne wo k.
In o de o a oid ambigui ies and isola e he analysis o he di e en ain-
ing s a egy used o op imize he 3D de ec ion ne wo k, a e alua ion ime I
adop he exac same se o 2D de ec ions used o e alua e he o iginal sys em,
which a e made publicly a ailable by he au ho s: his way, no pe o mance
change can be a ibu ed o a di e ence in beha io o he unde lying 2D objec
de ec o .
The models ained wi h he GIoU loss o mula ion s ongly ou pe o m
hei co esponding o iginal e sions, especially o he 3D ask which is he
one ha he loss explici ly aims o op imize. Pa icula ly no able is he ac
ha he 1 model, when ained using he p oposed loss, dis inc ly ou pe -
o ms he 2 e sion op imized wi h he o iginal loss unc ion, despi e being a
conside ably simple model. This is indica i e o he ac ha an app op ia e
op imiza ion s a egy migh be mo e impo an o he inal esul han he
model a chi ec u e.
To u he abla e he p oposed o mula ion, in Table 3.5 I also show he
esul s o he ained models o he Cyclis class. E en in his case, he 1
sys em ained wi h he GIoU loss as ly ou pe o m i s 2 coun e pa wi h
he s anda d aining p ocedu e. Mos su p ising, howe e , is he ac ha he
GIoU-op imized 2 model, while s ill ou pe o ming he o iginal model, pe -
o ms wo se han he GIoU-op imized 1 on he 3D me ic. This beha io migh
be he esul o o e i ing: as al eady s a ed in Sec. 3.5.1, a ailable Cyclis da a
is e y limi ed compa ed o Ca da a. As a esul , he isk o o e i ing such
da a is highe , especially when adop ing mo e complex models like he e sion
2 o F us um-Poin Ne s. Con e sely, he much simple a chi ec u e o e sion 1
leads o simple ea u es being lea ned, which in u n imp o es gene aliza ion
on unseen samples. On he BEV localiza ion ask, on he o he hand, e sion
2 i mly ou pe o ms e sion 1: his is possibly due o he ac ha he ask
is simple , and he e o e he noisie p edic ions induced by o e i ing do no
a ec pe o mance as much.
110
Chap e 3. Monocula 3D Objec De ec ion ia Gene alized
In e sec ion-o e -Union Minimiza ion
APBEV / AP3D @ 0.7 IoU
Cyclis Easy Mode a e Ha d
O iginal 2 81.82 / 77.15 60.03 / 56.49 56.32 / 53.37
GIoU 1 83.44 / 81.04 62.09 / 59.88 58.60 / 55.76
GIoU 2 84.77 / 78.53 63.75 / 57.42 59.44 / 53.96
Table 3.5: Resul s o he p oposed sys em on he Cyclis class, compa ed
agains he o iginal esul s.
3.7 Discussion
In his chap e , I p oposed an ex en ion o he 2D objec de ec o Fas e R-
CNN consis ing o a simple module esponsible o monocula 3D de ec ion
which is ained using a disjoin o mula ion o he Gene alized In e sec ion-
o e -Union (GIoU) loss unc ion. To ensu e he exis ence o an analy ical solu-
ion, I disen angled he es ima ion o he o ien a ion om ha o posi ion and
dimensions, o a ing he boxes o a canonical angle be o e compu ing hei
GIoU. Mo eo e , o a oid an anomalous beha io o he GIoU ea ly du ing
aining, I op ed o op imizing each deg ee o eedom sepa a ely by adop ing
a dedica ed loss unc ion o each.
The esul ing sys em exhibi ed ema kable pe o mance, su passing mo e
complex and model-d i en pipelines on he au onomous d i ing KITTI da ase .
The app oach is also simple, as he 3D de ec ion module only consis s in a
hand ul o ully-connec ed laye s, and hus could be inco po a ed s aigh o -
wa dly in o o he exis ing 2D de ec ion me hods.
Mos o he pe o mance gain achie ed by he p oposed sys em is o be
a ibu ed o he way ha he 3D de ec ion module is op imized, as shown by
he s udy conduc ed by u ilizing mo e adi ional op imiza ion unc ions. To
u he alida e his claim, I adop ed he GIoU loss unc ion o op imizing he
LiDAR-based 3D de ec o F us um-Poin Ne s, and showed ha he p oposed
o mula ion can lead o signi ican imp o emen s e en o en i ely di e en
3.7. Discussion 111
neu al a chi ec u es.
Chap e 4
3D Objec De ec ion on LiDAR
scans ia Vo ing and
Sel -A en ion Mechanisms
4.1 P io A and Mo i a ion
Came a-based pe cep ion solu ions, whils being cu en ly conside ably cheape
han LiDAR-based ones, su e om in e io pe o mance, mos ly due o he
ac ha images do no encode dep h in o ma ion di ec ly. S e eo se ups can
be used o es ima e dep h h ough dispa i y calcula ion; such dep h, howe e ,
ends o be conside ably noisie , especially a high dis ances, and ends o ail
in p esence o speci ic pa e ns o lack o ex u es. Simila ly, ne wo ks o 3D
de ec ion ained on monocula images equi e an eno mous amoun o da a
in o de o gene alize p ope ly o unobse ed scenes and, e en hen, hey s ill
s ongly unde pe o m s e eo and LiDAR-based pipelines.
LiDAR senso s, on he o he hand, p o ide ex emely accu a e, albei
spa se , dep h in o ma ion in he o m o a poin cloud o he su ounding
en i onmen . While cu en ly being conside ably mo e expensi e han cam-
e as, ad ancemen s in senso echnology a e consis en ly educing he cos s o
114
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
LiDAR solu ions, and p ices a e quickly app oaching hose o s e eo-came a
se ups. This makes LiDAR based pe cep ion echniques wo h in es iga ing,
as consume - iendly LiDAR op ions a e likely o eme ge in he nea u u e.
As al eady men ioned in Sec. 1.6.1, s a e-o - he-a neu al ne wo ks ope a -
ing on poin cloud da a a e gene ally di ided in o wo mac o ca ego ies: oxel-
based me hods and poin -based me hods. Voxel-based me hods [70, 71, 13, 23]
i s ans o m he poin cloud in o a egula g id o oxels, and hen p ocess
he ans o med da a using h ee-dimensional con olu ional ope a o s. Poin -
based me hods [14, 72, 26, 73], on he o he hand, ope a e di ec ly on he aw
poin cloud by le e aging Poin Ne -like [24, 25] models. Voxel-based me hods
end o be as e han poin -based app oaches, mos ly due o he ac ha a
egula g id o oxels is easie o p ocess compa ed o poin s and he e exis
lib a ies [68, 69] ha allow o e icien con olu ion compu a ion by igno ing
emp y oxels. Con e sely, aw poin clouds a e mo e challenging o p ocess
di ec ly, as hey a e se s wi h no in insic in e nal s uc u e, which makes
poin -based me hods usually less e icien . Howe e , p ocessing poin s di ec ly
a oids he quan iza ion in oduced by he oxeliza ion p ocess, allowing o
mo e accu a e localiza ion.
In his chap e I in es iga e sel -a an ion [27] as a way o s eng hening
in e media e ea u e ep esen a ions o poin -based me hods, which all ely on
Poin Ne -like s uc u es o ex ac ea u es. Mo e speci ically, I build upon
he de ec ion pipeline Vo ene , in oduced in [26] o pe o ming 3D objec
de ec ion on dense poin clouds o con olled scena ios, such as oom scenes
[103, 104]. As his me hod does no adap well o noisie scena ios, such as au-
onomous d i ing scenes ob ained om LiDAR scans, I in oduce some simple
modi ica ions in o de o boos pe o mance. Then, I in oduce sel -a en ion
as an in eg al pa o he Se Abs ac ion (SA) laye s, wi h he aim o mak-
ing poin s awa e o each o he when compu ing ea u es, which should lead o
s onge ep esen a ions and, ul ima ely, be e de ec ion pe o mance.
4.2. Vo ene o 3D Objec De ec ion on D i ing Scena ios 115
4.2 Vo ene o 3D Objec De ec ion on D i ing Sce-
na ios
Be o e del ing in o he de ails o he model, I b ie ly e iew Poin Ne models,
as hey ep esen a cen al componen in poin p ocessing pipelines.
4.2.1 Poin Ne s o poin cloud p ocessing
Di e en ly om images, poin clouds a e se s, ha is hey a e no cha ac e ized
by any explici in e nal s uc u e. As se s, one o he p ope ies hey ha e
is ha hey a e in a ian o pe mu a ions o hei elemen s, meaning ha
shu ling he o de o he poin s does no a ec he cloud i sel . Due o his
p ope y, pa icula ca e mus be aken when p ocessing his kind o da a using
neu al ne wo ks, as he esul ing model mus app oxima e a unc ion ha is
symme ic by cons uc ion, meaning ha i s ou pu mus be una ec ed by he
o de o he inpu elemen s.
Cu en ly, he dominan app oach in li e a u e o dealing wi h se inpu s
is ep esen ed by Poin Ne s [24]. The idea behind his class o models is e y
simple: in o de o he model o app oxima e a symme ic unc ion, i s each
elemen o he se is p ocessed indi idually in o de o ex ac ea u es, and
hen a global desc ip o o he se is ob ained by agg ega ing he in o ma ion
abou each indi idual elemen ia a symme ic unc ion:
({x1, . . . , xn}) = g(h(x1), . . . , h (xn)) .(4.1)
He e, ep esen s he model, hcan be any unc ion and gis a symme ic
unc ion. Commonly, in Poin Ne s he unc ion his app oxima ed by a s ack
o ully-connec ed laye s, also called in li e a u e as a Mul i-Laye Pe cep on
(MLP), while gby max pooling, as he maximum ope a ion is symme ic.
The e o e, he unc ion as a whole is symme ic, and e u ns as ou pu a
single elemen , o en called signa u e, ha encodes a global ep esen a ion
o he inpu se . Such signa u e can hen be used o u he p ocessing: o
122
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
Fo simplici y and e iciency, I modi y his o mula ion sligh ly, ou pu ing
3+4×NS + 2 ×NH alues ins ead:
• i s , I no ice ha in mos si ua ions he e exis s exac ly one size empla e
pe class, meaning ha NS =NC. When his is he case, he p edic o o
he size empla e and he p edic o o he objec class ca y ou exac ly
he same ask, ende ing one o he wo edundan . As a esul , I op o
emo ing he size empla e p edic o , and a es ime I choose he igh
empla e acco ding o he p edic ed objec class;
•second, ins ead o p edic ing he clus e objec ness (2) as well as a dis-
c e e p obabili y dis ibu ion o e he classes (NS), I p edic NS indi id-
ual pe -class objec ness alues ins ead.
This al e na i e o mula ion pe o ms on-pa wi h he o iginal, whils being
mo e compac and equi ing 2 + NS less ou pu s.
Op imiza ion o box cen e s, dimensions and o ien a ions ollows closely
ha o F us um-Poin Ne s, wi h he di e ence ha in his case hese losses
a e compu ed o each posi i e clus e (i.e. clus e s whose cen e is con ained in
a g ound u h box), and a e aged o e hei numbe . To ain he objec ness
p edic o , NS indi idual bina y c oss-en opy losses (Eq. 2.8) a e applied o
each clus e , one o each class, whe e he g ound u h p obabili y is equal o
1 i he clus e cen e is inside an objec o ha class, 0 o he wise. Again, he
inal classi ica ion loss is ob ained by a e aging all loss alues o all clus e s.
A in e ence ime, duplica e de ec ions a e handled by using Non-Maximum
Supp ession, a o ing hose ha ing highe objec ness sco e in case o high o e -
lap wi h o he de ec ions.
4.2.3 Modi ica ions o he baseline o Au onomous D i ing
Scena ios
The Vo ene baseline p esen ed abo e was o iginally hough o pe o m 3D
objec de ec ion on con olled scena ios, such as he indoo scenes depic ed in
4.2. Vo ene o 3D Objec De ec ion on D i ing Scena ios 123
he ScanNe [104] and SUN RGB-D [103] da ase s. He e, poin clouds end
o be qui e dense, and mos poin s usually belong o objec s, wi h a e y
small pe cen age o hem ha a e backg ound (i.e. poin s on he oom loo
o walls). D i ing scena ios cap u ed by LiDAR senso s, on he o he hand,
a e subs an ially di e en : objec s o in e es (e.g. ca s, pedes ians, cyclis s
e c...) a e e y ew and a be ween, which lea es mos o he poin cloud o
be backg ound. Mo eo e , LiDAR clouds a e gene ally noisie and spa ses. Fo
hese easons, he o iginal o mula ion o Vo ene does no adap well o his
new domain, losing pe o mance and o en missing de ec ions.
I iden i y he oo cause o his pe o mance loss in he way ha clus e
cen e selec ion is pe o med. To ecall om he p e ious sec ion, clus e c e-
a ion is accomplished h ough he use o a SA laye , which i s de e mines
a se o cen e s ia he FPS algo i hm and hen agg ega es in o ma ion nea
each cen e . In case o con olled scenes, whe e mos poin s belong o objec s,
i is likely ha he clus e s o med by he o es a e a apa om each o he ,
wi h ela i ely ew isola ed poin s in be ween. In d i ing scena ios, on he
o he hand, mos poin s a e likely o be backg ound, which leads he clus e s
ha o m in co espondence o objec s o be su ounded by nume ous isola ed
o es. Since FPS de e mines he cen e s by s a ing om a andom poin and
i e a i ely selec ing he nex ones such ha hey he a hes apa om he
al eady sampled se , i some noise poin s nea a clus e happen o be selec ed
by FPS be o e any o he poin s belonging o he clus e , hen all poin s o
he clus e migh be igno ed. This beha io is displayed in Fig. 4.2. When his
happens, no ea u es a e ex ac ed by he SA laye o ha clus e , which is
subsequen ly igno ed by he de ec ion module. This has wo majo implica-
ions: due o some objec s being missed, less posi i e samples a e p opaga ed
o he de ec o du ing aining, educing he amoun o eedback and slow-
ing down op imiza ion. Mo eo e , he same phenomenon migh happen a es
ime, leading o alse nega i es ha a e no caused by misclassi ica ions, bu
a he by objec s being skipped when sampling. Bo h o hese aspec s se e ely
unde mine he o e all pe o mance o he model. To limi he pe o mance
124
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
loss induced by missed clus e s, I es wo di e en solu ions o clus e cen e
sampling: seed-based sampling and ea u e dis ance-based sampling [73].
(a)
1
2
?
(b) (c)
Figu e 4.2: Examples o alse nega i es o igina ed by he sampling s a egy.
(a) Vo e poin s, displayed in ed, sampled clus e cen e s, shown in blue and
g ound u h boxes o a ca and a cyclis . (b) Schema ic ep esen a ion o
wha is happening: i noise poin s a ound he objec a e sampled i s , he
en i e clus e migh be skipped due o icini y. (c) Resul ing p edic ed boxes:
no objec is de ec ed due o no clus e cen e s being sampled.
Seed-based sampling Men ioned in he o iginal wo k [26], he idea behind
seed-based sampling is a he simple: when de e mining he se o clus e cen-
e s om he o es using FPS, ins ead o using he coo dina es o he o es,
he coo dina es o hei co esponding seeds a e used ins ead. Since seeds a e
e enly sp ead ou in space, as hey a e he esul o applying FPS mul iple
imes on he o iginal poin cloud, i is unlikely ha all seeds ha co espond
o a clus e o o es a e skipped when sampling he clus e cen e s. Whils
con ibu ing o almos no pe o mance gain when he sys em is applied on
he ScanNe and SUN RGB-D da ase s, his al e na i e sampling s a egy ac-
coun s o mos o he pe o mance gain when ope a ing on he noisie KITTI
d i ing scena ios. An example o esul is displayed in Fig. 4.3.
4.2. Vo ene o 3D Objec De ec ion on D i ing Scena ios 125
(a) (b) (c)
Figu e 4.3: De ec ion esul s om he seed-based sampling s a egy. (a) Loca-
ions o he seed poin s used o sampling, displayed in g een along wi h he
g ound u h boxes. (b) Vo e poin s, in ed, and sampled clus e cen e s, in
blue, along wi h he g ound u h boxes. I can be obse ed ha mul iple
cen e s pe clus e a e sampled. (c) Resul ing p edic ed boxes.
Fea u e dis ance-based sampling In he ecen wo k 3DSSD [73], he
au ho s p opose F-FPS, a a ian o he FPS algo i hm in which he me ic is
gi en by he sum o euclidean and ea u e dis ance be ween he se elemen s, as
opposed o he adi ional o mula ion which only conside s euclidean dis ance
and is he e o e labelled as D-FPS. They show ha using a combina ion o
F-FPS and D-FPS, which hey call FS (Fusion Sampling), inside he SA laye s
o he backbone leads o mo e objec poin s being kep , which esul s in highe
ecall and o e all be e de ec ion pe o mance. This imp o emen is caused by
he ac ha F-FPS allows o poin s ha a e close o each o he o be sampled,
p o ided ha hese poin s encode di e en en i ies in space, such as di e en
objec pa s. Con e sely, backg ound poin s such as poin s on he oad a e
less likely o be chosen, as hey end o ha e simila ea u e embeddings. As
al e na i e solu ion o seed sampling, I p opose o use F-FPS ins ead o D-FPS
in he SA laye o he de ec ion module: his a oids missed clus e s, as poin s
belonging o said clus e s encode di e en in o ma ion om he su ounding
backg ound poin s, and hus a e likely o be selec ed e en i he la e a e
126
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
chosen i s . An example o esul is displayed in Fig. 4.4.
(a) (b)
Figu e 4.4: De ec ion esul s om he ea u e dis ance-based sampling s a egy.
(a) Vo e poin s, in ed, and sampled clus e cen e s, in blue, along wi h he
g ound u h boxes. This me hod, on a e age, leads o mo e objec s poin s
o be selec ed when compa ed o seed-based sampling. (b) Resul ing p edic ed
boxes.
Gi en he inc ease in ecall shown in 3DSSD, I also op o adop ing FS
in he backbone module, ollowing he o iginal implemen a ion. Mo eo e , in-
s ead o p edic ing objec ness like in he o iginal sys em, I choose o p edic
cen e ness ins ead, due o i s syne gy wi h he sampling s a egies abo e.
Cen e ness es ima ion Whils sampling clus e cen e s ia D-FPS esul s
in a mos one cen e pe clus e , when using ei he seed o F-FPS based sam-
pling i is likely ha mul iple cen e s pe clus e a e sampled. As a esul ,
each objec is likely o be de ec ed mul iple imes by he de ec ion module.
While duplica e de ec ions a e handled by NMS, whe e he es ima ed con i-
dence/objec ness is he de e mining ac o in choosing which o he mul iple
de ec ions is kep , he e is no eal co ela ion be ween said con idence and he
ue quali y o he p edic ed box. Clus e cen e s ha a e close o hei co -
esponding objec cen e s, howe e , o en esul in highe quali y p edic ions.
As a esul , ollowing [73], ins ead o p edic ing a pe -class objec ness o each
4.3. Enhancing SA laye s ia Sel -A en ion 127
clus e cen e like in Sec. 4.2.2, I p edic i s pe -class cen e ness alues ins ead.
Mo e speci ically, i a clus e cen e is inside a g ound u h box o a speci ic
class, i is ained o p edic he ollowing alue
pc =3
smin ( , b)
max ( , b)·min (l, )
max (l, )·min ( , d)
max ( , d)(4.3)
and 0 o he wise. In he abo e equa ion, { , b, l, , , d}s and o indica e he
dis ances o he clus e cen e om he on , back, le , igh , op and bo om
aces o he assigned g ound u h box espec i ely. By p edic ing cen e ness
ins ead o objec ness, he subsequen NMS p io i izes boxes o igina ed om
clus e cen e s ha a e close o he objec cen e , and hus ha ing on a e age
highe quali y, which leads o o e all be e pe o mance o he sys em. To
ain o cen e ness p edic ion, I adop he s anda d bina y c oss-en opy loss
unc ion using as a ge , ins ead o a bina y 0/1 alue like o he objec ness,
he g ound u h cen e ness alue.
4.3 Enhancing SA laye s ia Sel -A en ion
While Poin Ne ++ ep esen s a s ong model o ex ac ing ea u es om poin
clouds ha a e local and hie a chical, i s s uc u e su e s om a limi a ion. In
SA laye s, he ea u e agg ega ion o each g oup is pe o med h ough he use
o a sha ed Poin Ne . To achie e pe mu a ion in a iance, Poin Ne p ocesses
each poin in each g oup independen ly using a MLP and hen agg ega es hei
in o ma ion ia a max pooling ope a ion. As a esul , he poin s in each g oup
a e unawa e o each o he while being p ocessed by he MLP, which likely
leads o subop imal ea u e ep esen a ions.
A possible solu ion o ci cum en ing his p oblem has been p oposed in
Poin Web [105], whe e a ea u e adjus men module is in oduced be o e he
MLP o ecalib a e he ea u es o each poin acco ding o i s ela ionship
wi h all o he poin s. Ins ead, I p opose o exploi he a en ion mechanism o
explici ly model in e -poin ela ionships wi hin each g oup.
128
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
4.3.1 The a en ion mechanism
A en ion can be de ined as a unc ion mapping h ee se s o elemen s, called
que ies,keys and alues in o a se o ou pu elemen s. Mo e speci ically, each
que y elemen co esponds o an ou pu , which is gi en by a weigh ed sum
o he alues whe e he weigh associa ed o each alue depends on he com-
pa ibili y o he que y wi h he key co esponding o ha alue. Fo mally, le
Q∈Rnq×dkbe he se o que ies, K∈Rnk ×dk he se o keys and V∈Rnk ×d
he se o alues, whe e nqindica es he numbe o que y elemen s, nk indi-
ca es he numbe o key/ alue pai s, dkindica es he dimensionali y o que ies
and keys and d he dimensionali y o he alues. Then, he mos commonly
adop ed e sion o a en ion [27] is as ollows:
A (Q, K, V )=ΩQ·KT·V. (4.4)
He e, he pai wise do p oduc Q·KT∈Rnq×nk measu es how compa ible
each que y is wi h each key, and Ω (·)is a unc ion mapping he compa ibili y
alues in o weigh s. Usually, Ω (·) akes he o m o a So max unc ion, applied
o each ow o Q·KTindependen ly.
The se s o que ies,keys and alues can be used o model any quan i y, so
long as hey can be encoded in he o m o ec o s. Commonly, in deep lea ning
hese quan i ies a e ob ained by p ojec ing ea u e ep esen a ions ex ac ed
by one o mo e neu al ne wo ks:
Q=PQ(xq), K =PK(xk), V =PV(x ).(4.5)
He e, xq,xk,x could be, o ins ance, ea u e maps e u ned by CNNs o
poin s ex ac ed by a Poin Ne ++ backbone. In he o me case, each pixel
ep esen s an elemen o he se ; in he la e , each poin ep esen an elemen
o he se . PQ(·),PK(·),PV(·) ep esen he p ojec ion unc ions used o map
each o hese ep esen a ions in o he se s o que ies,keys and alues. Such
unc ions a e usually op imized oge he wi h he es o he models, and a e
commonly implemen ed using an MLP applied concu en ly o each elemen
4.3. Enhancing SA laye s ia Sel -A en ion 129
o , in he simples case, as a ma ix mul iplica ion, whe e he pa ame e s
o he ma ices a e ainable. In mos cases, keys and alues a e ob ained
om he same se o ea u es, ha is xk =xk=x . Also, by imposing he
dimen ionali y o he alues o be he same as ha o he keys and que ies
(i.e. d=dk=d ), he ou pu s e u ned by he a en ion unc ion a e o en
used o upda e he o iginal se o que y ea u es:
y=A (PQ(xq), PK(xk ), PV(xk )) + xq.(4.6)
The esul ing se o ea u es ycan he e o e be in e p e ed as an upda ed
e sion o he o iginal se o ea u es xq, in which he upda e depends on he
ela ionship be ween each elemen in he se xqand all he elemen s in he se
xk . No e ha his ope a ion is in a ian o he pe mu a ion o he elemen s
in he se xk and equi a ian o he pe mu a ion o he elemen s in he se
xq, meaning ha pe mu ing he elemen s in xqinduces he same pe mu a ion
on he se y, bu he indi idual alues do no change. Sel -a en ion is a
a ian o a en ion in which que ies,keys and alues a e all de i ed om he
same se o ea u es, ha is x=xq=xk=x . Sel -a en ion can he e o e be
in e p e ed as an ope a ion ha upda es each elemen in xdepending on i s
ela ionship wi h he es o he elemen s in same se .
A guably he i s success ul applica ion o sel -a en ion is ep esen ed by
he seminal wo k in [27], in which he au ho s p opose an en i ely no el neu-
al a chi ec u e o pe o m he ask o machine ansla ion e ol ing en i ely
a ound he a en ion mechanism. Mo e speci ically, hey in oduce he T ans-
o me , an Encode -Decode s uc u e: in he Encode , sel -a en ion is used o
model he ela ionship be ween he di e en wo ds in he inpu sen ence. In he
Decode , sel -a en ion is i s used o ex ac in o ma ion om he ansla ed
sen ence gene a ed so a ; hen, by using hese ea u es as que ies, and he
ea u es gene a ed by he Encode as keys/ alues, addi ional a en ion blocks
a e used o gene a e he se o ea u es used o nex wo d p edic ion. Pe haps
one o he mos signi ican inno a ions in oduced in he T ans o me model is,
howe e , Mul i-Headed A en ion. The idea behind his concep is as ollows:
130
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
ins ead o p ojec ing he ea u e se s xqand xk in o a single se o que ies,
keys and alues, hey a e p ojec ed in o ho hese se s ins ead:
PQ(xq) = Q=hQ(1), . . . , Q(h)i∈Rnq×(h·dh),
PK(xk ) = K=hK(1), . . . , K(h)i∈Rnk×(h·dh),(4.7)
PV(xk ) = V=hV(1), . . . , V (h)i∈Rnk×(h·dh).
Each o hese se s is hen p ocessed independen ly and in pa allel using a en-
ion, and hei ou pu s a e used oge he ia an addi ional p ojec ion unc ion
PO(·), yielding he ou pu o he Mul i-Headed A en ion:
MHA (Q, K, V ) = PO([A 1,...,A h]) ∈Rnq×d.(4.8)
He e, [A 1,...,A h]∈Rnq×(h·dh)and A i=A Q(i), K(i), V (i). Again, he
p ojec ion unc ion PO(·)is implemen ed ei he as an MLP o , mo e commonly,
as a ma ix mul iplica ion wi h lea nable ma ix pa ame e s. Like be o e, he
esul o he Mul i-Headed A en ion is inally used o upda e he inpu se o
que y ea u es:
y=MHA(PQ(xq), PK(xk ), PV(xk )) + xq.(4.9)
The au ho s a gue ha using mul iple, independen a en ion heads allows he
model o concen a e on di e en aspec s o he inpu a di e en posi ions, im-
p o ing he o e all pe o mance o he sys em. No e ha by choosing dh=d/h
and pa allelizing he compu a ion o he se e al a en ion componen s, Mul i-
Headed A en ion in oduces no compu a ional o e head compa ed o he a-
di ional o mula ion. To be e use he inpu ea u es wi h hose compu ed
by he Mul i-Headed A en ion and imp o e aining, each Mul i-Headed A -
en ion Block (MAB) in he T ans o me encode adop s Laye No maliza ion
[106] o ea u e no maliza ion, as well as an addi ional MLP o pe o ming
addi ional ea u e usion, leading o he ollowing inal o mula ion:
MAB (xq,xk ) = LN (e
y+MLP (e
y)) ,(4.10)
4.3. Enhancing SA laye s ia Sel -A en ion 131
whe e e
y=LN (y),yis ob ained using Eq. 4.9, MLP indica es he MLP used o
ea u e usion and LN ep esen he Laye No maliza ion ope a o . A schema ic
ep esen a ion o he MAB block is depic ed in Fig. 4.5.
A en ion
Block
A en ion
Block
A en ion
Block
A en ion
Mask
LN MLP LN
Figu e 4.5: Schema ic ep esen a ion o a MAB block.
4.3.2 Sel -a en ion applied o SA laye s
I p opose o eplace bo h he MLP o ea u e ex ac ion and he max-pooling
o ea u e agg ega ion inside SA laye s wi h a en ion-based p ocessing, in
o de o enhance he ex ac ed ea u e ep esen a ions by explici ly modeling
he ela ionships be ween he poin s wi hin each g oup. In o de o be e ec i e,
such a en ion-based mechanisms should p ese e he p ope y o in a iance
o pe mu a ions o he o iginal o mula ion. Mo e speci ically, I in oduce wo
di e en o mula ions o eplace he MLP.
Full A en ion He e, I p opose o eplace each laye o he SA MLP h(x)
(see Eq. 4.1) wi h a sel -a en i e MAB block MAB (x,x), whe e x∈RK×C
ep esen s any single g oup o poin s e u ned by he sampling and g oup-
ing s ages o he SA laye , Kis he numbe o poin s in he g oup and Cis
he dimensionali y o each poin . This can be done di ec ly, as MAB (x,x)is
pe mu a ion equi a ian . Commonly, mos models p og essi ely inc ease he
dimensionali y Co he ea u es hey ex ac as hei dep h inc eases. To al-
low o he same lexibili y when using a en ion, I in oduce an addi ional
138
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
4.4.3 T aining Rou ine
All es ed models a e ained on he KITTI aining spli o a o al o 80700
i e a ions, using ADAM as op imize and ba ch size 4. The lea ning a e is
ini ially se o 5·10−4and educed by a ac o o 10 a e 64560 i e a ions.
Gi en he ela i ely limi ed da ase , I employ hea y da a augmen a ion in
o de o a o gene aliza ion and imp o e pe o mance. Fi s o all, I adop he
so-called mixup augmen a ion [71], an augmen a ion echnique o en adop ed
in ecen li e a u e when aining poin -based and oxel-based de ec o s which
consis s in inse ing in o he cu en scene objec s om o he scenes in o de o
p o ide iche aining examples o he model. P ac ically, a da abase con ain-
ing all objec s in he aining spli is gene a ed. Then, a aining ime, e e y
ime a new aining sample is loaded, a ce ain amoun o andomly sampled
objec s om his da abase is pas ed in o i , aking ca e ha no collisions be-
ween objec s a e gene a ed in he p ocess. An example o a mixup-augmen ed
image is shown in Fig. 4.9.
(a) (b)
Figu e 4.9: Example o mixup augmen a ion. (a) o iginal sample wi h he as-
socia ed g ound u h boxes. (b) Sample a e mixup augmen a ion.
A e mixup augmen a ion is pe o med, addi ional da a augmen a ion is
ca ied ou o u he inc ease he a iabili y o he aining da a:
• i s , he poin cloud is lipped wi h espec o he came a xz-plane wi h
4.5. Expe imen al Resul s 139
p obabili y 0.5;
•second, he poin cloud is scaled by a ac o andomly sampled wi hin
he in e al [0.9,1.1];
• hi d, he poin cloud is o a ed abou he came a y-axis by an angle
andomly sampled wi hin he in e al [−π/4, π/4];
• ou h, each objec , and i s associa ed poin s, is independen ly o a ed
abou i s y-axis by an angle andomly sampled wi hin he in e al [−π/3, π/3];
• i h, each objec , and i s associa ed poin s, is shi ed along he came a x
and z di ec ions by wo quan i ies andomly sampled wi hin he in e al
[−1,1].
4.5 Expe imen al Resul s
In his sec ion, I p esen he esul s ob ained om he expe imen s pe o med
on he p oposed sys em. Fi s , I pe o m an abla ion s udy on he a chi ec u al
design choices p esen ed in sec ions 4.2.3 and 4.3.2, alida ing hem. Then, I
ca y ou a mo e in-dep h s udy on a en ion, compa ing di e en o mula ions
agains he anilla model wi h no a en ion mechanisms. Finally, I pe o m a
compa ison be ween he p oposed sys em and o he s a e-o - he-a LiDAR-
based 3D de ec ion models.
I no e ha , despi e he ac ha he sys em is ained on all h ee main
KITTI classes (i.e. ca , pedes ian and cyclis ), mos compa isons a e pe -
o med conside ing only he ca class since, gi en he limi ed size o he ain-
ing and e alua ion spli s (3712 and 3769 samples espec i ely), i is he only
class ha p o ides nume ous enough g ound u h samples o enable eliable
analyses.
140
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
4.5.1 Abla ion s udy on he design choice
The imp o emen s in pe o mance b ough abou by he modi ica ions o he
o iginal Vo ene a chi ec u e (Sec. 4.2.3) and by he in eg a ion o he a en ion
mechanism (Sec. 4.3.2) a e displayed in Tab. 4.1.
I can be seen ha , indeed, he main cause o Vo ene poo pe o mance
on noisie au onomous d i ing scena ios is i s inadequa e clus e cen e sam-
pling s a egy: jus swi ching o seed-based cen e sampling leads o a ela i e
imp o emen o abou ∼14%. No e ha his pe o mance is no he esul o
a di e en neu al a chi ec u e, bu a he is due o he ac ha seed-based
sampling mos ly a oids missing clus e s, inc easing es - ime ecall. Mo eo e ,
his sampling s a egy o en leads o he same clus e being p ocessed mul iple
imes by he box p edic io , as i is likely ha mul iple poin s inside i end
up being sampled; as a esul , he ne wo k bene i s om inc eased eedback
du ing aining, which leads o as e and be e con e gence and he e o e
highe quali y de ec ions.
Mixup augmen a ion con ibu es o a u he boos in pe o mance, despi e
i being qui e limi ed o he ca class and mos ly isola ed o he ha d examples.
This is likely due o he ac ha ca s a e by a he mos ep esen ed class
in he da ase , while also being he mos sp ead ou among di e en samples.
Con e sely, pedes ians and cyclis s a e a less in numbe (abou 1/5 and 1/10
compa ed o ca s, espec i ely) and end o concen a e on a ew selec samples
[22]. The e o e, o cing each inpu sample o con ain mul iple ins ances o hose
classes s abilizes aining, which is no longe domina ed by ca s. This leads o
a conside able gain in pe o mance o pedes ian and cyclis de ec ion, as
displayed in Tab. 4.2.
Es ima ing cen e ness o e objec ness u he inc eases de ec ion accu acy,
since in his case he sco e assigned by he model o each p edic ed box be e
co ela es wi h he quali y o he box i sel . This has wo implica ions: on he
one hand, noisy de ec ions and ou igh alse posi i es a e mo e likely o ha e
low sco es, which causes he P ecision-Recall cu e o ha e g ea e a ea, im-
p o ing he AP me ic. On he o he hand, i is less likely ha NMS disca ds
4.5. Expe imen al Resul s 141
Ca AP 3D
Seed
Sampling
Fea u e
Sampling
Mixup
Augmen a ion Cen e ness A en ion Easy Mode a e Ha d
Vo ene [26] 76.85 65.99 64.69
X87.81 76.91 72.81
X X 87,37 76.96 74,73
X X X 87.86 78.25 77.14
P oposed X X X 88.49 78.55 77.38
P oposed X X X Full 89.20 79.23 78.35
P oposed X X X Induced (8) 89.37 79.40 78.51
Table 4.1: Abla ion s udy on di e en a ian s o he p oposed sys em o e he KITTI e alua ion se o
he Ca class.
142
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
be e boxes in a o o mo e inaccu a e ones. Unsu p isingly, cen e ness es i-
ma ion leads o be e imp o emen s o ha de samples: since hese samples
mos ly consis o unca ed o a away objec s, hey end o be ep esen ed by
e y ew poin s and he e o e a e mo e likely o esul in noisie clus e cen e
p edic ions.
Expe imen ally, using F-FPS on o es o sample clus e cen e s pe o ms
be e han seed-based sampling, and hus i cons i u es he de aul choice
o he p oposed anilla model. Ex ending such model wi h a en ion-based
blocks leads o u he boos s in pe o mance. A mo e de ailed s udy ela ed
o a en ion ollows in he nex sec ion.
AP3D)
Model class Easy Mode a e Ha d
Vanilla Ped 36.89 35.03 32.03
SS Ped 53.28 49.96 45.72
SS+MIX Ped 61.84 56.31 51.46
Vanilla Cyc 39.29 27.50 26.79
SS Cyc 67.13 52.52 50.52
SS+MIX Cyc 83.33 64.63 60.50
Table 4.2: Compa ison be ween he anilla sys em, he sys em wi h seed-based
sampling (SS), and he sys em wi h seed-based sampling and mixup augmen-
a ion (SS+MIX) on he KITTI e alua ion se o he Pedes ian (Ped) and
Cyclis (Cyc) classes.
4.5.2 Abla ion s udy on a en ion ype
Poin -based 3D de ec ion me hods (i.e. models based o o Poin Ne ++) a e,
by cons uc ion, non-de e minis ic: he esul o each FPS ope a o , in ac ,
depends on he i s selec ed poin in he cloud, which is chosen andomly.
While his phenomenon was shown o ha e a negligible e ec in e ms o es
4.5. Expe imen al Resul s 143
esul s s abili y [25], when coupled wi h he hea y da a augmen a ion in ol ed
du ing aining and he ac ha he aining se is ela i ely limi ed, i migh
lead o non- i ial di e ences du ing op imiza ion. Consequen ly, in o de o
pe o m a mo e accu a e compa ison be ween he anilla model and he a ious
a en ion-based o mula ions and o educe he impac o such andomness on
he esul s, I ain 5 iden ical models o each ype, showing he bes model ou
o he 5 as well as hei mean pe o mance and hei s anda d de ia ion.
The esul s o his s udy a e shown in Tab. 4.3. As can be obse ed, all
a en ion-based models pe o m, on a e age, be e han hei anilla coun e -
pa , despi e being cha ac e ized by highe aining ins abili y as indica ed by
hei highe s anda d de ia ion alues.
Su p isingly, he induced-a en ion based models ou pe o m hei ull-
a en ion coun e pa s. The eason o his migh be wo old: on he one
hand, induced-a en ion esul s in a highe numbe o ainable pa ame e s,
and he e o e in a model ha ing highe ep esen a ional capaci y, compa ed o
ull-a en ion, as each induced-a en ion laye is comp ised o wo mul i-headed
a en ion blocks ins ead o one (see Fig. 4.7). On he o he hand, he use o
inducing poin s migh allow he model o be e handle po en ial ou lying el-
emen s wi hin each g oup, educing hei con ibu ion o he ou pu .
Pe o ming ea u e agg ega ion ia a en ion ins ead o max-pooling also
leads o ele an imp o emen s, as i allows he model o speci ically modula e
he con ibu ion o he ou pu o each componen o he g oup, allowing he
sys em o ep esen a much wide amily o unc ions compa ed o simply
pe o ming a channel-wise max-pooling. This s onge ep esen a ional powe ,
howe e , also leads o u he aining ins abili y, as displayed by he inc ease
in he s anda d de ia ion alue o he alida ion AP.
I also es bo h ull-a en ion and induced-a en ion applied o he SA laye
o he de ec ion module (SA-C), which is esponsible o clus e c ea ion and
g ouping. In bo h cases, despi e he highe ep esen a ional capaci y, he e-
sul s a e in e io o hei coun e pa s using a s anda d MLP ollowed by max-
pooling. While induced-a en ion pe o ms only sligh ly wo se, ull-a en ion
144
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
leads o a mo e p onounced d op in pe o mance, esul ing o a solu ion ha
on a e age pe o ms wo se han he anilla model. This solu ion is also consid-
e ably mo e uns able du ing aining, as shown by he s anda d de ia ion o
i s pe o mance, which is almos ou imes ha o he model wi h no a en ion
in SA-C. This deg ada ion in pe o mance is likely o be a ibu ed o he ac
ha SA-C u ilizes BQS wi h low adius as g ouping algo i hm, which leads o
many clus e s ha ing epea ed poin s due o padding. This is especially ue
o o es ha do no belong o objec s, which a e likely o be isola ed in space.
Repea ed poin s lead o ins abili ies when used in conjuc ion wi h a en ion,
as each one ac i ely con ibu es o he end esul . The ob ious solu ion o
a oiding his p oblem would be o eplace BQS wi h kNN; his, howe e , leads
o addi ional noise in he clus e s, as each one is a mo e likely o include back-
g ound noise poin s o e en poin s om ex e nal objec s. Conside ing ha he
esul ing sys em pe o ms wo se han i s BQS-based coun e pa (see Tab. 4.4)
while being less e icien , I op agains his kind o solu ion.
4.5.3 Compa ison wi h S a e-o - he-A Sys ems
Tab. 4.5 shows a compa ison be ween he p oposed solu ion and o he LiDAR
and usion-based s a e-o - he-a 3D de ec ion sys ems on he KITTI alida ion
se .
The anilla sys em pe o ms conside ably be e han all o he app oaches
ba ing Poin RCNN, which sligh ly ou pe o ms i a all di icul ies. Poin R-
CNN, howe e , is o e all slowe , exhibi ing an in e ence ime o abou 100ms
agains he 70ms o he anilla model.
In eg a ing a en ion in he backbone ne wo k conside ably inc eases pe -
o mance, allowing he esul ing models o con incingly su pass Poin RCNN
while only equi ing an ex a 15ms o ull a en ion and 19ms o induced-
a en ion. This imp o emen is pa icula ly ele an o ha de objec s, which
ad oca es o he e ec i eness o a en ion mechanisms o deal wi h ha de
and noisie da a.
4.5. Expe imen al Resul s 145
Ca AP3D: mean ±s .d. (max)
Model Easy Mode a e Ha d
Vanilla 88.45 ±0.17 (88.49) 78.43 ±0.17 (78.55) 77.29 ±0.11 (77.38)
F-23-MP 88.35 ±0.23 (88.65) 78.59 ±0.22 (78.90) 77.67 ±0.19 (77.99)
F-23-AP 88.83 ±0.24 (89.20) 78.93 ±0.23 (79.23) 78.03 ±0.25 (78.35)
I(8)-23-MP 88.61 ±0.23 (88.72) 78.77 ±0.15 (78.93) 77.79 ±0.20 (77.89)
I(8)-23-AP 89.01 ±0.32 (89.37) 79.00 ±0.26 (79.40) 78.10 ±0.24 (78.51)
F-23C-AP 88.29 ±0.72 (89.07) 78.37 ±0.82 (79.09) 77.34 ±0.98 (78.15)
I(8)-23C-AP 88.92 ±0.15 (89.15) 78.93 ±0.20 (79.25) 78.03 ±0.26 (78.43)
Table 4.3: Abla ion s udy on he a en ion mechanism. A en ion-based models
a e shown in he o ma A-B-C:A ep esen s he a en ion ype, whe e F
indica es ull-a en ion and I(n) indica es induced a en ion wi h n inducing
poin s; B ep esen s he SA laye s he a en ion is applied o, whe e 2, 3
and C indica es he laye s SA-2, SA-3 and SA-C epsec i ely (see Fig. 4.8);
C ep esen s he echnique adop ed o pe o ming ea u e agg ega ion: MP is
he s anda d channel-wise max-pooling, AP is he a en ion-based agg ega ion.
The esul s a e shown in he o ma mean ±s .d. (max), o e a o al o 5
expe imen s.
Ca AP3D: mean ±s .d. (max)
SA-C Easy Mode a e Ha d
BQS 88.45 ±0.17 (88.49) 78.43 ±0.17 (78.55) 77.29 ±0.11 (77.38)
kNN 88.19 ±0.23 (88.48) 78.11 ±0.19 (78.33) 76.82 ±0.30 (77.19)
Table 4.4: Pe o mance compa ison be ween using BQS and kNN in he SA-C
laye . Bo h models a e wi hou a en ion.
146
Chap e 4. 3D Objec De ec ion on LiDAR scans ia Vo ing and
Sel -A en ion Mechanisms
Me hod APBEV @ 0.7 IoU AP3D @ 0.7 IoU
Easy Mode a e Ha d Easy Mode a e Ha d
MV3D [67] 86.55 78.10 76.67 71.29 62.68 56.56
F-Poin Ne [23] 88.16 84.02 76.44 83.76 70.92 63.65
AVOD [12] - - - 84.41 74.44 68.65
VoxelNe [70] 89.60 84.81 78.57 81.97 65.46 62.85
SECOND [71] 89.96 87.07 79.66 87.43 76.48 69.10
Poin Pilla s [13] - - - - 77.98 -
Poin RCNN [13] - - - 88.88 78.63 77.38
Vanilla (mean) 90.08 87.95 84.73 88.45 78.43 77.29
I(8)-23-AP (mean) 90.26 88.34 87.36 89.01 79.00 78.10
F-23-AP (mean) 90.22 88.26 87.30 88.83 78.93 78.03
Vanilla (max) 90.17 88.18 86.03 88.49 78.55 77.38
I(8)-23-AP (max) 90.46 88.58 87.53 89.37 79.40 78.51
F-23-AP (max) 90.36 88.56 87.52 89.20 79.23 78.35
Table 4.5: Pe o mance compa ison be ween he p oposed models and s a e-o -
he-a 3D de ec ion sys ems on he KITTI alida ion se o he Ca class. I
highligh he bes pe o ming models, conside ing bo h mean pe o mance and
bes pe o mance.
4.5. Expe imen al Resul s 147
4.5.4 Quali a i e esul s
In Figs. 4.10 and 4.11 I show he esul s o he bes pe o ming induced
a en ion-based model on some KITTI alida ion samples, displaying all h ee
main classes: Ca s, Pedes ians and Cyclis s.
As can be obse ed, he model is capable o de ec ing objec s wi h high po-
si ional accu acy, while being obus o occlusions and uncommon o ien a ions
(Figs. 4.10a, 4.10b, 4.10c, 4.11a). A common cause o de ec ion inaccu acies
is hea y objec unca ion, such as in Fig. 4.10d, in which case he p edic ed
o ien a ion migh be noisy. High dis ances migh also be p oblema ic, as hey
could lead o alse nega i es (Fig. 4.10e) especially o he smalle classes, as
he numbe o poin s migh be insu icien o success ul objec iden i ica ion.
O all he classes, pedes ians a e easily he mos challenging o a LiDAR-
based sys em due o he ac ha hey a e small, non- igid and ha e no s an-
da d s uc u e like ca s o bicycles. As a esul , hey a e di icul o es ima e,
especially in e ms o o ien a ion (Fig. 4.11a), and migh gi e ise o alse
posi i es as hey a e easily con used wi h o he en i ies, such as a ic sign,
poles o small ees. An example o his phenomenon is displayed i Fig. 4.10 ,
whe e a a ic sign is mis akenly in e p e ed as a pedes ian. Ano he chal-
lenging scena io o his kind o model is ep esen ed by si ua ions in which
many small objec s a e close oge he in space, such as g oups o pedes ians
(Fig. 4.11d). In hese si ua ions, o e poin s migh e oneously ga he a ound
objec s ha a e di e en om he ones hey belong o: when his happens,
ce ain objec s migh be le wi h e y small clus e s whose ea u es a e e y
simila o hose o neighbo ing clus e s, and he e o e un he isk o being
igno ed du ing sampling, o igina ing a alse nega i e.
The examples shown in Fig. 4.11, in pa icula , highligh cases in which
he model co ec ly de ec s objec s ha a e clea ly isible, bu no anno a ed:
• he pedes ian on he igh in Fig. 4.11a, dis inc ly isible in bo h he
image and he LiDAR scan;
• he pedes ian and he cyclis on he igh in Fig. 4.11b;