scieee Science in your language
[en] (orig)

Automatic 3D Object Recognition and Localization for Robotic Grasping

Abstract

This project aims to design a detection and pose estimation pipeline for objects, to be used in an industrial environment in a robotic grasping setting.

Read accessible full text

Automatic 3D Object Recognition and Localization for Robotic Grasping

Author: Bruno Miguel Silva Espírito Santo
Year: 2020
DOI: 10.34626/57h9-mr89
Source: https://repositorio-aberto.up.pt/bitstream/10216/132916/2/419215.pdf
FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO
Au oma ic 3D Objec Recogni ion and
Localiza ion o Robo ic G asping
B uno Miguel Sil a Espí i o San o
WORKING VERSION
Mes ado In eg ado em Engenha ia Ele o écnica e de Compu ado es
Supe iso : Gil Manuel Magalhães de And ade Gonçal es
Second Supe iso : Liliana Pa ícia Saldanha An ão
July 6, 2020
Abs ac
Wi h he ad en o Indus y 4.0 and i s highly econ igu able manu ac u ing con ex , he ypical
ixed-posi ion g asping sys ems a e no longe usable. This eali y unde lined he necessi y o ully
au oma ic and adap able obo ic g asping sys ems. The de elopmen o a Compu e Vision sys em
capable o de ec ing, iden i ying and es ima ing he 6D pose o an objec would b ing us close o
ha eali y. Wi h ha in mind, he main pu pose o his hesis is o join Machine Lea ning models
o de ec ion and pose es ima ion in o an au oma ic sys em o be used in a g asping en i onmen .
To achie e his, ex ensi e esea ch was ca ied ou on he cu en s a e-o - he-a app oaches.
The de eloped sys em uses Mask-RCNN and Dense usion models o he ecogni ion and pose
es ima ion o objec s, espec i ely. The g asping is execu ed aking in o conside a ion bo h he
pose and he objec ’s ID, as well as allowing o use and applica ion adap abili y h ough an ini ial
con igu a ion. The sys em was es ed bo h on a alida ion da ase and in a eal wo ld en i onmen .
The main esul s show ha he sys em has mo e di icul y wi h complex objec s, howe e , i shows
p omising esul s o simple objec s, e en wi h aining on a educed da ase . I is also able o
gene alize o objec s sligh ly di e en han he ones seen in aining. In g asping expe imen s,
he e is a 60% success a e in he bes cases, o simple g asping a emp s.
i
ii
Acknowledgemen s
My deepes hanks o bo h my supe iso s. I was an absolu e pleasu e o wo k on his hesis wi h
such g ea people and in a e y enjoyable en i onmen , despi e he di icul wo king si ua ions his
yea . A lo was lea ned in de eloping his hesis and I could no ha e done i wi hou hei i eless
help.
B uno Miguel Sil a Espí i o San o
iii

i
Con en s
Lis o Figu es ii
Lis o Tables ix
Abb e ia ions xi
1 In oduc ion 1
1.1 Con ex ....................................... 1
1.2 Mo i a ion...................................... 2
1.3 P oblemDe ini ion ................................. 3
1.4 Objec i es...................................... 4
1.5 ThesisS uc u e................................... 4
2 Li e a u e Re iew 7
2.1 MachineLea ning.................................. 7
2.1.1 A i icial Neu al Ne wo ks . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.2 PoseandRo a ions ................................. 13
2.3 Rela edWo k .................................... 13
2.3.1 Da ase s................................... 14
2.3.2 Machine Lea ning in Objec De ec ion and Recogni ion . . . . . . . . . 15
2.3.3 Machine Lea ning in Pose Es ima ion . . . . . . . . . . . . . . . . . . . 21
3 3D Objec Recogni ion and Localiza ion Sys em 31
3.1 Sys emO e iew.................................. 31
3.1.1 Ope a ionModes.............................. 32
3.2 Sys em’sModules.................................. 35
3.2.1 Da a Acquisi ion Module . . . . . . . . . . . . . . . . . . . . . . . . . . 35
3.2.2 Compu e Vision Module . . . . . . . . . . . . . . . . . . . . . . . . . . 36
4 Expe imen s and Resul s 41
4.1 Da ase ....................................... 41
4.2 O lineTes ingPhase ................................ 43
4.2.1 T aining he Objec De ec ion Model . . . . . . . . . . . . . . . . . . . 43
4.2.2 Pose Es ima ion Model T aining . . . . . . . . . . . . . . . . . . . . . . 45
4.2.3 Compu e Vision Module Tes ing . . . . . . . . . . . . . . . . . . . . . 46
4.3 OnlineTes ingPhase ................................ 47
4.3.1 Da a Acquisi ion Module Tes ing and Analysis . . . . . . . . . . . . . . 48
4.3.2 Objec De ec ion and Segmen a ion Tes ing . . . . . . . . . . . . . . . . 48
4.3.3 G aspingExpe imen s ........................... 54
i CONTENTS
4.4 Limi a ions ..................................... 56
5 Conclusions and Fu u e Wo k 59
5.1 Conclusions..................................... 59
5.2 Fu u eWo k..................................... 60
Re e ences 61
Lis o Figu es
2.1 Rep esen a ion o a pe cep on . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.2 Rep esen a iono aae................................ 10
2.3 NNVSCNN .................................... 11
2.4 Illus a ion o he a chi ec u e o Fas e R-CNN . . . . . . . . . . . . . . . . . . 15
2.5 Illus a ion o he a chi ec u e o R-FCN . . . . . . . . . . . . . . . . . . . . . . 17
2.6 Anexampleo asco emap............................. 17
2.7 O e iew o he SSD a chi ec u e . . . . . . . . . . . . . . . . . . . . . . . . . 18
2.8 O e iew o he o iginal YOLO ne wo k . . . . . . . . . . . . . . . . . . . . . . 19
2.9 O e iew o he SSD-6D model . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.10 Illus a ion o e e y iewpoin conside ed . . . . . . . . . . . . . . . . . . . . . 22
2.11 O e iew o he PoseCNN sys em . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.12 O e iew o he DPOD sys em . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
2.13 Illus a ion o he pose e inemen block . . . . . . . . . . . . . . . . . . . . . . 25
2.14 O e iew o he DenseFusion a chi ec u e . . . . . . . . . . . . . . . . . . . . . 26
2.15 Illus a ion o he pose e inemen ne wo k . . . . . . . . . . . . . . . . . . . . . 27
2.16 Illus a ion o he Hyb idPose model . . . . . . . . . . . . . . . . . . . . . . . . 28
2.17 O e iew o he DeepIM model . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.1 Illus a ion o he da a low in he sys em . . . . . . . . . . . . . . . . . . . . . . 31
3.2 O e iew o all ope a ion modes . . . . . . . . . . . . . . . . . . . . . . . . . . 33
3.3 S a e machine o he pose es ima ion module . . . . . . . . . . . . . . . . . . . . 34
3.4 ZED S e eo Came a and dimensions illus a ion . . . . . . . . . . . . . . . . . . 35
3.5 O e iew o he Mask-RCNN model . . . . . . . . . . . . . . . . . . . . . . . . 37
4.1 Dis ibu ion o numbe o ins ances pe objec . . . . . . . . . . . . . . . . . . . 41
4.2 Selec ion o illus a i e images om he da ase . . . . . . . . . . . . . . . . . . 42
4.3 O e iew o all he scena ios used in de ec ion es ing . . . . . . . . . . . . . . . 49
4.4 Illus a ion o one g asping posi ion pe objec . . . . . . . . . . . . . . . . . . . 54
4.5 RMS e o in dep h da a o 720p esolu ion . . . . . . . . . . . . . . . . . . . . 56
ii
2In oduc ion
ML is a ield o esea ch dedica ed o compu e s’ abili y o in e ac wi h hei en i onmen
and lea n om he da a collec ed. I is based on ma hema ical models, employed o enable he
lea ning o pa e ns in da a, and is o en used o inc ease he sys em’s in elligence/ lexibili y o
when p oblems canno be sol ed using adi ional explici p og amming.
On he o he hand, Deep Lea ning is a b anch o ML ha desc ibes a se o modi ied ML
echniques inspi ed by he biological ne ous sys em. Composed o ne wo ks o pa allel and
concu en con olu ions, as well as o he ma hema ical ope a ions pe o med di ec ly on he inpu
da a, i acqui es a g oup o ep esen a i e heu is ics be ween inpu and ou pu da a.
Due o he con incing esul s ha ha e been achie ed in he scope o compu e ision, he e is
an inc easing end owa ds he implemen a ion o DL algo i hms in obo ic g asping applica ions.
Wi h ML and DL, obo s a e able o in e ac wi h hei en i onmen and espond o a ious s imuli
in a comple ely au oma ed manne , enabling obo s o use came as o pe cei e hei en i onmen
and e en "unde s and" i when using hese echniques.
The main challenges in he ield o CV co e ed in his hesis a e objec de ec ion/ ecogni ion
and pose es ima ion, which will be ackled wi h he use o ML models. In objec de ec ion and
ecogni ion, he ask is o iden i y objec s in an image and a ibu e a seman ic classi ica ion (iden-
i y he class i belongs o). As o pose es ima ion, he e is a need o es ima e he objec ’s posi ion
and o ien a ion ela i e o he came a.
1.2 Mo i a ion
As p e iously men ioned, objec de ec ion/ ecogni ion and pose es ima ion is o g ea impo ance
o obo ic sys ems in se e al asks. One o hese asks is obo ic g asping: he ac o a obo ic
manipula o handling objec s in an au oma ed manne .
Fo ypical obo ic g asping o be accomplished, he obo needs o be able o know whe e and
how an objec is posi ioned in i s coo dina e sys em. 6-dimensional (6D) poses, composed o 3
DOF o he posi ion and 3 o o ien a ion, a e o en used o his pu pose. Con a y o ypical 3D
pose, 6D poses p o ide he o a ion, gi ing inpu no only on he "whe e" bu also on he "how."
Ne e heless, when compa ed o humans, obo s p esen signi ican ly lowe success a es in
g asping objec s in dynamic en i onmen s, as one would expec . Humans a e inhe en ly good a
pe cei ing hei en i onmen and he objec s ha su ound hem, iden i ying hei cha ac e is ics,
posi ion, and deciding na u ally whe e o g asp an objec and how much o ce o exe on i .
Wi h he eali y o inc easing Human-Robo Collabo a ion solu ions in he indus y, humans,
and obo s no only sha e physical space bu wo k oge he as a eam. I is o g ea impo ance o
a eam o ha e he same pe cep ion and conside a ions o he asks hey will ul ill.
Gi en his, ha ing g asping solu ions ha do no ake in o conside a ion he same aspec s as
humans do, would lead o wo se eam pe o mance due o misma ches in g asping posi ions o
hando e s, o e en due o g ippe damages in g asped ma e ial.
Inspi ed by his, mo i a ion is p o ided o de elop solu ions ha imp o e manipula ion asks
in lexible and collabo a i e indus ial en i onmen s, cha ac e is ic o he new indus ial eali y

1.3 P oblem De ini ion 3
imposed by Indus y 4.0. By p o iding obo s, in an indus ial en i onmen , wi h he abili y o
lea n hei g asping asks au onomously in a simila manne o humans, collabo a i e obo ics is
also empowe ed.
This way, obo s a e gi en mo e pe cep ion o he objec s in a manipula ion ask. O e all,
obo s a e able o wo k mo e in ui i ely wi h human ope a o s and e en assis hem in a ious
asks in a sa e and sha ed en i onmen .
1.3 P oblem De ini ion
The p oblem o objec de ec ion and pose es ima ion is one o he ac i e esea ch ha has su e ed
inc emen al imp o emen s, as can be seen in chap e 2. Achie ing accu a e pose es ima ion o an
objec would enable obo s o pe o m g asping in a ully au oma ic and e icien manne .
Howe e , in o de o imp o e obo ic g asping, besides localiza ion, p o iding objec iden i i-
ca ion can also be use ul. Fo his, ML is o en used, classi ying he seman ic class o he objec
ha can hen be u ilized o adap ha g asping o he objec ’s cha ac e is ics. Fo ins ance, he
g ipping s eng h ha he manipula o can apply wi hou causing damage o he objec can be in-
e ed om i s classi ica ion, o e en choosing which objec o pick ou in a g oup o objec s, by
i s cha ac e is ics.
The e a e al eady many challenges wi h de ec ing and es ima ing he pose o objec s, a ising
wi h he a iabili y o he su ounding en i onmen , such as a ying ligh condi ions, occlusion o
objec s, and di e si y o objec dimensions. O he issues can also occu in e ms o he accu acy
o he senso s used and he ypes o noise in oduced.
Gi en his, he sys em o be de eloped needs o be accu a e and as enough o wo k in a eal
indus ial en i onmen and be able o handle mul iple objec s o di e en dimensions, in si ua ion
o occlusions o pa o he objec s. In sum, challenges ha a ise in objec de ec ion and pose
es ima ion sys ems, by hemsel es, a e:
•The ligh ing condi ions o he en i onmen ;
•Occlusion and unca ion o ce ain objec s;
•The need o models o be able o pe o m in eal- ime wi h low p ocessing speeds;
•The need o high accu acy in indus ial en i onmen s whe e sys ems need o be sa e o
ope a e.
The e a e al eady many objec de ec ion/classi ica ion sys ems, as well as pose es ima ion
models o g asping. Howe e , e y ew showing he bene i s o joining bo h app oaches o a
mo e comple e and au onomous g asping, much less ully au oma ed o based on only came a
ame in o ma ion (bo h RGB and/o dep h).
The e o e, he e is a challenge o joining localiza ion and ca ego iza ion o a bi a y objec s
in 3D o g asping, in eal- ime, using li e- eed came a moni o ing. This would ul ima ely b ing
us a s ep close o seamless Human-Robo Collabo a ion.
4In oduc ion
1.4 Objec i es
Gi en he p oblem de ined abo e, he pu pose o his hesis is o explo e and de elop an in elligen
and ully au oma ed sys em o eal- ime objec ecogni ion and localiza ion de i ed om he
analysis o da a om a s e eo ision came a.
The de ec ion/ ecogni ion o a ious objec s and hei localiza ion wi h 6-DOF (posi ion and
o ien a ion in 3D space) will be e ie ed o hen pe o m g asping asks acco ding o ha da a.
This sys em should be able no only o ca ego ize and accu a ely loca e mul iple objec s bu also
o be obus o occlusions.
Finally, wi h his solu ion, a highe -le el obo g asping is hoped o be achie ed, imp o ing
lexible and au onomous obo ic solu ions. In sum, he ollowing objec i es a e expec ed o be
achie ed by he end o he hesis:
1. Ca y ou ex ensi e esea ch o he cu en s a e-o - he-a app oaches o objec de ec ion
and ecogni ion, as well as pose es ima ion;
2. Gain an unde s anding o he e olu ion o models in his ield;
3. C ea e a sys em, based on ML, capable o de ec ing and classi ying an objec and es ima ing
i s pose wi h 6-DOF;
4. The sys em needs o be ully au oma ed;
5. The sys em needs o be able o p ocess da a in eal- ime;
6. The ML model should be able o handle si ua ions o occlusion;
7. A obo ic manipula o should success ully and app op ia ely pe o m manipula ion asks o
a ious objec s when ecei ing inpu s om he sys em;
1.5 Thesis S uc u e
The ollowing ex is composed o ou chap e s: Li e a u e Re iew (chap e 2), 3D Objec Recog-
ni ion and Localiza ion Sys em (chap e 3), Expe imen s and Resul s (chap e 4) and, inally, Con-
clusions and Fu u e Wo k (chap e 5). Each chap e , as well as i s sec ions a e explained in his
sec ion.
In chap e 2, he main goal is o ca y ou ex ensi e esea ch o he ield in ques ion. In
his chap e , we will i s b ie ly desc ibe Machine Lea ning and i s ca ego ies, ollowed by a
mo e de ailed look in o A i icial Neu al Ne wo ks and some o i s a chi ec u es used in Compu e
Vision. A sho o e iew o he de ini ion and cha ac e is ics o he objec ’s pose and o a ion will
also be gi en.
In chap e 3, i s , an o e iew o he comple e sys em o Au oma ic Objec Recogni ion and
Localiza ion o obo ic g asping is gi en, as well as a look a i s da a low. A mo e in-dep h
1.5 Thesis S uc u e 5
desc ip ion is p o ided o he implemen a ion o his sys em, gi ing de ail no only on each ma-
chine lea ning a chi ec u e bu also on each module and he sys em’s ope a ing modes as a S a e
Machine.
In chap e 4, he a ious expe imen s and esul s a e p esen ed, as well as a discussion o hose
esul s and wha hey mean o he limi a ions o he sys em de eloped.
In chap e 5, he conclusions a e made and a ew ideas o u u e wo k a e also p esen ed.
6In oduc ion
Chap e 2
Li e a u e Re iew
In his chap e , we will i s b ie ly desc ibe Machine Lea ning and i s ca ego ies, ollowed by a
mo e de ailed look in o A i icial Neu al Ne wo ks and some o i s a chi ec u es used in Compu e
Vision. A sho o e iew o he de ini ion and cha ac e is ics o he objec ’s pose and o a ion will
also be gi en.
Finally, he ela ed wo k ega ding objec de ec ion and pose es ima ion is p esen ed, de ailing
commonly used da ase s, as well as Machine Lea ning me hods o bo h objec de ec ion and
iden i ica ion and pose es ima ion.
2.1 Machine Lea ning
Machine Lea ning is a subse o A i icial In elligence, and i deals wi h he abili y o compu e
sys ems o ecognize pa e ns and in e om hem. This means ha , ins ead o asks ha ing o be
di ec ly p og ammed, a ML model can be ained on sample da a and make decisions based on he
gene aliza ions i makes om ha da a.
ML is o g ea impo ance in au oma ion, enabling sys ems o pe o m bo h ou ine and com-
plex asks (such as analyzing la ge se s o da a), wi h he added ad an age o making he sys em
lexible o changes. I has p o en o be capable o sol ing a a ie y o asks, om image de ec ion
o speech ecogni ion, ha can some imes be conside ed ha d o humans.
This b anch o AI can be di ided in o h ee majo ca ego ies: 1) Unsupe ised lea ning, 2)
Supe ised lea ning, and 3) Rein o cemen lea ning. O hose, only he i s wo will be e iewed,
as hey a e he mos ele an o he opic in ques ion[2].
In Unsupe ised lea ning, he da a ed in o he model is no labeled. This means ha he ML
model ies o ind he pa e ns o cha ac e is ic ea u es in he da a i ecei es, c ea ing a so o
summa y wi hou being o e ed he eedback ha labeled aining da a would o e [2].
Me hods om his ca ego y gene ally can no be di ec ly applied o classi ica ion o eg ession
p oblems, gi en ha he e is no p ecise knowledge o wha he ou pu da a migh be. Basically,
his ca ego y exploi s he unlabeled da a a iance and sepa abili y o e alua e ea u e ele ance.
Some o he mos known applica ions o unsupe ised lea ning include:
7

8Li e a u e Re iew
•Clus e ing, which allows spli ing he da ase au oma ically in o g oups ("clus e s") acco d-
ing o simila i y, o en ailing o ea da a poin s as indi iduals
•Anomaly de ec ion, whe e he goal is o de ec a ypical da a poin s in he da ase au oma i-
cally;
•Associa ion mining, whe e se s o i ems ha occu oge he o en, a e iden i ied;
•La en a iable models ha a e usually u ilized as me hods o p e-p ocessing da a (such as
dimensionali y educ ion) o acili a ing da a isualiza ion ools.
The main algo i hms in Unsupe ised Lea ning include K-means and P incipal Componen
Analysis (PCA).
On he o he hand, wi h Supe ised lea ning, he models ecei e labeled da a, e e ed o as
" aining da a". This means ha he inpu consis s o he "objec " o be analyzed and he desi ed
ou pu alue. This ype o lea ning is usually done in he con ex o classi ica ion (map inpu da a
o ou pu class) o eg ession (map inpu da a o con inuous ou pu ).
Usual algo i hms in supe ised lea ning include Logis ic Reg ession, Suppo Vec o Ma-
chines, o A i icial Neu al Ne wo ks. The objec i e is he same in classi ica ion and eg ession:
o ind pa icula ela ionships in he inpu da a ha enable he model o compu e he ou pu co -
ec ly. In o he wo ds, he labels a e used o calcula e he co ec mapping unc ions o he inpu s
ecei ed, which will enable he model o classi y new unlabeled inpu da a [3].
When pe o ming supe ised lea ning, model complexi y should be conside ed. The complex-
i y o he model usually is dependen on he na u e o he aining da a. I he da ase is small, o
i he da a be ween classes is clea ly dis inc , one should choose a low-complexi y model (wi h
a high-complexi y model, i would likely o e - i , i.e., i would no be able o gene alize o o he
da a poin s).
The p oblem is ha labeled da a is ha d o ob ain since human anno a ion is bo ing. Labeling
may equi e expe s o e en special de ices and can be e y ime-consuming. Gi en his, a possible
solu ion is using semi-supe ised lea ning.
This ype o lea ning is di e en om he me hods men ioned abo e since he inpu da a is
a mix o labeled and unlabeled da a. Basically, i is an in e media e be ween supe ised and
unsupe ised lea ning.
The s anda d p ocess in ol es i s clus e ing simila da a using an unsupe ised lea ning al-
go i hm, and hen, using he exis ing labeled da a, he unlabeled da a is labeled. The use o his
me hod depends mainly on he applica ion and he se ing [3].
Wi h his b ie explana ion o he majo ca ego ies o ML and hei cha ac e is ics, a mo e
in-dep h look a A i icial Neu al Ne wo ks, a e sa ile and widely used ML model, e y common
in he opic o CV and objec de ec ion, will be gi en.
2.1 Machine Lea ning 9
2.1.1 A i icial Neu al Ne wo ks
A sho in oduc ion o a i icial neu al ne wo ks and i s cha ac e is ics a e p o ided he e in b ie .
This ype o model is one o he mos p e alen in ML, due o he ac ha ANNs a e some o
he mos e sa ile and e ec i e algo i hms, and, as a esul , hey a e one o he mos signi ican
subjec s in he ield o AI.
Mos ANNs ha e, a hei basic le el, a i icial neu ons called "pe cep ons" (Fig. 2.1).
Figu e 2.1: Rep esen a ion o a pe cep on. Image adap ed om [4]
A pe cep on (o neu on) ecei es a ious inpu s (xi) mul iplied by i s co esponding weigh s
(wi, numbe s exp essing he impo ance o he espec i e inpu s o he ou pu ) and akes he sum
o hese esul s o p oduce a single ac i a ion (Eq. 2.1). The e can also be an addi ional e m,
e e ed o as bias (b), ha ansla es how easy i is o ge he pe cep on o i e.
a=
n
∑
i=1
xi×wi+b(2.1)
The pa ame e ain Eq. 2.1 is known as ac i a ion. The neu on’s ou pu (y) esul s in he
ac i a ion o he pe cep on and is dependen on he alue aand he ac i a ion unc ion used.
The ac i a ion unc ion is used o in oduce non-linea i y in o he ou pu o he pe cep on. This
unc ion is chosen acco dingly o he na u e o he da a and he dis ibu ion o a ge a iables.
Each ype o ac i a ion unc ion en ails a di e en ixed ma hema ical calcula ion.
A adi ional A i icial Neu al Ne wo k is made up o in e connec ed laye s o pe cep ons
(neu ons). The i s laye is known as he inpu laye , he las as he ou pu laye and he middle
laye s a e e e ed o as hidden laye s.
The inpu laye usually is equal o he numbe o di e en labels in a da ase , while he numbe
o hidden laye s can ange om one o hund eds and a ies acco ding o he numbe o labels and
compu a ion capabili y.
As aining da a is ed h ough he ANN, he e o be ween he ac ual and he desi ed ou pu s
o he ne wo k a e calcula ed using a cos /loss unc ion; his e o is hen used o adjus he weigh s
and biases o he ANN. How hese adjus men s a e made depends on he op imiza ion algo i hm
used, and on he pa ame e s o he ne wo k ha de ine he a ious ypes o a chi ec u es.
10 Li e a u e Re iew
Mo eo e , he connec ions be ween neu ons can also be bi-di ec ional, leading o o he di -
e en a chi ec u es. All o his con ibu es o he as amoun o ANNs in exis ence and he
complexi y o his subjec [5]. In he nex subsec ions, some o he mos used ANN a chi ec u es
in CV will be p esen ed b ie ly.
2.1.1.1 Au oencode s
An Au oencode is an unsupe ised ANN, ha i s lea ns o comp ess and encode da a, and hen
lea ns o econs uc back he da a om he encoded e sion. In a nu shell, an au oencode educes
da a dimensions by lea ning how o igno e noise in he da a. Fo his o happen au oencode s ha e
ou main pa s: encode ;bo lleneck,decode and econs uc ion loss.
An example o an au oencode a chi ec u e is p esen ed in Fig. 2.2, whe e ˆxis he econs uc-
ion o he o iginal inpu x. The encode akes i s inpu and ou pu s a ec o wi h i s encoded
in o ma ion, comp essing and educing he inpu da a dimensions. The bo leneck is a laye in
which he inpu ’s comp ession ep esen a ion is con ained ( he lowes possible dimensions o he
inpu ).
The decode ’s job is o hen ake his comp essed ec o and ou pu he closes ma ch o he
o iginal inpu . The econs uc ion loss unc ion is hen minimized, which helps bo h he encode
and decode lea n wi hou he need o labeled da a. This loss measu es how well he decode
pe o ms by calcula ing how close he ou pu is o he o iginal da a.
Howe e , i an encode lea ns o ma ch he o iginal inpu exac ly, he e is no use ulness o he
ne wo k. Fo ha eason, es ic ions a e applied o limi how much o he inpu can be copied.
This esul s in a unc ion ha p io i izes and only ou pu s he mos impo an cha ac e is ics o i s
inpu [6].
Figu e 2.2: Gene al a chi ec u e o an Au oencode .
2.1 Machine Lea ning 11
2.1.1.2 Con olu ional Neu al Ne wo ks
Ano he ype o ANN is he Con olu ional Neu al Ne wo k (CNN). Since, in a adi ional neu al
ne wo k, all laye s a e ully connec ed, hey a e no able o ake in o accoun he image’s spa ial
na u e (s uc u e and pa e n). In o he wo ds, all inpu pixels in he image would be ea ed he
same way, independen ly o being a apa o close oge he .
This is whe e CNNs a e mos use ul. They a e a class o deep neu al ne wo ks (an ANN wi h
mul iple hidden laye s) op imized o da a in g id o m, which makes i especially use ul o image
p ocessing. Fo his eason, CNNs ha e e ol ed o be he mos used me hod o sol ing a ious
compu e ision p oblems [7] and will appea in he majo i y o Sec ion 2.3.
CNNs a e iden ical o egula NN in e ms o being composed by neu ons, whe e he goal
is o es ima e weigh s and bias. Howe e , CNNs ha e a much mo e e icien s uc u e o image
p ocessing han NNs, gi en ha each neu on is only connec ed o a speci ic egion o he p e ious
laye . The gene al a chi ec u e o a NN and CNN a e p esen ed in Fig. 2.3 o compa ison. Each
laye o a CNN is ep esen ed as 3D olume [8].
Figu e 2.3: NN a chi ec u e VS CNN a chi ec u e. A s anda d NN is ep esen ed on he le , and
a CNN on he igh . Each ci cle ep esen s a neu on [8].
This class o ne wo k uses a ian s o he ma hema ical ope a ion con olu ion. Con olu ion
(deno ed as ∗) can be seen as he ma ix mul iplica ion o an inpu Iby a ke nel K(as seen in Eq.
2.2). Using a ke nel wi h a size signi ican ly in e io o he inpu size, a CNN can ex ac small
ea u es o an image. This also esul s in ewe calcula ions and, he e o e, an inc ease in speed.
C(i,j) = (I∗K)(i,j) =
m
∑
i=1
n
∑
i=1
I(i−m)(j−n)K(m,n)(2.2)
Typically, he laye s o hese ne wo ks a e made up o h ee s ages: con olu ion s age, de ec o
s age, and pooling s age. In he con olu ion s age, mul iple pa allel con olu ions a e pe o med
and esul in a collec ion o linea ac i a ions. A e ha , in he de ec o s age, a non-linea ac i a-
ion unc ion is used on he linea ac i a ions.
The p oblem a ises due o he ac ha hese laye s encode he in o ma ion in he p ecise posi-
ion i is in, meaning ha i hese laye s p ocess he same inpu wi h a sligh change, i will esul
in a di e en ou pu . This is whe e pooling unc ions a e use ul, making he ne wo k esis an , o
in a ian , o small changes in he inpu .
18 Li e a u e Re iew
•Using bo h VOC and COCO aining da a leads o inc eased accu acy: 83.6% and 82% s.
80.5% and 77.6%;
•Using he i s 101 laye s o ResNe esul s in he bes accu acy: sa u a ion occu s a 80.5%;
•The use o a RPN achie es supe io esul s han o he al e na i es: 79.5% s. 77.8% ( o
he closes al e na i e).
The a chi ec u e o he R-FCN emo es he expensi e compu a ion ha app oaches such as
Fas e R-CNN use on each ROI (by passing each ROI h ough CNNs), in a o o a much mo e
s aigh o wa d calcula ion o o e lap, enabling he model o be as e bu s ill main ain he accu-
acy o e ed by a egion-based design.
2.3.2.3 SSD: Single Sho Mul ibox De ec o
SSD (Fig. 2.7) [23] is one o he single-sho app oaches o objec de ec ion wi h he in en o being
as and accu a e enough o be used in eal- ime applica ions. This de ec o emo es he need o
egion p oposals, while s ill main aining accu acy.
Figu e 2.7: O e iew o he SSD a chi ec u e. [23].
The i s s age in SSD is he gene a ion o ea u e maps a mul iple scales, which allows o
he de ec ion o objec s o di e en sizes - ea u e maps wi h highe esolu ion a e able o de ec
smalle objec s, while ones wi h lowe esolu ion can de ec bigge objec s. The backbone a chi-
ec u e used is based on VGG, and o he con olu ional laye s a e appended a he end o allow o
di e en scales.
In he second s age, de ec ion is based on small con olu ional ke nels applied o each cell in he
ea u e maps. Fo e e y cell, ou de aul bounding-boxes o di e en aspec a ios a e applied.
These aspec a ios a e chosen o accommoda e objec s o di e en shapes and sizes (Eq. 4 in
[23]). Fo each box, a class p obabili y sco e is compu ed, as well as di e en o se s o he size
o he box (in o de o be e adjus he shape o he box o he shape o he objec ). This app oach
is simila o ha o Fas e R-CNN bu applied o di e en esolu ions o ea u e maps.
A pa icula ly in e es ing aspec in [22] is he use o "da a augmen a ion" o aining. This
means ha , o each aining image, ei he he en i e image is used o a andom pa ch is ob ained

2.3 Rela ed Wo k 19
om i . In he case o a andom pa ch, i can also su e image dis o ions. The objec i e o da a
augmen a ion is o enable he model o handle a ious shapes and sizes o objec s mo e obus ly.
The model was e alua ed on he Pascal VOC and MS COCO da ase s, whe e he main conclu-
sions we e he ollowing:
•As wi h p e ious models, aining on bo h he VOC and COCO da ase s inc eases accu acy
(Tables 1, 4 and 5 in [22]);
•Using mul iple ou pu s laye s o mul iple ea u e map esolu ions esul s in highe accu acy
(Table 3 in [22]);
•Da a augmen a ion achie es be e esul s, as expec ed Table 6 in [22]);
•SSD can achie e highe mAP han i s compe i o s using a smalle esolu ion inpu image
Table 7 in [22]).
2.3.2.4 YOLO: You Only Look Once
YOLO was i s in oduced in [24], su e ing subsequen al e a ions in [25] and some mino
changes in [26]. I was in oduced as a single-sho app oach o objec de ec ion, by using a single
ne wo k o sol e a eg ession p oblem o p edic bounding boxes and class p obabili ies o objec s
in an image.
The i s e sion o YOLO di ides i s inpu image in o a 7x7 g id, and each cell in he g id
is esponsible o iden i ying an objec i i s cen e is loca ed wi hin he cell. Each cell p edic s
a se o 2 bounding-boxes wi h associa ed con idence sco es, as well as class p obabili ies. The
con idence sco es ake in o accoun he p obabili y o he boxes ha ing an objec wi hin hem and
also he accu acy o he box in e ms o he shape o he objec .
One p oblem ha can a ise om his app oach is he same objec being de ec ed by mo e han
one cell. In his case, he cell wi h he highes con idence o a speci ic objec is chosen. The
p edic ions o each cell a e made by classi ie s applied a e ea u e ex ac ion (which is done on
he en i e image by he ne wo k in Fig. 2.8).
Figu e 2.8: O e iew o he o iginal YOLO ne wo k. [24].
20 Li e a u e Re iew
In he second e sion o YOLO [25], bounding-boxes we e eplaced by he same ancho boxes
used in Fas e R-CNN, since i made lea ning easie o he ne wo k and i also enabled each cell
o p edic mo e han one objec . This had a small dec ease in accu acy, bu o e all ecall imp o ed.
The numbe o ancho boxes chosen o his app oach was 5.
Ano he change was eplacing he o iginal backbone ne wo k (Fig. 2.8) wi h a cus om ne wo k
called Da kne -19, whose pu pose was o educe complexi y and imp o e accu acy. A signi ican
imp o emen in his e sion was ha his model was able o de ec mo e han 9000 objec ca e-
go ies, compa ed o he o iginal 20.
The hi d e sion [26] came wi h mino changes, he mos impo an o which was he back-
bone CNN being expanded o Da kne -53, a CNN wi h 53 con olu ional laye s ha is mo e pow-
e ul han i s p edecesso , bu mo e e ec i e han o he used a ian s o ResNe . The mos ecen
es esul s o YOLO we e p esen ed in [26]. They a e he esul s o es ing on he COCO da ase ,
and he ollowing conclusions we e made:
•When compa ing mAP o he p ocessing speed (Fig. 3 in [26]), YOLO achie es he bes
esul s;
•In e ms o accu acy, YOLO is mo e accu a e han SSD, bu s ill in e io o wo-s age ap-
p oaches, such as Fas e R-CNN.
2.3.2.5 Compa ison o Resul s
The ollowing ables ea u e he esul s o mean a e age p ecision (in pe cen age) o each o he
me hods in each da ase . The alues a e aken om he pape s ha in oduce each me hod. When
a pape does no p esen esul s o a pa icula da ase , he alue is omi ed.
Tables 2.1 and 2.2 p esen he esul s o he VOC07 and VOC12 da ase s, espec i ely.
Table 2.1: Compa ison be ween me hods on he Pascal VOC07 da ase
Au ho /Re e ence Yea Me hod mAP (%)
Ren e al. [20] 2017 Fas e R-CNN 78.8
Dai e al. [22] 2016 R-FCN 83.6
Liu e al. [23] 2016 SSD 81.6
Redmon and Fa hadi [26] 2018 YOLO −
Table 2.2: Compa ison be ween me hods on he Pascal VOC12 da ase
Au ho /Re e ence Yea Me hod mAP (%)
Ren e al. [20] 2017 Fas e R-CNN 75.9
Dai e al. [22] 2016 R-FCN 82.0
Liu e al. [23] 2016 SSD 80.0
Redmon and Fa hadi [26] 2018 YOLO −
2.3 Rela ed Wo k 21
Table 2.3 p esen s he esul s o he MS COCO da ase . On his da ase , mAP can be di ided
in o wo me ics. The i s is he s anda d mAP when IoU be ween he objec bounding-box and
he g ound- u h is 0.5. The second, is a mo e s ic me ic, o he mean a e age p ecision o e
a ious h esholds, o an IoU be ween 0.5 and 0.95.
Table 2.3: Compa ison be ween me hods on he MS COCO da ase
Au ho /Re e ence Yea Me hod [email p o ec ed] (%) mAP@[.5,.95] (%)
Ren e al. [20] 2017 Fas e R-CNN 42.7 21.9
Dai e al. [22] 2016 R-FCN 53.2 31.5
Liu e al. [23] 2016 SSD 46.5 26.8
Redmon and Fa hadi [26] 2018 YOLO 57.9 33.0
As can be seen, R-FCN and SSD a e he mos accu a e me hods on he VOC da ase s. SSD is
he as es o he wo, being a one-sho app oach. On he MS COCO da ase , YOLO achie es he
bes esul s. When IoU is be ween 0.5 and 0.95, he e is a signi ican d op in accu acy, since i is
di icul o ob ain a high IoU.
2.3.3 Machine Lea ning in Pose Es ima ion
2.3.3.1 SSD-6D
The objec i e o his ne wo k (Fig. 2.9), in oduced in [27], is o build upon he s uc u e o SSD
o ob ain a collec ion o ou pu s: class p obabili ies, coo dina es o a 2D bounding-box, sco es o
possible iewpoin s and in-plane o a ions.
Figu e 2.9: O e iew o he SSD-6D model. [27].
The inpu image is passed h ough a backbone CNN based on Incep ionV4 o ou pu a se o
ea u e maps a di e en scales. The same s uc u e as SSD is used o p edic ion o class and
22 Li e a u e Re iew
eg ession o he bounding-box. The main changes occu in he sco ing o possible iewpoin s
and in-plane o a ions.
As is shown in Fig. 2.9, each map is con ol ed wi h a ke nel o shape (4 + C + V + R),
whe e C ep esen s he numbe o objec classes, V he numbe o iewpoin s and R he numbe
o in-plane o a ions. The iewpoin is he posi ion in 3D space om which he objec is iewed,
which in luences he aspec o he objec as seen om he came a, while in-plane o a ion is seen
as a ans o ma ion o he same iewpoin .
As can be seen in Fig. 2.10, he iewpoin s a e sampled om a hal -sphe e a ound he objec .
Fo symme ical objec s, only he a c in g een is sampled, and, o semi-symme ical objec s, only
he poin s in ed a e used. In [27], i is a gued ha con olu ional laye s a e mo e e ec i e a
sco ing a iewpoin and in-plane o a ion han using eg ession o p edic a se o ansla ions and
o a ions. Fu he mo e, by sco ing iewpoin s wi h a con idence alue, all ha is abo e a ce ain
h eshold can be accep ed, he eby dealing well wi h symme ical objec s.
Figu e 2.10: Illus a ion o e e y iewpoin conside ed. [27].
Knowing he pa ame e s o he came a, he mos con iden sco es can be pooled o calcula e
he ansla ion and he o a ion o he objec , esul ing in a se o 6-DOF pose hypo heses. Each
hypo hesis goes h ough a pose e inemen s ep, which uses he I e a i e Closes Poin (ICP) al-
go i hm (2.3.3.7). The bes pose is chosen by compa ing he e ined poses o he dep h da a om
RGB-D.
SSD-6D was e alua ed on bo h he Linemod and he Tejani da ase s. Howe e , since none o
he o he me hods a e e alua ed on he las da ase , he e a e no compa isons o be made. Tejani
da ase esul s will, he e o e, be omi ed.
2.3.3.2 PoseCNN
PoseCNN [19] e olu ionized pose es ima ion algo i hms by decoupling pose es ima ion in o h ee
sepa a e asks: seman ic labeling, 3D ansla ion es ima ion and 3D o a ion eg ession (as can be
seen in Fig. 2.11). This app oach acili a es he job o he ne wo k by enabling i o model how
each ask ela es o o he s. The backbone CNN akes an inpu image and ou pu s ea u e maps
o di e en scales. The b anches co esponding o each ask hen use hese ea u e maps o
calcula ions.
2.3 Rela ed Wo k 23
Figu e 2.11: O e iew o he PoseCNN sys em. [19].
In he seman ic labeling ask, wo 512 channel maps a e p ocessed in o de o ob ain one o he
same scale as he o iginal image. He e, objec de ec ion is done by applying con olu ional ke nels
o each pixel in he image and a ibu ing a seman ic label: each pixel is classi ied in o an objec
class. This gi es be e esul s han he bounding-box app oaches used in SSD-6D, o example,
since i handles occlusions be e .
Fo 3D ansla ion es ima ion, he ne wo k needs o ou pu a ec o o coo dina es (Tx,Ty,Tz)
o he objec cen e (in he came a coo dina e sys em). To ha ex en , each pixel ha belongs o
an objec needs o o e o he cen e o he objec in he image coo dina e sys em, (cx,cy). A e
ha , Txand Tycan be de i ed using Eq. 1 in [19]. Vo ing is implemen ed by a Hough o ing laye
in he ne wo k and is done as ollows:
1. A eg ession ne wo k p edic s an a ay (nx,ny,Tz) o each pixel;
2. Each pixel o es o o he pixels (on whe he hey a e an objec cen e ) along he di ec ion
o he ec o de ined by (nx,ny);
3. The pixel wi h he mos o es is chosen as he objec cen e .
The mean o all alues o Tz o each pixel in an objec is chosen as he ue alue o Tz.
Addi ionally, all pixels inside an objec a e known as inlie s. The bounding-box ha con ains all
inlie s is also gene a ed in his s ep and is used o he 3D o a ion ask.
In his ask, he objec i e is o ou pu a qua e nion, which ep esen s he es ima ed o a ion.
The i s s ep is o apply 2 ROI pooling laye s using he bounding-boxes gene a ed in he p e ious
ask. In ha way, he ea u e maps o each ROI can be ob ained. These ea u e maps a e hen
added and ed in o 3 FC laye s, he las o which ou pu s a qua e nion.
PoseCNN is e alua ed on he YCB-Video and he Occlusion LineMod da ase s. The ollowing
conclusions can be made om ables 2 and 3 in [19]:
•This model ou pe o ms coo dina e eg ession algo i hms;
•Resul s a e a supe io when using pose e inemen ;
•On he Occlusion LineMod da ase , PoseCNN wi h a pose e inemen algo i hm (ICP)
achie es highe accu acy han o he me hods;

24 Li e a u e Re iew
2.3.3.3 DPOD: Dense Pose Objec De ec o
DPOD [28] akes a di e en app oach o o he s men ioned p e iously. I s model (Fig. 2.12) is
comp ised o h ee blocks: co espondence block, pose block, and a inal pose e inemen block.
The co espondence block akes as inpu he desi ed image and uses an encode (based on he i s
12 laye s o ResNe ) and h ee decode s ha ou pu a co espondence map ( i s wo decode s) and
he ID masks o each objec ( hi d decode ).
Figu e 2.12: O e iew o he DPOD sys em. [28].
A co espondence map is a 2-channel image wi h alues om 0-255 ha maps a pixel in
he image o a e ex in he objec 3D model. This esul s in mo e s aigh o wa d aining o he
ne wo k and an o e all inc ease in quali y since i jus needs o sol e a colo classi ica ion p oblem
o ma ch he image co espondence map o he 3D model co espondence map, ins ead o ha ing
o eg ess he coo dina es o he objec .
The ID masks a e a esul o he p obabili y ha an objec pixel belongs o a speci ic class
(such as p e ious app oaches ha e used). Using he ID masks, he 3D model o each class can
be ob ained. Bo h he co espondence map and he 3D models a e ed in o he pose block, which
uses PnP+RANSAC o ou pu a o a ion ma ix Rand a ansla ion ec o T.
As opposed o o he o ms o pose e inemen (2.3.3.7), he DPOD model u ilizes a pose
e inemen block (Fig. 2.13) based on he ResNe a chi ec u e. This block akes in a pa ch o he
o iginal image con aining he objec and a 3D ende ing o he objec in he es ima ed pose. These
inpu s a e ed sepa a ely h ough wo b anches composed o he i s i e laye s o ResNe (E11
and E12), and he ou pu s a e sub ac ed and ed in o ano he ResNe -like ne wo k (E2).
The di e ence ea u e ec o ha is ou pu ed is used o compu e he e o be ween he p e-
dic ed pose and he ac ual pose. This is done by eeding i in o h ee sepa a e eg ession laye s
(XY head, Z head and R head) ha co ec he coo dina es o ansla ion and he o a ion ma ix
(as is seen in Fig. 2.13).
2.3 Rela ed Wo k 25
Figu e 2.13: Illus a ion o he pose e inemen block. [28].
The model was e alua ed on he LineMod and he Occlusion LineMod da ase s, ha ing ob-
ained he ollowing conclusions:
•Compa a i ely, DPOD achie es s a e-o - he-a esul s, wi h only PVNe (explained in Sec-
ion 2.3.3.4) ha ing be e accu acy in pose es ima ion (Table 1 in [28]);
•Pose e inemen imp o es accu acy by almos 10% (Table 1 in [28]);
•DPOD’s no el pose e inemen echnique achie es be e esul s han DeepIM (Tables 1 and
5 in [28]).
2.3.3.4 PVNe : Pixel-wise Vo ing Ne wo k
PVNe looks o imp o e upon he wo-s age app oach o keypoin de ec ion and pose es ima ion
by using a model mo e obus o occlusion and unca ion. The main idea is o p edic , o each
pixel in an objec , he uni ec o s om ha pixel o all he keypoin s in he objec - simila o wha
was done in PoseCNN o he localiza ion o an objec cen e . This app oach is mo e obus o
si ua ions whe e he keypoin s a e no isible in he image since i s posi ion can be in e ed om
he o he isible pixels.
In he i s s age o PVNe , a backbone CNN based on ResNe is used o ob ain class p ob-
abili ies o each pixel (seman ic segmen a ion) and he uni ec o s ha ep esen he di ec ion
om ha pixel o e e y keypoin . The o ing o each pixel is cha ac e ized in Eq. 1 and 2 o
[29]. To educe a iance in localiza ion, he model needs o choose he keypoin s loca ed on he
su ace o he objec . The ollowing algo i hm (Fa hes Poin Sampling) is used o choose a se o
8 keypoin s:
1. Add he cen e o he objec o he se o keypoin s;
2. Choose he a hes keypoin om he se and add i o he se ;
3. Repea he second s ep un il eigh keypoin s a e chosen.
Finally, he model needs o compu e he 6-DOF pose using he se o keypoin s. This is done
using an al e ed e sion o he PnP algo i hm ha akes in o accoun he mean and co a iance
(unce ain y) associa ed wi h each keypoin (which a e calcula ed wi h Eq. 3 and 4 o [29]).
26 Li e a u e Re iew
This me hod is e alua ed on he LineMod, Occlusion LineMod, and YCB-Video da ase s. I
is concluded ha PVNe achie es he bes esul s in e ms o he ADD(-S) me ic when compa ed
o o he app oaches (Table 3, 5, and 7 in [29]).
2.3.3.5 DenseFusion
Mos app oaches so a ha e used 2D ea u es om RGB images and used dep h da a only o pose
e inemen . DenseFusion’s [30] model (Fig. 2.14) akes ad an age o dep h in o ma ion om he
i s s age by using RGB da a and poin cloud alues on a pe -pixel basis. Since bo h RGB and
poin cloud da a a e e y di e en da a ypes, his model ea s each one sepa a ely and uses a
pixel-wise dense usion algo i hm o combine hem.
Figu e 2.14: O e iew o he DenseFusion a chi ec u e. [30].
In he i s s age, he same seman ic segmen a ion algo i hm applied in PoseCNN is used o
ob ain segmen a ion masks o each objec . A bounding-box ha includes each mask is calcula ed
and is used o ob ain an image c op in i s loca ion. The dep h da a o pixels inside he bounding
box is also used in o de o con e i in o a 3D poin cloud ep esen a ion.
The c opped image is ed in o an au oencode ne wo k (based on ResNe ) ha ou pu s colo
in o ma ion. On he o he hand, poin cloud da a is p ocessed by a Poin Ne -like s uc u e o ob ain
he geome ic in o ma ion o he objec .
The usion algo i hm used combines colo (colo embeddings) and geome ic in o ma ion (ge-
ome y embeddings) in o ea u e ec o s o each pixel in he objec . E e y ea u e ec o is hen
p ocessed by a MLP wi h a e age pooling o ob ain a global ea u e ec o , which is hen appended
o each pe -pixel ec o , c ea ing a pixel-wise ea u e (as shown in Fig. 2.14). The pixel-wise ea-
u e is ed in o he pose p edic o ne wo k which ou pu s an es ima ion o he pose o e e y pixel
wi h an associa ed con idence alue: Ri, i, and ci.
2.3 Rela ed Wo k 27
Figu e 2.15: Illus a ion o he pose e inemen ne wo k. [30].
DenseFusion also p esen s an i e a i e app oach o pose e inemen (Fig. 2.15) u ilizing a
co ec ion ne wo k ha is composed o 4 FC laye s. Howe e , his ne wo k needs o lea n o
co ec he pose ins ead o p edic ing a new one. The e o e, i s inpu s a e he colo embeddings o
he image and he geome y embeddings o a ans o med poin cloud.
This ans o med poin cloud is he esul o al e ing he inpu poin cloud by he ac o s ∆R and
∆ p edic ed by he pose esidual es ima o . This can be done i e a i ely, e ining he es ima ion
u he and u he in each s ep. The inal pose p edic ed is ob ained by conca ena ing e e y pose
ob ained in each i e a ion.
E alua ion o he model is done on he LineMod and YCB-Video da ase s. The main conclu-
sions ob ained in [30] a e he ollowing:
1. Wi h e inemen , DenseFusion achie es he highes accu acy when compa ed o he second
bes me hod, PoseCNN+DeepIM ( ables 1 and 2 in [30]);
2. Pose e inemen inc eases pose accu acy by 4-8% ( ables 1 and 2 in [30]);
3. In Table 3 o [30], i can be seen ha DenseFusion is close o 200x as e han PoseCNN
plus ICP;
A pa icula ly ele an aspec ( o his hesis) in [30] is he use o DenseFusion in a obo ic
g asping si ua ion. Ou o 60 a emp s, he obo has a 73% success a e while using DenseFusion
as a pose p edic o .
2.3.3.6 Hyb idPose
Mos app oaches ha e only used only one ype o in e media e ep esen a ion o an objec , he
mos popula o which is a keypoin ep esen a ion. In [31], he au ho s a gue ha , in a eal-
wo ld se ing, i is di icul o p edic keypoin s accu a ely based on a RGB image. Hyb idPose
(Fig. 2.16) looks o sol e his p oblem by adding mo e in e media e ep esen a ions, enabling i s
model o wo k be e in occlusion si ua ions and ha ing mo e in o ma ion abou he geome ical
cha ac e is ics o he objec .
34 3D Objec Recogni ion and Localiza ion Sys em
The Da a Acquisi ion Module has wo main unc ions. The i s unc ion deals wi h da a
e ie al om he ZED came a; i uns on an in ini e loop, con inuously e ie ing he RGB image
ame and he dep h map. The o he unc ion calcula es di e ences be ween consecu i e ames,
using he S uc u al Simila i y (SSIM) algo i hm. This algo i hm ou pu s a alue, om 0 o 1, ha
indica es how simila wo images a e, and he unc ion compa es his alue o a h eshold alue
(calcula ions o his alue will be shown in Sec ion 4.3.1), in o de o indica e i he e we e changes
be ween ames (such as an objec mo ing). In case o changes, he unc ion aises a global lag
(indica ing ha a new ame is a ailable) and passes he new ame and i s co esponding dep h
map o he Compu e Vision Module.
Figu e 3.3: S a e machine o he pose es ima ion module
In he Compu e Vision Module, a s a e machine (Fig. 3.3) is implemen ed. The e is one
Ini ial s a e and ou main s a es:
•De ec ion: implemen s he Objec De ec ion and Segmen a ion sys em explained in 3.2.2.1;
•Localiza ion: implemen s he 6D Objec Pose Es ima ion sys em explained in 3.2.2.2;
•Ready o G asp: chooses which objec o g asp and sends coo dina es o obo ic manipula-
o ;

3.2 Sys em’s Modules 35
•G asping: wai s o g asping o inish and ies o de ec changes in ame.
The ansi ion be ween s a es depends on he changes be ween ames ha a e de ec ed in module
1. I any objec in he ame mo es o disappea s, he whole s a e machine has o ese o he
De ec ion s a e.
3.2 Sys em’s Modules
3.2.1 Da a Acquisi ion Module
In he Da a Acquisi ion Module, he goal is o acqui e usable da a om he g asping wo kspace
o he Compu e Vision module o p ocess and ex ac in o ma ion needed o g asping asks.
In his case, he in o ma ion wan ed is 6D pose es ima ion and objec classi ica ion. Gi en his,
and ega ding he equi emen o a 3D based solu ion in ou app oach, a dep h/3D came a was
needed. In ou solu ion, gi en se e al epo ed uses and a ailabili y, he ZED s e eo ision came a
and i s SKD we e used.
The ZED came a was de eloped by S e eolabs and is composed o wo side-by-side RGB
came as, which ou pu synch onized RGB images. The wo wide-angle lenses ha e 110◦ ield-o -
iew, spaced a a baseline o 120 mm, allowing accu a e dep h es ima ion in he ange o 0.7 o 20
me e s. Came a pic u e, along wi h p oduc dimensions, a e shown in Fig. 3.4. The mos ele an
cha ac e is ics o he ZED came a a e summa ized in Table 3.1.
Figu e 3.4: ZED S e eo Came a and dimensions illus a ion
Fo dep h acquisi ion, he ZED came a imi a es human ision, whe e each eye has a sligh ly
di e en iew o he wo ld a ound, and by compa ing he wo iews, dep h can be in e ed. Like-
wise, his came a, wi h i s sepa a e lens, can es ima e dep h by compa ing he pixels’ displacemen
be ween he le and igh RGB images. The came a hen p o ides a dep h map (image con aining
in o ma ion ega ding he dis ance o he su aces o scene objec s (z), o e e y pixel i,j) in he
came a coo dina e sys em. This dep h is exp essed in me ic uni s and calcula ed om he le
came a’s eye back o he g asping objec .
36 3D Objec Recogni ion and Localiza ion Sys em
Table 3.1: ZED Came a ele an ea u es [1]
Size and Weigh
Dimensions:
175x30x33 mm
Weigh : 159 g
Indi idual image and dep h esolu ion
HD2K: 2208 x 1242 (15 FPS)
HD1080: 1920 x 1080 (30, 15 FPS)
HD720: 1280 x 720 (60, 30, 15 FPS)
WVGA: 672 x 376 (100, 60, 30, 15 FPS)
Dep h
Range: 1-20 m
Fo ma : 32 bi s
Baseline: 120 mm
Lens Field o View: 110◦
/2.0 ape u e
Connec i i y USB 3.0 (5 V / 380 mA)
0◦C o +45◦C
SDK Sys em Requi emen s
Windows o Linux
Dual-co e 2.3 GHz
4 GB RAM
N idia GPU
The came a also p o ides a SDK o acili a ing came a con ol, bo h in C++ and Py hon.
Fo ou implemen a ion, only he Py hon e sion was used. The ZED SDK is designed a ound
OpenCV and CUDA lib a ies, wi h i s calib a ion and dep h es ima ion ou ines exploi ing CUDA’s
pa allel GPU compu ing capabili ies. The SDK also o e s op ional dep h map p ocessing, includ-
ing occlusion illing and edge sha pening. Un o una ely, hese buil -in pos -p ocessing echniques
come wi h high compu a ional cos s, so hey could no be used in ou solu ion since eal- ime is a
c ucial equi emen .
The came a is con igu ed o un a 15 ps, wi h 720p esolu ion. The minimum dis ance o
dep h es ima ion is 0.3m, and he dep h es ima ion was se oquali y mode, meaning i ou pu s
mo e accu a e dep h alues a a small compu a ional cos .
3.2.2 Compu e Vision Module
As shown in Fig. 3.1, he Compu e Vision Module is he mos impo an in he p oposed app oach,
since i is h ough i s p ocessing capabili ies ha he objec ’s poses and ca ego ies a e iden i ied
o pe o ming g asping ope a ions. This module is composed o wo di e en bu in e connec ed
models: he Objec De ec ion and Segmen a ion model, and he Objec Pose es ima ion model. In
he ollowing Subsec ions, de ails on hei implemen a ion will be p o ided.
3.2.2.1 De ec ion and Segmen a ion Model
The objec de ec ion and segmen a ion model is composed o a p e-p ocessing s age and an in e -
ence s age. In p e-p ocessing, he da a ecei ed om he Da a Acquisi ion module is p ocessed, so
i can be ed in o he in e ence s age, which is made up o a Mask-RCNN model [34]. This s age
3.2 Sys em’s Modules 37
hen ou pu s a se o objec iden i ica ions, bounding-boxes, segmen a ion masks, and con idence
alues o each objec in he came a ame.
Figu e 3.5: O e iew o he Mask-RCNN model [34]
In he p e-p ocessing s age, h ee RGB image ames and co esponding dep h maps a e e-
cei ed by he pipeline, and an a e age o he pixel alues o bo h he RGB ames and he dep h
maps a e calcula ed. This is o help educe he impac o small illumina ion di e ences, pixel
shi ing, and andom noise when a ame is acqui ed, which in u n educes he licke ing e ec in
he objec masks. This is discussed a g ea e leng h in Sec ion 4.3.2. A e wa d, he image ame
is esized o i he inpu size o he in e ence s age.
The Mask-RCNN model (o e iewed in Fig. ) is an ex ension o he Fas e R-CNN model
2.3.2.1, by applying a small FCN each ROI and ou pu ing a segmen a ion mask. As can be seen
in he s a e-o - he-a analysis in sec ion 2.3.2.5, Fas e R-CNN is no he op-pe o ming model
when i comes o pu e objec de ec ion, so why is i he one used? The answe is ha pu e objec
de ec ion is no he only ac o in his p ojec . The inal ou pu should be he 6D pose o he objec ,
which means ha an addi ional model (in his case, he 6D pose es ima ion model discussed u he
on, in Sec ion 3.2.2.2) is needed. One o he model equi emen s is he inpu o a segmen a ion
mask o each objec and YOLO - he op-pe o ming model (in objec de ec ion), bo h in e ms
o accu acy and speed - does no ou pu segmen a ion masks.
Two di e en al e na i es we e analyzed: using YOLO wi h a segmen a ion mask ne wo k
o e head (much like how Mask-RCNN imp o es upon Fas e R-CNN) o using Mask-RCNN
ins ead. The conclusions we e ha Mask R-CNN is ac ually mo e accu a e and ma ginally as e
han he o he al e na i e; he e o e, i is ac ually a be e choice o his implemen a ion.
3.2.2.2 Pose Es ima ion Model
This sys em, as opposed o objec de ec ion and segmen a ion, wo ks on objec s one by one. I
begins wi h a p e-p ocessing s age, which p ocesses he da a ecei ed om he objec de ec ion
model o op imize he inpu s o he pose es ima ion s age. In pose es ima ion, a Dense usion
2.3.3.5 a chi ec u e is used o in e ence, and i ou pu s a se o o a ion qua e nions and ansla ion
ec o s o each objec . A pose e inemen s age is hen applied o gi e a be e pose es ima ion.
38 3D Objec Recogni ion and Localiza ion Sys em
Finally, a pos -p ocessing s age is execu ed o calcula e he inal op imal g asping poin coo dina es
and g ippe pose o e e y objec in he ame.
Due o he ZED came a limi a ions, he dep h map migh e u n alues such as in ini e and
NaN. This occu s when poin s a e oo close/ a om he came a (in he case o in ini e alues), o
when dep h canno be es ima ed due o occlusions (in he case o NaN). To make su e hese alues
do no con ibu e o poo calcula ions, a mask o he dep h map is c ea ed whe e only eal alues
a e masked.
A e he dep h map mask is ob ained, i is mul iplied by he segmen a ion mask o ob ain all
he eal dep h alues o only he poin s ha belong o he objec . Since he came a in he eal-wo ld
en i onmen is posi ioned close o he objec s han he came a in he da ase , his means mo e
poin s in he image will ep esen an objec , so choosing mo e poin s gi es us mo e geome ical
in o ma ion abou he objec . Using mo e poin s is also a good way o making ou lie s (poin s in
he mask wi h poo dep h alues) less in luen ial in he calcula ions.
A se o 2000 poin s is hen andomly chosen om he objec dep h alues. The alue o 2000
poin s, a e some es ing, was chosen because i is a good ade-o be ween accu acy and speed.
The ac ha he poin s a e chosen a andom also means ha he se o poin s will ha e a p ope
dis ibu ion h oughou he whole objec . Wi h he poin s chosen, he poin cloud alue o each
poin is calcula ed, ollowing Equa ions 3.1,3.2 and 3.3.
z=dep h
scale (3.1)
x=(a−cx)×z
x (3.2)
y=(b−cy)×z
y (3.3)
In hese equa ions, dep h ep esen s he dep h alue o a poin , scale depends on he measu e-
men uni s used by he came a, aand ba e he pixel coo dina es o a poin and cx,cy, x and y
a e in insic came a pa ame e s. This is he same way he ZED came a calcula es i s poin cloud,
howe e , he SDK calcula es a poin cloud o he whole image and each poin can be accessed
indi idually, which inc eases compu a ion imes signi ican ly mo e.
Following all his p e-p ocessing, he Dense usion model akes as inpu a c op o he came a
image ame (in he dimensions o he objec ’s bounding-box), he calcula ed poin cloud, he 2000
poin s om he mask, and he objec ’s ID. The pose es ima ion is ou pu ed in he o m o o a ion
qua e nions and a ansla ion ec o , which hen goes h ough wo i e a ions o pose e inemen , o
p oduce be e esul s. Doing wo i e a ions is as e and has compa able accu acy o doing mo e
i e a ions.
The ou pu o Dense usion is no he inal pose es ima ion, howe e . These alues c ea e a
ans o ma ion ma ix ha ans o ms he poin s in he objec 3D model om he canonical ame
3.2 Sys em’s Modules 39
de ined in he da ase o he came a coo dina e ame. To ob ain a inal pose es ima ion, pos -
p ocessing is c ucial. In pos -p ocessing, he se o op imal g asping poin s (which we e loaded
om one o he con igu a ion iles), a e ans o med acco ding o he ans o ma ion gi en by he
Dense usion model. The cen e o his g oup o poin s is hen calcula ed o gi e he op imal poin
o con ac o g asping.

40 3D Objec Recogni ion and Localiza ion Sys em
Chap e 4
Expe imen s and Resul s
4.1 Da ase
In o de o es and alida e ou p oposed sys em, and gi en ha a machine lea ning model is only
as good as he da a i is ed, a da ase was chosen o ain and es he compu e ision module.
The da ase used o ain bo h he objec de ec ion and he pose es ima ion models was he YCB-
Video da ase , which uses 20 objec s om he YCB Objec Model Se [35]. Using a RGB-D
came a, di e en g oups o he 20 objec s we e ilmed in di e en en i onmen s, wi h a ying
backg ounds, di e en ligh ing condi ions, and si ua ions o occlusion. The came a mo ed in
ela ion o he objec s o c ea e di e en iewpoin s o e e y objec , which esul ed in a o al o
133827 ames o ideo. Fu he mo e, he 3D poin cloud models o e e y objec we e ob ained
wi h a scanning ig.
F om a quan i a i e s andpoin , his da ase has 133827 RGB images (wi h 480x640 pixels)
and hei co esponding dep h maps, as well as a collec ion o he poin cloud 3D models o all
he objec s. Fo e e y ame and each o he objec s in he ame, he e a e anno a ions o class
label, g ound- u h segmen a ion mask, and pose alues o he objec (gi en as a ans o ma ion,
om he canonical ame o he objec o he came a coo dina e ame). The came a’s in insic
pa ame e s and i s posi ions in he wo ld a e also gi en o each ame.
Figu e 4.1: Dis ibu ion o numbe o ins ances pe objec
41
42 Expe imen s and Resul s
In o de o e alua e i he da ase was imbalanced, he numbe o ins ances pe ype o objec in
he da ase was plo ed. As can be seen in Fig. 4.1, he dis ibu ion o ins ances o each objec is
ai ly balanced, and each objec has, a leas , 15000 ins ances. This con ibu es o a mo e balanced
aining o he machine lea ning models and helps o educe bias. This cha ac e is ic explains why
his da ase is one o he mos impo an and one he mos used o he aining and alida ion o
he pose es ima ion models e iewed in Sec ion 2.3.3.
F om a quali a i e s andpoin , he YCB-Video da ase p esen s a lo o ad an ages in e ms o
a iabili y. One o hem is he di e ence in backg ound, which can be bo h unclu e ed (Fig. 4.2a)
and clu e ed (Fig. 4.2e), as well as ligh o da k, and simple o complex (as is shown in all he
images o Fig. 4.2).
(a) (b) (c)
(d) (e) ( )
Figu e 4.2: Selec ion o illus a i e images om he da ase
When compa ing Figs. 4.2a and 4.2b, i can be seen ha he same objec s, in he same scene,
a e shown in a a ie y o iewpoin s. This is mainly due o he mo emen o he came a, which
helps o c ea e a subs an ial ange o poses o he aining o he pose es ima ion model. Ano he
impo an aspec o he da ase is he change in ligh ing ha occu s be ween scenes. Compa ing
Figs. 4.2b and 4.2d, he e a e examples o an o e exposed and an unde exposed image, espec-
i ely. This a iabili y can help o make any model mo e obus o ligh ing changes.
Finally, i is also impo an o men ion he way occlusions a e c ea ed o an objec . The
a ying iewpoin s help c ea e na u al occlusions - when objec s o e lap - and unca ion o an
objec - when i lea es he ield o iew o he came a (such as in Fig. 4.2c). Mo eo e , he
a angemen o he objec s in he scene also helps o c ea e complex iewpoin s o an objec ,
whe e o e laps may occu ( o example, in igs. 4.2a,4.2d and 4.2e). All hese aspec s we e
ele an o he choice o da ase , gi en ha wi h mo e a iabili y in di e en aspec s o each
objec , he obus ness o he ained model will imp o e.
4.2 O line Tes ing Phase 43
4.2 O line Tes ing Phase
In his sec ion, he aining and alida ion o each ML model in he Compu e Vision Module, will
be discussed. The models we e ained and alida ed on he da ase analyzed in he sec ion abo e
(Sec ion 4.1). A e his, he pose es ima ion model was es ed using as inpu s he ou pu s gi en
by he objec de ec ion model.
4.2.1 T aining he Objec De ec ion Model
The aining o his model was hea ily based on he idea o ans e lea ning. This me hod consis s
o using p e- ained weigh s om a model - ob ained on one da ase - as a s a ing poin o aining
on a di e en da ase . This app oach speeds up aining ime and inc eases pe o mance, bu only
when he p e- ained model lea ned ele an gene al ea u es on a balanced da ase . In ou i s
app oach, he p e- ained weigh s we e used o e e y b anch, excep he ne wo k heads; on he
second app oach, he p e- ained weigh s we e u ilized only o he backbone ea u e ex ac ion
ne wo k, excep o he inal laye s. The i s app oach esul ed in a 14% d op in mask gene a ion
accu acy and a 5% d op in class label accu acy, in ela ion o he second app oach. Al hough
he i s app oach only akes 33 hou s o ain compa ed o he 46 hou s o aining in he second
app oach, he accu acy o he segmen a ion masks is pa amoun , so his is an accep able ade-o
be ween speed and accu acy.
Table 4.1: Compa ison o o iginal and ine- uned hype -pa ame e alues
Hype -pa ame e O iginal
alue
Fine- uned
alue
Numbe _classes NA 21
Max_g _ins ances 100 10
De ec ion_min_con idence 0.8 0.7
ROI_miniba ch_size 512 128
RPN_nms_ h eshold 0.7 0.5
Lea ning_ a e 0.02 0.002
RPN_class_loss 1.0 1.0
RPN_bbox_loss 1.0 0.01
M cnn_class_loss 1.0 1.0
M cnn_bbox_loss 1.0 0.01
M cnn_mask_loss 1.0 10.0
Gi en ha he o iginal hype -pa ame e s o he model we e op imized o he COCO da ase ,
he hype -pa ame e s o he objec de ec ion model we e uned. Fo ha , Tenso boa d (a isu-
aliza ion and op imiza ion ool o uning pa ame e s in Tenso low) was used. In Table 4.1, a
compa ison be ween he o iginal and he ine- uned model pa ame e s is shown.
The numbe _classes pa ame e akes he alue o 21, since he e a e 20 objec s in he da ase
and one backg ound class. Max_g _classes was changed o 10 om 100 due o he ac ha each
image in he da ase has a maximum o 10 objec ’s ins ances. To allow o mo e p edic ions
50 Expe imen s and Resul s
da ase used in aining he model, i did no lea n how o di e en ia e hese ea u es in a complex
objec co ec ly.
In scene 2 (Fig. 4.3b), he h ee main changes we e he a ia ion o iewpoin o he bowl,
he o e lap be ween one o he gela in boxes and he mug, and he pa ial occlusion o he ma ke
by he smalle clamp. Like is shown in 4.7, his had an immedia e impac on he de ec ion o he
gela in box. The e is only 50% accu a e de ec ions which co espond o he non-occluded box.
The inaccu a e de ec ions o he bowl co espond o i being iden i ied as a mug, due o he shape
in his iewpoin being simila o ha o a mug. The pa ial occlusion o he ma ke causes he
ma ke o be unde ec ed.
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 153 50 0 50 0.85
Bowl 13 1 153 81.7 18.3 0 0.72
Mug 14 1 153 100 0 0 0.98
La ge ma ke 18 1 153 0 0 100 NA
Clamp 19 2 153 0 74.5 25.5 NA
Table 4.7: De ec ion esul s on scene 2
In scene 3 (Fig. 4.3c), he smalle gela in box is o e lapped on op o he bowl, c ea ing a
pa ial occlusion o he bowl. This occlusion c ea es a iewpoin e y simila o he mug iewpoin ,
which explains he inaccu a e de ec ion a e (which can be seen in Table 4.8). The smalle gela in
box is s ill well de ec ed e en on op o he bowl; howe e , he occluded gela in box emains
unde ec ed, while i is occluded by he clamp. The ma ke , in an up igh posi ion, p esen s good
de ec ion esul s.
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 116 50.0 0 50.0 0.89
Bowl 13 1 116 88.8 11.2 0 0.90
Mug 14 1 116 100 0 0 0.97
La ge ma ke 18 1 116 89.7 0 10.3 0.76
Clamp 19 3 116 0 66.7 33.3 NA
Table 4.8: De ec ion esul s on scene 3
Scene 4 (Fig. 4.3d) p esen s all he objec s close oge he , wi h pa ial occlusions o he
smalle gela in box. This p oximi y be ween objec s makes i ha d o he de ec ion model o

4.3 Online Tes ing Phase 51
co ec ly iden i y he objec s. In ac , he wo bes pe o ming objec s (acco ding o he esul s in
Table 4.9) - he bowl and he mug - a e he mo e isola ed in he g oup.
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 115 4.3 0 95.7 0.72
Bowl 13 1 115 100 0 0 0.95
Mug 14 1 115 100 0 0 0.97
La ge ma ke 18 1 115 0 0 100 NA
Clamp 19 2 115 0 44.8 55.2 NA
Table 4.9: De ec ion esul s on scene 4
Scene 5 (Fig. 4.3e) p esen s a on al iew o he op o he bowl. Since his iewpoin is e y
unique o his ype o objec , he de ec ion is 100% accu a e (Table 4.10). In a simila si ua ion o
scene 4, he ma ke is unde ec ed due o i s p oximi y o he bowl and he wood block behind.
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 113 100 0 0 0.87
Bowl 13 1 113 100 0 0 0.87
Mug 14 1 113 72.6 0 27.4 0.74
La ge ma ke 18 1 113 0 0 100 NA
Clamp 19 3 113 0 0 100 NA
Table 4.10: De ec ion esul s on scene 5
As can be seen in Table 4.11, e en wi h he hea y occlusion o he bowl in scene 6 (Fig. 4.3 ),
is s ill p esen s a easonable de ec ion a e. In his scene, he powe d ill was in oduced, howe e ,
due o i s occlusion, he iewpoin gene a ed is no ep esen a i e enough o he shape o he
objec , in o de o he model o de ec i .
52 Expe imen s and Resul s
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 140 91.4 0 8.6 0.81
Bowl 13 1 140 88.6 0 11.4 0.83
Mug 14 1 140 100 0 0 0.97
Powe d ill 15 1 140 0 0 100 NA
La ge ma ke 18 1 140 100 0 0 0.92
Clamp 19 2 140 0 0 100 NA
Table 4.11: De ec ion esul s on scene 6
Fo scene 7, he esul s o Table 4.12 show ha e en when occlusion does no occu , ce ain
posi ions o he gela in boxes cause he model no o be able o de ec i . This is due o he small
amoun o iewpoin s in he educed da ase used o online es ing. All o he objec s, excep he
clamp, ha e a 100% de ec ion a e.
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 209 50 0 50 0.83
Bowl 13 1 209 100 0 0 0.83
Mug 14 1 209 100 0 0 0.96
Powe d ill 15 1 209 100 0 0 0.83
La ge ma ke 18 1 209 100 0 0 0.92
Clamp 19 1 209 0 100 0 NA
Table 4.12: De ec ion esul s on scene 7
Scene 8 (shown in Fig. 4.3h) p esen s a ai ly s anda d iew o he objec s, his ime wi h he
emo al o he clamp om he wo kspace. Table 4.13 shows ha mos o he objec s a e well
de ec ed, wi h he excep ion o he ma ke and he powe d ill.
4.3 Online Tes ing Phase 53
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 143 97.2 0 2.8 0.83
Bowl 13 1 143 100 0 0 0.88
Mug 14 1 143 100 0 0 0.99
Powe d ill 15 1 143 24.5 0 75.5 0.73
La ge ma ke 18 1 143 0 0 100 NA
Table 4.13: De ec ion esul s on scene 8
Finally, in scene 9 (Fig. 4.3i), e en wi h e y clea iewpoin s o he objec s, de ec ion esul s
a e poo o he gela in box, he powe d ill and he clamp.
Name ID Coun
Numbe
o
ames
Accu a e
De ec ion
(%)
Inaccu a e
De ec ion
(%)
No
de ec ion
(%)
A e age
Con idence
(0-1)
Gela in box 8 2 161 50 0 50 0.8
Bowl 13 1 161 100 0 0 0.91
Mug 14 1 161 100 0 0 0.99
Powe d ill 15 1 161 0 0 100 NA
La ge ma ke 18 1 161 91.9 0 8.1 0.76
Clamp 19 2 161 0 50 50 NA
Table 4.14: De ec ion esul s on scene 9
As can be seen, he educed da ase has a big impac on de ec ion pe o mance. The main
issues wi h he use o so li le ins ances o each objec esul s in he ollowing p oblems:
•Fo complex objec s (such as he clamp), he model needs mo e da a o be able o dis inguish
ea u es be ween objec s, esul ing in si ua ions o inco ec classi ica ion;
•E en o objec s wi h simple ea u es (like he gela in box, he bowl, and he ma ke ), he
model o e - i ed o ce ain iewpoin s (which we e mo e p e alen in he educed da ase ),
which esul ed in no de ec ion in ce ain objec posi ions;
•When objec s we e oo close oge he , he model had ouble in co ec ly de ec ing hem,
since his abili y is a mo e complex ea u e which needed mo e aining da a;
Howe e , as in any o he p oblem, besides he issues, he e a e some posi i e aspec s wo h
men ioning ega ding he model’s pe o mance:
•The model is able o gene alize, and co ec ly iden i y objec s which a e di e en om he
da ase , e en wi h limi ed da a. This is he case o he gela in boxes - which ha e di e en
54 Expe imen s and Resul s
heigh and wid h in ela ion o he o iginal objec -, he bowl - which is alle and less wide
han he o iginal model - and he mug - which has a di e en colo and is mo e na ow han
he da ase model;
•Wi h he limi ed amoun o aining da a, he model is s ill able o de ec accu a ely and wi h
g ea con idence, he less complex objec s, like he gela in boxes, he bowl, and he mug;
•E en wi h hea y occlusion (as seen in Fig. 4.3 ), he model can co ec ly de ec he bowl;
One p oblem ha is common o bo h he educed da ase and he o iginal da ase is he lack
o a ie y in he numbe o objec s pe image. Since his da ase was o iginally made o pose
es ima ion, he main objec i e was o c ea e a complex en i onmen wi h p oximi y and occlusion
be ween objec s. This is he igh app oach o pose es ima ion; howe e , in he objec de ec ion
model’s aining, i can lead o o e - i ing.
In es ing, when an objec was posi ioned alone in he wo kspace, he de ec ion a e d opped
signi ican ly and was e y inconsis en . Howe e , when o he objec s we e posi ioned close o i ,
he de ec ion a e imp o ed and was mo e s able. This issue is due o o e - i ing o si ua ions
wi h a g oup o objec s.
4.3.3 G asping Expe imen s
In he hi d and inal es , he compu e ision module was alida ed. The es ing me hodology
used was a g asping expe imen on he 5 o he 6 objec s ha we e chosen o es ing. This was
due o he a ailable g ippe being oo small o be able o g asp he powe d ill.
Figu e 4.4: Illus a ion o one g asping posi ion pe objec
Each objec was placed in a di e en wo kspace posi ion, su ounded by o he objec s. A
g asping a emp was made o each posi ion, o a o al o i e a emp s. Illus a ions o one o he
4.3 Online Tes ing Phase 55
posi ions o each objec a e p esen ed in Fig. 4.4. The esul s ob ained we e eco ded in Table
4.15.
Scene Objec ID
Success ul
a emp s
(%)
A e age
e o
(cm)
1 8 60 2
2 8 60 5
3 13 20 2.7
4 14 80 0.5
5 18 60 3.5
6 19 40 0.75
Table 4.15: G asping esul s
Fo each objec , he Compu e Vision Module an de ec ion, gene a ion o he segmen a ion
mask, and pose es ima ion. The coo dina es o he cen e poin o he objec we e ob ained, and
a ans o ma ion was applied om he cen e poin o he op imal g asping poin o enable co ec
g asping.
Due o he g ippe model no being eadily a ailable and also ime- es ic ions o es ing,
he ans o ma ion o he se o op imal g asping poin s (discussed in Sec ion 3.2.2.2) was no
implemen ed. Howe e , his es se es as a p oo o concep , since he ans o ma ion o all
objec poin s is equi alen o he ans o ma ion o only a subse o op imal g asping poin s.
Fu he mo e, since he p ima y pu pose o his p ojec is o de elop a sys em o objec ecog-
ni ion and localiza ion, he main ocus was on he c ea ion o a pipeline wi h hose capabili ies
whe e he pos -p ocessing s ep in ol ed in g asping was no he main ocus.
The wo s pe o ming objec s a e he bowl and he clamp. These a e expec ed esul s when
looking a bo h he esul s in Table 4.3 and in Sec ion 4.3.2. The clamp, being a complex objec
wi h he wo s esul in Table 4.3 and wi h poo de ec ion esul s, was unlikely o p esen success ul
g asping a emp s. As o he bowl, i was he second-wo s pe o ming objec in Table 4.15, due
o being a symme ical objec . Bo h he gela in boxes ha e compa able success a es; howe e ,
he a e age e o is mo e signi ican o he la ge box, since small e o s in he pose’s es ima ion
ha e a mo e subs an ial in luence in la ge objec s.
One o he main sou ces o pose es ima ion e o du ing his phase o es ing was he incon-
sis ency and inaccu acy o he segmen a ion masks ou pu ed by he objec de ec ion model. This
was al eady an issue in Sec ion 4.2.2, howe e , he small acquisi ion e o s om he came a also
con ibu e o he o e all de ec ion e o . Addi ionally, as is concluded in [1], he ZED came a has
an associa ed e o o dep h da a (Fig. 4.5). Due o he wo kspace being a a maximum dis ance
o 1.1m om he came a, his dep h es ima ion e o is e y small, howe e , i is s ill in luen ial.

56 Expe imen s and Resul s
Figu e 4.5: RMS e o in dep h da a o 720p esolu ion [1]
In his phase o es ing, gi en he eal- ime cons ains o g asping applica ions, he e was
also a measu emen o he p ocessing ime o each componen : da a acquisi ion module, objec
de ec ion, and segmen a ion model, and 6D objec pose es ima ion model. On a e age, ame
acquisi ion and analysis ook 10ms. The objec de ec ion and objec localiza ion models ook on
a e age 346ms and 29.6ms, espec i ely. Since he compu e ision module uns pa allel wi h
ame acquisi ion, he whole p og am akes abou 375ms pe ame on a e age. This means i can
un a abou 2.6 ps.
A ame- a e o 2.6 ps is e y slow when applied o no mal ideo-playback o complex asks,
such as eal- ime au onomous d i ing. Howe e , conside ing he no e y dynamic en i onmen
on which his sys em aims o wo k in and conside ing he slow speed o a ull g asping a emp by
he obo ic manipula o used o es ing (abou 5s), his ame- a e is adequa e.
4.4 Limi a ions
Wi h he expe imen s and es s pe o med, some limi a ions in ou app oach we e ound. One o
he limi a ions o he p oposed sys em is ha he ou pu o ideal g ippe o a ions is based on a
p e iously c ea ed con igu a ion ile. While his is a alid solu ion ha esol es many e o s ha
can occu when calcula ing hese o a ions au oma ically, while also allowing he use o choose
he ideal g asping acco ding o his own p e e ence and applica ion, i also asks much o e head
wo k om anyone who ies o implemen he sys em. Howe e , since he sys em’s main objec i e
is o ou pu objec de ec ion and pose es ima ions o execu ing he g asping asks, i is no a
e y impac ul limi a ion. O e all, he ade-o be ween allowing o use adap abili y and use
o e head wo k may be bene icial.
The de eloped sys em also a ec ed e o s in he acquisi ion o isual da a and ligh ing con-
di ions due o he cha ac e is ics o he ML models used. Ligh ing is always a big pa o e o s
4.4 Limi a ions 57
in oduced in CV models, so his is an accep able limi a ion. The main way o sol ing his limi-
a ion would be o ain he model on an e en bigge da ase wi h a ange o ligh ing condi ions,
which would be oo cos ly.
A se e e limi a ion is he ac ha he de ec ion a e d ops when an objec is by i sel in he
wo kspace due o o e - i ing o si ua ions wi h a g oup o objec s. This is an issue o aining he
de ec ion model on a pose es ima ion da ase . E en hough, when i comes o de ec ion, he sys em
shows ha i gene alizes o objec s wi h di e en dimensions and cha ac e is ics om he da ase ,
in he case o pose es ima ion, objec s ha di e oo much om he da ase models will no be
co ec ly es ima ed since he pose es ima ion pipeline akes in o accoun he canonical objec 3D
model om he da ase . Ano he limi a ion is he ac ha he sys em canno choose he absolu e
op imal objec o g asp in a clu e ed si ua ion. I only chooses he objec s ha i has he highes
con idence o , wi hou aking in o accoun i s posi ion in ela ion o he o he objec s.
58 Expe imen s and Resul s
Chap e 5
Conclusions and Fu u e Wo k
5.1 Conclusions
In his hesis, he p oblem o de eloping an au oma ic objec ecogni ion and pose es ima ion sys-
em was add essed. This sys em combines objec localiza ion and iden i ica ion o imp o e obo ic
g asping. The main con ibu ion o his wo k was he de elopmen o a amewo k in eg a ing wo
di e en ML models o c ea e a sys em o mo e comple e and au onomous g asping asks.
Ex ensi e esea ch was ca ied ou on he cu en s a e o Machine Lea ning me hods, pa icu-
la ly Con olu ional Neu al Ne wo ks. This esea ch esul ed in a deepe unde s anding o his ield
and i s applica ions in Compu e Vision: om he mos basic o he mo e ad anced de elopmen s
ha ha e occu ed h oughou he yea s. The main s a e-o - he-a app oaches o objec de ec-
ion/ ecogni ion and pose es ima ion we e e iewed and compa ed, which led o mo e in o med
de elopmen o he p oposed sys em.
This hesis’ main ocus was on he in eg a ion o wo di e en ML models o de ec ion and
pose es ima ion. Bo h Mask-RCNN and Dense usion we e co ec ly in eg a ed and show adequa e
esul s in a eal es ing en i onmen . The de eloped sys em is capable o de ec ing mo emen s
and di e ences be ween ideo ames, he e o e educing he need o cons an calcula ion and
in e ence on each ideo ame. The p oposed solu ion uns en i ely au oma ically, a e an ini ial
use pa ame e s’ selec ion and can gene a e g asping ou pu s om ideo da a a a a e o 2.6 ps.
An in e es ing posi i e aspec o he p oposed solu ion is he abili y o a human o be in con ol
o choosing he op imal g asping poin s o an objec and c ea ing a con igu a ion ile wi h hese
poin s, hus allowing o use and applica ion adap abili y, while inc easing accu acy and speed o
he sys em, since i does no es ima e hese poin s du ing un ime.
The main limi a ions ound in he p oposed solu ion we e he o e head wo k needed o ou pu
co ec g ippe o a ions o ce ain objec posi ions, he impac o ligh ing condi ions, and he
dependency on p e iously scanned 3D models o he eal objec s.
59