scieee Open visual document viewer

Automatic 3D Object Recognition and Localization for Robotic Grasping

Bruno Miguel Silva Espírito Santo

Abstract

This project aims to design a detection and pose estimation pipeline for objects, to be used in an industrial environment in a robotic grasping setting.

Full text

FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO Au oma ic 3D Objec Recogni ion and Localiza ion o Robo ic G asping B uno Miguel Sil a Espí i o San o WORKING VERSION Mes ado In eg ado em Engenha ia Ele o écnica e de Compu ado es Supe iso : Gil Manuel Magalhães de And ade Gonçal es Second Supe iso : Liliana Pa ícia Saldanha An ão July 6, 2020 Abs ac Wi h he ad en o Indus y 4.0 and i s highly econ igu able manu ac u ing con ex , he ypical ixed-posi ion g asping sys ems a e no longe usable. This eali y unde lined he necessi y o ully au oma ic and adap able obo ic g asping sys ems. The de elopmen o a Compu e Vision sys em capable o de ec ing, iden i ying and es ima ing he 6D pose o an objec would b ing us close o ha eali y. Wi h ha in mind, he main pu pose o his hesis is o join Machine Lea ning models o de ec ion and pose es ima ion in o an au oma ic sys em o be used in a g asping en i onmen . To achie e his, ex ensi e esea ch was ca ied ou on he cu en s a e-o - he-a app oaches. The de eloped sys em uses Mask-RCNN and Dense usion models o he ecogni ion and pose es ima ion o objec s, espec i ely. The g asping is execu ed aking in o conside a ion bo h he pose and he objec ’s ID, as well as allowing o use and applica ion adap abili y h ough an ini ial con igu a ion. The sys em was es ed bo h on a alida ion da ase and in a eal wo ld en i onmen . The main esul s show ha he sys em has mo e di icul y wi h complex objec s, howe e , i shows p omising esul s o simple objec s, e en wi h aining on a educed da ase . I is also able o gene alize o objec s sligh ly di e en han he ones seen in aining. In g asping expe imen s, he e is a 60% success a e in he bes cases, o simple g asping a emp s. i ii Acknowledgemen s My deepes hanks o bo h my supe iso s. I was an absolu e pleasu e o wo k on his hesis wi h such g ea people and in a e y enjoyable en i onmen , despi e he di icul wo king si ua ions his yea . A lo was lea ned in de eloping his hesis and I could no ha e done i wi hou hei i eless help. B uno Miguel Sil a Espí i o San o iii i Con en s Lis o Figu es ii Lis o Tables ix Abb e ia ions xi 1 In oduc ion 1 1.1 Con ex ....................................... 1 1.2 Mo i a ion...................................... 2 1.3 P oblemDe ini ion ................................. 3 1.4 Objec i es...................................... 4 1.5 ThesisS uc u e................................... 4 2 Li e a u e Re iew 7 2.1 MachineLea ning.................................. 7 2.1.1 A i icial Neu al Ne wo ks . . . . . . . . . . . . . . . . . . . . . . . . . 9 2.2 PoseandRo a ions ................................. 13 2.3 Rela edWo k .................................... 13 2.3.1 Da ase s................................... 14 2.3.2 Machine Lea ning in Objec De ec ion and Recogni ion . . . . . . . . . 15 2.3.3 Machine Lea ning in Pose Es ima ion . . . . . . . . . . . . . . . . . . . 21 3 3D Objec Recogni ion and Localiza ion Sys em 31 3.1 Sys emO e iew.................................. 31 3.1.1 Ope a ionModes.............................. 32 3.2 Sys em’sModules.................................. 35 3.2.1 Da a Acquisi ion Module . . . . . . . . . . . . . . . . . . . . . . . . . . 35 3.2.2 Compu e Vision Module . . . . . . . . . . . . . . . . . . . . . . . . . . 36 4 Expe imen s and Resul s 41 4.1 Da ase ....................................... 41 4.2 O lineTes ingPhase ................................ 43 4.2.1 T aining he Objec De ec ion Model . . . . . . . . . . . . . . . . . . . 43 4.2.2 Pose Es ima ion Model T aining . . . . . . . . . . . . . . . . . . . . . . 45 4.2.3 Compu e Vision Module Tes ing . . . . . . . . . . . . . . . . . . . . . 46 4.3 OnlineTes ingPhase ................................ 47 4.3.1 Da a Acquisi ion Module Tes ing and Analysis . . . . . . . . . . . . . . 48 4.3.2 Objec De ec ion and Segmen a ion Tes ing . . . . . . . . . . . . . . . . 48 4.3.3 G aspingExpe imen s ........................... 54 i CONTENTS 4.4 Limi a ions ..................................... 56 5 Conclusions and Fu u e Wo k 59 5.1 Conclusions..................................... 59 5.2 Fu u eWo k..................................... 60 Re e ences 61 Lis o Figu es 2.1 Rep esen a ion o a pe cep on . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 2.2 Rep esen a iono aae................................ 10 2.3 NNVSCNN .................................... 11 2.4 Illus a ion o he a chi ec u e o Fas e R-CNN . . . . . . . . . . . . . . . . . . 15 2.5 Illus a ion o he a chi ec u e o R-FCN . . . . . . . . . . . . . . . . . . . . . . 17 2.6 Anexampleo asco emap............................. 17 2.7 O e iew o he SSD a chi ec u e . . . . . . . . . . . . . . . . . . . . . . . . . 18 2.8 O e iew o he o iginal YOLO ne wo k . . . . . . . . . . . . . . . . . . . . . . 19 2.9 O e iew o he SSD-6D model . . . . . . . . . . . . . . . . . . . . . . . . . . 21 2.10 Illus a ion o e e y iewpoin conside ed . . . . . . . . . . . . . . . . . . . . . 22 2.11 O e iew o he PoseCNN sys em . . . . . . . . . . . . . . . . . . . . . . . . . 23 2.12 O e iew o he DPOD sys em . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 2.13 Illus a ion o he pose e inemen block . . . . . . . . . . . . . . . . . . . . . . 25 2.14 O e iew o he DenseFusion a chi ec u e . . . . . . . . . . . . . . . . . . . . . 26 2.15 Illus a ion o he pose e inemen ne wo k . . . . . . . . . . . . . . . . . . . . . 27 2.16 Illus a ion o he Hyb idPose model . . . . . . . . . . . . . . . . . . . . . . . . 28 2.17 O e iew o he DeepIM model . . . . . . . . . . . . . . . . . . . . . . . . . . 29 3.1 Illus a ion o he da a low in he sys em . . . . . . . . . . . . . . . . . . . . . . 31 3.2 O e iew o all ope a ion modes . . . . . . . . . . . . . . . . . . . . . . . . . . 33 3.3 S a e machine o he pose es ima ion module . . . . . . . . . . . . . . . . . . . . 34 3.4 ZED S e eo Came a and dimensions illus a ion . . . . . . . . . . . . . . . . . . 35 3.5 O e iew o he Mask-RCNN model . . . . . . . . . . . . . . . . . . . . . . . . 37 4.1 Dis ibu ion o numbe o ins ances pe objec . . . . . . . . . . . . . . . . . . . 41 4.2 Selec ion o illus a i e images om he da ase . . . . . . . . . . . . . . . . . . 42 4.3 O e iew o all he scena ios used in de ec ion es ing . . . . . . . . . . . . . . . 49 4.4 Illus a ion o one g asping posi ion pe objec . . . . . . . . . . . . . . . . . . . 54 4.5 RMS e o in dep h da a o 720p esolu ion . . . . . . . . . . . . . . . . . . . . 56 ii 2In oduc ion ML is a ield o esea ch dedica ed o compu e s’ abili y o in e ac wi h hei en i onmen and lea n om he da a collec ed. I is based on ma hema ical models, employed o enable he lea ning o pa e ns in da a, and is o en used o inc ease he sys em’s in elligence/ lexibili y o when p oblems canno be sol ed using adi ional explici p og amming. On he o he hand, Deep Lea ning is a b anch o ML ha desc ibes a se o modi ied ML echniques inspi ed by he biological ne ous sys em. Composed o ne wo ks o pa allel and concu en con olu ions, as well as o he ma hema ical ope a ions pe o med di ec ly on he inpu da a, i acqui es a g oup o ep esen a i e heu is ics be ween inpu and ou pu da a. Due o he con incing esul s ha ha e been achie ed in he scope o compu e ision, he e is an inc easing end owa ds he implemen a ion o DL algo i hms in obo ic g asping applica ions. Wi h ML and DL, obo s a e able o in e ac wi h hei en i onmen and espond o a ious s imuli in a comple ely au oma ed manne , enabling obo s o use came as o pe cei e hei en i onmen and e en "unde s and" i when using hese echniques. The main challenges in he ield o CV co e ed in his hesis a e objec de ec ion/ ecogni ion and pose es ima ion, which will be ackled wi h he use o ML models. In objec de ec ion and ecogni ion, he ask is o iden i y objec s in an image and a ibu e a seman ic classi ica ion (iden- i y he class i belongs o). As o pose es ima ion, he e is a need o es ima e he objec ’s posi ion and o ien a ion ela i e o he came a. 1.2 Mo i a ion As p e iously men ioned, objec de ec ion/ ecogni ion and pose es ima ion is o g ea impo ance o obo ic sys ems in se e al asks. One o hese asks is obo ic g asping: he ac o a obo ic manipula o handling objec s in an au oma ed manne . Fo ypical obo ic g asping o be accomplished, he obo needs o be able o know whe e and how an objec is posi ioned in i s coo dina e sys em. 6-dimensional (6D) poses, composed o 3 DOF o he posi ion and 3 o o ien a ion, a e o en used o his pu pose. Con a y o ypical 3D pose, 6D poses p o ide he o a ion, gi ing inpu no only on he "whe e" bu also on he "how." Ne e heless, when compa ed o humans, obo s p esen signi ican ly lowe success a es in g asping objec s in dynamic en i onmen s, as one would expec . Humans a e inhe en ly good a pe cei ing hei en i onmen and he objec s ha su ound hem, iden i ying hei cha ac e is ics, posi ion, and deciding na u ally whe e o g asp an objec and how much o ce o exe on i . Wi h he eali y o inc easing Human-Robo Collabo a ion solu ions in he indus y, humans, and obo s no only sha e physical space bu wo k oge he as a eam. I is o g ea impo ance o a eam o ha e he same pe cep ion and conside a ions o he asks hey will ul ill. Gi en his, ha ing g asping solu ions ha do no ake in o conside a ion he same aspec s as humans do, would lead o wo se eam pe o mance due o misma ches in g asping posi ions o hando e s, o e en due o g ippe damages in g asped ma e ial. Inspi ed by his, mo i a ion is p o ided o de elop solu ions ha imp o e manipula ion asks in lexible and collabo a i e indus ial en i onmen s, cha ac e is ic o he new indus ial eali y 1.3 P oblem De ini ion 3 imposed by Indus y 4.0. By p o iding obo s, in an indus ial en i onmen , wi h he abili y o lea n hei g asping asks au onomously in a simila manne o humans, collabo a i e obo ics is also empowe ed. This way, obo s a e gi en mo e pe cep ion o he objec s in a manipula ion ask. O e all, obo s a e able o wo k mo e in ui i ely wi h human ope a o s and e en assis hem in a ious asks in a sa e and sha ed en i onmen . 1.3 P oblem De ini ion The p oblem o objec de ec ion and pose es ima ion is one o he ac i e esea ch ha has su e ed inc emen al imp o emen s, as can be seen in chap e 2. Achie ing accu a e pose es ima ion o an objec would enable obo s o pe o m g asping in a ully au oma ic and e icien manne . Howe e , in o de o imp o e obo ic g asping, besides localiza ion, p o iding objec iden i i- ca ion can also be use ul. Fo his, ML is o en used, classi ying he seman ic class o he objec ha can hen be u ilized o adap ha g asping o he objec ’s cha ac e is ics. Fo ins ance, he g ipping s eng h ha he manipula o can apply wi hou causing damage o he objec can be in- e ed om i s classi ica ion, o e en choosing which objec o pick ou in a g oup o objec s, by i s cha ac e is ics. The e a e al eady many challenges wi h de ec ing and es ima ing he pose o objec s, a ising wi h he a iabili y o he su ounding en i onmen , such as a ying ligh condi ions, occlusion o objec s, and di e si y o objec dimensions. O he issues can also occu in e ms o he accu acy o he senso s used and he ypes o noise in oduced. Gi en his, he sys em o be de eloped needs o be accu a e and as enough o wo k in a eal indus ial en i onmen and be able o handle mul iple objec s o di e en dimensions, in si ua ion o occlusions o pa o he objec s. In sum, challenges ha a ise in objec de ec ion and pose es ima ion sys ems, by hemsel es, a e: •The ligh ing condi ions o he en i onmen ; •Occlusion and unca ion o ce ain objec s; •The need o models o be able o pe o m in eal- ime wi h low p ocessing speeds; •The need o high accu acy in indus ial en i onmen s whe e sys ems need o be sa e o ope a e. The e a e al eady many objec de ec ion/classi ica ion sys ems, as well as pose es ima ion models o g asping. Howe e , e y ew showing he bene i s o joining bo h app oaches o a mo e comple e and au onomous g asping, much less ully au oma ed o based on only came a ame in o ma ion (bo h RGB and/o dep h). The e o e, he e is a challenge o joining localiza ion and ca ego iza ion o a bi a y objec s in 3D o g asping, in eal- ime, using li e- eed came a moni o ing. This would ul ima ely b ing us a s ep close o seamless Human-Robo Collabo a ion. 4In oduc ion 1.4 Objec i es Gi en he p oblem de ined abo e, he pu pose o his hesis is o explo e and de elop an in elligen and ully au oma ed sys em o eal- ime objec ecogni ion and localiza ion de i ed om he analysis o da a om a s e eo ision came a. The de ec ion/ ecogni ion o a ious objec s and hei localiza ion wi h 6-DOF (posi ion and o ien a ion in 3D space) will be e ie ed o hen pe o m g asping asks acco ding o ha da a. This sys em should be able no only o ca ego ize and accu a ely loca e mul iple objec s bu also o be obus o occlusions. Finally, wi h his solu ion, a highe -le el obo g asping is hoped o be achie ed, imp o ing lexible and au onomous obo ic solu ions. In sum, he ollowing objec i es a e expec ed o be achie ed by he end o he hesis: 1. Ca y ou ex ensi e esea ch o he cu en s a e-o - he-a app oaches o objec de ec ion and ecogni ion, as well as pose es ima ion; 2. Gain an unde s anding o he e olu ion o models in his ield; 3. C ea e a sys em, based on ML, capable o de ec ing and classi ying an objec and es ima ing i s pose wi h 6-DOF; 4. The sys em needs o be ully au oma ed; 5. The sys em needs o be able o p ocess da a in eal- ime; 6. The ML model should be able o handle si ua ions o occlusion; 7. A obo ic manipula o should success ully and app op ia ely pe o m manipula ion asks o a ious objec s when ecei ing inpu s om he sys em; 1.5 Thesis S uc u e The ollowing ex is composed o ou chap e s: Li e a u e Re iew (chap e 2), 3D Objec Recog- ni ion and Localiza ion Sys em (chap e 3), Expe imen s and Resul s (chap e 4) and, inally, Con- clusions and Fu u e Wo k (chap e 5). Each chap e , as well as i s sec ions a e explained in his sec ion. In chap e 2, he main goal is o ca y ou ex ensi e esea ch o he ield in ques ion. In his chap e , we will i s b ie ly desc ibe Machine Lea ning and i s ca ego ies, ollowed by a mo e de ailed look in o A i icial Neu al Ne wo ks and some o i s a chi ec u es used in Compu e Vision. A sho o e iew o he de ini ion and cha ac e is ics o he objec ’s pose and o a ion will also be gi en. In chap e 3, i s , an o e iew o he comple e sys em o Au oma ic Objec Recogni ion and Localiza ion o obo ic g asping is gi en, as well as a look a i s da a low. A mo e in-dep h 1.5 Thesis S uc u e 5 desc ip ion is p o ided o he implemen a ion o his sys em, gi ing de ail no only on each ma- chine lea ning a chi ec u e bu also on each module and he sys em’s ope a ing modes as a S a e Machine. In chap e 4, he a ious expe imen s and esul s a e p esen ed, as well as a discussion o hose esul s and wha hey mean o he limi a ions o he sys em de eloped. In chap e 5, he conclusions a e made and a ew ideas o u u e wo k a e also p esen ed. 6In oduc ion Chap e 2 Li e a u e Re iew In his chap e , we will i s b ie ly desc ibe Machine Lea ning and i s ca ego ies, ollowed by a mo e de ailed look in o A i icial Neu al Ne wo ks and some o i s a chi ec u es used in Compu e Vision. A sho o e iew o he de ini ion and cha ac e is ics o he objec ’s pose and o a ion will also be gi en. Finally, he ela ed wo k ega ding objec de ec ion and pose es ima ion is p esen ed, de ailing commonly used da ase s, as well as Machine Lea ning me hods o bo h objec de ec ion and iden i ica ion and pose es ima ion. 2.1 Machine Lea ning Machine Lea ning is a subse o A i icial In elligence, and i deals wi h he abili y o compu e sys ems o ecognize pa e ns and in e om hem. This means ha , ins ead o asks ha ing o be di ec ly p og ammed, a ML model can be ained on sample da a and make decisions based on he gene aliza ions i makes om ha da a. ML is o g ea impo ance in au oma ion, enabling sys ems o pe o m bo h ou ine and com- plex asks (such as analyzing la ge se s o da a), wi h he added ad an age o making he sys em lexible o changes. I has p o en o be capable o sol ing a a ie y o asks, om image de ec ion o speech ecogni ion, ha can some imes be conside ed ha d o humans. This b anch o AI can be di ided in o h ee majo ca ego ies: 1) Unsupe ised lea ning, 2) Supe ised lea ning, and 3) Rein o cemen lea ning. O hose, only he i s wo will be e iewed, as hey a e he mos ele an o he opic in ques ion[2]. In Unsupe ised lea ning, he da a ed in o he model is no labeled. This means ha he ML model ies o ind he pa e ns o cha ac e is ic ea u es in he da a i ecei es, c ea ing a so o summa y wi hou being o e ed he eedback ha labeled aining da a would o e [2]. Me hods om his ca ego y gene ally can no be di ec ly applied o classi ica ion o eg ession p oblems, gi en ha he e is no p ecise knowledge o wha he ou pu da a migh be. Basically, his ca ego y exploi s he unlabeled da a a iance and sepa abili y o e alua e ea u e ele ance. Some o he mos known applica ions o unsupe ised lea ning include: 7 8Li e a u e Re iew •Clus e ing, which allows spli ing he da ase au oma ically in o g oups ("clus e s") acco d- ing o simila i y, o en ailing o ea da a poin s as indi iduals •Anomaly de ec ion, whe e he goal is o de ec a ypical da a poin s in he da ase au oma i- cally; •Associa ion mining, whe e se s o i ems ha occu oge he o en, a e iden i ied; •La en a iable models ha a e usually u ilized as me hods o p e-p ocessing da a (such as dimensionali y educ ion) o acili a ing da a isualiza ion ools. The main algo i hms in Unsupe ised Lea ning include K-means and P incipal Componen Analysis (PCA). On he o he hand, wi h Supe ised lea ning, he models ecei e labeled da a, e e ed o as " aining da a". This means ha he inpu consis s o he "objec " o be analyzed and he desi ed ou pu alue. This ype o lea ning is usually done in he con ex o classi ica ion (map inpu da a o ou pu class) o eg ession (map inpu da a o con inuous ou pu ). Usual algo i hms in supe ised lea ning include Logis ic Reg ession, Suppo Vec o Ma- chines, o A i icial Neu al Ne wo ks. The objec i e is he same in classi ica ion and eg ession: o ind pa icula ela ionships in he inpu da a ha enable he model o compu e he ou pu co - ec ly. In o he wo ds, he labels a e used o calcula e he co ec mapping unc ions o he inpu s ecei ed, which will enable he model o classi y new unlabeled inpu da a [3]. When pe o ming supe ised lea ning, model complexi y should be conside ed. The complex- i y o he model usually is dependen on he na u e o he aining da a. I he da ase is small, o i he da a be ween classes is clea ly dis inc , one should choose a low-complexi y model (wi h a high-complexi y model, i would likely o e - i , i.e., i would no be able o gene alize o o he da a poin s). The p oblem is ha labeled da a is ha d o ob ain since human anno a ion is bo ing. Labeling may equi e expe s o e en special de ices and can be e y ime-consuming. Gi en his, a possible solu ion is using semi-supe ised lea ning. This ype o lea ning is di e en om he me hods men ioned abo e since he inpu da a is a mix o labeled and unlabeled da a. Basically, i is an in e media e be ween supe ised and unsupe ised lea ning. The s anda d p ocess in ol es i s clus e ing simila da a using an unsupe ised lea ning al- go i hm, and hen, using he exis ing labeled da a, he unlabeled da a is labeled. The use o his me hod depends mainly on he applica ion and he se ing [3]. Wi h his b ie explana ion o he majo ca ego ies o ML and hei cha ac e is ics, a mo e in-dep h look a A i icial Neu al Ne wo ks, a e sa ile and widely used ML model, e y common in he opic o CV and objec de ec ion, will be gi en. 2.1 Machine Lea ning 9 2.1.1 A i icial Neu al Ne wo ks A sho in oduc ion o a i icial neu al ne wo ks and i s cha ac e is ics a e p o ided he e in b ie . This ype o model is one o he mos p e alen in ML, due o he ac ha ANNs a e some o he mos e sa ile and e ec i e algo i hms, and, as a esul , hey a e one o he mos signi ican subjec s in he ield o AI. Mos ANNs ha e, a hei basic le el, a i icial neu ons called "pe cep ons" (Fig. 2.1). Figu e 2.1: Rep esen a ion o a pe cep on. Image adap ed om [4] A pe cep on (o neu on) ecei es a ious inpu s (xi) mul iplied by i s co esponding weigh s (wi, numbe s exp essing he impo ance o he espec i e inpu s o he ou pu ) and akes he sum o hese esul s o p oduce a single ac i a ion (Eq. 2.1). The e can also be an addi ional e m, e e ed o as bias (b), ha ansla es how easy i is o ge he pe cep on o i e. a= n ∑ i=1 xi×wi+b(2.1) The pa ame e ain Eq. 2.1 is known as ac i a ion. The neu on’s ou pu (y) esul s in he ac i a ion o he pe cep on and is dependen on he alue aand he ac i a ion unc ion used. The ac i a ion unc ion is used o in oduce non-linea i y in o he ou pu o he pe cep on. This unc ion is chosen acco dingly o he na u e o he da a and he dis ibu ion o a ge a iables. Each ype o ac i a ion unc ion en ails a di e en ixed ma hema ical calcula ion. A adi ional A i icial Neu al Ne wo k is made up o in e connec ed laye s o pe cep ons (neu ons). The i s laye is known as he inpu laye , he las as he ou pu laye and he middle laye s a e e e ed o as hidden laye s. The inpu laye usually is equal o he numbe o di e en labels in a da ase , while he numbe o hidden laye s can ange om one o hund eds and a ies acco ding o he numbe o labels and compu a ion capabili y. As aining da a is ed h ough he ANN, he e o be ween he ac ual and he desi ed ou pu s o he ne wo k a e calcula ed using a cos /loss unc ion; his e o is hen used o adjus he weigh s and biases o he ANN. How hese adjus men s a e made depends on he op imiza ion algo i hm used, and on he pa ame e s o he ne wo k ha de ine he a ious ypes o a chi ec u es. 10 Li e a u e Re iew Mo eo e , he connec ions be ween neu ons can also be bi-di ec ional, leading o o he di - e en a chi ec u es. All o his con ibu es o he as amoun o ANNs in exis ence and he complexi y o his subjec [5]. In he nex subsec ions, some o he mos used ANN a chi ec u es in CV will be p esen ed b ie ly. 2.1.1.1 Au oencode s An Au oencode is an unsupe ised ANN, ha i s lea ns o comp ess and encode da a, and hen lea ns o econs uc back he da a om he encoded e sion. In a nu shell, an au oencode educes da a dimensions by lea ning how o igno e noise in he da a. Fo his o happen au oencode s ha e ou main pa s: encode ;bo lleneck,decode and econs uc ion loss. An example o an au oencode a chi ec u e is p esen ed in Fig. 2.2, whe e ˆxis he econs uc- ion o he o iginal inpu x. The encode akes i s inpu and ou pu s a ec o wi h i s encoded in o ma ion, comp essing and educing he inpu da a dimensions. The bo leneck is a laye in which he inpu ’s comp ession ep esen a ion is con ained ( he lowes possible dimensions o he inpu ). The decode ’s job is o hen ake his comp essed ec o and ou pu he closes ma ch o he o iginal inpu . The econs uc ion loss unc ion is hen minimized, which helps bo h he encode and decode lea n wi hou he need o labeled da a. This loss measu es how well he decode pe o ms by calcula ing how close he ou pu is o he o iginal da a. Howe e , i an encode lea ns o ma ch he o iginal inpu exac ly, he e is no use ulness o he ne wo k. Fo ha eason, es ic ions a e applied o limi how much o he inpu can be copied. This esul s in a unc ion ha p io i izes and only ou pu s he mos impo an cha ac e is ics o i s inpu [6]. Figu e 2.2: Gene al a chi ec u e o an Au oencode . 2.1 Machine Lea ning 11 2.1.1.2 Con olu ional Neu al Ne wo ks Ano he ype o ANN is he Con olu ional Neu al Ne wo k (CNN). Since, in a adi ional neu al ne wo k, all laye s a e ully connec ed, hey a e no able o ake in o accoun he image’s spa ial na u e (s uc u e and pa e n). In o he wo ds, all inpu pixels in he image would be ea ed he same way, independen ly o being a apa o close oge he . This is whe e CNNs a e mos use ul. They a e a class o deep neu al ne wo ks (an ANN wi h mul iple hidden laye s) op imized o da a in g id o m, which makes i especially use ul o image p ocessing. Fo his eason, CNNs ha e e ol ed o be he mos used me hod o sol ing a ious compu e ision p oblems [7] and will appea in he majo i y o Sec ion 2.3. CNNs a e iden ical o egula NN in e ms o being composed by neu ons, whe e he goal is o es ima e weigh s and bias. Howe e , CNNs ha e a much mo e e icien s uc u e o image p ocessing han NNs, gi en ha each neu on is only connec ed o a speci ic egion o he p e ious laye . The gene al a chi ec u e o a NN and CNN a e p esen ed in Fig. 2.3 o compa ison. Each laye o a CNN is ep esen ed as 3D olume [8]. Figu e 2.3: NN a chi ec u e VS CNN a chi ec u e. A s anda d NN is ep esen ed on he le , and a CNN on he igh . Each ci cle ep esen s a neu on [8]. This class o ne wo k uses a ian s o he ma hema ical ope a ion con olu ion. Con olu ion (deno ed as ∗) can be seen as he ma ix mul iplica ion o an inpu Iby a ke nel K(as seen in Eq. 2.2). Using a ke nel wi h a size signi ican ly in e io o he inpu size, a CNN can ex ac small ea u es o an image. This also esul s in ewe calcula ions and, he e o e, an inc ease in speed. C(i,j) = (I∗K)(i,j) = m ∑ i=1 n ∑ i=1 I(i−m)(j−n)K(m,n)(2.2) Typically, he laye s o hese ne wo ks a e made up o h ee s ages: con olu ion s age, de ec o s age, and pooling s age. In he con olu ion s age, mul iple pa allel con olu ions a e pe o med and esul in a collec ion o linea ac i a ions. A e ha , in he de ec o s age, a non-linea ac i a- ion unc ion is used on he linea ac i a ions. The p oblem a ises due o he ac ha hese laye s encode he in o ma ion in he p ecise posi- ion i is in, meaning ha i hese laye s p ocess he same inpu wi h a sligh change, i will esul in a di e en ou pu . This is whe e pooling unc ions a e use ul, making he ne wo k esis an , o in a ian , o small changes in he inpu . 18 Li e a u e Re iew •Using bo h VOC and COCO aining da a leads o inc eased accu acy: 83.6% and 82% s. 80.5% and 77.6%; •Using he i s 101 laye s o ResNe esul s in he bes accu acy: sa u a ion occu s a 80.5%; •The use o a RPN achie es supe io esul s han o he al e na i es: 79.5% s. 77.8% ( o he closes al e na i e). The a chi ec u e o he R-FCN emo es he expensi e compu a ion ha app oaches such as Fas e R-CNN use on each ROI (by passing each ROI h ough CNNs), in a o o a much mo e s aigh o wa d calcula ion o o e lap, enabling he model o be as e bu s ill main ain he accu- acy o e ed by a egion-based design. 2.3.2.3 SSD: Single Sho Mul ibox De ec o SSD (Fig. 2.7) [23] is one o he single-sho app oaches o objec de ec ion wi h he in en o being as and accu a e enough o be used in eal- ime applica ions. This de ec o emo es he need o egion p oposals, while s ill main aining accu acy. Figu e 2.7: O e iew o he SSD a chi ec u e. [23]. The i s s age in SSD is he gene a ion o ea u e maps a mul iple scales, which allows o he de ec ion o objec s o di e en sizes - ea u e maps wi h highe esolu ion a e able o de ec smalle objec s, while ones wi h lowe esolu ion can de ec bigge objec s. The backbone a chi- ec u e used is based on VGG, and o he con olu ional laye s a e appended a he end o allow o di e en scales. In he second s age, de ec ion is based on small con olu ional ke nels applied o each cell in he ea u e maps. Fo e e y cell, ou de aul bounding-boxes o di e en aspec a ios a e applied. These aspec a ios a e chosen o accommoda e objec s o di e en shapes and sizes (Eq. 4 in [23]). Fo each box, a class p obabili y sco e is compu ed, as well as di e en o se s o he size o he box (in o de o be e adjus he shape o he box o he shape o he objec ). This app oach is simila o ha o Fas e R-CNN bu applied o di e en esolu ions o ea u e maps. A pa icula ly in e es ing aspec in [22] is he use o "da a augmen a ion" o aining. This means ha , o each aining image, ei he he en i e image is used o a andom pa ch is ob ained 2.3 Rela ed Wo k 19 om i . In he case o a andom pa ch, i can also su e image dis o ions. The objec i e o da a augmen a ion is o enable he model o handle a ious shapes and sizes o objec s mo e obus ly. The model was e alua ed on he Pascal VOC and MS COCO da ase s, whe e he main conclu- sions we e he ollowing: •As wi h p e ious models, aining on bo h he VOC and COCO da ase s inc eases accu acy (Tables 1, 4 and 5 in [22]); •Using mul iple ou pu s laye s o mul iple ea u e map esolu ions esul s in highe accu acy (Table 3 in [22]); •Da a augmen a ion achie es be e esul s, as expec ed Table 6 in [22]); •SSD can achie e highe mAP han i s compe i o s using a smalle esolu ion inpu image Table 7 in [22]). 2.3.2.4 YOLO: You Only Look Once YOLO was i s in oduced in [24], su e ing subsequen al e a ions in [25] and some mino changes in [26]. I was in oduced as a single-sho app oach o objec de ec ion, by using a single ne wo k o sol e a eg ession p oblem o p edic bounding boxes and class p obabili ies o objec s in an image. The i s e sion o YOLO di ides i s inpu image in o a 7x7 g id, and each cell in he g id is esponsible o iden i ying an objec i i s cen e is loca ed wi hin he cell. Each cell p edic s a se o 2 bounding-boxes wi h associa ed con idence sco es, as well as class p obabili ies. The con idence sco es ake in o accoun he p obabili y o he boxes ha ing an objec wi hin hem and also he accu acy o he box in e ms o he shape o he objec . One p oblem ha can a ise om his app oach is he same objec being de ec ed by mo e han one cell. In his case, he cell wi h he highes con idence o a speci ic objec is chosen. The p edic ions o each cell a e made by classi ie s applied a e ea u e ex ac ion (which is done on he en i e image by he ne wo k in Fig. 2.8). Figu e 2.8: O e iew o he o iginal YOLO ne wo k. [24]. 20 Li e a u e Re iew In he second e sion o YOLO [25], bounding-boxes we e eplaced by he same ancho boxes used in Fas e R-CNN, since i made lea ning easie o he ne wo k and i also enabled each cell o p edic mo e han one objec . This had a small dec ease in accu acy, bu o e all ecall imp o ed. The numbe o ancho boxes chosen o his app oach was 5. Ano he change was eplacing he o iginal backbone ne wo k (Fig. 2.8) wi h a cus om ne wo k called Da kne -19, whose pu pose was o educe complexi y and imp o e accu acy. A signi ican imp o emen in his e sion was ha his model was able o de ec mo e han 9000 objec ca e- go ies, compa ed o he o iginal 20. The hi d e sion [26] came wi h mino changes, he mos impo an o which was he back- bone CNN being expanded o Da kne -53, a CNN wi h 53 con olu ional laye s ha is mo e pow- e ul han i s p edecesso , bu mo e e ec i e han o he used a ian s o ResNe . The mos ecen es esul s o YOLO we e p esen ed in [26]. They a e he esul s o es ing on he COCO da ase , and he ollowing conclusions we e made: •When compa ing mAP o he p ocessing speed (Fig. 3 in [26]), YOLO achie es he bes esul s; •In e ms o accu acy, YOLO is mo e accu a e han SSD, bu s ill in e io o wo-s age ap- p oaches, such as Fas e R-CNN. 2.3.2.5 Compa ison o Resul s The ollowing ables ea u e he esul s o mean a e age p ecision (in pe cen age) o each o he me hods in each da ase . The alues a e aken om he pape s ha in oduce each me hod. When a pape does no p esen esul s o a pa icula da ase , he alue is omi ed. Tables 2.1 and 2.2 p esen he esul s o he VOC07 and VOC12 da ase s, espec i ely. Table 2.1: Compa ison be ween me hods on he Pascal VOC07 da ase Au ho /Re e ence Yea Me hod mAP (%) Ren e al. [20] 2017 Fas e R-CNN 78.8 Dai e al. [22] 2016 R-FCN 83.6 Liu e al. [23] 2016 SSD 81.6 Redmon and Fa hadi [26] 2018 YOLO − Table 2.2: Compa ison be ween me hods on he Pascal VOC12 da ase Au ho /Re e ence Yea Me hod mAP (%) Ren e al. [20] 2017 Fas e R-CNN 75.9 Dai e al. [22] 2016 R-FCN 82.0 Liu e al. [23] 2016 SSD 80.0 Redmon and Fa hadi [26] 2018 YOLO − 2.3 Rela ed Wo k 21 Table 2.3 p esen s he esul s o he MS COCO da ase . On his da ase , mAP can be di ided in o wo me ics. The i s is he s anda d mAP when IoU be ween he objec bounding-box and he g ound- u h is 0.5. The second, is a mo e s ic me ic, o he mean a e age p ecision o e a ious h esholds, o an IoU be ween 0.5 and 0.95. Table 2.3: Compa ison be ween me hods on he MS COCO da ase Au ho /Re e ence Yea Me hod [email p o ec ed] (%) mAP@[.5,.95] (%) Ren e al. [20] 2017 Fas e R-CNN 42.7 21.9 Dai e al. [22] 2016 R-FCN 53.2 31.5 Liu e al. [23] 2016 SSD 46.5 26.8 Redmon and Fa hadi [26] 2018 YOLO 57.9 33.0 As can be seen, R-FCN and SSD a e he mos accu a e me hods on he VOC da ase s. SSD is he as es o he wo, being a one-sho app oach. On he MS COCO da ase , YOLO achie es he bes esul s. When IoU is be ween 0.5 and 0.95, he e is a signi ican d op in accu acy, since i is di icul o ob ain a high IoU. 2.3.3 Machine Lea ning in Pose Es ima ion 2.3.3.1 SSD-6D The objec i e o his ne wo k (Fig. 2.9), in oduced in [27], is o build upon he s uc u e o SSD o ob ain a collec ion o ou pu s: class p obabili ies, coo dina es o a 2D bounding-box, sco es o possible iewpoin s and in-plane o a ions. Figu e 2.9: O e iew o he SSD-6D model. [27]. The inpu image is passed h ough a backbone CNN based on Incep ionV4 o ou pu a se o ea u e maps a di e en scales. The same s uc u e as SSD is used o p edic ion o class and 22 Li e a u e Re iew eg ession o he bounding-box. The main changes occu in he sco ing o possible iewpoin s and in-plane o a ions. As is shown in Fig. 2.9, each map is con ol ed wi h a ke nel o shape (4 + C + V + R), whe e C ep esen s he numbe o objec classes, V he numbe o iewpoin s and R he numbe o in-plane o a ions. The iewpoin is he posi ion in 3D space om which he objec is iewed, which in luences he aspec o he objec as seen om he came a, while in-plane o a ion is seen as a ans o ma ion o he same iewpoin . As can be seen in Fig. 2.10, he iewpoin s a e sampled om a hal -sphe e a ound he objec . Fo symme ical objec s, only he a c in g een is sampled, and, o semi-symme ical objec s, only he poin s in ed a e used. In [27], i is a gued ha con olu ional laye s a e mo e e ec i e a sco ing a iewpoin and in-plane o a ion han using eg ession o p edic a se o ansla ions and o a ions. Fu he mo e, by sco ing iewpoin s wi h a con idence alue, all ha is abo e a ce ain h eshold can be accep ed, he eby dealing well wi h symme ical objec s. Figu e 2.10: Illus a ion o e e y iewpoin conside ed. [27]. Knowing he pa ame e s o he came a, he mos con iden sco es can be pooled o calcula e he ansla ion and he o a ion o he objec , esul ing in a se o 6-DOF pose hypo heses. Each hypo hesis goes h ough a pose e inemen s ep, which uses he I e a i e Closes Poin (ICP) al- go i hm (2.3.3.7). The bes pose is chosen by compa ing he e ined poses o he dep h da a om RGB-D. SSD-6D was e alua ed on bo h he Linemod and he Tejani da ase s. Howe e , since none o he o he me hods a e e alua ed on he las da ase , he e a e no compa isons o be made. Tejani da ase esul s will, he e o e, be omi ed. 2.3.3.2 PoseCNN PoseCNN [19] e olu ionized pose es ima ion algo i hms by decoupling pose es ima ion in o h ee sepa a e asks: seman ic labeling, 3D ansla ion es ima ion and 3D o a ion eg ession (as can be seen in Fig. 2.11). This app oach acili a es he job o he ne wo k by enabling i o model how each ask ela es o o he s. The backbone CNN akes an inpu image and ou pu s ea u e maps o di e en scales. The b anches co esponding o each ask hen use hese ea u e maps o calcula ions. 2.3 Rela ed Wo k 23 Figu e 2.11: O e iew o he PoseCNN sys em. [19]. In he seman ic labeling ask, wo 512 channel maps a e p ocessed in o de o ob ain one o he same scale as he o iginal image. He e, objec de ec ion is done by applying con olu ional ke nels o each pixel in he image and a ibu ing a seman ic label: each pixel is classi ied in o an objec class. This gi es be e esul s han he bounding-box app oaches used in SSD-6D, o example, since i handles occlusions be e . Fo 3D ansla ion es ima ion, he ne wo k needs o ou pu a ec o o coo dina es (Tx,Ty,Tz) o he objec cen e (in he came a coo dina e sys em). To ha ex en , each pixel ha belongs o an objec needs o o e o he cen e o he objec in he image coo dina e sys em, (cx,cy). A e ha , Txand Tycan be de i ed using Eq. 1 in [19]. Vo ing is implemen ed by a Hough o ing laye in he ne wo k and is done as ollows: 1. A eg ession ne wo k p edic s an a ay (nx,ny,Tz) o each pixel; 2. Each pixel o es o o he pixels (on whe he hey a e an objec cen e ) along he di ec ion o he ec o de ined by (nx,ny); 3. The pixel wi h he mos o es is chosen as he objec cen e . The mean o all alues o Tz o each pixel in an objec is chosen as he ue alue o Tz. Addi ionally, all pixels inside an objec a e known as inlie s. The bounding-box ha con ains all inlie s is also gene a ed in his s ep and is used o he 3D o a ion ask. In his ask, he objec i e is o ou pu a qua e nion, which ep esen s he es ima ed o a ion. The i s s ep is o apply 2 ROI pooling laye s using he bounding-boxes gene a ed in he p e ious ask. In ha way, he ea u e maps o each ROI can be ob ained. These ea u e maps a e hen added and ed in o 3 FC laye s, he las o which ou pu s a qua e nion. PoseCNN is e alua ed on he YCB-Video and he Occlusion LineMod da ase s. The ollowing conclusions can be made om ables 2 and 3 in [19]: •This model ou pe o ms coo dina e eg ession algo i hms; •Resul s a e a supe io when using pose e inemen ; •On he Occlusion LineMod da ase , PoseCNN wi h a pose e inemen algo i hm (ICP) achie es highe accu acy han o he me hods; 24 Li e a u e Re iew 2.3.3.3 DPOD: Dense Pose Objec De ec o DPOD [28] akes a di e en app oach o o he s men ioned p e iously. I s model (Fig. 2.12) is comp ised o h ee blocks: co espondence block, pose block, and a inal pose e inemen block. The co espondence block akes as inpu he desi ed image and uses an encode (based on he i s 12 laye s o ResNe ) and h ee decode s ha ou pu a co espondence map ( i s wo decode s) and he ID masks o each objec ( hi d decode ). Figu e 2.12: O e iew o he DPOD sys em. [28]. A co espondence map is a 2-channel image wi h alues om 0-255 ha maps a pixel in he image o a e ex in he objec 3D model. This esul s in mo e s aigh o wa d aining o he ne wo k and an o e all inc ease in quali y since i jus needs o sol e a colo classi ica ion p oblem o ma ch he image co espondence map o he 3D model co espondence map, ins ead o ha ing o eg ess he coo dina es o he objec . The ID masks a e a esul o he p obabili y ha an objec pixel belongs o a speci ic class (such as p e ious app oaches ha e used). Using he ID masks, he 3D model o each class can be ob ained. Bo h he co espondence map and he 3D models a e ed in o he pose block, which uses PnP+RANSAC o ou pu a o a ion ma ix Rand a ansla ion ec o T. As opposed o o he o ms o pose e inemen (2.3.3.7), he DPOD model u ilizes a pose e inemen block (Fig. 2.13) based on he ResNe a chi ec u e. This block akes in a pa ch o he o iginal image con aining he objec and a 3D ende ing o he objec in he es ima ed pose. These inpu s a e ed sepa a ely h ough wo b anches composed o he i s i e laye s o ResNe (E11 and E12), and he ou pu s a e sub ac ed and ed in o ano he ResNe -like ne wo k (E2). The di e ence ea u e ec o ha is ou pu ed is used o compu e he e o be ween he p e- dic ed pose and he ac ual pose. This is done by eeding i in o h ee sepa a e eg ession laye s (XY head, Z head and R head) ha co ec he coo dina es o ansla ion and he o a ion ma ix (as is seen in Fig. 2.13). 2.3 Rela ed Wo k 25 Figu e 2.13: Illus a ion o he pose e inemen block. [28]. The model was e alua ed on he LineMod and he Occlusion LineMod da ase s, ha ing ob- ained he ollowing conclusions: •Compa a i ely, DPOD achie es s a e-o - he-a esul s, wi h only PVNe (explained in Sec- ion 2.3.3.4) ha ing be e accu acy in pose es ima ion (Table 1 in [28]); •Pose e inemen imp o es accu acy by almos 10% (Table 1 in [28]); •DPOD’s no el pose e inemen echnique achie es be e esul s han DeepIM (Tables 1 and 5 in [28]). 2.3.3.4 PVNe : Pixel-wise Vo ing Ne wo k PVNe looks o imp o e upon he wo-s age app oach o keypoin de ec ion and pose es ima ion by using a model mo e obus o occlusion and unca ion. The main idea is o p edic , o each pixel in an objec , he uni ec o s om ha pixel o all he keypoin s in he objec - simila o wha was done in PoseCNN o he localiza ion o an objec cen e . This app oach is mo e obus o si ua ions whe e he keypoin s a e no isible in he image since i s posi ion can be in e ed om he o he isible pixels. In he i s s age o PVNe , a backbone CNN based on ResNe is used o ob ain class p ob- abili ies o each pixel (seman ic segmen a ion) and he uni ec o s ha ep esen he di ec ion om ha pixel o e e y keypoin . The o ing o each pixel is cha ac e ized in Eq. 1 and 2 o [29]. To educe a iance in localiza ion, he model needs o choose he keypoin s loca ed on he su ace o he objec . The ollowing algo i hm (Fa hes Poin Sampling) is used o choose a se o 8 keypoin s: 1. Add he cen e o he objec o he se o keypoin s; 2. Choose he a hes keypoin om he se and add i o he se ; 3. Repea he second s ep un il eigh keypoin s a e chosen. Finally, he model needs o compu e he 6-DOF pose using he se o keypoin s. This is done using an al e ed e sion o he PnP algo i hm ha akes in o accoun he mean and co a iance (unce ain y) associa ed wi h each keypoin (which a e calcula ed wi h Eq. 3 and 4 o [29]). 26 Li e a u e Re iew This me hod is e alua ed on he LineMod, Occlusion LineMod, and YCB-Video da ase s. I is concluded ha PVNe achie es he bes esul s in e ms o he ADD(-S) me ic when compa ed o o he app oaches (Table 3, 5, and 7 in [29]). 2.3.3.5 DenseFusion Mos app oaches so a ha e used 2D ea u es om RGB images and used dep h da a only o pose e inemen . DenseFusion’s [30] model (Fig. 2.14) akes ad an age o dep h in o ma ion om he i s s age by using RGB da a and poin cloud alues on a pe -pixel basis. Since bo h RGB and poin cloud da a a e e y di e en da a ypes, his model ea s each one sepa a ely and uses a pixel-wise dense usion algo i hm o combine hem. Figu e 2.14: O e iew o he DenseFusion a chi ec u e. [30]. In he i s s age, he same seman ic segmen a ion algo i hm applied in PoseCNN is used o ob ain segmen a ion masks o each objec . A bounding-box ha includes each mask is calcula ed and is used o ob ain an image c op in i s loca ion. The dep h da a o pixels inside he bounding box is also used in o de o con e i in o a 3D poin cloud ep esen a ion. The c opped image is ed in o an au oencode ne wo k (based on ResNe ) ha ou pu s colo in o ma ion. On he o he hand, poin cloud da a is p ocessed by a Poin Ne -like s uc u e o ob ain he geome ic in o ma ion o he objec . The usion algo i hm used combines colo (colo embeddings) and geome ic in o ma ion (ge- ome y embeddings) in o ea u e ec o s o each pixel in he objec . E e y ea u e ec o is hen p ocessed by a MLP wi h a e age pooling o ob ain a global ea u e ec o , which is hen appended o each pe -pixel ec o , c ea ing a pixel-wise ea u e (as shown in Fig. 2.14). The pixel-wise ea- u e is ed in o he pose p edic o ne wo k which ou pu s an es ima ion o he pose o e e y pixel wi h an associa ed con idence alue: Ri, i, and ci. 2.3 Rela ed Wo k 27 Figu e 2.15: Illus a ion o he pose e inemen ne wo k. [30]. DenseFusion also p esen s an i e a i e app oach o pose e inemen (Fig. 2.15) u ilizing a co ec ion ne wo k ha is composed o 4 FC laye s. Howe e , his ne wo k needs o lea n o co ec he pose ins ead o p edic ing a new one. The e o e, i s inpu s a e he colo embeddings o he image and he geome y embeddings o a ans o med poin cloud. This ans o med poin cloud is he esul o al e ing he inpu poin cloud by he ac o s ∆R and ∆ p edic ed by he pose esidual es ima o . This can be done i e a i ely, e ining he es ima ion u he and u he in each s ep. The inal pose p edic ed is ob ained by conca ena ing e e y pose ob ained in each i e a ion. E alua ion o he model is done on he LineMod and YCB-Video da ase s. The main conclu- sions ob ained in [30] a e he ollowing: 1. Wi h e inemen , DenseFusion achie es he highes accu acy when compa ed o he second bes me hod, PoseCNN+DeepIM ( ables 1 and 2 in [30]); 2. Pose e inemen inc eases pose accu acy by 4-8% ( ables 1 and 2 in [30]); 3. In Table 3 o [30], i can be seen ha DenseFusion is close o 200x as e han PoseCNN plus ICP; A pa icula ly ele an aspec ( o his hesis) in [30] is he use o DenseFusion in a obo ic g asping si ua ion. Ou o 60 a emp s, he obo has a 73% success a e while using DenseFusion as a pose p edic o . 2.3.3.6 Hyb idPose Mos app oaches ha e only used only one ype o in e media e ep esen a ion o an objec , he mos popula o which is a keypoin ep esen a ion. In [31], he au ho s a gue ha , in a eal- wo ld se ing, i is di icul o p edic keypoin s accu a ely based on a RGB image. Hyb idPose (Fig. 2.16) looks o sol e his p oblem by adding mo e in e media e ep esen a ions, enabling i s model o wo k be e in occlusion si ua ions and ha ing mo e in o ma ion abou he geome ical cha ac e is ics o he objec . 34 3D Objec Recogni ion and Localiza ion Sys em The Da a Acquisi ion Module has wo main unc ions. The i s unc ion deals wi h da a e ie al om he ZED came a; i uns on an in ini e loop, con inuously e ie ing he RGB image ame and he dep h map. The o he unc ion calcula es di e ences be ween consecu i e ames, using he S uc u al Simila i y (SSIM) algo i hm. This algo i hm ou pu s a alue, om 0 o 1, ha indica es how simila wo images a e, and he unc ion compa es his alue o a h eshold alue (calcula ions o his alue will be shown in Sec ion 4.3.1), in o de o indica e i he e we e changes be ween ames (such as an objec mo ing). In case o changes, he unc ion aises a global lag (indica ing ha a new ame is a ailable) and passes he new ame and i s co esponding dep h map o he Compu e Vision Module. Figu e 3.3: S a e machine o he pose es ima ion module In he Compu e Vision Module, a s a e machine (Fig. 3.3) is implemen ed. The e is one Ini ial s a e and ou main s a es: •De ec ion: implemen s he Objec De ec ion and Segmen a ion sys em explained in 3.2.2.1; •Localiza ion: implemen s he 6D Objec Pose Es ima ion sys em explained in 3.2.2.2; •Ready o G asp: chooses which objec o g asp and sends coo dina es o obo ic manipula- o ; 3.2 Sys em’s Modules 35 •G asping: wai s o g asping o inish and ies o de ec changes in ame. The ansi ion be ween s a es depends on he changes be ween ames ha a e de ec ed in module 1. I any objec in he ame mo es o disappea s, he whole s a e machine has o ese o he De ec ion s a e. 3.2 Sys em’s Modules 3.2.1 Da a Acquisi ion Module In he Da a Acquisi ion Module, he goal is o acqui e usable da a om he g asping wo kspace o he Compu e Vision module o p ocess and ex ac in o ma ion needed o g asping asks. In his case, he in o ma ion wan ed is 6D pose es ima ion and objec classi ica ion. Gi en his, and ega ding he equi emen o a 3D based solu ion in ou app oach, a dep h/3D came a was needed. In ou solu ion, gi en se e al epo ed uses and a ailabili y, he ZED s e eo ision came a and i s SKD we e used. The ZED came a was de eloped by S e eolabs and is composed o wo side-by-side RGB came as, which ou pu synch onized RGB images. The wo wide-angle lenses ha e 110◦ ield-o - iew, spaced a a baseline o 120 mm, allowing accu a e dep h es ima ion in he ange o 0.7 o 20 me e s. Came a pic u e, along wi h p oduc dimensions, a e shown in Fig. 3.4. The mos ele an cha ac e is ics o he ZED came a a e summa ized in Table 3.1. Figu e 3.4: ZED S e eo Came a and dimensions illus a ion Fo dep h acquisi ion, he ZED came a imi a es human ision, whe e each eye has a sligh ly di e en iew o he wo ld a ound, and by compa ing he wo iews, dep h can be in e ed. Like- wise, his came a, wi h i s sepa a e lens, can es ima e dep h by compa ing he pixels’ displacemen be ween he le and igh RGB images. The came a hen p o ides a dep h map (image con aining in o ma ion ega ding he dis ance o he su aces o scene objec s (z), o e e y pixel i,j) in he came a coo dina e sys em. This dep h is exp essed in me ic uni s and calcula ed om he le came a’s eye back o he g asping objec . 36 3D Objec Recogni ion and Localiza ion Sys em Table 3.1: ZED Came a ele an ea u es [1] Size and Weigh Dimensions: 175x30x33 mm Weigh : 159 g Indi idual image and dep h esolu ion HD2K: 2208 x 1242 (15 FPS) HD1080: 1920 x 1080 (30, 15 FPS) HD720: 1280 x 720 (60, 30, 15 FPS) WVGA: 672 x 376 (100, 60, 30, 15 FPS) Dep h Range: 1-20 m Fo ma : 32 bi s Baseline: 120 mm Lens Field o View: 110◦ /2.0 ape u e Connec i i y USB 3.0 (5 V / 380 mA) 0◦C o +45◦C SDK Sys em Requi emen s Windows o Linux Dual-co e 2.3 GHz 4 GB RAM N idia GPU The came a also p o ides a SDK o acili a ing came a con ol, bo h in C++ and Py hon. Fo ou implemen a ion, only he Py hon e sion was used. The ZED SDK is designed a ound OpenCV and CUDA lib a ies, wi h i s calib a ion and dep h es ima ion ou ines exploi ing CUDA’s pa allel GPU compu ing capabili ies. The SDK also o e s op ional dep h map p ocessing, includ- ing occlusion illing and edge sha pening. Un o una ely, hese buil -in pos -p ocessing echniques come wi h high compu a ional cos s, so hey could no be used in ou solu ion since eal- ime is a c ucial equi emen . The came a is con igu ed o un a 15 ps, wi h 720p esolu ion. The minimum dis ance o dep h es ima ion is 0.3m, and he dep h es ima ion was se oquali y mode, meaning i ou pu s mo e accu a e dep h alues a a small compu a ional cos . 3.2.2 Compu e Vision Module As shown in Fig. 3.1, he Compu e Vision Module is he mos impo an in he p oposed app oach, since i is h ough i s p ocessing capabili ies ha he objec ’s poses and ca ego ies a e iden i ied o pe o ming g asping ope a ions. This module is composed o wo di e en bu in e connec ed models: he Objec De ec ion and Segmen a ion model, and he Objec Pose es ima ion model. In he ollowing Subsec ions, de ails on hei implemen a ion will be p o ided. 3.2.2.1 De ec ion and Segmen a ion Model The objec de ec ion and segmen a ion model is composed o a p e-p ocessing s age and an in e - ence s age. In p e-p ocessing, he da a ecei ed om he Da a Acquisi ion module is p ocessed, so i can be ed in o he in e ence s age, which is made up o a Mask-RCNN model [34]. This s age 3.2 Sys em’s Modules 37 hen ou pu s a se o objec iden i ica ions, bounding-boxes, segmen a ion masks, and con idence alues o each objec in he came a ame. Figu e 3.5: O e iew o he Mask-RCNN model [34] In he p e-p ocessing s age, h ee RGB image ames and co esponding dep h maps a e e- cei ed by he pipeline, and an a e age o he pixel alues o bo h he RGB ames and he dep h maps a e calcula ed. This is o help educe he impac o small illumina ion di e ences, pixel shi ing, and andom noise when a ame is acqui ed, which in u n educes he licke ing e ec in he objec masks. This is discussed a g ea e leng h in Sec ion 4.3.2. A e wa d, he image ame is esized o i he inpu size o he in e ence s age. The Mask-RCNN model (o e iewed in Fig. ) is an ex ension o he Fas e R-CNN model 2.3.2.1, by applying a small FCN each ROI and ou pu ing a segmen a ion mask. As can be seen in he s a e-o - he-a analysis in sec ion 2.3.2.5, Fas e R-CNN is no he op-pe o ming model when i comes o pu e objec de ec ion, so why is i he one used? The answe is ha pu e objec de ec ion is no he only ac o in his p ojec . The inal ou pu should be he 6D pose o he objec , which means ha an addi ional model (in his case, he 6D pose es ima ion model discussed u he on, in Sec ion 3.2.2.2) is needed. One o he model equi emen s is he inpu o a segmen a ion mask o each objec and YOLO - he op-pe o ming model (in objec de ec ion), bo h in e ms o accu acy and speed - does no ou pu segmen a ion masks. Two di e en al e na i es we e analyzed: using YOLO wi h a segmen a ion mask ne wo k o e head (much like how Mask-RCNN imp o es upon Fas e R-CNN) o using Mask-RCNN ins ead. The conclusions we e ha Mask R-CNN is ac ually mo e accu a e and ma ginally as e han he o he al e na i e; he e o e, i is ac ually a be e choice o his implemen a ion. 3.2.2.2 Pose Es ima ion Model This sys em, as opposed o objec de ec ion and segmen a ion, wo ks on objec s one by one. I begins wi h a p e-p ocessing s age, which p ocesses he da a ecei ed om he objec de ec ion model o op imize he inpu s o he pose es ima ion s age. In pose es ima ion, a Dense usion 2.3.3.5 a chi ec u e is used o in e ence, and i ou pu s a se o o a ion qua e nions and ansla ion ec o s o each objec . A pose e inemen s age is hen applied o gi e a be e pose es ima ion. 38 3D Objec Recogni ion and Localiza ion Sys em Finally, a pos -p ocessing s age is execu ed o calcula e he inal op imal g asping poin coo dina es and g ippe pose o e e y objec in he ame. Due o he ZED came a limi a ions, he dep h map migh e u n alues such as in ini e and NaN. This occu s when poin s a e oo close/ a om he came a (in he case o in ini e alues), o when dep h canno be es ima ed due o occlusions (in he case o NaN). To make su e hese alues do no con ibu e o poo calcula ions, a mask o he dep h map is c ea ed whe e only eal alues a e masked. A e he dep h map mask is ob ained, i is mul iplied by he segmen a ion mask o ob ain all he eal dep h alues o only he poin s ha belong o he objec . Since he came a in he eal-wo ld en i onmen is posi ioned close o he objec s han he came a in he da ase , his means mo e poin s in he image will ep esen an objec , so choosing mo e poin s gi es us mo e geome ical in o ma ion abou he objec . Using mo e poin s is also a good way o making ou lie s (poin s in he mask wi h poo dep h alues) less in luen ial in he calcula ions. A se o 2000 poin s is hen andomly chosen om he objec dep h alues. The alue o 2000 poin s, a e some es ing, was chosen because i is a good ade-o be ween accu acy and speed. The ac ha he poin s a e chosen a andom also means ha he se o poin s will ha e a p ope dis ibu ion h oughou he whole objec . Wi h he poin s chosen, he poin cloud alue o each poin is calcula ed, ollowing Equa ions 3.1,3.2 and 3.3. z=dep h scale (3.1) x=(a−cx)×z x (3.2) y=(b−cy)×z y (3.3) In hese equa ions, dep h ep esen s he dep h alue o a poin , scale depends on he measu e- men uni s used by he came a, aand ba e he pixel coo dina es o a poin and cx,cy, x and y a e in insic came a pa ame e s. This is he same way he ZED came a calcula es i s poin cloud, howe e , he SDK calcula es a poin cloud o he whole image and each poin can be accessed indi idually, which inc eases compu a ion imes signi ican ly mo e. Following all his p e-p ocessing, he Dense usion model akes as inpu a c op o he came a image ame (in he dimensions o he objec ’s bounding-box), he calcula ed poin cloud, he 2000 poin s om he mask, and he objec ’s ID. The pose es ima ion is ou pu ed in he o m o o a ion qua e nions and a ansla ion ec o , which hen goes h ough wo i e a ions o pose e inemen , o p oduce be e esul s. Doing wo i e a ions is as e and has compa able accu acy o doing mo e i e a ions. The ou pu o Dense usion is no he inal pose es ima ion, howe e . These alues c ea e a ans o ma ion ma ix ha ans o ms he poin s in he objec 3D model om he canonical ame 3.2 Sys em’s Modules 39 de ined in he da ase o he came a coo dina e ame. To ob ain a inal pose es ima ion, pos - p ocessing is c ucial. In pos -p ocessing, he se o op imal g asping poin s (which we e loaded om one o he con igu a ion iles), a e ans o med acco ding o he ans o ma ion gi en by he Dense usion model. The cen e o his g oup o poin s is hen calcula ed o gi e he op imal poin o con ac o g asping. 40 3D Objec Recogni ion and Localiza ion Sys em Chap e 4 Expe imen s and Resul s 4.1 Da ase In o de o es and alida e ou p oposed sys em, and gi en ha a machine lea ning model is only as good as he da a i is ed, a da ase was chosen o ain and es he compu e ision module. The da ase used o ain bo h he objec de ec ion and he pose es ima ion models was he YCB- Video da ase , which uses 20 objec s om he YCB Objec Model Se [35]. Using a RGB-D came a, di e en g oups o he 20 objec s we e ilmed in di e en en i onmen s, wi h a ying backg ounds, di e en ligh ing condi ions, and si ua ions o occlusion. The came a mo ed in ela ion o he objec s o c ea e di e en iewpoin s o e e y objec , which esul ed in a o al o 133827 ames o ideo. Fu he mo e, he 3D poin cloud models o e e y objec we e ob ained wi h a scanning ig. F om a quan i a i e s andpoin , his da ase has 133827 RGB images (wi h 480x640 pixels) and hei co esponding dep h maps, as well as a collec ion o he poin cloud 3D models o all he objec s. Fo e e y ame and each o he objec s in he ame, he e a e anno a ions o class label, g ound- u h segmen a ion mask, and pose alues o he objec (gi en as a ans o ma ion, om he canonical ame o he objec o he came a coo dina e ame). The came a’s in insic pa ame e s and i s posi ions in he wo ld a e also gi en o each ame. Figu e 4.1: Dis ibu ion o numbe o ins ances pe objec 41 42 Expe imen s and Resul s In o de o e alua e i he da ase was imbalanced, he numbe o ins ances pe ype o objec in he da ase was plo ed. As can be seen in Fig. 4.1, he dis ibu ion o ins ances o each objec is ai ly balanced, and each objec has, a leas , 15000 ins ances. This con ibu es o a mo e balanced aining o he machine lea ning models and helps o educe bias. This cha ac e is ic explains why his da ase is one o he mos impo an and one he mos used o he aining and alida ion o he pose es ima ion models e iewed in Sec ion 2.3.3. F om a quali a i e s andpoin , he YCB-Video da ase p esen s a lo o ad an ages in e ms o a iabili y. One o hem is he di e ence in backg ound, which can be bo h unclu e ed (Fig. 4.2a) and clu e ed (Fig. 4.2e), as well as ligh o da k, and simple o complex (as is shown in all he images o Fig. 4.2). (a) (b) (c) (d) (e) ( ) Figu e 4.2: Selec ion o illus a i e images om he da ase When compa ing Figs. 4.2a and 4.2b, i can be seen ha he same objec s, in he same scene, a e shown in a a ie y o iewpoin s. This is mainly due o he mo emen o he came a, which helps o c ea e a subs an ial ange o poses o he aining o he pose es ima ion model. Ano he impo an aspec o he da ase is he change in ligh ing ha occu s be ween scenes. Compa ing Figs. 4.2b and 4.2d, he e a e examples o an o e exposed and an unde exposed image, espec- i ely. This a iabili y can help o make any model mo e obus o ligh ing changes. Finally, i is also impo an o men ion he way occlusions a e c ea ed o an objec . The a ying iewpoin s help c ea e na u al occlusions - when objec s o e lap - and unca ion o an objec - when i lea es he ield o iew o he came a (such as in Fig. 4.2c). Mo eo e , he a angemen o he objec s in he scene also helps o c ea e complex iewpoin s o an objec , whe e o e laps may occu ( o example, in igs. 4.2a,4.2d and 4.2e). All hese aspec s we e ele an o he choice o da ase , gi en ha wi h mo e a iabili y in di e en aspec s o each objec , he obus ness o he ained model will imp o e. 4.2 O line Tes ing Phase 43 4.2 O line Tes ing Phase In his sec ion, he aining and alida ion o each ML model in he Compu e Vision Module, will be discussed. The models we e ained and alida ed on he da ase analyzed in he sec ion abo e (Sec ion 4.1). A e his, he pose es ima ion model was es ed using as inpu s he ou pu s gi en by he objec de ec ion model. 4.2.1 T aining he Objec De ec ion Model The aining o his model was hea ily based on he idea o ans e lea ning. This me hod consis s o using p e- ained weigh s om a model - ob ained on one da ase - as a s a ing poin o aining on a di e en da ase . This app oach speeds up aining ime and inc eases pe o mance, bu only when he p e- ained model lea ned ele an gene al ea u es on a balanced da ase . In ou i s app oach, he p e- ained weigh s we e used o e e y b anch, excep he ne wo k heads; on he second app oach, he p e- ained weigh s we e u ilized only o he backbone ea u e ex ac ion ne wo k, excep o he inal laye s. The i s app oach esul ed in a 14% d op in mask gene a ion accu acy and a 5% d op in class label accu acy, in ela ion o he second app oach. Al hough he i s app oach only akes 33 hou s o ain compa ed o he 46 hou s o aining in he second app oach, he accu acy o he segmen a ion masks is pa amoun , so his is an accep able ade-o be ween speed and accu acy. Table 4.1: Compa ison o o iginal and ine- uned hype -pa ame e alues Hype -pa ame e O iginal alue Fine- uned alue Numbe _classes NA 21 Max_g _ins ances 100 10 De ec ion_min_con idence 0.8 0.7 ROI_miniba ch_size 512 128 RPN_nms_ h eshold 0.7 0.5 Lea ning_ a e 0.02 0.002 RPN_class_loss 1.0 1.0 RPN_bbox_loss 1.0 0.01 M cnn_class_loss 1.0 1.0 M cnn_bbox_loss 1.0 0.01 M cnn_mask_loss 1.0 10.0 Gi en ha he o iginal hype -pa ame e s o he model we e op imized o he COCO da ase , he hype -pa ame e s o he objec de ec ion model we e uned. Fo ha , Tenso boa d (a isu- aliza ion and op imiza ion ool o uning pa ame e s in Tenso low) was used. In Table 4.1, a compa ison be ween he o iginal and he ine- uned model pa ame e s is shown. The numbe _classes pa ame e akes he alue o 21, since he e a e 20 objec s in he da ase and one backg ound class. Max_g _classes was changed o 10 om 100 due o he ac ha each image in he da ase has a maximum o 10 objec ’s ins ances. To allow o mo e p edic ions 50 Expe imen s and Resul s da ase used in aining he model, i did no lea n how o di e en ia e hese ea u es in a complex objec co ec ly. In scene 2 (Fig. 4.3b), he h ee main changes we e he a ia ion o iewpoin o he bowl, he o e lap be ween one o he gela in boxes and he mug, and he pa ial occlusion o he ma ke by he smalle clamp. Like is shown in 4.7, his had an immedia e impac on he de ec ion o he gela in box. The e is only 50% accu a e de ec ions which co espond o he non-occluded box. The inaccu a e de ec ions o he bowl co espond o i being iden i ied as a mug, due o he shape in his iewpoin being simila o ha o a mug. The pa ial occlusion o he ma ke causes he ma ke o be unde ec ed. Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 153 50 0 50 0.85 Bowl 13 1 153 81.7 18.3 0 0.72 Mug 14 1 153 100 0 0 0.98 La ge ma ke 18 1 153 0 0 100 NA Clamp 19 2 153 0 74.5 25.5 NA Table 4.7: De ec ion esul s on scene 2 In scene 3 (Fig. 4.3c), he smalle gela in box is o e lapped on op o he bowl, c ea ing a pa ial occlusion o he bowl. This occlusion c ea es a iewpoin e y simila o he mug iewpoin , which explains he inaccu a e de ec ion a e (which can be seen in Table 4.8). The smalle gela in box is s ill well de ec ed e en on op o he bowl; howe e , he occluded gela in box emains unde ec ed, while i is occluded by he clamp. The ma ke , in an up igh posi ion, p esen s good de ec ion esul s. Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 116 50.0 0 50.0 0.89 Bowl 13 1 116 88.8 11.2 0 0.90 Mug 14 1 116 100 0 0 0.97 La ge ma ke 18 1 116 89.7 0 10.3 0.76 Clamp 19 3 116 0 66.7 33.3 NA Table 4.8: De ec ion esul s on scene 3 Scene 4 (Fig. 4.3d) p esen s all he objec s close oge he , wi h pa ial occlusions o he smalle gela in box. This p oximi y be ween objec s makes i ha d o he de ec ion model o 4.3 Online Tes ing Phase 51 co ec ly iden i y he objec s. In ac , he wo bes pe o ming objec s (acco ding o he esul s in Table 4.9) - he bowl and he mug - a e he mo e isola ed in he g oup. Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 115 4.3 0 95.7 0.72 Bowl 13 1 115 100 0 0 0.95 Mug 14 1 115 100 0 0 0.97 La ge ma ke 18 1 115 0 0 100 NA Clamp 19 2 115 0 44.8 55.2 NA Table 4.9: De ec ion esul s on scene 4 Scene 5 (Fig. 4.3e) p esen s a on al iew o he op o he bowl. Since his iewpoin is e y unique o his ype o objec , he de ec ion is 100% accu a e (Table 4.10). In a simila si ua ion o scene 4, he ma ke is unde ec ed due o i s p oximi y o he bowl and he wood block behind. Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 113 100 0 0 0.87 Bowl 13 1 113 100 0 0 0.87 Mug 14 1 113 72.6 0 27.4 0.74 La ge ma ke 18 1 113 0 0 100 NA Clamp 19 3 113 0 0 100 NA Table 4.10: De ec ion esul s on scene 5 As can be seen in Table 4.11, e en wi h he hea y occlusion o he bowl in scene 6 (Fig. 4.3 ), is s ill p esen s a easonable de ec ion a e. In his scene, he powe d ill was in oduced, howe e , due o i s occlusion, he iewpoin gene a ed is no ep esen a i e enough o he shape o he objec , in o de o he model o de ec i . 52 Expe imen s and Resul s Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 140 91.4 0 8.6 0.81 Bowl 13 1 140 88.6 0 11.4 0.83 Mug 14 1 140 100 0 0 0.97 Powe d ill 15 1 140 0 0 100 NA La ge ma ke 18 1 140 100 0 0 0.92 Clamp 19 2 140 0 0 100 NA Table 4.11: De ec ion esul s on scene 6 Fo scene 7, he esul s o Table 4.12 show ha e en when occlusion does no occu , ce ain posi ions o he gela in boxes cause he model no o be able o de ec i . This is due o he small amoun o iewpoin s in he educed da ase used o online es ing. All o he objec s, excep he clamp, ha e a 100% de ec ion a e. Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 209 50 0 50 0.83 Bowl 13 1 209 100 0 0 0.83 Mug 14 1 209 100 0 0 0.96 Powe d ill 15 1 209 100 0 0 0.83 La ge ma ke 18 1 209 100 0 0 0.92 Clamp 19 1 209 0 100 0 NA Table 4.12: De ec ion esul s on scene 7 Scene 8 (shown in Fig. 4.3h) p esen s a ai ly s anda d iew o he objec s, his ime wi h he emo al o he clamp om he wo kspace. Table 4.13 shows ha mos o he objec s a e well de ec ed, wi h he excep ion o he ma ke and he powe d ill. 4.3 Online Tes ing Phase 53 Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 143 97.2 0 2.8 0.83 Bowl 13 1 143 100 0 0 0.88 Mug 14 1 143 100 0 0 0.99 Powe d ill 15 1 143 24.5 0 75.5 0.73 La ge ma ke 18 1 143 0 0 100 NA Table 4.13: De ec ion esul s on scene 8 Finally, in scene 9 (Fig. 4.3i), e en wi h e y clea iewpoin s o he objec s, de ec ion esul s a e poo o he gela in box, he powe d ill and he clamp. Name ID Coun Numbe o ames Accu a e De ec ion (%) Inaccu a e De ec ion (%) No de ec ion (%) A e age Con idence (0-1) Gela in box 8 2 161 50 0 50 0.8 Bowl 13 1 161 100 0 0 0.91 Mug 14 1 161 100 0 0 0.99 Powe d ill 15 1 161 0 0 100 NA La ge ma ke 18 1 161 91.9 0 8.1 0.76 Clamp 19 2 161 0 50 50 NA Table 4.14: De ec ion esul s on scene 9 As can be seen, he educed da ase has a big impac on de ec ion pe o mance. The main issues wi h he use o so li le ins ances o each objec esul s in he ollowing p oblems: •Fo complex objec s (such as he clamp), he model needs mo e da a o be able o dis inguish ea u es be ween objec s, esul ing in si ua ions o inco ec classi ica ion; •E en o objec s wi h simple ea u es (like he gela in box, he bowl, and he ma ke ), he model o e - i ed o ce ain iewpoin s (which we e mo e p e alen in he educed da ase ), which esul ed in no de ec ion in ce ain objec posi ions; •When objec s we e oo close oge he , he model had ouble in co ec ly de ec ing hem, since his abili y is a mo e complex ea u e which needed mo e aining da a; Howe e , as in any o he p oblem, besides he issues, he e a e some posi i e aspec s wo h men ioning ega ding he model’s pe o mance: •The model is able o gene alize, and co ec ly iden i y objec s which a e di e en om he da ase , e en wi h limi ed da a. This is he case o he gela in boxes - which ha e di e en 54 Expe imen s and Resul s heigh and wid h in ela ion o he o iginal objec -, he bowl - which is alle and less wide han he o iginal model - and he mug - which has a di e en colo and is mo e na ow han he da ase model; •Wi h he limi ed amoun o aining da a, he model is s ill able o de ec accu a ely and wi h g ea con idence, he less complex objec s, like he gela in boxes, he bowl, and he mug; •E en wi h hea y occlusion (as seen in Fig. 4.3 ), he model can co ec ly de ec he bowl; One p oblem ha is common o bo h he educed da ase and he o iginal da ase is he lack o a ie y in he numbe o objec s pe image. Since his da ase was o iginally made o pose es ima ion, he main objec i e was o c ea e a complex en i onmen wi h p oximi y and occlusion be ween objec s. This is he igh app oach o pose es ima ion; howe e , in he objec de ec ion model’s aining, i can lead o o e - i ing. In es ing, when an objec was posi ioned alone in he wo kspace, he de ec ion a e d opped signi ican ly and was e y inconsis en . Howe e , when o he objec s we e posi ioned close o i , he de ec ion a e imp o ed and was mo e s able. This issue is due o o e - i ing o si ua ions wi h a g oup o objec s. 4.3.3 G asping Expe imen s In he hi d and inal es , he compu e ision module was alida ed. The es ing me hodology used was a g asping expe imen on he 5 o he 6 objec s ha we e chosen o es ing. This was due o he a ailable g ippe being oo small o be able o g asp he powe d ill. Figu e 4.4: Illus a ion o one g asping posi ion pe objec Each objec was placed in a di e en wo kspace posi ion, su ounded by o he objec s. A g asping a emp was made o each posi ion, o a o al o i e a emp s. Illus a ions o one o he 4.3 Online Tes ing Phase 55 posi ions o each objec a e p esen ed in Fig. 4.4. The esul s ob ained we e eco ded in Table 4.15. Scene Objec ID Success ul a emp s (%) A e age e o (cm) 1 8 60 2 2 8 60 5 3 13 20 2.7 4 14 80 0.5 5 18 60 3.5 6 19 40 0.75 Table 4.15: G asping esul s Fo each objec , he Compu e Vision Module an de ec ion, gene a ion o he segmen a ion mask, and pose es ima ion. The coo dina es o he cen e poin o he objec we e ob ained, and a ans o ma ion was applied om he cen e poin o he op imal g asping poin o enable co ec g asping. Due o he g ippe model no being eadily a ailable and also ime- es ic ions o es ing, he ans o ma ion o he se o op imal g asping poin s (discussed in Sec ion 3.2.2.2) was no implemen ed. Howe e , his es se es as a p oo o concep , since he ans o ma ion o all objec poin s is equi alen o he ans o ma ion o only a subse o op imal g asping poin s. Fu he mo e, since he p ima y pu pose o his p ojec is o de elop a sys em o objec ecog- ni ion and localiza ion, he main ocus was on he c ea ion o a pipeline wi h hose capabili ies whe e he pos -p ocessing s ep in ol ed in g asping was no he main ocus. The wo s pe o ming objec s a e he bowl and he clamp. These a e expec ed esul s when looking a bo h he esul s in Table 4.3 and in Sec ion 4.3.2. The clamp, being a complex objec wi h he wo s esul in Table 4.3 and wi h poo de ec ion esul s, was unlikely o p esen success ul g asping a emp s. As o he bowl, i was he second-wo s pe o ming objec in Table 4.15, due o being a symme ical objec . Bo h he gela in boxes ha e compa able success a es; howe e , he a e age e o is mo e signi ican o he la ge box, since small e o s in he pose’s es ima ion ha e a mo e subs an ial in luence in la ge objec s. One o he main sou ces o pose es ima ion e o du ing his phase o es ing was he incon- sis ency and inaccu acy o he segmen a ion masks ou pu ed by he objec de ec ion model. This was al eady an issue in Sec ion 4.2.2, howe e , he small acquisi ion e o s om he came a also con ibu e o he o e all de ec ion e o . Addi ionally, as is concluded in [1], he ZED came a has an associa ed e o o dep h da a (Fig. 4.5). Due o he wo kspace being a a maximum dis ance o 1.1m om he came a, his dep h es ima ion e o is e y small, howe e , i is s ill in luen ial. 56 Expe imen s and Resul s Figu e 4.5: RMS e o in dep h da a o 720p esolu ion [1] In his phase o es ing, gi en he eal- ime cons ains o g asping applica ions, he e was also a measu emen o he p ocessing ime o each componen : da a acquisi ion module, objec de ec ion, and segmen a ion model, and 6D objec pose es ima ion model. On a e age, ame acquisi ion and analysis ook 10ms. The objec de ec ion and objec localiza ion models ook on a e age 346ms and 29.6ms, espec i ely. Since he compu e ision module uns pa allel wi h ame acquisi ion, he whole p og am akes abou 375ms pe ame on a e age. This means i can un a abou 2.6 ps. A ame- a e o 2.6 ps is e y slow when applied o no mal ideo-playback o complex asks, such as eal- ime au onomous d i ing. Howe e , conside ing he no e y dynamic en i onmen on which his sys em aims o wo k in and conside ing he slow speed o a ull g asping a emp by he obo ic manipula o used o es ing (abou 5s), his ame- a e is adequa e. 4.4 Limi a ions Wi h he expe imen s and es s pe o med, some limi a ions in ou app oach we e ound. One o he limi a ions o he p oposed sys em is ha he ou pu o ideal g ippe o a ions is based on a p e iously c ea ed con igu a ion ile. While his is a alid solu ion ha esol es many e o s ha can occu when calcula ing hese o a ions au oma ically, while also allowing he use o choose he ideal g asping acco ding o his own p e e ence and applica ion, i also asks much o e head wo k om anyone who ies o implemen he sys em. Howe e , since he sys em’s main objec i e is o ou pu objec de ec ion and pose es ima ions o execu ing he g asping asks, i is no a e y impac ul limi a ion. O e all, he ade-o be ween allowing o use adap abili y and use o e head wo k may be bene icial. The de eloped sys em also a ec ed e o s in he acquisi ion o isual da a and ligh ing con- di ions due o he cha ac e is ics o he ML models used. Ligh ing is always a big pa o e o s 4.4 Limi a ions 57 in oduced in CV models, so his is an accep able limi a ion. The main way o sol ing his limi- a ion would be o ain he model on an e en bigge da ase wi h a ange o ligh ing condi ions, which would be oo cos ly. A se e e limi a ion is he ac ha he de ec ion a e d ops when an objec is by i sel in he wo kspace due o o e - i ing o si ua ions wi h a g oup o objec s. This is an issue o aining he de ec ion model on a pose es ima ion da ase . E en hough, when i comes o de ec ion, he sys em shows ha i gene alizes o objec s wi h di e en dimensions and cha ac e is ics om he da ase , in he case o pose es ima ion, objec s ha di e oo much om he da ase models will no be co ec ly es ima ed since he pose es ima ion pipeline akes in o accoun he canonical objec 3D model om he da ase . Ano he limi a ion is he ac ha he sys em canno choose he absolu e op imal objec o g asp in a clu e ed si ua ion. I only chooses he objec s ha i has he highes con idence o , wi hou aking in o accoun i s posi ion in ela ion o he o he objec s. 58 Expe imen s and Resul s Chap e 5 Conclusions and Fu u e Wo k 5.1 Conclusions In his hesis, he p oblem o de eloping an au oma ic objec ecogni ion and pose es ima ion sys- em was add essed. This sys em combines objec localiza ion and iden i ica ion o imp o e obo ic g asping. The main con ibu ion o his wo k was he de elopmen o a amewo k in eg a ing wo di e en ML models o c ea e a sys em o mo e comple e and au onomous g asping asks. Ex ensi e esea ch was ca ied ou on he cu en s a e o Machine Lea ning me hods, pa icu- la ly Con olu ional Neu al Ne wo ks. This esea ch esul ed in a deepe unde s anding o his ield and i s applica ions in Compu e Vision: om he mos basic o he mo e ad anced de elopmen s ha ha e occu ed h oughou he yea s. The main s a e-o - he-a app oaches o objec de ec- ion/ ecogni ion and pose es ima ion we e e iewed and compa ed, which led o mo e in o med de elopmen o he p oposed sys em. This hesis’ main ocus was on he in eg a ion o wo di e en ML models o de ec ion and pose es ima ion. Bo h Mask-RCNN and Dense usion we e co ec ly in eg a ed and show adequa e esul s in a eal es ing en i onmen . The de eloped sys em is capable o de ec ing mo emen s and di e ences be ween ideo ames, he e o e educing he need o cons an calcula ion and in e ence on each ideo ame. The p oposed solu ion uns en i ely au oma ically, a e an ini ial use pa ame e s’ selec ion and can gene a e g asping ou pu s om ideo da a a a a e o 2.6 ps. An in e es ing posi i e aspec o he p oposed solu ion is he abili y o a human o be in con ol o choosing he op imal g asping poin s o an objec and c ea ing a con igu a ion ile wi h hese poin s, hus allowing o use and applica ion adap abili y, while inc easing accu acy and speed o he sys em, since i does no es ima e hese poin s du ing un ime. The main limi a ions ound in he p oposed solu ion we e he o e head wo k needed o ou pu co ec g ippe o a ions o ce ain objec posi ions, he impac o ligh ing condi ions, and he dependency on p e iously scanned 3D models o he eal objec s. 59