______________________________________________
IMPLEMENTATION AND PERFORMANCE EVALUATION OF A
SEMANTIC IMAGE SEGMENTATION SYSTEM ON A MOBILE DEVICE
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _
IMPLEMENTACIÓN Y EVALUACIÓN DE RENDIMIENTO DE UN SISTEMA DE
SEGMENTACIÓN SEMÁNTICA DE IMÁGENES SOBRE UN DISPOSITIVO MÓVIL
______________________________________________
TRABAJO FIN DE GRADO
CURSO 2020-2021
AUTORA
ESTHER CARREÑO ALOCÉN
DIRECTORES
LUIS PIÑUEL MORENO
FRANCISCO D. IGUAL
GRADO EN INGENIERÍA INFORMÁTICA
FACULTAD DE INFORMÁTICA
UNIVERSIDAD COMPLUTENSE DE MADRID
MADRID,SEPTIEMBRE DE 2021
ACKNOWLEDGEMENTS
To he p ojec di ec o s Luis Piñuel Mo eno and F ancisco D. Igual, o all hei us
and suppo . To my good iend Ma ín Bá ez, who has always been he e o me. To my
boy iend, o helping me no o all in o he quicksand o my mind. [55]
ABSTRACT
The e has been a g ow h o in e es in seman ic segmen a ion in ecen imes,
and i s employmen in asks such as au onomous d i ing, medical diagnosis o ideo
su eillance is now c ucial. The aining and in e ence p ocesses in Deep Neu al
Ne wo ks (DNNs) a e pe o med in da a cen es, which causes unbea able la ency.
Edge Compu ing is a esponse o his limi a ion. Ne e heless, i is es ic ed by he
compu ing powe and ene gy consump ion in de ices.
This p ojec p oposes he implemen a ion o algo i hms o seman ic
segmen a ion o images using DeepLab and Tenso Flow as a basis, along wi h i s
adap a ion and h oughpu e alua ion in e ms o esponse ime in a mobile de ice
among di e en seman ic segmen a ion models.
A Raspbe y Pi was used along wi h a Co al USB Accele a o by Google, which
p o ides an Edge TPU o accele a e in e ence in Machine Lea ning by quan izing he
models. The inal goal is o p o e an e icien implemen a ion in his low ene gy
consump ion a chi ec u e.
Keywo ds
Seman ic segmen a ion, DNNs, Edge Compu ing, DeepLab, Co al, Ene gy
E iciency, La ency, Raspbe y Pi
RESUMEN
En los úl imos iempos ha habido un aumen o en el in e és po la segmen ación
semán ica, y su uso en a eas como la conducción au ónoma, el diagnós ico médico
o la ideo igilancia es c ucial. Los p ocesos de en enamien o e in e encia en Redes
Neu onales P o undas (DNNs) se ealizan en cen os de da os, lo que causa una
la encia insos enible. El Edge Compu ing es una espues a a es a limi ación, pe o es á
es ingido po la po encia compu acional y el consumo de ene gía de los disposi i os.
Es e p oyec o p opone la implemen ación de algo i mos de segmen ación
semán ica de imágenes usando como base DeepLab y Tenso Flow, además de su
adap ación y e aluación de endimien o según el iempo de espues a en disposi i os
mó iles en e di e en es modelos de segmen ación semán ica.
Pa a ello se ha u ilizado una Raspbe y Pi y se ha op ado po el acele ado USB
Co al, de Google, que o ece un Edge TPU pa a acele a la in e encia en Machine
Lea ning median e la cuan ización. El obje i o inal es demos a que una
implemen ación e icien e en una a qui ec u a de bajo consumo ene gé ico es posible.
Palab as cla e
Segmen ación semán ica, DNNs, Edge Compu ing, DeepLab, Co al, E iciencia
Ene gé ica, La encia, Raspbe y Pi
CONTENT INDEX
Acknowledgemen s III
Abs ac IV
Resumen V
Con en Index VII
Figu e Index XI
Table Index XIII
Capí ulo 1 - In oduc ion 1
1.1 Mo i a ion 1
1.1.1 Seman ic segmen a ion 1
1.2 Goals 3
1.3 P ojec plan 3
Capí ulo 2 - S a e o a 5
2.1 Con olu ional Neu al Ne wo ks 6
2.2 T ans e Lea ning 9
2.2.1 T ans e Lea ning in Edge-TPU 9
2.3 Quan iza ion 10
2.3.1 How o quan ize 10
2.3.3 Risks o quan iza ion 12
Capí ulo 3 - Ha dwa e de ices and a chi ec u es used 13
3.1 Lap op i7 NVIDIA® GeFo ce® GTX 13
3.2 Vi ual Machines 13
3.3 Raspbe y Pi 4 Model B 8GB RAM 13
3.4 USB Accele a o Google Co al 15
3.4.1 Edge TPU 16
Capí ulo 4 - F amewo ks used 17
4.1 Tenso Flow 17
4.2 Tenso Flow Li e 17
4.2.1 Pos - aining quan iza ion 18
4.2.1.1 Full in ege pos - aining quan iza ion 19
4.2.1.2 Tenso Flow Li e con e e 20
4.3 DeepLab 21
4.3.1 How does DeepLab wo k 21
4.3.1.1 Spa ial py amid pooling 22
4.3.1.2 A ous Con olu ions 22
4.3.1.3 Dep hwise Sepa able Con olu ions 23
4.3.2 Da ase s 23
4.3.3 Model Zoo 24
4.3.4 E alua ion me ics 26
4.3.5 Ad an ages o DeepLab 26
4.4 OpenCV 26
Capí ulo 5 - Implemen a ion and expe imen a ion 27
5.1 DeepLab local ins alla ion 27
5.1.1 Lap op ins alla ion 27
5.1.1.1 T aining 28
5.1.1.2 E alua ion 30
5.1.1.3 Visualiza ion 31
5.1.1.4 Sc ip (d a ) 32
5.2 Vi ual Machine 34
7
5.2.1 Vi ual Machine con igu a ion 35
5.2.2 Upda ed sc ip 35
5.2.3 Measu ed imes o each model 36
5.3 DeepLab in Raspbe y Pi 4 37
5.3.1 Raspbe y Pi con igu a ion 37
5.3.2 Measu ed imes o each model 38
5.4 Co al USB Accele a o in Raspbe y Pi 4 44
5.4.1 Co al ins alla ion 44
5.4.2 How o un a model on he Edge TPU 44
5.4.3 How o un TFLi e objec de ec ion models on he Raspbe y Pi 45
5.4.4 How o un Edge TPU objec de ec ion models on he Raspbe y Pi using
Co al USB Accele a o 46
5.4.5 How o quan ize a model o un i on Edge TPU 48
Capí ulo 6 - Resul s Compa ison and Analysis 53
6.1 Lap op s. Vi ual Machine execu ion imes 53
6.2 Vi ual Machine image inpu s. Vi ual Machine webcam inpu 53
6.3 Vi ual Machine s. Raspbe y Pi 54
6.4 Raspbe y Pi: egula model s. Edge TPU model 55
Capí ulo 7 - Conclusions and u u e wo k 57
Bibliog a ía 59
Apéndice A - Con igu a ion Guidelines 65
8
Figu e 1-1. Example o seman ic image segmen a ion. [18]
As he Figu e 1-1 shows, seman ic segmen a ion labels each pixel in he image
wi h a ca ego y label, howe e i does no di e en ia e ins ances. This means ha i
does no sepa a e he objec s o he same ca ego y. In his example, we can see ha
bo h cows in he image a e conside ed as a whole uni o pixels ha co espond o he
cow ca ego y.
The e has been a g ow h o in e es in his ield, mo i a ed by he e i al o he
Deep Neu al Ne wo ks (DNNs) and he imp o emen o p ocesso s’ compu ing
capaci y. Howe e , i has o be aken in o conside a ion ha he aining and in e ence
p ocesses in DNNs a e pe o med in da a cen es (in cloud), which o en leads o an
unbea able esponse ime o la ency. In esponse o his limi a ion, Edge Compu ing is
being used, which consis s in pe o ming he ope a ions nea he de ice which needs
he calcula ions. This app oach educes he la ency, bu i is limi ed by he compu ing
powe as well as he ene gy consump ion in hese de ices.
No e ha in e ence e e s o he p ocess o making p edic ions on unseen da a
using ained DNNs model.
2
1.2 Goals
This p ojec aims o s udy a neu al ne wo k model and o compa e i s me ics, like
p ecision and la ency, depending on he ained models used, hei da ase s and he
a chi ec u e by c ea ing a seman ic segmen a ion applica ion.
The inal goal is o p o e an e icien implemen a ion in a low ene gy
consump ion a chi ec u e.
1.3 P ojec plan
Table 1-1. Gan diag am wi h an o e iew he asks done
3
4
Capí ulo 2 - S a e o a
T aining and in e ence p ocesses in DNNs a e usually pe o med in da a cen es.
Al hough some cloud se ices, such as AWS, p o ide in e ence se ices like Amazon
Polly, Rekogni ion and Lex, i is no enough. When de ices send da a o he cloud o
p ocessing, he ecep ion o he inal esul s is o ally dependen on he conges ion as
well as he la ency o in e ne connec ion. Mo eo e , he e is a lack o e iciency in
ega ds o cos and ene gy when s eaming dense in o ma ion, like images and ideo.
I should be no ed ha when wo king wi h eal- ime applica ions, esul s need o be
e u ned soon o main ain a posi i e use expe ience and o mee he equi emen s o
he applica ion. Because o his, an adequa e in e ence la ency is c ucial. This means
ha he as e he p edic ion is, he highe he numbe o p edic ions pe ime uni we
ge is, and hence a gene ally educed cos . [23]
In esponse o hese limi a ions o cloud, Edge Compu ing comes in o play; i
consis s in pe o ming he ope a ions nea he de ice which needs hose calcula ions.
This app oach educes he la ency, bu i is limi ed by he compu ing powe as well as
he ene gy consump ion in hese de ices.
5
Table 2-1. Compa ed in e ence imes o DNNs models o edge and cloud compu ing. [23]
Table 2-1 shows a compa ison be ween he in e ence imes ha we can ind o
di e en DNNs models o bo h cloud and on-p emise (edge) compu ing. F om his
in o ma ion i can be ga he ed ha he e is a balance be ween accu acy and la ency.
In addi ion, i also compa es GPU and CPU. La e on in his p ojec TPU will also be
discussed, as he USB Accele a o Google Co al p o ides an Edge TPU cop ocesso o
he Raspbe y Pi 4 Model B, which will educe he la ency e en mo e.
2.1 Con olu ional Neu al Ne wo ks
Compu e Vision (CV) is an a ea o A i icial In elligence ha imi a es he human
ision sys em, in such a manne ha compu e s can iden i y and p ocess i ems in
images and ideos jus like humans do. [25]
6
Figu e 2-1. Example o a CNN sequence o classi y handw i en numbe s. [36]
Con olu ional Neu al Ne wo ks (CNNs) a e Deep Lea ning algo i hms used in he
ield o CV. CNNs use an image as an inpu and gi e a alue o di e en he objec s in
ha image in o de o dis inguish all o hem. This kind o ne wo k is equi alen o he
neu on connec ions in he human b ain. [54]
Figu e 2-2. 4x4x3 RGB Image. [36]
7
Figu e 2-2 ep esen s he h ee colo planes o a RGB image. In his case, i is only
4x4 pixels, bu e en when he image has bigge dimensions, he CNN educes he
p ocessing complexi y while keeping he necessa y in o ma ion o a good p edic ion.
In o de o implemen he i s con olu ion ope a ion in he con olu ion laye , o
a 5x5x1 inpu image he e is a Fil e (K), which is selec ed o be 3x3x1.
K = 101, 010, 101 ( )
Figu e 2-3. Con olu ioning p ocess. [36]
Then, he il e shi s 9 imes pe o ming a ma ix mul iplica ion by he
co esponding po ion (P) o he image. Wi h he con olu ion ope a ion i is aimed o
ex ac high-le el a ibu es om he image.
A e he con olu ional laye , he pooling laye educes as well he size o he
image. This way, he p ocessing o he da a does no need o employ such a high
compu a ional powe . Max Pooling is he bes ype o pooling, gi en ha i e u ns he
maximum alue om P and i is a noise supp essan , which means ha i ejec s he
noisy da a ha does no p o ide any use ul in o ma ion.
As s a ed in [36], “by adding a Fully-Connec ed laye he ne wo k lea ns
non-linea combina ions o high-le el ea u es”. Then he image ma ix is la ened in o
a column ec o , which is used as an inpu o a eed- o wa d neu al ne wo k and
backp opaga ion is p ac iced in each i e a ion o he aining p ocess o compu e he
g adien and loss unc ion.
8
Finally, he ea u es o he images a e classi ied using he So max Classi ica ion
echnique [45] by il e ing he alues which a e below a maximum es ablished alue.
[36]
2.2 T ans e Lea ning
T ans e Lea ning (some imes also called " ine uning") ocuses on applying he
knowledge p e iously ob ained o one ask o sol e simila p oblems. [37] In his way, i
is no necessa y o lea n e e y hing om sc a ch, which would consume ime and
compu ing esou ces.
Hence, he key mo i a ion o using ans e lea ning is ha ge ing a la ge
quan i y o labeled da a is necessa y o models which sol e complex asks, bu
labeling ha da a can be ha d in ega ds o he ime and labo employed. Mo eo e ,
ans e lea ning may be highly needed in case he da a can be eadily ou da ed
because he p e iously ob ained labeled da a may no ollow he same dis ibu ion
la e . [29]
In seman ic segmen a ion, and gene ally in compu e ision, p e- ained models
based on big CNNs a e used in ans e lea ning. [24]
To sum up, we can use a model ha is al eady ained o pe o m ex a aining
using a smalle aining da ase o each he model new classi ica ions.
2.2.1 T ans e Lea ning in Edge-TPU
Using ans e lea ning, i is possible o e ain an exis ing model compa ible wi h
Edge TPU. This can be done by wo me hods, which a e explained a [44]:
●“Re aining he whole model by adjus ing he weigh s ac oss he whole ne wo k.”
●“Remo ing he inal laye ha pe o ms classi ica ion, and aining a new laye
on op ha ecognizes you new classes. Once you' e happy wi h he model's
pe o mance, simply con e i o Tenso Flow Li e and hen compile i o he
Edge TPU. This will be explained la e in he wo k.”
9
2.3 Quan iza ion
As explained p e iously, NNs employ high compu a ional cos s and consume a
lo o memo y. Because o his, i is impo an o op imize i s aining and in e ence.
Nowadays, mo e models mo e om se e s o edge because o he ad an ages ha
edge o e s ega ding la ency. When unning models on he edge, ne wo k
op imiza ion is e en mo e c ucial, aking in o accoun hei limi a ions in compu ing
powe . One echnique o educing he complexi y o CNNs is quan iza ion.
The main idea o quan iza ion is o con e loa ing poin weigh s and inpu s in o
he nea es in ege s. This helps o consume less memo y and may lead o as e
calcula ions depending on he ha dwa e. In o he wo ds, “i makes he model smalle
and as e ”. [32] E en i he 8-bi in ege depic ion can be less p ecise, he in e ence
accu acy o he NN is no ema kably comp omised. [31]
The e a e se e al ways o ep esen ing loa ing poin s and in ege s, bu loa 32
and in 8 (which anges om -128 o 127) a e he mos common o quan iza ion.[11]
Al hough he calcula ion speed depends on he ha dwa e used, in gene al, in 8
is ypically as e han loa 32. Ne e heless, i mus be aken in o accoun ha loa 32 is
used by de aul o aining and in e ence o NNs. [3]
2.3.1 How o quan ize
We can pe o m quan iza ion h ough ma ix mul iplica ions. This p ocedu e is an
app oxima ion; hence, we lose in o ma ion in he p ocess. Howe e , his may be
accep able some imes. [11]
The e a e wo main ways o pe o m quan iza ion:
●Pos - aining quan iza ion: I consis s in aining he model using he de aul
loa 32 weigh s and inpu s, and la e quan izing he weigh s. I is easy o pe o m,
bu he disad an age is ha i can lead o accu acy loss.
●Quan iza ion-awa e aining: The weigh s a e quan ized du ing he aining
phase. in 8 quan iza ion o e s be e esul s, bu i is a mo e complex op ion.
10
Table 2-2. Quan iza ion me hods and hei pe o mance in Tenso Flow Li e. [26]
As we can see in Table 2.2, he e a e mul iple quan iza ion echniques a ailable
o Tenso Flow Li e.
Figu e 2-4. Quan iza ion echnique decision ee. [26]
Figu e 2-4 shows a decision ee o help selec he mos adequa e quan iza ion
me hod. I akes in o accoun he size and he expec ed p ecision o he model.
11
Figu e 4-1. Wo k low o c ea e a model o Edge TPU. [13]
Figu e 4-1 explains he wo ways o gene a ing he compa ible Edge TPU model:
quan iza ion-awa e aining, and pos - aining quan iza ion. These we e p e iously
in oduced. As shown in he igu e, quan iza ion-awa e aining is a much b oade
p ocess. Fo he sake o simplici y, I will ocus on pos - aining quan iza ion. [13]
4.2.1 Pos - aining quan iza ion
“Pos - aining quan iza ion [31] is a con e sion echnique ha is able o educe
he model size as well as imp o ing he [...] la ency”, wi hou comp omising he
accu acy oo much. The quan iza ion can be implemen ed by con e ing a ained
Tenso Flow model o TFLi e o ma using he TFLi e Con e e . As men ioned be o e, i is
necessa y o simpli y models o be able o b ing p e ained models o a chi ec u es wi h
less esou ces.
Table 4-1. Pos - aining quan iza ion echniques. [31]
18
4.2.1.1 Full in ege pos - aining quan iza ion
As Table 4-1 indica es, he e a e se e al echniques o choose om. Gi en ha
Co al USB Accele a o will be used, he chosen echnique is Full in ege quan iza ion,
since i suppo s Edge TPU. I s main bene i s a e ha i makes he model 4x smalle and
p o ides 3x mo e speedup. [31]
This quan iza ion echnique is used o ans o m an al eady ained ne wo k in o
a quan ized model, hence i does no equi e modi ica ions o he ne wo k. [32]
“In addi ion, i o e s u he la ency imp o emen s, cu s down peak memo y
usage, and is compa ible wi h in ege -only ha dwa e de ices o accele a o s” [31]
(Co al in his case).
I is equi ed o calib a e he ange o all loa ing-poin enso s in he model. To
calib a e a iable enso s like model inpu o ou pu s, i is necessa y o un a ew
in e ence cycles wi h a small ep esen a i e da ase . [16] Fu he in his p ojec , I
explained ha o my own quan iza ion I used dummy images and webcam images as
ep esen a i e da ase s,
19
4.2.1.2 Tenso Flow Li e con e e
This con e e ans o ms a TF model in o a TFLi e model (an op imized Fla Bu e
. li e). The e a e wo ways o use he con e e : by he Py hon API o by he command
line.
Figu e 4-2. TFLi e con e sion diag am. [41]
To c ea e a TFLi e model o he Edge TPU, he i s s ep is o con e he model o
TFLi e. Bewa e ha o c ea e a compa ible model wi h pos - aining quan iza ion,
Tenso Flow 1.15 mus be he ins alled e sion ( o se inpu and ou pu ype o in 8); i is
no possible o use Tenso Flow 2.0 because i only suppo s loa inpu s and ou pu s.
Finally, i is necessa y o compile he model o compa ibili y wi h he Edge TPU. [19]
20
4.3 DeepLab
“DeepLab [35] is a s a e-o -a deep lea ning model o seman ic segmen a ion
o images ha aims o assign meaning ul labels o all pixels o he inpu image.” I was
designed and open-sou ced by Google in 2016. I has se e al ea u es and many
imp o emen s ha e been made since hen, coun ing DeepLab V2 [4], DeepLab V3 [5]
and DeepLab V3+. [6]
Figu e 4-3. Segmen a ion esul example on Flick image. [53]
DeepLab allows he use s o ain he model, e alua e he esul s in e ms o mIoU
(mean in e sec ion-o e -union) and isualize he segmen a ion esul s.
4.3.1 How does DeepLab wo k
To unde s and DeepLab we will ocus on DeepLab 3+ [6], i s la es e sion. The
model consis s o wo s eps: [28]
●Encoding phase: The goal is o ge he key in o ma ion om he image. This will
be done by a p e- ained CNN.
●Decoding phase: The p e iously ex ac ed in o ma ion is used o ebuild ou pu o
co ec dimensions.
21
4.3.1.1 Spa ial py amid pooling
To gua an ee ha he model is obus o changes in objec sizes, spa ial py amid
pooling (SPP) ne wo ks a e employed, which ake mul i-scale in o ma ion because hey
use di e en scaled a ian s o he inpu du ing aining. Then, he c ucial ea u es ha
can ep esen mos in o ma ion a e combined.
Figu e 4-4. Spa ial py amid pooling. [28]
“Encode -Decode ne wo ks ans o m he inpu in o a dense o m ha can
ep esen all he inpu in o ma ion.” [28]
4.3.1.2 A ous Con olu ions
●To mi iga e he enhancemen in he compu a ional and memo y
equi emen s o aining caused by SPP,au ous con olu ions a e in oduced.
●A ous Con olu ions ge in o ma ion om a b oade e ec i e ield o iew,
bu hey keep equal compu a ional complexi y. [52] I s gene alized o m (in which
no mal con olu ion has a io =1) is:
As ous Spa ial Py amid Pooling (ASPP) is he combina ion o SPP wi h au ous
con olu ions. They no mally consis o 4 pa allel ope a ions.
22
4.3.1.3 Dep hwise Sepa able Con olu ions
This echnique allows o educe he compu a ion numbe when pe o ming
con olu ions. Fo ins ance, he inpu is 12 x 12 x 3 and a con olu ion o 5 x 5 is desi ed,
which gi es an ou pu 8 x 8 x 1. Fo his pu pose, he con olu ion is di ided in wo s eps:
●Dep hwise con olu ion: In his s ep he con olu ion 5 x 5 x 1 is pe o med, and he
ob ained ou pu is 8 x 8 x 3.
●Poin wise con olu ion: Then, he channel numbe mus inc ease. 1 x 1 ke nels a e
used wi h 3 as he dep h o he inpu . Wi h his 1 x 1 x 3 con olu ion, he ou pu is
8 x 8 x 1. To inc ease he numbe o channels, he con olu ions 1 x 1 x 3 can be
applied as much as wan ed.
4.3.2 Da ase s
I ha e s udied mainly hese hese da ase s o ain he models, which ha e
di e en cha ac e is ics:
●PASCAL VOC 2012 [30]: I includes abou 1400 images o aining and alida ion
and 20 objec ca ego ies such as animals, ehicles and o he daily li e objec s.
●ADE20k [2]: This da ase con ains mo e han 27K images and o e 3K objec
ca ego ies. I is e y dense, as he e a e many objec s in he images; his does no
happen, o example, in PASCAL VOC 2012, which includes ew objec s in each
image.
●Ci yscapes [7]: I is ocused on s e eo ideo sequences eco ded in s ee s and
oads om 50 ci ies. The e a e abou 3000 images classi ied in 30 ca ego ies and
di ided in 8 g oups (na u e, sky, humans…). These images ha e been chosen om
key ames o he ideos.
23
4.3.3 Model Zoo
This model zoo [51] o e s DeepLab models ained on PASCAL VOC 2012,
Ci yscapes and ADE20K.
I ha e wo k he mos wi h PASCAL VOC 2012, whose di ec o y includes:
●A ozen in e ence g aph: ozen_in e ence_g aph.pb
●A checkpoin : model.ckp .da a-00000-o -00001,model.ckp .index
These checkpoin s used we e:
●mobilene 2_coco_ oc_ ainaug
●mobilene 2_coco_ oc_ ain al
●xcep ion65_coco_ oc_ ainaug
●xcep ion65_coco_ oc_ ain al
The checkpoin s ha e been p e- ained on he PASCAL VOC 2012 ain_aug se
o ain_aug + ain al se .
Become awa e ha MobileNe - 2 based models do no use ASPP (seen be o e)
and decode modules o as compu a ion. This implies ha a a glance we could
deduce ha i s accu acy is lowe han in Xcep ion_68.
Table 4-2. Compu a ion complexi y (in e ms o Mul iply-Adds and CPU Run ime) and segmen a ion
pe o mance (in e ms o mIOU) on he PASCAL VOC al o es se . [51]
24
Looking a he PASCAL mIOU [10] i shows ha he p ecision is indeed lowe in
MobileNe - 2 han in Xcep ion_68. Howe e , he imes a e as e in MobileNe - 2.
Le ’s y he model wi h he lowes mIOU (mobilene 2_coco_ oc_ ainaug) and
he one wi h he highes mIOU (xcep ion65_coco_ oc_ ain al) wi h his inpu [1].
Figu e 4-5. Example mobilene 2_coco_ oc_ ainaug.
Figu e 4-6. Example xcep ion65_coco_ oc_ ain al
As said be o e, a a glance i is possible o see ha i s accu acy is highe in
Xcep ion_68.
25
4.3.4 E alua ion me ics
The me ic in which he accu acy is measu ed is MIoU (Mean In e sec ion o e
Union).
I consis s in calcula ing he IoU be ween he g ound u h and he ou pu
p edic ed by he NN. This conside s he egion common o bo h and calcula ed he
simila i y pe cen age in compa ison o he ac ual one. [10]
4.3.5 Ad an ages o DeepLab
The e a e h ee main ad an ages o using DeepLab: [4]
●Speed: hanks o a ous con olu ion.
●Accu acy: s a e-o - he-a esul s a e ob ained on complex da ase s, such as
PASCAL VOC 2012.
●Simplici y: his sys em consis s o wo ixed modules, DCNNs and CRFs.
4.4 OpenCV
OpenCV [27] is an Open sou ce Compu e Vision lib a y. I ha e used i in he Py hon
p og am ha I w o e o be able o cap u e mo emen and ecognize objec s h ough
my webcam.
26
Capí ulo 5 - Implemen a ion and expe imen a ion
As seen in he p ojec plan, a e esea ching seman ic segmen a ion and
DeepLab, and expe imen ing wi h a DeepLab demo, I p oceeded o ins all DeepLab
and es models locally, bo h in my lap op and in a i ual machine.
5.1 DeepLab local ins alla ion
5.1.1 Lap op ins alla ion
As a i s s ep, I cloned he gi eposi o y o DeepLap [35] in my Windows lap op
wi h he help o Gi Bash in o de o implemen a local ins alla ion. I ins alled a
Tenso low e sion compa ible wi h GPU (py hon -m pip ins all enso low-gpu==1.15.3)
as well as some lib a ies and d i e s equi ed o i s co ec unc ioning in GPU, such as
CUDA 9. The so wa e dependencies able can be ound he e [9].
When e e y hing is ins alled, y he model_ es .py o es ha i is wo king
co ec ly.
Figu e 5-1. Console ou pu when unning model_ es .py
27
As s a ed p e iously,MobileNe - 2 models do no use ASPP and decode
modules o as compu a ion. Table 5-1 shows ha - al models ha e sligh ly highe
in e ence execu ion imes han i s -aug e sion. I mus be aken in o conside a ion ha
he models ha e been p e- ained on PASCAL VOC 2012 ain_aug se o ain_aug +
ain al se , which means ha - al models ha e highe p e aining han -aug models.
This could explain he highe accu acy o xcep ion_coco_ oc ain al and
mobilene 2_coco_ oc ain al ha can be obse ed in he igu es abo e.
5.2 Vi ual Machine
In he i ual machine I ins alled DeepLab locally jus like I did in my lap op.
Howe e , a e expe imen ing wi h he local e sion and ob aining mo e p oblems han
ele an ou comes,
I will now p esen he execu ion imes o di e en models which ha e now been
un in he Vi ual Machine wi h he ea lie discussed sc ip . I ha e used he exac same
inpu image as be o e.
Model
Execu ion ime
mobilene 2_coco_ oc ainaug
max: 1.37438035 s
min: 1.037979 s
mobilene 2_coco_ oc ain al
max: 1.2677140 s
min: 1.101976156 s
xcep ion_coco_ oc ainaug
max: 10.826671s
min: 8.763001 s
xcep ion_coco_ oc ain al
max: 12.645364 s
min: 9.4737105 s
Table 5-2. Execu ion ime (in seconds) o seman ic segmen a ion using di e en
models in he Vi ual Machine
34
Table 5-2 displays ha he execu ion imes in he Vi ual Machine a e sligh ly
highe han he imes o he models un di ec ly in he lap op, as well as p opo ional.
Howe e , i seems ha he e is no much di e ence be ween he i s wo models.
mobilene 2_coco_ oc ainaug has a highe maximum alue han
mobilene 2_coco_ oc ain al, unlike wha happened in he lap op; bu i s minimum
alue is p opo ional o he esul s in he lap op.
I con inued de eloping he sc ip o implemen seman ic segmen a ion in
webcam images and measu ing di e en execu ion imes. This new e sion o he sc ip
will be discussed la e .
5.2.1 Vi ual Machine con igu a ion
The con igu a ion o he i ual machine VM Vi ualBox wi h all he se ings
necessa y o un he sc ip ha I implemen ed o execu e DeepLab models can be
ound a Apéndice A. I used Ubun u 20.04 as he Ope a ing Sys em, and I ac i a ed a
i ual en i onmen wi h py hon 3.7 and ins alled Tenso Flow 1.15.
5.2.2 Upda ed sc ip
As an imp o emen o he p og am, wi h he help o he lib a y OpenCV [27] he
inpu image is cap u ed h ough a webcam. Mo eo e , ins ead o compu e he
execu ion ime as a whole, di e en imes will be measu ed:
●Image cap u e ime
●Image edimension ime
●In e ence ime
35
5.2.3 Measu ed imes o each model
Models
Cap u e ime
In e ence ime
Redimension ime
mobilene 2_coco_ oc
ainaug
max_1: 0.744371
max2: 0.026295
min: 0.004427
max_1: 1.566632
max2: 1.184288
min: 1.079641
max: 0.042289
max2: 0.028708
min: 0.007805
mobilene 2_coco_ oc
ain al
max_1: 0.712017
max2: 0.018788
min: 0.003908
max_1: 1.785160
max2: 1.590308
min: 1.144132
max: 0.033032
max2: 0.0235562
min: 0.012508
xcep ion_coco_ oc ain
aug
max_1: 0.716405
max2: 0.013771
min: 0.004082
max_1: 11.369783
max2: 10.720958
min: 8.296429
max: 0.019892
max2: 0.015657
min: 0.005462
xcep ion_coco_ oc ain
al
max_1: 0.742603
max2: 0.015859
min: 0.005048
max_1: 16.157057
max2: 9.322600
min: 8.121144
max: 0.025819
max2: 0.023149
min: 0.008753
Table 5-3. Seman ic segmen a ion execu ion imes (in seconds)o webcam images in he Vi ual Machine
These a e he maximums and minimum alues o he cap u e, in e ace and
edimensions imes ex ac ed om 20 i e a ions. Two maximums a e p esen ed because
he i s measu emen has always he highes alue (max_1) o cap u e and in e ence
imes. This i s i e a ion is so high ha i is no ep esen a i e o he eal imes, because o
ha , max2 is needed, which co esponds o a alue independen om he i s
measu emen . In edimension ime his does no happen, bu o he sake o cohe ence
I ha e also included wo max alues.
Table 5-2 shows ha he in e ence ime is much highe o he Xcep ion models,
and, in pa icula , xcep ion_coco_ oc ainaug has he highes in e ence ime ange.
Rega ding Mobilene 2, mobilene 2_coco_ oc ain al appea s o ha e he highes
imes.
36
5.3 DeepLab in Raspbe y Pi 4
As a u he s ep, I used a Raspbe y Pi as a chi ec u e o execu e he p og am.
To do so, I con igu ed he de ice as explained in he nex sec ion. In addi ion, I
cap u ed he execu ion imes as done be o e in he Vi ual Machine.
5.3.1 Raspbe y Pi con igu a ion
Fi s ly, I ins alled Ubun u 20.10 on a mic oSD, in o de o use i as he Ope a ing
Sys em o he Raspbe y Pi. This p ocess was e y use iendly hanks o he ins alla ion
in e ace. I ollowed his [20] guide.
Figu e 5-12. Raspbe y Pi ins alla ion in e ace. [20]
I ins alled he Ope a ing Sys em success ully. Howe e , I had ce ain issues ega ding he
so wa e and lib a ies dependencies. I inally decided o ins all Ubun u 20.04 on he
mic oSD, which was de ini ely a mo e complex p ocess. I ins alled Ubun u Se e 20.04
LTS on Raspbe y Pi 4, as well as Ubun u GNOME 3 desk op en i onmen on i ollowing
his u o ial. [39]
Then I had o ins all all he so wa e needed, o ins ance Tenso low and
OpenCV. In his case I ins alled Tenso low 2.4 and he las numpy e sion. [21]
Finally, I connec ed he came a and execu ed my sc ip , which an success ully.
37
5.3.2 Measu ed imes o each model
Models
Cap u e ime
In e ence ime
Redimension ime
mobilene 2_coco_ oc
ainaug
max_1: 0.265220
max2: 0.003352
min: 0.001896
max_1: 3.361115
max2: 2.792313
min: 2.239232
max: 0.022485
max2: 0.022106
min: 0.019185
mobilene 2_coco_ oc
ain al
max_1: 0.265913
max2: 0.003488
min: 0.0019512
max_1: 3.518164
max2: 2.715697
min: 2.364746
max: 0.023496
max2: 0.023123
min: 0.018469
xcep ion_coco_ oc ain
aug
max_1: 0.261291
max2: 0.009885
max3: 0.003544
min: 0.001897
max_1: 33.316450
max2: 25.020431
min: 23.625263
max: 0.020824
max2: 0.020595
min: 0.018410
xcep ion_coco_ oc ain
al
max_1: 0.263705
max2: 0.007603
max3: 0.002768
min: 0.002304
max_1: 32.254001
max2: 23.877033
min: 21.974183
max: 0.024083
max2: 0.020609
min: 0.018630
Table 5-4. Seman ic segmen a ion execu ion imes (in seconds) o webcam images in he Raspbe y Pi 4
These esul s ha e also been ex ac ed om 20 i e a ions. In his case, i happens
he same as in he imes measu ed in he Vi ual Machine, which ha e been p e iously
discussed. The i s measu emen has he highes alues in Cap u e and In e ence ime,
so we ha e max_1 o ha alue, and max2 o he nex highes alue. In some cases I
ha e also included max3 because max2 was ex emely high.
I can conclude ha he seman ic segmen a ion imes execu ed in he Raspbe y
Pi a e app oxima ely wice as high as he esul s ob ained in he Vi ual Machine.
Mo eo e , xcep ion_coco_ oc ainaug s ill has he highes alues.
I will now show he segmen a ion esul s I cap u ed h ough he webcam in
di e en models and i s cha ac e is ics.
38
Fi s ly, I used he model xcep ion_coco_ oc ain al; he one wi h he highes
MIoU (87%), as seen in Table 4-2. I concluded ha his model needs o ha e clea
images as inpu , wi h good ligh ning and clea igu es, o make a p edic ion. I he
image is no o ally clea , he ou pu o he p edic ion is all “backg ound”. Mo eo e ,
when seman ically segmen ing an objec , he bo de s a e e y accu a ely speci ied in
he image, unlike in he models xcep ion_coco_ oc ainaug and
mobilene 2_coco_ oc ainaug, which I measu ed la e .I mus be no ed ha when I
pe o med he segmen a ion in bo h Xcep ion models, he Raspbe y Pi hea ed o a
e y high empe a u e and I had o e ige a e i . This happened due o he high
in e ence ime. I also concluded ha xcep ion_coco_ oc ainaug is no wo h he ime
i akes o he in e ence in ela ion o i s accu acy (82-83% MIoU), which is almos he
same as mobilene 2_coco_ oc ain al (80%) and has a e y low in e ence ime.
When using mobilene 2_coco_ oc ainaug, he model wi h he lowes MIoU, I
no ed ha in spi e o no ha ing a e y clea inpu image, his model makes a
p edic ion. Some imes he in e ence is co ec , and some imes i is no . Mo eo e , his
model does no o e e y accu a e segmen a ion o he objec s. The pe ime e ha i
co e s is gene ally la ge han he ac ual objec . Howe e , he in e ence ime is much
lowe in Mobilene han in Xcep ion. On he o he hand, he accu acy is no much o a
p oblem in mobilene 2_coco_ oc ain al (80% o MIoU e sus 75 - 77% o
mobilene 2_coco_ oc ainaug), and i doesn no need o ally clea images o make a
igh p edic ion.
I will now compa e he in e ence beha iou o he models in simila si ua ions.
Figu e 5-13. RasPi Unclea inpu , xcep ion_coco_ oc ain al 7
39
Figu e 5-14. RasPi Unclea inpu , xcep ion_coco_ oc ainaug 26
Figu e 5-15. RasPi Unclea inpu , mobilene 2_coco_ oc ain al 29
Figu e 5-16. RasPi Unclea inpu , al e na i e mobilene 2_coco_ oc ain al 31
Figu e 5-17. RasPi Unclea inpu , mobilene 2_coco_ oc ainaug 39
When he image is no o ally clea (is oo ligh o oo da k), igu es 5-13 o 5-17
show how he Xcep ion models ha e mo e ouble segmen ing he images when hey
do no know o su e which ag he objec has. On he con a y, Mobilene models
40
pe o m in e ence e en when he image image is no clea . This can ha e igh (Figu e
15) o w ong (Figu e 16) esul s.
Figu e 5-18. RasPi Clea inpu , xcep ion_coco_ oc ain al 9
Figu e 5-19. RasPi Clea inpu , xcep ion_coco_ oc ainaug
Figu e 5-20. RasPi Clea inpu , mobilene 2_coco_ oc ain al 23
Figu e 5-21. RasPi Clea inpu , mobilene 2_coco_ oc ainaug 30
41
Figu e 5-22. RasPi Clea inpu , al e na i e mobilene 2_coco_ oc ainaug 63
When he inpu image has good ligh ing, Xcep ion models gene ally beha e
well, hough xcep ion_coco_ oc ainaug (Figu e 19) may no be as p ecise as
xcep ion_coco_ oc ain al (Figu e 18).mobilene 2_coco_ oc ain al has co ec
beha iou oo (Figu e 5-20). Howe e , mobilene 2_coco_ oc ainaug may iden i y he
co ec objec , bu highligh o he a eas in he image in which he objec does no
appea (Figu e 5-21), o i can e en indica e a di e en ag ha does no co espond
o he objec in he inpu image (Figu e 5-22).
Figu e 5-23. RasPi Se e al objec s, xcep ion_coco_ oc ain al 47
Figu e 5-24. RasPi Se e al objec s, xcep ion_coco_ oc ainaug
42
Figu e 5-25. RasPi Se e al objec s, mobilene 2_coco_ oc ain al 51
Figu e 5-26. RasPi Se e al objec s, mobilene 2_coco_ oc ainaug 9
In his si ua ion he e a e a ious objec s, and he ligh ning is no ideal. I is a e
ha in xcep ion_coco_ oc ain al he second moni o is no iden i ied. Howe e , he
in e ence o he o he objec s is e y accu a e. The o he hing o be no ed is ha , in
his case, mobilene 2_coco_ oc ain al looks less accu a e han
mobilene 2_coco_ oc ainaug.
To conclude, using he model mobilene 2_coco_ oc ain al is a good solu ion
because o i s accu acy and i s in e ence ime. This model is no as accu a e as
xcep ion_coco_ oc ain al, bu i is much as e (as i can be no ed in he ables),
which makes mobilene 2_coco_ oc ain al a good op ion i he p ecision is no
ex emely c ucial.
43
Then, I c ea ed a Google Colab p og am o pe o m he quan iza ion, which
beha ed p ope ly. I needed o use a ep esen a i e da ase o pe o m he ull in ege
pos - aining quan iza ion and be able o un he esul ing model in an Edge TPU
a chi ec u e ( he Co al Accele a o ).
Fo his pu pose, I i s gene a ed a da ase o dummy images, which a e
andomly gene a ed, o es ha i wo ked p ope ly. I la e c ea ed my own da ase o
30 images aken om my own webcam o use hem as he ep esen a i e da ase .
Once his s ep was comple ed, he mobilene 2_dm05_coco_ oc_ ain al model was
ans o med in o a TF Li e model.
I compa ed he by es sa ed wi h he quan iza ion using he dummy images:
Tenso Flow Model is 3042785 by es
TFLi e Model is 963896 by es
Pos aining in 8 quan iza ion sa es 2078889 by es
And my webcam images da ase :
Tenso Flow Model is 3042785 by es
TFLi e Model is 2789284 by es
Pos aining in 8 quan iza ion sa es 253501 by es
The highe he quali y o he ep esen a i e da ase , he smalle numbe o by es
sa ed in he quan iza ion.
The nex s ep is o compile he TF Li e model o Edge TPU ( o ge a _edge pu. li e
model) and o download i o pe o m in e ence on he Raspbe y Pi using he Co al
Accele a o . I mus be no ed ha he only model ha I ound whose ope a ions can
un o EdgeTPU is mobilene 2_dm05_coco_ oc_ ain al.
50
Figu e 5-32. In e ence esul o dummy_edge pu. li e model wi h inpu image
When I used he dummy_edge pu. li e model ( he model quan ized using
dummy images as a ep esen a i e da ase ) using a downloaded image as an inpu , I
go he esul s displayed in Figu e 5-32.
Figu e 5-33. In e ence esul o sma y_edge pu. li e model wi h inpu image
And when I used he sma y_edge pu. li e model ( he model quan ized using
webcam images as a ep esen a i e da ase ) using a downloaded image as an inpu , I
go he esul o Figu e 5-33. I can be no ed ha nei he a e e y accu a e.
To pe o m in e ence wi h hese models I c ea ed a sc ip , which can be
execu ed as ollows:
py hon3 sem_seg_edge pu_pics.py --model dummy_edge pu. li e
--inpu bi d.bmp --keep_aspec _ a io --ou pu
segmen a ion_ esul _dummy.jpg
51
Figu e 5-34. In e ence esul o sma y_edge pu. li e model wi h webcam inpu
Howe e , when using he dummy model wi h a webcam image inpu , he
in e ence esul s we e almos always labeled as “backg ound” e en hough he e we e
se e al objec s p esen in hose images. Ne e heless, sma y models p edic ed some
objec labels e en i hey we e no always co ec o accu a e, as we can see in his
example in Figu e 5-34.
Rega ding in e ence ime, he dummy model akes a li le longe han he sma y
model. A compa ison be ween he EdgeTPU models and he non-quan ized models will
be p esen ed la e .
-
52
Capí ulo 6 - Resul s Compa ison and Analysis
The in e ence, cap u e and edimension imes o he di e en models in a ious
a chi ec u es ha e al eady been al eady shown. In his sec ion he esul s will be
compa ed and analyzed o comp ehend be e hei meaning.
6.1 Lap op s. Vi ual Machine execu ion imes
In he o e all execu ion ime esul s o he seman ic segmen a ion execu ed in
he lap op (Table 5-1) a e sligh ly lowe han he execu ion imes in he Vi ual Machine
(Table 5-2). The e is no a big di e ence be ween he models’ beha iou depending on
he a chi ec u e used; he imes in he lap op a e p opo ional o he imes in he i ual
machine. These esul s we e compu ed wi h he same p og am using he same inpu
images.
6.2 Vi ual Machine image inpu s. Vi ual Machine webcam inpu
La e , I implemen ed some modi ica ions in he p og am o measu e he
in e ence, cap u e and edimension imes on hei own, unlike he ables be o e whe e
all hese imes we e conside ed he execu ion ime. This ime he p og am did no
pe o m he seman ic segmen a ion on pic u es gi en as inpu , bu on images cap u ed
h ough he webcam, hence he cap u e ime was also measu ed. (Table 5-3)
In his able i can be no ed ha he cap u e ime has no signi ican a iance
be ween he models. Mo eo e , he i s cap u e ime o each model is signi ican ly
highe han he es . This phenomenon happens du ing he in e ence as well, bu no so
much du ing he edimension ime.
When compa ing he in e ence o he webcam images wi h he execu ion ime
o he inpu images, i can be seen ha he imes a e lowe in he lap op han in he
i ual machine o he Mobilene V2 models. All his aking in o accoun ha he
execu ion ime o he lap op includes he in e ence and edimension ime. Howe e , o
he Xcep ion models his is mo e ambiguous, bu we can see ha he minimum alues
a e somewha smalle in he i ual machine han in he lap op.
53
Finally, he edimension ime has he lowes alues o he Xcep ion models,
speci ically he xcep ion_coco_ oc ainaug model; while he Mobilene V2 models
ha e highe edimension imes.
6.3 Vi ual Machine s. Raspbe y Pi
When i comes o he Raspbe y Pi, he cap u e imes a e signi ican ly smalle
han in he Vi ual Machine, and he e is no much ime di e ence be ween he
models. (Table 5-4)
Models
Vi ual Machine
in e ence ime
Raspbe y Pi
in e ence ime
mobilene 2_coco_ oc ainaug
max_1: 1.566632
max2: 1.184288
min: 1.079641
max_1: 3.361115
max2: 2.792313
min: 2.239232
mobilene 2_coco_ oc ain al
max_1: 1.785160
max2: 1.590308
min: 1.144132
max_1: 3.518164
max2: 2.715697
min: 2.364746
xcep ion_coco_ oc ainaug
max_1: 11.369783
max2: 10.720958
min: 8.296429
max_1: 33.316450
max2: 25.020431
min: 23.625263
xcep ion_coco_ oc ain al
max_1: 16.157057
max2: 9.322600
min: 8.121144
max_1: 32.254001
max2: 23.877033
min: 21.974183
Table 6-1. In e ence ime compa ison (in seconds) be ween Vi ual Machine and Raspbe y Pi
Howe e , he in e ence ime in he RasPi is app oxima ely wo imes highe han
in he Vi ual Machine. Table 6-1 shows he in e ence ime compa ison be ween bo h
a chi ec u es.
Rega ding he edimension ime, he Xcep ion models ha e smalle alues in he
Vi ual Machine han in he Raspbe y Pi. In pa icula , he minimum alues a e
ex emely small, while he maximum alues a e e y simila o he ones in he RasPi.
54
6.4 Raspbe y Pi: egula model s. Edge TPU model
The quan iza ion and EdgeTPU compila ion we e pe o med on a
mobilene 2_dm05_coco_ oc_ ain al model. This will be a compa ison o he cap u e,
be ween he non-quan ized ( egula ) model and he _edge pu. li e model.
Model
Cap u e ime
In e ence ime
Redimension ime
egula
max1: 0.303826
max2: 0.0177173
min: 0.001894
max1: 5.633136
max2: 2.095738
min: 1.611738
max: 0.058004
min: 0.020867
dummy_edge pu. li e
max: 0.269128
min: 0.266331
max: 0.008964
min: 0.008219
max: 0.033319
min: 0.029464
sma y_edge pu. li e
max: 0.273280
min: 0.268579
max: 0.007211
min: 0.007089
max: 0.031751
min: 0.030280
Table 6-2. Time compa ison (in seconds) be ween egula and EdgeTPU seman ic segmen a ion model in
Raspbe y Pi
The able shows, as we ha e seen in p e ious egula models, ha
mobilene 2_dm05_coco_ oc_ ain al egula model always has a highe cap u e and
in e ence ime on he i s i e a ion (max1), bu his pa e n does no happen o he
edimension ime.
In he _edge pu. li e models I implemen ed a sc ip which pe o ms jus one
ope a ion each ime i is un, unlike he sc ip o he egula model, which has a loop
ha cap u es webcam images un il i is speci ied ha i should s op. Because o his, he
cap u e imes o he _edge pu. li e models a e highe (all he imes he cap u e ime is
measu ed, i is conside ed he i s (and only) ime du ing ha execu ion, and he e o e
he highes ). O e all, i seems like he cap u e ime o he egula model may be a li le
highe .
Wi h ega ds o he in e ence, we see ha in he _edge pu. li e models he ime
is ex emely lowe hanks o he Co al Accele a o . E en in he case ha i happened
he same as wi h he cap u e ime, ha means ha he in e ence ime could be e en
lowe in he nex i e a ions i he e was a loop in he sc ip .
55
On he o he hand, i is in e es ing o see ha he imes o he _edge pu. li e
models ha e e y simila maximum and minimum alues. The edimension ime seems
e y simila o all models, bu as said p e iously, he di e ence be ween maximum and
minimum alues is e y small o _edge pu. li e models.
Mo eo e , he dummy model has sligh ly highe in e ence ime han he sma y
model.
56
Capí ulo 7 - Conclusions and u u e wo k
Seman ic segmen a ion has become essen ial nowadays. E en e e yday
people use his echnology on hei phones, and i should be possible o e e yone o
be able o ake ad an age o i . Thanks o he ad ances in Edge Compu ing, his is now
possible.
This p ojec shows ha he esul s ob ained a e posi i e. By simpli ying
(quan iza ion) a seman ic segmen a ion model and using an USB Co al Accele a o I
was able o pe o m in e ence in eal- ime on images cap u ed h ough he webcam
on a de ice es ic ed by compu ing powe and ene gy consump ion (Raspbe y Pi).
Human-eye ision canno di e en ia e a succession o images in 24 FPS, which is
equi alen o 0.0416 seconds. This means ha we pe cei e i as a ideo, no as
independen images. I ob ained an a e age in e ence ime o 0.0085 and 0.0071
seconds in di e en EdgeTPU models, hence i he sc ip which uns he in e ence had a
loop, he seman ic segmen a ion would ope a e like a ideo, in eal ime. Howe e , o
ge hese esul s, we ha e o sac i ice accu acy.
None heless, using a egula (non-quan ized) seman ic segmen a ion model on
he Raspbe y Pi, o e en a Linux i ual machine, and using DeepLab as a basis, an
accu a e implemen a ion is possible al hough in e ence imes a e no e icien .
The nex s eps would be o deploy DeepLab o o he pla o ms such as In el NCS
and Imagina ion Technologies Powe VR.
57
58
BIBLIOGRAFÍA
1. (n.d.). picsum.
h ps://i.picsum.pho os/id/1012/3973/2639.jpg?hmac=s2eybz51lnKy2ZHkE2wsgc6S81
VD1W2NKYOSh8bzDc
2. ADE20K. (n.d.). MIT CSAIL. h ps://g oups.csail.mi .edu/ ision/da ase s/ADE20K/
3. B io , A., Viswana h, P., & Yogamani, S. (2018). Analysis o E icien CNN Design
Techniques o Seman ic Segmen a ion. IEEE/CVF Con e ence on Compu e Vision
and Pa e n Recogni ion Wo kshops (CVPRW). 10.1109/CVPRW.2018.00109
4. Chen, L.-C., Papand eou, G., Kokkinos, I., Mu phy, K., & Yuille, A. L. (2016). DeepLab:
Seman ic Image Segmen a ion wi h Deep Con olu ional Ne s, A ous Con olu ion,
and Fully Connec ed CRFs. IEEE T ansac ions on Pa e n Analysis and Machine
In elligence,40, pp. 834-848. 10.1109/TPAMI.2017.2699184
5. Chen, L.-C., Papand eou, G., Sch o , F., & Adam, H. (2017). Re hinking A ous
Con olu ion o Seman ic Image Segmen a ion. (a Xi :1706.05587).
6. Chen, L.-C., Zhu, Y., Papand eou, G., Sch o , F., & Adam, H. (2018).
Encode -Decode wi h A ous Sepa able Con olu ion o Seman ic Image
Segmen a ion. Eu opean Con e ence on Compu e Vision (ECCV), pp 833-851.
7. The Ci yscapes Da ase . (n.d.). h ps://www.ci yscapes-da ase .com
8. Cloud Tenso P ocessing Uni s (TPUs). (n.d.). Google Cloud.
h ps://cloud.google.com/ pu/docs/ pus
9. Con igu aciones de compilación p obadas. (n.d.). Tenso Flow.
h ps://www. enso low.o g/ins all/sou ce?hl=es-419# es ed_build_con igu a ions
10. CYBORG NITR. (2020, May 9). MIoU Calcula ion. Medium.
h ps://medium.com/@cybo g. eam.ni /miou-calcula ion-4875 918 4cb
59
Tools ins alla ion
The nex s ep is o download he co esponding Tenso low e sion acco ding o
he p e iously ins alled Py hon e sion. This can be done by using pip:
pip ins all enso low==1.15
In o de o use images cap u ed om he webcam as inpu o he p og am,
OpenCV needs o be ins alled:
sudo ap upda e
py hon -m pip ins all openc -py hon
To e i y ha he co ec OpenCV e sion was ins alled, he ollowing command
e u ns he cu en e sion ins alled in ou machine. Ou pu example: 4.2.0
py hon3 -c "impo c 2; p in (c 2.__ e sion__)"
Webcam con igu a ion
Now ha all he necessa y ools a e ins alled, i is ime o connec he came a o
he i ual machine.
When we un ou i ual machine, he uppe pa o he menu shows he window
‘De ices’ o ‘Disposi i os’, which should include he op ion ‘Webcams’ o ‘Cáma as
web’ once we click on i . A his poin , we jus need o selec ou webcam, which
should appea lis ed in ha op ion. I , on he con a y, he op ion ‘Webcams’ does no
appea in he menu, ha means ha we need o ins all he adequa e Vi ualBox
Ex ension pack co esponding o he e sion o ou Vi ualBox.
The inal s ep
Once he webcam is selec ed, he sc ip should un co ec ly and ob ain he
inpu images h ough he webcam and seman ically segmen hem.
66
67