i
Da a Labeling ools o Compu e Vision:
Ped o Miguel Lima de Sousa Reis
a Re iew
Disse a ion p esen ed as pa ial equi emen o ob aining
he Mas e 's deg ee in Da a Science and Ad anced Analy ics
i
NOVA In o ma ion Managemen School
Ins i u o Supe io de Es a ís ica e Ges ão de In o mação
Uni e sidade No a de Lisboa
DATA LABELING TOOLS FOR COMPUTER VISION: A REVIEW
by
Ped o Miguel Lima de Sousa Reis
Disse a ion p esen ed as pa ial equi emen o ob aining he Mas e 's deg ee in Da a Science
and Ad anced Analy ics
Ad iso : Robe o And é Pe ei a Hen iques, PhD
No embe 2021
ii
To my belo ed Ca olina, who is he pu es and kindes human being I know, wi h whom my li e is
happie e e y day. To Paula and Miguel o hei endless lo e and suppo . To Rui and Ri a ha
always wo y and wonde abou wha and how I am cu en ly doing. To Manuel and Conceição Ma ia
o always being he e uncondi ionally, pushing me u he and belie ing in me since I was a child,
whose wo ds o encou agemen and enaci y s ill echo in my ea s. To Rui, who has always le me
s and on his shoulde s. To João, Alexand ino, Miguel, Pinhal, B i es and Sa abando – each one wi h
i s supe powe . This disse a ion is dedica ed o all my amily and iends. Finally, I’d also like o
acknowledge NTT DATA, o me ly known as e e is, and he Da a & Analy ics eam, o he suppo
and inspi a ion h oughou all my mas e ’s deg ee as a ull- ime wo ke -s uden .
iii
ABSTRACT
La ge olumes o labeled da a a e equi ed o ain Machine Lea ning models in o de o sol e
oday’s compu e ision challenges. The ecen exace ba ed hype and in es men in Da a Labeling
ools and se ices has led o many ad-hoc labeling ools. In his e iew, a de ailed compa ison be ween
a selec ion o da a labeling ools is amed o ensu e he bes so wa e choice o holis ically op imize
he da a labeling p ocess in a Compu e Vision p oblem. This analysis is buil on mul iple domains o
ea u es and unc ionali ies ela ed o Compu e Vision, Na u al Language P ocessing, Au oma ion,
and Quali y Assu ance, enabling i s applica ion o he mos p e alen da a labeling use cases ac oss
he scien i ic communi y and global ma ke .
KEYWORDS
Re iew; Compu e Vision; Image Anno a ion; Da a Labeling so wa e; Supe ised Machine
Lea ning; Me hodologies and Tools
i
INDEX
1. In oduc ion ........................................................................................................ 1
1.1. Da a Labeling ......................................................................................................... 2
2. His o ical O e iew ............................................................................................. 4
3. Li e a u e Re iew ............................................................................................... 9
4. Da a Labeling Tools ........................................................................................... 13
4.1. Compu e Vision ea u es.................................................................................... 16
4.2. Na u al Language P ocessing ea u es ................................................................ 18
4.3. Au oma ion and De elope - iendly ea u es ..................................................... 20
4.4. Managemen and Quali y Assu ance ea u es .................................................... 22
4.5. Gene al Compa a i e Analysis............................................................................. 24
5. Discussion ......................................................................................................... 25
6. Conclusions ....................................................................................................... 27
7. Limi a ions and ecommenda ions o u u e wo ks ....................................... 28
8. Bibliog aphy ...................................................................................................... 29
LIST OF FIGURES
Figu e 1 – E olu ion o he image gene a ion capabili ies by Gene a i e Ad e sa ial Ne wo ks
(GANs) om 2014 o 2018 (Saxena & Cao, 2020) .............................................................. 4
Figu e 2 – E olu ion o he mos ac i e esea ch opics in he Compu e Vision ield o e ime
(Szeliski, 2010) .................................................................................................................... 7
Figu e 3 – F equency ha he e m “Da a Science” was sea ched o on Google, in he wo ld
and in mul iple languages, om 2004 un il he p esen (Google, 2021) ........................... 8
Figu e 4 – Da a Labeling anno a ion ool ImageTagge (Fiedle e al., 2019) aming an
Objec De ec ion ask ......................................................................................................... 9
Figu e 5 – E olu ion o Image Classi ica ion models on ImageNe da ase : Top 1 Accu acy
(le ) and Top 5 Accu acy ( igh ) (Facebook AI Resea ch, 2021) ...................................... 10
Figu e 6 – Hype Cycle o A i icial In elligence 2021 (Ga ne Inc., 2021), as o July 2021..... 11
Figu e 7 – F equency ha he e m “Supe ised Machine Lea ning” was sea ched o on
Google, in he wo ld and in mul iple languages, om 2016 un il he p esen ................ 13
i
LIST OF TABLES
Table 1 – Manually selec ed da a anno a ion ools and espec i e de elope eams and
e e ences, i applicable ................................................................................................... 14
Table 2 – Compu e Vision ela ed ea u es in he selec ed da a anno a ion ools ................ 17
Table 3 – NLP- ela ed ea u es in he selec ed da a anno a ion ools .................................... 18
Table 4 – Au oma ion and de elope - iendly ea u es in he selec ed da a anno a ion ools20
Table 5 – Managemen and QA ea u es in he selec ed da a anno a ion ools ..................... 22
Table 6 – Compa a i e summa y o he ea u es g ouped by ca ego y ac oss he analyzed
da a anno a ion ools ....................................................................................................... 24
ii
LIST OF ABBREVIATIONS AND ACRONYMS
AGI A i icial Gene al In elligence
AI A i icial In elligence
API Applica ion P og amming In e ace
CAGR Compound Annual G ow h Ra e
CDO Chie Da a O ice
CI/CD Con inuous In eg a ion/Con inuous De elopmen
CNN Con olu ional Neu al Ne wo k
COVID-19 Co ona i us Disease 2019
CV Compu e Vision
DARPA De ense Ad anced Resea ch P ojec s Agency
FTE Full- ime Equi alen employee
GAN Gene a i e Ad e sa ial Ne wo k
GPU G aphics P ocessing Uni
HDR High Dynamic Range
HITL Human-In-The-Loop
ILSVRC ImageNe La ge Scale Visual Recogni ion Challenge
IT In o ma ion Technology
MIT Massachuse s Ins i u e o Technology
ML Machine Lea ning
NER Named En i y Recogni ion
NLG Na u al Language Gene a ion
NLP Na u al Language P ocessing
QA Quali y Assu ance
RPA Robo ic P ocess Au oma ion
SaaS So wa e-as-a-Se ice
SDK So wa e De elopmen Ki
SVR Single-View Recons uc ion
UI Use In e ace
UX Use Expe ience
1
1. INTRODUCTION
The concep o Compu e Vision (CV) has unde gone many changes since he beginning o he 21s
cen u y. In 2003, Compu e Vision was jus en isioned as an exci ing bu diso ganized ield (Fo sy h &
Ponce, 2003), bu en yea s la e his de ini ion apidly u ned in o a ield ha ocused on he
au oma ed ex ac ion o in o ma ion om images o in e some hing abou he wo ld (P ince, 2012;
Solem, 2012). Cu en ly, scien i ic and business communi ies a e al eady used o hea on he concep
o CV in a newspape headline highligh ing an in es men o millions o dolla s on he esolu ion o
cu en global challenges, such as igh ing Co ona i us Disease 2019 (COVID-19) (Ulhaq e al., 2020) o
p e en ing clima e changes (Ramachand a, 2019).
As a ield ha seeks o ex ac in o ma ion om image da a au oma ically, cu en in-use solu ions
o sys ems wo k wi h Compu e Vision h ough wo di e en app oaches. Handc a ed app oaches,
o example, apply CV echniques by using se s o ules o sol e a speci ic challenge. Some use-cases
o i a e 1) using a su eillance came a o de ec mo emen in a oom; 2) de ec ing g een o de o es ed
a eas in sa elli e images, and; 3) c opping some pa o a documen om a pic u e. On he o he hand,
Compu e Vision solu ions may also be associa ed wi h machine lea ning models o o he A i icial
In elligence (AI) applica ions, such as objec ecogni ion and de ec ion, image classi ica ion o anomaly
de ec ion solu ions. These Compu e Vision sys ems a e o en inspi ed by he p ope ies and
cha ac e is ics o human ision. Con e sely, hese algo i hms can also o e insigh s in o how he
in o ma ion ex ac ed om images is in e p e ed in he human b ain.
The abili y o a i icially in elligen sys ems o see like humans has been a subjec o inc easing
in e es and does no appea o be slowing down any ime soon. Howe e , he p ocess o deciphe ing
images, due o he mo e signi ican amoun o da a ha needs analysis, is mo e complex han
unde s anding o he o ms o bina y in o ma ion. None heless, he usage o a i icial neu al ne wo ks
is making compu e ision mo e capable o iden i ying pa e ns om images han he human isual
cogni i e sys em (Schei e e al., 2014).
Also, compu e ision echnologies will no only be less demanding o ain, bu also be capable o
pe cei ing mo e om images han hey a e doing wi hin he p esen . Toge he wi h o he echnologies
o o he subse s o AI, hese can be used in o de o build e en mo e powe ul and obus applica ions.
Fo ins ance, image cap ioning echniques can be combined wi h Na u al Language Gene a ion (NLG)
applica ions a e used o deciphe he su ounding objec s o isually impai ed indi iduals (Kim, 2020).
In a nea u u e, compu e ision will play a i al ole wi hin he de elopmen o A i icial Gene al
In elligence (AGI) and A i icial Supe in elligence by g an ing he capaci y o handle da a as simila ly o
e en be e han he human isual sys em (Pueyo, 2018). Taking his in o conside a ion, i can be ha d
o he scien i ic communi y o accep ha oday’s compu e ision capabili ies oge he wi h i s
8
(CNN), and i s b eak h ough consis ed in i s abili y o use a GPU o ain he compu e ision model
signi ican ly as e and o longe – AlexNe was ained o e 6 days on wo GPUs ha we e accessible
o he consume . Since hen, Da a Science and all i s subdomains ha e e ol ed ema kably, bo h
ma hema ically, in e ms o he in as uc u e i uses, and in e ms o compu ing and la ge-scale
p ocessing. Finally, he global ma ke has unde s ood he ad an ages o applying he da a- ela ed
echnologies and me hodological app oaches in en ed o da e, and has been able o ma e ialize hem
ei he in p oduc s o se ices, o in he in e nal p ocesses o o ganiza ions, such as in chu n p edic ion
and sales o ecas ing (Google, 2021). E en C-le el execu i e posi ions such as Chie Da a O ice s
(CDOs) ha e been c ea ed o manage he c ea ion and go e nance o da a p ocesses and o ensu e a
da a-d i en cul u e in o ganiza ions.
The Compu e Vision ield has inhe en ly bene i ed om his e olu ion, and is aced, pe haps o
he i s ime, wi h he challenge o no only con inuing o imp o e i s pe o mance, bu also o
op imizing he cos s ela ed o he necessa y e o . This is whe e he da a labeling componen comes
in, which may well be he d i ing o ce behind all u u e de elopmen s.
Figu e 3 – F equency ha he e m “Da a Science” was sea ched o on Google, in he wo ld and in
mul iple languages, om 2004 un il he p esen (Google, 2021)
9
3. LITERATURE REVIEW
(Du a & Zisse man, 2019) de ined da a labeling in he con ex o machine lea ning as he p ocess
o de ec ing and agging da a samples while a aching meaning and/o con ex o digi al da a. This
p ocess can be manual bu is usually pe o med o assis ed by so wa e. Da a labeling is an impo an
pa o da a p ep ocessing o Machine Lea ning (ML), pa icula ly o Supe ised Lea ning. Bo h inpu
and ou pu da a a e labeled o classi ica ion o p o ide a lea ning basis o u u e da a p ocessing. Fo
example, a sys em aining o iden i y animals in images migh be p o ided wi h mul iple images o
a ious ypes o animals om which i would lea n he s anda d ea u es o each, enabling i o
co ec ly iden i y he animals in unlabeled images (Why ock e al., 2021).
Machine Lea ning and Deep Lea ning sys ems o en equi e massi e amoun s o da a o es ablish
a ounda ion o eliable lea ning pa e ns. Da a acili a ed o he lea ning p ocess mus be labeled o
anno a ed, which means ha e e y hing, o some imes only he mos impo an hings, mus be
iden i ied and localized in he image. I mus also be labeled based on da a ea u es ha help he
model o ganize he da a in o pa e ns ha p oduce a desi ed answe . A p ope ly labeled da ase
p o ides a g ound u h ha he ML model uses o check i s p edic ions o accu acy and o con inue
e ining i s algo i hm. E o s in his p ocedu e impai he quali y o he aining da ase and he
pe o mance o any p edic i e models i is used o (Kshe i, 2021). To mi iga e his, many
o ganiza ions ake a Human-In-The-Loop (HITL) app oach, which is called a “da a labele ” main aining
human in ol emen in aining and es ing da a models h oughou hei i e a i e g ow h (Mona ch,
2021). The e a e se e al p ocedu es o s uc u e and label da a while main aining human in ol emen
(Fiedle e al., 2019). Ei he by using c owdsou cing, whe e a hi d-pa y pla o m gi es an en e p ise
Figu e 4 – Da a Labeling anno a ion ool ImageTagge (Fiedle e al., 2019)
aming an Objec De ec ion ask
10
access o many wo ke s a once, and/o by using con ac o s, whe e an en e p ise can hi e empo a y
eelance wo ke s o p ocess and label da a. A ecen epo om AI esea ch and ad iso y i m
Cognily ica ound ha o e 80% o he ime en e p ises spend on AI p ojec s goes owa d p epa ing,
cleaning, and labeling da a (C. Resea ch, 2019). Manual da a labeling is he mos ime-consuming and
expensi e me hod, bu i migh be wa an ed o impo an applica ions (Fiedle e al., 2019). Some
expe s do belie e ha da a labeling may p esen a new low-skilled job oppo uni y o eplace he
ones ha a e nulli ied by au oma ion, because he e is an e e -g owing su plus o da a and machines
ha need o p ocess i o pe o m he asks necessa y o ad anced ML and AI, which will c ea e mo e
and mo e low-skilled jobs and needs o hi e mo e ope a ional p o iles (Kshe i, 2021).
Apa om ha , he e olu ion o image classi ica ion models shows a clea upwa d end in he
eme gence o new image classi ica ion models, s emming om an in es men in esea ch (Facebook
AI Resea ch, 2021). This is also ue o seman ic segmen a ion, language modelling, ime se ies
o ecas ing, speech ecogni ion, among o he me hods (Wason, 2018).
E iden ly, he e olu ion in da a anno a ion echniques and so wa e ollows a simila end in
ecen yea s, since ha dwa e limi a ions a e a hind ance in inc easing he pe o mance o he me hods
lis ed abo e, and esea ch and de elopmen e o is di ec ed owa ds his a ea. Acco ding o (G. V.
Resea ch, 2021), he global da a anno a ion ma ke was alued a US$ 695.5 million in 2019, is
cu en ly alued a US$ 1.66 billion and is expec ed o each US$ 8.22 billion by 2028. The g owing da a
anno a ion indus y, which is expec ed o g ow a a Compound Annual G ow h Ra e (CAGR) o 25.6%
om 2021 o 2028 (G. V. Resea ch, 2021), is expec ed o expe ience eno mous expansion in he nea
u u e.
Plus, Ga ne classi ied Da a Labeling and Anno a ion Se ices as en e ing he T ough o
Disillusionmen (Figu e 6) as his end jus walked by he Peak o In la ed Expec a ions (Ga ne Inc.,
2021). This means ha his ield has jus su ed i s wa e o exace ba ed hype and in es men and i is
inally slowing down while implemen a ions ail o deli e . I ypically happens be o e mo e ins ances
Figu e 5 – E olu ion o Image Classi ica ion models on ImageNe da ase : Top 1 Accu acy (le )
and Top 5 Accu acy ( igh ) (Facebook AI Resea ch, 2021)
11
o how hese se ices can bene i he en e p ise s a o ake shape and become mo e b oadly
ecognized, and is c ucial o s a ga he ing mo e ma u i y in he Da a and AI ma ke . Ga ne es ima es
i o each he Pla eau o P oduc i i y in 5 o 10 yea s (Ga ne Inc., 2021). Thus, da a labeling is a
p ocess ha will cons an ly e ol e and change o mee he business and echnical objec i es, ha is,
labeling asks oday a e e y p one o be di e en in h ee mon hs. Th ough he ime, a da a labeling
eam e ol es and inds be e ways o label aining da a o imp o ed quali y and model pe o mance,
c ea ing guidelines and sha ing in o ma ion on how o deal wi h he a es use cases o scena ios.
Rela i ely o he so wa e, he bes da a labeling ools mus be use - iendly in e ms o Use
In e ace and Use Expe ience (UI/UX) and b eak he wo k down in o a omic and smalle asks o
maximize labeling quali y (Du a & Zisse man, 2019). When a complex ask is ans o med in o a se o
a omic componen s, i is easy o measu e and quan i y each o hose asks. I also allows he
iden i ica ion o which asks a e bes sui ed o humans and which ones can be au oma ed. To op imize
bo h da a quali y and he wo k o ce in es men , he e a e plen y aspec s o conside when choosing
he ideal da a labeling ool. In he ollowing chap e , an in-dep h analysis and compa ison o a hand-
picked selec ion o ools is p esen ed, ocusing on he mul iple unc ional and echnical issues ela ed
o he opics o Compu e Vision, Na u al Language P ocessing (NLP), Au oma ion, Quali y Assu ance
(QA) and Managemen (Du a & Zisse man, 2019; Said e al., 2017). These opics a e deal wi h as
clus e s o ea u es o unc ionali ies ha a e appa en ly p e alen wi hin he da a anno a ion ools
and se ices ma ke and migh de ia e sligh ly om ou ocus on he Compu e Vision ield.
Figu e 6 – Hype Cycle o A i icial In elligence 2021 (Ga ne Inc., 2021), as o July
2021
12
Las ly, (Gau e al., 2018) conduc ed a simila wo k o e iew he s a e-o - he-a in ideo
anno a ion. Howe e , gi en he la es ma ke upda es in Da a Labeling se ices, his e iew lacks a
iew o he p esen and i doesn’ analyze he p e alen concep s o ea u es ha a e common o da a
labeling ools in s uc u ed e ms.
13
4. DATA LABELING TOOLS
The ool selec ion ha ollows was pe o med based on da a ha was manually collec ed un il July
2021. These ools we e selec ed based on hei gene al epu a ion, ma ke adop ion, and he ea u es
hey p o ide o speed up and sol e he da a labeling ask in Machine Lea ning p oblems. E en hough
his disse a ion consis s o a de ailed and s uc u ed analysis, he de ini ion o a c i e ion o he
selec ion o he ools o be s udied is no i ial, o se e al easons. Fi s , a panoply o new da a
labeling ools has been c ea ed and made a ailable la ely, gi en hei demand, which makes choosing
hem di icul , as hey a e o en eleased wi h e y imma u e documen a ion. On he o he hand, he e
a e ools p oduced by la ge echnological companies wo ldwide, and he e a e o he s ha a e
de eloped by pa icula s as side-p ojec s o e en as hobbies, which makes hei compa ison
imp ac icable due o he lack o esou ces associa ed wi h he la e . The inc easing ma ke p essu e
o de eloping new da a labeling ools can be explained by i s own needs and ma e ialized in Figu e 7,
using Google T ends (Google, 2021) o he e m “Supe ised Machine Lea ning”, ha is immedia ely
associa ed o he da a labeling p ocess because o i s dependency on labeled da a. Finally, pe sonal
expe ience and p e e ence migh ha e an undesi able impac on he me iculous de ini ion o he s udy
po en ial ha ools may hold. Thus, he selec ion o he da a labeling ools was based on he
knowledge acqui ed h oughou he wo k expe ience, om he sha ing o people and o ums o
e e ence in he domain, om publica ions in scien i ic jou nals, om code sha ing pla o ms, om
news on ma ke adop ion, and om men ions in ele an con e ences in he subjec .
As such, Table 1 enume a es he conside ed ools ha we e analyzed in his disse a ion, along wi h
he espec i e e e ence and wi h each co esponding de eloping en i y o b and, i applicable.
Figu e 7 – F equency ha he e m “Supe ised Machine Lea ning” was sea ched o
on Google, in he wo ld and in mul iple languages, om 2016 un il he p esen
14
Tool Name
De elope /B and
Re e ence
Colabele
Colabele
N/A
CVAT
In el
A cos -e ec i e, as , and
obus anno a ion ool (Said
e al., 2017)
di g am
Di g am
N/A
ImageTagge
Hambu g Bi -Bo s, Uni e si y o
Hambu g
ImageTagge : An Open
Sou ce Online Pla o m o
Collabo a i e Image Labeling
(Fiedle e al., 2019)
Label S udio
Hea ex
N/A
Labelbox
LabelBox
N/A
LabelD
N/A
N/A
LabelImg
N/A
N/A
LabelMe
Compu e Science and A i icial
In elligence Labo a o y, MIT
LabelMe: A Da abase and
Web-Based Tool o Image
Anno a ion (Russell e al.,
2008)
makesense.ai
makesense.ai
N/A
Playmen
Playmen , TELUS In e na ional
N/A
Ra snake
N/A
Ra snake: A Ve sa ile Image
Anno a ion Tool wi h
Applica ion o Compu e -
Aided Diagnosis (Iako idis e
al., 2014)
Rec Label
N/A
N/A
Remo.ai
Redisco e y.io
N/A
V7 Da win
V7
N/A
VGG Image Anno a ion
Visual Geome y G oup, Uni e si y
o Ox o d
The VIA anno a ion so wa e
o images, audio and ideo
(Du a & Zisse man, 2019)
VoTT
Mic oso
N/A
COCO Anno a o
N/A
N/A
EVA
N/A
N/A
Supe Anno a e
Supe Anno a e
N/A
Table 1 – Manually selec ed da a anno a ion ools and espec i e de elope eams and e e ences,
i applicable
Du ing his e ision, i was ealized ha a la ge pa o he s udied so wa e ools does no ha e an
associa ed a icle o published da a. In o ma ion abou hei au ho s is also di icul o each, which
demons a es ha hese ools a e e iden ly di ided in o 3 g oups: 1) ools ha we e de eloped
speci ically o comme cial pu poses; 2) ools ha we e de eloped acco ding o he au ho s' own
needs, and; 3) ools ha a e ocused on scien i ic esea ch and imp o emen in o de o make he da a
anno a ion p ocess as agile as possible. Fo hese easons, Table 1 displays de elope eam/b and as
“N/A” when in o ma ion abou he so wa e’s au ho s is no publicly a ailable o when hey we e a
15
dynamic and changeable eam o con ibu o s h oughou ime and whe e e e y de elope had
speci ic In o ma ion Technology (IT) knowledge and con ibu ed o an Open-Sou ce p ojec .
Taking his in o conside a ion, his i s analysis demons a es in a glance ha he main de elope
eams o b ands in ol ed in Da a Anno a ion ools o amewo ks a e big ech companies, specialized
ech s a -ups mainly based on Silicon Valley in he Uni ed S a es o Ame ica, o globally enowned
academic esea ch g oups. I p o es ha nowadays, he da a anno a ion ools a e a p io i y o he
mos ecognized IT companies and scien i ic en i ies a ound he wo ld and hei in es men s. These
ools a e o en sold as SaaS (So wa e-as-a-Se ice), mone izing no only he p oduc i sel bu also he
op ional se ice o ou sou cing da a labele s, cons i u ing an ex emely aluable asse o p ese e,
gi en he high demand ha only ends o inc ease e en u he (Moulik, 2020).
Aiming a deep analysis o he echnical ea u es o he selec ed s a e-o - he-a da a anno a ion
ools, mul iple clus e s o unc ionali ies coupled wi h hei espec i e desc ip ions, g ouped by
unc ional ields o ac i i y a e ollowed. Each subchap e co esponds o a speci ic clus e , whe e he
analyzed ea u es help o peel o he mul iple ools and a able is p esen ed o ha pu pose, whe e
he “✓” ma k indica es he p esence o a gi en unc ionali y, “X” deno es i s absence, and he “?”
shows ha he e was no in o ma ion a ailable on ha gi en subjec a he ime o his disse a ion.
16
4.1. COMPUTER VISION FEATURES
Rega ding CV- ela ed unc ionali ies, unc ionali ies as Bounding Boxes, Polygons, Lines, Key-poin s,
Cuboids, Image classi ica ion and Video labeling we e selec ed o enable a ull compa ison be ween
he selec ed ools.
▪ Bounding Boxes: Func ionali y ha allows he use o d aw ec angula Bounding
Boxes ha de ine he loca ion o he a ge objec s in s a ic images, ypically sui able o objec
de ec ion asks.
▪ Polygons: Tool le s he use o d aw polygons o delimi objec s in s a ic images, usually
needed o ins ance segmen a ion asks.
▪ Lines: Func ionali y o d aw lines o ec o s, usually employed in au onomous d i ing
applica ions in lane when anno a ing he lanes on he highways, o ins ance.
▪ Key-poin s: Pe mi s labeling h ough he connec ion o key-poin s o build a skele on
and o unde s and mo e easily wha is labeled, widely known o mo ion acking, acial
landma k de ec ion and hand ges u e ecogni ion.
▪ Cuboids: Allows 3D da a anno a ion, sa ing he dep h and heigh o each objec o
in e es . Usually applied on Objec De ec ion o sel -d i ing ehicles.
▪ Image Classi ica ion: Func ionali y ha associa es an image as a whole o a speci ic
ca ego y.
▪ Video Labeling: Tool pe mi s an easy na iga ion be ween ames o a ideo and make
anno a ions o each sequen ial image. Can also deal wi h ideos mo e complexly, using a
model o es ima e he posi ion o a p e iously anno a ed objec in he ollowing ames.
To pe mi a be e unde s anding abou hese ools’ unc ionali ies on Compu e Vision, Table 2
shows how hey a e used h ough he p e iously selec ed da a labeling ools.
Tool Name
Bounding
Boxes
Polygons
Lines
Key-poin s
Cuboids
Image
Classi ica ion
Video
Labeling
Colabele
✓
✓
✓
X
X
X
✓
CVAT
✓
✓
✓
✓
✓
X
✓
di g am
✓
✓
✓
✓
✓
✓
✓
ImageTagge
✓
✓
✓
✓
X
X
X
Label S udio
✓
✓
✓
✓
X
✓
X
Labelbox
✓
✓
✓
✓
X
✓
✓
LabelD
✓
?
?
?
?
✓
?
LabelImg
✓
X
X
X
X
X
✓
LabelMe
✓
✓
✓
✓
X
✓
✓
makesense.ai
✓
✓
✓
✓
X
✓
✓
17
Playmen
✓
✓
✓
✓
✓
X
✓
Ra snake
✓
✓
✓
✓
X
✓
✓
Rec Label
✓
✓
✓
✓
✓
X
X
Remo.ai
✓
✓
✓
X
X
✓
X
V7 Da win
✓
✓
✓
✓
✓
✓
✓
VGG Image
Anno a ion
✓
✓
✓
✓
✓
✓
✓
VoTT
✓
✓
X
X
X
✓
✓
COCO Anno a o
✓
✓
?
✓
?
?
?
EVA
✓
X
X
X
X
X
✓
Supe Anno a e
✓
✓
✓
✓
✓
✓
✓
Table 2 – Compu e Vision ela ed ea u es in he selec ed da a anno a ion ools
Table 2 shows ha all he selec ed da a anno a ion ools con empla e he possibili y o d awing
Bounding Boxes as a labeling ask. Besides, polygons d awing o ins ance segmen a ion asks is also
e y common. On he o he hand, Cuboids, Video Labeling and Image Classi ica ion seem o be he
leas exis ing ea u es in he da a anno a ion ools. The lack o unc ionali y on Cuboids d awing and
Video Labeling can be explained by he ac ha i s use is o ien ed owa ds uncommon and e y
speci ic use cases which need in o ma ion on dep h o he objec s o eal- ime ideo p ocessing,
espec i ely. Plus, companies ha de elop da a anno a ion ools o mee his equi emen end o use
hem in-house only in o de o ge a compe i i e ad an age. The lack o he unc ionali y ha pe mi s
Image classi ica ion can mean ha he de elope b ands conside his ask as a ainable by a manual
p ocess, ha ing nea ly no cos bene i o de elop i .
24
4.5. GENERAL COMPARATIVE ANALYSIS
Table 6 is shown below as a w ap up o all he pe o med analysis, whe e he ca ego ies o ea u es
p e alen in each da a labeling ool a e easily poin ed ou . The p esen ed alues we e calcula ed based
on he a io o co e ed unc ionali ies unde each o he ou clus e s (Compu e Vision, NLP,
Au oma ion and QA) scope o each analyzed ool.
Tool Name
Compu e
Vision
Na u al
Language
P ocessing
Au oma ion
& De elope
Managemen
and QA
Colabele
57%
100%
50%
0%
CVAT
86%
0%
50%
17%
di g am
100%
100%
100%
83%
ImageTagge
57%
0%
0%
0%
Label S udio
71%
100%
100%
17%
Labelbox
86%
100%
100%
67%
LabelD
29%
0%
0%
0%
LabelImg
29%
0%
0%
0%
LabelMe
86%
0%
0%
17%
makesense.ai
86%
0%
50%
0%
Playmen
86%
0%
0%
17%
Ra snake
86%
0%
50%
17%
Rec Label
71%
0%
50%
0%
Remo.ai
57%
0%
100%
33%
V7 Da win
100%
0%
100%
67%
VGG Image
Anno a ion
100%
0%
50%
0%
VoTT
57%
0%
50%
33%
COCO Anno a o
43%
0%
50%
0%
EVA
29%
0%
0%
0%
Supe Anno a e
100%
0%
50%
0%
Table 6 – Compa a i e summa y o he ea u es g ouped by ca ego y ac oss he analyzed da a
anno a ion ools
The compa ison be ween he ela i e co e age o he unc ionali y ca ego ies by ool demons a es
only a mino i y o he ools ocus on he NLP- ela ed ea u es, unlike he unc ionali ies ela ed o
Compu e Vision. Mo eo e , he selec ed Au oma ion and De elope - iendly ea u es a e appa en ly
ep esen a i e along mos o he analyzed ools, and Managemen and QA unc ionali ies a e mo e
p esen in he ools ha we e de eloped by a company.
25
5. DISCUSSION
The objec i e o his disse a ion is o compa e he selec ed ools in a s uc u ed way, wi h a iew
o hei po en ial use o da a labeling pu poses o Compu e Vision challenges. In o de o e alua e
he ools in hei o ali y, his compa a i e analysis was ex ended o ea u es ha a e unnecessa y o
CV opics, bu ha migh e en ually suppo an o ganiza ion's decision in selec ing a ool. As such,
mul iple unc ionali ies we e highligh ed ha ocus no only on image, ideo, and ex da a, bu also
on he au oma ion and op imiza ion o he labeling p ocess and i s quali y assu ance. Howe e , Table
6 shows ha he selec ion o ools migh be biased owa ds he ocus on Compu e Vision, which
means ha his analysis has mo e signi icance in his ield.
Ini ially, i was planned o c ea e a poin sys em o e alua e each da a labeling ool ullness, whe e
he p esence o each ea u e would add up one poin o a ool, and he decision on he bes ool
would be judged only by he maximum numbe o poin s ob ained h oughou he analysis. Howe e ,
his compa ison would no be ai o ealis ic since, o example, he weigh ing o impo ance o NLP-
ela ed ea u es is likely o be lowe i he o ganiza ion unde conside a ion is a so wa e house ha
de elops deep lea ning models o in e p e images. In addi ion, he inancial ac o may also ha e an
impac on he choice and he e o e he choice o he da a labeling ool p esen ed in his disse a ion
will only be acco ding o he au ho 's pe spec i e and migh di e acco ding o he ci cums ances o a
eade who belongs o a speci ic o ganiza ion o de elopmen eam.
Gi en he di e si y o pu poses o he analyzed ools, one o mo e ools will be chosen o each
analysis conduc ed. S a ing wi h he main analysis o he unc ionali ies wi h Compu e Vision, he
ools ha p esen all he analyzed unc ionali ies s and ou e iden ly: di g am, V7 Da win, VGG Image
Anno a ion and Supe Anno a e. Rega ding he pilla o NLP unc ionali ies, he ools ha allow ex
classi ica ion and he iden i ica ion o named en i ies a e Colabele , di g am, Label S udio and
Labelbox. I should be no ed ha he pool o ools ha o e hese wo unc ionali ies is qui e di e en
when compa ed o he Compu e Vision o ien ed one, as ex ea u es a e ex emely
unde ep esen ed. This is explained by he scope es ic ions o he di e en p oduc s, which a e
p obably aimed a di e en niche ma ke s o academic acks, as p e iously explained. Conce ning
HITL and being open o he inclusion o cus omizable add-ins, di g am, Label S udio, Labelbox, Remo.ai
and V7 Da win a e he winne s, meaning ha hese a e he ools wi h he mos e sa ili y o mul iple
use cases, while ha ing a educed e o a e needed o achie e he da ase anno a ion goal. As o he
pilla ha ela es o ea u es o Managemen and QA, he ools ha s and ou he mos a e di g am,
Labelbox, Remo.ai and V7 Da win. The unc ionali ies o ien ed o p ojec managemen , ask planning
and assignmen and da a managemen a e ce ainly unde de eloped in he labeling ools ma ke
o e all, bu a e mos ly p esen in he co esponding de elopmen oadmaps, which o esee ha hey
26
will be olled ou du ing 2021 o 2022. F om he en i e pale e o e iewed ools, he e a e wo ools
ha s and ou om he compe i ion – V7 Da win and di g am.
V7 Da win has all he necessa y ea u es o da a anno a ion o a Compu e Vision ML p ojec ,
which a e de eloped in a obus , quali y and ex emely use - iendly way. I s majo ocus is on ease o
anno a ion and au oma ion o labeling using machine lea ning models, while being simplis ic and
ha ing an appa en ly sho lea ning cu e. Thus, i is one o he mos p omising ools o image and
ideo labeling, despi e being expensi e and no being open-sou ce which is a g ea disad an age pe
se o no elying on i s own communi y o e ol e.
On he o he hand, di g am is he mos comple e open-sou ce choice ha holis ically co e s he
en i e da a li ecycle, om da a inges ion and mining o in eg a ion wi h cloud o on-p emise machine
lea ning pipelines. In addi ion, i is a pla o m which is only paid o eams consis ing o mo e han 20
use s, which ensu es da a s o age and e sioning, da a labeling o mul iple asks, wo k low
managemen , and da a secu i y - all hese ea u es a e accessible di ec ly om i s pla o m o i s
Applica ion P og amming In e ace (API). Mo eo e , i is based on a e y ac i e Gi Hub eposi o y and
has mo e han 500 s a s. Plus, i ensu es in e pola ion inside ideos o absol e he labele s’ e o in
ha ing o label e e y single ame in e e y imes amp o he ideo, using an objec acke and sma
ame compa ison heu is ics. Un o una ely, i does no ye co e he unc ionali y needed o NLP
p oblems, al hough hese will be on nex yea 's de elopmen oadmap, as well as audio- ela ed
unc ionali y.
27
6. CONCLUSIONS
This disse a ion ames he e olu ion o he domain o Compu e Vision, om i s incep ion o he
p esen . Pas all he ups and downs o he ield, da a eme ges as he new oil o he 21s cen u y, so
he e is a global shi in he bleeding-edge ma ke ends o da a-d i en cul u e. As such, he need
a ises o s a op imizing he da a labeling p ocess, en isioning he gene a ion o da ase s wi h mo e
and be e da a e en as e , in o de o minimize ime e o and inancial cos s, wi hou penalizing he
labeling p ocess. Howe e , a panoply o ools was c ea ed wi h hese in en s, c ea ing he need o
cons an ly e iew he s a e-o - he-a and o s a egically pick he igh so wa e o one’s needs. Fo
his, a selec ion o he ools wi h he mos men ions in he scien i ic and indus ial communi y is
p oposed, on which a compa a i e analysis is made o elec he mos e olu iona y and comple e ools
o pe o m da a labeling and espond o Compu e Vision o Machine Lea ning p oblems in gene al, in
any o ganiza ion. Howe e , he labeling p ocess is ypically expensi e and edious and migh e en
ha m he easibili y o a p ojec . Thus, he scope o his compa ison seeks o mi iga e and add ess he
bo leneck associa ed wi h he da a labeling phase in a Machine Lea ning p ojec .
F om he echnical poin o iew, he compa a i e analysis o exis ing da a labeling amewo ks
poin ed o he ic o y o di g am so wa e, whose unc ionali ies co e p odigiously he en i e pipeline
o a da a p ojec , ea u ing da a inges ion, anno a ion, in eg a ion, explo a ion and p oduc ion in he
o m o Machine Lea ning models. Di g am is open-sou ce and main ains a public de elopmen
oadmap, le e aging he communi y pa icipa ion in i s e olu ion.
In e ms o mo e business- ela ed insigh s abou his ool, di g am has a ee e sion o eams
wi h less han 20 use s and can p e en mul iple e o s and da a edundancy while a oiding he daily
impo s and expo s o da a be ween di e en ools. By cen alizing he en i e pipeline in one ool, a
eam can: 1) sho en i s own lea ning cu e, which is a huge ad an age gi en he ma ke p essu e o
quickly ain Full- ime Equi alen employees (FTEs) in new echnologies; 2) minimize po en ial secu i y
p oblems; 3) educe licensing cos s, and; 4) a oid edundancy o s o ed da a. Mo eo e , he ac o
being open-sou ce opens a panoply o possibili ies o quickly adding and es ing new ea u es o he
p oduc .
28
7. LIMITATIONS AND RECOMMENDATIONS FOR FUTURE WORKS
A p ac ical componen was planned o accompany his disse a ion – he wo k was in ended o
consis in he c ea ion o a new da a labeling ool, which was which was quickly seen as oo ambi ious
o a mas e hesis disse a ion. Then, he p ac ical componen shaped i sel in o appending a new
ea u e/ unc ionali y o an al eady exis ing da a labeling ool. Howe e , as his wo k ad anced
h oughou he exis ing documen a ion on he mul iple analyzed ools, he mo e i allowed o ealize
ha a unc ionali y designed and de eloped in less han one yea would no be able o compe e wi h
a ool de eloped by a niche company o a la ge IT leade , since i would ine i ably all sho o quali y
and complexi y o he ea u es al eady p esen in o he ools as hey a e olled-ou and p oduc ized
by la ge, specialized, de elope eams wi h much mo e c i ical mass.
Also, he lack o p emium/en e p ise licenses in non-open-sou ce ools was one o he limi a ions
ound du ing his wo k. As a ecommenda ion o u u e wo k, i is sugges ed o con ac he owne s o
hese ools o eques and ob ain p emium/en e p ise licenses o academic pu poses. The usage o
such licenses may allow his wo k o go in o mo e de ail a he echnical le el and mo e ob ious
in e ence o how he backends o he ools wo k.
Once a license is gi en, one o he possible essen ial aspec s o explo e is he compa ison o objec
acke s ega ding ideo labeling. An objec acke ha p esen s a be e pe o mance can also
encou age he choice o i s ool, as he ideo labeling p ocess can be hugely op imized, as he acke
i sel can spa e he da a labele o anno a ing all he ames in a ideo while using in o ma ion collec ed
in p e ious ideo ames o help he consequen ones. Fu he mo e, u u e wo ks could conside using
mul iple ools o achie e a labeled da ase while main aining he same eam o da a labele s, in o de
o measu e and compa e he e iciency o each labeling ool in e ms o e o and ime consumed in
a eal-li e scena io.
29
8. BIBLIOGRAPHY
Aa s, E., & Ko s , J. (1987). Simula ed Annealing: Theo y and Applica ion. Simula ed Annealing: Theo y
and Applica ion, 7. h ps://link-sp inge -
com.ezp oxy2.lib a y.colos a e.edu/con en /pd /10.1007%2F978-94-015-7744-1_2.pd
Abbas, A., Su e , D., Zou al, C., Lucchi, A., Figalli, A., & Woe ne , S. (2021). The powe o quan um
neu al ne wo ks. Na u e Compu a ional Science, 1(6), 403–409. h ps://doi.o g/10.1038/s43588-
021-00084-1
Aga , J. O. N. (2020). Wha is science o ? The Ligh hill epo on a i icial in elligence ein e p e ed.
B i ish Jou nal o he His o y o Science, 53(3), 289–310.
h ps://doi.o g/10.1017/S0007087420000230
Aga wala, A., Don che a, M., Ag awala, M., D ucke , S., Colbu n, A., Cu less, B., Salesin, D., & Cohen,
M. (2004). In e ac i e digi al pho omon age. ACM SIGGRAPH 2004 Pape s, SIGGRAPH 2004, 294–
302. h ps://doi.o g/10.1145/1186562.1015718
Belongie, S., Malik, J., & Puzicha, J. (2002). Shape ma ching and objec ecogni ion using shape
con ex s. IEEE T ansac ions on Pa e n Analysis and Machine In elligence, 24(4), 509–522.
h ps://doi.o g/10.1109/34.993558
Be e o, M., Poggio, T. A., & To e, V. (1988). Ill-Posed P oblems in Ea ly Vision. P oceedings o he
IEEE, 76(8), 869–889. h ps://doi.o g/10.1109/5.5962
Besl, P. J., & Jain, R. C. (1985). Th ee-dimensional objec ecogni ion. ACM Compu ing Su eys (CSUR),
17(1), 75–145. h ps://doi.o g/10.1145/4078.4081
Blake, A., & Isa d, M. (1998). Ac i e Con ou s. In Ac i e Con ou s. Sp inge London.
h ps://doi.o g/10.1007/978-1-4471-1555-7
B own, M., & Lowe, D. G. (2007). Au oma ic pano amic image s i ching using in a ian ea u es.
In e na ional Jou nal o Compu e Vision, 74(1), 59–73. h ps://doi.o g/10.1007/s11263-006-
0002-3
Bu ka , N., & Hube , M. F. (2021). A Su ey on he Explainabili y o Supe ised Machine Lea ning.
Jou nal o A i icial In elligence Resea ch, 70, 245–317. h ps://doi.o g/10.1613/jai .1.12228
Chang, J. C., Ame shi, S., & Kama , E. (2017). Re ol : Collabo a i e c owdsou cing o labeling machine
lea ning da ase s. Con e ence on Human Fac o s in Compu ing Sys ems - P oceedings, 2017-May,
2334–2346. h ps://doi.o g/10.1145/3025453.3026044
Comaniciu, D., & Mee , P. (2002). Mean shi : a obus app oach owa d ea u e space analysis. IEEE
T ansac ions on Pa e n Analysis and Machine In elligence, 24(5), 603–619.
h ps://doi.o g/10.1109/34.1000236
Da id, M., & Jayan , S. (1989). Op imal App oxima ions by Piecewise Smoo h Func ions and Associa ed
30
Va ia ional P oblems. Communica ions on Pu e and Applied Ma hema ics, 42, 577–685.
Debe ec, P. E., & Malik, J. (1997). Reco e ing high dynamic ange adiance maps om pho og aphs.
P oceedings o he 24 h Annual Con e ence on Compu e G aphics and In e ac i e Techniques -
SIGGRAPH ’97, 3(1), 369–378. h ps://doi.o g/10.1145/258734.258884
DodgeSpecial, J. (2001). D ape P ize Hono s Fou “Fa he s o he In e ne .”
h ps://www.wsj.com/a icles/SB982004616905008338
Du a, A., & Zisse man, A. (2019). The VIA anno a ion so wa e o images, audio and ideo. MM 2019
- P oceedings o he 27 h ACM In e na ional Con e ence on Mul imedia, 2276–2279.
h ps://doi.o g/10.1145/3343031.3350535
E e h, J. (2018). Da aOps – Towa ds a de ini ion. CEUR Wo kshop P oceedings, 2191, 104–112.
Facebook AI Resea ch. (2021). Pape sWi hCode.com.
Fauge as, O. D., & Hebe , M. (1986). The Rep esen a ion, Recogni ion, and Loca ing o 3-D Objec s.
The In e na ional Jou nal o Robo ics Resea ch, 5(3), 27–52.
h ps://doi.o g/10.1177/027836498600500302
Fe gus, R., Pe ona, P., & Zisse man, A. (2007). Weakly supe ised scale-in a ian lea ning o models o
isual ecogni ion. In e na ional Jou nal o Compu e Vision, 71(3), 273–303.
h ps://doi.o g/10.1007/s11263-006-8707-x
Fiedle , N., Bes mann, M., & Hend ich, N. (2019). ImageTagge : An Open Sou ce Online Pla o m o
Collabo a i e Image Labeling. In Lec u e No es in Compu e Science (including subse ies Lec u e
No es in A i icial In elligence and Lec u e No es in Bioin o ma ics): Vol. 11374 LNAI (pp. 162–
169). h ps://doi.o g/10.1007/978-3-030-27544-0_13
Fo sy h, D., & Ponce, J. (2003). Compu e Vision: A Mode n App oach. P en ice Hall P o essional
Technical Re e ence.
Ga ne Inc. (2021). Ga ne . Hype Cycle o A i icial In elligence 2021.
h ps://www.ga ne .com/en/in o ma ion- echnology/insigh s/a i icial-in elligence
Gau , E., Saxena, V., & Singh, S. K. (2018). Video anno a ion ools: A Re iew. P oceedings - IEEE 2018
In e na ional Con e ence on Ad ances in Compu ing, Communica ion Con ol and Ne wo king,
ICACCCN 2018, 911–914. h ps://doi.o g/10.1109/ICACCCN.2018.8748669
Google. (2021). Google T ends. h ps://www.google.com/ ends
Hil on, A., Fua, P., & Ron a d, R. (2006). Modeling people: Vision-based unde s anding o a pe son’s
shape, appea ance, mo emen , and beha iou . Compu e Vision and Image Unde s anding,
104(2–3), 87–89. h ps://doi.o g/10.1016/j.c iu.2006.09.002
Iako idis, D. K., Goudas, T., Smailis, C., & Maglogiannis, I. (2014). Ra snake: A Ve sa ile Image
Anno a ion Tool wi h Applica ion o Compu e -Aided Diagnosis. The Scien i ic Wo ld Jou nal,
2014, 1–12. h ps://doi.o g/10.1155/2014/286856
31
Jianbo Shi, & Malik, J. (2000). No malized cu s and image segmen a ion. IEEE T ansac ions on Pa e n
Analysis and Machine In elligence, 22(8), 888–905. h ps://doi.o g/10.1109/34.868688
Jianbo Shi, & Tomasi. (1994). Good ea u es o ack. P oceedings o IEEE Con e ence on Compu e
Vision and Pa e n Recogni ion CVPR-94, 593–600. h ps://doi.o g/10.1109/CVPR.1994.323794
Kass, M., Wi kin, A., & Te zopoulos, D. (1988). Snakes: Ac i e con ou models. In e na ional Jou nal o
Compu e Vision, 1(4), 321–331. h ps://doi.o g/10.1007/BF00133570
Kim, J. (2020). Applica ion on cha ac e ecogni ion sys em on oad sign o isually impai ed: Case
s udy app oach and u u e. In e na ional Jou nal o Elec ical and Compu e Enginee ing, 10(1),
778–785. h ps://doi.o g/10.11591/ijece. 10i1.pp778-785
Kopuklu, O., Kose, N., Gunduz, A., & Rigoll, G. (2019). Resou ce e icien 3D con olu ional neu al
ne wo ks. P oceedings - 2019 In e na ional Con e ence on Compu e Vision Wo kshop, ICCVW
2019, 1910–1919. h ps://doi.o g/10.1109/ICCVW.2019.00240
K izhe sky, A., Su ske e , I., & Hin on, G. E. (2017). ImageNe classi ica ion wi h deep con olu ional
neu al ne wo ks. Communica ions o he ACM, 60(6), 84–90. h ps://doi.o g/10.1145/3065386
Kshe i, N. (2021). Da a Labeling o he A i icial In elligence Indus y: Economic Impac s in De eloping
Coun ies. IT P o essional, 23(2), 96–99. h ps://doi.o g/10.1109/MITP.2020.2967905
Lani is, A., Taylo , C. J., & Coo es, T. F. (1997). Au oma ic in e p e a ion and coding o ace images using
lexible models. IEEE T ansac ions on Pa e n Analysis and Machine In elligence, 19(7), 743–756.
h ps://doi.o g/10.1109/34.598231
Lecle c, Y. G. (1989). Cons uc ing simple s able desc ip ions o image pa i ioning. In e na ional
Jou nal o Compu e Vision, 3(1), 73–102. h ps://doi.o g/10.1007/BF00054839
Ligoza , A.-L., Le è e, J., Bugeau, A., & Combaz, J. (2021). Un a eling he hidden en i onmen al
impac s o AI solu ions o en i onmen . In P oceedings o ACM Con e ence (Con e ence’17) (Vol.
1, Issue 1). Associa ion o Compu ing Machine y. h p://a xi .o g/abs/2110.11822
Lucas, B. D., & Kanade, T. (1981). I e a i e Image Regis a ion Technique Wi h an Applica ion To S e eo
Vision. 2(Ap il 1981), 674–679.
Malladi, R., Se hian, J. A., & Vemu i, B. C. (1995). Shape Modeling wi h F on P opaga ion: A Le el Se
App oach. IEEE T ansac ions on Pa e n Analysis and Machine In elligence, 17(2), 158–175.
h ps://doi.o g/10.1109/34.368173
Ma , D. (1982). Vision: a compu a ional in es iga ion in o he human ep esen a ion and p ocessing
o isual in o ma ion. In Vision: a compu a ional in es iga ion in o he human ep esen a ion and
p ocessing o isual in o ma ion. W. H. F eeman.
Ma hews, I., Xiao, J., & Bake , S. (2007). 2D s. 3D De o mable Face Models: Rep esen a ional Powe ,
Cons uc ion, and Real-Time Fi ing. In e na ional Jou nal o Compu e Vision, 75(1), 93–113.
h ps://doi.o g/10.1007/s11263-007-0043-2
32
Ma hews, J., & Bake , S. (2004). Ac i e appea ance models e isi ed. In e na ional Jou nal o
Compu e Vision, 60(2), 135–164. h ps://doi.o g/10.1023/B:VISI.0000029666.37597.d3
Moeslund, T. B., Hil on, A., & K üge , V. (2006). A su ey o ad ances in ision-based human mo ion
cap u e and analysis. Compu e Vision and Image Unde s anding, 104(2-3 SPEC. ISS.), 90–126.
h ps://doi.o g/10.1016/j.c iu.2006.08.002
Mona ch, R. (Mun o). (2021). Human-in- he-Loop Machine Lea ning: Ac i e lea ning and anno a ion
o human-cen e ed AI. Manning Publica ions Co.
Mo i, G., Xiao eng Ren, E os, A. A., & Malik, J. (2004). Reco e ing human body con igu a ions:
combining segmen a ion and ecogni ion. P oceedings o he 2004 IEEE Compu e Socie y
Con e ence on Compu e Vision and Pa e n Recogni ion, 2004. CVPR 2004., 2, 326–333.
h ps://doi.o g/10.1109/CVPR.2004.1315182
Moulik, S. (2020). Da a as he New Cu ency—How Open Sou ce Toolki s Ha e Made Labeled Da a he
Co e Value in he AI Ma ke place. Academic Radiology, 27(1), 140–142.
h ps://doi.o g/10.1016/j.ac a.2019.09.016
Mundy, J. L. (2006). Towa d Ca ego y-Le el Objec Recogni ion (J. Ponce, M. Hebe , C. Schmid, & A.
Zisse man (eds.); Vol. 4170). Sp inge Be lin Heidelbe g. h ps://doi.o g/10.1007/11957959
Nguyen, T. Q., & Salaza , J. (2019). T ans o me s wi hou Tea s: Imp o ing he No maliza ion o Sel -
A en ion. 1. h ps://doi.o g/10.5281/zenodo.3525484
No ig, S. R. and P. (2019). A i icial In elligence A Mode n App oach 4 h Ed. In Jou nal o Chemical
In o ma ion and Modeling (Vol. 53, Issue 9).
Pape , S. (1966). The summe ision p ojec (pp. 1–6). h p://dspace.mi .edu/handle/1721.1/6125
Pe schnigg, G., Szeliski, R., Ag awala, M., Cohen, M., Hoppe, H., & Toyama, K. (2004). Digi al
pho og aphy wi h lash and no- lash image pai s. ACM T ansac ions on G aphics, 23(3), 664–672.
h ps://doi.o g/10.1145/1015706.1015777
Poggio, T., To e, V., & Koch, C. (1987). Compu a ional ision and egula iza ion heo y. Readings in
Compu e Vision, 638–643. h ps://doi.o g/10.1016/b978-0-08-051581-6.50061-1
P ince, S. D. J. (Uni e si y C. L. (2012). Compu e Vision: Models, Lea ning and In e ence.
h p://www.camb idge.o g/9781107011793
Pueyo, S. (2018). G ow h, deg ow h, and he challenge o a i icial supe in elligence. Jou nal o Cleane
P oduc ion, 197, 1731–1736. h ps://doi.o g/10.1016/j.jclep o.2016.12.138
Pul o d, G. W. (2005). Taxonomy o mul iple a ge acking me hods. IEE P oceedings: Rada , Sona
and Na iga ion, 152(5), 291–304. h ps://doi.o g/10.1049/ip- sn:20045064
Ramachand a, V. (2019). Causal in e ence o clima e change e en s om sa elli e image ime se ies
using compu e ision and deep lea ning. h p://a xi .o g/abs/1910.11492
Rehg, J. M., & Kanade, T. (1994). Visual acking o high DOF a icula ed s uc u es: An applica ion o
33
human hand acking. In Lec u e No es in Compu e Science (including subse ies Lec u e No es in
A i icial In elligence and Lec u e No es in Bioin o ma ics): Vol. 801 LNCS (pp. 35–46).
h ps://doi.o g/10.1007/BFb0028333
Resea ch, C. (2019). Da a Enginee ing, P epa a ion, and Labeling o AI 2019: Technical Repo .
Resea ch, G. V. (2021). Da a Collec ion And Labeling Ma ke Size, Sha e & T ends Analysis Repo .
h ps://doi.o g/10.1109/CVPR.2004.1315182
Robe s, L. G. (1963). Machine pe cep ion o h ee-dimensional solids. No embe .
h p://dspace.mi .edu/handle/1721.1/11589
Rosen eld, A. (1984). Some Use ul P ope ies o Py amids. 2–5. h ps://doi.o g/10.1007/978-3-642-
51590-3_1
Rosen eld, Az iel, & Kak, A. C. (2019). Digi al pic u e p ocessing 2nd edi ion. In Jou nal o Chemical
In o ma ion and Modeling (Vol. 1, Issue 1).
Rosen eld, Az iel, & P al z, J. L. (1966). Sequen ial Ope a ions in Digi al Pic u e P ocessing. Jou nal o
he ACM (JACM), 13(4), 471–494. h ps://doi.o g/10.1145/321356.321357
Russell, B. C., To alba, A., Mu phy, K. P., & F eeman, W. T. (2008). LabelMe: A Da abase and Web-
Based Tool o Image Anno a ion. In e na ional Jou nal o Compu e Vision, 77(1–3), 157–173.
h ps://doi.o g/10.1007/s11263-007-0090-8
S., L. L., Blake, A., & Zisse man, A. (1987). Visual Recons uc ion. In Ma hema ics o Compu a ion (Vol.
53, Issue 188). h ps://doi.o g/10.2307/2008745
Said, A. F., Kashyap, V., Choudhu y, N., & Akhba i, F. (2017). A cos -e ec i e, as , and obus
anno a ion ool. P oceedings - Applied Image y Pa e n Recogni ion Wo kshop, 2017-Oc ob, 1–6.
h ps://doi.o g/10.1109/AIPR.2017.8457958
Saxena, D., & Cao, J. (2020). Gene a i e Ad e sa ial Ne wo ks (GANs): Challenges, Solu ions, and Fu u e
Di ec ions. 54(3). h p://a xi .o g/abs/2005.00065
Schei e , W. J., An hony, S. E., Nakayama, K., & Cox, D. D. (2014). Pe cep ual anno a ion: Measu ing
human ision o imp o e compu e ision. IEEE T ansac ions on Pa e n Analysis and Machine
In elligence, 36(8), 1679–1686. h ps://doi.o g/10.1109/TPAMI.2013.2297711
Shapi o, L. G. (2020). Compu e ision: he las 50 yea s. In e na ional Jou nal o Pa allel, Eme gen
and Dis ibu ed Sys ems, 35(2), 112–117. h ps://doi.o g/10.1080/17445760.2018.1469018
Sidenbladh, H., Black, M. J., & Flee , D. J. (2000). S ochas ic T acking o 3D Human Figu es Using 2D
Image Mo ion. In Jou nal o Explosi es Enginee ing (Vol. 31, Issue 6, pp. 702–718).
h ps://doi.o g/10.1007/3-540-45053-X_45
Solem, J. E. (2012). P og amming Compu e Vision wi h Py hon: Tools and algo i hms o analyzing
images (Vol. 1, Issue 1). O’Reilly Media, Inc.
Szeliski, R. (2010). Compu e Vision: Algo i hms and Applica ions (Vol. 1). Sp inge -Ve lag London.