scieee Open visual document viewer

Tackling automatic audience experience measurement in online environments

Villalobos Sánchez, Pablo; Rivero Rodríguez, Eduardo

Abstract

The availability of automatic and personalized feedback is a large advantage when facing an audience. An effective way to give such feedback is to analyze the audience experience, which provides valuable information about the quality of a speech or performance. In this document, we present the design and implementation of a computer vision system to automatically measure audience experience. This includes the definition of a theoretical and practical framework grounded on the theatrical perspective to quantify this concept, the development of an artificial intelligence system which serves as a proof-of-concept of our approach, and the creation of a dataset to train our system. To facilitate the data collection step, we have also created a custom video conferencing tool. Additionally, we present the evaluation of our artificial intelligence system and the final conclusions.

Full text

Tackling au oma ic audience expe ience measu emen in online en i onmen s Abo dando la medición au omá ica de la expe iencia de la audiencia en línea Po Pablo Villalobos Sánchez y Edua do Ri e o Rod íguez T abajo de in de g ado del Doble G ado en Ingenie ía In o má ica y Ma emá icas Facul ad de In o má ica Di igido po : Bo ja Mane o Iglesias Me iem El Yam i El Kha ibi Mad id, 2020–2021 Abs ac The a ailabili y o au oma ic and pe sonalized eedback is a la ge ad- an age when acing an audience. An e ec i e way o gi e such eedback is o analyze he audience expe ience, which p o ides aluable in o ma ion abou he quali y o a speech o pe o mance. In his documen , we p esen he design and implemen a ion o a compu e ision sys em o au oma ically measu e audience expe ience. This includes he de ini ion o a heo e ical and p ac ical amewo k g ounded on he hea ical pe spec i e o quan i y his concep , he de elopmen o an a i icial in elligence sys em which se es as a p oo -o -concep o ou app oach, and he c ea ion o a da ase o ain ou sys em. To acili a e he da a collec ion s ep, we ha e also c ea ed a cus om ideo con e encing ool. Addi ionally, we p esen he e alua ion o ou a i icial in elligence sys em and he inal conclusions. Keywo ds –compu e ision, machine lea ning, sen imen analysis, emo- ion ecogni ion, objec acking, a ec i e compu ing, WebRTC Resumen La disponibilidad de eedback au omá ico y pe sonalizado supone una g an en aja a la ho a de en en a se a un público. Una o ma e ec i a de da es e ipo de eedback es analiza la expe iencia de la audiencia, que p opo - ciona in o mación undamen al sob e la calidad de una ponencia o ac uación. En es e documen o exponemos el diseño e implemen ación de un sis ema au- omá ico de medición de la expe iencia de la audiencia basado en la isión po compu ado . Es o incluye la de inición de un ma co eó ico y p ác ico undamen ado en la pe spec i a del mundo del ea o pa a cuan i ica el con- cep o de expe iencia de la audiencia, el desa ollo de un sis ema basado en in eligencia a i icial que si e como p o o ipo de nues a ap oximación y la ecopilación un conjun o de da os pa a en ena el sis ema. Pa a acili a es e úl imo paso hemos desa olado una aplicación de ideocon e encias pe sonal- izada. Además, en es e abajo p esen amos la e aluación de nues o sis ema de in eligencia a i icial y las conclusiones ex aídas. Palab as cla e – isión po compu ado , ap endizaje au omá ico, análi- sis de sen imien o, econocimien o de emociones, seguimien o de obje os, com- pu ación a ec i a, WebRTC Acknowledgemen s We woud like o hank Bo ja Mane o Iglesias, Alejand o Rome o He nán- dez, and Me iem El Yam i El Kha ibi o hei di ec con ibu ions, eedback, and con inuous suppo h oughou he en i e yea . Wi hou hem his would no ha e been possible. We also hank Sco , who helped us co ec he pa- pe ; ou amilies and iends, who ha e been suppo ing us un il now; and he olun ee s in ou expe imen , who gene ously len us hei ime and pa ience. Con en s 1 In oduc ion 1 1.1 Wo kplan ................................... 2 1.2 Indi idual con ibu ions . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.2.1 Edua do Ri e o Rod íguez . . . . . . . . . . . . . . . . . . . . . . 4 1.2.2 Pablo Villalobos Sánchez . . . . . . . . . . . . . . . . . . . . . . . 5 2 S a e o he A 7 2.1 Objec de ec ion ............................... 7 2.1.1 R-CNN ................................ 8 2.1.2 Fas R-CNN ............................. 10 2.1.3 Fas e R-CNN ............................ 11 2.1.4 YOLO................................. 12 2.2 Mul iple Objec acking . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.2.1 ROLO................................. 15 2.2.2 SORT................................. 15 2.2.3 DeepSORT.............................. 16 2.3 Public speaking and a ec i e compu ing . . . . . . . . . . . . . . . . . . 16 2.3.1 Public speaking aining sys ems . . . . . . . . . . . . . . . . . . 16 2.3.2 Cha ac e izing and quan i ying audience expe ience . . . . . . . . 18 2.3.3 Public speaking da ase s . . . . . . . . . . . . . . . . . . . . . . . 20 2.4 Emo ion ecogni ion ............................. 21 2.4.1 Measu ing he engagemen le el o TV iewe s . . . . . . . . . . 21 2.4.2 Measu ing he engagemen le el o s uden s in he class oom . . . 22 2.4.3 Recognizing emo ions in mo ie audiences using a ia ional au oen- code s................................. 23 2.5 Videocon e ence sys ems . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 2.5.1 GoogleMee ............................. 24 2.5.2 BbCollabo a e............................ 24 2.5.3 Zoom ................................. 24 2.5.4 Ji si .................................. 24 3 The ERVF Da ase 25 3.1 Quan i ying audience expe ience . . . . . . . . . . . . . . . . . . . . . . . 25 3.2 Expe imen aldesign ............................. 27 3.2.1 Toolsused............................... 27 3.2.2 Me hodology ............................. 31 3.3 Expe imen al esul s ............................. 32 3.4 Limi a ions .................................. 34 CONTENTS 4 The isual acking module 35 4.1 Modulea chi ec u e ............................. 35 4.1.1 Single ame objec de ec ion wi h YOLO 4 . . . . . . . . . . . . 35 4.1.2 The acking algo i hm . . . . . . . . . . . . . . . . . . . . . . . . 40 4.2 Implemen a ion................................ 44 5 The emo ion ecogni ion module 49 5.1 A chi ec u e.................................. 50 5.1.1 MobileNe V3 ............................. 50 5.1.2 Reg esso s............................... 53 5.2 Implemen a ion................................ 53 5.2.1 Ne wo k a chi ec u e . . . . . . . . . . . . . . . . . . . . . . . . . 54 5.2.2 Da ase loading............................ 54 5.2.3 T aining, alida ion, and es ing . . . . . . . . . . . . . . . . . . . 57 5.3 Resul s and limi a ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 6 Conclusions 63 Appendices 65 A Code 67 B Pape and submission con i ma ion 83 Glossa y 95 Chap e 1 In oduc ion Public speaking is a co e skill in he mode n wo ld. Bo h educa o s and lea ne s sha e a common desi e owa ds be e eaching me hods o his elusi e skill. In pa icu- la , au oma ed eedback sys ems o public speaking aining ha e spa ked he in e es o esea che s due o hei p omise o objec i e and pe sonalized ad ice a a massi e scale. The e ha e been a ious p oposals o his kind o sys em o e he las yea s and, wi h he ecen ad ances in machine lea ning (ML) and a ec i e compu ing, he capabili ies o hese sys ems ha e inc eased. The majo i y o he p oposed sys ems di ec ly e alua e he e bal and non e bal beha io o he speake and hey a e usually ained om expe a ings o public speaking pe o mances. Howe e , he ue judge o public speaking skill is he audience, which collec i ely de- cides which speake s a e inspi ing and which a e no . Needless o say, hese audience e alua ions a e no explici , bu emain implici in hei expe ience o he pe o mance. Fu he mo e, he p ocess ha de e mines he na u e o ha expe ience is poo ly unde - s ood, e en by he audience membe s hemsel es. I ollows ha , i we had quan i ied and au oma ic access o he inne expe ience o audience membe s, his would cons i u e a ue gold s anda d o measu ing speake pe o mance. Thus, he ocus o ou wo k is achie ing an au oma ed assessmen o audience expe ience. How o use his documen In his documen , we will e e ence se e al concep s om machine lea ning and a ec i e compu ing. While we p o ide a glossa y o e e ence, we will assume he eade is amilia wi h hese opics. We will documen ou wo k wi h as much de ail as possible, and will p o ide wo king code o e e y so wa e componen used. The code is published unde he MIT license, and he eade is welcome o use i and imp o e upon i . Ou goals As men ioned be o e, ou end goal is p oducing an au oma ed sys em o accu a ely measu ing audience expe ience. Howe e , his is no small ask, and a ully unc ional, eady o use sys em is ou o each o his expe imen al wo k. The e o e, we cla i ied ou pu pose by di iding i in o h ee goals. 1. De e mining a p ac ical amewo k o quan i y audience expe ience. This ame- wo k should be g ounded in exis ing heo y as well as in empi ical da a. Ideally, 1 8CHAPTER 2. STATE OF THE ART 2.1.1 R-CNN Regions wi h CNN Fea u es, commonly known as R-CNN, is a 3-s ep objec ecogni ion me hod ha elies on CNNs o ex ac ea u es om di e en egions in he image which a e hen used o p edic ion. Ini ially, he image is p ocessed h ough a egion p oposal sys em, ha ex ac s egions o he image ha may con ain an objec . Then, he image is passed h ough a CNN ha ans o ms each egion in o a 4096-dimensional ep esen- a ion. Finally, he e is a Suppo Vec o Machine (SVM) o each objec class ha indica es whe he he e is an objec in a gi en egion and, in he a i ma i e case, a linea eg ession model indica es he posi ion and dimensions o he bounding box. CNN ca ? yes plan ? no backg ound? no Inpu image 1. Region p oposal 2. Con olu ional ea u e ex ac ion 3. Classi ica ion Figu e 2.1: Classi ica ion p ocess wi h R-CNN Region p oposal The egion p oposal model elies on a pa icula egion ex ac ion algo i hm, he mos commonly used one being selec i e sea ch. Selec i e sea ch wo ks in wo-s eps. Fi s ly, i makes use o Felzenszwalb and Hu enloche ’s algo i hm o ob ain an ini ial se o egions. Then i p oceeds o i e a i ely educe he amoun o such egions by me ging neighbo ing egions wi h he highes deg ee o simila i y. Algo i hm 1: Selec i e sea ch [64] Da a: RGB Image Resul : Objec loca ion hypo hesis L={l1, ..., lN} Ob ain ini ial egions R={ 1, ..., n} h ough Felzenszwalb and Hu enloche ’s algo i hm Ini ialise he simila i y se S=ϕ o ( i, j)neighbo ing egion pai do Calcula e simila i y s( i, j) S=S∪ {s( i, j)} end while S=ϕdo Ge highes simila i y pai ( i, j)such ha s( i, j) = max(S) Me ge co esponding egions = i∪ j Remo e simila i ies ega ding iS=S s( i, ∗) Remo e simila i ies ega ding jS=S s( j, ∗) Calcula e simila i y se S be ween and i s neighbo s S=S∪S R=R∪ { } end Ex ac objec loca ion bounding boxes L om all egions R 2.1. OBJECT DETECTION 9 The ini ial egion p oposal is calcula ed h ough Felzenswalb and Hu enloche ’s algo- i hm, which ea s images as g aphs and de e mines a p elimina y componen segmen- a ion. A monoch ome (in ensi y) image is a g aph whe e each pixel pihas an associa ed e ex i∈V. Each one o hose e ices is connec ed o he e ices associa ed wi h neighbo ing pixels. The weigh o hose edges is gi en by w(e i, j) = |I(pi)−I(pj)| whe e I(pk)is he in ensi y o pixel pk. We also need an ope a o o compa e wo sepa a e componen s and, he e o e, de e mine whe he hey should be me ged. We conside he ollowing unc ions named, espec i ely, he in e nal di e ence o a componen C⊂V and he minimum in e nal di e ence be ween wo componen s In (C) = max e∈MST (C,E)w(e) MIn (C1, C2) = min(In (C1) + τ(C1), In (C2) + τ(C2)) whe e MST (C, E)is he minimum spanning ee o he componen Cand τis a non- nega i e h eshold unc ion, usually τ(C) = k Cwi h kbeing a pa icula cons an ha de e mines (in p ac ice) how la ge he inal componen s will be. I is conside ed ha wo componen s should be me ged when he weigh o an edge joining hem is smalle han hei hei minimum in e nal di e ence. Algo i hm 2: Felzenszwalb and Hu enloche ’s algo i hm [21] Da a: G= (V, E) Resul : Segmen a ion componen s S= (C1, ..., C ) ha pa i ion he g aph So E in o π= (o1, ..., om)by non-dec easing weigh S a wi h segmen a ion S0whe e each e ex is in i s own componen o q= 1, .., m do Le Vi, Vjdeno e he e ices connec ed by he q- h edge in he o de ing i Viand Vja e in disjoin componen s and ω(q)≤MIn (Cq−1 i, Cq−1 j) hen Sqis ob ained by me ging Cq−1 iand Cq−1 j else Sq=Sq−1 end end Re u n S=Sm Fea u e ex ac ion In he o iginal pape [26][43], ea u e ex ac ion is done h ough a e y simple CNN a chi- ec u e consis ing o mean subs ac ion and egion wa ping (227 ×227) as p ep ocessing, 5 con olu ional laye s (using ReLU ac i a ion) and 2 Fully Connec ed (FC) laye s. Each egion is inally ep esen ed by a 4096-dimensional ec o . The e a e addi ional max pooling and local esponse no maliza ion [43] ope a ions applied a e con olu ional lay- e s CONV 1,CONV 2, and CONV 5. 10 CHAPTER 2. STATE OF THE ART P edic ion: classes and bounding boxes An SVM ained o each class sco es he gi en egion ea u e ec o s and egions ha ing high in e sec ion-o e -union (IoU) wi h a highe sco e egion (la ge han a ce ain lea ned h eshold) a e disca ded. The inal e sion o he model also includes a linea eg ession model ha p edic s a new de ec ion window gi en he max pooled esul s a e CONV 5, calcula ing new bounding boxes o he gi en objec s [22]. 2.1.2 Fas R-CNN Fas R-CNN appea s as an i e a ion o R-CNN ha achie es a conside able speedup on i s p edecesso [25]. Fo his pu pose, Fas R-CNN a oids compu ing he CNN ea u es o each o he egions by ex ac ing a global ea u e map om he image using a CNN. Fas R-CNN s ill uses Selec i e Sea ch as a egion p oposal ne wo k be o e p ocessing he image. Each egion p oposal, o Region o In e es (RoI), is p ojec ed on o he ea u e map and a specialized laye called RoI pooling ex ac s a lowe -dimensional ea u e ec o ha will be la e used o p edic ion. The RoI pooling laye wo ks by applying max pooling on he ea u e map o he RoI. I we wished o ge an H×Wdimensional ou pu and we had a RoI wi h op-le co ne coo dina es ( , c), heigh h, and wid h w; he pooling p ocess would p oceed as ollows: he p ojec ed ea u e map ( he sec ion o he image co esponding o he RoI) is di ided in o a g id o w W×h H-dimensional ec angles and hen max pooling is applied on each o he ec angles. Figu e 2.2: Illus a ion o RoI pooling (H=W= 2) The las pa o he a chi ec u e consis s o wo sibling FC laye s. One o hem applies so max o e K+ 1 classes (whe e he e a e Kpossible objec ca ego ies) and he o he one is connec ed o class-speci ic bounding box eg esso s. I is impo an o no e ha he choice o Wand Hmus be compa ible wi h he size o hese FC laye s. The ou pu o hese wo p edic ion laye s is, espec i ely, p= (p0, ..., pk)p obabili y map o e K+ 1 ca ego ies o he objec de ec ion laye and k= ( k x, k y, k w, k h) o each o he Kobjec classes, which indica es he bounding box eg ession o se s. Ano he one o he key componen s in Fas R-CNN’s p oposal is he use o a mul i- ask loss o aining, which allows o end- o-end single s age aining. Each aining RoI has a g ound- u h class uand a g ound- u h bounding-box eg ession a ge . Fo such RoI, he mul i- ask loss is L(p, u, u, ) = Lcls(p, u) + λ[u≥1]Lloc( u, ) whe e Lcls(p, u)and Lloc( u, )a e he classi ica ion and bounding-box eg ession losses gi en by 2.1. OBJECT DETECTION 11 Lcls(p, u) = −log puLloc =X i∈{x,y,w,h} smoo hL1( u i− i) Fu he mo e, [u≥1] e alua es o 1only when u≥1and o 0o he wise. The smoo h L1 loss is a a ia ion o he L1loss (see Fig. 2.3) aimed a educing sensi i i y o ou lie s o he L2loss while main aining is p ope ies and is gi en by smoo hL1(x) = (1 2x2|x|<1 |x| − 1 2o he wise Figu e 2.3: L1loss and smoo hL1loss compa ison. The abo e loss is a pa icula case o he Hube loss [37] Lδ(x) = (1 2x2|x|< δ δ(|x| − 1 2δ)o he wise wi h δ= 1. 2.1.3 Fas e R-CNN Fas e R-CNN is ye ano he i e a ion o R-CNN ha add esses ano he one one o i s p oblems (which i sha es wi h Fas R-CNN): he selec i e sea ch bo leneck [58]. Ins ead o using selec i e sea ch, Fas e R-CNN p oposes he use o a dedica ed CNN called he Region P oposal Ne wo k (RPN). An RPN ecei es an image as inpu and ou pu s a se o egion p oposals ( ec angula in ou case). This is achie ed in p ac ice h ough a CNN. Fu he mo e, bo h he RPN and he CNN o ea u e ex ac ion sha e con olu ional laye s. The RPN wo ks in a sliding window ashion, by sliding he RPN o e an n×nspa ial window o he con olu ional ea u e map ou pu ed by he las sha ed con olu ional laye . This window is mapped on o a smalle dimension ea u e ec o (256-dimensional in he 12 CHAPTER 2. STATE OF THE ART o iginal pape ) which is passed on o he wo p edic ion laye s: he classi ica ion laye (cls) and he eg ession laye ( eg). The ou pu s o hese laye s a e he same as in Fas R-CNN. Ano he impo an pa o Fas e R-CNN’s app oach is he use o ancho boxes. Ancho boxes a e p ede ined shapes (usually ec angles) used o mimic he scale and aspec a ios o he objec s o he class o be p edic ed. The o al numbe o ancho boxes is K=S·A whe e Sis he numbe o scales and A he numbe o aspec a ios (in he o iginal pape S=A= 3). Fas e R-CNN uses he ancho boxes o p edic K egion p oposals o each one o he sliding-window loca ions. This gi es us he size o he ou pu s o he cls laye (2Ksco es, p(objec )and p(no objec )) and he eg (4Ksco es, 4 o each p edic ed bounding box). This whole pa o Fas e R-CNNs a chi ec u e is ansla ion in a ian . The aining o Fas e R-CNN uses a mul i- ask loss (jus as Fas R-CNN) o achie e one s age lea ning. This unc ion akes as inpu he ou pu o he cls and eg laye s o he i h ancho and is gi en by L({pi},{ i}) = 1 Ncls X j Lcls(pj, p∗ j) + λ1 N eg X j p∗ jL eg( j, ∗ j) whe e λis a balancing weigh , Ncls and N eg a e no maliza ion weigh s, pj he j− h alue o {pi}and jis he j− h alue o { ∗ i}whe e p∗ jand ∗ ja e he espec i e g ound- u h alues. Lcls is he log loss and L eg( j, ∗ j) = smoo hL1( j− ∗ j). As o he eg ession alues, he coo dina es a e epa ame ized as ollows x=x−xa wa y=y−ya wa w=log w wa h=log h ha o he eg ession coo dina es x, y, h, w deno ing he midpoin coo dina es (x, y), heigh and wid h espec i ely. The same epa ame iza ion is applied o he g ound- u h al- ues. The eg ession i sel is also di e en om ha o p e ious e sions. A se o Kbounding box eg esso s a e lea ned, each eg esso being being esponsible o one scale and one aspec a io. These eg esso s do no sha e weigh s and ecei e as inpu ea u es o he same spa ial size (n×n). 2.1.4 YOLO You Only Look Once (YOLO) is an objec de ec ion algo i hm ocused on p edic ion speed. While ela i ely accu a e and e y as , i s ill has issues wi h he de ec ion o ce ain objec s, especially small ones [57]. YOLO’s a chi ec u e is s uc u ed as ollows. Fi s , an inpu image is ecei ed and di ided in o an S×Sg id, whe e each cell is esponsible o p edic ing Bbounding boxes. Then each bounding box is p edic ed by he ne wo k wi h a con idence sco e de ined by con idence(p ed) = p(Objec )·IoU u h p ed , 2.2. MULTIPLE OBJECT TRACKING 13 whe e p(Objec )is he p obabili y o any objec being in he ame and IoU u h p ed is he in e sec ion o e union o he p edic ed bounding box and he g ound u h one (see Fig. 2.4). Fo each bounding box, 5 alues a e p edic ed: x, y, w, h and con idence. The i s wo alues de e mine he cen e posi ion o he bounding box and he las wo i s wid h and heigh . The e a e also Cp edic ions o each bounding box (whe e Cis he numbe o classes being conside ed) ha indica e he p oabili y o he bounding box belonging o each pa icula class. This p edic ion is achie ed h ough a CNN and is encoded as an S×S×(5B+C)ou pu enso . P (Objec ) = 0.98 IoU = 0.9 con idence(p ed) = 0.882 Figu e 2.4: YOLO con idence sco e calcula ion. On he igh , p edic ed bounding box ( ed) e sus g ound- u h (blue) o one o he g id squa es. Once he bounding boxes ha e been p edic ed, YOLO makes use o Non-Max Supp ession (NMS), whe e o e lapping bounding boxes a e emo ed i hey p edic he same ype o objec and ha e a high IoU wi h he bounding box ha has he highes con idence sco e. This helps p e en duplica e de ec ions and selec he mos adequa e bounding box. YOLO has unde gone se e al i e a ions and imp o emen s since i was i s p oposed. One o i s mos ecen e sions [9] makes use o a sophis ica ed ea u e ex ac ion p ocess. A p e ained o ine- uned backbone ne wo k (VGG, ResNe , Da kne ...) is connec ed o a neck ne wo k (FPN, Bi-FPN, PANe ) in o de o ex ac hie a chical con olu ional ea u es om he inpu image. These ea u es inally eed a p edic ion ne wo k (RPN, YOLO, SSD, Re inaNe ...) o ou pu he bounding boxes. 2.2 Mul iple Objec acking When we alk abou he Mul iple Objec T acking (MOT) p oblem we mus speci y whe he we a e dealing wi h online o o line acking. In online acking, he sys em only has in o ma ion abou he cu en ame and p e ious ones. In o line acking, he sys em has in o ma ion abou all ames in a eco ding, hus being able o use u u e in o ma ion o adjus p edic ions. Fo he pu poses o ou wo k, we will es ic ou sel es o online MOT. In he ollowing sec ions we will desc ibe some common app oaches o online MOT. Classically, objec acking schemes usually all in one o ou ca ego ies: 14 CHAPTER 2. STATE OF THE ART 1. Fil e ing schemes: he name e e s o he use o a Kalman il e, an algo i hm used o es ima e he s a e o dynamic sys ems while simul aneously keeping ack o he a iance o he es ima e. 2. Mean-shi me hods: a se o me hods elying on he mean-shi algo i hm, a mode-seeking algo i hm used o ind maxima in p obabili y densi y unc ions, o pe o m objec acking. 3. Templa e ma ching: a se o echniques used o ind pa s o an image ha ma ch a gi en empla e. 4. Op ical low es ima ion: a se o echniques o de e mine appa en mo ion ac oss adjacen ames. Op ical low may e e o dense op ical low, when low ec o s a e calcula ed o he en i e image, o spa se op ical low, whe e only he “mos in e es ing” low ec o s a e conside ed. Mo e ecen ly, howe e , se e al al e na i es ha e eme ged wi h a ious deg ees o success in ei he ackling objec acking by hemsel es o enhancing exis ing acke s. Some o he mos ele an a e: 1. T ack-by-de ec ion me hods: hey wo k in wo s eps. Fi s , a de ec ion mod- ule is applied o each ame in o de o loca e objec s. Then, a acking module associa es exis ing objec iden i ies o he new de ec ions. This p ocess esembles he il e ing ap oach men ioned ea lie , and Kalman il e s a e o en used in ack- by-de ec ion sys ems [59][74][7]. 2. Pa icle swa m op imiza ion (PSO): hey use a se o pa icles whose mo e- men s a e adjus ed depending on bo h hei posi ion and hei neighbo s’ posi ion o y o loca e global op ima. In mul i objec acking, his ansla es o de ining adequa e simila i y unc ions ha a e o be minimized o he objec s o be acked. Some p oposed op ions include di iding he pa icle swa m in o se e al “species” ha keep ack o a speci ic objec [76] o applying PSO successi ely o loca e each objec and hen associa e each one o hem wi h p e ious de ec ions [45] as in ack- by-de ec ion sys ems. I is also common ha he simila i y me ic compa es he co a iance ma ix o he image pa ches de e mined by he swa m wi h he image empla es ha ep esen each objec [38]. 3. In eg a ion o con ex in o ma ion: e e s o echniques ha y o exploi he pa icula con ex in which he sys em is going o be used. The e o e, hese echniques a e mainly aimed a enhancing an exis ing acking sys em by inco po- a ing use ul con ex in o ma ion. An example o con ex in o ma ion being used o enhance a acking sys em could be an objec acking sys em ha s o es he appea ance p o iles o acked objec s in o de o be able o eacqui e a ge s i hey a e e e los . 4. Ensemble acking: ensemble acking sys ems combine one o se e al o he abo e echniques (o o he s ha ha e no been men ioned) o achie e a mo e obus sys em. The goal o ensemble sys ems is o d aw on he s eng hs o di e en echniques while simul aneously co e ing up hei weaknesses. Ou o he abo e me hods, ack-by-de ec ion me hods a e he mos common and can be bo h accu a e and as . PSO me hods ha e seen mild success in specialized sys ems bu 2.2. MULTIPLE OBJECT TRACKING 15 a e gene ally ou pe o med by o he echniques. In eg a ion o con ex in o ma ion has become s anda d and is almos always used in some way o ano he . Finally, ensemble acking should be able o p oduce accu a e esul s, bu may be bo lenecked by he pe o mance o each indi idual acke . Fo he pu pose o his sec ion, we a e pa icula ly in e es ed in sys ems ha achie e high p ecision sco es while main aining high ame a es. The e o e, we ha e selec ed 3 well- known acking sys ems ha achie e s a e o he a esul s bo h in e ms o p ecision and in e ence speed: ROLO, SORT, and DeepSORT. 2.2.1 ROLO ROLO s ands o Recu en YOLO and, as i name sugges s, i make use o he YOLO ne wo k and and Long Sho -Te m Memo y (LSTM) cells in o de o pe o m mul i-objec acking. An LSTM [33] is a speci ic Recu en Neu al Ne wo k (RNN) a chi ec u e ha is able o cap u e bo h sho and long dis ance dependencies in sequence da a. Figu e 2.5: ROLO a chi ec u e [53]. ROLO adds a laye o LSTMs on op o he YOLO de ec ion model (see Fig. 2.5). These LSTMs ake as inpu bo h he ea u es ex ac ed by YOLO and he de ec ion in o ma ion. The acking p oblem is hen ea ed as a eg ession p oblem, and he loss unc ion used is he Mean Squa ed E o (MSE). Depending on he s ep size, he numbe o p e ious ames ROLO akes in o accoun o p edic ion, he FPS coun anges om 270 o abou 33 o s ep size alues be ween 1 and 9 espec i ely, and he a e age accu acy on he OTB-30 da ase is as high as 0.45, measu ed in a e age IoU wi h he g ound- u h, o a s ep size o 6. 2.2.2 SORT Simple Online Real ime acking (SORT) is a ack-by-de ec ion sys em ha combines CNN ea u e ex ac ion wi h a Kalman il e and he Hunga ian algo i hm in o de o ack objec s a up o 260 FPS. The CNN a chi ec u e used o ea u e ex ac ion in he o iginal pape is Fas e R- CNN [58] and he Kalman il e , based on a linea cons an eloci y model, is used o 16 CHAPTER 2. STATE OF THE ART es ima e in e - ame displacemen . As o associa ion be ween exis ing iden i ies and new de ec ions, an op imal ma ching is pe o med h ough he Hunga ian algo i hm [44] by minimizing he pai wise IoU me ic. I new objec s en e he scene o exis ing ones lea e i , wo handling s a egies a e applied. Fo objec de ec ions wi h oo low an o e lap wi h exis ing objec s, a new iden i y is gene a ed. Upon disappea ance o an objec , he iden i y will be main ained o TLos i he objec is no de ec ed again. An obse a ion abou SORT is ha , i an objec is no de ec ed ( o example, because i is occluded by ano he one), he linea cons an eloci y model will be he one ha de e mines he de ac o posi ion o he objec in he nex ame. Since his model is a poo p edic o (in gene al), he au ho s a gue ha TLos should be se o 1. Fu he mo e, his has he added bene i o a lowe compu a ional o e head when agen s exi he scene. 2.2.3 Deep SORT Deep SORT is simply an ex ension o SORT ha inco po a es appea ance in o ma ion (in he o m o an image embedding) o acili a e iden i y acking. The main goal o his i e a ion o SORT is o educe he amoun o iden i y swi ches. Fo his pu pose, ap- pea ance in o ma ion is used o e-iden i y objec s ha ha e been empo a ily los . Ano he inno a ion o Deep SORT is he simul aneous use o 2 me ics: he squa ed Mahalanobis dis ance o associa ion be ween Kalman s a es and he coo dina e-wise smalles cosine dis ance in appea ance space. The Mahalanobis dis ance wo ks well when he unce ain y is low (e.g: sho - e m p edic ions) bu , when his unce ain y is in- c eased, he cosine dis ance o e s a be e simila i y indica o by aking in o accoun appea ance in o ma ion. Deep SORT uns a app oxima ely 33 FPS o 32 bounding boxes on an N idia GeFo ce GTX 1050 mobile GPU. 2.3 Public speaking and a ec i e compu ing Public speaking has been ex ensi ely s udied o millenia, bu only in he las decades ha e compu e ized me hods been applied o his s udy. In pa icula , he ield has been e i alized by he echniques o a ec i e compu ing, a e m which e e s o he s udy and de elopmen o sys ems o ecognize and in e p e human emo ions. In his sec ion, we will o e iew wo k ha has been made o quan i y audience expe ience, as well as he exis ing public speaking aining sys ems and da ase s ha we ound. 2.3.1 Public speaking aining sys ems The use o echnological ools o imp o e public speaking skills has ecei ed a lo o a en ion om esea che s. In his sec ion, we’ll b ie ly e iew some o he sys ems ha ha e been p oposed, including some i ual audience sys ems, in which an audience o a ying deg ees o esponsi eness is simula ed and p esen ed o he speake . While mos wo k on i ual audiences has ocused on educing speake anxie y, some ha e ackled pe o mance quali y. 2.3. PUBLIC SPEAKING AND AFFECTIVE COMPUTING 17 Cice o Cice o [4] is an in e ac i e i ual audience sys em, whe e he speake pe o ms in on o a simula ed audience ha also eac s o he pe o mance, p o iding eal- ime eedback. Using h ee senso s (mic ophone, came a, and Mic oso Kinec ), he sys em ex ac s a se o desc ip o s o quali y ha a e hen combined o c ea e an o e all sco e. This sco e con ols he membe s o he i ual audience, which can change pos u e, head o ien a ion and eye gaze o con ey di e en deg ees o in e es in he p esen a ion. Figu e 2.6: Vi ual audience snapsho [4]. Figu e 2.7: Ra ed beha io s and associa ed de- sc ip o s [4]. The g ound da a used o ain he sys em is a se o expe e iews by wo senio membe s o he public speaking o ganiza ion Toas mas e s. The expe s wa ched eco dings o each p esen a ion once and we e asked o a e 21 cha ac e is ics o he pe o mance (some o hem a e shown in Figu e 2.7), as well as o gi e an o e all imp ession. The a ings use 7-poin Like scales. Nex , he au ho s a emp o iden i y au oma ic desc ip o s ha co ela e well wi h each o he a ed beha io s. They use a amewo k called Mul iSense o in eg a e mul imodal da a om hei h ee senso s, and he esul can be seen in Figu e 2.7. Eigh o hese desc ip o s ( i e oice ea u es, wo pos u e ea u es, and one gaze ea u e) a e chosen as inpu s o an SVM which is ained o app oxima e he o e all expe a ing. The sys em was es ed in a pos e io expe imen [16] wi h 51 pa icipan s di ided in h ee g oups: a g oup wi h no eedback (passi e i ual audience), a g oup wi h di ec eedback (passi e i ual audience and a isual indica o o he pe o mance sco e), and a g oup wi h indi ec eedback (in e ac i e audience, no isual indica o ). ROC Speak ROC Speak [23] is a web ool ha allows use s o p ac ice speaking and ge eedback. The use s a e eco ded wi h a webcam and mic ophone and hen a e gi en an au oma ic sco e based on hei non e bal beha io . They also ha e an op ion o ge a c owdsou ced human a ing. To p oduce he au oma ic sco e, he ollowing ea u es a e ex ac ed om he eco d- ing: •Smile in ensi y, cap u ed om a s anda d acial ea u e de ec ion lib a y. •Mo emen , de ined as no malized pixel di e ences. 24 CHAPTER 2. STATE OF THE ART 2.5.1 Google Mee Google Mee is ee, widely a ailable, and p o ides a eco ding unc ionali y. Videos can be s eamed by sc een sha ing. By de aul , i only eco ds he cu en speake and displayed con en a any gi en momen . Howe e , by using he Google Mee G id View plugin [24], i ’s possible o eco d e e yone in mosaic iew. Thus, his app ul ills all o ou essen ial equi emen s bu none o he addi ional ones. 2.5.2 Bb Collabo a e Blackboa d Collabo a e is commonly used o online classes in highe educa ion and while i ’s no ee, ou ins i u ion p o ides us access o i . I can play ideos using sc een sha ing, and i also allows eco ding sessions. Howe e , hese eco ding only include ac i e speake ideo o displayed con en . The e o e, his app is unsui able o us. 2.5.3 Zoom Ano he widely known ideocon e encing applica ion, Zoom is simila o Google Mee in ha i allows playing ideos and eco ding o e e yone in he session, bu no eco ding pa icipan s indi idually no p e en ing hem om seeing each o he , he e o e i ’s a iable candida e bu does no sa is y ou ex a equi emen s. 2.5.4 Ji si A somewha less known al e na i e, Ji si is ee, open sou ce, and o e s ideo s eaming and mosaic eco ding unc ionali y. I is a iable candida e bu does no sa is y ou ex a equi emen s. Howe e , i migh be possible o c ea e a plugin o modi y he sou ce code so ha i does. Chap e 3 The ERVF Da ase In o de o ain he sys em, a undamen al equi emen is access o adequa e aining da a. Speci ically, a da ase o audience eco dings du ing p esen a ions o public speak- ing e en s, labeled wi h he emo ional s a e and le el o engagemen o each indi idual. While we ound some da a se s ha i hese wo condi ions (see sec ion 2.3.3), hey p e- sen ed wo main p oblems. Fi s , hey we e no a ailable o he gene al public. Second, hey we e eco ded in pe son and so we e a poo i o online en i onmen s. As a consequence, i became necessa y o ga he ou own aining da a. This en ailed p ecisely de ining he a iables we wan ed o measu e, designing an expe imen al se up o ob ain hose measu emen s, and p ocessing and cu a ing he esul s in o a usable o m. In his chap e we will explain his p ocess and he esul s we ob ained. The i s s ep was de ining a quan i a i e measu e o audience expe ience, which is in o- duced in Sec ion 3.1. A e wa ds, we pe o med wo expe imen s, desc ibed in Sec ion 3.2, and hen we ex ac ed he esul s, explained in Sec ion 3.3. Finally, we will cla i y some limi a ions o ou p ocess. 3.1 Quan i ying audience expe ience The s udy o communica ion and public speaking is a huge and e y ac i e academic ield. As comple e ou side s o his ield, i was ha d o us o ind ele an li e a u e on he subjec o audience expe ience. In addi ion, he numbe o publica ions on audience expe ience is dwa ed by hose on o he aspec s o public speaking and pe o mances, such as speake anxie y. Despi e hese issues, we ound some p e ious wo k on he subjec , which is documen ed in Sec ion 2.3.2. These wo ks concep ualize audience expe ience using a combina ion o dimensions, d awing inspi a ion om psychological heo y and quali a i e in e iews, and also p o ide expe imen al alida ion o some o hese me ics. In pa icula , in Cap u ing he audience expe ience: A handbook o he hea e [13] a i e-dimension amewo k is p oposed. These i e dimensions a e: 1. Engagemen and concen a ion: The ex en o which he pe o mance cap u es and main ains he audience’s a en ion. 25 26 CHAPTER 3. THE ERVF DATASET 2. Lea ning and challenge: The challenge can be on knowledge, expec a ion, o a i udes. 3. Ene gy and ension: Physiological eac ions o he pe o mance, such as exci e- men o anxie y. 4. Sha ed expe ience and a mosphe e: The sense o collec i e expe ience a o ded by a pe o mance. 5. Pe sonal esonance and emo ional connec ion: The ex en o which membe o he audience can eel empa hy o iden i y hemsel es in he pe o mance. We ini ially conside ed applying his amewo k di ec ly. Howe e , dimension 4 does no play a big ole in he online se ing, whe e membe s o he audience a e ypically isola ed om each o he . The e o e, we decided o d op i and ocus on he o he ou dimensions. The nex s ep a e ha ing speci ied a concep ual amewo k was p o iding a conc e e me ic o each o he ou componen s, a p ocess usually called ope a ionaliza ion. Con- enien ly, he au ho s o he epo also p o ide a se o guidelines and example ques ions o measu e hese componen s om sel - epo ques ionnai es. To complemen his guidelines, we e iewed he da a-ga he ing p ocess used o c ea e he audience expe ience da ase s ound ea lie , especially he wo k o Cu is e al. [17]. F om his we inco po a ed hei wo k in engagemen , as well as hei comp ehension me ic, combining i wi h dimension 2 o he concep ual amewo k. In addi ion, we make use o he abundan li e a u e on a ec i e esponse [10] and use i as an ope a ionaliza ion o ene gy and ension. Thus, ou inal amewo k is composed o he ollowing ou dimensions: a ec i e esponse (A ), engagemen (En), emo ional connec ion (Ec), and lea ning (Le). Wi h espec o ac ually measu ing hese componen s in an expe imen , he ollowing op ions we e a ailable: 1. Sel - epo s: he audience ills in a ques ionnai e a e wa ching he pe o mance, wi h ques ions abou hei subjec i e expe ience. The main ad an ages o his me hod a e ha i ’s as , cheap, ela i ely unbiased, and does no dis u b he expe ience. The main disad an age is ha i only p o ides one da a poin o each indi idual and pe o mance, and he e o e has e y low empo al esolu ion. This app oach was he one ollowed in Cap u ing he audience expe ience: A handbook o he hea e [13]. 2. Ex e nal a ing: an ex e nal obse e e alua es each componen o e he du a- ion o he pe o mance. The main ad an ages a e ha i p o ides much g ea e empo al esolu ion, since we can anno a e, o example, each 1-minu e in e al wi h a di e en sco e. The main disad an age is ha i equi es ime-consuming manual labo and is p one o bias on he pa o he anno a o . This app oach was ollowed by Cu is e al. [17]. 3. Physiological measu emen s: Physiological senso s a e connec ed o he pa ici- pan s, which measu e gal anic skin esponse, hea bea , and o he ele an signals. This is p obably he mos di ec way o measu ing he physical and emo ional s a e 3.2. EXPERIMENTAL DESIGN 27 o he audience, bu i equi es specialized machine y and equipmen ha was no a ailable o us. Fu he mo e, i is in usi e and may ha e a di ec in luence on he audience expe ience. 3.2 Expe imen al design Fo ou expe imen al design we i e a ed h ough se e al p oposals, spo ing and co ec - ing he p oblems we ound un il we con e ged o he inal design. The basic idea always emained he same: eco d an audience as hey eac o a pe o mance, ei he li e o ideo aped. In addi ion, egis e he audience expe ience using ou amewo k and ei he sel - epo o ex e nal anno a ions. In he i s i e a ion we conside ed wo di e en se ups: an in-pe son one, on campus, wi h a igh ly con olled en i onmen o emo e any con ounding a iables and alida e ou me hodology; and a c owdsou ced e sion, using he olun ee s’ webcams, o ga he mo e da a in a scalable way. Howe e , we soon ealized ha he public heal h si ua ion was no a o able o any in-pe son expe imen s and hus decided o pe o m ou expe imen en i ely online. Fo he second i e a ion, since we knew he expe imen would be online, we de e mined ha he pe o mances would need o ha e he o m o ideo eco dings ha we could s eam o e he In e ne . Thus, we s a ed selec ing a pool o ideos o he expe imen , de ailed below. As o da a collec ion, we conside ed a wo-p onged app oach. We would use ex e nal anno a ions o engagemen , in combina ion wi h a sel - epo ques ionnai e o all o he componen s. This way, we would ha e high- esolu ion da a o one o he componen s, and we would be able o pe o m c oss-checks wi h he sel - epo da a o ensu e he soundness o ou me hod. A his s age, we de eloped he ques ionnai e, which will be explained below. Due o ime cons ain s, we de e mined i would be in easible o us o anno a e he da ase . So, o he hi d i e a ion, we we e o ced o se le on sel - epo s o assess audience expe ience. 3.2.1 Tools used Ha ing decided ha he se ing would be online, he pe o mances would be eco ded, and he da a collec ion me hod would be sel - epo , we needed h ee ex a pieces o pe o m he expe imen : a ideocon e encing ool, a selec ion o ideo pe o mances, and a sel - epo ques ionnai e. Video selec ion To make ou expe imen lexible in e ms o ime commi men , we decided o use sho ideos, o abou 10 minu es. The ideos had o be a ied in con en and s yle, and in Spanish language, since ou expe imen al audience would likely be na i e Spanish speake s. E en ually, we chose en ideos, which can be ound in Table 3.1. These emo ional, poli ical, comedic and di ulga i e ideos. Sel - epo ques ionnai e The inal ques ionnai e included he ollowing ques ions, di ided in o sec ions o each componen o he audience expe ience: 28 CHAPTER 3. THE ERVF DATASET Ti le Speake Link “La soledad del adic o: Comp ende al o o puede sal a le la ida” Lau a Ve ga a h ps://www.you ube.com/ wa ch? =z6FoUeSohnk “Lo imposible a eces sólo cues a un poco más” Edua do Llano h ps://www.you ube.com/ wa ch? =9K ZsJuEN 0 Pablo Casado’s esponse o San iago Abascal a he 2020 mo ion o censu e Pablo Casado h ps://www.you ube.com/ wa ch? =9Ehh 94YDG09 In e en ion o Gab iel Ru ián a Ma i- ano Rajoy’s in es i u e deba e Gab iel Ru ián h ps://www. e. es/alaca a/ ideos/ especiales-in o ma i os/ la1- u ian-020916/ 3709343/ “Tengo un sueño” Dani Ro i a h ps://www.you ube.com/ wa ch? =8JWsg4Psm 0 “¿Po qué ‘ unne ’?” Ana Mo gade h ps://www.you ube.com/ wa ch? =N-NdyyHc_Gk “La selección la inoame icana de ce e- b os” Juan En íquez h ps://www.you ube.com/ wa ch? =GglVs9scY6I “A nadie le impo a la e dad” Rocío Vidal h ps://www.you ube.com/ wa ch? =b_I6Wma S2o “¿Po qué me igilan, si no soy nadie?’ Ma a Pei ano h ps://www.you ube.com/ wa ch? =NPE7i8wuupk Spain’s 2015 Royal Ch is mas message Felipe VI h ps://www.you ube.com/ wa ch? =P Vm83 2 mk Table 3.1: Video selec ion. A ec i e Response This componen is measu ed using he Sel Assessmen Manikin [10]. Fo each o he h ee images in Figu e 3.1, he pa icipan s selec which cell hey eel mos iden i ied wi h. The images ep esen he h ee ac o s o a ec i e esponse ound in he li e a u e. Those a e, in o de , alence (how good you eel), a ousal (how in ense you eelings a e), and dominance ( o wha ex en you eel in con ol). The images a e labeled om 1 o 9, he e o e being analogous o a 9-poin Like scale, and he inal sco e is a no malized sum o he h ee answe s. Conc e ely, we add he h ee answe s, sub ac hei a e age, and escale so ha he end esul is wi hin he in e al [0,1]. Engagemen This sec ion consis s o ou 5-poin Like scale ques ions, wi h wo la- bels pe ques ion, shown in Table 3.2. The pa icipan s a e asked o selec hei ag eemen le el wi h hese wo labels. The o al engagemen sco e is gi en by he no malized sum o all he answe s. 1-poin label 5-poin label My mind wande ed I was comple ely ocused on wha I saw Time seemed o pass e y slowly I ha dly no iced he passage o ime The ideo didn’ ca ch my a en ion I couldn’ keep my eyes o he sc een I don’ wan o alk abou he ideo I wan o sha e my expe ience wa ching he ideo Table 3.2: Engagemen sec ion ques ions. Emo ional Connec ion This sec ion consis s o h ee 5-poin Like scale ques ions, wi h wo labels pe ques ion, shown in Table 3.3. The pa icipan s a e asked o selec 3.2. EXPERIMENTAL DESIGN 29 Figu e 3.1: Sel Assessmen Manikin. hei ag eemen le el wi h hese wo labels. As be o e, he sco e o his sec ion is he no malized sum o all he answe s. 1-poin label 5-poin label I was no mo ed by he ideo The ideo a ec ed me emo ionally I wouldn’ wa ch mo e con en om he same speake I would like o wa ch mo e con en om he same speake The ideo did no say much abou my pe - sonal expe iences I el iden i ied a a pe sonal le el wi h some pa s o he ideo Table 3.3: Emo ional connec ion sec ion ques ions. Comp ehension and Lea ning This sec ion consis s o six 5-poin Like scale ques- ions, wi h wo labels pe ques ion, shown in Table 3.4. The pa icipan s a e asked o selec hei ag eemen le el wi h hese wo labels. In he i s h ee i ems, he o e a ching ques ion is “How unde s andable would you say i ’s been?”, while o he las h ee i ems he o e a ching ques ion is “How do you eel wi h espec o he heme o he ideo?”. Once again, he sco e o his sec ion is he no malized sum o all he answe s. 1-poin label 5-poin label I didn’ unde s and any hing E e y hing was pe ec ly clea I was ha d o ollow A all imes I knew wha he speake was alking abou I wouldn’ be able o explain he con en o someone else I could explain he con en o someone else wi hou p oblems The e was no hing new o me I opened my mind o new ideas o iew- poin s I knew he opic in-dep h be o e wa ching he ideo I did no know any hing abou he opic be- o e wa ching he ideo My ideas abou he opic ha en’ changed Now I ha e a comple ely di e en iew Table 3.4: Comp ehension and lea ning sec ion ques ions. The o m was c ea ed and adminis e ed using Google Fo ms. 30 CHAPTER 3. THE ERVF DATASET Video con e encing ool To pe o m he expe imen , we would need a ool sa is ying a leas he ollowing equi e- men s: • Abili y o play ideos h ough he app • Abili y o eco d he pa icipan s while doing so • Abili y o ex ac om he eco ding an exclusi e ideo s eam o each indi idual pa icipan . This en ails ei he eco ding each pa icipan sepa a ely h ough hei webcam, o eco ding e e yone in a mosaic iew and hen manually demul iplexing he indi idual ideos. The i s is a mo e e icien and na u al app oach, bu is no always a ailable. In addi ion, we ound a ac i e he idea o p e en ing pa icipan s om seeing each o he du ing he expe imen , which would p e en hem om biasing each o he . While we could achie e his making mul iple expe imen s wi h a single indi idual each, his app oach is e y slow. A e e iewing he mos commonly used ideo con e encing apps, we ound some ha we e su icien , bu none o hem we e ideal. Mo e conc e ely, we ound applica ions sa - is ying he manda o y equi emen s, bu none o hem allowed eco ding each pa icipan sepa a ely, which o ced us o use he manual demul iplexing me hod. In addi ion, none o hem allowed us o p e en pa icipan s om seeing each o he . So, a e some delibe a ion, we decided o c ea e ou own ideo con e encing ool, which would allow us o p ecisely con ol he expe imen al se ing. This ool is able o eco d indi idual pa icipan s, and p e en ing hem om seeing and hea ing each o he , in addi ion o ul illing he o he equi emen s. The applica ion p o ides wo di e en use in e aces: one o he subjec s o he expe i- men and one o he hos . As can be seen in Figu e 3.2, he hos can see all he subjec s in a mosaic iew o he igh , and a lis o s eams o he le . The hos can c ea e, cas , and emo e s eams om he lis , and he s eams can be cap u ed om a webcam, sc een sha ing, o a YouTube ideo. When he hos chooses o cas an s eam, i will appea in he subjec iew. In he case o a YouTube ideo, he hos can con ol he playback s a e using he no mal con ols o he embedded playe . When he s a e changes, he change will be b oadcas o all he subjec s, in such a way ha i he hos s a s, s ops, o changes he ideo imes amp, he subjec s will see he same changes. In he in e io pane he e is a bu on o add a new s eam, a bu on o s a o s op eco ding, and he common cha . The subjec s only iew hei own image, as well as any hing he hos chooses o cas . They can con ol hei came a and mic ophone, and w i e h ough a common cha . When he hos cas s a YouTube ideo, he subjec s can’ con ol he playback o he ideo in any way, bu ha e he op ion o ep oduce i in ull sc een. When he hos s a s eco ding a sepa a e ile is c ea ed o each subjec . In addi ion, wo mo e iles a e c ea ed: one is an iden i ica ion ile which links he iden i ie s o he subjec s o hei espec i e ideo iles, and he o he one is a synch oniza ion ile. This ile s o es he imes amps o all eco dings when he playback s a us o a YouTube ideo is modi ied, as well as he ideo imes amp and ype o e en ha igge ed he synch oniza ion. This 3.2. EXPERIMENTAL DESIGN 31 Figu e 3.2: A sc eensho o he hos iew aken du ing he pilo expe imen . allows he expe imen e s o iden i y which pa s o he eco dings co espond o a gi en ideo segmen , e en i he e is some delay in he eco dings due o ne wo k la ency. The ool was c ea ed using he WebRTC API [29], which o e s he wo essen ial unc- ionali ies needed o ideocon e encing: cap u e and con ol o audio isual s eams and signaling o eal ime communica ion. We used he Pion WebRTC s ack [54] o he backend, adding sligh modi ica ions o adap i o ou needs. The web in e ace was c ea ed using he Pion SDK and anilla Ja aSc ip 1. 3.2.2 Me hodology Ha ing in oduced ou design p ocess and ou ools, we can now explain ou comple e me hodology, which is in oduced diag amma ically in Figu e 3.3. We an wo expe i- men s, he i s o which was a pilo es . Pa icipan s we e olun ee s, ei he s uden s om ou acul y o acquain ances o he expe imen e s. A e a sui able pool o pa ici- pan s was ound, we emailed hem asking o olun a y pa icipa ion in he expe imen , along wi h a su ey o ind a ailable da es. The pilo expe imen was scheduled o ha e a du a ion o wo hou s bu , a e ge ing eedback om he i s one, we educed he du a ion o he ac ual expe imen o one hou . Figu e 3.3: Diag am o ou expe imen al me hodology. Once we go enough eplies and ound an adequa e da e, we asked pa icipan s o sign a consen o m o ake pa in he expe imen . We assigned nume ic iden i ie s o each pa icipan and decided which ideos om ou selec ion we would play du ing he session. 1All he code is a ailable in h ps://gi hub.com/pablo- s/ c, and he Ja aSc ip RTC lib a y which o ms he bulk o he code we de eloped can be seen in Appendix A. 32 CHAPTER 3. THE ERVF DATASET This was done andomly o educe any biases om he o de ing o he ideos. A ew hou s be o e he expe imen , we emailed pa icipan s hei iden i ie s and he URL hey would use o connec o he ideo con e encing ool. The ac ual expe imen s we e conduc ed as ollows: a e a b ie p esen a ion and sol ing any echnical issues, we explained how he expe imen would wo k o he pa icipan s. In he pilo we did no in oduce he ques ionnai e a he beginning o he expe imen bu , a e no icing ha he subjec s ound pa o he ques ionnai e con using, in he ac ual expe imen we decided o le pa icipan s openly explo e he ques ionnai e a he beginning and answe ed any ques ions hey had. Then we played each one o he selec ed ideos h ough he app, emo ing any o he con en om he sc een so ha he pa icipan s could only see hemsel es and he ideo. A e he ideo inished, pa icipan s illed in he o m, and hen we mo ed on o he nex ideo. The o m used in he expe imen s con ained h ee sec ions: one o he beginning o he ideo, ano he o he middle, and he las one o he ending. Each sec ion con ained he ull se o ques ions ou lined abo e. Tha way, we would ha e h ee da a poin s pe subjec and ideo, ins ead o one. The esponses o he ques ionnai e we e associa ed wi h he eco dings using he unique iden i ie o each subjec . Du ing he pilo expe imen , we an in o a echnical issue wi h he ideocon e encing ool ha p e en ed us om eco ding any hing. Conc e ely, he se e we we e using o hos he applica ion hi a esou ce limi a ion o which we we e no awa e, and d ama ically educed i s pe o mance. Thus, we decided o all back o Google Mee and con inue he expe imen he e. Un o una ely, due o a mis ake on ou pa , he eco ding om Google Mee was unusable. Howe e , a e expe iencing his issues we we e able o sol e he unde lying cause and use ou ideo con e encing ool success ully in he nex expe imen . The pilo expe imen was conduc ed wi h 10 people and 5 ideos, while he nex one was conduc ed wi h 8 people and 3 ideos. 3.3 Expe imen al esul s As men ioned in he p e ious sec ion, he pilo expe imen did no p o ide any iable esul s. Howe e , he second expe imen was mo e success ul. A e a p elimina y explo- a ion o he da a, i seems ha ou expe imen was able o success ully cap u e a ia ion in expe ience ac oss subjec s, ideos, and pa s. In addi ion, he ou componen s a e no comple ely co ela ed wi h each o he . In Figu e 3.4, we can see his a ia ion ac oss ideos, as well as some common pa e ns: comp ehension ends o be highe han he o he componen s, while emo ional connec ion ends o be lowe . 3.3. EXPERIMENTAL RESULTS 33 Figu e 3.4: A e age o he dimensional sco es o all pa icipan s in h ee di e en ideos. The da ase To c ea e he ac ual da ase , we synch onized he eco dings wi h he ideos and cu hem o ma ch he du a ion o he pe o mances. Then, he eco dings we e p ocessed by ou acking module, p oducing a s eam o cons an esolu ion acking each pa icipan . Du ing his s ep, we had o disca d he h ee eco dings o one o he pa icipan s because o an un o eseen p oblem: in he backg ound he e we e a se ies o anime pos e s, which 40 CHAPTER 4. THE VISUAL TRACKING MODULE a bounding box wi h highe p edic ion sco e Msuch ha IoU − RDIoU (M,Bi)≥ε o a ce ain h eshold ε. The e m RDIoU (M,Bi)is gi en by RDIoU (M,Bi) = ρ2(pM,pBi) c2(4.4) which ma ches he hi d e m in equa ion (4.1). 4.1.2 The acking algo i hm Gi en he con ex desc ibed a he beginning o he chap e , i is clea ha he acking algo i hm does no need o be excessi ely complica ed. Fo his, we designed a cus om acking algo i hm o iden i y agen s ac oss ames and o keep ack o hem. The gene al idea is o ma ch exis ing iden i ies wi h iden i ied objec s in a “mos likely” ashion. Fo his pu pose, we keep ack o he mo emen o objec s in adjacen ames and use i o p edic an “expec ed nex posi ion” o each o he iden i ied agen s. Each iden i ied agen s is hen ma ched wi h he iden i y whose “expec ed nex posi ion” is closes o he de ec ed one. This app oach is illus a ed in Figu e 4.5. The main ca ea s o his app oach a e how o deal wi h missiden i ica ions, agen s mo ing o -sc een, o new agen s mo ing in om ou side he scene. T acking algo i hm De ec ed objec s Iden i ied objec s Expec ed iden i ica ions Figu e 4.5: Gene al s uc u e o he acking algo i hm. Iden i ying de ec ed objec s To speci y how he algo i hm wo ks exac ly, le O1, ..., On∈R4be he bounding boxes o nde ec ed objec s a ime such ha Oi= (pi, wi, hi)∈R2×R×Rwhe e pi speci ies he coo dina es o he objec cen e and wi, hispeci y, espec i ely, he wid h and heigh o he bounding box. Le s also conside a se o mbounding boxes o he expec ed iden i ica ions P1, ..., Pm∈R4a ime in he same ashion. Each one o he iden i ica ions is canonically associa ed wi h some i∈ {1, ..., m}and we say ha Pi speci ies he expec ed bounding box o agen i. 4.1. MODULE ARCHITECTURE 41 The dissimila i y be ween a de ec ed objec wi h bounding box Oiand an expec ed iden i ica ion wi h bounding box Pjis measu ed as m(Oi,Pj) = ||cOi−cPj|| + log(1 + |hOi−hPj|) + log(1 + |wOi−wPj|)(4.5) whe e he Oiand Pjsupe sc ip s a e used o disambigua e. The unc ion abo e de e - mines a me ic in R4. I is easily obse ed ha m(Oi,Pj)≥0and ha m(Oi,Pj) = 0 i and only i Oi=Pj. Symme y is ob ious and he iangula inequali y can be deduced om he ac ha each o he e ms e i ies i . I is also in e es ing o no e ha he i s e m will gene ally be much mo e ele an han he o he wo, which a e only impo - an (in p ac ice) when he cen e o he de ec ed objec and he cen e o he expec ed de ec ion a e e y close. Ou algo i hm ma ches each de ec ed objec , ep esen ed by i s bounding box, Oi o an expec ed de ec ion, ep esen ed by Pj, by using Gale-Shapley’s (GS) algo i hm. Fo his pu pose, p e e ence lis s a e c ea ed o each de ec ed objec and o each expec ed de ec ion. A de ec ed objec Oiwill p e e being ma ched wi h Pjo e Pki m(Oi,Pj)< m(Oi,Pk). The p e e ence lis is buil analogously o each expec ed de ec ion Pi. We decided o use GS ins ead o he mo e common Hunga ian algo i hm o 3 easons. Fi s , he GS algo i hm e u ns ma chings ha a e close o he op imal solu ion ound ia he Hunga ian algo i hm [46]. Second, we eel like he objec i e o be op imized is mo e na u al in he con ex o ma ching de ec ions and iden i ies. In GS, we aim o ma ch iden i ies and de ec ions in pai s such ha any de ec ion Oi( espec i ely, iden i y Ei) o a gi en pai can’ be ma ched o an iden i y Ej(de ec ion Oj) o ano he one such ha iden i y Ej(de ec ion Oj) is a be e ma ch o de ec ion Oi(iden i y Ei) han i s cu en iden i y (de ec ion) and ice e sa. On he o he hand, he Hunga ian algo i hm aims o minimize he sum o he dis ances o each pai . Finally, he GS algo i hm is mo e e icien han he Hunga ian algo i hm, unning in O(nm)whe e nis he numbe o de ec ions and mis he numbe o iden i ies. Calcula ing expec ed iden i ica ions Expec ed iden i ica ions ep esen he posi ion and bounding box a which we expec an objec o appea in he nex ame. They co espond o objec s ha ha e al eady been iden i ied and should he e o e appea in he ollowing ames. The posi ion is a p edic ion made based on he his o y o posi ions and bounding boxes o ha speci ic iden i y (see Fig. 4.6). Suppose ha , o agen j, we ha e a his o y o de ec ions D1, ..., D ∈R4 o ames 1, ..., ∈N, in ch onological o de . Then, a ame , we de ine he a ia ion in he posi ion o agen jas ∆D=D −D −1= (∆p,∆h, ∆w)whe e ∆p,∆hand ∆wa e, espec i ely, he a ia ions in cen e posi ion, heigh and wid h o he bounding box. Wi h his in o ma ion, he expec ed iden i ica ion Pj o agen ja ame + 1 will be Pj( + 1) = D + ∆D. I ≤1 hen ∆D= 0. We obse e ha he p ocess abo e is simila o he idea o a Kalman il e ollowing a cons an eloci y model. The main di e ence is ha his p ocess is no p obabilis ic and does no ake unce ain y o e o s in measu emen s in o accoun . 42 CHAPTER 4. THE VISUAL TRACKING MODULE Figu e 4.6: The expec ed posi ion in e ence p ocess illus a ed. In ed, he de ec ions o he cu en ame. In blue, he de ec ions in he las ame. As a do ed line, he expec ed bounding box and in a con inuous line he bounding box o cu en de ec ions. Handling misiden i ica ions, misde ec ions, and o he p oblems in p ac ice Misde ec ions consis ei he o ins ances whe e he sys em de ec s objec s ha a e no eally in he ame o ins ances whe e i ails o iden i y objec s ha a e ac ually in i (see Fig. 4.7). Misiden i ica ions, on he o he hand, ep esen ailu e o associa e an objec wi h i s ac ual iden i y (see Fig. 4.8). Misde ec ions and misiden i ica ions can po en ially ha e a la ge impac on he pe o mance o he sys em as a whole. Fo example, i he emo ion ecogni ion module akes in o accoun he whole his o y o ames o a speci ic agen , hen a misiden i ica ion could make subsequen p edic ions unusable. Figu e 4.7: Two examples o misde ec ions. Figu e 4.8: A misiden i ica ion in 2successi e ames. The sys em ails o p ope ly ecognize agen 2in he second ame. In p ac ice, whe e g ound u h is no a ailable, i is i ually impossible o de e mine whe he we a e acing a misiden i ica ion o a misde ec ion. Howe e , we can s ill mi iga e 4.1. MODULE ARCHITECTURE 43 he e ec o hese e o s and make he sys em mo e obus wi h espec o hem. I he numbe o objec s de ec ed nand he amoun o expec ed objec s mdo no ma ch, hen we a e in one o he ollowing scena ios 1. The e has been a leas a misiden i ica ion o misde ec ion. 2. An agen has walked ou o he ield o iew o he came a. 3. A new agen has walked in o he ield o iew o he came a. Since we do no ha e a means o disce n which one o he scena ios abo e bes desc ibes he si ua ions ha may appea in eal- ime, we need o p opose a uni o m handling s a egy. Fo ha eason, i he e is a misma ch in he numbe o objec s de ec ed and he expec ed amoun o objec s ( ha is, n=m), we p oceed as ollows 1. I n > m (mo e de ec ed agen s han expec ed de ec ions) we assign he de ec ed objec s o he mos sui able iden i y (as explained in he p e ious sec ion) un il he e a e no unassigned iden i ies and we c ea e new iden i ies o he emaining n−magen s. 2. I n < m (less de ec ed agen s han expec ed de ec ions) we i s assign each o he de ec ed objec s o he mos sui able iden i y. Then, we use a ole ance ime pe iden i y τj,j= 1, ..., m, o decide whe he o dele e each o he excess iden i ies in o de o p o ide obus ness o misde ec ion. I τj= 0 and iden i y jis unused hen we dele e i . O he wise we dec ease τj. Algo i hm 3: Objec acking algo i hm Da a: Objec de ec ions O1, .., Onand expec ed de ec ions P1, ..., Pm Resul : A s able ma ching pai ing each de ec ed agen o an iden i y gi en by M={(Oi, ji)|i= 1, ..., n, j ∈ {1, ..., max (n, m)}, jl=jmi l=m} Build p e e ence lis s o each de ec ed objec Build p e e ence lis s o each expec ed de ec ion Ma ch each agen o an iden i y ha op imizes he me ic m(·,·)using he GS algo i hm i n > m hen C ea e new iden i ies o ex a objec s. end i m > n hen o each unused iden i y jdo i τj= 0 hen Dele e iden i y j else τj=τj−1 end end end Re u n ma ching M The ime and space complexi ies o Gale-Shapley a e O(nm), which is also he complex- i y o he ull acking algo i hm. 44 CHAPTER 4. THE VISUAL TRACKING MODULE 4.2 Implemen a ion In o de o minimize he amoun o ime necessa y o implemen his pa o ou sys em, we decided o use a eadily a ailable implemen a ion o YOLO 4, p esen in he da kne lib a y. We used Py hon 3.7.3 o build he acking algo i hm p oposed in sec ion 4.1.2 and o implemen Gale-Shapley. We also make use o numpy o e icien a ay ope a ions and openc o in e ac wi h webcams and o wo k wi h images and ideo. Ou implemen a ion o Gale-Shapley is comple ely gene al and wo ks o a bi a y ca- paci ies o indi idual p oposan s and p oposees. de p opose(p oposan , p opose, cap_p oposes, p e _p oposes, m_p oposan s, m_p oposes):,→ """ Re u ns T ue i he p oposal is success ul and alse o he wise. :p oposan in :p opose in :cap_p oposes lis (in ) :p e _p oposes lis (in ) :m_p oposan s lis (se ()) :m_p oposes lis (se ()) """ success =False # I he e is place we simply accep i (cap_p oposes[p opose] >len(m_p oposes[p opose])): m_p oposes[p opose].add(p oposan ) success =T ue else:# I he e is no place he e a e wo op ions i=len(p e _p oposes)-1 #las elemen in p io i y con =T ue while con : i p e _p oposes[i] in m_p oposes[p opose]: # 1. Candida e be e han wo s ma ch -> Accep idx_wo s =p e _p oposes[i] m_p oposan s[idx_wo s ]. emo e(p opose) m_p oposes[p opose]. emo e(idx_wo s ) m_p oposes[p opose].add(p oposan ) success =T ue con =False eli p e _p oposes[i] == p oposan : # 2. Candida e wo se han wo s ma ch -> Rejec con =False else: i-= 1 e u n success 4.2. IMPLEMENTATION 45 de gale_shapley(n_p oposan s, n_p oposes, cap_p oposan s, cap_p oposes, p e _p oposan s,,→ p e _p oposes): """ Recei es a map o s ing o in ha indica es, o each p oposan o p oposee,→ he index a i s p e e ence ma ix. Assumes alues a e indexed in he p e e ence,→ ma ices. :p oposan s in :p oposes in :cap_p oposan s lis (in ) :cap_p oposes lis (in ) :p e _p oposan s lis (lis (in )) :p e _p oposes lsi (lis (in )). """ m_p oposan s =[se () o xin ange(n_p oposan s)] m_p oposes =[se () o xin ange(n_p oposes)] con ,cu en =(T ue,0) cu _p op =[0 o xin ange(n_p oposan s)] i (n_p oposan s >0and n_p oposes > 0): while con : while (len(m_p oposan s[cu en ]) == cap_p oposan s[cu en ]):,→ cu en =(cu en +1)%n_p oposan s # E e yone mus p opose once i, p oposal =cu _p op[cu en ], T ue # P oposes un il accep ed,→ while (p oposal and i<len(p e _p oposan s[cu en ])): p oposee_idx =p e _p oposan s[cu en ][i] i (p opose(cu en , p oposee_idx, cap_p oposes, p e _p oposes[p oposee_idx],,→ m_p oposan s, m_p oposes)): m_p oposan s[cu en ].add(p oposee_idx) # Upda e p oposan ,→ i (len(m_p oposan s[cu en ]) == cap_p oposan s[cu en ]):,→ p oposal =False # We end he p oposal p ocess i=i+1 # A leas a p oposan is ee con = educe(lambda x,y: (x o len(y[0]) <y[1]), zip(m_p oposan s,cap_p oposan s), False),→ # A leas a p oposee is ee 46 CHAPTER 4. THE VISUAL TRACKING MODULE con =con and educe(lambda x,y: (x o len(y[0]) <y[1]), zip(m_p oposes, cap_p oposes), False),→ e u n (m_p oposan s, m_p oposes) To c ea e he acking unc ion, we also need an implemen a ion o he me ic. This is done by using numpy’s buil -in unc ions and ia slicing, as he bounding boxes a e ep esen ed by 4-dimensional a ays (x, y, w, h). de me ic(B1, B2): """ Gi en wo bounding boxes, e u ns he alue o he me ic. """ e u n (np.linalg.no m(B1[0:2]-B2[0:2]) + np.log(1+abs(B1[2]-B2[2])) + np.log(1+abs(B1[3]-B2[3]))) Finally, he acking unc ion i sel ans o ms he inpu in o da a s uc u es sui able o apply Gale-Shapley and calls he ma ching algo i hm. Ma ches a e ans o med om lis (se ()) o lis (in )and he ole ance and iden i y upda es explained in algo i hm 3 a e applied. The pa ame e Aindica es he las assigned agen numbe o gua an ee uniqueness while he ol pa ame e is used o adjus he maximum ole ance alue. de nai e_ ack(assig, expec ed, cen e s_ , ole ance, ol=2, A=0): """ T acks he gi en agen s: associa es each poin o cen e s_ wi h an agen iden i y gi en he p e ious assignmen , he expec ed de ec ions, and he cu en ole ance alue o each agen . The e u n alues a e a map om iden i y o poin , which s o es he las known posi ion o each agen ; and a lis o s ings, whe e he s ing in index i is he iden i y o he i h poin in cen e s_ . :assig dic (s ->poin ) :expec ed dic (s ->poin ) :cen e s_ lis (poin ) : ole ance dic (s ->in ) : ol in :A in : e u n (dic (s ->poin ),lis (s )) """ # Hea ily a ec ed by g anula i y in ime disc e iza ion n,m =len(expec ed), len(cen e s_ ) p e ious_cen e s =lis (expec ed.i ems()) #[(agen , al) o agen , al in expec ed.i ems()],→ # C ea e "dis ance" ma ix (len(expec ed) x len(cen e s_ )) d_ma ix =np.a ay([[me ic(bb_p e ,bb_de ) o bb_de in cen e s_ ] o _,bb_p e in expec ed.i ems()]),→ p e s_p e =np.a ay([so ed(lis ( ange(m)), key=lambda j:d_ma ix[i][j]) o iin ange(n)]) # so by dis ance ( ows),→ 4.2. IMPLEMENTATION 47 p e s_de =np.a ay([so ed(lis ( ange(n)), key=lambda i:d_ma ix[i][j]) o jin ange(m)]) # so by dis ance (columns) ,→ ,→ # Call GS algo i hm ( e u ns lis o size 1 se s) ma ch_p e , ma ch_de =gale_shapley(n, m, [1 o _in ange(n)], [1 o _in ange(m)], p e s_p e , p e s_de ),→ ma ch_p e =[nex (i e (x)) i len(x) != 0 else None o xin ma ch_p e ],→ ma ch_de =[nex (i e (x)) i len(x) != 0 else None o xin ma ch_de ],→ # T acking will wo k p ope ly i he agen s' ange o mo emen is smalle han,→ # 1/(2* ol) imes hei sepa a ion assignmen _dic =dic () assignmen _lis =[] o iin ange(m): i ma ch_de [i] is None:# New agen s iden i y =gen_iden i y(A) assignmen _dic [iden i y] =cen e s_ [i] assignmen _lis .append(iden i y) A+= 1 # uppe bound o las assigned agen numbe else:# Exis ing agen iden i y =p e ious_cen e s[ma ch_de [i]][0] assignmen _dic [iden i y] =cen e s_ [i] assignmen _lis .append(iden i y) ole ance[iden i y] = ol o iin ange(n): i ma ch_p e [i] is None:# Misde ec ion o agen le iden i y =p e ious_cen e s[i][0] i ole ance[iden i y] > 0:# Add o ma ch wi h las posi ion,→ assignmen _dic [iden i y] =assig[iden i y] assignmen _lis .append(iden i y) # These will appea a e all de ec ions,→ ole ance[iden i y] -= 1 else: del ole ance[iden i y] p in ("Dele ing agen ",iden i y) p in ("Cu en assignmen is", assignmen _dic ) e u n assignmen _dic , assignmen _lis , A 48 CHAPTER 4. THE VISUAL TRACKING MODULE Chap e 5 The emo ion ecogni ion module The las building block o ou sys em is he emo ion ecogni ion module. This module is in cha ge o in e ing he dimensional sco es explained in Chap e 3 by using he acking eeds ex ac ed by he acking module (see Chap e 4). A each poin in ime , he emo ion ecogni ion module p ocesses he iden i ied ames gene a ed by he acking module and ou pu s he dimensional sco es o each agen (see Fig. 5.1). We ea emo ion ecogni ion as a eg ession p oblem, whe e we wan o es ima e each one o he dimensional sco es (A ,En,Le,Ec) in he ange [0,1] o a gi en inpu ideo. We decided o u he simpli y his by conside ing ha a poin wise es ima ion is good enough. Tha is, we calcula e he sco es only o he ames gene a ed a ime wi hou using in o ma ion om p e ious ames. Figu e 5.1: The emo ion ecogni ion pipeline illus a ed. T acking eeds p o ide he sys em wi h con inuous in o ma ion o each one o he agen s. Howe e , as we will discuss u he in his chap e , making use o he empo al in o ma ion a ailable added an addi ional laye o complexi y o he module design ha did no make sense o a p oo -o -concep . As we ha e done in Chap e 3 wi h he acking module, we op ed o a simple implemen a ion ha allowed us o ob ain esul s quickly. In he ollowing sec ions he emo ion ecogni ion module will be p esen ed in de ail. Fi s , we will in oduce he sys em a chi ec u e. Then, we will de ail i s implemen a- ion. Finally, we will p esen he esul s ob ained on he ERVF da ase and he sys em limi a ions. 49 56 CHAPTER 5. THE EMOTION RECOGNITION MODULE ans o ms.ToTenso (), ans o ms.No malize(mean=mean, s d=s d)]) de __len__(sel ): """Re u ns he o al numbe o ames in he Da ase .""" e u n len(sel .lis _IDs) de __ge i em__(sel , index): """Gene a es a p ocessed ame wi h i s associa ed sco es.""" # Selec sample ID =sel .lis _IDs[index] # Ex ac ideo name om ID _name, ame =ID.spli ('.') _name += '.webm' ame =in ( ame.spli ('_')[-1]) # Load ame and label cap =c 2.VideoCap u e(DATA_DIR+VIDEO_DIR+ _name) cap.se (1, ame) _, ame =cap. ead() cap. elease() ame=Image. oma ay( ame,'RGB') label =sel .labels[ID] e u n sel .p ep ocess( ame), o ch. enso (label). loa () The gene a o s a e hen buil wi h he ollowing code. # C ea ing da ase in o ( ain, es , alida ion) ain_ids, ain_sco es =gene a e_da ase _in o(' iles_ ain. x ', d op=20) , → al_ids, al_sco es =gene a e_da ase _in o(' iles_ alida ion. x ', d op=20),→ es _ids, es _sco es =gene a e_da ase _in o(' iles_ es . x ',d op=20) # Pa i ion ini ializa ion pa i ion =dic () labels =dic () pa i ion[' ain']= ain_ids pa i ion[' alida ion']= al_ids pa i ion[' es ']= es _ids o x,y in zip( ain_ids, ain_sco es): labels[x] =y o x,y in zip( al_ids, al_sco es): labels[x] =y o x,y in zip( es _ids, es _sco es): labels[x] =y pa ams ={'ba ch_size':256, 'shu le':T ue, 'pin_memo y':T ue, 5.2. IMPLEMENTATION 57 'num_wo ke s':0} ain_da ase =Da ase (pa i ion[' ain'], labels) al_da ase =Da ase ( al_ids,labels) es _da ase =Da ase ( es _ids,labels) ain_gene a o = o ch.u ils.da a.Da aLoade ( ain_da ase , **pa ams) al_gene a o = o ch.u ils.da a.Da aLoade ( al_da ase , **pa ams) es _gene a o = o ch.u ils.da a.Da aLoade ( es _da ase , **pa ams) 5.2.3 T aining, alida ion, and es ing To ain a py o ch ne wo k i is necessa y o p og am he aining loop explici ly. We ained ou sys em using he Adam op imiza ion me hod using he ollowing code. # CUDA o PyTo ch use_cuda = o ch.cuda.is_a ailable() de ice = o ch.de ice("cuda:0" i use_cuda else "cpu") o ch.backends.cudnn.benchma k =T ue p in ("USE_CUDA:",use_cuda) # T aining pa ame e s max_epochs = 20 model =Ne () i use_cuda: model =model.cuda() c i e ion = o ch.nn.L1Loss( educe="sum") op imize = o ch.op im.Adam(model.pa ame e s(),l =0.0005) N_AVG = 4 ain_his o y, al_his o y =[],[] # T aining and alida ing he ne wo k o epoch in ange(max_epochs): ############################ T aining ############################# model. ain() o al_loss, unning_loss,i = 0.0,0.0,1 o local_ba ch, local_labels in ain_gene a o : # T ans e o GPU and model compu a ions local_ba ch, local_labels =local_ba ch. o(de ice), local_labels. o(de ice),→ p in ('T aining ba ch %i/%i'%(i , len( ain_gene a o ))) op imize .ze o_g ad() ou =model(local_ba ch) loss =c i e ion(ou , local_labels) loss.backwa d() op imize .s ep() unning_loss += loa (loss) o al_loss += loa (loss) i += 1 i i %N_AVG == 0: 58 CHAPTER 5. THE EMOTION RECOGNITION MODULE p in ('<<T aining>> [%d,%5d] loss: %.3 '%(epoch + 1, i + 1, unning_loss /N_AVG)),→ unning_loss = 0.0 p in ("<<T aining>> Accumula ed loss: %.3 A e age loss: %.3 "% ( o al_loss, o al_loss/len( ain_gene a o ))),→ ain_his o y.append( o al_loss/len( ain_gene a o )) ################################## Valida ion #################### model.e al() o al_loss, unning_loss,i = 0.0,0.0,0 i = 1 wi h o ch.se _g ad_enabled(False): o local_ba ch, local_labels in al_gene a o : # T ans e o GPU and model compu a ions local_ba ch, local_labels =local_ba ch. o(de ice), local_labels. o(de ice),→ p in ('Valida ion ba ch %i/%i'%(i , len( al_gene a o ))) ou =model(local_ba ch) loss =c i e ion(ou , local_labels) unning_loss += loa (loss) o al_loss += loa (loss) i += 1 i i %N_AVG == 0: p in ('<<Valida ion>> [%d,%5d] loss: %.3 '%(epoch + 1, i + 1, unning_loss /N_AVG)),→ unning_loss = 0.0 p in ("<<Valida ion>> Accumula ed loss: %.3 A e age loss: %.3 "% ( o al_loss, o al_loss/len( al_gene a o ))),→ al_his o y.append( o al_loss/len( al_gene a o )) p in ('T aining loss (epochwise):', ain_his o y) p in ('Valida ion loss (epochwise):', al_his o y) No e ha he code abo e epo s s a is ics e e y N_AVG ba ches. Fo each epoch, he ne wo k is ained on he aining se and subsequen ly alida ed on he alida ion se . We moni o ed he aining o 18 epochs. A e 8epochs, we ound ha alida ion and aining losses s a ed o di e ge (see Fig. 5.6). The e o e, he ne wo k was e ained om sc a ch o 8epochs on he combined ain and alida ion da ase s. Figu e 5.6: E olu ion o aining and alida ion losses. 5.2. IMPLEMENTATION 59 The 3 ideos co esponding o a andomly chosen agen we e held ou o es ing pu poses. We an ou model on he es ing se and calcula ed he a e age in e ence ime, o abou 37 ms o a ound 27 FPS. de es _model(model, gene a o , use_cuda=False): """Re u ns he esul s o he model on he gi en gene a o o each dimension as a dic iona y,,→ as well as he p edic ions, g ound u h and a e age in e ence ime.""",→ model.e al() p edic ed,labels,in _ imes =[],[],[] i = 1 wi h o ch.se _g ad_enabled(False): o local_ba ch, local_labels in gene a o : # T ans e o GPU and model compu a ions local_ba ch, local_labels =local_ba ch. o(de ice), local_labels. o(de ice),→ p in ('Tes ing ba ch %i/%i'%(i , len(gene a o ))) i use_cuda: s a = o ch.cuda.E en (enable_ iming=T ue) end = o ch.cuda.E en (enable_ iming=T ue) # Reco d in e ence ime s a . eco d() ou =model(local_ba ch) end. eco d() # Sync o ch.cuda.synch onize() in _ imes.append(s a .elapsed_ ime(end)) else: # Reco d in e ence ime s a = ime. ime() ou =model(local_ba ch) end = ime. ime() in _ imes.append(end-s a ) p edic ed.append(ou ) labels.append(local_labels) i += 1 p ed, ue=[],[] p ed_np, ue_np =[],[] i use_cuda: p ed_np =[x.cpu().numpy() o xin p edic ed][:-1] ue_np =[x.cpu().numpy() o xin labels][:-1] else: p ed_np =p edic ed[:-1] ue_np =labels[:-1] p ed =np.s ack(p ed_np, axis=0). eshape(-1,4) ue =np.s ack( ue_np, axis=0). eshape(-1,4) cols =['A ','En','Ec','Le'] 60 CHAPTER 5. THE EMOTION RECOGNITION MODULE sco es_by_col =dic () o col in [0,1,2,3]: sco es_by_col[cols[col]] ={"MSE": mean_squa ed_e o ( ue[:,col], p ed[:,col]),,→ "MAE": mean_absolu e_e o ( ue[:,col], p ed[:,col]), "MAPE": mean_absolu e_pe cen age_e o ( ue[:,col], p ed[:,col]), "R2": 2_sco e( ue[:,col], p ed[:,col])} e u n p ed, ue,sco es_by_col,sum(in _ imes)/len(in _ imes) 5.3 Resul s and limi a ions The e a e se e al me ics ha measu e how e ec i ely a eg ession sys em is. To e alua e ou sys em, we op ed o use he ollowing 4 o each dimension o he audience expe ience: Mean Squa ed E o (MSE), Mean Absolu e E o (MAE), Mean Absolu e Pe cen age E o (MAPE), and R2sco e. I Y={y1, y2, ..., yn} ⊂ Ris he se o p edic ions and ˆ Y={ˆy1,ˆy2, ..., ˆyn} ⊂ Ris he se o g ound u h alues, he me ics abo e a e calcula ed as ollows MSE(ˆ Y , Y ) = 1 n n X i=1 (ˆyi−yi)2MAE(ˆ Y , Y ) = 1 n n X i=1 |ˆyi−yi| MAPE =1 n n X i=1  ˆyi−yi max(ˆyi, ε) R2(ˆ Y , Y ) = 1 −Pn i=1 (ˆyi−yi)2 Pn i=1 (ˆyi−¯y)2 whe e ¯yis he mean o ˆ Y. Ideally, we wish o minimize MSE, MAE, and MAPE and maximize R2. We mus make a couple obse a ions abou ou me ics • MSE is p one o anishing when wo king wi h e y small di e ences (because o he squa ed alues in he sum). I is also exp esses dissimila i y in e ms o he absolu e di e ence be ween alues and i s in e p e a ion mus be made ca e ully. • MAE also exp esses dissimila i y in e ms o he absolu e di e ence and i s in e - p e a ion mus also be ca e ul. • MAPE sco e may epo high alues due o low g ound u h sco es. •R2may epo low alues due o g ound u h alues being oo close o he mean. Since ou dimensional sco es a e in he [0,1] ange, we mus be pa icula ly ca e ul when assessing MAPE and R2. To e alua e ou p oblem, we ake pa icula in e es a ha ing low MAE alues, which we belie e bes e lec s ou model’s pe o mance. Table 5.1: Resul s on he es da ase . MSE MAE MAPE R2 A ec i e Response 0,0861 0,2557 1,2240 -0,1142 Engagemen 0,1556 0,3093 14,0547 -0,3560 Emo ional Connec ion 0,0757 0,2324 0,7135 -0,4725 Lea ning 0,1775 0,2906 6.3201 -0,4725 Combined 0,1099 0,2642 1,5040 -0,2330 5.3. RESULTS AND LIMITATIONS 61 We es ed ou model on 3unseen ideos o a andomly chosen agen ha he ne wo k had no p e ious in o ma ion on. The esul s o each dimension and hei combined alues a e shown in Table 5.1. As we can see, MSE has ela i ely low alues while MAPE is e y la ge. R2being nega i e indica es ha a cons an p edic ion ma ching he mean o he es da ase would pe o m be e han ou sys em. While he esul s in Table 5.1 a e no conclusi e, a ending o MAE, he sys em seems o be able o disc imina e be ween ex eme cases (e.g: e y low o e y high Engagemen ). Fu he mo e, he poo R2sco e may be in luenced by he linea in e ence used o de e mine he sco es and he MAPE alues may be in luenced by he scale o he esidues. The esul s may also be in luenced by some o he limi a ions o ou sys em and ou da ase : • The da ase is composed o highly co ela ed da a since all o he ideos co espond o 7pa icipan s. A la ge -scale expe imen would undoub edly p o ide us wi h a iche and mo e a ied da ase and, p esumably, a mo e obus sys em. • The da ase is ela i ely small once he downsampling has aken place. To a ce ain ex en , i is possible ha da a augmen a ion could be used o palia e his issue. • The p oposed model is e y simple. The main assump ion ha may no hold is ha we can exp ess each o he dimensional sco es as he sigmoid o a linea unc ion o he con olu ional ea u es. Addi ional laye s in he eg esso s may allow hem o p o ide a be e es ima ion, al hough he sys em may be mo e p one o o e i ing. • The p oposed model does no ake in o accoun empo al dependencies. We expec he dimensional sco es o a ce ain indi idual o e ol e “con inuously” in ime. Tha is, we expec ha pe son o ha e simila dimensional sco es a close poin s in ime. An RNN model may be able o sol e his p oblem, o example, by sha ing he p e ious dimensional sco es wi h he eg esso s. Besides es me ics and esul s, wo impo an lines o wo k o imp o e his sys em a e ha o explainabili y and ai ness. Explainabili y is help ul in unde s anding exac ly how o sol e exis ing issues in he sys em and is key in ensu ing ha he sys em is eally wo king as expec ed. Fai ness s udies a e necessa y in o de o de ec po en ial bias and noise in a i icial in elligence sys ems, and poo ai ness esul s may be he di ec cause o pe o mance issues. Add essing he abo e limi a ions should be he nex s ep in he pa h o building a obus and accu a e sys em o au oma ic audience expe ience measu emen . 62 CHAPTER 5. THE EMOTION RECOGNITION MODULE Chap e 6 Conclusions In his documen we ha e p esen ed ou 3 main con ibu ions: a comp ehensi e amewo k o quan i y audience expe ience, he Emo ion Recogni ion om Video Feeds (ERVF) da ase o ain machine lea ning sys ems using he p e ious amewo k, and a p oo - o -concep machine lea ning sys em1 ha es ima es audience expe ience om ideo in eal- ime. We ha e also been able o imp o e he ma ching s ep commonly used in acking-by-de ec ion sys ems, educing he complexi y om O(n3) o O(n2). While he e a e p oposals on how o measu e he audience expe ience, as explained in Chap e 2, cu en ly he e is no s anda d me hodology o doing so. Ou hea e -based amewo k is mean o be a i s s ep in sol ing he issue. Simila ly, using his amewo k o ain ML sys ems equi es he a ailabili y o sui able da a. The expe imen desc ibed in Chap e 3 allows o he cons uc ion o sui able da ase s a any desi ed scale. The ERVF da ase p o ides a s a ing poin in his ega d. Finally, he objec acking and emo ion ecogni ion sys ems, desc ibed in Chap e 4 and Chap e 5 espec i ely, se e as a baseline in he design o sys ems o au oma ic audience expe ience es ima ion om ideo. Addi ionally, we ha e ound ha mos o he widely a ailable ideo con e encing so wa e only allows o e y limi ed con ol. To add ess his issue, we ha e p oposed ou own ideo con e encing ool, whose code is p esen ed in Appendix A, ailo ed speci ically o ou needs. Ou wo k can be applied and ha e an impac in a b oad se o ields. In educa ion, ha ing a sys em o au oma ically assess he s a e o s uden s allows eache s o adjus hei discou se on he ly. In comedy, i allows he comedian o assess he quali y o hei ac . Mo e gene ally, in any e en whe e public speaking is in ol ed, i allows he speake o ecei e immedia e eedback and adjus acco dingly. In o de o his impac o be meaning ul, some o he majo limi a ions mus be ad- d essed. Some o he dimensions o ou amewo k a e s ill e y abs ac and ha d o measu e. Fu he mo e, measu ing he dimensional sco es o he amewo k expe imen- ally is an in usi e p ocess, in he sense ha illing a su ey in e up s he audience’s expe ience. The in o ma ion ob ained as desc ibed in Chap e 3 is also highly edundan , 1The code o he ull ML sys em is a ailable in ou Gi Hub eposi o y h ps://gi hub.com/ pablo- s/emo ion- ecogni ion 63 64 CHAPTER 6. CONCLUSIONS bo h due o he in e ence p ocess and due o he audience expe ience gene ally being con- inuous (simila o close poin s in ime). This edundancy hu s he pe o mance o ML sys ems. When i comes o ou ML sys em, he e is also oom o imp o emen . The acking module is sensi i e o occlusion and o he de ec o ailing o a pe iod o ime. I se e al ames a e los , hen he iden i y will p obably be los oo. The de ec o in he objec acking module is also sensi i e o backg ound objec s esembling humans. The emo ion de ec ion sys em, on he o he hand, is e y simple and does no ake in o accoun empo al dependencies. The e a e, he e o e, se e al lines o wo k ha ma k he nex s eps in he de elopmen o a obus au oma ic audience expe ience es ima o . Fi s ly, ou amewo k’s dimensions mus be s udied mo e in-dep h in o de o acili a e measu emen . The combined me ic Smus be alida ed, since a mo e ca e ully chosen combina ion o dimensions may be e es ima e audience expe ience. A la ge and mo e a ied da ase mus also be c ea ed h ough a la ge-scale expe imen . T acking sys em imp o emen s should be ocused on educing he sensi i i y o he de ec o and he numbe o iden i y swi ches while min- imizing he impac on la ency. Finally, imp o emen s o he emo ion de ec ion sys em should ocus on inco po a ing in o ma ion om p e ious p edic ions. Appendices 65 72 APPENDIX A. CODE gum =awai IonSDK.LocalS eam.ge Use Media(cons ain s).ca ch( (e o ) => { ale ("Could no access local s eam: " +e o ); } ); i (!gum) e u n null; localS eams[gum.id] =gum; localS eamId =gum.id; e u n gum.id; } cons ge DisplayS eam =async (cons ain s =null) => { gum =awai IonSDK.LocalS eam.ge DisplayMedia({ ideo: ue, audio: ue}).ca ch(,→ (e o ) => { ale ("Could no access sc een: " +e o ); } ) console.log("Go display s eam " +gum.id); i (!gum) e u n null; localS eams[gum.id] =gum; console.log("Go display s eam " +gum.id); e u n gum.id; }; unc ion se upClien Hos () { clien Hos .on ack =( ack, s eam) => { console.log("go ack", ack.id, " o s eam", s eam.id); s eam.on emo e ack =() => { console.log("T ack ended"); emo eRemo eS eamElemen (s eam.id); } s Elem =ge Remo eS eamElemen (s eam.id); i (s Elem.s cObjec === null) { s Elem.s cObjec =s eam; }else { s Elem.s cObjec .addT ack( ack); } 73 }; } unc ion addLocalS eam(id =null, hos = alse) { ge LocalS eamElemen (id, hos ).s cObjec =localS eams[id]; //ge LocalS eamElemen (id).mu ed = ue; console.log("Local ideo ack added"); } unc ion emo eLocalS eam(id) { localS eams[id].ge T acks(). o Each(( ) => .s op()); dele e localS eams[id] emo eLocalS eamElemen (id); } unc ion se upClien Sub() { clien Sub.on ack =( ack, s eam) => { console.log("go ack", ack.id, " o s eam", s eam.id); s eam.on emo e ack =() => { console.log("T ack ended"); emo eRemo eS eamElemen (s eam.id); } s Elem =ge Remo eS eamElemen (s eam.id); i (!(s eam.id in subsc ibe s)) { console.log("New subsc ibe "); subsc ibe s[s eam.id] =null; o (id in ideos) { i ( ideos[id].ge I ame().pa en Elemen .child en[1].child en[0].on) { code = ideos[id].ge Playe S a e() == 1 ? " -4 " :" -3 " con ol.send(id +code + ideos[id].ge Cu en Time()); console.log(id +code + ideos[id].ge Cu en Time()); } } } i (s Elem.s cObjec === null) { s Elem.s cObjec =s eam; }else { s Elem.s cObjec .addT ack( ack); } }; } /* * * UI 74 APPENDIX A. CODE * */ cons emo eS eamCon aine =documen .ge Elemen ById(" emo e-s eams"); le localS eamCon aine ; unc ion se LocalS eamCon aine (hos = alse) { i (hos ) localS eamCon aine = documen .ge Elemen ById("local-s eams"); , → else localS eamCon aine = emo eS eamCon aine ; } unc ion ge Remo eS eamElemen (id, hos = alse) { le elem =documen .ge Elemen ById(" emo e-s eam-"+id) i (elem === null) { le di =documen .c ea eElemen ("di "); i (hos ) di .classLis .add(" ideo"); else di .classLis .add("s eam-con aine "); le news =documen .c ea eElemen (" ideo"); news .au oplay = ue; news .id =" emo e-s eam-"+id; news .playsinline = ue; news .mu ed = alse; di .appendChild(news ); emo eS eamCon aine .appendChild(di ); elem =news ; } e u n elem } unc ion emo eRemo eS eamElemen (id) { emo eS eamCon aine . emo eChild( ge Remo eS eamElemen (id).pa en Elemen ); } unc ion ge LocalS eamElemen (id, hos = alse) { le elem =documen .ge Elemen ById("local-s eam-"+id) i (elem === null) { le di =documen .c ea eElemen ("di "); i (hos ) di .classLis .add("local","s eam-con aine "); else 75 di .classLis .add("s eam-con aine "); le news =documen .c ea eElemen (" ideo"); news .classLis .add("local"," ideo"); news .au oplay = ue; news .mu ed = ue; news .id ="local-s eam-"+id; news .playsinline = ue; i (hos ) { le con Di =c ea eS eamCon ols(news , id); di .appendChild(news ); di .appendChild(con Di ); }else { di .appendChild(news ); } localS eamCon aine .appendChild(di ); elem =news ; } e u n elem } unc ion c ea eS eamCon ols(s Elem, id) { le con =documen .c ea eElemen ("di "); con .classLis .add("s eam-con ols"); con .inne HTML =` <di class="publish bu on"> <di class="icon"> <i class=" a a-wi i" a ia-hidden=" ue"></i> </di </di >`; con .inne HTML += ` <di class="came a bu on"> <di class="icon"> <i class=" a a- ideo-came a" a ia-hidden=" ue"></i> </di </di >`; con .inne HTML += ` <di class="mu e bu on"> <di class="icon"> <i class=" a a-mic ophone" a ia-hidden=" ue"></i> </di </di >`; 76 APPENDIX A. CODE con .inne HTML += ` <di class=" emo e bu on"> <di class="icon"> <i class=" a a- imes" a ia-hidden=" ue"></i> </di </di >`; le publish =con .child en[0]; publish.on = alse; publish.onclick =() => publishS eam(s Elem, publish); le disable =con .child en[1]; disable.on = alse; disable.onclick =() => disableS eam(s Elem, disable); le mu e =con .child en[2]; mu e.on = alse; mu e.onclick =() => mu eS eam(s Elem, mu e); le emo e =con .child en[3]; emo e.onclick =() => { i (s Elem.s cObjec != unde ined) { s Elem.s cObjec .unpublish(); emo eLocalS eam(id); }else { ideos[id].s opVideo(); dele e ideos[id] localS eamCon aine . emo eChild( documen .ge Elemen ById("playe ="+id).pa en Elemen ); } } e u n con ; } unc ion publishS eam(s Elem, publish) { i (s Elem.s cObjec == unde ined) { publishVideo(s Elem.id, publish); e u n; } i (publish.on) { s Elem.s cObjec .unpublish(); publish.on = alse; publish.s yle.backg oundColo ="black"; }else { clien Hos .publish(s Elem.s cObjec ); 77 publish.on = ue; publish.s yle.backg oundColo ="ligh blue"; } } unc ion publishVideo(id, publish) { id =id.spli ("=")[1]; i (publish.on) { console.log("unpublish "+id); con ol.send(id +" 0"); publish.on = alse; publish.s yle.backg oundColo ="black"; }else { console.log("publish "+id +" " + ideos[id].ge Cu en Time()); i ( ideos[id].ge Playe S a e() == 1) { con ol.send(id +" -2 " + ideos[id].ge Cu en Time()); }else { con ol.send(id +" -1"); } publish.on = ue; publish.s yle.backg oundColo ="ligh blue"; } } unc ion disableS eam(s Elem, disable) { i (disable.on) { s Elem.s cObjec .ge VideoT acks()[0].enabled = ue; s Elem.s cObjec .unmu e(' ideo'); disable.on = alse; disable.s yle.backg oundColo ="ligh blue"; }else { s Elem.s cObjec .ge VideoT acks()[0].enabled = alse; s Elem.s cObjec .mu e(' ideo'); disable.on = ue; disable.s yle.backg oundColo ="black"; } } unc ion mu eS eam(s Elem, mu e) { i (mu e.on) { s Elem.s cObjec .unmu e('audio'); mu e.on = alse; mu e.s yle.backg oundColo ="ligh blue"; }else { s Elem.s cObjec .mu e('audio'); mu e.on = ue; mu e.s yle.backg oundColo ="black"; } 78 APPENDIX A. CODE } unc ion emo eLocalS eamElemen (id) { localS eamCon aine . emo eChild( ge LocalS eamElemen (id).pa en Elemen ); } le ideos ={}; unc ion ge VideoSub(id, playS a e, s) { le di =documen .c ea eElemen ("di "); di .classLis .add("s eam-con aine "); le news =documen .c ea eElemen ("di "); news .id ="playe ="+id; news .classLis .add("local"," ideo"); le o e lay =documen .c ea eElemen ("di "); o e lay.classLis .add("playe -o e lay"); le o e lay2 =documen .c ea eElemen ("di "); o e lay.classLis .add("playe -o e lay"); o e lay2.onclick =() => { ideos[id].playVideo(); ideos[id].s opVideo();};,→ le s =documen .c ea eElemen ("di "); s.classLis .add(" ullsc een-bu on"); le icon =documen .c ea eElemen ("i"); icon.classLis .add(" a"," a-squa e-o"); icon.a iaHidden = ue; di . esizeObs =new ResizeObse e ((en ies) => { ideos[id].se Size(en ies[0].con en Rec .wid h, en ies[0].con en Rec .heigh );,→ }); di . esizeObs.obse e(di ); s.appendChild(icon); o e lay.appendChild( s); o e lay.appendChild(o e lay2); di .appendChild(o e lay); s.onclick =() => { i (!documen . ullsc eenElemen ) { 79 di . eques Fullsc een(); }else { documen .exi Fullsc een(); } } di .appendChild(news ); emo eS eamCon aine .appendChild(di ); unc ion onPlaye Ready(e en ) { e en . a ge .se PlaybackQuali y('hd720'); e en . a ge .seekTo( s); e en . a ge .playVideo(); i (playS a e != -2) e en . a ge .s opVideo(); } playe =new YT.Playe ('playe ='+id, { s yle:"heigh : au o; wid h: 100%;", ideoId:id, playe Va s:{'con ols':0}, disablekb: 1, e en s:{ 'onReady':onPlaye Ready, //'onS a eChange': onPlaye S a eChange } }); ideos[id] =playe ; } async unc ion ge VideoHos () { ideoId =documen .ge Elemen ById(" ideo-id"). alue; le di =documen .c ea eElemen ("di "); di .classLis .add("local","s eam-con aine "); le news =documen .c ea eElemen ("di "); news .id ="playe ="+ ideoId; news .classLis .add("local"," ideo"); le con Di =c ea eS eamCon ols(news , ideoId); di .appendChild(news ); di .appendChild(con Di ); localS eamCon aine .appendChild(di ); 80 APPENDIX A. CODE unc ion onPlaye Ready(e en ) { //e en . a ge .playVideo(); } unc ion onPlaye S a eChange(e en ) { i (con Di .child en[0].on) con ol.send( ideoId +" " +e en .da a +" " + ideos[ ideoId].ge Cu en Time());,→ i ( eco dBu on.on) sendReco dTs( ideoId +" " +e en .da a +" " + ideos[ ideoId].ge Cu en Time());,→ } playe =new YT.Playe ('playe ='+ ideoId, { s yle:"heigh : au o; wid h: 80%;", ideoId: ideoId, //playe Va s: { 'au oplay': 1, 'con ols': 0 }, e en s:{ 'onReady':onPlaye Ready, 'onS a eChange':onPlaye S a eChange } }); ideos[ ideoId] =playe ; } /* * * CHAT * */ le cha ; async unc ion s a Cha (hos = alse) { awai wai Fo Clien (); cha =clien Hos .c ea eDa aChannel("cha "); cha .onmessage =(msg) => { cha Elemen . alue += msg.da a; } messageElemen .onkeyup =(e ) => { i (e .keyCode == 13) { cha .send((hos ?"Hos : n" :UNIQUE_ID +': n')+ messageElemen . alue);,→ cha Elemen . alue += "Tú: n"; cha Elemen . alue += messageElemen . alue; 81 messageElemen . alue =""; } } } a ag =documen .c ea eElemen ('sc ip '); ag.s c ="h ps://www.you ube.com/i ame_api"; a i s Sc ip Tag =documen .ge Elemen sByTagName('sc ip ')[0]; i s Sc ip Tag.pa en Node.inse Be o e( ag, i s Sc ip Tag); // 3. This unc ion c ea es an <i ame> (and YouTube playe ) // a e he API code downloads. le you ubeReady = alse; a playe ; unc ion onYouTubeI ameAPIReady() { you ubeReady = ue; } Figu e 4. O e iew o he sys em a chi ec u e (OT – Objec T acking, ER – Emo ion Recogni ion, A g - A e age). Then, he s eams associa ed wi h each iden i y en e he emo ion ecogni ion module, which akes a ba ch o ames and ou pu s he alue o he 4 me ics o each ame in he ba ch. These me ics ep esen he sco ing de e mined by he emo ion ecogni ion module o each o he 4 dimensions o audience expe ience: a ec i e esponse, engagemen , emo ional connec ion, and lea ning. Finally, he las s ep is a pooling s ep ha e u ns he alue o hese me ics o he whole ba ch. 2.2.1 A chi ec u al implemen a ion The objec acking module wo ks in 2 s eps. Fi s , we apply s anda d YOLO 4 (Bochko skiy, Wang, & Liao, 2020) o de ec pa icipan s in he ame. We ha e c ea ed a cus om algo i hm ha is hen used o ack speci ic pa icipan s. This algo i hm wo ks by keeping ack o he posi ions o he objec cen e s in successi e ames as well as he bounding box sizes. Th ough his his o y, we can p edic he nex objec ’s bounding box cen e and size. Le 𝒄𝑛 and 𝒃𝒃𝑛 be, espec i ely, he las obse ed bounding box cen e and size (wid h and heigh ) o a pa icula objec . Then, he a ia ion in cen e posi ion and bounding box size Δ𝑐 and Δ𝑏 a e de ined, espec i ely, as Δ𝒄=𝒄𝑛−𝒄𝑛−1 Δ𝒃𝒃 =𝒃𝒃𝑛−𝒃𝒃𝑛−1 (3) The alues ob ained by applying (3) can hen be used o p edic he expec ed cen e posi ion 𝒄𝑛+1 and bounding box size 𝒃𝒃𝑛+1 𝒄𝑛+1 = 𝒄𝑛+Δ𝒄 𝒃𝒃𝑛+1 =𝒃𝒃𝑛+Δ𝒃𝒃 (4) Fo a single obse a ion, he a ia ions ob ained in equa ion (3) a e ze o. This p ocess is essen ially a Kalman il e (Kalman, 1960). Wi h he p edic ed alues and wi h each bounding box 𝒃𝒃 being ep esen ed as a pai o heigh and wid h (ℎ,𝑤) in pixels, we ma ch exis ing iden i ies o he obse ed pa icipan s wi h he mos simila bounding box cha ac e is ics, minimizing he me ic m(𝒄,𝒃𝒃)=||𝒄−𝒄𝑛+1||+𝑓(ℎ,ℎ𝑛+1)+𝑓(𝑤,𝑤𝑛+1) (5) whe e 𝑓(𝑥,𝑦)=log(1+|𝑥−𝑦|). Mos o he ime, he i s e m in equa ion (5) will be conside ably la ge han he o he wo, as he wid h and heigh a e only ele an when he e a e a ious pa icipan s wi h e y simila cen e s. To ob ain he ma ching e icien ly, many acking sys ems (Bewley, Ge, O , Ramos, & Upc o , 2016; A un Kuma , Laxmanan, Ram Kuma , S inidh, & Ramana han, 2021; Luo, Xing, Milan, Zhang, Liu, & Kim, 2021) p opose he usage o he Hunga ian algo i hm, which uns in ime complexi y 𝒪(𝑛3) (Kuhn, 1955; Munk es, 1957). We, on he o he hand, ha e ound an al e na i e app oach ha emains la gely unexplo ed in he li e a u e and is only implemen ed by a ew selec sys ems (Godbehe e & Goldbe g, 2014; Oh e al, 2020). I we o mula e he ma ching p oblem abo e as a s able ma ching p oblem (SMP) by con e ing he dis ance ma ix de e mined by me ic 𝑚(⋅,⋅) in o p e e ence lis s o he p e ious de ec ions and he de ec ed objec , we can use he Gale-Shapley (GS) algo i hm (Gale & Shapley, 1962) o pe o m he ma ching. This algo i hm has a lowe ime complexi y han he Hunga ian algo i hm, being capable o unning in 𝒪(𝑛2). The las issue o sol e in he acking module is how o deal wi h si ua ions whe e he numbe o expec ed de ec ions (which is he same as he numbe o iden i ies in he las ame) and he numbe o de ec ions does no ma ch. I he e a e mo e de ec ions han expec ed, new iden i ies a e c ea ed o hose ha emain unma ched a e unning he GS algo i hm. Addi ionally, we conside a ole ance alue τ𝑘 o each exis ing iden i y which indica es how many ames he sys em will ole a e no inding a sui able ma ch o said agen . I he e a e mo e iden i ies han de ec ions, he ole ance o unma ched iden i ies is dec eased and iden i ies wi h 0 ole ance a e dele ed. I he e is a su icien ly high ame a e wi h espec o he speed a which he pa icipan s mo e, he de ec ion accu acy is high. The emo ion ecogni ion module also wo ks in 2 s eps. Fi s , we use a con olu ional neu al ne wo k (CNN) (LeCun, Ha ne , Bo ou, & Bengio, 1999) o ob ain a low-dimensional embedding o he inpu ames. We a o a chi ec u es p e- ained on ImageNe (Deng e al, 2010), as ans e lea ning is s anda d o imp o e pe o mance in compu e ision models (Weiss, Khoshgo aa , & Wang, 2016; Oquab, Bo ou, Lap e , & Si ic, 2014; Hussain, Bi d, & Fa ia, 2019). In his pape , we p opose he use o MobileNe V3 (Howa d e al., 2019), bu he e a e o he a chi ec u es (Zoph, Vasude an, Shlens, & Le, 2017; Simonyan & Zisse man, 2015; He, Zhang, Ren, & Sun, 2016) ha can ul ill his pu pose equally well. This embedding is hen passed on o 4 ully connec ed (FC) laye s wi h linea ac i a ion, each ained o p edic a pa icula dimension-speci ic sco e (𝐴𝑓,𝐸𝑛,𝐿𝑒,𝐸𝑐). Al e na i ely, i should also be possible o use suppo ec o machines (SVMs) (Co es & Vapnik, 1995). Finally, he pooling module ex ac s he global sco es o each dimension by a e aging ac oss he numbe o agen s as indica ed by (1). I is also possible o hen calcula e he summa y me ic as shown in (2). This module mus adap o changes in he numbe o pa icipan s in successi e p edic ions. 3. Expe imen , da ase , and esul s 3.1.1 Expe imen al design and pa icipan s While he e a e some audio isual da ase s on audience expe ience and a ec i e esponse (Cu is e al., 2015; Soleymani e al., 2012), hey a e no a ailable o he gene al public and do no con ain da a in online se ings. Fo his eason, we an a da a collec ion expe imen and c ea ed a da ase o online audience esponse. A o al o 8 las yea Spanish college s uden s (6 males and 2 emales) olun ee ed o ake pa in he s udy. All o hem we e a ound 20 yea s old. The expe imen was conduc ed as ollows: he olun ee s connec ed o ou cus om ideo con e encing ool wi h hei pe sonal compu e s and webcams. We s eamed h ee sho (10min) ideo pe o mances while we eco ded he pa icipan s wi h hei webcams. A e wa ching each pe o mance, he pa icipan s illed ou a sho ques ionnai e. The ideo pe o mances we e selec ed om a pool o YouTube ideos o a ied con en , including TED alks, monologues, and poli ical speeches. The ques ionnai e was di ided in o h ee sec ions, measu ing he i s , second, and hi d pa s o each ideo, espec i ely. Each sec ion included 16 like -scale ques ions o sepa a ely measu e he ou di e en componen s o audience expe ience. E en hough we had ew pa icipan s and ew ideo pe o mances, a p elimina y analysis o he ques ionnai e esponses showed ha ou me hodology was able o cap u e some a ia ion in he ou dimensions, bo h ac oss ideos and h oughou each ideo. 3.1.2 Da ase A e he da a collec ion phase, he eco dings we e cleaned up and sco e labels we e gene a ed o each ame in he ollowing manne : Fi s , each eco ding is synch onized wi h he pe o mance using imes amps gene a ed by ou cus om ool. Then he eco dings a e cu o he leng h o he pe o mance, and a e p ocessed by he acking pipeline, which p oduces a se o ixed-size ideo iles acking each o he subjec s. Finally, ame-le el labels a e gene a ed by linea in e pola ion om he h ee samples gi en by he ques ionnai e esul s o each ideo. The esul ing da ase consis s o 224,206 indi idually anno a ed ames, o ming 24 ideos (8 pa icipan s wa ching 3 pe o mances). A e an ini ial explo a ion and some aining uns, we de e mined ha he da a had a high deg ee o edundancy, so we applied a 20x downsampling, d opping 19 ou o e e y 20 ames. 3.1.3 Expe imen al esul s Fo he inal aining, we used an Adam op imize (Kingma & Ba, 2015) wi h a lea ning a e o 0.0005, L1 loss unc ion, and ba ch size o 256. The model was ained o 2 epochs using 72% o he da a, while he emaining da a, eco dings belonging o 2 pa icipan s, we e used o alida ion (14%) and es ing (14%). The aining was pe o med using a single N idia RTX 3070 GPU, unning o abou 20 minu es, and he me ics o he lea ning sys em on he es da ase a e shown in Table 1. We ha e acked Mean Squa ed E o (MSE), Mean Absolu e E o (MAE), Mean Absolu e Pe cen age E o (MAPE) and 𝑅2 sco e in o de o assess he pe o mance o he model on he es da ase . Table 1. Tes ing Me ics MSE MAE MAPE 𝑅2 A . Resp. 0,0861 0,2557 1,2240 -0,1142 Engagemen 0,1556 0,3093 14,0547 -0,3560 Em. Con. 0,0757 0,2324 0,7135 -0,4725 Lea ning 0,1775 0,2906 6.3201 -0,4725 Combined 0,1099 0,2642 1,5040 -0,2330 As we can see, he MSE is ela i ely low, while he MAPE is e y la ge. Howe e , since he e o o each ame is be ween 0 and 1, his is o be expec ed and i is mos ly an a i ac o wo king wi h small numbe s. The 𝑅2 is nega i e, which implies ha ou model pe o ms wo se han a cons an p edic ion ha ma ches he mean o he es da ase . Howe e , we mus in e p e his esul in he ligh o he da a gene a ion p ocess: since o each ideo he sco es we e in e pola ed om h ee poin s, he a iance in he da ase is e y low, which migh explain such a low R2 sco e. All in all, i seems he MAE is he me ic ha bes e lec s he ac ual pe o mance o he sys em. Wha his me ic is elling us is ha ou sys em is making a e age e o s o a ound 0.3. This is no e y p ecise bu may allow he speake o ule ou ex eme si ua ions ( e y low o e y high engagemen , o example). 4. Conclusions and Fu u e Wo k In his pape we p oposed and e alua ed a gene al amewo k ha can be used o design in elligen sys ems o au oma ically e alua e audience expe ience in i ual se ings. The amewo k is based on how he hea e wo ld e alua es hei audiences. I goes beyond he one-dimensional (engagemen - based) cu en end and speci ies ou dimensions: a ec i e esponse (𝐴𝑓), engagemen (𝐸𝑛), emo ional connec ion (Ec), and lea ning (Le). Besides, we speci ied a pa icula implemen a ion using YOLO 4, a cus om acking algo i hm, and a combina ion o ine- uned MobileNe V3 image embeddings and FC laye s o p edic audience expe ience sco es. We also desc ibed he expe imen we ca ied ou o ob ain a da ase and es ou sys em and p esen ed he inal esul s. On one hand, he p oposed 4-dimensional amewo k cap u es aspec s o he audience expe ience ha we e no conside ed in he one-dimensional measu emen s. An audience may be highly engaged bu all sho in e ms o lea ning. In he same manne , an audience migh no be ully engaged bu s ill ha e a high deg ee o a ec i e esponse. The abili y o cap u e hese sub le ies makes he p oposed amewo k a be e choice han engagemen -based ones o see he ull pic u e when i comes o measu ing audience expe ience. A he same ime, while exis ing a chi ec u es we e designed o be used in-pe son, he p oposed a chi ec u e is designed o i he cha ac e is ics o i ual se ings. E en i we we e o adap p e ious sys ems o wo k coupled wi h ideo con e ence so wa e, we would s ill need o adap he cues hey use o de e mine he deg ee o engagemen so ha hey would be ully unc ional in hese en i onmen s. To mi iga e he limi a ions o popula ideo con e encing so wa e, we c ea ed an ad-hoc con e ence ool. We a e awa e ha he inal esul s a e no conclusi e, bu he p oposed a chi ec u e and ML pipeline a e jus mean as a p oo o concep o es he iabili y o using an in elligen sys em o moni o audience expe ience. While he gene ic building blocks ha cons i u e he objec acking (YOLO 4) and emo ion ecogni ion (MobileNe V3 and FC laye s) sys ems a e eliable and well-p o en in hei espec i e asks, only he acking module pe o med as expec ed. We belie e ha he main obs acle o he emo ion ecogni ion module is he insu icien amoun o da a and i s edundancy. The e o e, he s a ing poin o u u e esea ch should be a la ge scale expe imen o expand he da ase . On he o he hand, he e a e some limi a ions o he p oposed amewo k and sys em ha we also plan o add ess in u u e esea ch: • We selec ed 4 me ics o e alua e audience expe ience based on he li e a u e and p e ious esea ch, bu he e migh be o he dimensions ha esul in be e accu acy. • Misiden i ica ions and missing ames in he da a collec ion p ocess pose a p oblem o he p ope unc ioning o he ML model. • We do no ha e in o ma ion abou how well he amewo k will pe o m wi h a la ge numbe o pa icipan s. • We do no ully comp ehend he ela ionship be ween inpu image ea u es and neu on ac i a ions in he FC laye s. Finally, ano he impo an line o esea ch would be he design o an emo ion ecogni ion a chi ec u e ha conside s empo al dependencies. I is o be expec ed ha , i any o he dimensions o he audience expe ience is posi i e (o nega i e) a a pa icula momen in ime, i will be simila in bo h he p eceding and successi e momen s. Bo h he p oposed amewo k and he implemen a ion o objec acking ha e he ools necessa y o ackle his p oblem (since hey keep ack o speci ic iden i ies o e ime). Howe e , he emo ion ecogni ion module does no conside pas p edic ions. Likely, modi ying he a chi ec u e o inco po a e such p edic ions would esul in a mo e obus and accu a e model. Acknowledgemen s We would like o hank all he olun ee s who helped us build ou da ase and, he e o e, enabled us o wo k on he p edic i e sys em. This p ojec has been ounded by he Minis y o Science, Inno a ion and Uni e si ies o Spain (Didascalias, RTI2018-096401-A-I00). Re e ences Whi ehill, J., Se pell, Z., Lin, Y. C., Fos e , A., & Mo ellan, J. R. (2014). The aces o engagemen : Au oma ic ecogni ion o s uden engagemen om acial exp essions. IEEE T ansac ions on A ec i e Compu ing, 5(1). h ps://doi.o g/10.1109/TAFFC.2014.2316163 Sun, W., Li, Y., Tian, F., Fan, X., & Wang, H. (2019). How P esen e s Pe cei e and Reac o Audience Flow P edic ion In-si u: An explo a i e s udy o li e online lec u es. P oceedings o he ACM on Human- Compu e In e ac ion, 3(CSCW). h ps://doi.o g/10.1145/3359264 Goldbe g, P., Süme , Ö., S ü me , K., Wagne , W., Göllne , R., Ge je s, P., … T au wein, U. (2019). A en i e o No ? Towa d a Machine Lea ning App oach o Assessing S uden s’ Visible Engagemen in Class oom Ins uc ion. Educa ional Psychology Re iew. h ps://doi.o g/10.1007/s10648-019-09514-z Cu is, K., Jones, G. J. F., & Campbell, N. (2015). E ec s o good speaking echniques on audience engagemen . ICMI 2015 - P oceedings o he 2015 ACM In e na ional Con e ence on Mul imodal In e ac ion. h ps://doi.o g/10.1145/2818346.2820766 Bochko skiy, A., Wang, C.-Y., & Liao, H.-Y. M. (2020). YOLO 4: Op imal Speed and Accu acy o Objec De ec ion. A Xi . Re ie ed om h p://a xi .o g/abs/2004.10934 Independen Thea e Council, The Socie y o London Thea e, Thea ical Managemen Associa ion, & The New Economics Founda ion. (2005). Cap u ing he audience expe ience: A handbook o he hea e. h ps://i c- a s-s3.s udiocoucou.com/uploads/helpshee _a achmen / ile/23/Thea e_handbook.pd WebRTC. (2011). [So wa e]. h ps://web c.o g She no , D. J., Csikszen mihalyi, M., Schneide , B., & She no , E. S. (2003, June). S uden engagemen in high school class ooms om he pe spec i e o low heo y. School Psychology Qua e ly, Vol. 18, pp. 158–176. h ps://doi.o g/10.1521/scpq.18.2.158.21860 Cu is, K., Jones, G. J. F., & Campbell, N. (2016). Speake impac on audience comp ehension o academic p esen a ions. ICMI 2016 - P oceedings o he 18 h ACM In e na ional Con e ence on Mul imodal In e ac ion. h ps://doi.o g/10.1145/2993148.2993194 Webs e , J., & Ho, H. (1997). Audience engagemen in mul imedia p esen a ions. ACM SIGMIS Da abase: The DATABASE o Ad ances in In o ma ion Sys ems, 28(2), 63–77. h ps://doi.o g/10.1145/264701.264706 Simonyan, K., & Zisse man, A. (2015). Ve y deep con olu ional ne wo ks o la ge-scale image ecogni ion. 3 d In e na ional Con e ence on Lea ning Rep esen a ions, ICLR 2015 - Con e ence T ack P oceedings. Re ie ed om h p://www. obo s.ox.ac.uk/ He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep esidual lea ning o image ecogni ion. P oceedings o he IEEE Compu e Socie y Con e ence on Compu e Vision and Pa e n Recogni ion, 2016-Decem, 770–778. h ps://doi.o g/10.1109/CVPR.2016.90 Zoph, B., Vasude an, V., Shlens, J., & Le, Q. V. (2017). Lea ning T ans e able A chi ec u es o Scalable Image Recogni ion. P oceedings o he IEEE Compu e Socie y Con e ence on Compu e Vision and Pa e n Recogni ion, 8697–8710. Re ie ed om h p://a xi .o g/abs/1707.07012 Weiss, K., Khoshgo aa , T. M., & Wang, D. D. (2016). A su ey o ans e lea ning. Jou nal o Big Da a, 3(1), 9. h ps://doi.o g/10.1186/s40537-016-0043-6 Oquab, M., Bo ou, L., Lap e , I., & Si ic, J. (2014). Lea ning and ans e ing mid-le el image ep esen a ions using con olu ional neu al ne wo ks. P oceedings o he IEEE Compu e Socie y Con e ence on Compu e Vision and Pa e n Recogni ion, 1717–1724. h ps://doi.o g/10.1109/CVPR.2014.222 Hussain, M., Bi d, J. J., & Fa ia, D. R. (2019). A s udy on CNN ans e lea ning o image classi ica ion. Ad ances in In elligen Sys ems and Compu ing, 840, 191–202. h ps://doi.o g/10.1007/978-3-319-97982- 3_16 Deng, J., Dong, W., Soche , R., Li, L.-J., Kai Li, & Li Fei-Fei. (2010, Ma ch 1). ImageNe : A la ge-scale hie a chical image da abase. 248–255. h ps://doi.o g/10.1109/c p .2009.5206848 LeCun, Y., Ha ne , P., Bo ou, L., & Bengio, Y. (1999). Objec ecogni ion wi h g adien -based lea ning. Lec u e No es in Compu e Science (Including Subse ies Lec u e No es in A i icial In elligence and Lec u e No es in Bioin o ma ics), 1681, 319–345. h ps://doi.o g/10.1007/3-540-46805-6_19 Co es, Co inna (AT&TBellLabs., Hohndel, NJ07733, U., & Vladimi , Vapnik (AT&TBellLabs., Hohndel, NJ07733, U. (1995). Suppo -Vec o Ne wo ks. Machine Lea ning, 297(20), 273–297. B adley, M. M., & Lang, P. J. (1994). Measu ing emo ion: The sel -assessmen manikin and he seman ic di e en ial. Jou nal o Beha io The apy and Expe imen al Psychia y, 25(1), 49–59. h ps://doi.o g/10.1016/0005-7916(94)90063-9 McKinney, J. D., Mason, J., Pe ke son, K., & Cli o d, M. (1975). Rela ionship be ween class oom beha io and academic achie emen . Jou nal o Educa ional Psychology, 67(2), 198–203. h ps://doi.o g/10.1037/h0077012 Lei, H., Cui, Y., & Zhou, W. (2018). Rela ionships be ween s uden engagemen and academic achie emen : A me a-analysis. Social Beha io and Pe sonali y, 46(3), 517–528. h ps://doi.o g/10.2224/sbp.7054 Pian a, R. C., & Ham e, B. K. (2009). Concep ualiza ion, Measu emen , and Imp o emen o Class oom P ocesses: S anda dized Obse a ion Can Le e age Capaci y. Educa ional Resea che , 38(2), 109–119. h ps://doi.o g/10.3102/0013189X09332374 Kalman, R. E. (1960). A new app oach o linea il e ing and p edic ion p oblems. Jou nal o Fluids Enginee ing, T ansac ions o he ASME, 82(1), 35–45. h ps://doi.o g/10.1115/1.3662552 Kuhn, H. W. (1955). The Hunga ian me hod o he assignmen p oblem. Na al Resea ch Logis ics Qua e ly, 2(1–2), 83–97. h ps://doi.o g/10.1002/na .3800020109 Bewley, A., Ge, Z., O , L., Ramos, F., & Upc o , B. (2016). Simple online and eal ime acking. P oceedings - In e na ional Con e ence on Image P ocessing, ICIP, 2016-Augus , 3464–3468. h ps://doi.o g/10.1109/ICIP.2016.7533003 A un Kuma , N. P., Laxmanan, R., Ram Kuma , S., S inidh, V., & Ramana han, R. (2021). Pe o mance S udy o Mul i- a ge T acking Using Kalman Fil e and Hunga ian Algo i hm. Communica ions in Compu e and In o ma ion Science, 1364, 213–227. h ps://doi.o g/10.1007/978-981-16-0422-5_15 Munk es, J. (1957). Algo i hms o he Assignmen and T anspo a ion P oblems. Jou nal o he Socie y o Indus ial and Applied Ma hema ics, 5(1), 32–38. h ps://doi.o g/10.1137/0105003 Luo, W., Xing, J., Milan, A., Zhang, X., Liu, W., & Kim, T. K. (2021). Mul iple objec acking: A li e a u e e iew. A i icial In elligence, 293. h ps://doi.o g/10.1016/j.a in .2020.103448 Godbehe e, A. B., & Goldbe g, K. (2014). Algo i hms o isual acking o isi o s unde a iable-ligh ing condi ions o a esponsi e audio a ins alla ion. Con ols and A : Inqui ies a he In e sec ion o he Subjec i e and he Objec i e, 181–204. h ps://doi.o g/10.1007/978-3-319-03904-6_8 Oh, A. R., Lee, J., Lee, J. S., Moon, S. W., Nam, D. W., & Yoo, W. (2020). Mul i-objec acking sys em using dissimila appa a us in ideo sequence. In e na ional Con e ence on ICT Con e gence, 2020-Oc obe , 1528– 1530. h ps://doi.o g/10.1109/ICTC49870.2020.9289264 Gale, D., & Shapley, L. S. (1962). College Admissions and he S abili y o Ma iage. The Ame ican Ma hema ical Mon hly, 69(1), 9. h ps://doi.o g/10.2307/2312726 Howa d, A., Sandle , M., Chu, G., Chen, L. C., Chen, B., Tan, M., … Adam, H. (2019). Sea ching o MobileNe V3. A Xi . Kingma, D. P., & Ba, J. L. (2015). Adam: A me hod o s ochas ic op imiza ion. 3 d In e na ional Con e ence on Lea ning Rep esen a ions, ICLR 2015 - Con e ence T ack P oceedings. 94 APPENDIX B. PAPER AND SUBMISSION CONFIRMATION Glossa y A A ec i e e sponse. 26 BoF Bag o F eebies. 36 BoS Bag o Specials. 36 CIoU Comple e-IoU. 36 CmBN C oss mini-ba chno maliza ion. 36 CNN Con olu ional Neu al Ne wo k. 7 CSP C oss-S age Pa ial connec ions. 36 DIoU-NMS Dis ance-IoU loss wi h NMS. 36 Ec Emo ional Connec ion. 26 En Engagemen . 26 ERT Ensemble Reg ession T ee. 23 ERVF Emo ion Recogni ion om Video Feeds. 34 FAU Facial Ac ion Uni . 22 FC Fully Connec ed. 9 FPN Fea u e Py amid Ne wo k. 37 FPS F ames Pe Second. 15 FVAE Fac o ized VAE. 23 GS Gale-Shapley. 41 IoU In e sec ion o e Union. 10 Le Lea ning. 26 LSTM Long Sho -Te m Memo y. 15 MAE Mean Absolu e E o . 60 95 96 Glossa y MAPE Mean Absolu e Pe cen age E o . 60 MiWRC Mul i-inpu Weigh ed Residual Connec ions. 36 ML Machine Lea ning. 1 MLP Mul iLaye Pe cep on. 18 MOT Mul iple Objec T acking. 7 MSE Mean Squa ed E o . 15 NAS Ne wo k A chi ec u e Sea ch. 51 NMS Non-Max Supp ession. 13 PAN Pa h Agg ega ion Ne wo k. 36 PSO Pa icle Swa m Op imiza ion. 14 R-CNN Regions wi h CNN ea u es. 8 ReLU Rec i ied Linea Uni . 9 RNN Recu en Neu al Ne wo k. 15 RoI Region o In e es . 10 ROLO Recu en YOLO. 15 RPN Region P oposal Ne wo k. 11 SAE Sum o Absolu e E o s. 50 SAM Spa ial A en ion Module. 39 SAT Sel -Ad e sa ial T aining. 36 SORT Simple Online and Real ime T acking. 15 SPP Spa ial Py amid Pooling. 39 SVM Suppo Vec o Machine. 8, 17 VAE Va ia ional Au oEncode . 23 YOLO You Only Look Once. 12 Bibliog aphy [1] Kelson R.T. Ai es, And e M. San ana, and Adela do A.D. Medei os. “Op ical low using colo in o ma ion: P elimina y esul s”. In: P oceedings o he ACM Sym- posium on Applied Compu ing (2008), pp. 1607–1611. doi:10.1145/1363686. 1364064. [2] Ahmad Ali e al. “Visual objec acking—classical and con empo a y app oaches”. In: F on ie s o Compu e Science 10.1 (2016), pp. 167–188. issn: 20952236. doi: 10.1007/s11704-015-4246-3. [3] Amidi, A shine and Amidi, She ine. A de ailed example o how o gene a e you da a in pa allel wi h PyTo ch.u l:h ps://s an o d.edu/~she ine/blog/ py o ch-how- o-gene a e-da a-pa allel#. [4] Ligia Ba inca e al. Cice o - Towa ds a mul imodal i ual audience pla o m o public speaking aining. Tech. ep. 2013, pp. 116–128. doi:10.1007/978-3-642- 40415-3_10.u l:h p://www. oas mas e s.o g/ ips.asp. [5] He be Bay, Tinne Tuy elaa s, and Luc Van Gool. “SURF: Speeded Up Robus Fea u es”. In: Compu e Vision – ECCV 2006. Ed. by Aleš Leona dis, Ho s Bischo , and Axel Pinz. Be lin, Heidelbe g: Sp inge Be lin Heidelbe g, 2006, pp. 404–417. isbn: 978-3-540-33833-8. [6] S. S. Beauchemin and J. L. Ba on. “The Compu a ion o Op ical Flow”. In: ACM Compu ing Su eys (CSUR) 27.3 (1995), pp. 433–466. issn: 15577341. doi:10. 1145/212094.212141. [7] Alex Bewley e al. “Simple online and eal ime acking”. In: P oceedings - In e na- ional Con e ence on Image P ocessing, ICIP 2016-Augus (2016), pp. 3464–3468. issn: 15224880. doi:10.1109/ICIP.2016.7533003. a Xi : 1602.00763. [8] E ik Blasch e al. “O e iew o con ex ual acking app oaches in in o ma ion u- sion”. In: Geospa ial In oFusion III 8747 (2013), 87470B. issn: 0277786X. doi: 10.1117/12.2016312. [9] Alexey Bochko skiy, Chien-Yao Wang, and Hong-Yuan Ma k Liao. YOLO 4: Op- imal Speed and Accu acy o Objec De ec ion. 2020. a Xi : 2004.10934 [cs.CV]. [10] Ma ga e M. B adley and Pe e J. Lang. “Measu ing emo ion: The sel -assessmen manikin and he seman ic di e en ial”. In: Jou nal o Beha io The apy and Ex- pe imen al Psychia y 25.1 (Ma . 1994), pp. 49–59. issn: 00057916. doi:10.1016/ 0005-7916(94)90063-9. [11] Robe o B unelli. Templa e Ma ching Techniques in Compu e Vision: Theo y and P ac ice. 2009, pp. 1–338. isbn: 9780470517062. doi:10.1002/9780470744055. 97