scieee Open visual document viewer

Resource-Efficient GRU for Real-Time Gesture Recognition

Kaseris, Michail; Malassiotis, Sotiris; Kostavelis, Ioannis

Abstract

Time-of-flight sensors deliver distance measurements through light propagation timing, facilitating object detection, gesture recognition, and three-dimensional mapping. These sensors maintain functionality across lighting conditions while ensuring response rates required for real-time applications. Based on these properties, combined with their cost and energy consumption parameters, we propose their exploitation in gesture recognition systems for human-robot interaction (HRI). This work presents two contributions: First, we introduce a dataset for gesture recognition tasks. Second, we develop a model for sequence classification optimized for resource constraints. The experimental results demonstrate the feasibility of the approach through quantitative evaluation metrics. Finally, we present classification results on the dataset, analyze the model's performance on gesture categories, and discuss future challenges and directions in this field of research.

Full text

Resou ce-E icien GRU o Real-Time Ges u e Recogni ion Michail Kase is* In e na ional Hellenic Uni e si y Ka e ini, G eece [email p o ec ed] *Co esponding au ho So i is Malassio is In o ma ion Technologies Ins i u e CERTH Thessaloniki, G eece [email p o ec ed] Ioannis Kos a elis In e na ional Hellenic Uni e si y Ka e ini, G eece gkos a[email p o ec ed] Abs ac —Time-o - ligh senso s deli e dis ance measu e- men s h ough ligh p opaga ion iming, acili a ing objec de- ec ion, ges u e ecogni ion, and h ee-dimensional mapping. These senso s main ain unc ionali y ac oss ligh ing condi ions while ensu ing esponse a es equi ed o eal- ime applica ions. Based on hese p ope ies, combined wi h hei cos and ene gy consump ion pa ame e s, we p opose hei exploi a ion in ges u e ecogni ion sys ems o human- obo in e ac ion (HRI). This wo k p esen s wo con ibu ions: Fi s , we in oduce a da ase o ges u e ecogni ion asks. Second, we de elop a model o sequence classi ica ion op imized o esou ce cons ain s. The expe imen al esul s demons a e he easibili y o he app oach h ough quan i a i e e alua ion me ics. Finally, we p esen clas- si ica ion esul s on he da ase , analyze he model’s pe o mance on ges u e ca ego ies, and discuss u u e challenges and di ec ions in his ield o esea ch. Index Te ms—Time-o -Fligh Senso s, Ges u e Recogni ion, Human-Robo In e ac ion I. INTRODUCTION Hand ges u e ecogni ion has expanded beyond e e yday applica ions in o di e se ields including i ual eali y, med- ical sys ems, educa ion, and au omo i e in e aces. Among he h ee main ecogni ion echnologies—da a glo e-based, ision-based, and ada -based—da a glo es o e p ecise ac- ile sensing bu p esen limi a ions in p ac icali y. Despi e hei ad an age in no equi ing ges u e ex ac ion om backg ound scenes, da a glo es ace adop ion ba ie s due o hei bulk, high cos s, and calib a ion equi emen s, making hem less widely applicable han ision-based al e na i es [1], [2]. Time-o -Fligh (ToF) came as p oduce dep h images by encoding scene dis ances in indi idual pixels h ough phase- delay measu emen s o e lec ed in a ed ligh . This echnol- ogy enables di ec 3D s uc u e es ima ion while exhibi ing obus ness o adiome ic, geome ic, and illumina ion ac o s. O he bene i s include compac design, CMOS-based cos - e ec i eness, high ame a es, and measu emen accu acy. ToF came as ha e p o en aluable ac oss di e se applica ions including obo na iga ion, 3D econs uc ion, and human- machine in e ac ion, o e ing a complemen a y app oach o This esea ch was unded by he Eu opean Union’s Ho izon Eu ope P ojec “Ses osenso” (h p://ses osenso.eu/ (accessed on 27 No embe 2024), HORIZON-CL4-Digi al-Eme ging) unde G an 101070310. o he dep h-sensing me hods such as s uc u ed-ligh sys ems [3]. In he con ex o human compu e in e ac ion, ges u e ecogni ion by means o 3D senso s (such as Kinec ) has been al eady in es iga ed [9], also in obo ics applica ions [10]. The main downside o such sys ems is ha hey equi e signi ican compu a ional esou ces e.g. o human skele on ex ac ion o hand pose es ima ion. They also equi e mos o he use body o be isible by he came a. Thus hey a e sui able in use cases whe e a ixed came a is looking di ec ly owa ds a use wi h limi ed mo emen and minimal occlusions. On he con a y ou applica ion a ge a e use cases such as collabo a i e assembly in au omo i e indus y whe e ixing he came a is no possible e.g. due o signi ican occlusion o he wo kspace and he wo ke s mo emen is signi ican . In such use cases we a e expe imen ing wi h wo solu ions one based on a ne wo k o body wa n IMU senso s and he second, p esen ed in his pape , based on a ing o en low-cos /low- esolu ion ToF senso s in a mul i- iew se up in eg a ed on he collabo a i e obo manipula o . By exploi ing mul iple senso s wi h di e en iewpoin s we achie e obus ness wi h espec o obo and use mo emen . In addi ion he p oposed app oach has e y low la ency and compu a ional equi emen s while achie ing high accu acy wi h a ela i ely small numbe o ges u es despi e he ac ha he senso s ha e a e y low esolu ion (8×8). Use eedback may also be con enien ly p o ided by means o colo ed LEDs o audio in eg a ed on he ing. In summa y his wo k con ibu ion o he ield o human- obo in e ac ion a e: •Demons a ion and e alua ion o ges u e ecogni ion based on a no el mul i- iew low esolu ion ToF senso se up. •Cap u ing a new da ase o se e al ges u e sequences cap u ed by he a o men ioned ToF senso s a ay. •Building upon his da ase , we p opose a esou ce- e icien classi ica ion model o sequence classi ica ion ha add esses compu a ional cons ain s. 66 2025 11 h In e na ional Con e ence on Au oma ion, Robo ics, and Applica ions 979-8-3315-0923-1/25/$31.00 ©2025 IEEE 2025 11 h In e na ional Con e ence on Au oma ion, Robo ics, and Applica ions (ICARA) | 979-8-3315-0923-1/25/$31.00 ©2025 IEEE | DOI: 10.1109/ICARA64554.2025.10977711 Au ho ized licensed use limi ed o: Uni e sidad de Za agoza. Downloaded on Sep embe 30,2025 a 11:10:41 UTC om IEEE Xplo e. Res ic ions apply. II. RELATED WORK The wo k p oposed by [5] p esen s a ges u e ecogni ion sys em o UAV con ol using a da a ep esen a ion model ha ans o ms 4D spa io empo al da a in o 2D ma ices and 1D a ays. Thei sys em p ocesses skele on da a om a Leap Mo ion Con olle and employs h ee neu al ne wo k a chi ec u es: a 2-laye ully connec ed ne wo k, a 5-laye ully connec ed ne wo k, and an 8-laye con olu ional ne wo k. The app oach was alida ed bo h in simula ion and on physical d one pla o ms, demons a ing he easibili y o ges u e- based d one con ol h ough hei p oposed da a ep esen a- ion model. [6] add esses eal- ime pe o mance limi a ions in UAV con ol h ough a hyb id sys em combining IMU and ision-based ges u e ecogni ion. Thei app oach uses a humb-moun ed IMU o con inuous commands and ision- based de ec ion o disc e e inpu s, o e coming challenges in dynamic ges u e ecogni ion ha ypically a ec pu ely ision-based sys ems. The amewo k demons a es imp o ed pe o mance o e adi ional con ol in e aces in simula ion en i onmen s. [7] p opose a Poin Ne -based deep neu al ne - wo k o hand ges u e ecogni ion using ToF senso da a. Thei me hod includes a mul is age hand segmen a ion p ocess and demons a es ha 3D poin cloud p ocessing ou pe o ms 2D app oaches. The wo k includes a cus om da ase c ea ion and compa a i e analysis be ween 2D and 3D me hodologies. [12] de eloped a eal- ime ges u e ecogni ion sys em u ilizing ime-o - ligh came a da a o Windows applica ion con ol. Thei app oach combines mo phological analysis o hand silhoue es wi h mo ion pa e n es ima ion o ecognize s a ic and dynamic ges u es, demons a ing an e ec i e al e na i e o con en ional in e aces. III. PROPOSED APPROACH A. Ha dwa e Se up The expe imen al se up employs a Time-o -Fligh (ToF) sen- so ing a ay comp ising 10 ToF came as [11], s a egically posi ioned a a ying dis ances (20cm o 1m) om he subjec , as shown in igu e 1. This dis ance ange was speci ically chosen o e lec ealis ic human- obo collabo a ion scena ios, whe e ope a o s ypically wo k in close p oximi y o he obo du ing ask execu ion. The came as a e a anged in a uni o m ci cula dis ibu ion wi hin a cus om-designed chassis, wi h each senso posi ioned a equal in e als o 36 deg ees (360°/10) o ensu e comple e co e age o he su ounding space, as depic ed in igu e 2. This 360-deg ee ecep i e ield is essen ial as he ope a o ’s posi ion ela i e o he obo is no ixed and can a y du ing ope a ion. Each ToF senso cap u es dep h da a a a esolu ion o 8x8 pixels, c ea ing a compac ye in o ma i e ep esen a ion o he ges u e space. The synch oniza ion o he mul i-came a sys em is handled h ough a bu e -based app oach: he sys em accumula es ames om all senso s, and once a comple e se is ecei ed, a uni ied imes amp is assigned o he en i e ame collec ion by he publishe be o e p oceeding o he nex cap u e cycle. This con igu a ion enables comp ehensi e cap u e o ges u e da a om mul iple iewpoin s simul aneously. Fig. 1. Ou expe imen al se up. Time-o -Fligh Senso Fig. 2. A op iew o he senso a ay. Each small ci cle ep esen s a Time- o -Fligh senso . All senso s a e dis ibu ed equally on he ci cum e ence o he ci cula base. The ed iangles ep esen he ield o iew o each senso . A ges u e can ake place in on o any ime-o - ligh senso . B. Da ase Acquisi ion and P ep ocessing The ges u e ecogni ion sys em employs a deep lea ning app oach o p ocess sequen ial ges u e da a, implemen ing and compa ing bo h Long Sho -Te m Memo y (LSTM) and Ga ed Recu en Uni (GRU) neu al ne wo k a chi ec u es. Each ame o inpu consis s o a 640-dimensional ea u e ec o , de i ed di ec ly om he aw senso a ay da a (10 senso s × 8×8 pixels). The ne wo k a chi ec u e consis s o a single-laye ecu en ne wo k wi h 128 hidden uni s, connec ed o a linea ou pu laye ha maps o ges u e classes. The hidden uni size o 128 was empi ically de e mined du ing ini ial expe imen s and demons a ed s able pe o mance ac oss di e en ges u e pa e ns. 67 Au ho ized licensed use limi ed o: Uni e sidad de Za agoza. Downloaded on Sep embe 30,2025 a 11:10:41 UTC om IEEE Xplo e. Res ic ions apply. Da a p ep ocessing ollows a s aigh o wa d app oach: aw senso eadings a e escaled om hei o iginal ange [0, 4000] o [0, 1] h ough min-max no maliza ion. To acili a e e icien ba ch p ocessing, sequences a e ze o-padded o a uni o m leng h, ensu ing consis en inpu dimensions ac oss all samples. C. Neu al Ne wo k A chi ec u e The ne wo k a chi ec u e consis s o a Ga ed Recu en Uni (GRU) [8] ollowed by a linea classi ica ion laye . The inpu o he ne wo k is a sequence o 640-dimensional ec o s, whe e each ec o ep esen s he conca ena ed eadings om all ToF senso s a a single ime s ep. These sequences ha e a iable leng hs depending on he du a ion o he pe o med ges u e. To handle a iable-leng h sequences e icien ly, he inpu is i s p ocessed using packed sequences, which allows he GRU o ope a e only on ac ual da a poin s while igno ing padding. The GRU p ocesses hese sequences wi h hhidden uni s ac oss llaye s, main aining a hidden s a e ha cap u es empo al dependencies in he ges u e da a. The ini ial hidden s a e h0is ini ialized as a ze o ec o . Fo classi ica ion, he ne wo k u ilizes only he inal ele an ou pu o each sequence (co esponding o he las ac ual ime s ep, excluding padding). This ou pu is hen passed h ough a linea laye ha maps he h-dimensional hidden ep esen a ion o a p obabili y dis ibu ion o e he ges u e classes. Fo mally, o an inpu sequence (X= (x1, ..., xT)) whe e (x ∈R640), he ne wo k compu es: h =GRU(x , h −1)(1) y=W hT+b(2) whe e h ep esen s he hidden s a e a ime ,Wand ba e he weigh s and bias o he inal linea laye , and yis he p obabili y ec o used o classi ica ion. IV. EXPERIMENTS A. Da ase A da ase o hand ges u es was collec ed using a Time-o - Fligh senso a ay. Two pa icipan s pe o med h ee dis inc ges u es: ”come he e,” ”s op,” and ”spin.” To in oduce a i- abili y in he da a collec ion, he ges u es we e execu ed om bo h s anding and sea ed posi ions a a ying dis ances om he senso a ay. The da ase includes a ia ions whe e pa - icipan s al e na ed be ween hei dominan and non-dominan hands, as well as ins ances o wo-handed ges u e execu ion. In o al, 80 sequen ial hand ges u e samples we e eco ded, p o iding a di e se collec ion o spa io empo al da a o ges u e ecogni ion analysis. B. Implemen a ion De ails The expe imen al implemen a ion u ilized PyTo ch o model de elopmen and aining. The GRU-based classi ie was ained o 50 epochs wi h a ba ch size o 32 and a lea ning a e o 0.001 using he Adam op imize . The model a chi ec u e consis ed o wo GRU laye s wi h 128 hidden uni s, ollowed by a ully connec ed laye o classi ica ion. The inpu dimension was se o ma ch he conca ena ed ToF senso da a (64 dimensions ×10 senso s), and he ou pu dimension co esponded o he h ee ges u e classes: ”come he e,” ”spin,” and ”s op.” The da ase was sys ema ically di ided in o aining and es ing se s, wi h h ee samples pe ges u e class ese ed o es ing o ensu e a balanced e alua ion. To handle a iable-leng h sequences e icien ly, we implemen ed a cus om colla e unc ion ha so ed sequences by leng h and applied padding whe e necessa y. This app oach op imized ba ch p ocessing while p ese ing he empo al cha ac e is ics o he ges u e da a. C. T aining P o ocol The aining p ocess employed c oss-en opy loss o classi- ica ion and inco po a ed dynamic sequence packing o accom- moda e a iable-leng h ges u es. We moni o ed bo h aining loss and alida ion accu acy h oughou he aining p ocess, sa ing model checkpoin s based on bes alida ion pe o - mance. The ba ch- i s da a o ma was u ilized o op imize memo y usage and compu a ional e iciency du ing aining. D. Resul s The expe imen al esul s demons a ed he e ec i eness o ou app oach in ecognizing he h ee dis inc ges u e classes. The model achie ed consis en pe o mance ac oss di e en ges u e ypes, wi h a ia ions in ecogni ion accu acy co e- la ed wi h ges u e complexi y. The con usion ma ix isual- iza ion e ealed pa e ns in classi ica ion e o s, p o iding in- sigh s in o po en ial a eas o imp o emen in bo h da a collec- ion and model a chi ec u e, as illus a ed in igu e ??. T aining me ics showed s eady con e gence, wi h he loss unc ion dec easing mono onically and alida ion accu acy s abilizing a e app oxima ely 30 epochs. This beha io sugges ed ha he chosen hype pa ame e s and model a chi ec u e we e well- sui ed o he ges u e ecogni ion ask. The inal model achie ed obus pe o mance on he es se , demons a ing i s capabili y o gene alize o unseen ges u e ins ances while main aining eal- ime p ocessing capabili ies necessa y o human- obo in e ac ion applica ions. E. Model A chi ec u e E alua ion We pe o med a sys ema ic e alua ion o a ious neu al ne wo k a chi ec u es o de e mine he op imal con igu a ion o ges u e ecogni ion. The analysis examined he e ec s o ecu en uni ype (GRU, LSTM), hidden s a e dimensionali y (64, 128, 256), ne wo k dep h (1-3 laye s), and egula iza ion s eng h (d opou a es: 0.0, 0.2, 0.4). Figu e 3 p esen s he quan i a i e esul s o his in es iga ion. 1) Pe o mance Analysis: The empi ical esul s indica e supe io pe o mance o GRU-based a chi ec u es o e LSTM a ian s, wi h mean alida ion accu acy o 88-90% and 67% espec i ely. GRU con igu a ions also exhibi ed educed a i- ance ac oss expe imen al condi ions, indica ing enhanced s a- bili y. Analysis o hidden s a e dimensionali y e ealed a 68 Au ho ized licensed use limi ed o: Uni e sidad de Za agoza. Downloaded on Sep embe 30,2025 a 11:10:41 UTC om IEEE Xplo e. Res ic ions apply. Fig. 3. Hype pa ame e pe o mance analysis o he classi ie . TABLE I PERFORMANCE ANALYSIS OF GRU CLASSIFIER Sequence CPU Pe o mance GPU Pe o mance Leng h Time (ms) Memo y (MB) Time (ms) Memo y (MB) 50 2.86 0.89 0.70 0.85 100 5.08 0.93 1.08 8.98 150 7.43 0.96 1.68 8.98 200 9.95 1.00 1.97 8.98 250 12.42 1.04 2.45 8.98 300 14.87 1.07 2.88 8.98 350 17.17 1.11 3.32 8.98 400 19.36 1.15 3.87 8.98 posi i e co ela ion wi h model pe o mance, whe e ep e- sen a ions o dimension 256 achie ed op imal accu acy wi h minimal a iance. The in es iga ion o a chi ec u al dep h p o- duced an unexpec ed inding: single-laye con igu a ions sys- ema ically ou pe o med deepe a chi ec u es in bo h accu acy and consis ency. This obse a ion sugges s ha he empo al dynamics o ges u e ecogni ion may be su icien ly cap u ed by a single ecu en laye . Fu he mo e, he applica ion o d opou egula iza ion demons a ed an in e se ela ionship wi h model pe o mance. The absence o d opou yielded su- pe io esul s, indica ing ha he cons ain o ne wo k capaci y may be de imen al o his pa icula ecogni ion ask. As shown in Table I, he model demons a es e icien scaling wi h sequence leng h, wi h GPU in e ence ime inc easing linea ly om 0.70ms a 50 ames o 3.87ms a 400 ames, while main aining consis en memo y usage o app oxima ely 9MB. CPU in e ence exhibi s simila linea scaling bu a highe absolu e imes, anging om 2.86ms o 19.36ms. These esul s indica e ha eal- ime ges u e ecogni ion is easible on bo h pla o ms, wi h GPU accele a ion p o iding a 4-5x speedup while main aining minimal memo y o e head. Th ough his sys ema ic e alua ion, we de e mined ha op imal pe o - mance is achie ed using a single-laye GRU a chi ec u e wi h hidden s a e dimension 256 and no d opou egula iza ion. This con igu a ion demons a es ha a chi ec u al simplici y, combined wi h su icien ep esen a ional capaci y, p o ides obus pe o mance o ges u e ecogni ion. The esul s sugges ha he empo al s uc u e o ges u e da a may be e ec i ely modeled wi hou equi ing deep a chi ec u al hie a chies. V. CONCLUSION This wo k p esen ed a ges u e ecogni ion sys em using Time-o -Fligh senso s o human- obo in e ac ion applica- ions. Th ough expe imen al e alua ion, we demons a ed he e ec i eness o a ci cula ToF senso a ay con igu a ion coupled wi h a single-laye GRU a chi ec u e, achie ing 88- 90% accu acy ac oss h ee ges u e classes. Ou pe o mance analysis showed ha he sys em can p ocess ges u e sequences e icien ly on bo h CPU and GPU pla o ms, wi h in e ence imes scaling linea ly wi h sequence leng h. Key indings include he supe io i y o GRU o e LSTM a chi ec u es and he unexpec ed e ec i eness o single-laye con igu a ions. Fu u e wo k should ocus on expanding he ges u e da ase and alida ing he sys em in eal-wo ld human- obo in e ac ion scena ios. This wo k con ibu es o he ield by demons a ing ha esou ce-e icien a chi ec u es can achie e obus pe o - mance when combined wi h s a egically posi ioned Time-o - Fligh senso s. REFERENCES [1] Sua ez, J., Mu phy, R. R. (2012, Sep embe ). “Hand ges u e ecogni ion wi h dep h images: A e iew”. In 2012 IEEE RO-MAN: he 21s IEEE in e na ional symposium on obo and human in e ac i e communica ion (pp. 411-417). IEEE. [2] Khan, R. Z., Ib aheem, N. A. (2012). “Hand ges u e ecogni ion: a li e a u e e iew”. In e na ional jou nal o a i icial In elligence & Applica ions, 3(4), 161. [3] Hansa d M., Lee S., Choi O., Ho aud R. “Time o Fligh Came as: P inciples, Me hods, and Applica ions.” Sp inge , pp.95, 2012, Sp inge - B ie s in Compu e Science, ISBN 978-1-4471-4658- 2. 10.1007/978- 1-4471-4658-2 [4] Li, L. “Time-o - ligh came a—an in oduc ion.” Technical whi e pape SLOA190B (2014). [5] Hu, B., Wang, J. (2020). “Deep lea ning based hand ges u e ecogni ion and UAV ligh con ols.” In e na ional Jou nal o Au oma ion and Compu ing, 17(1), 17-29. [6] Yoo, M., Na, Y., Song, H., Kim, G., Yun, J., Kim, S., Jo, K. (2022). Mo ion es ima ion and hand ges u e ecogni ion-based human–UAV in e ac ion app oach in eal ime. Senso s, 22(7), 2513. [7] Mi su, R., Simion, G., Caleanu, C. D., Pop-Calimanu, I. M. (2020). “A poin ne -based solu ion o 3D hand ges u e ecogni ion.” Senso s, 20(11), 3226. [8] Cho, K. (2014). “Lea ning ph ase ep esen a ions using RNN encode -decode o s a is ical machine ansla ion”. a Xi p ep in a Xi :1406.1078. [9] S. Malassio is, N. Ai an i and M. G. S in zis, ”A ges u e ecogni ion sys em using 3D da a,” P oceedings. Fi s In e na ional Symposium on 3D Da a P ocessing Visualiza ion and T ansmission, Padua, I aly, 2002, pp. 190-193 [10] Alexee , Alexande e al. “An O e iew o Kinec Based Ges u e Recog- ni ion Me hods.” P oceedings o In e na ional Con e ence on A i icial Li e and Robo ics (2024) [11] VL53L8CX - Low-powe high-pe o mance 8x8 mul i- zone Time-o -Fligh senso (ToF) h ps://www.s .com/en/ imaging-and-pho onics-solu ions/ l53l8cx.h ml [12] Molina, J., Escude o-Vi˜ nolo, M., Signo iello, A., Pa d` as, M., Fe ´ an, C., Besc´ os, J., ... & Ma ´ ınez, J. M. (2013). Real- ime use independen hand ges u e ecogni ion om ime-o - ligh came a ideo using s a ic and dynamic models. Machine ision and applica ions, 24, 187-204. 69 Au ho ized licensed use limi ed o: Uni e sidad de Za agoza. Downloaded on Sep embe 30,2025 a 11:10:41 UTC om IEEE Xplo e. Res ic ions apply.