scieee Open visual document viewer

CNN Architecture for Posture Classification on Small Data

Hečková, Michaela; Mihálik, Ondrej; Jirgl, Miroslav

Abstract

A convolutional neural network is often mentioned as one of the deep learning methods that requires a large amount of training data. Questioning this belief, this paper explores the applicability of classification based on a shallow net structure trained on a small data set in the~context of patient posture classification based on data from a pressure mattress. Designing a CNN often presents a complex problem, especially without a universally applicable approach, allowing many diverse structural possibilities and training settings. We tested various training options and layer configurations to provide an overview of influential parameters for posture classification. Experiments show encouraging results with the leave-one-out cross-validation accuracy of 93.1% of one of the evaluated CNN structures and its hyperparameter settings.

Full text

IFAC Pape sOnLine 58-9 (2024) 299–304 ScienceDi ec ScienceDi ec A ailable online a www.sciencedi ec .com 2405-8963 Copy igh © 2024 The Au ho s. This is an open access a icle unde he CC BY-NC-ND license . Pee e iew unde esponsibili y o In e na ional Fede a ion o Au oma ic Con ol. 10.1016/j.i acol.2024.07.413 10.1016/j.i acol.2024.07.413 2405-8963 CNN A chi ec u e o Pos u e Classi ica ion on Small Da a Michaela Mes´a oˇso ´a ∗Ond ej Mih´alik ∗Mi osla Ji gl ∗ ∗Depa men o Con ol and Ins umen a ion, Facul y o Elec ical Enginee ing, B no Uni e si y o Technology, B no, Czech epublic (e-mail: Michaela.Mesa oso a@ u .cz). Abs ac : A con olu ional neu al ne wo k is o en men ioned as one o he deep lea ning me hods ha equi es a la ge amoun o aining da a. Ques ioning his belie , his pape explo es he applicabili y o classi ica ion based on a shallow ne s uc u e ained on a small da a se in he con ex o pa ien pos u e classi ica ion based on da a om a p essu e ma ess. Designing a CNN o en p esen s a complex p oblem, especially wi hou a uni e sally applicable app oach, allowing many di e se s uc u al possibili ies and aining se ings. We es ed a ious aining op ions and laye con igu a ions o p o ide an o e iew o in luen ial pa ame e s o pos u e classi ica ion. Expe imen s show encou aging esul s wi h he lea e-one-ou c oss- alida ion accu acy o 93.1% o one o he e alua ed CNN s uc u es and i s hype pa ame e se ings. Keywo ds: CNN, ine uning, ne wo k s uc u e, op imiza ion, pos u e classi ica ion 1. INTRODUCTION Deep lea ning me hods bene i g ea ly om e y la ge da a se s. Howe e , in some si ua ions, i migh be di icul o ga he enough da a, i he da a-acquisi ion p ocess is ime-consuming o labo ious. Choosing he s uc u e and pa ame e s o a CNN ha ep esen s images and hei s o ed in o ma ion uniquely and co ec ly is one o he mos impo an aspec s o achie ing a high CNN pe - o mance. Typically he CNN con igu a ion pa ame e s a e di ided in o wo ca ego ies: hose conce ning he a - chi ec u e and hose conce ning he aining p ocess, see Fe ey a-Rami ez e al. (2019). The speci ic ini ializa ion o he pa ame e s o en has a signi ican impac on how long i akes o he aining p ocess o ind a solu ion and on he gene aliza ion abili y o he esul ing ained ne - wo k. Reg e ully, as s a ed in Bishop and Bishop (2023), he e is no much heo y o help de e mine an ini ializa ion s a egy. Fo his eason, we o e a compac o e iew o a selec ion o se ings and modi iable pa ame e s and hei impac on he classi ica ion esul s. 2. NETWORK TRAINING OPTIONS The igh design o he ne wo k is c ucial o he applica- ion whe e he CNN is o be applied. Howe e he same im- po ance lies in he ask o disco e ing app op ia e aining pa ame e s ha lead o he bes -pe o ming classi ie . We in oduce some o hese pa ame e s along wi h an o e iew o hei ole in he aining p ocess. 2.1 Op imiza ion algo i hms In MATLAB 2023, he e a e a ailable he ollowing ou op imiza ion algo i hms: • SGDM, •RMSP op, •Adam, •L-BGFS Al hough he e exis many mo e algo i hms, hey a e o en de i a i es o hose lis ed abo e, hence we will lea e i o he eade o e alua e hei pe o mance a e he mos sui able base app oach was selec ed. Only he i s h ee will be aken in o conside a ion in his pape . The S ochas ic G adien Descen wi h Momen um o SGDM is an ex ension o he S ochas ic g adien descen (SGD) algo i hm ha upda es he ne wo k pa ame e s o minimize he loss unc ion by aking small s eps a each i e a ion in he di ec ion o he nega i e g adien o he loss, θℓ+1 =θℓ−α∇E(θℓ)+γ(θℓ−θℓ−1),(1) whe e αis he lea ning a e,γis he momen um alue, ℓis he i e a ion numbe , θis he pa ame e ec o , and E(θ) is he loss unc ion. The g adien is e alua ed using a subse o he aining da a, also called a mini-ba ch, whe e a di e en subse is used a each i e a ion, o de ailed explana ion see Bishop and Bishop (2023) and Beale e al. (2023). The en i e pass o e he aining da a is called an epoch. The con ibu ion o he p e ious g adien s ep o he cu en i e a ion is de e mined by he lea ning a e αand he momen um e m γ, which in addi ion educes he oscilla ion a ound he op imum ha migh occu in he case o a e y s eep nega i e g adien . The unpublished Roo Mean Squa e P opaga ion o RM- SP op is designed o accele a e he op imiza ion p ocess, and hus, dec ease he numbe o unc ion e alua ions o each he op imum. This is achie ed by using di e gen lea ning a es o each weigh ha a e able o adap au oma ically o he op imiza ion o he loss unc ion. The lea ning a e o a pa icula weigh is di ided by a CNN A chi ec u e o Pos u e Classi ica ion on Small Da a Michaela Mes´a oˇso ´a ∗Ond ej Mih´alik ∗Mi osla Ji gl ∗ ∗Depa men o Con ol and Ins umen a ion, Facul y o Elec ical Enginee ing, B no Uni e si y o Technology, B no, Czech epublic (e-mail: Michaela.Mesa oso a@ u .cz). Abs ac : A con olu ional neu al ne wo k is o en men ioned as one o he deep lea ning me hods ha equi es a la ge amoun o aining da a. Ques ioning his belie , his pape explo es he applicabili y o classi ica ion based on a shallow ne s uc u e ained on a small da a se in he con ex o pa ien pos u e classi ica ion based on da a om a p essu e ma ess. Designing a CNN o en p esen s a complex p oblem, especially wi hou a uni e sally applicable app oach, allowing many di e se s uc u al possibili ies and aining se ings. We es ed a ious aining op ions and laye con igu a ions o p o ide an o e iew o in luen ial pa ame e s o pos u e classi ica ion. Expe imen s show encou aging esul s wi h he lea e-one-ou c oss- alida ion accu acy o 93.1% o one o he e alua ed CNN s uc u es and i s hype pa ame e se ings. Keywo ds: CNN, ine uning, ne wo k s uc u e, op imiza ion, pos u e classi ica ion 1. INTRODUCTION Deep lea ning me hods bene i g ea ly om e y la ge da a se s. Howe e , in some si ua ions, i migh be di icul o ga he enough da a, i he da a-acquisi ion p ocess is ime-consuming o labo ious. Choosing he s uc u e and pa ame e s o a CNN ha ep esen s images and hei s o ed in o ma ion uniquely and co ec ly is one o he mos impo an aspec s o achie ing a high CNN pe - o mance. Typically he CNN con igu a ion pa ame e s a e di ided in o wo ca ego ies: hose conce ning he a - chi ec u e and hose conce ning he aining p ocess, see Fe ey a-Rami ez e al. (2019). The speci ic ini ializa ion o he pa ame e s o en has a signi ican impac on how long i akes o he aining p ocess o ind a solu ion and on he gene aliza ion abili y o he esul ing ained ne - wo k. Reg e ully, as s a ed in Bishop and Bishop (2023), he e is no much heo y o help de e mine an ini ializa ion s a egy. Fo his eason, we o e a compac o e iew o a selec ion o se ings and modi iable pa ame e s and hei impac on he classi ica ion esul s. 2. NETWORK TRAINING OPTIONS The igh design o he ne wo k is c ucial o he applica- ion whe e he CNN is o be applied. Howe e he same im- po ance lies in he ask o disco e ing app op ia e aining pa ame e s ha lead o he bes -pe o ming classi ie . We in oduce some o hese pa ame e s along wi h an o e iew o hei ole in he aining p ocess. 2.1 Op imiza ion algo i hms In MATLAB 2023, he e a e a ailable he ollowing ou op imiza ion algo i hms: •SGDM, •RMSP op, •Adam, •L-BGFS Al hough he e exis many mo e algo i hms, hey a e o en de i a i es o hose lis ed abo e, hence we will lea e i o he eade o e alua e hei pe o mance a e he mos sui able base app oach was selec ed. Only he i s h ee will be aken in o conside a ion in his pape . The S ochas ic G adien Descen wi h Momen um o SGDM is an ex ension o he S ochas ic g adien descen (SGD) algo i hm ha upda es he ne wo k pa ame e s o minimize he loss unc ion by aking small s eps a each i e a ion in he di ec ion o he nega i e g adien o he loss, θℓ+1 =θℓ−α∇E(θℓ)+γ(θℓ−θℓ−1),(1) whe e αis he lea ning a e,γis he momen um alue, ℓis he i e a ion numbe , θis he pa ame e ec o , and E(θ) is he loss unc ion. The g adien is e alua ed using a subse o he aining da a, also called a mini-ba ch, whe e a di e en subse is used a each i e a ion, o de ailed explana ion see Bishop and Bishop (2023) and Beale e al. (2023). The en i e pass o e he aining da a is called an epoch. The con ibu ion o he p e ious g adien s ep o he cu en i e a ion is de e mined by he lea ning a e αand he momen um e m γ, which in addi ion educes he oscilla ion a ound he op imum ha migh occu in he case o a e y s eep nega i e g adien . The unpublished Roo Mean Squa e P opaga ion o RM- SP op is designed o accele a e he op imiza ion p ocess, and hus, dec ease he numbe o unc ion e alua ions o each he op imum. This is achie ed by using di e gen lea ning a es o each weigh ha a e able o adap au oma ically o he op imiza ion o he loss unc ion. The lea ning a e o a pa icula weigh is di ided by a CNN A chi ec u e o Pos u e Classi ica ion on Small Da a Michaela Mes´a oˇso ´a ∗Ond ej Mih´alik ∗Mi osla Ji gl ∗ ∗Depa men o Con ol and Ins umen a ion, Facul y o Elec ical Enginee ing, B no Uni e si y o Technology, B no, Czech epublic (e-mail: Michaela.Mesa oso a@ u .cz). Abs ac : A con olu ional neu al ne wo k is o en men ioned as one o he deep lea ning me hods ha equi es a la ge amoun o aining da a. Ques ioning his belie , his pape explo es he applicabili y o classi ica ion based on a shallow ne s uc u e ained on a small da a se in he con ex o pa ien pos u e classi ica ion based on da a om a p essu e ma ess. Designing a CNN o en p esen s a complex p oblem, especially wi hou a uni e sally applicable app oach, allowing many di e se s uc u al possibili ies and aining se ings. We es ed a ious aining op ions and laye con igu a ions o p o ide an o e iew o in luen ial pa ame e s o pos u e classi ica ion. Expe imen s show encou aging esul s wi h he lea e-one-ou c oss- alida ion accu acy o 93.1% o one o he e alua ed CNN s uc u es and i s hype pa ame e se ings. Keywo ds: CNN, ine uning, ne wo k s uc u e, op imiza ion, pos u e classi ica ion 1. INTRODUCTION Deep lea ning me hods bene i g ea ly om e y la ge da a se s. Howe e , in some si ua ions, i migh be di icul o ga he enough da a, i he da a-acquisi ion p ocess is ime-consuming o labo ious. Choosing he s uc u e and pa ame e s o a CNN ha ep esen s images and hei s o ed in o ma ion uniquely and co ec ly is one o he mos impo an aspec s o achie ing a high CNN pe - o mance. Typically he CNN con igu a ion pa ame e s a e di ided in o wo ca ego ies: hose conce ning he a - chi ec u e and hose conce ning he aining p ocess, see Fe ey a-Rami ez e al. (2019). The speci ic ini ializa ion o he pa ame e s o en has a signi ican impac on how long i akes o he aining p ocess o ind a solu ion and on he gene aliza ion abili y o he esul ing ained ne - wo k. Reg e ully, as s a ed in Bishop and Bishop (2023), he e is no much heo y o help de e mine an ini ializa ion s a egy. Fo his eason, we o e a compac o e iew o a selec ion o se ings and modi iable pa ame e s and hei impac on he classi ica ion esul s. 2. NETWORK TRAINING OPTIONS The igh design o he ne wo k is c ucial o he applica- ion whe e he CNN is o be applied. Howe e he same im- po ance lies in he ask o disco e ing app op ia e aining pa ame e s ha lead o he bes -pe o ming classi ie . We in oduce some o hese pa ame e s along wi h an o e iew o hei ole in he aining p ocess. 2.1 Op imiza ion algo i hms In MATLAB 2023, he e a e a ailable he ollowing ou op imiza ion algo i hms: •SGDM, •RMSP op, •Adam, •L-BGFS Al hough he e exis many mo e algo i hms, hey a e o en de i a i es o hose lis ed abo e, hence we will lea e i o he eade o e alua e hei pe o mance a e he mos sui able base app oach was selec ed. Only he i s h ee will be aken in o conside a ion in his pape . The S ochas ic G adien Descen wi h Momen um o SGDM is an ex ension o he S ochas ic g adien descen (SGD) algo i hm ha upda es he ne wo k pa ame e s o minimize he loss unc ion by aking small s eps a each i e a ion in he di ec ion o he nega i e g adien o he loss, θℓ+1 =θℓ−α∇E(θℓ)+γ(θℓ−θℓ−1),(1) whe e αis he lea ning a e,γis he momen um alue, ℓis he i e a ion numbe , θis he pa ame e ec o , and E(θ) is he loss unc ion. The g adien is e alua ed using a subse o he aining da a, also called a mini-ba ch, whe e a di e en subse is used a each i e a ion, o de ailed explana ion see Bishop and Bishop (2023) and Beale e al. (2023). The en i e pass o e he aining da a is called an epoch. The con ibu ion o he p e ious g adien s ep o he cu en i e a ion is de e mined by he lea ning a e αand he momen um e m γ, which in addi ion educes he oscilla ion a ound he op imum ha migh occu in he case o a e y s eep nega i e g adien . The unpublished Roo Mean Squa e P opaga ion o RM- SP op is designed o accele a e he op imiza ion p ocess, and hus, dec ease he numbe o unc ion e alua ions o each he op imum. This is achie ed by using di e gen lea ning a es o each weigh ha a e able o adap au oma ically o he op imiza ion o he loss unc ion. The lea ning a e o a pa icula weigh is di ided by a CNN A chi ec u e o Pos u e Classi ica ion on Small Da a Michaela Mes´a oˇso ´a ∗Ond ej Mih´alik ∗Mi osla Ji gl ∗ ∗ Depa men o Con ol and Ins umen a ion, Facul y o Elec ical Enginee ing, B no Uni e si y o Technology, B no, Czech epublic (e-mail: Michaela.Mesa os[email p o ec ed]). Abs ac : A con olu ional neu al ne wo k is o en men ioned as one o he deep lea ning me hods ha equi es a la ge amoun o aining da a. Ques ioning his belie , his pape explo es he applicabili y o classi ica ion based on a shallow ne s uc u e ained on a small da a se in he con ex o pa ien pos u e classi ica ion based on da a om a p essu e ma ess. Designing a CNN o en p esen s a complex p oblem, especially wi hou a uni e sally applicable app oach, allowing many di e se s uc u al possibili ies and aining se ings. We es ed a ious aining op ions and laye con igu a ions o p o ide an o e iew o in luen ial pa ame e s o pos u e classi ica ion. Expe imen s show encou aging esul s wi h he lea e-one-ou c oss- alida ion accu acy o 93.1% o one o he e alua ed CNN s uc u es and i s hype pa ame e se ings. Keywo ds: CNN, ine uning, ne wo k s uc u e, op imiza ion, pos u e classi ica ion 1. INTRODUCTION Deep lea ning me hods bene i g ea ly om e y la ge da a se s. Howe e , in some si ua ions, i migh be di icul o ga he enough da a, i he da a-acquisi ion p ocess is ime-consuming o labo ious. Choosing he s uc u e and pa ame e s o a CNN ha ep esen s images and hei s o ed in o ma ion uniquely and co ec ly is one o he mos impo an aspec s o achie ing a high CNN pe - o mance. Typically he CNN con igu a ion pa ame e s a e di ided in o wo ca ego ies: hose conce ning he a - chi ec u e and hose conce ning he aining p ocess, see Fe ey a-Rami ez e al. (2019). The speci ic ini ializa ion o he pa ame e s o en has a signi ican impac on how long i akes o he aining p ocess o ind a solu ion and on he gene aliza ion abili y o he esul ing ained ne - wo k. Reg e ully, as s a ed in Bishop and Bishop (2023), he e is no much heo y o help de e mine an ini ializa ion s a egy. Fo his eason, we o e a compac o e iew o a selec ion o se ings and modi iable pa ame e s and hei impac on he classi ica ion esul s. 2. NETWORK TRAINING OPTIONS The igh design o he ne wo k is c ucial o he applica- ion whe e he CNN is o be applied. Howe e he same im- po ance lies in he ask o disco e ing app op ia e aining pa ame e s ha lead o he bes -pe o ming classi ie . We in oduce some o hese pa ame e s along wi h an o e iew o hei ole in he aining p ocess. 2.1 Op imiza ion algo i hms In MATLAB 2023, he e a e a ailable he ollowing ou op imiza ion algo i hms: •SGDM, •RMSP op, •Adam, •L-BGFS Al hough he e exis many mo e algo i hms, hey a e o en de i a i es o hose lis ed abo e, hence we will lea e i o he eade o e alua e hei pe o mance a e he mos sui able base app oach was selec ed. Only he i s h ee will be aken in o conside a ion in his pape . The S ochas ic G adien Descen wi h Momen um o SGDM is an ex ension o he S ochas ic g adien descen (SGD) algo i hm ha upda es he ne wo k pa ame e s o minimize he loss unc ion by aking small s eps a each i e a ion in he di ec ion o he nega i e g adien o he loss, θℓ+1 =θℓ−α∇E(θℓ)+γ(θℓ−θℓ−1),(1) whe e αis he lea ning a e,γis he momen um alue, ℓis he i e a ion numbe , θis he pa ame e ec o , and E(θ) is he loss unc ion. The g adien is e alua ed using a subse o he aining da a, also called a mini-ba ch, whe e a di e en subse is used a each i e a ion, o de ailed explana ion see Bishop and Bishop (2023) and Beale e al. (2023). The en i e pass o e he aining da a is called an epoch. The con ibu ion o he p e ious g adien s ep o he cu en i e a ion is de e mined by he lea ning a e αand he momen um e m γ, which in addi ion educes he oscilla ion a ound he op imum ha migh occu in he case o a e y s eep nega i e g adien . The unpublished Roo Mean Squa e P opaga ion o RM- SP op is designed o accele a e he op imiza ion p ocess, and hus, dec ease he numbe o unc ion e alua ions o each he op imum. This is achie ed by using di e gen lea ning a es o each weigh ha a e able o adap au oma ically o he op imiza ion o he loss unc ion. The lea ning a e o a pa icula weigh is di ided by a CNN A chi ec u e o Pos u e Classi ica ion on Small Da a Michaela Mes´a oˇso ´a ∗ Ond ej Mih´alik ∗ Mi osla Ji gl ∗ ∗Depa men o Con ol and Ins umen a ion, Facul y o Elec ical Enginee ing, B no Uni e si y o Technology, B no, Czech epublic (e-mail: Michaela.Mesa oso a@ u .cz). Abs ac : A con olu ional neu al ne wo k is o en men ioned as one o he deep lea ning me hods ha equi es a la ge amoun o aining da a. Ques ioning his belie , his pape explo es he applicabili y o classi ica ion based on a shallow ne s uc u e ained on a small da a se in he con ex o pa ien pos u e classi ica ion based on da a om a p essu e ma ess. Designing a CNN o en p esen s a complex p oblem, especially wi hou a uni e sally applicable app oach, allowing many di e se s uc u al possibili ies and aining se ings. We es ed a ious aining op ions and laye con igu a ions o p o ide an o e iew o in luen ial pa ame e s o pos u e classi ica ion. Expe imen s show encou aging esul s wi h he lea e-one-ou c oss- alida ion accu acy o 93.1% o one o he e alua ed CNN s uc u es and i s hype pa ame e se ings. Keywo ds: CNN, ine uning, ne wo k s uc u e, op imiza ion, pos u e classi ica ion 1. INTRODUCTION Deep lea ning me hods bene i g ea ly om e y la ge da a se s. Howe e , in some si ua ions, i migh be di icul o ga he enough da a, i he da a-acquisi ion p ocess is ime-consuming o labo ious. Choosing he s uc u e and pa ame e s o a CNN ha ep esen s images and hei s o ed in o ma ion uniquely and co ec ly is one o he mos impo an aspec s o achie ing a high CNN pe - o mance. Typically he CNN con igu a ion pa ame e s a e di ided in o wo ca ego ies: hose conce ning he a - chi ec u e and hose conce ning he aining p ocess, see Fe ey a-Rami ez e al. (2019). The speci ic ini ializa ion o he pa ame e s o en has a signi ican impac on how long i akes o he aining p ocess o ind a solu ion and on he gene aliza ion abili y o he esul ing ained ne - wo k. Reg e ully, as s a ed in Bishop and Bishop (2023), he e is no much heo y o help de e mine an ini ializa ion s a egy. Fo his eason, we o e a compac o e iew o a selec ion o se ings and modi iable pa ame e s and hei impac on he classi ica ion esul s. 2. NETWORK TRAINING OPTIONS The igh design o he ne wo k is c ucial o he applica- ion whe e he CNN is o be applied. Howe e he same im- po ance lies in he ask o disco e ing app op ia e aining pa ame e s ha lead o he bes -pe o ming classi ie . We in oduce some o hese pa ame e s along wi h an o e iew o hei ole in he aining p ocess. 2.1 Op imiza ion algo i hms In MATLAB 2023, he e a e a ailable he ollowing ou op imiza ion algo i hms: •SGDM, •RMSP op, •Adam, •L-BGFS Al hough he e exis many mo e algo i hms, hey a e o en de i a i es o hose lis ed abo e, hence we will lea e i o he eade o e alua e hei pe o mance a e he mos sui able base app oach was selec ed. Only he i s h ee will be aken in o conside a ion in his pape . The S ochas ic G adien Descen wi h Momen um o SGDM is an ex ension o he S ochas ic g adien descen (SGD) algo i hm ha upda es he ne wo k pa ame e s o minimize he loss unc ion by aking small s eps a each i e a ion in he di ec ion o he nega i e g adien o he loss, θℓ+1 =θℓ−α∇E(θℓ)+γ(θℓ−θℓ−1),(1) whe e αis he lea ning a e,γis he momen um alue, ℓis he i e a ion numbe , θis he pa ame e ec o , and E(θ) is he loss unc ion. The g adien is e alua ed using a subse o he aining da a, also called a mini-ba ch, whe e a di e en subse is used a each i e a ion, o de ailed explana ion see Bishop and Bishop (2023) and Beale e al. (2023). The en i e pass o e he aining da a is called an epoch. The con ibu ion o he p e ious g adien s ep o he cu en i e a ion is de e mined by he lea ning a e αand he momen um e m γ, which in addi ion educes he oscilla ion a ound he op imum ha migh occu in he case o a e y s eep nega i e g adien . The unpublished Roo Mean Squa e P opaga ion o RM- SP op is designed o accele a e he op imiza ion p ocess, and hus, dec ease he numbe o unc ion e alua ions o each he op imum. This is achie ed by using di e gen lea ning a es o each weigh ha a e able o adap au oma ically o he op imiza ion o he loss unc ion. The lea ning a e o a pa icula weigh is di ided by a CNN A chi ec u e o Pos u e Classi ica ion on Small Da a Michaela Mes´a oˇso ´a ∗Ond ej Mih´alik ∗Mi osla Ji gl ∗ ∗Depa men o Con ol and Ins umen a ion, Facul y o Elec ical Enginee ing, B no Uni e si y o Technology, B no, Czech epublic (e-mail: Michaela.Mesa oso a@ u .cz). Abs ac : A con olu ional neu al ne wo k is o en men ioned as one o he deep lea ning me hods ha equi es a la ge amoun o aining da a. Ques ioning his belie , his pape explo es he applicabili y o classi ica ion based on a shallow ne s uc u e ained on a small da a se in he con ex o pa ien pos u e classi ica ion based on da a om a p essu e ma ess. Designing a CNN o en p esen s a complex p oblem, especially wi hou a uni e sally applicable app oach, allowing many di e se s uc u al possibili ies and aining se ings. We es ed a ious aining op ions and laye con igu a ions o p o ide an o e iew o in luen ial pa ame e s o pos u e classi ica ion. Expe imen s show encou aging esul s wi h he lea e-one-ou c oss- alida ion accu acy o 93.1% o one o he e alua ed CNN s uc u es and i s hype pa ame e se ings. Keywo ds: CNN, ine uning, ne wo k s uc u e, op imiza ion, pos u e classi ica ion 1. INTRODUCTION Deep lea ning me hods bene i g ea ly om e y la ge da a se s. Howe e , in some si ua ions, i migh be di icul o ga he enough da a, i he da a-acquisi ion p ocess is ime-consuming o labo ious. Choosing he s uc u e and pa ame e s o a CNN ha ep esen s images and hei s o ed in o ma ion uniquely and co ec ly is one o he mos impo an aspec s o achie ing a high CNN pe - o mance. Typically he CNN con igu a ion pa ame e s a e di ided in o wo ca ego ies: hose conce ning he a - chi ec u e and hose conce ning he aining p ocess, see Fe ey a-Rami ez e al. (2019). The speci ic ini ializa ion o he pa ame e s o en has a signi ican impac on how long i akes o he aining p ocess o ind a solu ion and on he gene aliza ion abili y o he esul ing ained ne - wo k. Reg e ully, as s a ed in Bishop and Bishop (2023), he e is no much heo y o help de e mine an ini ializa ion s a egy. Fo his eason, we o e a compac o e iew o a selec ion o se ings and modi iable pa ame e s and hei impac on he classi ica ion esul s. 2. NETWORK TRAINING OPTIONS The igh design o he ne wo k is c ucial o he applica- ion whe e he CNN is o be applied. Howe e he same im- po ance lies in he ask o disco e ing app op ia e aining pa ame e s ha lead o he bes -pe o ming classi ie . We in oduce some o hese pa ame e s along wi h an o e iew o hei ole in he aining p ocess. 2.1 Op imiza ion algo i hms In MATLAB 2023, he e a e a ailable he ollowing ou op imiza ion algo i hms: •SGDM, •RMSP op, •Adam, •L-BGFS Al hough he e exis many mo e algo i hms, hey a e o en de i a i es o hose lis ed abo e, hence we will lea e i o he eade o e alua e hei pe o mance a e he mos sui able base app oach was selec ed. Only he i s h ee will be aken in o conside a ion in his pape . The S ochas ic G adien Descen wi h Momen um o SGDM is an ex ension o he S ochas ic g adien descen (SGD) algo i hm ha upda es he ne wo k pa ame e s o minimize he loss unc ion by aking small s eps a each i e a ion in he di ec ion o he nega i e g adien o he loss, θℓ+1 =θℓ−α∇E(θℓ)+γ(θℓ−θℓ−1),(1) whe e αis he lea ning a e,γis he momen um alue, ℓis he i e a ion numbe , θis he pa ame e ec o , and E(θ) is he loss unc ion. The g adien is e alua ed using a subse o he aining da a, also called a mini-ba ch, whe e a di e en subse is used a each i e a ion, o de ailed explana ion see Bishop and Bishop (2023) and Beale e al. (2023). The en i e pass o e he aining da a is called an epoch. The con ibu ion o he p e ious g adien s ep o he cu en i e a ion is de e mined by he lea ning a e αand he momen um e m γ, which in addi ion educes he oscilla ion a ound he op imum ha migh occu in he case o a e y s eep nega i e g adien . The unpublished Roo Mean Squa e P opaga ion o RM- SP op is designed o accele a e he op imiza ion p ocess, and hus, dec ease he numbe o unc ion e alua ions o each he op imum. This is achie ed by using di e gen lea ning a es o each weigh ha a e able o adap au oma ically o he op imiza ion o he loss unc ion. The lea ning a e o a pa icula weigh is di ided by a Copy igh © 2024 The Au ho s. This is an open access a icle unde he CC BY-NC-ND license ( h ps://c ea i ecommons.o g/licenses/by-nc-nd/4.0/ ) 300 Michaela Mesá ošo á e al. / IFAC Pape sOnLine 58-9 (2024) 299–304 unning a e age o he magni udes o ecen g adien s o he co esponding weigh . ℓ=β2 ℓ−1+ (1 −β2)[∇E(θℓ)]2(2) θℓ+1 =θℓ−α∇E(θℓ) √ ℓ+ϵ,(3) whe e ℓis he mo ing a e age, β2is he squa ed g adien decay ac o o he mo ing a e age, and ϵis a small cons an o a oid di ision by ze o, o mo e de ail see Tieleman and Hin on (2012). As s a ed by Kingma and Ba (2015), RMSP op is sui ed o online and non-s a iona y se ings. Adap i e momen es ima ion, also called Adam, was i s in oduced in Kingma and Ba (2015). I is an e icien s ochas ic op imiza ion ha equi es i s -o de g adien s. Indi idual adap i e lea ning a es a e again main ained o all pa ame e s sepa a ely. Upda es a e di ec ly es ima ed by using a unning a e age o he i s and he second momen s o he g adien s. Mo ing a e ages a e calcula ed as ollows: mℓ=β1mℓ−1+ (1 −β1)∇E(θℓ) (4) ℓ=β2 ℓ−1+ (1 −β2)[∇E(θℓ)]2(5) mℓ=mℓ/(1 −βℓ 1) (6)  ℓ= ℓ/(1 −βℓ 2) (7) whe e β1and β2a e he exponen ial decay a es o he momen es ima es. The ne wo k pa ame e s hen a e upda ed as θℓ+1 =θℓ−αmℓ √ ℓ+ϵ(8) I g adien s a e simila h oughou many i e a ions up- da es o he weigh s a e able o gain momen um in a ce ain di ec ion by using a mo ing a e age o he g adien . The Adam algo i hm is sui able o wo king wi h la ge amoun s o da a o i s small memo y equi emen s. 2.2 Lea ning a e Lea ning a e, also e e ed o as s ep size, de e mines he a e o change o he weigh s. The ini ial lea ning a e has a signi ican impac on he lea ning speed and he o e all aining de elopmen . When he lea ning a e is oo la ge, he aining e o may inad e en ly inc ease a he han dec ease, ye wi h a alue oo small, aining is subs an ially slowe and may become s uck wi h a high aining e o , o mo e de ail see Good ellow e al. (2016). Fu he ema ks om Bengio (2012) and Reed and Ma ks (1999) iden i y he lea ning a e as he mos impo an hype pa ame e . I s ypical de aul alue is se o 0.01. Howe e , addi ional uning is always ecommended. 2.3 Regula iza ion A easonable app oach o p oblem-sol ing would be o adap he model complexi y acco ding o he complexi y o he p oblem. One way o con ol gene aliza ion and a oid he o e i ing phenomenon is by egula iza ion. To include egula iza ion in model aining, a penal y unc ion is added o he loss unc ion E(θ), discou aging he pa ame e s om ha ing la ge magni udes.  E(θ)=E(θ)+λΩ (9) whe e Ω is he egula iza ion unc ion and λis he mul i- plica i e pa ame e . An op imal ne wo k is one in which a comp omise be ween he bes i o he aining da a and a smoo hness o he i is ound, i.e. minimum o e all e o  E, see Zaknich (2003). 2.4 Momen um Du ing aining, he p oblem o widely di e ing eigen al- ues o en occu s. One echnique o deal wi h his incon- enience is o add momen um γ o he weigh upda e o mula. This limi s he oscilla ions and adds ine ia o he mo emen h ough weigh space. As may be ound in Bishop and Bishop (2023), he e ec o he momen um is inc easing he e ec i e lea ning a e along he cu e in he weigh space, whe e he g adien emains unchanged. On he con a y, in high cu a u e egions whe e he g adien changes signi ican ly, he con ibu ion o he momen um is supp essed and has li le o no e ec on he e ec i e lea ning a e alue. Fig. 1. Di e ence be ween he g adien descen wi h he momen um e m on he le side wi h as e con- e gence owa ds he op imum in compa ison o g a- dien descen wi hou he momen um e m oscilla ing along he pa h on he igh side. 2.5 Mini-Ba ch In i s p inciple, SGD uses only one da a poin o calcula e g adien es ima ion o he e o unc ion leading o a e y noisy es ima e, while he compu a ion on he en i e da a se yields accu a e g adien esul s. Howe e , o de e mine g adien es ima ion o each da a poin is compu a ionally demanding, hence a mini-ba ch—a small subse o da a poin s—is used o e alua e he g adien a each i e a ion. Acco ding o Bishop and Bishop (2023), an impo an ac o o bea in mind when using mini-ba ches is ha cons i uen da a poin s should be chosen andomly om he aining se . This is due o possible co ela ions be- ween consecu i e da a poin s a ising om he way da a was collec ed o s o ed, o example, i is so ed in an al- phabe ical o ch onological o de . 3. NET STRUCTURE A g ea di e si y o ne s uc u es can be obse ed ac oss hei applica ions as a consequence o di e en classi ica- ion p oblems ha ing da a se s ha di e in hei o ma , and hus, equi ing a co esponding ne wo k a chi ec u e. In his pape we ocus on applying CNNs in an a ypical se ing: designing a classi ie using a small da a se o 290 samples composed o a low esolu ion images, see Fig 2. A e ca e ul conside a ion o he da ase ’s p ope ies, a simple s uc u e wi h a conside ably smalle numbe o Michaela Mesá ošo á e al. / IFAC Pape sOnLine 58-9 (2024) 299–304 301 unning a e age o he magni udes o ecen g adien s o he co esponding weigh . ℓ=β2 ℓ−1+ (1 −β2)[∇E(θℓ)]2(2) θℓ+1 =θℓ−α∇E(θℓ) √ ℓ+ϵ,(3) whe e ℓis he mo ing a e age, β2is he squa ed g adien decay ac o o he mo ing a e age, and ϵis a small cons an o a oid di ision by ze o, o mo e de ail see Tieleman and Hin on (2012). As s a ed by Kingma and Ba (2015), RMSP op is sui ed o online and non-s a iona y se ings. Adap i e momen es ima ion, also called Adam, was i s in oduced in Kingma and Ba (2015). I is an e icien s ochas ic op imiza ion ha equi es i s -o de g adien s. Indi idual adap i e lea ning a es a e again main ained o all pa ame e s sepa a ely. Upda es a e di ec ly es ima ed by using a unning a e age o he i s and he second momen s o he g adien s. Mo ing a e ages a e calcula ed as ollows: mℓ=β1mℓ−1+ (1 −β1)∇E(θℓ) (4) ℓ=β2 ℓ−1+ (1 −β2)[∇E(θℓ)]2(5) mℓ=mℓ/(1 −βℓ 1) (6)  ℓ= ℓ/(1 −βℓ 2) (7) whe e β1and β2a e he exponen ial decay a es o he momen es ima es. The ne wo k pa ame e s hen a e upda ed as θℓ+1 =θℓ−αmℓ √ ℓ+ϵ(8) I g adien s a e simila h oughou many i e a ions up- da es o he weigh s a e able o gain momen um in a ce ain di ec ion by using a mo ing a e age o he g adien . The Adam algo i hm is sui able o wo king wi h la ge amoun s o da a o i s small memo y equi emen s. 2.2 Lea ning a e Lea ning a e, also e e ed o as s ep size, de e mines he a e o change o he weigh s. The ini ial lea ning a e has a signi ican impac on he lea ning speed and he o e all aining de elopmen . When he lea ning a e is oo la ge, he aining e o may inad e en ly inc ease a he han dec ease, ye wi h a alue oo small, aining is subs an ially slowe and may become s uck wi h a high aining e o , o mo e de ail see Good ellow e al. (2016). Fu he ema ks om Bengio (2012) and Reed and Ma ks (1999) iden i y he lea ning a e as he mos impo an hype pa ame e . I s ypical de aul alue is se o 0.01. Howe e , addi ional uning is always ecommended. 2.3 Regula iza ion A easonable app oach o p oblem-sol ing would be o adap he model complexi y acco ding o he complexi y o he p oblem. One way o con ol gene aliza ion and a oid he o e i ing phenomenon is by egula iza ion. To include egula iza ion in model aining, a penal y unc ion is added o he loss unc ion E(θ), discou aging he pa ame e s om ha ing la ge magni udes.  E(θ)=E(θ)+λΩ (9) whe e Ω is he egula iza ion unc ion and λis he mul i- plica i e pa ame e . An op imal ne wo k is one in which a comp omise be ween he bes i o he aining da a and a smoo hness o he i is ound, i.e. minimum o e all e o  E, see Zaknich (2003). 2.4 Momen um Du ing aining, he p oblem o widely di e ing eigen al- ues o en occu s. One echnique o deal wi h his incon- enience is o add momen um γ o he weigh upda e o mula. This limi s he oscilla ions and adds ine ia o he mo emen h ough weigh space. As may be ound in Bishop and Bishop (2023), he e ec o he momen um is inc easing he e ec i e lea ning a e along he cu e in he weigh space, whe e he g adien emains unchanged. On he con a y, in high cu a u e egions whe e he g adien changes signi ican ly, he con ibu ion o he momen um is supp essed and has li le o no e ec on he e ec i e lea ning a e alue. Fig. 1. Di e ence be ween he g adien descen wi h he momen um e m on he le side wi h as e con- e gence owa ds he op imum in compa ison o g a- dien descen wi hou he momen um e m oscilla ing along he pa h on he igh side. 2.5 Mini-Ba ch In i s p inciple, SGD uses only one da a poin o calcula e g adien es ima ion o he e o unc ion leading o a e y noisy es ima e, while he compu a ion on he en i e da a se yields accu a e g adien esul s. Howe e , o de e mine g adien es ima ion o each da a poin is compu a ionally demanding, hence a mini-ba ch—a small subse o da a poin s—is used o e alua e he g adien a each i e a ion. Acco ding o Bishop and Bishop (2023), an impo an ac o o bea in mind when using mini-ba ches is ha cons i uen da a poin s should be chosen andomly om he aining se . This is due o possible co ela ions be- ween consecu i e da a poin s a ising om he way da a was collec ed o s o ed, o example, i is so ed in an al- phabe ical o ch onological o de . 3. NET STRUCTURE A g ea di e si y o ne s uc u es can be obse ed ac oss hei applica ions as a consequence o di e en classi ica- ion p oblems ha ing da a se s ha di e in hei o ma , and hus, equi ing a co esponding ne wo k a chi ec u e. In his pape we ocus on applying CNNs in an a ypical se ing: designing a classi ie using a small da a se o 290 samples composed o a low esolu ion images, see Fig 2. A e ca e ul conside a ion o he da ase ’s p ope ies, a simple s uc u e wi h a conside ably smalle numbe o lea ning pa ame e s was selec ed. This decision is based on he gene ally obse ed ac , ha la ge CNNs end o o e i he model wi h such a small amoun o a ailable da a. One app oach o handle his issue s a ed by Kesha i e al. (2018) is o educe he numbe o lea nable pa am- e e s. 3.1 Con olu ional laye One o he i s s uc u al p ope ies we conside a e con olu ional laye s, numbe o ke nels and hei size. Con olu ional laye ex ac s in o ma ion o ea u es om he inpu signal using con olu ion il e s. Gi en he low esolu ion o he images in he da a se , he commonly used il e sizes sensible o his applica ion a e 3 ×3 and 5×5. Hence, we can desc ibe he con olu ional laye s o he ne wo k as a uple (l1,l 2), whe e li×liis he size o he con olu ion il e in i- h laye , o i∈{1,2}. These pa ame e combina ions we e es ed: (l1,l 2); l1,l 2∈{3,5}.(10) 3.2 Pooling laye Pooling is esponsible o downsizing he spa ial size o he ou pu om he p e ious laye by compu ing an a e age alue in he il e window—a e age pool—o he popula app oach o picking he maximal alue in he il e egion—max pool —see in Zhou and Chellappa (1988). This educes compu a ional complexi y and allows he subsequen con olu ional laye o ex ac ea u es a a di e en scale. Simila ly o con olu ional laye s we can desc ibe ou selec ion o pooling laye s. I we ake in o accoun he image esolu ion and he ac ha key ea u es o he lying posi ion a e o en concen a ed on smalle su aces, i will also be app op ia e o adjus he il e dimensions acco dingly. Hence he con empla ed sizes we e 2×2and3×3 wi h wo possible s ep sizes 1 and 2. The indi idual con igu a ions compa ed in he expe imen s can be hen desc ibed by a pa ame e ec o (p1,s 1,p 2,s 2) whe e he dimension o he i- h pooling laye is pi×piand siis he s ep size o he i- h il e o i∈{1,2}. Tes ed combina ions we e (p1,s 1,p 2,s 2)∈({2,3}×{1,2})2.(11) 2 4 6 8 10 5 10 15 20 25 30 2 4 6 8 10 5 10 15 20 25 30 2 4 6 8 10 5 10 15 20 25 30 2 4 6 8 10 5 10 15 20 25 30 Fig. 2. Da a se example o one subjec in ou lying posi ions. The chosen pooling me hod is he max pool. 2The simple easoning behind his choice is o maximize he p essu e indica ions in key a eas such as shoulde s, hips, knees and heels ha a e impo an o a success ul classi ica ion. Using he a e age pool would blu hese a eas, causing a dec eased di e en iabili y be ween he ea u es o classi- ica ion classes. 3.3 Ac i a ion unc ion The ac i a ion unc ion de ines he ou pu o a node o a neu on o a gi en inpu . I can be hough o as an e alua ion o whe he o no a neu on should be ac i a ed upon he a i al o a speci ic inpu . Ne wo ks a e capable o sol ing non i ial p oblems when nonlinea ac i a ion unc ions a e included in he s uc u e, see Hinkelmann (2018). Some o he adi ional and mode n ac i a ion unc ions wi h hei applica ions a e he sigmoid unc ion, hype bolic angen , bina y s ep in oduced in McCulloch and Pi s (1943), Rec i ied Linea Uni (ReLU) desc ibed in Nai and Hin on (2010), Leaky ReLU desc ibed in Maas (2013), Exponen ial Linea Uni (ELU) desc ibed in Cle e e al. (2016), o name a ew. Sigmoid, hype bolic angen , ReLU and ELU ac i a ion unc ions we e used in he p esen ed expe imen s. 4. DATA Fo aining pu poses p essu e map images we e used. These maps ep esen he alues measu ed by he p essu e- sensi i e ma ess wi h a pe son si ua ed in lying posi ions, see Fig. 2. The da a se consis s o 290 images o size 30 ×11. Each image has an associa ed subjec and class numbe . All classes 1–4 illus a ed in Fig. 2 co espond o hese posi ions: •on he back (28.2%), •on he igh side (23.7%), •on he le side (21.7%), •on he s omach (26.2%), espec i ely. Numbe s ep esen he pe cen age ep esen- a ion o a class in he da ase . Wi hin each o he posi- ions, sligh a ia ions may occu such as di e en posi- ioning o he a ms and legs o loca ion on he ma ess, i.e., lying in he middle, on he edge o diagonally. The o al numbe o measu ed subjec s is 18. Images a e s o ed as ma ices ha con ain alues om 0 o 1, whe e 0 ep esen s he maximum p essu e and 1 is no p essu e on he ma ess. 5. DESIGN AND PERFORMANCE A chi ec u es we e ained using MATLAB 2023 Deep Lea ning Toolbox. I is necessa y o s a e ha he esul s p esen ed in his pape a e only op imal o pa ame e space ea ma ked in Sec ions 3 and 5. The e may exis ne s uc u es ha would achie e simila o be e esul s, howe e , i is no possible o sea ch he whole pa ame e space o ind he globally bes -pe o ming CNN. Ou esul s a e quan i ied using he Accu acy measu e (Acc) ob ained by c oss- alida ion wi h 18 olds (numbe o subjec s). 5.1 S uc u e Be o e we a e able o une he ne wo k’s pa ame e s, he ne wo k s uc u e needs o be es ablished. The s uc u al 302 Michaela Mesá ošo á e al. / IFAC Pape sOnLine 58-9 (2024) 299–304 Fig. 3. Dependence o accu acy on he il e con igu a ion and egula iza ion. Con olu ion il e sizes we e se o (l1,l 2)=(3,3) wi h s ide 1. s a ing poin was se o wo con olu ional laye s be ween which he pooling laye s we e placed. Th ee o en il e s we e es ed in each con olu ional laye o il e sizes s a ed in (10). As o he pa ame e s, he only one esol ed a he beginning is he op imiza ion algo i hm. A e conside a- ion o algo i hm p ope ies, Adam was chosen as i com- bines he ad an ages o he g adien descen op imiza ion and RMSP op. The lea ning a e was se o 0.05 o all aining scena ios and he ange o egula iza ion alues is om 0.01 o 7 ·10−5. O he aining op ions emain a de aul alues un il he inal s uc u e is chosen. Figs. 3 and 4 depic wo examples o il e size combina- ions. Fig. 3 shows ha dec easing egula iza ion c ea es a a he uns able and luc ua ing su ace accompanied by descending a e age Acc. Fo his eason, u he es s did no include egula iza ion wi h alues 10−4and 7 ·10−5. None o he combina ions shown in Fig. 4 we e able o su pass he bes esul s o il e sizes 3 ×3, which is also he case o he emaining il e size a ia ions. The e o e, he inal s uc u e o he con olu ional laye s is six il e s in he i s laye and 7 il e s in he second laye . Weigh s in he il e s we e ini ialized om he no mal dis ibu ion wi h ze o mean and s anda d de ia ion 0.01. The nex pa ame e analysed was he ac i a ion unc- ion. The expe imen included sigmoid, hype bolic an- gen , leaky ReLU and ELU. A e e alua ion o a ious combina ions, we concluded ha he ac i a ion unc ion Fig. 4. Dependence o Acc on il e con igu a ion and egula iza ion. Con olu ion il e sizes we e se o (l1,l 2)=(5,3) wi h s ide 1. should be consis en h oughou he ne wo k. The i s ejec ed ac i a ion was sigmoid o i s poo esul s, only a ound 25% Acc, which is compa able wi h andom guess- ing. Mo eo e , i is compu a ionally demanding. Al hough hype bolic angen is also a complex unc ion o compu e, i yields he bes esul s closely ollowed by ReLU and ELU ac i a ion unc ions wi h c oss- alida ion Acc ≈90%. The inal examined s uc u al aspec s a e he pooling laye s. Fo he easons men ioned in Sec ion 3 max pooling me hod was applied in con igu a ion (11). Di e ences in he classi ica ion accu acy o he di e en con igu a ions a ied only sligh ly (±2%). Wi h smalle sliding s eps he accu acy inc eased, howe e , lowe ing he s ep size escala es he ime equi ed o aining. I s ide 1 is used e en in one laye , he aining ime almos doubles. The size o he used da ase implies ha such an inc ease in he aining ime s ill does no p esen a p oblem, bu in he case o di e en da ase s (numbe o measu emen s o da a ype), i is a ac o ha needs o be aken in o accoun . Con igu a ion (p1,s 1,p 2,s 2)=(2,1,2,2) achie ed he highes accu acy. A e he examina ion o he pa ial success a es inal ne wo k s uc u e was selec ed: •6 con olu ion il e 3 ×3, •hype bolic angen , •max pool 2 ×2, •7 con olu ion il e 3 ×3, •hype bolic angen , •max pool 2 ×2, • ully connec ed laye , •so max wi h 4 classes. This ne wo k con ains only 1,700 lea nable pa ame e s, which is a no ably smalle numbe in compa ison wi h Michaela Mesá ošo á e al. / IFAC Pape sOnLine 58-9 (2024) 299–304 303 Fig. 3. Dependence o accu acy on he il e con igu a ion and egula iza ion. Con olu ion il e sizes we e se o (l1,l 2)=(3,3) wi h s ide 1. s a ing poin was se o wo con olu ional laye s be ween which he pooling laye s we e placed. Th ee o en il e s we e es ed in each con olu ional laye o il e sizes s a ed in (10). As o he pa ame e s, he only one esol ed a he beginning is he op imiza ion algo i hm. A e conside a- ion o algo i hm p ope ies, Adam was chosen as i com- bines he ad an ages o he g adien descen op imiza ion and RMSP op. The lea ning a e was se o 0.05 o all aining scena ios and he ange o egula iza ion alues is om 0.01 o 7 ·10−5. O he aining op ions emain a de aul alues un il he inal s uc u e is chosen. Figs. 3 and 4 depic wo examples o il e size combina- ions. Fig. 3 shows ha dec easing egula iza ion c ea es a a he uns able and luc ua ing su ace accompanied by descending a e age Acc. Fo his eason, u he es s did no include egula iza ion wi h alues 10−4and 7 ·10−5. None o he combina ions shown in Fig. 4 we e able o su pass he bes esul s o il e sizes 3 ×3, which is also he case o he emaining il e size a ia ions. The e o e, he inal s uc u e o he con olu ional laye s is six il e s in he i s laye and 7 il e s in he second laye . Weigh s in he il e s we e ini ialized om he no mal dis ibu ion wi h ze o mean and s anda d de ia ion 0.01. The nex pa ame e analysed was he ac i a ion unc- ion. The expe imen included sigmoid, hype bolic an- gen , leaky ReLU and ELU. A e e alua ion o a ious combina ions, we concluded ha he ac i a ion unc ion Fig. 4. Dependence o Acc on il e con igu a ion and egula iza ion. Con olu ion il e sizes we e se o (l1,l 2)=(5,3) wi h s ide 1. should be consis en h oughou he ne wo k. The i s ejec ed ac i a ion was sigmoid o i s poo esul s, only a ound 25% Acc, which is compa able wi h andom guess- ing. Mo eo e , i is compu a ionally demanding. Al hough hype bolic angen is also a complex unc ion o compu e, i yields he bes esul s closely ollowed by ReLU and ELU ac i a ion unc ions wi h c oss- alida ion Acc ≈90%. The inal examined s uc u al aspec s a e he pooling laye s. Fo he easons men ioned in Sec ion 3 max pooling me hod was applied in con igu a ion (11). Di e ences in he classi ica ion accu acy o he di e en con igu a ions a ied only sligh ly (±2%). Wi h smalle sliding s eps he accu acy inc eased, howe e , lowe ing he s ep size escala es he ime equi ed o aining. I s ide 1 is used e en in one laye , he aining ime almos doubles. The size o he used da ase implies ha such an inc ease in he aining ime s ill does no p esen a p oblem, bu in he case o di e en da ase s (numbe o measu emen s o da a ype), i is a ac o ha needs o be aken in o accoun . Con igu a ion (p1,s 1,p 2,s 2)=(2,1,2,2) achie ed he highes accu acy. A e he examina ion o he pa ial success a es inal ne wo k s uc u e was selec ed: •6 con olu ion il e 3 ×3, •hype bolic angen , •max pool 2 ×2, •7 con olu ion il e 3 ×3, •hype bolic angen , •max pool 2 ×2, • ully connec ed laye , •so max wi h 4 classes. This ne wo k con ains only 1,700 lea nable pa ame e s, which is a no ably smalle numbe in compa ison wi h 1 Inpu 6 con + anh 6 max pool 6 7 con + anh 42 max pool 42 ully conn. So max Ou pu Fig. 5. The esul ing CNN s uc u e. Bigge blocks ep esen ou pu om he p e ious laye wi h newly ob ained dimensions while small boxes illus a e il e s wi h hei numbe o channels and unc ionali y. Fig. 6. Success a es, po ayed by colo ba , based on he uned hype pa ame e s o mini-ba ch size 128. usually used CNNs o image p ocessing. Fo isualiza ion o he inal s uc u e see Fig. 5. 5.2 Fine uning To a ain op imal esul s i is essen ial o pay a en ion o ine uning. Fo a clea summa y o he uned pa ame e s and hei ange e e o Table 1. Fo each pa ame e , he de aul alue and i s su ounding in e al we e es ed. Fo ins ance, in mini-ba ch size powe s o wo a e commonly used o op imize he wo k wi h memo y, excep he las alue, which is he whole size o a aining se . Had he alue been se o 256, he aining se would ha e been passed emp y, as he algo i hm au oma ically o ms ba ches only i enough da a is a ailable. In Table 1, he in luence o he pa ame e s on he Acc dec eases om le o igh . Table 1. O e iew o uned hype pa ame e s Lea n a e Regula iza ion Epochs Mini-ba ch G adien decay 0.05 0.001 200 32 0.9 0.035 0.0002 700 64 0.8 0.02 0.0007 2000 128 0.7 0.01 4000 247 0.005 304 Michaela Mesá ošo á e al. / IFAC Pape sOnLine 58-9 (2024) 299–304 Fig. 6 shows ha he majo i y o esul s a e in he ange o 75–85%. This is also he case o he mini-ba ch o size 247. The bes -pe o ming con igu a ion eached 93.1% wi h a lea ning a e o 0.005, egula iza ion o 0.001, g adien decay o 0.8 and mini-ba ch size o 247 h oughou 4,000 epochs. The ou h plo in Fig. 6 is o e all he one wi h he highes c oss- alida ion Acc in he whole pa ame e space anging om nea ly 80–93%, while he es epoch combina ions go as low as 35%. Ano he pe cep ible phenomenon is he ac s a ed in Sec ion 2 ha he lea ning a e has he la ges impac on he classi ica ion esul s— he Acc alues a e mos ly changing along he lea ning a e axis. The second mos in luen ial pa ame e is egula iza ion and in combina ion wi h he lea ning a e i is ob ious, ha hei mu ual dec easing also wo sens he Acc. Hence i is impo an o no e ha wi h a small lea ning a e and a he low egula iza ion he Acc alue always signi ican ly d ops, due o insu icien aining o excessi ely complex models con aining high weigh alues. I we compa e he Acc ha di e only in hei g adien decay we obse e ha hey a e p ac ically he same, implying ha g adien decay has a li le e ec on he esul ing Acc. Fo u u e expe imen s, his pa ame e may be le ou un il mo e e ec ual se ings a e es ablished. I may be es ed as one o he las uning op ions. The pe o mance o he p oposed ne wo k has been com- pa ed o he esul s ob ained by ans e lea ning employ- ing GoogLeNe , ResNe and Squeezene . In all cases, he p oposed small-scale ne wo k pe o med be e achie ing highe accu acy. This con i ms he conside a ion om Sec- ion 3 ha a simple ne wo k is be e sui ed o his ype o ask. Conside able di e ence also lies in he ime needed o ain he ne wo k, whe e he shallow s uc u e wi h 1,700 lea nable pa ame e s is ained much as e in compa ison o ans e lea ning. 6. CONCLUSION This pape discusses he possibili ies o aining a small con olu ional neu al ne wo k while ha ing a ailable only a small da a se . Al hough CNNs ypically need housands o da a samples o ope a e eliably o o p o ide sa is ac o y accu acy, we ha e been able o ind an app op ia e ne - wo k s uc u e design and disco e co esponding op imal pa ame e s h ough ex ensi e ine- uning expe imen s. Lea ning a e and egula iza ion ha e shown he s onges impac on he o e all accu acy. Numbe o epochs also has a s ong e ec . Howe e , his a ec s only whe he he model had enough ime o be ained o i he model s a s o o e - i on gi en da a. Fo his eason, i is ap o obse e he aining p ocess in he ea ly s ages o design. Mini-ba ch has shown a mode a e in luence and g adien descen had only a mino in luence on he inal accu acy. The e o e, hese wo hype pa ame e s could be uned a he end o he ne wo k design. Wi h he highes eached accu acy being 93.1%, we ha e shown he applicabili y o CNNs o his ype o classi ica ion p oblem. As a pa o u u e wo k, he in luence o o he pa ame e s such as lea ning a e scheduling o squa ed g adien decay ac o could be examined. Addi ional expansion o he pa ame e space such as he numbe o il e s used in con olu ional laye s may also un eil ne wo k s uc u e wi h imp o ed classi ica ion pe o mance. ACKNOWLEDGEMENTS The comple ion o his pape was made possible by he g an No. FEKT-S-23-8451 – “Resea ch on ad anced me hods and echnologies in cybe ne ics, obo ics, a i i- cial in elligence, au oma ion and measu emen ” inancially suppo ed by he In e nal science und o B no Uni e si y o Technology. REFERENCES Beale, M.H., Hagan, M.T., and Demu h, H.B. (2023). Deep lea ning oolbox e e ence. Bengio, Y. (2012). P ac ical ecommenda- ions o g adien -based aining o deep a chi ec u es. CoRR, abs/1206.5533. URL h p://a xi .o g/abs/1206.5533. Bishop, C.M. and Bishop, H. (2023). Deep Lea ning: Founda ions and Concep s. Sp inge , Cham, 1s ed. edi ion. Cle e , D.A., Un e hine , T., and Hoch ei e , S. (2016). Fas and accu a e deep ne wo k lea ning by exponen ial linea uni s (elus). Fe ey a-Rami ez, A., A iles-C uz, C., Rod iguez- Ma inez, E., Villegas-Co ez, J., and Zu˜niga-Lopez, A. (2019). An imp o ed con olu ional neu al ne wo k a chi ec u e o image classi ica ion. In Lec u e No es in Compu e Science, olume 11524, 89–101. Sp inge In e na ional Publishing, Cham. Good ellow, I., Bengio, Y., and Cou ille, A. (2016). Deep Lea ning. The MIT P ess. Hinkelmann, K. (2018). Neu al ne wo ks, p. 7. Kesha i, R., Va sa, M., Singh, R., and Noo e, A. (2018). Lea ning s uc u e and s eng h o cnn il e s o small sample size aining. In 2018 IEEE/CVF Con e ence on Compu e Vision and Pa e n Recogni ion, 9349–9358. doi:10.1109/CVPR.2018.00974. Kingma, D.P. and Ba, J. (2015). Adam: A me hod o s ochas ic op imiza ion. URL h p://a xi .o g/abs/1412.6980. Maas, A.L. (2013). Rec i ie nonlinea i ies imp o e neu al ne wo k acous ic models. McCulloch, W.S. and Pi s, W. (1943). A logical calculus o he ideas immanen in ne ous ac i i y. Bulle in o Ma hema ical Biology, 5, 115–133. Nai , V. and Hin on, G.E. (2010). Rec i ied linea uni s imp o e es ic ed bol zmann machines. In P oceedings o he 27 h In e na ional Con e ence on Machine Lea n- ing, 807–814. Reed, R. and Ma ks, R.J. (1999). Neu al Smi hing: Supe ised Lea ning in Feed o wa d A i icial Neu al Ne wo ks. The MIT P ess. doi: 10.7551/mi p ess/4937.001.0001. Tieleman, T. and Hin on, G. (2012). Lec u e 6.5- msp op, cou se a: Neu al ne wo ks o machine lea ning. Uni- e si y o To on o, Technical Repo . Zaknich, A. (2003). Neu al ne wo ks o in elligen signal p ocessing. Wo ld Scien i ic, New Je sey. Zhou, Y.T. and Chellappa, R. (1988). Compu a ion o op ical low using a neu al ne wo k. IEEE 1988 In e na ional Con e ence on Neu al Ne wo ks, 2, 71–78. doi:10.1109/ICNN.1988.23914.