scieee Open visual document viewer

Automatic segmentation of infant cry signals using hidden Markov models

Naithani, Gaurav,Kivinummi, Jaana,Virtanen, Tuomas,Tammela, Outi,Peltola, Mikko J,Leppänen, Jukka M

Full text

Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 DOI 10.1186/s13636-018-0124-x RESEARCH Open Access Au oma ic segmen a ion o in an c y signals using hidden Ma ko models Gau a Nai hani1*†, Jaana Ki inummi2†, Tuomas Vi anen1, Ou i Tammela3, Mikko J. Pel ola4 and Jukka M. Leppänen2 Abs ac Au oma ic ex ac ion o acous ic egions o in e es om eco dings cap u ed in ealis ic clinical en i onmen s is a necessa y p ep ocessing s ep in any c y analysis sys em. In his s udy, we p opose a hidden Ma ko model (HMM) based audio segmen a ion me hod o iden i y he ele an acous ic pa s o he c y signal (i.e., expi a o y and inspi a o y phases) om eco dings made in na u al en i onmen s wi h a ious in e e ing acous ic sou ces. We examine and op imize he pe o mance o he sys em by using di e en audio ea u es and HMM opologies. In pa icula , we p opose using undamen al equency and ape iodici y ea u es. We also p opose a me hod o adap ing he segmen a ion sys em ained on acous ic ma e ial cap u ed in a pa icula acous ic en i onmen o a di e en acous ic en i onmen by using ea u e no maliza ion and semi-supe ised lea ning (SSL). The pe o mance o he sys em was e alua ed by analyzing a o al o 3 h and 10 min o audio ma e ial om 109 in an s, cap u ed in a a ie y o eco ding condi ions in hospi al wa ds and clinics. The p oposed sys em yields ame-based accu acy up o 89.2%. We conclude ha he p oposed sys em o e s a solu ion o au oma ed segmen a ion o c y signals in c y analysis applica ions. Keywo ds: In an c y analysis, Acous ic analysis, Audio segmen a ion, Hidden Ma ko models, Model adap a ion 1 In oduc ion Fo se e al decades, he e has been an ongoing in e - es in he connec ion o acous ic cha ac e is ics o in an c y ocaliza ions wi h in an heal h and de elopmen- al issues [1–4]. A ypicali ies in speci ic ea u es o c y (e.g., undamen al equency) ha e been linked wi h diag- nosed condi ions such as au ism, de elopmen al delays, and ch omosome abno mali ies [5,6]andwi h isk ac- o s such as p ema u i y and p ena al d ug exposu e [7,8]. These indings ha e gene a ed hope ha c y analysis may o e a cos -e ec i e [9], low- isk, and non-in asi e [10,11] diagnos ic echnique o ea ly iden i ica ion o child en wi h de elopmen al and heal h p oblems. The need o de ec ing heal h p oblems and isks (e.g., as poin ed ou by [6])asea lyaspossibleisimpo an because he plas ici y o he de eloping b ain and he sensi i e pe iods o skill o ma ion a he e y ea ly age *Co espondence: [email p o ec ed] †Equal con ibu o s 1Depa men o Signal P ocessing, Tampe e Uni e si y o Technology, Ko keakoulunka u 10, Tampe e, Finland Full lis o au ho in o ma ion is a ailable a he end o he a icle o e he bes chances o suppo op imal de elopmen by ehabili a ion and medical ca e [12–16]. An in an c y signal consis s o a se ies o expi a ions and inspi a ions sepa a ed by bou s o silence. These will be e e ed o as expi a o y and inspi a o y phases in his pape . A c y signal cap u ed in a ealis ic en i onmen (e.g., pedia ic wa d o a hospi al) may con ain ex aneous sounds (e.g., non-c y ocals p oduced by he in an , ocals o o he people p esen in he oom, and backg ound noise con ibu ed by he su ounding en i onmen o by he eco ding equipmen i sel ). A c y signal eco ding can hus be hough o being composed o wha we call he egions o in e es , namely, expi a o y and inspi a o y phases, and ex aneous egions consis ing he es o he audio ac i i y con ained in he eco ding, e med as esid- ual in his pape . Figu e 1is an example o a chunk o a c y eco ding cap u ed in hospi al en i onmen . In ealis ic clinical si ua ions, he eco dings a e a ec ed by he acous ic en i onmen including he oom acous- ics and o he sound sou ces p esen du ing he eco ding. The c y signal i sel is a ec ed by se e al ac o s ela ed o © The Au ho (s). 2018 Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0 In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e Commons license, and indica e i changes we e made. Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 2 o 14 0 0.5 1 1.5 2 2.5 3 3.5 44.5 5 −1 −0.8 −0.6 −0.4 −0.2 0 0.2 0.4 0.6 0.8 1 Time (s) Ampli ude Expi a o y phase Inspi a o y phase Non c y ocals Fig. 1 An example o a chunk o in an c y signal showing expi a o y and inspi a o y phases and non-c y ocals p esen in he eco ding (ca ego ized as esidual) he s a e o in an s’ heal h and de elopmen [6], age [17], size [18], eason o c ying [19], and a ousal [20]. In a c y analysis sys em mean o wo k wi h eco dings cap u ed in ealis ic en i onmen s, he e is o en a need o a p e-p ocessing sys em which is able o di e en i- a e he egions o in e es (i.e., expi a o y and inspi a o y phases) om ex aneous acous ic egions (i.e., esidual). The need o iden i ying he expi a o y and inspi a o y phases as sepa a e classes a ises om he ac ha hey di e in hei p ope ies, e.g., undamen al equency, ha - monici y, and ime du a ion. Success ul ex ac ion o hese a e o signi ican in e es when he sys em ou pu is used as a diagnos ic ool whe e he ela ion o hese p ope ies o in an neu o-de elopmen al ou comes can be explo ed. Manual anno a ion o c y eco dings is p one o e o s and is ende ed un easible when he numbe o eco dings o be anno a ed is la ge. The segmen a ion mechanism in any such c y analysis sys em hus needs o be au o- ma ic and should be able o wo k wi h ma e ial cap u ed in di e se eco ding con ex s. Va ious me hods ha e been p e iously used in he ield o in an c y analysis o deal wi h he p oblem o iden- i ying he use ul acous ic egions om c y eco dings, o example, manual selec ion o oiced pa o eco d- ings [4,21], ceps al analysis o oicing de e mina ion [22], ha monic p oduc spec um (HPS) based me hods [23], sho - e m ene gy (STE) his og am based me hods [24,25], and k-nea es neighbo algo i hm based de ec- ion [26]. Mos o hese me hods ha e ea ed inspi a o y phases as noise, and p ima y a en ion has been ocused on ex ac ion o expi a o y phases. The ele ance o ana omical and physiological bases o inspi a o y phona- ion has been poin ed ou by G au e al. [27]. P e iously, Aucou u ie e al. [28] ha e used he hidden Ma ko model (HMM) based me hod o au oma ic segmen a ion o c y signals while ea ing inspi a o y ocaliza ion as a sepa a e class. They u ilized s anda d mel- equency cep- s al coe icien s (MFCC) as audio ea u es and employed HMM opology consis ing o a single s a e o each a - ge class. Abou-Abbas e .al [29]p oposedasimila HMM based me hod u ilizing del a and del a-del a ea u es along wi h MFCCs and expe imen ing wi h mo e numbe o HMM s a es o each class. Simila ly, Abou-Abbas e .al. [30] p oposed c y segmen a ion using di e en signal decomposi ion echniques. Hidden Ma ko models o c y classi ica ion ins ead o de ec ion ha e been s udied by Lede man e .al. [31,32]. In his pape , we p opose an HMM based me hod o iden i ying use ul acous ic egions om c y eco dings cap u ed unde di e se eco ding condi ions. The di e - si y o eco ding condi ions includes acous ic condi ions o eco ding, ypes o c y igge , and in an - ela ed ac- o s which a e known o a ec acous ic cha ac e is ics o c y. Sec ions 4.1 and 4.2 desc ibe his in de ail. The wo k p esen ed he e dis inguishes i sel om simila p e ious e o s by p oposing he use o undamen al equency and ape iodici y (see Sec ion 2.1) as audio ea u es in addi ion o con en ionally used ea u es, e.g., MFCCs and hei i s and second o de de i a i es. We show ha his yields an imp o emen in segmen a ion pe o mance. Mo eo e , we show ha he p oposed sys em is able o adap o ma e- ial eco ded in unseen acous ic en i onmen s o which Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 3 o 14 i has no been ained. We use a combina ion o ea- u e no maliza ion and semi-supe ised lea ning o his adap a ion p oblem. The pape ollows he ollowing s uc u e. Sec ion 2 explains he implemen a ion o he p oposed sys em, Sec ion 3explains he model adap a ion echniques, Sec ion 4desc ibes he da a used in expe imen s, Sec ion 5desc ibes he e alua ion and p esen s he ob ained esul s, and, inally, Sec ion 6p o ides some concluding ema ks wi h sugges ions o u u e di ec ions o his wo k. 2 P oposed me hod In o de o analyze in an c y eco dings cap u ed in ealis ic en i onmen s con aining in e e ing sou ces, he goal is o segmen c y eco dings in o h ee classes, namely, expi a o y phases, inspi a o y phases, and esid- ual. The esidual class consis s o all acous ic egions in he c y eco ding excep he ones co e ed by he o he wo classes. A supe ised pa e n ecognize based on hidden Ma ko models (HMM) wi h Gaussian mix- u e model (GMM) densi ies [33] is used o segmen- a ion. An HMM is a s a is ical model which models a gene a i e ime sequence cha ac e ized by an unde ly- ing hidden s ochas ic p ocess gene a ing an obse able sequence [34]. HMMs ha e been widely used in au o- ma ic speech ecogni ion (e.g., [35]) o model a iabili y in speech caused by di e en speake s, speaking s yles, ocabula ies, and en i onmen s. Figu e 2depic s he block diag am o he segmen a ion p ocess. Each c y eco ding unde in es iga ion is di ided in o windowed o e lapping sho ime ames. Fo each such ame, he HMM pa e n ecognize ou pu s a se o obse a ion p obabili ies o he h ee classes being ac i e in ha ame. These p obabili ies a e decoded using he Vi e bi algo i hm. Decoding he e e e s o he p ocess o inding he bes pa h in he sea ch space o unde lying HMM s a es ha gi es maximum likelihood o he acous- ic ea u e ec o s om he c y signal unde in es iga ion. I ou pu s a class label o each ame o he signal, and his in o ma ion is used o iden i y he egions o in e es (i.e., expi a o y and inspi a o y phases) in he c y sig- nal. The o e all implemen a ion can hus be desc ibed in h ee s ages, namely, ea u e ex ac ion, HMM aining, and Vi e bi decoding. These s ages a e desc ibed in he ollowing subsec ions. 2.1 Fea u e ex ac ion Mel- equency ceps al coe icien s (MFCC) a e used as he p ima y audio ea u es. They ha e been widely used in audio signal p ocessing p oblems, o example, speech ecogni ion [36], audio e ie al [37], and emo ion ecog- ni ion in speech [38]. F ame du a ion o 25 ms wi h 50% o e lap be ween consecu i e ames and Hamming window unc ion was used o ex ac ing he MFCCs. Fo each ame o he signal, a 13-dimensional MFCC ea u e ec o , x=[x1,x2....., x13]T,isex ac ed,which includes he ze o h MFCC coe icien ; he e, T ep esen s he ma ix anspose. The sampling equency o each audio signal is 48 kHz. In conjunc ion wi h MFCCs, he ollowing addi ional ea u es a e in es iga ed. 1. Del as and del a-del as : MFCCs a e s a ic ea u es and p o ide a compac spec al ep esen a ion o only he co esponding ame. Tempo al e olu ion o hese ea u es migh be use ul o segmen a ion pu poses since HMMs assume each ame o be condi ionally independen o he p e ious ones gi en he p esen s a e. This empo al dynamics is cap u ed by compu ing he ime de i a i es o MFCCs, known as del a ea u es. Simila ly, empo al dynamics o del a ea u es can be cap u ed by compu ing hei ime de i a i es, known as del a-del a ea u es. Fo 13 MFCCs pe ame, we ha e 13 del a coe icien s and 13 del a-del a coe icien s. The use o hese ime de i a i es also means ha he sys em is non-causal. 2. Fundamen al equency (F0): Inspi a o y phases a e known o ha e highe undamen al equency (F0) han expi a o y phases [27]. This p ope y can be exploi ed o segmen a ion pu poses by including F0 Fig. 2 Block diag am o he audio segmen a ion sys em Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 4 o 14 as an audio ea u e. The YIN algo i hm [39]isa popula pi ch es ima ion algo i hm, which has been ound o pe o m well in he con ex s o speech [40] and music [41] signals. A eely a ailable MATLAB implemen a ion o he algo i hm is used in he p oposed sys em [42], and one F0 alue is ob ained o each ame. We ound YIN algo i hm o be sui able o his da ase as F0 alues we e empi ically ound o be be ween 200 and 800 Hz and e y ew ins ances o hype phona ion (F0<1000 Hz) [6] we e obse ed. 3. Ape iodici y : Ape iodici y in his s udy e e s o he p opo ion o ape iodic powe in he signal ame and is compu ed h ough he YIN algo i hm. In o de o compu e an F0es ima e, he YIN algo i hm employs a unc ion known as cumula i e mean no malized di e ence unc ion . The minima o his unc ion ha subsc ibes o ce ain condi ions gi es an es ima e o he undamen al pe iod o he signal ame. The alue o he unc ion a his minima is p opo ional o he ape iodic powe con ained in he signal ame. A de ailed ma hema ical ea men can be ound om he o iginal pape [39]. One ape iodici y alue is ob ained co esponding o each ame. 2.2 C y modeling using HMMs The a ailable da ase is manually anno a ed and di ided in o aining and es se s, as will be desc ibed in de ail in Sec ion 4. Fea u es ex ac ed om all audio iles in he aining da ase o a pa icula a ge class a e conca e- na ed o gi e aining ea u e ma ix Xi,ibeing he index o he a ge class. Using hese ea u e ma ices, h ee sep- a a e HMM models a e ained co esponding o he h ee a ge classes: expi a o y phases, inspi a o y phases, and esidual. The p obabili y densi y unc ion (pd ) o each HMM s a e is modeled wi h Gaussian mix u e models (GMMs). T aining in ol es es ima ing HMM pa ame e s, λi(i.e, weigh , mean, and co a iance o componen Gaussians and s a e ansi ion p obabili ies), which bes i s he aining da a Xi. P obabilis ically, i is amed as p oblem o maximizing p obabili y o an HMM model gi en he aining da a Xi, which in u n can be amed as maximum likelihood es ima ion p oblem, i.e., λop i=a g max λ P(Xi|λi)(1) whe e λop iindica es he op imal model o i h class. Fo his, he s anda d Baum-Welch algo i hm [43], an expec- a ion maximiza ion algo i hm used o es ima e HMM pa ame e s, is used. AHTO oolbox o he Audio Resea ch G oup, Tampe e Uni e si y o Technology, is used o his pu pose. Fully connec ed HMMs a e used wi h each s a e ha ing equal ini ial s a e p obabili y. I also means all en ies o ini ial s a e ansi ion p obabili y ma ix a e non-ze o and equal. Fo s a e means and co a iances, k-means clus e ing ini ializa ion is used. Twopa ame e sha e obechosen o eachHMM:S, he numbe o s a es used o adequa ely model he class, and C, he numbe o Gaussian componen s in he co e- sponding GMM used o model each s a e o he HMM. The e ec o bo h hese pa ame e s on sys em pe o - mance has been in es iga ed and will be discussed in Sec ion 5. The numbe o s a es and componen Gaussians in he h ee HMMs a e deno ed by Sexp and Cexp,Sins and Cins,andS es and C es o expi a o y phase, inspi- a o y phase, and esidual, espec i ely. HMMs ained o he h ee a ge classes a e hen combined o o m a single HMM ha ing a combined s a e space and ansi- ion p obabili y ma ix. S a e ansi ions om any s a e o one model o any s a e o ano he model a e possible, in o he wo ds, he combined HMM model is ully con- nec ed. The combined model has a ansi ion p obabili y ma ix ha ing dimensions (Sexp +Sins +S es)×(Sexp + Sins +S es). The p obabili y o ansi ion om one model o ano he depends upon model p io s and in e -model ansi ion penal y, a pa ame e simila o HTK oolki ’s [44]wo d ansi ion penal y pa ame e . In e -model an- si ion penal y penalizes model ansi ion om one model o ano he and has o be empi ically de e mined (we ha e used a alue o −1 in his pape ). The model p io s a e calcula ed simply by coun ing he occu ences o he co esponding class om he anno a ed da a. HMM pa ame e s o his combined model a e used o Vi e bi decoding o obse a ion p obabili y ou pu s in he ollowing sec ion. Figu e 3depic s he combined HMM. 2.3 Vi e bi decoding Fea u esex ac ed om hec y eco ding obesegmen ed a e ed o he h ee HMM models, each ained o a pa icula a ge class. Fo each ame o he eco ding, Fig. 3 HMM models o indi idual classes a e combined in o a single ully connec ed HMM Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 5 o 14 he HMM ou pu s he p obabili ies o i s cons i uen s a es being ac i e in ha ame. Th ee obse a ion p ob- abili y ma ices a e gene a ed co esponding o h ee HMMs which a e combined in o a single ma ix Ocomb as depic ed in Fig. 4. The Vi e bi algo i hm is employed upon his combined obse a ion p obabili y ma ix using he pa ame e s lea ned o combined HMM model in he p e- ious sec ion. The algo i hm maximizes he p obabili y o occu ence o s a e sequence qgi en a lea ned HMM λcomb and obse a ion p obabili y ma ix Ocomb, i.e., qop =a g max q P(q|Ocomb,λcomb)(2) whe e qop is he s a e sequence gi ing maximum likeli- hood h ough he combined HMM s a e space which now consis s o (Sexp +Sins +S es)s a es. This ou pu sequence consis s o a s a e assignmen o each ame o he eco d- ing, which can u he be used o gi e co esponding class assignmen o each ame. I is done by iden i ying he con ibu ing HMM co esponding o he chosen s a e o ha ame. AHTO oolbox is used o Vi e bi decoding. Figu e 4shows he implemen a ion o he audio segmen- a ion sys em. The segmen a ion esul s o a 5-s chunk o a c y signal is depic ed in Fig. 5. 3 Model adap a ion The p oposed audio segmen a ion sys em is ained on a da ase eco ded in a pa icula acous ic en i onmen and may no necessa ily be able o gene alize o a da ase cap- u ed in a di e en acous ic en i onmen . In his sec ion, we will show ha he p oposed sys em can be made o wo k o da a eco ded in unseen acous ic en i onmen s as well. We will ain ou sys em on da a eco ded in a known acous ic en i onmen and use i p edic class labels on da a eco ded in an unseen acous ic en i onmen . Ou p oposed solu ion consis s o wo s ages: ea u e no mal- iza ion [45] and semi-supe ised lea ning. These will be desc ibedinde ailin he ollowingsubsec ions. 3.1 Fea u e no maliza ion Fea u es ex ac ed om an audio ile a e no malized by sub ac ing he mean and di iding i by he s anda d de ia ion be o e eeding i o he HMM. The mean and s anda d de ia ion ec o s a e de i ed o each audio ile sepa a ely. This is epea ed o each audio ile p esen in he aining da a ( om known en i onmen ) as well as he es da a ( om unknown en i onmen ). Fo a ea u e ec o Fjn ex ac ed om j h ame o n h audio ile, he no malized ea u e ec o is gi en by Fs d jn =Fjn −μn σn (3) whe e μnand σna emean ec o ands anda dde ia ion ec o , espec i ely, de i ed o n h audio ile. The di ide ope a ion he e is elemen -wise. 3.2 Semi-supe ised lea ning A semi-supe ised lea ning (SSL) me hod, known as sel aining [46], is used o u he adap he HMM models o an unseen acous ic en i onmen . In a classical SSL p ob- lem, we ha e wo da ase s: labeled da a ( om a known acous ic en i onmen ) and unlabeled da a ( om a new acous ic en i onmen ). The idea behind his me hod is o gene a e addi ional labeled aining da a using he unla- beled da a comp ised o audio iles eco ded in he unseen acous ic en i onmen . The ou pu labels gene a ed by he model o he unlabeled da a a e ea ed as ue labels, and he models a e e ained using he combina ion o o ig- inal aining da a and his newly gene a ed labeled da a. Figu e 6depic s his p ocess. Al e na i ely, ins ead o using he en i e unlabeled da a, a selec ion o only hose ames can be made o which we a e con iden o he assigned label being ue. The likeli- hoods ou pu ed by HMMs co esponding o h ee a ge classes may be used o de ise a con idence c i e ion. In Fig. 4, we ha e h ee likelihood ma ices co esponding o each a ge class o each es ile. The maximum likeli- hood o each column o he h ee ma ices is calcula ed. The a io be ween he maximum and second la ges alue oughly ep esen s how con iden we can be abou he classi ica ion esul o a pa icula ame. We will e e i as he con idence h eshold. Only hose ames o which his a io exceeds a ce ain h eshold a e chosen. A con- idence h eshold o 2 was used in his wo k. Figu e 7 shows he p ocedu e o selec ing da a based on con idence h eshold. Da a selec ed his way can be used as addi- ional aining da a o HMMs co esponding o he h ee classes. 4 Acous ic ma e ial Fo his s udy, we collec ed c y eco dings om wo coho s o in an s in Tampe e, Finland, and in Cape Town, Sou h A ica. The ollowing subsec ions desc ibe hese wo da abases and e alua ion o he pe o mance o he audio segmen a ion sys em on hem. 4.1 Da abase: Tampe e coho In Tampe e, Finland, we cap u ed he eco dings a Ma e - ni y Wa d Uni s and Neona al Wa d Uni o Tampe e Uni- e si y Hospi al. The eco ding pe iod was om Ap il 13 o Augus 3, 2014. The s udy ollowed he s ipula ed e hi- cal guidelines and was app o ed by he E hical Commi ee o Tampe e Uni e si y Hospi al. The coho consis ed o a he e ogeneous g oup o 57 neona es whose ch onological ages (i.e., he ime elapsed since bi h) a eco ding we e om0 o5daysasdepic edinTable1. The coho was no Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 6 o 14 Fig. 4 Implemen a ion o he audio segmen a ion sys em. The inpu o he sys em is a ea u e ma ix de i ed om a es audio ile. The ou pu is a class label assigned o each ame o he c y signal 0 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 −0.04 −0.02 0 0.02 0.04 Ampli ude 0 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 Exp Ins Res Time (s) Class Labels Fig. 5 Audio segmen a ion esul s o a chunk o c y signal shown in he op panel. The bo om panel depic s he ac ual and p edic ed class labels wi h blue and ed plo s, espec i ely Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 7 o 14 Fig. 6 Semi-supe ised lea ning block diag am s anda dized because he a ge o he p esen s udy was o de elop a obus ool o iden i ying in an c y sounds in he cap u ed eco dings o gene al neona e popula ion. In o de o minimize he in luence o lea ning and ma u- a ion on c y cha ac e is ics, he age o he in an s was he only s anda dized a iable in he coho . Thec ysampleswe ecap u edina a ie yo eco d- ing condi ions. Fi s ly, he place o eco ding and he associa ed acous ic en i onmen a ied signi ican ly. I included he hospi al co ido , no mal pedia ic wa d, in ensi e ca e uni (ICU), wai ing oom, and nu se’s o ice. Wi hin each oom, eco dings we e cap u ed a di e en places (e.g., mo he ’s bed, weighing scales, and in an ’s bed). Secondly, he backg ound sounds p esen in he eco ding consis ed o human oices (e.g., coughing and speaking) and mechanical sounds (e.g., sound o unning wa e , ai condi ioning, and diape ape being opened). Thi dly, in an - ela ed ac o s (e.g., weigh o he in an and p ema u i y o bi h) ha a e known o in luence he acous ic quali ies o c y a ied . Apa om he eco d- ing condi ions, he c y-ini ia ing igge also a ied. I included in asi e (e.g., enipunc u e) and non-in asi e (e.g., changing diape s and measu ing body empe a u e) ope a ions, as well as spon aneous c ies (e.g., due o hunge o a igue). All Tampe e eco dings we e s o ed as 48 kHz sam- pling a e, wo-channel audio in a 24-bi Wa e o m audio ile (WAV) o ma . The audio eco de used was Tascam DR-100MK II wi h RØDE M3 ca dioid mic ophone. Fo u he compu a ion, he mean o he wo channels was aken o yield he signal o be segmen ed. The dis ance be ween he in an ’s mou h and he eco de was kep a Fig. 7 Selec ion o da a based on con idence h eshold o each unlabeled audio ile. The model ou pu s p o ide he labels o semi-supe ised lea ning Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 8 o 14 Table 1 The ch onological ages o in an subjec s in he Tampe e coho No. o in an s Ch onological age (day) 10 11 1 29 2 10 3 34 25 1 Missing in o. app oxima ely 30 cm. Each eco ding was gi en a sepa- a e numbe code. The eco dings we e manually anno- a ed using Audaci y [47] applica ion o gene a e labels o aining he HMM models. Figu e 8is a snapsho o he Audaci y applica ion showing an example o a chunk o he labeled c y eco ding. The da abase o 57 manually anno a ed audio eco d- ings spans a ound 115 min in du a ion. A o al o 1529 expi a o y phases we e ound wi h a mean du a ion o 0.95 s and a s anda d de ia ion o 0.65 s. Simila ly, 1005 inspi a o y phases we e ound wi h a mean du a ion o 0.17 s and a s anda d de ia ion o 0.06 s. Figu e 9( op) illus a es he dis ibu ion o he ime du a ions o expi- a o y and inspi a o y phases o he Tampe e coho . No e ha inspi a o y phases we e ewe in numbe and sho e in du a ion as compa ed o expi a o y phases. Hence, less da a we e a ailable o aining he HMM o inspi a o y phases as compa ed o expi a o y phases. Mo eo e , i needs o be emphasized he e ha inspi a- o y phases exhibi ed mo e a ia ions h oughou he da a in compa ison o expi a o y phases. Fo example, on he one hand, we had eco dings wi h e y sho o almos no disce nible inspi a o y phases, and on he o he hand, we had eco dings which ha e unusually p ominen inspi a- o y phases as compa ed o expi a o y phases. I is also possible o obse e bo h hese ex eme cases wi hin he same eco ding. 4.2 Da abase: Cape Town coho The o he coho used o his s udy is being in es i- ga ed unde a la ge esea ch p ojec in coope a ion wi h he Depa men o Psychia y, Uni e si y o S ellenbosch, Cape Town. The da a we e collec ed in 2014 and consis ed o c y eco dings o 52 in an s whose age was less han 7 weeks (mean 33.5 days, s anda d de ia ion 3.5 days). The c y eco dings in his da abase we e also manually anno a ed using he Audaci y applica ion. The da abase o 52 manually anno a ed audio eco dings spans a ound 75 min in du a ion. A o al o 1307 expi a o y phases we e ound wi h a mean du a ion o 1.1 s and a s anda d de ia- ion o 0.76 s. Simila ly, 680 inspi a o y phases we e ound wi h a mean du a ion o 0.25 s and a s anda d de ia ion o 0.07 s. Figu e 9(bo om ) illus a es he dis ibu ion o he du a ions o expi a o y and inspi a o y phases o he Cape Town coho . Fig. 8 Snapsho o he Audaci y applica ion showing a manually anno a ed chunk o he c y eco ding. Expi a o y and inspi a o y phases a e coded by names exp_c y and insp_c y, espec i ely Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 9 o 14 0 1 2 3 4 0 50 100 150 200 Du a ion (s) No. o segmen s 00.1 0.2 0.3 0.4 0.5 0 50 100 150 200 250 Du a ion (s) No. o segmen s 0 1 2 3 4 0 50 100 150 200 Du a ion (s) No. o segmen s 00.1 0.2 0.3 0.4 0.5 0 50 100 150 Du a ion (s) No. o segmen s Fig. 9 Dis ibu ion o du a ions o expi a o y (le ) and inspi a o y ( igh ) phases o Tampe e ( op) and Cape Town (bo om) coho s In he Cape Town coho , he loca ion and p ocedu e o eco ding we e somewha mo e s anda dized han in he Tampe e coho (i.e., he eco dings we e cap u ed while conduc ing ou ine examina ions in he same nu sing oom). The c y igge used was accina ion (i.e., in a- si e) o measu emen o in an weigh a a weighing scale (i.e., non-in asi e). All Cape Town eco dings we e s o ed as 48-kHz sampling a e, wo-channel audio in a 24-bi Wa e o m audio ile (WAV) o ma . The audio eco de used was Zoom H4n eco de wi h buil -in condense mic ophones. The dis ance be ween in an ’s mou h and he eco de was app oxima ely 1.3 m o in an s being accina ed and 70 cm o in an s being weighed. Ou da a collec ion was conjoined wi h ano he s udy whose p o- ocol equi ed he mic o be a bi a and hence he la ge dis ance be ween he in an and he eco de as compa ed o he Tampe e coho . Due o guidelines o he p ojec conce ning p o ec ion o p i acy o he in ol ed pa ici- pan s, we a e no able o publish he audio da a used in his p ojec . 5 E alua ion The segmen a ion pe o mance was e alua ed using a i e- old c oss- alida ion amewo k. In he case o Tampe e coho , he a ailable da ase o 57 c y eco d- ings was di ided in o i e pa i ions: ou pa i ions o 12 eco dings each and one pa i ion o nine eco d- ings. In a simila manne , o he Cape Town coho , he da ase o 52 c y eco dings was di ided in o i e pa i- ions: ou pa i ions o 10 eco dings and one pa i ion o 12 eco dings. The di ision was done acco ding o c y codes assigned o he eco dings which co espond o he ch onological o de in which hey we e cap u ed. In each old, one o he pa i ions was used as he es se and he es o he pa i ions we e used o aining. Fi e such olds we e pe o med wi h each old ha ing a di e en pa i ion as he es se . The ou pu labels gene a ed by he sys em we e compa ed agains he manually anno a ed g ound u h. Fo each es ile unde in es iga ion, he ou pu labels p oduced by he model we e compa ed agains he g ound u h (i.e., manual anno a ions) o calcula e he pe o - mance me ics. Two me ics ha e been used in his s udy o e alua e he pe o mance o he sys em, namely, ame- based accu acy and ame-based Fsco e. The ame- based accu acy is de ined as accu acy =numbe o co ec ly labeled ames o al numbe o ames .(4) The ame-based Fsco eisde inedas heha monic mean o p ecision and ecall alues. P ecision is he a io o ue posi i e alue o he es ou come posi i es o a pa icula class. T ue posi i e alue is he numbe o ames co ec ly labeled by he sys em o a pa icula class, and es ou come posi i e alue is he numbe o ames de ec ed by he sys em belonging o ha class. Recall is he a io o ue posi i e alues o o al posi i e alues o any class. To al posi i e alues a e numbe o ames in he es se belonging o ha pa icula class. The ame-based Fsco e is hus gi en by Fsco e =2P·R P+R,(5)