Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music
P ocessing (2018) 2018:1
DOI 10.1186/s13636-018-0124-x
RESEARCH Open Access
Au oma ic segmen a ion o in an c y
signals using hidden Ma ko models
Gau a Nai hani1*†, Jaana Ki inummi2†, Tuomas Vi anen1, Ou i Tammela3, Mikko J. Pel ola4
and Jukka M. Leppänen2
Abs ac
Au oma ic ex ac ion o acous ic egions o in e es om eco dings cap u ed in ealis ic clinical en i onmen s is a
necessa y p ep ocessing s ep in any c y analysis sys em. In his s udy, we p opose a hidden Ma ko model (HMM)
based audio segmen a ion me hod o iden i y he ele an acous ic pa s o he c y signal (i.e., expi a o y and
inspi a o y phases) om eco dings made in na u al en i onmen s wi h a ious in e e ing acous ic sou ces. We
examine and op imize he pe o mance o he sys em by using di e en audio ea u es and HMM opologies. In
pa icula , we p opose using undamen al equency and ape iodici y ea u es. We also p opose a me hod o
adap ing he segmen a ion sys em ained on acous ic ma e ial cap u ed in a pa icula acous ic en i onmen o a
di e en acous ic en i onmen by using ea u e no maliza ion and semi-supe ised lea ning (SSL). The pe o mance
o he sys em was e alua ed by analyzing a o al o 3 h and 10 min o audio ma e ial om 109 in an s, cap u ed in a
a ie y o eco ding condi ions in hospi al wa ds and clinics. The p oposed sys em yields ame-based accu acy up o
89.2%. We conclude ha he p oposed sys em o e s a solu ion o au oma ed segmen a ion o c y signals in c y
analysis applica ions.
Keywo ds: In an c y analysis, Acous ic analysis, Audio segmen a ion, Hidden Ma ko models, Model adap a ion
1 In oduc ion
Fo se e al decades, he e has been an ongoing in e -
es in he connec ion o acous ic cha ac e is ics o in an
c y ocaliza ions wi h in an heal h and de elopmen-
al issues [1–4]. A ypicali ies in speci ic ea u es o c y
(e.g., undamen al equency) ha e been linked wi h diag-
nosed condi ions such as au ism, de elopmen al delays,
and ch omosome abno mali ies [5,6]andwi h isk ac-
o s such as p ema u i y and p ena al d ug exposu e
[7,8]. These indings ha e gene a ed hope ha c y analysis
may o e a cos -e ec i e [9], low- isk, and non-in asi e
[10,11] diagnos ic echnique o ea ly iden i ica ion o
child en wi h de elopmen al and heal h p oblems. The
need o de ec ing heal h p oblems and isks (e.g., as
poin ed ou by [6])asea lyaspossibleisimpo an
because he plas ici y o he de eloping b ain and he
sensi i e pe iods o skill o ma ion a he e y ea ly age
*Co espondence: [email p o ec ed]
†Equal con ibu o s
1Depa men o Signal P ocessing, Tampe e Uni e si y o Technology,
Ko keakoulunka u 10, Tampe e, Finland
Full lis o au ho in o ma ion is a ailable a he end o he a icle
o e he bes chances o suppo op imal de elopmen by
ehabili a ion and medical ca e [12–16].
An in an c y signal consis s o a se ies o expi a ions
and inspi a ions sepa a ed by bou s o silence. These will
be e e ed o as expi a o y and inspi a o y phases in his
pape . A c y signal cap u ed in a ealis ic en i onmen
(e.g., pedia ic wa d o a hospi al) may con ain ex aneous
sounds (e.g., non-c y ocals p oduced by he in an , ocals
o o he people p esen in he oom, and backg ound
noise con ibu ed by he su ounding en i onmen o by
he eco ding equipmen i sel ). A c y signal eco ding
can hus be hough o being composed o wha we call
he egions o in e es , namely, expi a o y and inspi a o y
phases, and ex aneous egions consis ing he es o he
audio ac i i y con ained in he eco ding, e med as esid-
ual in his pape . Figu e 1is an example o a chunk o a c y
eco ding cap u ed in hospi al en i onmen .
In ealis ic clinical si ua ions, he eco dings a e a ec ed
by he acous ic en i onmen including he oom acous-
ics and o he sound sou ces p esen du ing he eco ding.
The c y signal i sel is a ec ed by se e al ac o s ela ed o
© The Au ho (s). 2018 Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0
In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he
C ea i e Commons license, and indica e i changes we e made.
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 2 o 14
0 0.5 1 1.5 2 2.5 3 3.5 44.5 5
−1
−0.8
−0.6
−0.4
−0.2
0
0.2
0.4
0.6
0.8
1
Time (s)
Ampli ude
Expi a o y
phase
Inspi a o y
phase
Non c y
ocals
Fig. 1 An example o a chunk o in an c y signal showing expi a o y and inspi a o y phases and non-c y ocals p esen in he eco ding
(ca ego ized as esidual)
he s a e o in an s’ heal h and de elopmen [6], age [17],
size [18], eason o c ying [19], and a ousal [20].
In a c y analysis sys em mean o wo k wi h eco dings
cap u ed in ealis ic en i onmen s, he e is o en a need
o a p e-p ocessing sys em which is able o di e en i-
a e he egions o in e es (i.e., expi a o y and inspi a o y
phases) om ex aneous acous ic egions (i.e., esidual).
The need o iden i ying he expi a o y and inspi a o y
phases as sepa a e classes a ises om he ac ha hey
di e in hei p ope ies, e.g., undamen al equency, ha -
monici y, and ime du a ion. Success ul ex ac ion o hese
a e o signi ican in e es when he sys em ou pu is used
as a diagnos ic ool whe e he ela ion o hese p ope ies
o in an neu o-de elopmen al ou comes can be explo ed.
Manual anno a ion o c y eco dings is p one o e o s
and is ende ed un easible when he numbe o eco dings
o be anno a ed is la ge. The segmen a ion mechanism
in any such c y analysis sys em hus needs o be au o-
ma ic and should be able o wo k wi h ma e ial cap u ed
in di e se eco ding con ex s.
Va ious me hods ha e been p e iously used in he ield
o in an c y analysis o deal wi h he p oblem o iden-
i ying he use ul acous ic egions om c y eco dings,
o example, manual selec ion o oiced pa o eco d-
ings [4,21], ceps al analysis o oicing de e mina ion
[22], ha monic p oduc spec um (HPS) based me hods
[23], sho - e m ene gy (STE) his og am based me hods
[24,25], and k-nea es neighbo algo i hm based de ec-
ion [26]. Mos o hese me hods ha e ea ed inspi a o y
phases as noise, and p ima y a en ion has been ocused
on ex ac ion o expi a o y phases. The ele ance o
ana omical and physiological bases o inspi a o y phona-
ion has been poin ed ou by G au e al. [27]. P e iously,
Aucou u ie e al. [28] ha e used he hidden Ma ko
model (HMM) based me hod o au oma ic segmen a ion
o c y signals while ea ing inspi a o y ocaliza ion as a
sepa a e class. They u ilized s anda d mel- equency cep-
s al coe icien s (MFCC) as audio ea u es and employed
HMM opology consis ing o a single s a e o each a -
ge class. Abou-Abbas e .al [29]p oposedasimila HMM
based me hod u ilizing del a and del a-del a ea u es
along wi h MFCCs and expe imen ing wi h mo e numbe
o HMM s a es o each class. Simila ly, Abou-Abbas e .al.
[30] p oposed c y segmen a ion using di e en signal
decomposi ion echniques. Hidden Ma ko models o
c y classi ica ion ins ead o de ec ion ha e been s udied
by Lede man e .al. [31,32].
In his pape , we p opose an HMM based me hod o
iden i ying use ul acous ic egions om c y eco dings
cap u ed unde di e se eco ding condi ions. The di e -
si y o eco ding condi ions includes acous ic condi ions
o eco ding, ypes o c y igge , and in an - ela ed ac-
o s which a e known o a ec acous ic cha ac e is ics o
c y. Sec ions 4.1 and 4.2 desc ibe his in de ail. The wo k
p esen ed he e dis inguishes i sel om simila p e ious
e o s by p oposing he use o undamen al equency and
ape iodici y (see Sec ion 2.1) as audio ea u es in addi ion
o con en ionally used ea u es, e.g., MFCCs and hei
i s and second o de de i a i es. We show ha his yields
an imp o emen in segmen a ion pe o mance. Mo eo e ,
we show ha he p oposed sys em is able o adap o ma e-
ial eco ded in unseen acous ic en i onmen s o which
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 3 o 14
i has no been ained. We use a combina ion o ea-
u e no maliza ion and semi-supe ised lea ning o his
adap a ion p oblem.
The pape ollows he ollowing s uc u e. Sec ion 2
explains he implemen a ion o he p oposed sys em,
Sec ion 3explains he model adap a ion echniques,
Sec ion 4desc ibes he da a used in expe imen s,
Sec ion 5desc ibes he e alua ion and p esen s he
ob ained esul s, and, inally, Sec ion 6p o ides some
concluding ema ks wi h sugges ions o u u e di ec ions
o his wo k.
2 P oposed me hod
In o de o analyze in an c y eco dings cap u ed in
ealis ic en i onmen s con aining in e e ing sou ces, he
goal is o segmen c y eco dings in o h ee classes,
namely, expi a o y phases, inspi a o y phases, and esid-
ual. The esidual class consis s o all acous ic egions
in he c y eco ding excep he ones co e ed by he
o he wo classes. A supe ised pa e n ecognize based
on hidden Ma ko models (HMM) wi h Gaussian mix-
u e model (GMM) densi ies [33] is used o segmen-
a ion. An HMM is a s a is ical model which models a
gene a i e ime sequence cha ac e ized by an unde ly-
ing hidden s ochas ic p ocess gene a ing an obse able
sequence [34]. HMMs ha e been widely used in au o-
ma ic speech ecogni ion (e.g., [35]) o model a iabili y
in speech caused by di e en speake s, speaking s yles,
ocabula ies, and en i onmen s.
Figu e 2depic s he block diag am o he segmen a ion
p ocess. Each c y eco ding unde in es iga ion is di ided
in o windowed o e lapping sho ime ames. Fo each
such ame, he HMM pa e n ecognize ou pu s a se o
obse a ion p obabili ies o he h ee classes being ac i e
in ha ame. These p obabili ies a e decoded using he
Vi e bi algo i hm. Decoding he e e e s o he p ocess o
inding he bes pa h in he sea ch space o unde lying
HMM s a es ha gi es maximum likelihood o he acous-
ic ea u e ec o s om he c y signal unde in es iga ion.
I ou pu s a class label o each ame o he signal, and
his in o ma ion is used o iden i y he egions o in e es
(i.e., expi a o y and inspi a o y phases) in he c y sig-
nal. The o e all implemen a ion can hus be desc ibed in
h ee s ages, namely, ea u e ex ac ion, HMM aining,
and Vi e bi decoding. These s ages a e desc ibed in he
ollowing subsec ions.
2.1 Fea u e ex ac ion
Mel- equency ceps al coe icien s (MFCC) a e used as
he p ima y audio ea u es. They ha e been widely used
in audio signal p ocessing p oblems, o example, speech
ecogni ion [36], audio e ie al [37], and emo ion ecog-
ni ion in speech [38]. F ame du a ion o 25 ms wi h
50% o e lap be ween consecu i e ames and Hamming
window unc ion was used o ex ac ing he MFCCs.
Fo each ame o he signal, a 13-dimensional MFCC
ea u e ec o , x=[x1,x2....., x13]T,isex ac ed,which
includes he ze o h MFCC coe icien ; he e, T ep esen s
he ma ix anspose. The sampling equency o each
audio signal is 48 kHz. In conjunc ion wi h MFCCs, he
ollowing addi ional ea u es a e in es iga ed.
1.
Del as and del a-del as
: MFCCs a e s a ic ea u es
and p o ide a compac spec al ep esen a ion o
only he co esponding ame. Tempo al e olu ion o
hese ea u es migh be use ul o segmen a ion
pu poses since HMMs assume each ame o be
condi ionally independen o he p e ious ones gi en
he p esen s a e. This empo al dynamics is cap u ed
by compu ing he ime de i a i es o MFCCs, known
as del a ea u es. Simila ly, empo al dynamics o
del a ea u es can be cap u ed by compu ing hei
ime de i a i es, known as del a-del a ea u es. Fo
13 MFCCs pe ame, we ha e 13 del a coe icien s
and 13 del a-del a coe icien s. The use o hese ime
de i a i es also means ha he sys em is non-causal.
2.
Fundamen al equency
(F0): Inspi a o y phases a e
known o ha e highe undamen al equency (F0)
han expi a o y phases [27]. This p ope y can be
exploi ed o segmen a ion pu poses by including F0
Fig. 2 Block diag am o he audio segmen a ion sys em
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 4 o 14
as an audio ea u e. The YIN algo i hm [39]isa
popula pi ch es ima ion algo i hm, which has been
ound o pe o m well in he con ex s o speech [40]
and music [41] signals. A eely a ailable MATLAB
implemen a ion o he algo i hm is used in he
p oposed sys em [42], and one F0 alue is ob ained
o each ame. We ound YIN algo i hm o be
sui able o his da ase as F0 alues we e empi ically
ound o be be ween 200 and 800 Hz and e y ew
ins ances o hype phona ion (F0<1000 Hz) [6]
we e obse ed.
3.
Ape iodici y
: Ape iodici y in his s udy e e s o he
p opo ion o ape iodic powe in he signal ame and
is compu ed h ough he YIN algo i hm. In o de o
compu e an F0es ima e, he YIN algo i hm employs
a unc ion known as
cumula i e mean no malized
di e ence unc ion
. The minima o his unc ion ha
subsc ibes o ce ain condi ions gi es an es ima e o
he undamen al pe iod o he signal ame. The
alue o he unc ion a his minima is p opo ional
o he ape iodic powe con ained in he signal ame.
A de ailed ma hema ical ea men can be ound
om he o iginal pape [39]. One ape iodici y alue
is ob ained co esponding o each ame.
2.2 C y modeling using HMMs
The a ailable da ase is manually anno a ed and di ided
in o aining and es se s, as will be desc ibed in de ail
in Sec ion 4. Fea u es ex ac ed om all audio iles in he
aining da ase o a pa icula a ge class a e conca e-
na ed o gi e aining ea u e ma ix Xi,ibeing he index
o he a ge class. Using hese ea u e ma ices, h ee sep-
a a e HMM models a e ained co esponding o he h ee
a ge classes: expi a o y phases, inspi a o y phases, and
esidual. The p obabili y densi y unc ion (pd ) o each
HMM s a e is modeled wi h Gaussian mix u e models
(GMMs).
T aining in ol es es ima ing HMM pa ame e s, λi(i.e,
weigh , mean, and co a iance o componen Gaussians
and s a e ansi ion p obabili ies), which bes i s he
aining da a Xi. P obabilis ically, i is amed as p oblem
o maximizing p obabili y o an HMM model gi en he
aining da a Xi, which in u n can be amed as maximum
likelihood es ima ion p oblem, i.e.,
λop
i=a g max
λ
P(Xi|λi)(1)
whe e λop
iindica es he op imal model o i h class. Fo
his, he s anda d Baum-Welch algo i hm [43], an expec-
a ion maximiza ion algo i hm used o es ima e HMM
pa ame e s, is used. AHTO oolbox o he Audio Resea ch
G oup, Tampe e Uni e si y o Technology, is used o
his pu pose. Fully connec ed HMMs a e used wi h each
s a e ha ing equal ini ial s a e p obabili y. I also means
all en ies o ini ial s a e ansi ion p obabili y ma ix a e
non-ze o and equal. Fo s a e means and co a iances,
k-means clus e ing ini ializa ion is used.
Twopa ame e sha e obechosen o eachHMM:S,
he numbe o s a es used o adequa ely model he class,
and C, he numbe o Gaussian componen s in he co e-
sponding GMM used o model each s a e o he HMM.
The e ec o bo h hese pa ame e s on sys em pe o -
mance has been in es iga ed and will be discussed in
Sec ion 5. The numbe o s a es and componen Gaussians
in he h ee HMMs a e deno ed by Sexp and Cexp,Sins
and Cins,andS es and C es o expi a o y phase, inspi-
a o y phase, and esidual, espec i ely. HMMs ained
o he h ee a ge classes a e hen combined o o m a
single HMM ha ing a combined s a e space and ansi-
ion p obabili y ma ix. S a e ansi ions om any s a e
o one model o any s a e o ano he model a e possible,
in o he wo ds, he combined HMM model is ully con-
nec ed. The combined model has a ansi ion p obabili y
ma ix ha ing dimensions (Sexp +Sins +S es)×(Sexp +
Sins +S es). The p obabili y o ansi ion om one model
o ano he depends upon model p io s and in e -model
ansi ion penal y, a pa ame e simila o HTK oolki ’s
[44]wo d ansi ion penal y pa ame e . In e -model an-
si ion penal y penalizes model ansi ion om one model
o ano he and has o be empi ically de e mined (we
ha e used a alue o −1 in his pape ). The model p io s
a e calcula ed simply by coun ing he occu ences o he
co esponding class om he anno a ed da a.
HMM pa ame e s o his combined model a e used o
Vi e bi decoding o obse a ion p obabili y ou pu s in he
ollowing sec ion. Figu e 3depic s he combined HMM.
2.3 Vi e bi decoding
Fea u esex ac ed om hec y eco ding obesegmen ed
a e ed o he h ee HMM models, each ained o a
pa icula a ge class. Fo each ame o he eco ding,
Fig. 3 HMM models o indi idual classes a e combined in o a single
ully connec ed HMM
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 5 o 14
he HMM ou pu s he p obabili ies o i s cons i uen
s a es being ac i e in ha ame. Th ee obse a ion p ob-
abili y ma ices a e gene a ed co esponding o h ee
HMMs which a e combined in o a single ma ix Ocomb as
depic ed in Fig. 4. The Vi e bi algo i hm is employed upon
his combined obse a ion p obabili y ma ix using he
pa ame e s lea ned o combined HMM model in he p e-
ious sec ion. The algo i hm maximizes he p obabili y
o occu ence o s a e sequence qgi en a lea ned HMM
λcomb and obse a ion p obabili y ma ix Ocomb, i.e.,
qop =a g max
q
P(q|Ocomb,λcomb)(2)
whe e qop is he s a e sequence gi ing maximum likeli-
hood h ough he combined HMM s a e space which now
consis s o (Sexp +Sins +S es)s a es. This ou pu sequence
consis s o a s a e assignmen o each ame o he eco d-
ing, which can u he be used o gi e co esponding class
assignmen o each ame. I is done by iden i ying he
con ibu ing HMM co esponding o he chosen s a e o
ha ame. AHTO oolbox is used o Vi e bi decoding.
Figu e 4shows he implemen a ion o he audio segmen-
a ion sys em. The segmen a ion esul s o a 5-s chunk o
a c y signal is depic ed in Fig. 5.
3 Model adap a ion
The p oposed audio segmen a ion sys em is ained on a
da ase eco ded in a pa icula acous ic en i onmen and
may no necessa ily be able o gene alize o a da ase cap-
u ed in a di e en acous ic en i onmen . In his sec ion,
we will show ha he p oposed sys em can be made o
wo k o da a eco ded in unseen acous ic en i onmen s
as well. We will ain ou sys em on da a eco ded in a
known acous ic en i onmen and use i p edic class labels
on da a eco ded in an unseen acous ic en i onmen . Ou
p oposed solu ion consis s o wo s ages: ea u e no mal-
iza ion [45] and semi-supe ised lea ning. These will be
desc ibedinde ailin he ollowingsubsec ions.
3.1 Fea u e no maliza ion
Fea u es ex ac ed om an audio ile a e no malized by
sub ac ing he mean and di iding i by he s anda d
de ia ion be o e eeding i o he HMM. The mean and
s anda d de ia ion ec o s a e de i ed o each audio ile
sepa a ely. This is epea ed o each audio ile p esen in
he aining da a ( om known en i onmen ) as well as
he es da a ( om unknown en i onmen ). Fo a ea u e
ec o Fjn ex ac ed om j h ame o n h audio ile, he
no malized ea u e ec o is gi en by
Fs d
jn =Fjn −μn
σn
(3)
whe e μnand σna emean ec o ands anda dde ia ion
ec o , espec i ely, de i ed o n h audio ile. The di ide
ope a ion he e is elemen -wise.
3.2 Semi-supe ised lea ning
A semi-supe ised lea ning (SSL) me hod, known as sel
aining [46], is used o u he adap he HMM models o
an unseen acous ic en i onmen . In a classical SSL p ob-
lem, we ha e wo da ase s: labeled da a ( om a known
acous ic en i onmen ) and unlabeled da a ( om a new
acous ic en i onmen ). The idea behind his me hod is o
gene a e addi ional labeled aining da a using he unla-
beled da a comp ised o audio iles eco ded in he unseen
acous ic en i onmen . The ou pu labels gene a ed by he
model o he unlabeled da a a e ea ed as ue labels, and
he models a e e ained using he combina ion o o ig-
inal aining da a and his newly gene a ed labeled da a.
Figu e 6depic s his p ocess.
Al e na i ely, ins ead o using he en i e unlabeled da a,
a selec ion o only hose ames can be made o which we
a e con iden o he assigned label being ue. The likeli-
hoods ou pu ed by HMMs co esponding o h ee a ge
classes may be used o de ise a con idence c i e ion. In
Fig. 4, we ha e h ee likelihood ma ices co esponding o
each a ge class o each es ile. The maximum likeli-
hood o each column o he h ee ma ices is calcula ed.
The a io be ween he maximum and second la ges alue
oughly ep esen s how con iden we can be abou he
classi ica ion esul o a pa icula ame. We will e e i
as he con idence h eshold. Only hose ames o which
his a io exceeds a ce ain h eshold a e chosen. A con-
idence h eshold o 2 was used in his wo k. Figu e 7
shows he p ocedu e o selec ing da a based on con idence
h eshold. Da a selec ed his way can be used as addi-
ional aining da a o HMMs co esponding o he h ee
classes.
4 Acous ic ma e ial
Fo his s udy, we collec ed c y eco dings om wo
coho s o in an s in Tampe e, Finland, and in Cape Town,
Sou h A ica. The ollowing subsec ions desc ibe hese
wo da abases and e alua ion o he pe o mance o he
audio segmen a ion sys em on hem.
4.1 Da abase: Tampe e coho
In Tampe e, Finland, we cap u ed he eco dings a Ma e -
ni y Wa d Uni s and Neona al Wa d Uni o Tampe e Uni-
e si y Hospi al. The eco ding pe iod was om Ap il 13
o Augus 3, 2014. The s udy ollowed he s ipula ed e hi-
cal guidelines and was app o ed by he E hical Commi ee
o Tampe e Uni e si y Hospi al. The coho consis ed o a
he e ogeneous g oup o 57 neona es whose ch onological
ages (i.e., he ime elapsed since bi h) a eco ding we e
om0 o5daysasdepic edinTable1. The coho was no
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 6 o 14
Fig. 4 Implemen a ion o he audio segmen a ion sys em. The inpu o he sys em is a ea u e ma ix de i ed om a es audio ile. The ou pu is a
class label assigned o each ame o he c y signal
0 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5
−0.04
−0.02
0
0.02
0.04
Ampli ude
0 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5
Exp
Ins
Res
Time (s)
Class Labels
Fig. 5 Audio segmen a ion esul s o a chunk o c y signal shown in he op panel. The bo om panel depic s he ac ual and p edic ed class labels
wi h blue and ed plo s, espec i ely
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 7 o 14
Fig. 6 Semi-supe ised lea ning block diag am
s anda dized because he a ge o he p esen s udy was
o de elop a obus ool o iden i ying in an c y sounds
in he cap u ed eco dings o gene al neona e popula ion.
In o de o minimize he in luence o lea ning and ma u-
a ion on c y cha ac e is ics, he age o he in an s was he
only s anda dized a iable in he coho .
Thec ysampleswe ecap u edina a ie yo eco d-
ing condi ions. Fi s ly, he place o eco ding and he
associa ed acous ic en i onmen a ied signi ican ly. I
included he hospi al co ido , no mal pedia ic wa d,
in ensi e ca e uni (ICU), wai ing oom, and nu se’s o ice.
Wi hin each oom, eco dings we e cap u ed a di e en
places (e.g., mo he ’s bed, weighing scales, and in an ’s
bed). Secondly, he backg ound sounds p esen in he
eco ding consis ed o human oices (e.g., coughing and
speaking) and mechanical sounds (e.g., sound o unning
wa e , ai condi ioning, and diape ape being opened).
Thi dly, in an - ela ed ac o s (e.g., weigh o he in an
and p ema u i y o bi h) ha a e known o in luence he
acous ic quali ies o c y a ied . Apa om he eco d-
ing condi ions, he c y-ini ia ing igge also a ied. I
included in asi e (e.g., enipunc u e) and non-in asi e
(e.g., changing diape s and measu ing body empe a u e)
ope a ions, as well as spon aneous c ies (e.g., due o
hunge o a igue).
All Tampe e eco dings we e s o ed as 48 kHz sam-
pling a e, wo-channel audio in a 24-bi Wa e o m audio
ile (WAV) o ma . The audio eco de used was Tascam
DR-100MK II wi h RØDE M3 ca dioid mic ophone. Fo
u he compu a ion, he mean o he wo channels was
aken o yield he signal o be segmen ed. The dis ance
be ween he in an ’s mou h and he eco de was kep a
Fig. 7 Selec ion o da a based on con idence h eshold o each unlabeled audio ile. The model ou pu s p o ide he labels o semi-supe ised
lea ning
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 8 o 14
Table 1 The ch onological ages o in an subjec s in he
Tampe e coho
No. o in an s Ch onological age (day)
10
11 1
29 2
10 3
34
25
1 Missing in o.
app oxima ely 30 cm. Each eco ding was gi en a sepa-
a e numbe code. The eco dings we e manually anno-
a ed using Audaci y [47] applica ion o gene a e labels o
aining he HMM models. Figu e 8is a snapsho o he
Audaci y applica ion showing an example o a chunk o
he labeled c y eco ding.
The da abase o 57 manually anno a ed audio eco d-
ings spans a ound 115 min in du a ion. A o al o 1529
expi a o y phases we e ound wi h a mean du a ion o
0.95 s and a s anda d de ia ion o 0.65 s. Simila ly, 1005
inspi a o y phases we e ound wi h a mean du a ion o
0.17 s and a s anda d de ia ion o 0.06 s. Figu e 9( op)
illus a es he dis ibu ion o he ime du a ions o expi-
a o y and inspi a o y phases o he Tampe e coho .
No e ha inspi a o y phases we e ewe in numbe and
sho e in du a ion as compa ed o expi a o y phases.
Hence, less da a we e a ailable o aining he HMM
o inspi a o y phases as compa ed o expi a o y phases.
Mo eo e , i needs o be emphasized he e ha inspi a-
o y phases exhibi ed mo e a ia ions h oughou he da a
in compa ison o expi a o y phases. Fo example, on he
one hand, we had eco dings wi h e y sho o almos no
disce nible inspi a o y phases, and on he o he hand, we
had eco dings which ha e unusually p ominen inspi a-
o y phases as compa ed o expi a o y phases. I is also
possible o obse e bo h hese ex eme cases wi hin he
same eco ding.
4.2 Da abase: Cape Town coho
The o he coho used o his s udy is being in es i-
ga ed unde a la ge esea ch p ojec in coope a ion wi h
he Depa men o Psychia y, Uni e si y o S ellenbosch,
Cape Town. The da a we e collec ed in 2014 and consis ed
o c y eco dings o 52 in an s whose age was less han
7 weeks (mean 33.5 days, s anda d de ia ion 3.5 days).
The c y eco dings in his da abase we e also manually
anno a ed using he Audaci y applica ion. The da abase
o 52 manually anno a ed audio eco dings spans a ound
75 min in du a ion. A o al o 1307 expi a o y phases we e
ound wi h a mean du a ion o 1.1 s and a s anda d de ia-
ion o 0.76 s. Simila ly, 680 inspi a o y phases we e ound
wi h a mean du a ion o 0.25 s and a s anda d de ia ion
o 0.07 s. Figu e 9(bo om ) illus a es he dis ibu ion o
he du a ions o expi a o y and inspi a o y phases o he
Cape Town coho .
Fig. 8 Snapsho o he Audaci y applica ion showing a manually anno a ed chunk o he c y eco ding. Expi a o y and inspi a o y phases a e coded
by names exp_c y and insp_c y, espec i ely
Nai hani e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2018) 2018:1 Page 9 o 14
0 1 2 3 4
0
50
100
150
200
Du a ion (s)
No. o segmen s
00.1 0.2 0.3 0.4 0.5
0
50
100
150
200
250
Du a ion (s)
No. o segmen s
0 1 2 3 4
0
50
100
150
200
Du a ion (s)
No. o segmen s
00.1 0.2 0.3 0.4 0.5
0
50
100
150
Du a ion (s)
No. o segmen s
Fig. 9 Dis ibu ion o du a ions o expi a o y (le ) and inspi a o y ( igh ) phases o Tampe e ( op) and Cape Town (bo om) coho s
In he Cape Town coho , he loca ion and p ocedu e o
eco ding we e somewha mo e s anda dized han in he
Tampe e coho (i.e., he eco dings we e cap u ed while
conduc ing ou ine examina ions in he same nu sing
oom). The c y igge used was accina ion (i.e., in a-
si e) o measu emen o in an weigh a a weighing scale
(i.e., non-in asi e). All Cape Town eco dings we e s o ed
as 48-kHz sampling a e, wo-channel audio in a 24-bi
Wa e o m audio ile (WAV) o ma . The audio eco de
used was Zoom H4n eco de wi h buil -in condense
mic ophones. The dis ance be ween in an ’s mou h and
he eco de was app oxima ely 1.3 m o in an s being
accina ed and 70 cm o in an s being weighed. Ou da a
collec ion was conjoined wi h ano he s udy whose p o-
ocol equi ed he mic o be a bi a and hence he la ge
dis ance be ween he in an and he eco de as compa ed
o he Tampe e coho . Due o guidelines o he p ojec
conce ning p o ec ion o p i acy o he in ol ed pa ici-
pan s, we a e no able o publish he audio da a used in
his p ojec .
5 E alua ion
The segmen a ion pe o mance was e alua ed using a
i e- old c oss- alida ion amewo k. In he case o
Tampe e coho , he a ailable da ase o 57 c y eco d-
ings was di ided in o i e pa i ions: ou pa i ions o
12 eco dings each and one pa i ion o nine eco d-
ings. In a simila manne , o he Cape Town coho , he
da ase o 52 c y eco dings was di ided in o i e pa i-
ions: ou pa i ions o 10 eco dings and one pa i ion
o 12 eco dings. The di ision was done acco ding o c y
codes assigned o he eco dings which co espond o he
ch onological o de in which hey we e cap u ed. In each
old, one o he pa i ions was used as he es se and he
es o he pa i ions we e used o aining. Fi e such olds
we e pe o med wi h each old ha ing a di e en pa i ion
as he es se . The ou pu labels gene a ed by he sys em
we e compa ed agains he manually anno a ed g ound
u h.
Fo each es ile unde in es iga ion, he ou pu labels
p oduced by he model we e compa ed agains he g ound
u h (i.e., manual anno a ions) o calcula e he pe o -
mance me ics. Two me ics ha e been used in his s udy
o e alua e he pe o mance o he sys em, namely, ame-
based accu acy and ame-based Fsco e. The ame-
based accu acy is de ined as
accu acy =numbe o co ec ly labeled ames
o al numbe o ames .(4)
The ame-based Fsco eisde inedas heha monic
mean o p ecision and ecall alues. P ecision is he a io
o ue posi i e alue o he es ou come posi i es o
a pa icula class. T ue posi i e alue is he numbe o
ames co ec ly labeled by he sys em o a pa icula
class, and es ou come posi i e alue is he numbe o
ames de ec ed by he sys em belonging o ha class.
Recall is he a io o ue posi i e alues o o al posi i e
alues o any class. To al posi i e alues a e numbe o
ames in he es se belonging o ha pa icula class.
The ame-based Fsco e is hus gi en by
Fsco e =2P·R
P+R,(5)