Llomba e al. EURASIP Jou nal on Audio, Speech, and Music
P ocessing (2021) 2021:1
h ps://doi.o g/10.1186/s13636-020-00191-3
RESEARCH Open Access
P og essi e loss unc ions o speech
enhancemen wi h deep neu al ne wo ks
Jo ge Llomba *, Dayana Ribas, An onio Miguel, Luis Vicen e, Al onso O ega and Edua do Lleida
Abs ac
The p og essi e pa adigm is a p omising s a egy o op imize ne wo k pe o mance o speech enhancemen
pu poses. Recen wo ks ha e shown di e en s a egies o imp o e he accu acy o speech enhancemen solu ions
based on his mechanism. This pape s udies he p og essi e speech enhancemen using con olu ional and esidual
neu al ne wo k a chi ec u es and explo es wo c i e ia o loss unc ion op imiza ion: weigh ed and uni o m
p og essi e. This wo k ca ies ou he e alua ion on simula ed and eal speech samples wi h e e be a ion and added
noise using REVERB and VoiceHome da ase s. Expe imen al esul s show a a ie y o achie emen s among he loss
unc ion op imiza ion c i e ia and he ne wo k a chi ec u es. Resul s show ha he p og essi e design s eng hens
he model and inc eases he obus ness o dis o ions due o e e be a ion and noise.
Keywo ds: P og essi e loss unc ion, Speech enhancemen , ResNe , CNN
1 In oduc ion
Mos deep neu al ne wo k speech enhancemen (DNN-
SE) me hods ac like a monoli hic block, whe e he noisy
signal is he inpu o he a chi ec u e and he enhanced
signal is he ou pu , while in e media e signals a e no
easily in e p e able. Howe e , SE can also be pe o med
as a g adual imp o emen p ocess, wi h a s ep-by-s ep
speech denoising. In his pa adigm, he signal is enhanced
p og essi ely a di e en sys em s ages, by inc emen ally
imp o ing he speech quali y a each s age in e ms o
noise educ ion, speech dis o ion, e c.
The inc emen al SE pa adigm has been ecen ly
app oached h ough he so-called p og essi e speech
enhancemen (PSE) [1–3]. In his mechanism, he ne -
wo k lea ning p ocess is decomposed in mul iple s ages,
such ha he a ge is p og essi ely op imized. This way,
he subp oblem sol ed a each s age can boos he sub-
sequen lea ning in he nex s ages. P e ious wo ks ol-
lowing his s a egy ha e shown imp o ed esul s o
he p og essi e a chi ec u es compa ed o usual DNN-SE
me hods.
*Co espondence: [email p o ec ed]
ViVoLab, A agón Ins i u e o Enginee ing Resea ch (I3A), Uni e si y o
Za agoza, C/ Ped o Ce buna 12, 50009 Za agoza, Spain
P e ious p og essi e p oposals ha e ocused on he
inc emen al signal- o-noise a io (SNR) econs uc ion
a di e en deg ees. In [2], a eed o wa d deep neu al
ne wo k implemen ed a eg ession scheme, whe e he ne -
wo k a ge was lea ning an ideal bina y mask esponsible
o imp o ing he SNR h ee imes in 10 dB. The same
example was used wi h di e en SNR o achie e he p o-
g essi e enhancemen . In [3], he au ho s ex ended his
wo k by es ing mo e ad anced a chi ec u es. Ini ially, a
ep oduc ion o he p ocedu e in [2]usingalongsho -
e m memo y cell (LSTM) showed a deg ada ion o he
SE pe o mance wi h he numbe o a ge laye s. Then, a
each cleaning s ep, hey used addi ional knowledge om
he p e ious s eps, inally achie ing an imp o emen in
pe o mance.
Mo e ecen ly and mo i a ed by he in e p e abili y o
he enhancemen p ocess, we ha e p esen ed a p og es-
si e a chi ec u e based on wide esidual ne wo ks [1]. Ou
main goal was o unde s and he enhancemen p ocess,
s ep by s ep, by using a isualiza ion p obe a each ne -
wo k block. Insigh s p o ided by he in e p e a ion o he
enhancemen p ocess led o he modi ica ion o he ne -
wo k a chi ec u e, which p o ided imp o ed esul s o
he SE p ocess. In he p oposed a chi ec u e, he mean
© The Au ho (s). 2021 Open Access This a icle is licensed unde a C ea i e Commons A ibu ion 4.0 In e na ional License, which
pe mi s use, sha ing, adap a ion, dis ibu ion and ep oduc ion in any medium o o ma , as long as you gi e app op ia e c edi
o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e Commons licence, and indica e i changes we e made. The
images o o he hi d pa y ma e ial in his a icle a e included in he a icle’s C ea i e Commons licence, unless indica ed
o he wise in a c edi line o he ma e ial. I ma e ial is no included in he a icle’s C ea i e Commons licence and you in ended
use is no pe mi ed by s a u o y egula ion o exceeds he pe mi ed use, you will need o ob ain pe mission di ec ly om he
copy igh holde . To iew a copy o his licence, isi h p://c ea i ecommons.o g/licenses/by/4.0/.
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 2o 16
squa e e o (MSE) o he log-spec al ampli ude (LSA)
be ween he enhanced signal and he e e ence is com-
pu ed a e e y ne wo k s age and e eshes he backp op-
aga ion g adien s. Fu he mo e, he econs uc ion e o
o each block con ibu es o he op imiza ion loss unc ion
wi h a weigh ed p og essi e mechanism.
Ou p elimina y app oach o his p oblem had he
in en ion o jus p esen ing a p og essi e app oach o
DNN speech enhancemen [1]. Now, his wo k deeply
s udies he p og essi e s a egy o DNN-SE. This pape
explo es he gene aliza ion o he aining me hod on
wo consolida ed DNN a chi ec u es used o SE asks:
a con olu ional neu al ne wo k (CNN) and a esidual
neu al ne wo k (ResNe ). This s udy analyzes wo di e -
en c i e ia o implemen he p og essi e pa adigm: he
weigh ed p og essi e (WP) c i e ion in [1]andanewly
p oposed uni o m p og essi e c i e ion (UP). The UP
c i e ion implemen s he inal op imiza ion o he loss
unc ion, conside ing ha he econs uc ion e o s om
all blocks con ibu e in he same way. Mo eo e , in his
wo k, we conside no only he de e e be a ion p oblem
bu he whole enhancemen p oblem. Also, a wide expe -
imen al se up is implemen ed, including simula ed and
eal da ase s.
Mo e ecen DNN a chi ec u es used o SE such as
gene a i e ad e sa ial ne wo ks (GAN) [4], U-Ne [5], o
esidual hou glass ecu en neu al ne wo ks (RHR-Ne )
[6] ha e demons a ed hei capabili ies and cu en ly hey
o e he bes esul s. Despi e hese a chi ec u es could
also bene i om he use o he p oposed me hod, in his
wo k, we concen a e on he pe o mance on a selec ed
se o e y well-known, simple, and es ablished a chi ec-
u es o show he bene i s in e ms o pe o mance wi h-
ou negligible inc ease in compu a ional complexi y ( e y
educed a aining ime and no compu a ional inc ease a
in e ence ime) o he p og essi e app oach dis ega ding
he speci ic me hod o ne wo k a chi ec u e.
The con ibu ions o his wo k a e:
•S udy o he PSE on wo consolida ed deep neu al
ne wo k (DNN) a chi ec u es: CNN and ResNe .
•Assessmen o wo c i e ia o p og essi e loss
unc ion op imiza ion: weigh ed and uni o m.
•Explo ing he space o inpu ea u es.
•Analysis o he p og essi e mechanism e ec on
g adien s and speech quali y measu es.
The es o he pape is o ganized as ollows. Sec ion 2
summa izes he an eceden s o his wo k. Sec ion 3goes
deepe in o he applica ion o he p og essi e pa adigm
o he loss unc ion. Sec ion 4desc ibes he expe imen al
condi ions. Sec ion 5p esen s some p elimina y esul s
on he anishing g adien p oblem, and Sec ion 6analyzes
he beha io o he CNN/ResNe a chi ec u es when hey
a e using he p og essi e pa adigm by p esen ing ob ained
esul s. Finally, Sec ion 7concludes he pape .
2 An eceden s
The a chi ec u es conside ed in his wo k a e CNN and
ResNe . In o de o adap hese a chi ec u es o he
p og essi e pa adigm, i is necessa y o add addi ional
es ic ions and modi y he loss unc ion. In he ollowing
subsec ions, we p o ide an o e iew o he a chi ec u e
design and he loss unc ion ha will be he base o his
wo k.
2.1 A chi ec u e
A chi ec u es based on CNN a e capable o exploi ing
local pa e ns in he spec um om bo h equency and
empo al domains [7,8]. The e ec o noise and e e be -
a ion appea s as a pe u ba ion o he signal spec al shape
ex ended h ough a speci ic ime- equency a ea. The
na u al s uc u e o he speech signal o he dis o ion pa -
e ns can show co ela ion in consecu i e ime- equency
bins in a con ex . CNN-based a chi ec u es e ec i ely
deal wi h his cha ac e is ic o he speech signal s uc u e,
wha makes hem app op ia e o speech enhancemen
pu poses. CNN has also appea ed combined wi h ecu -
en blocks o u he model he dynamic co ela ions
among consecu i e ames [9]. In Fig. 1,weshowa ypical
s uc u e o a CNN whe e each a chi ec u e block could
ha e di e en con igu a ions in e ms o con olu ional
laye s, ba ch no maliza ion, o non-linea i ies.
The inco po a ion o esidual connec ions b ough
a egula iza ion po en ial o he CNN app oach [10].
ResNe a chi ec u e makes use o sho cu connec ions
be ween neu al ne wo k laye s, allowing sys ems o han-
dle mo e dep h, wi h as e con e gence and a smalle
g adien anishing e ec . Since hey can manage deepe
ne wo ks, hey can be mo e exp essi e, p o ide mo e
de ailed ep esen a ions o he unde lying s uc u e o he
co up ed signal and manage longe con ex s. All o his
esul s in mo e accu a ely enhanced speech. We show his
modi ica ion in Fig. 2, whe e we desc ibe he connec ion
be ween con olu ional blocks in a esidual app oach.
In [1], we added o he ResNe an addi ional cons ain :
he a chi ec u e kep a cons an numbe o channels
along all he blocks o he DNN. The cons an num-
be o channels allowed he ou pu econs uc ion and
a isualiza ion p obe a any in e nal block. The manda-
o y p og essi e signal econs uc ion o ced an inc e-
men al p ocess o he SE ha ended o imp o e he
obus ness o he model. Besides, his a chi ec u e uses a
weigh ed composi ion o econs uc ion e o s by block
o pe o m he loss unc ion op imiza ion. This way, each
block makes pa ial econs uc ion, and he nex block
has as inpu a p e iously enhanced ep esen a ion o he
signal.
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 3 o 16
Fig. 1 Con olu ional neu al ne wo k (CNN). The con olu ional block
can ha e any dis ibu ion o con olu ional laye s and auxilia y laye s
such as he ba ch no maliza ion, o non-linea i ies
2.2 Loss unc ion
In [1], we p oposed an SE sys em based on he econ-
s uc ion o he LSA o a noisy signal: he audio sig-
nal was econs uc ed, by means o he o e lap-add
mechanism, using he enhanced loga i hmic ou pu spec-
um wi h he phase o he o iginal noisy speech. The
loss unc ion was he classical MSE be ween he LSA
o he e e ence and he LSA o he enhanced signal,
MSE yn,τ,ˆ
xn,τ=1
D
D−1
d=0yd,n,τ−ˆ
xd,n,τ2(1)
whe e Dis he signal inpu dimension, yd,n,τ,ˆ
xd,n,τa e he
equency bins o he loga i hmic spec um a he aining
example nand ame τ.yn,τis he a ge ec o o he clean
LSA e e ence, and ˆ
xn,τis he econs uc ed ec o o he
enhanced signal.
F om ou p e ious expe ience [1,11,12], ins ead o
using a ame-by- ame loss unc ion, his loss uses he
whole inpu as a sequence. Namely, he base loss unc ion
Fig. 2 Residual neu al ne wo k (ResNe ). The con olu ional block can
ha e any dis ibu ion o con olu ional laye s and auxilia y laye s such
as he ba ch no maliza ion, o non-linea i ies. The main di e ence
be ween CNN and ResNe is he esidual pa h in ResNe
is he MSE o he LSA o e all he examples and sequence
leng h o an upda e s ep,
JY,ˆ
X=1
N
N−1
n=0
1
T
T−1
τ=0
MSE yn,τ,ˆ
xn,τ(2)
whe e Yand ˆ
Xa e he LSA ep esen a ion o he aining
upda e.
Eachexampleisasequenceo all he ameso he
inpu signal, whe e Nis henumbe o examplesin he
aining p ocedu e s ep, and Tis he numbe o ames
o he example. In o de o simpli y he aining p oce-
du e, all he aining examples ha e he same numbe o
ames. The e o e, he aining keeps ixing he same seg-
men size, which is ob ained by andomly c opping he
inpu signals. This way, any example selec ed o a aining
upda e is an a bi a y segmen o he inpu example.
Finally, [1] implemen s he p og essi e pa adigm mod-
i ying he objec i e loss unc ion composing he MSE
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 4 o 16
be ween noisy inpu LSA and he enhanced LSA a di e -
en ne wo k le els o blocks. This p og essi e loss unc-
ion is a pa icula case o his pape p oposal, and i will
be s udied in de ail in he ollowing sec ion.
3 Speech enhancemen
This pape aims o s udy he unde lying po en ial o
he PSE pa adigm. P e ious wo ks ha e poin ed ou
he pe o mance imp o emen o he SE ask in p o-
g essi e a chi ec u e designs. Beyond hese esul s, his
pape b ings he hypo hesis ha he p og essi e pa adigm
ob ains be e SE pe o mance because hese mechanisms
also e esh g adien s du ing he neu al ne wo k ain-
ing. In he ollowing, we will desc ibe he PSE a chi-
ec u e p oposed in his pape , which is based on ou
p e ious wo k [1], bu addi ionally includes a se o no el-
ies/con ibu ions designed explici ly o his s udy.
3.1 A chi ec u e
This pape s udy will be based on wo DNN a chi ec u es:
p og essi e con olu ional neu al ne wo k (P-CNN) and
p og essi e esidual neu al ne wo k (P-ResNe ). Beyond
ou p e ious p oposal in [1] using he ResNe opology,
his pape includes he CNN opology wi h compa a-
i e pu poses and o ex end he s udy o gene alize he
p og essi e pa adigm o di e en a chi ec u es.
Figu e 3 ep esen s he on -end o bo h a chi ec u es.
The inpu signal, x( ), is i s windowed, and hen, we
ob ain he loga i hm o he absolu e alue o i s sho - e m
Fou ie ans o m (STFT), yielding he LSA X.Wealso
ob ain he Mel-scaled il e bank (FB), and Mel- equency
ceps al coe icien s (MFCC) wi h di e en windowing
p ocesses o p o ide addi ional in o ma ion o he ne -
wo k, XC.
Fig. 3 F on -end. P ep ocessing block a beginning o P-CNN and
P-ResNe a chi ec u es
Bo h a chi ec u es keep he same numbe o channels
along all hei con olu ional blocks. Also, hey use he
same basic con olu ional block (Fig. 4) o emain as com-
pa able as possible. This con olu ional block is composed
o wo successi e iden ical s uc u es. This s uc u e s a s
wi h ba ch no maliza ion, ollowed by a pa ame ic ec-
i ied linea uni (PReLU), and a 1D-Con olu ional laye
wi h hesamenumbe o channelsa heinpu and he
ou pu . In Fig. 4,Csis henumbe o channels.The
dimension o he ke nel (k) is 3 in all con olu ions o he
a chi ec u e. The ou pu o his s uc u e has he same
dimensions as he enhanced ou pu . Thus, we can ob ain
a pa ially enhanced signal a each block ou pu o P-CNN
and P-ResNe .
Fo his wo k, we used 1D-con olu ional laye s. Unlike
2D-con olu ional laye s ha combine empo al and e-
quency dimensions locally, 1D-con olu ional laye s pe -
o m a global combina ion o e all he equency dimen-
sions in a sho - e m empo al con ex . Recen wo ks sug-
ges ha when con olu ional a chi ec u es a e employed,
Fig. 4 Con olu ional block. Basic block o composing P-CNN and
P-ResNe a chi ec u es
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 5 o 16
he use o con olu ional laye s compu ed h ough he
single empo al dimension a e mo e app op ia e o he
speech enhancemen p ocessing [13,14].
The mul i esolu ion windowing p ocessing o he signal
con ibu es o he de e e be a ion ask, especially when
he impulse esponse is longe han he window leng h
used in he LSA analysis [15]. Xcou pu is only used as
inpu o he i s con olu ional block as shown in Figs. 1
and 2. The ollowing blocks ha e he same inpu and ou -
pu dimensions o allow he use o he p og essi e loss
unc ion. By p o iding he MFCC, he ne wo k has he
possibili y o calcula ing a e age ceps al ep esen a ions
o help wi h he channel iden i ica ion and imp o e he
de e e be a ion. The il e bank can also play a ole in he
iden i ica ion o use ul speech s uc u es in a pe cep ual
scale. As we can see in he expe imen s, hei combined
use makes a signi ican imp o emen .
3.2 Loss unc ion
In [1], we designed a neu al ne wo k o ha e he same
numbe o channels as he inpu signal a ce ain p obe
poin s. To induce he desi ed beha io , we o ced he
desi ed enhanced signal o be ob ained a hese poin s
by adding hei econs uc ion e o s o he aining loss,
which p o ided a p og essi e educ ion o he di e ence
be ween he e e ence signal and he econs uc ion a e
each block. Unlike he classical laye -wise aining, whe e
a s acking echnique is used, we ain he whole ne -
wo k agains he inal objec i e in he p oposed me hod
bu wi h he addi ional cons ain ha a ull econs uc-
ion a e each a chi ec u e block mus be ca ied ou .
Ou p e ious wo k demons a ed ha i we do no o ce
he econs uc ion a e each block, in e media e block
ou pu s a e en i ely di e en om ou objec i e and no
in e p e able. The inclusion o he econs uc ion con-
s ain h ough ou loss unc ion allows he isualiza ion
o he enhancemen p ocedu e. We can choose an in e -
media e esul o educe he e alua ion compu a ional
cos depending on he applica ion and help he aining
p ocedu e o ob ain be e esul s.
Wi h he p oposed loss unc ion, we add he ull econ-
s uc ion cons ain a e each con olu ional block min-
imizing he MSE be ween he clean e e ence Yand he
block ou pu ˆ
Xb(Fig. 5). Equa ion 3shows a gene al de -
ini ion o he p og essi e loss unc ion as a weigh ed sum
o e he econs uc ion loss o each con olu ional block
Fig. 5 PSE gene al a chi ec u e o P-CNN and P-ResNe . This igu e illus a es he applica ion o he p og essi e loss ha allows o di ec ly ep esen
he ou pu a e each block
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 6 o 16
Table 1 T aining da ase s desc ip ion
Da ase Timi Lib ispeech TedLium
Files 6299 292329 56704
Speake s 630 2484 698
Speech ype Read speech Con e ence
In e ace Close mic ophone Audi o ium mic ophone
JPY,ˆ
X=
B
b=1
Wb·JY,ˆ
Xb.(3)
Depending on he weigh s in Equa ion 3, i is possible o
de ine di e en p og essi e loss unc ion c i e ia. In [1],
we p oposed he WP loss unc ion and he e we also p o-
pose he UP c i e ion. In he nex sec ions, bo h c i e ia
a e expe imen ally e alua ed in combina ion wi h P-CNN
and P-ResNe .
•Weigh ed p og essi e (WP): The main weigh o he
loss unc ion is he inal cos , as usual in
app oxima ion asks. Then, he cos o all he
a chi ec u e blocks is uni o mly dis ibu ed and
added in a weigh ed sum,
JWP Y,ˆ
XB=JY,ˆ
XB+α1
B
B
b=1
JY,ˆ
Xb(4)
whe e B is he numbe o blocks o he a chi ec u e.
No e ha Equa ion 4is a pa icula case o he
gene al p og essi e loss unc ion in Equa ion 3,
whe e Wb=α/B o b=1, ...,B−1and
WB=1+α/B. This loss unc ion implemen s
p og essi e p ocessing along blocks, i.e., e e y
in e media e block econs uc s he enhanced signal.
This design o ces he enhancemen p ocess o be
inc emen al, om sligh ly o de ailed cleaning. In he
end, his p ocessing complemen s he adi ional
p ocess o ob ain he inal sys em ou pu , namely he
s anda d back-p opaga ion o g adien s h oughou
he ull a chi ec u e (ou pu -inpu ).
•Uni o m p og essi e (UP): This loss unc ion
p oposes a uni o m dis ibu ion o he block losses
along he a chi ec u e,
JUP Y,ˆ
XB=1
B
B
b=1
JY,ˆ
Xb,(5)
which is a special case o Equa ion 3whe e
Wb=1/B o b=1, ...,B.
Wi h his s a egy, all he ou pu s ha e he same
impac in he econs uc ion. This way, e e y block
can equally con ibu e o he inal loss, and he ull
a chi ec u e makes he same e o in he signal
econs uc ion.
4 Expe imen al se up
4.1 T aining da a
Fo DNN aining, we ha e used h ee di e en public
da ase s: Tedlium [16] om Ted alks; Lib ispeech [17],
audio-books; and Timi [18], a phone ically balanced dis-
ibu ed ead speech. These da ase s a e ully employed,
wi hou any pa i ion. See Table 1 o he cha ac e is ics
o he da ase s.
4.2 Da a augmen a ion: e e be a ed and noisy aining
da a
Da a augmen a ion using e e be a ion and addi i e noise
was pe o med a he aining se . Fo each andom ain-
ing example, he e a e h ee ans o ma ions (See Table 2
o u he de ails):
1 Impulse esponses: We simula ed andom ooms and
sou ce- ecei e dis ances desc ibed h ough he
oom impulse esponses (RIR) using he py hon
package i -gene a o 1[19]. Fo he da a
augmen a ion loop, he e a e h ee di e en kinds o
simula ed ooms: small, medium, and la ge, selec ed
wi h a p obabili y o 0.5, 0.3, and 0.2.
2 Addi i e noise: We add some noise, wi h SNR
uni o mly sampled be ween 5 and 25 dB, om he
music and noise iles in he Musan da ase [20]. No e
ha among he noise iles, he e is c owd noise, bu
he e is no any in elligible speech.
3 Time scaling: We andomly selec a scale be ween 0.8
and 1.2. The e a e signals wi h no scaling, i.e., he
o iginal speed. Some o he s a e slowed down o sped
up.
4.3 E alua ion da a
Fo e alua ion pu poses, we use wo da abases: (1)
REVERB [21] and (2) VoiceHome 0.2 [22]and 1.0
[23]. REVERB is di ided in a de elopmen se (REVERB-
De ), gene ally used o e alua ing in e media e esul s
du ing he s udy, and an e alua ion se (REVERB-E al),
o con i ming he esul s and e alua ion o he sys em.
VoiceHome e alua es he sys em in a ealis ic domes ic
en i onmen wi h noise and e e be a ion. So, wi h hese
woda abases,wecansepa a e wocondi ions:
1h ps://gi hub.com/Ma in182/ i -gene a o
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 7 o 16
Table 2 RIR and noise o aining da a augmen a ion
Room impulse esponses
Small Medium La ge
P obabili y 0.5 0.3 0.2
Size (x,y,z)[m] x∼U(1, 6),y∼U(1, 6),z∼U(2, 3.5)x∼U(6, 10),y∼U(6, 10),z∼U(3, 5)x∼U(10, 20),y∼U(10, 20),z∼U(4, 6)
RT60[s] RT60 ∼U(0.1, 0.25)
Dis ance[m] 0.5, 1.0, 1.5, 2.0, 2.5
Mic ophone ype Bidi ec ional, hype ca diodid, ca dioid, subca doid, omnidi ec ional
Noise
Music 659 iles
Noise 929 iles
SNR [dB] SNR ∼U(5, 25)
Simula ed da a Pa o he REVERB da ase co esponds
o simula ed condi ions. They a e speech samples om
he WSJCAM0 co pus [24]combinedwi h h eekindso
RIR: small, medium, and big oom (RT60 =0.25, 0.5, 0.7s).
Fo each one, he e a e wo sou ce-mic dis ances: a (2m)
and nea (0.5m). Also, a s a iona y noise was added om
he same ooms (SNR =20dB). Fo his s udy, we only
use he i s channel o he eigh a ailable. We also add
i e noises (SNR =0, 5, 10, 15, 20, and 25dB) o all signals
a he simula ed condi ion o REVERB. These a e babble
noise, ca e en i onmen noise, music, s ee en i onmen
wi h lo o a ic, and noise cap u ed inside a mo ing
am.
Real da a We used wo e alua ion se s wi h eal condi-
ions: he eal pa in REVERB and VoiceHome da ase
( 0.2 and 1.0). REVERB was eco ded in a mee ing oom
wi h RT60 =0.7sa wo dis ances: a (2.5 m)and nea
(1m), omMC-WSJ-AV[25]. VoiceHome co esponds o
a ealis ic domes ic en i onmen wi h e e yday noises like
a acuum cleane , dish-washing, o sound o TV shows.
4.4 Speech quali y measu es
To measu e he le el o denoising and de e e be a ion
achie ed by he PSE me hod, we es ima e he segmen-
al SNR [26] and he speech- o- e e be a ion modula-
ion ene gy a io (SRMR) [27,28]. In hese me ics, he
highe he alues, he be e speech quali y. Howe e ,
i is well-known ha he SE p ocessing migh gene a e
dis o ion on he ou pu speech. The e o e, o he sim-
ula ed da ase , we also measu e he dis o ion be ween
he clean e e ence and enhanced speech using he log-
likelihood a io (LLR) [29]. In his case, lowe alues mean
less dis o ion, so he be e quali y o he speech. The
combina ion o bo h speech quali y iewpoin s, i.e., he
ade-o be ween noise/ e e be a ion educ ion and dis-
o ion, p o ides a gene al assessmen o he SE me hod
pe o mance. This way, he bes enhancemen sys em is
he one which imp o es SNR o SRMR, bu e ains he
dis o ion, in his case, measu ed wi h LLR, as low as pos-
sible. Addi ionally we use he well-known PESQ measu e
[30] o simula ed da a. PESQ measu e is in ange 0-5
whe e he highe he be e pe o mance.
4.5 Neu al ne wo k con igu a ion
The inpu p o ided o he CNN, ResNe , P-CNN, and
P-ResNe a chi ec u es consis s o he loga i hm o he
magni ude o he 512-STFT o he co up ed signal, sam-
pled a 16 kHz. The STFT is compu ed e e y 10 ms o a
25 ms sliding Hamming window. We also conca ena e he
Mel-Scaled Fil e -bank and he MFCC as auxilia y inpu s,
wi h il e bank sizes 32, 50, and 100, e e y 10 ms. MFCC
a e compu ed using he disc e e cosine ans o m (DCT)
wi hou unca ion. Howe e , each equency esolu ion
has a di e en sliding Hamming window o 25 ms, 50 ms,
and 75 ms espec i ely. These auxilia y ea u es p o ide
di e en equency and empo al esolu ions, which can
bene i he speech enhancemen p ocess [15]. Taking in o
accoun ha he LSA dimension is 512, he o e all inpu
size is 876.
Fo all he expe imen s, we use adap i e momen es i-
ma o (Adam) as he upda e unc ion. Each laye has 512
neu ons o ollow he philosophy o main aining unal-
e ed he numbe o channels along he a chi ec u e. The
aining consis s o 900 epochs. Fo each epoch, 10,000
inpu iles a e andomly selec ed om he aining se . As
long as he e a e unused aining examples, no ile can
be selec ed mo e han once. Ba ch no maliza ion mo ing
pa ame e s a e blocked a e epoch 700. Fo he JWP loss
unc ion, we use α=0.1 as in [1], which p o ided he bes
SRMR alue on REVERB-De .
5 P elimina y g adien s udy
This sec ion p esen s a p elimina y s udy o he beha -
io o he g adien o explo e how he injec ion o new
esh g adien s a di e en a chi ec u e le els imp o e
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 8 o 16
he aining p ocedu e. When g adien s back-p opaga e
h ough a la ge numbe o laye s, hey end o lose ene gy.
Thus, hei abili y o mo e weigh s o he laye s nea o
he inpu is educed. The p oposed PSE me hod eeds
a esh and s onge g adien a e each block o mo e
he weigh s o each laye . In o de o check his, we
design an expe imen o obse e he ene gy o he g a-
dien s ha modi y he weigh s o he i s con olu ional
block du ing he 100 i s op imiza ion upda es. This
p ocedu e is epea ed 100 imes wi h di e en weigh
ini ializa ions o obse e he a iance among di e en
s a s and he a ia ion o his g adien ene gy du ing
op imiza ion.
Figu e 6p esen s he esul s ob ained o P-CNN and P-
ResNe a chi ec u es, o non-p og essi e baselines, and
o each p oposed p og essi e c i e ia. The e is a no ice-
able di e ence in he beha io o he wo s uc u es. In P-
CNN, he e is a signi ican di e ence among he g adien
ene gy o each compa ed sys em. The lowe ene gy co e-
sponds o he baseline a chi ec u e, he one wi hou any
p og essi e assump ion. On he o he hand, he p og es-
si e mechanisms show a signi ican li ing o he g adien
ene gy. These boos ed g adien s ha e mo e s eng h o
mo e he weigh s allowing a be e lea ning a inne laye s
o he whole a chi ec u e.
In con as , in P-ResNe , he e is no ele an di e ence
be ween he g adien ene gy o he p og essi e echniques
and ha o he no p og essi e baseline a he i s con o-
lu ional block. Conside ha P-ResNe is an a chi ec u e
designed o deal wi h he anishing p oblem, and hanks
o esidual connec ions, he g adien s ha e a sho cu o
p opaga e up o he i s laye s wi hou anishing. In his
case, injec ing new g adien s does no push much mo e
he p e ious g adien s. Howe e , he new g adien s a e
mo e accu a e because hey di ec ly come om he a ge
e alua ion a he ou pu o each a chi ec u e block.
6 Resul s and discussion
6.1 Analysis o al e na i es o he DNN inpu
In his sec ion, we p esen a s udy o asses ha he com-
bined use o complemen a y inpu s o he co up ed LSA
may imp o e he pe o mance o he sys em. We use mul-
i esolu ion in he MFCC and FB inpu s as desc ibed
in Sec ion 4.5, bu we pe o m an abla ion s udy abou
he use o each ea u e ype. Fo his s udy, we ocus on
he de e e be a ion pe o mance o he P-ResNe wi h
Fig. 6 Mean and s anda d de ia ion (shaded a ea) o he log-ene gy o he g adien s o 100 andom ne wo k ini ializa ion du ing he 100 i s
mini-ba ches in aining. The log-ene gy is measu ed on he nea es block o he inpu
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 9 o 16
Table 3 E alua ion o he use o complemen a y in o ma ion a he inpu o he P-ResNe wi h WP a chi ec u e o e he REVERB-De
da ase in e ms o e e be a ion measu ed wi h SRMR
Complemen a y in o. Real condi ion Simula ed condi ion
Wi hou 6.55 7.99
FB 7.25 8.31
MFCC 7.07 8.44
FB + MFCC 7.14 8.41
Bold ex ema ks on he bes esul pe condi ion and i alic ex he second bes
WP o e he REVERB-De da ase in eal and simula ed
condi ions.
Table 3shows ha he bes esul s in simula ed condi-
ions a e a ained using only MFCC, bu o eal condi-
ions hey a e ob ained wi h FB ea u es. On a e age, he
combined use o bo h ea u es, FB and MFCC, p o ides
he bes pe o mance, especially compa ed o he use o
LSA wi hou any auxilia y inpu s.
6.2 A chi ec u e dep h analysis
SE p og essi e me hods use a sequence o s eps o pe -
o m he enhancemen . We ha e o de e mine he numbe
o s eps o he numbe o blocks ha composes he a chi-
ec u e. Table 4shows he a chi ec u e dep h s udy in
e ms o SRMR o e he REVERB-De da ase . This s udy
shows he esul s o simula ed and eal condi ions and
he a e age o bo h.
Resul indica e ha he con igu a ion wi h 16 blocks
achie es he bes pe o mance o all he e alua ed con-
di ions. No e how p og essi e sys ems can achie e high
SRMR, bo h o simula ed and eal condi ions. This con-
sis ency among di e en condi ions demons a es how
he p og essi e s a egy can p o ide a be e gene aliza-
ion o he DNN aining.
Fo CNN opology, he e e ence sys em in eal condi-
ions quickly deg ades he pe o mance wi h he dep h o
he a chi ec u e. Besides, esul s o P-CNN wi h UP a e
be e han he CNN e e ence sys em, i.e., P-CNN wi h
UP does no deg ade as as as CNN e e ence sys em as
dep h inc eases.
Fo ResNe opology, he a ailabili y o esidual con-
nec ions wo ks well wi h a high numbe o blocks.
Fo ins ance, he esul s o he ResNe e e ence sys-
em achie e he bes pe o mance on simula ed condi-
ions wi h he deepe a chi ec u e (32 blocks). Howe e ,
no e ha in eal condi ions, he ResNe e e ence sys-
em achie es he bes esul wi h 8 blocks e sus he 32
blocks o simula ed condi ions. Ne e heless, P-ResNe
Table 4 Speech quali y in e ms o SRMR o simula ed and eal e e be a ed speech samples h ough a chi ec u e dep h o
REVERB-De da ase . The las ows ep esen s he mean and s anda d de ia ion along he expe imen s p esen ed o each column
Re e ence sys ems P og essi e sys ems
Condi ion Blocks dep h CNN ResNe P-CNN wi h WP P-CNN wi h UP P-ResNe wi h WP P-ResNe wi h UP
Simula ed
8 7.33 8.23 6.49 7.53 8.31 7.91
16 7.60 8.27 8.96 7.70 8.41 8.05
24 8.87 8.14 6.18 8.09 8.03 8.02
32 7.01 8.56 7.65 7.41 7.98 7.78
Real
8 6.05 6.82 4.90 6.32 7.06 6.91
16 5.98 5.81 3.74 7.26 7.14 6.85
24 4.76 5.77 2.07 6.90 6.53 6.91
32 3.35 6.33 2.33 6.34 5.97 6.62
AVG5±STD
86.69±0.64 7.52±0.70 5.69±0.79 6.92±0.60 7.68±0.62 7.41±0.50
16 6.79±0.81 7.04±1.23 6.35±2.61 7.48±0.22 7.77±0.63 7.45±0.60
24 6.81±2.05 6.97±1.16 4.12±2.05 7.49±0.59 7.28±0.75 7.46±0.55
32 5.18±1.83 7.44±1.11 4.99±2.66 6.87±0.53 6.97±1.00 7.20±0.58
Bold alues show he bes esul o each condi ion
Llomba e al. EURASIP Jou nal on Audio, Speech, and Music P ocessing (2021) 2021:1 Page 16o 16
21. K. Kinoshi a, M. Delc oix, T. Yoshioka, T. Naka ani, E. Habe s, R.
Haeb-Umbach, V. Leu nan , A. Seh , W. Kelle mann, R. Maas, S. Ganno , B.
Raj, in P oceedings o he IEEE Wo kshop on Applica ions o Signal P ocessing
o Audio and Acous ics (WASPAA-13). The REVERB challenge: a common
e alua ion amewo k o de e e be a ion and ecogni ion o e e be an
speech (IEEE, New Pal z, 2013), pp. 1–4
22. N. Be in, E. Cambe lein, E. Vincen , R. Leba benchon, S. Peillon, É.
Lamandé, S. Si asanka an, F. Bimbo , I. Illina, A. Tom, e al,inIn e speech
2016. A F ench co pus o dis an -mic ophone speech p ocessing in eal
homes, (San F ancisco, 2016), pp. 2781–2785
23. N. Be in, E. Cambe lein, R. Leba benchon, E. Vincen , S. Si asanka an, I.
Illina, F. Bimbo . VoiceHome-2, an ex ended co pus o mul ichannel
speech p ocessing in eal homes, ol. 106, (2019), pp. 68–78
24. T. Robinson, J. F ansen, D. Pye, J. Foo e, S. Renals, in IEEE In e na ional
Con e ence on Acous ic, Speech and Signal P ocessing (ICASSP).WSJCAM0:a
B i ish English speech co pus o la ge ocabula y con inuous speech
ecogni ion (IEEE, De oi , 1995), pp. 81–84
25. M. Lincoln, I. McCowan, J. Vepa, H. K. Magan i, in P oceedings o he 2005
IEEE Wo kshop on Au oma ic Speech Recogni ion and Unde s anding
(ASRU-05). The mul i-channel Wall S ee Jou nal audio isual co pus
(MC-WSJ-AV): speci ica ion and ini ial expe imen s (IEEE, Philadelphia,
2005), pp. 357–362
26. C. Kim, R. M. S e n, in Nin h Annual Con e ence o he In e na ional Speech
Communica ion Associa ion (In e speech 2008). Robus signal- o-noise a io
es ima ion based on wa e o m ampli ude dis ibu ion analysis, (B isbane,
2008)
27. T. H. Falk, C. Zheng, W. Y. Chan, A non-in usi e quali y and in elligibili y
measu e o e e be an and de e e be a ed speech. IEEE T ans. Audio
Speech Lang. P ocess. 18(7), 1766–1774 (2010)
28. J. F. San os, M. Senoussaoui, T. H. Falk, in P oc. In . Wo kshop Acous . Signal
Enhancemen (IWAENC 2014). An upda ed objec i e in elligibili y
es ima ion me ic o no mal hea ing lis ene s unde noise and
e e be a ion, (An ibes - Jaun les Pins, 2014), pp. 55–59
29. P. C. Loizou, Speech quali y asssessmen . in: mul imedia analysis, p ocessing
and communica ions. (Sp inge , Be lin, 2011), pp. 623–654
30. A. W. Rix, J. G. Bee ends, M. P. Hollie , A. P. Heks a, in 2001 IEEE
In e na ional Con e ence on Acous ics, Speech, and Signal P ocessing.
P oceedings (Ca . No. 01CH37221), ol. 2. Pe cep ual e alua ion o speech
quali y (PESQ)-a new me hod o speech quali y assessmen o elephone
ne wo ks and codecs (IEEE, Sal Lake Ci y, 2001), pp. 749–752
31. L. D ude, J. Heymann, C. Boeddeke , R. Haeb-Umbach, NARA-WPE: a
Py hon package o weigh ed p edic ion e o de e e be a ion in Numpy and
Tenso low o online and o line p ocessing, (S u ga , 2018), pp. 1–5
32. T. Naka ani, T. Yoshioka, K. Kinoshi a, M. Miyoshi, B.-H. Juang. Speech
de e e be a ion based on a iance-no malized delayed linea p edic ion,
ol. 18, (2010), pp. 1717–1731
Publishe ’s No e
Sp inge Na u e emains neu al wi h ega d o ju isdic ional claims in
published maps and ins i u ional a ilia ions.