ON THE INTERACTION BETWEEN TIME AND FREQUENCY FILTERING
OF SPEECH PARAMETERS FOR ROBUST SPEECH RECOGNITION
Dušan Macho* and Climen Nadeu**
*Dep . o Telecommunica ions, Slo ak Technical Uni e si y and Dep . o Speech Analysis and Syn hesis,
Slo ak Academy o Sciences, B a isla a, Slo akia
**Dep . o Signal Theo y and Communica ions, Uni e si a Poli ècnica de Ca alunya, Ba celona, Spain
ABSTRACT
One o he g ea oday’s challenges in speech ecogni ion is o
ensu e he obus ness o he used speech ep esen a ion. Usually,
he ecogni ion a e is s ongly educed when he speech is
co up ed, e.g. by con olu ional o addi i e noise, and he
speech ea u es a e no designed o be obus . In his pape we
s udy he e ec o addi i e noise on he loga i hmic il e -bank
ene gy ep esen a ion. We use ime and equency il e ing
echniques o emphasize he disc imina i e in o ma ion and o
educe he misma ch be ween noisy and clean speech
ep esen a ion. A 2-D spec al ep esen a ion is in oduced o see
he egions mos a ec ed by noise in he 2-D que ency-
modula ion equency domain and o help o design he
equency and ime il e shapes. Expe imen s wi h one and wo
dynamic ea u e se s show he use ulness o he combina ion o
ime and equency il e ing o bo h, whi e and low-pass noise
speech ecogni ion. A he end he powe ime and equency
il e ing echnique is p esen ed.
1. INTRODUCTION
Only a pa o he in o ma ion con ained in he speech signal is
used o speech ecogni ion. Mo eo e , a speech signal can be
dis o ed by non-speech componen s (e.g. channel o
mic ophone dis o ion, addi i e noise, e e be a ion…). I is
necessa y o ex ac phone ically impo an ea u es wi h good
disc imina i e p ope ies and obus ness when used in ad e se
en i onmen s.
Fo ecogni ion pu poses, speech is o en con e ed o a ime
sequence o log il e -bank ene gies (log FBE). In his way, he
conside ed speech uni is ep esen ed as a wo-dimensional (2-
D) ime- equency sequence. This sequence is u he p ocessed
in o de o ob ain mo e obus and disc imina i e ea u es (e.g.
ans o med o mel-ceps um, RASTA il e ed [1]…). Recen ly,
he au ho s in [2] showed ha a simple il e ing pe o med on
he equency dimension o e e y ame (F equency Fil e ing –
FF) gi es be e ecogni ion esul s o clean speech han
ceps al coe icien s. The FF can be seen as a li e ing ope a ion
pe o med in he spec al domain. The equency il e s in [2]
we e designed o equalize he a iance o ceps al coe icien s
and a simple, da abase independen , second-o de il e z-z-1
(he e deno ed as FF2) was ound as a good comp omise. In [3],
he FF ea u es appea ed mo e obus han ceps al coe icien s
when speech is dis o ed by addi i e whi e noise and a i s -
o de il e 1-z-1 (FF1) ga e good ecogni ion esul s.
The componen s o he speech ea u e ec o a y in ime,
acco ding o he changes o he speech signal, desc ibing ime
ajec o ies. The spec um o he ime ajec o y is called
modula ion spec um. The ypical speech modula ion spec um
dec eases along he modula ion equency axis [4]. Thus, he
low modula ion equencies gene ally domina e he dis ance
compu a ion in he classi ie (simila ly, as do he low que ency
componen s) bu hey do no ca y he mos disc imina i e
in o ma ion [5]. Mo eo e , when he speech is co up ed by
s a iona y con olu ional noise, he 0 h modula ion equency is
he mos a ec ed in he log FBE ep esen a ion. Thus, il e ing
on he ime dimension (Time Fil e ing – TF) can emo e
undesi able pa s o he modula ion spec um.
In [5], bo h ime and equency il e ing we e p esen ed join ly,
bu conside ing ha he e is no in e ac ion be ween hem.
Howe e , we ecen ly obse ed some ac s ha led us o
conside ha he in e ac ion exis s. Fi s ly, he no iceable be e
clean speech pe o mance o FF wi h espec o ceps um ha is
ob ained when only one s a ic ea u e se (wi hou TF) is used,
may be educed o a sligh di e ence i dynamic ea u es a e
included in he ep esen a ion. Second, FF looses i s good
pe o mance o noisy speech when he noise is colo ed.
In his wo k, we gain mo e insigh in o ha in e ac ion p oblem
by using he 2-D modula ion spec um ep esen a ion ob ained
om log FBE sequence. We obse ed, o example, ha he
mean alue o ha 2-D unc ion o noisy speech shows highe
alues a low indices han he co esponding unc ion o clean
speech. Thus, TF and FF can imp o e he ecogni ion a e by
a enua ing he mos dis o ed egions. Mo eo e , in he same
way ha i can be con enien o use sligh ly di e en ime il e s
in wo di e en equency bands [6], he use o di e en
equency il e s in di e en modula ion equency egions can
also inc ease he ecogni ion pe o mance o speech dis o ed
by addi i e noise. Fo designing he il e s, we can ake
ad an age o ha 2-D spec al ep esen a ion.
Fo he ecogni ion es s p esen ed in his pape , we used he
ollowing condi ions: single digi s om he adul po ion o he
TI da abase, decima ed om 20 kHz o 8 kHz sampling a e; no
p eemphasis; 30 ms long Hamming windowed ames wi h 10
ms shi ; 13-o de log FBE basic pa ame e iza ion scheme;
con inuous densi y HMMs wi h 8 s a es pe digi and 3 s a es o
he silence model; o noisy speech, ei he s a iona y whi e
addi i e noise o low-pass addi i e noise wi h cu -o equency
1100 Hz we e added o he clean speech o ob ain SNR equal o
20 dB and 10 dB. T aining was pe o med always wi h clean
speech and es ing wi h noisy speech.
This wo k was ca ied ou du ing he s ay o D. Macho a UPC
Ba celona and sponso ed by Spanish go e nmen and pa ially by
Slo ak Academy o Sciences.
5 h In e na ional Con e ence on Spoken
Language P ocessing (ICSLP 98)
Sydney, Aus alia
No embe 30 -- Decembe 4, 1998
ISCAA chi e
h p://www.isca-speech.o g/a chi e
2. THE 2-D MODULATION SPECTRUM
Fo be e analysis pu poses, we sp ead modula ion spec um
ep esen a ion [4] in wo dimensions, whe e he modula ion
spec um o e e y ceps al coe icien is p esen . Le log S(k,n)
be he sho - ime log FBE es ima e o he speech signal wi h k
deno ing he il e -bank ou pu and n he ame index. The 2-D
modula ion spec um (in [7], he modula ion spec og am has
been in oduced which displays he e olu ion o low modula ion
equencies in ime and equency) is hen es ima ed by
compu ing and a e aging unc ion ()
2
,
q
mC o e a speech
da abase. ()
2
,
q
mC is ob ained by in e se disc e e-Fou ie
ans o ming om he equency domain k o he que ency m
and by he Fou ie ans o ming om he ime domain n o he
modula ion equency domain
q
,
2
),(),(),(),(log 2
qq
mCmCnmcnkS nk FTIDFT
¾®¾¾¾®¾¾¾¾®¾
×
. (2)
The 2-D modula ion spec um es ima ed om clean isola ed
digi s da abase is shown on Figu e 1(a). The dec easing il in
bo h dimensions can be obse ed. Figu e 1(b) shows he 2-D
modula ion spec um o speech dis o ed by addi i e whi e
noise. The low indices in que ency and modula ion equency
seem o be he mos a ec ed by noise.
The misma ch be ween aining and es ing log FBE
ep esen a ion is he main eason o he poo ecogni ion esul s
ob ained when he speech co up ed by addi i e noise is used o
es ing. We compu ed he misma ch be ween he clean and noisy
speech ep esen a ion as
() ()
2
,,
qq
mCmC cleannoisy
-
o all m and
q
,(3)
and a e aging i o e many speake s and u e ances we
es ima ed he 2-D modula ion spec um o misma ch. Figu e 2
shows he 2-D modula ion spec a o misma ch o speech
co up ed by addi i e whi e noise (a) and addi i e low-pass
noise (b), bo h o SNR=10dB. The la ges misma ch is si ua ed
Figu e 1: 2-D modula ion spec a o (a) clean and (b) whi e
noise speech wi h SNR=10dB
Figu e 2: 2-D modula ion spec a o misma ch o (a) addi i e
whi e noise and (b) addi i e low-pass noise, bo h wi h
SNR=10dB
(a)
(b)
(a)
(b)
in low que encies and modula ion equencies wi h i s
maximum a he (0,0) poin . No e he di e ence along que ency
be ween bo h igu es (especially a low modula ion equencies),
while along modula ion equency hey a e simila . In bo h
igu es, when he modula ion equency inc eases, low and
middle que encies a e less a ec ed and can be used o
ecogni ion. Using FF and TF, we can emo e he dis o ed pa
o he 2-D modula ion spec um and e en be e disc imina i e
p ope ies o ea u es can be ob ained. I we emphasize wo
di e en egions in he modula ion equency dimension by
using wo di e en ime il e s, one o each o wo ea u e se s,
we can use a di e en equency il e o e e y egion in o de
o weigh di e en ly in he que ency dimension. In he
ollowing sec ions, he e ec o equency and ime il e ing on
he ecogni ion pe o mance is shown.
3. RECOGNITION TESTS
3.1 S a ic Fea u e Se
I no TF is used, we e e o he speech ep esen a ion as a s a ic
ea u e se . In he ollowing, only he e ec o he equency
il e ing is p esen ed. Fo his pu pose, we used 13 di e en
equency il e s o leng h 3 wi h sys em unc ion (z-1)(z+a),
whe e a changes om –0,2 o 1,0 wi h s ep 0,1 (no e, ha he
il e wi h a=0 is FF1 and a=1 is FF2). The ans o m esponse
o he il e is a que ency unc ion (a li e ). Changing he
pa ame e a o he il e s in he in e al <-0,2; 0,2>, he shape
o he li e s in low and middle que encies changes, while does
no change much in high que encies. When he pa ame e a
changes in he in e al <0,6; 1,0>, he li e shape in he high
que encies changes while in low and middle que encies i does
no . Since he i s and he las il e ed log FBE o each ame
con ain absolu e ene gy [2] hey can ca y much noise, so ha
hey we e no used in his ea u e se .
Figu e 3 shows he ecogni ion a es o all il e s. The clean
speech ecogni ion a e (Figu e 3(a)) inc eases when a inc eases
and FF2 gi es he bes esul s. Howe e , when he speech is
co up ed by addi i e whi e noise (Figu e 3(b)), il e s ha
a enua e low and middle que encies a e p e e able. This is due
o ac ha , al hough he low and middle que encies a e use ul
o clean speech ecogni ion, hey a e se e ally a ec ed by noise
[8].
3.2 One Time-Fil e ed Fea u e Se
In his case, ime il e ing is applied o he sequence o ea u es.
As ime il e s we used he wo di e en Slepian il e s ( he
same as hose in [4] wi h pa ame e s K=1, W=12, L=14,
deno ed as TF1 and K=2, W=12, L=14 deno ed as TF2) join
wi h equaliza ion 1-0,97z-1. TF1 p ese es he modula ion
equencies o speech oughly om 0 Hz o 3 Hz and TF2 om
2 Hz o 9 Hz.
The i s es we pe o med was wi hou equency il e ing.
F om he i s wo lines o Table 1 i seems ha he ea u es
om he TF1 egion yield mo e disc imina i e in o ma ion
(97,71% ecogni ion a e o clean speech) han hose om he
TF2 egion (95,13%). Howe e , when noisy speech is
ecognized, he TF2 ea u es gi e be e esul s and a e mo e
obus han he ea u es om he TF1 egion.
Technique Clean SNR=20dB SNR=10dB
Whi e noise
TF1, no FF 97,71 50,70 16,10
TF2, no FF 95,13 64,10 41,01
TF1, FF1 13/12 97,79 94,37 82,98
TF1, FF2 13/12 99,16 95,90 81,17
TF2, FF1 13/12 97,26 92,84 77,02
TF2, FF2 13/12 98,39 95,13 79,60
Low-pass noise
TF1, FF1 13/12 97,79 90,38 78,11
TF1, FF2 13/12 99,16 90,30 77,14
TF2, FF1 13/12 97,26 92,23 77,99
TF2, FF2 13/12 98,39 93,32 78,63
The si ua ion changes when FF is used in conjunc ion wi h TF.
Figu e 3 shows he beha io o bo h, TF1 and TF2 ea u e se s
in e ms o di e en FFs. Fo clean speech, he TF1 ea u es
gi e be e esul s o e e y FF han he TF2 ea u es (see Figu e
3(a)). In he noisy case, he equency il e ing pa ially educes
he high con en o noise in TF1 egion and he ecogni ion a es
e en ou pe o m he TF2 esul s (Figu e 3(b)). Mo eo e , a
di e en beha io o he ea u e se s om wo men ioned
modula ion equency egions can be obse ed. Fo he TF1
egion, he equency il e s which a enua e mo e he low and
Table 1: Recogni ion a es in % using TF1, TF2 wi hou
equency il e ing and wi h FF1 and FF2
90
91
92
93
94
95
96
97
98
99
100
-0,2 -0,1 0,0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1,0
F equency Fil e -> a
Recogni ion Ra e [%]
s a ic clean
TF1 clean
TF2 clean
Figu e 3: Recogni ion a e in e ms o he FF and TF used
in he pa ame e iza ion o (a) clean and (b) whi e noise
speech wi h SNR=10dB
(b)
40
45
50
55
60
65
70
75
80
-0,2 -0,1 0,0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1,0
F equency Fil e -> a
Recogni ion Ra e [%]
s a ic 10dB
TF1 10dB
TF2 10dB
(a)
middle que encies ( hose wi h a=<-0,2; 0,2>) gi e sligh ly
be e esul s han he o he s o noisy speech. A enua ing he
high que encies in his egion seems o imp o e he ecogni ion
oo ( il e s wi h a=<0,6; 1,0>). Fo he TF2 egion, an
inc easing endency in he ecogni ion a e can be obse ed
when he coe icien a inc eases and FF2 is he op imal il e .
We ied o include he i s and he las il e ed log FBE o he
ea u e ec o . Only including o he i s one imp o ed he
noisy speech ecogni ion. In gene al, he i s log FBE con ains
mo e speech ene gy and is no a ec ed by addi i e noise so
much as he las one, which con ains less speech ene gy. Table 1
shows he ecogni ion a es wi h he i s log FBE included in
he ea u e ec o .
In he addi i e low-pass noise case, he esul s om expe imen s
when only FF is used a e e y low (nea 17% o FF1 o FF2
and SNR=10dB). This is due o ac , ha he il e ed log FBEs
include he s ep o he ansi ion band o he noise spec um.
Since he s ep e ec is cons an in he ime, i can be almos
canceled by he ime il e . Resul s wi h di e en ime and
equency il e s a e in he Table 1.
3.3 Two Time-Fil e ed Fea u e Se s
The ecogni ion a e can be imp o ed using ea u es om bo h
ime- il e ed egions in wo di e en ea u e se s. We ha e
ound he s a ic ea u e se is a sou ce o e o s when used
oge he wi h ime- il e ed ea u es and we do no use i . In he
Table 2, he ecogni ion es s a e p esen ed o h ee
combina ions o ime and equency il e s. Fo clean speech, he
bes ecogni ion esul is ob ained when FF2 il e is used o
bo h ime- il e ed ea u e se s. When FF1 is used, he
ecogni ion o clean speech dec eases, bu inc eases o noisy
speech. F om Figu e 3(b) i can be obse ed (he e he ea u e
se s we e used sepa a ely), ha o he TF1 ea u e se he
equency il e s which a enua e low and middle que encies a e
p e e able and o TF2 ea u es, he FF2 is he bes il e . Using
his obse a ion, an addi ional imp o emen o noisy speech
ecogni ion was ob ained. A he end o Table 2 he bes
ecogni ion a es o low-pass noise speech a e men ioned.
Technique Clean SNR=20dB SNR=10dB
Whi e noise
FF1 13/12, TF1 &
FF1 13/12, TF2 98,15 96,18 86,68
FF2 13/12, TF1 &
FF2 13/12, TF2 99,48 97,06 84,59
FF1 13/12, TF1 &
FF2 13/12, TF2 99,12 96,74 88,01
Low-pass noise
FF1 13/12, TF1 &
FF2 13/12, TF2 99,12 94,16 81,41
3.4 Powe F equency and Time Fil e ing
In his echnique we assumed, ha in he log FBE ep esen a ion
o noisy speech he high-ene gy coe icien s a e less a ec ed by
noise han he coe icien s wi h low ene gy con en . Thus, we
use simple powe ope a ion on he log FBEs be o e hey en e o
he FF in o de o emphasize he high-ene gy coe icien s. In
gene al, he powe ope a ion can be exp essed as ()
g
nkS,log .
Table 3 shows he esul s om he same expe imen s as Table 2
bu using squa e-powe equency and ime il e ing ( 2
=
g
) o
wo ea u e se s. A clea imp o emen o noisy speech can be
ob ained while he ecogni ion a es o clean speech do no
dec ease.
Technique Clean SNR=20dB SNR=10dB
Whi e noise
FF1 13/12, TF1 &
FF1 13/12, TF2 98,43 97,22 91,51
FF2 13/12, TF1 &
FF2 13/12, TF2 99,28 98,03 90,95
FF1 13/12, TF1 &
FF2 13/12, TF2 99,16 97,63 92,31
Low-pass noise
FF1 13/12, TF1 &
FF2 13/12, TF2 99,16 96,62 89,66
4. CONCLUSIONS
So a , ime and equency il e ing ha e been s udied
sepa a ely. In his pape , we o e an in oduc ion o hei join
in es iga ion. We showed TF-FF ea u es a e obus agains
s a iona y, addi i e whi e and low-pass noises o isola ed digi
ecogni ion. A g ea ad an age o his echnique is ha i does
no dec ease clean speech ecogni ion esul s. In he u he
wo k, a 2-D il e can be designed, which will include di e en
FF in di e en TF egions in one ea u e se . Mo eo e , he
powe coe icien can be op imized. Also, we wan o ex end he
men ioned echniques o mo e di icul asks.
5. REFERENCES
1. He mansky, H., Mo gan, N. “RASTA P ocessing o
Speech”, IEEE T ans. on Speech and Audio P ocessing,
Vol.2, No. 4, 1-12, Oc obe 1994.
2. Nadeu, C., He nando, J., Go icho, M., “On he
Deco ela ion o Fil e -Bank Ene gies in Speech
Recogni ion”, P oc. Eu ospeech, 1381-84, 1995.
3. He nando, J., Nadeu, C. “Robus Speech Pa ame e s
Loca ed in he F equency Domain”, P oc. Eu ospeech,
417-20, 1997.
4. Nadeu, C., Paches-Leal, P., Juang, B. H. “Fil e ing he
Time Sequence o Spec al Pa ame e s o Speech
Recogni ion”, Speech Communica ion 22, 315-322, 1997.
5. Nadeu, C., Ma ino, J. B., He nando, J., Noguei as, A.
“F equency and Time Fil e ing o Fil e -Bank Ene gies o
HMM Speech Recogni ion”, P oc. ICSLP, 430-33, 1996.
6. A endaño, C., Vuu en, S., He mansky, H., “Da a Based
Fil e Design o RASTA-like Channel No maliza ion in
ASR”, P oc. ICSLP, 2087-90, 1996.
7. G eenbe g, S., Kingsbu y, B. E. D., “The Modula ion
Spec og am: in Pu sui o an In a ian Rep esen a ion o
Speech”, P oc. ICASSP, 1647-50, 1997.
8. He nando, J., Nadeu, C., “Linea P edic ion o he One-
Sided Au oco ela ion Sequence o Noisy Speech
Recogni ion”, IEEE T ans. on SAP, Vol. 5, N 1, 80-84,
1997.
Table 2: Recogni ion a es in % o wo ea u e se s
Table 3: Recogni ion a es in % o wo ea u e se s using squa e-
powe equency and ime il e ing