scieee Science in your language
[en] (orig)

On the interaction between time and frequency filtering of speech parameters

Abstract

One of the great today's challenges in speech recognition is to ensure the robustness of the used speech representation. Usually, the recognition rate is strongly reduced when the speech is corrupted, e.g. by convolutional or additive noise, and the speech features are not designed to be robust. In this paper we study the effect of additive noise on the logarithmic filter-bank energy representation. We use time and frequency filtering techniques to emphasize the discriminative information and to reduce the mismatch between noisy and clean speech representation. A 2-D spectral representation is introduced to see the regions most affected by noise in the 2-D quefrency-modulation frequency domain and to help to design the frequency and time filter shapes. Experiments with one and two dynamic feature sets show the usefulness of the combination of time and frequency filtering for both, white and low-pass noise speech recognition. At the end the power time and frequency filtering technique is presented.

Read accessible full text

On the interaction between time and frequency filtering of speech parameters

Author: Macho Ciena, Dusan,Nadeu Camprubí, Climent
Publisher: Robert H. Mannel and Jordi Robert-Ribes
Year: 1998
Source: https://upcommons.upc.edu/bitstream/2117/103582/1/i98_1137.pdf
ON THE INTERACTION BETWEEN TIME AND FREQUENCY FILTERING
OF SPEECH PARAMETERS FOR ROBUST SPEECH RECOGNITION
Dušan Macho* and Climen Nadeu**
*Dep . o Telecommunica ions, Slo ak Technical Uni e si y and Dep . o Speech Analysis and Syn hesis,
Slo ak Academy o Sciences, B a isla a, Slo akia
**Dep . o Signal Theo y and Communica ions, Uni e si a Poli ècnica de Ca alunya, Ba celona, Spain
ABSTRACT
One o he g ea oday’s challenges in speech ecogni ion is o
ensu e he obus ness o he used speech ep esen a ion. Usually,
he ecogni ion a e is s ongly educed when he speech is
co up ed, e.g. by con olu ional o addi i e noise, and he
speech ea u es a e no designed o be obus . In his pape we
s udy he e ec o addi i e noise on he loga i hmic il e -bank
ene gy ep esen a ion. We use ime and equency il e ing
echniques o emphasize he disc imina i e in o ma ion and o
educe he misma ch be ween noisy and clean speech
ep esen a ion. A 2-D spec al ep esen a ion is in oduced o see
he egions mos a ec ed by noise in he 2-D que ency-
modula ion equency domain and o help o design he
equency and ime il e shapes. Expe imen s wi h one and wo
dynamic ea u e se s show he use ulness o he combina ion o
ime and equency il e ing o bo h, whi e and low-pass noise
speech ecogni ion. A he end he powe ime and equency
il e ing echnique is p esen ed.
1. INTRODUCTION
Only a pa o he in o ma ion con ained in he speech signal is
used o speech ecogni ion. Mo eo e , a speech signal can be
dis o ed by non-speech componen s (e.g. channel o
mic ophone dis o ion, addi i e noise, e e be a ion…). I is
necessa y o ex ac phone ically impo an ea u es wi h good
disc imina i e p ope ies and obus ness when used in ad e se
en i onmen s.
Fo ecogni ion pu poses, speech is o en con e ed o a ime
sequence o log il e -bank ene gies (log FBE). In his way, he
conside ed speech uni is ep esen ed as a wo-dimensional (2-
D) ime- equency sequence. This sequence is u he p ocessed
in o de o ob ain mo e obus and disc imina i e ea u es (e.g.
ans o med o mel-ceps um, RASTA il e ed [1]…). Recen ly,
he au ho s in [2] showed ha a simple il e ing pe o med on
he equency dimension o e e y ame (F equency Fil e ing –
FF) gi es be e ecogni ion esul s o clean speech han
ceps al coe icien s. The FF can be seen as a li e ing ope a ion
pe o med in he spec al domain. The equency il e s in [2]
we e designed o equalize he a iance o ceps al coe icien s
and a simple, da abase independen , second-o de il e z-z-1
(he e deno ed as FF2) was ound as a good comp omise. In [3],
he FF ea u es appea ed mo e obus han ceps al coe icien s
when speech is dis o ed by addi i e whi e noise and a i s -
o de il e 1-z-1 (FF1) ga e good ecogni ion esul s.
The componen s o he speech ea u e ec o a y in ime,
acco ding o he changes o he speech signal, desc ibing ime
ajec o ies. The spec um o he ime ajec o y is called
modula ion spec um. The ypical speech modula ion spec um
dec eases along he modula ion equency axis [4]. Thus, he
low modula ion equencies gene ally domina e he dis ance
compu a ion in he classi ie (simila ly, as do he low que ency
componen s) bu hey do no ca y he mos disc imina i e
in o ma ion [5]. Mo eo e , when he speech is co up ed by
s a iona y con olu ional noise, he 0 h modula ion equency is
he mos a ec ed in he log FBE ep esen a ion. Thus, il e ing
on he ime dimension (Time Fil e ing – TF) can emo e
undesi able pa s o he modula ion spec um.
In [5], bo h ime and equency il e ing we e p esen ed join ly,
bu conside ing ha he e is no in e ac ion be ween hem.
Howe e , we ecen ly obse ed some ac s ha led us o
conside ha he in e ac ion exis s. Fi s ly, he no iceable be e
clean speech pe o mance o FF wi h espec o ceps um ha is
ob ained when only one s a ic ea u e se (wi hou TF) is used,
may be educed o a sligh di e ence i dynamic ea u es a e
included in he ep esen a ion. Second, FF looses i s good
pe o mance o noisy speech when he noise is colo ed.
In his wo k, we gain mo e insigh in o ha in e ac ion p oblem
by using he 2-D modula ion spec um ep esen a ion ob ained
om log FBE sequence. We obse ed, o example, ha he
mean alue o ha 2-D unc ion o noisy speech shows highe
alues a low indices han he co esponding unc ion o clean
speech. Thus, TF and FF can imp o e he ecogni ion a e by
a enua ing he mos dis o ed egions. Mo eo e , in he same
way ha i can be con enien o use sligh ly di e en ime il e s
in wo di e en equency bands [6], he use o di e en
equency il e s in di e en modula ion equency egions can
also inc ease he ecogni ion pe o mance o speech dis o ed
by addi i e noise. Fo designing he il e s, we can ake
ad an age o ha 2-D spec al ep esen a ion.
Fo he ecogni ion es s p esen ed in his pape , we used he
ollowing condi ions: single digi s om he adul po ion o he
TI da abase, decima ed om 20 kHz o 8 kHz sampling a e; no
p eemphasis; 30 ms long Hamming windowed ames wi h 10
ms shi ; 13-o de log FBE basic pa ame e iza ion scheme;
con inuous densi y HMMs wi h 8 s a es pe digi and 3 s a es o
he silence model; o noisy speech, ei he s a iona y whi e
addi i e noise o low-pass addi i e noise wi h cu -o equency
1100 Hz we e added o he clean speech o ob ain SNR equal o
20 dB and 10 dB. T aining was pe o med always wi h clean
speech and es ing wi h noisy speech.
This wo k was ca ied ou du ing he s ay o D. Macho a UPC
Ba celona and sponso ed by Spanish go e nmen and pa ially by
Slo ak Academy o Sciences.
5 h In e na ional Con e ence on Spoken
Language P ocessing (ICSLP 98)
Sydney, Aus alia
No embe 30 -- Decembe 4, 1998
ISCAA chi e
h p://www.isca-speech.o g/a chi e
2. THE 2-D MODULATION SPECTRUM
Fo be e analysis pu poses, we sp ead modula ion spec um
ep esen a ion [4] in wo dimensions, whe e he modula ion
spec um o e e y ceps al coe icien is p esen . Le log S(k,n)
be he sho - ime log FBE es ima e o he speech signal wi h k
deno ing he il e -bank ou pu and n he ame index. The 2-D
modula ion spec um (in [7], he modula ion spec og am has
been in oduced which displays he e olu ion o low modula ion
equencies in ime and equency) is hen es ima ed by
compu ing and a e aging unc ion ()
2
,
q
mC o e a speech
da abase. ()
2
,
q
mC is ob ained by in e se disc e e-Fou ie
ans o ming om he equency domain k o he que ency m
and by he Fou ie ans o ming om he ime domain n o he
modula ion equency domain
q
,
2
),(),(),(),(log 2
qq
mCmCnmcnkS nk FTIDFT
¾®¾¾¾®¾¾¾¾®¾
×
. (2)
The 2-D modula ion spec um es ima ed om clean isola ed
digi s da abase is shown on Figu e 1(a). The dec easing il in
bo h dimensions can be obse ed. Figu e 1(b) shows he 2-D
modula ion spec um o speech dis o ed by addi i e whi e
noise. The low indices in que ency and modula ion equency
seem o be he mos a ec ed by noise.
The misma ch be ween aining and es ing log FBE
ep esen a ion is he main eason o he poo ecogni ion esul s
ob ained when he speech co up ed by addi i e noise is used o
es ing. We compu ed he misma ch be ween he clean and noisy
speech ep esen a ion as
() ()
2
,,
qq
mCmC cleannoisy
-
o all m and
q
,(3)
and a e aging i o e many speake s and u e ances we
es ima ed he 2-D modula ion spec um o misma ch. Figu e 2
shows he 2-D modula ion spec a o misma ch o speech
co up ed by addi i e whi e noise (a) and addi i e low-pass
noise (b), bo h o SNR=10dB. The la ges misma ch is si ua ed
Figu e 1: 2-D modula ion spec a o (a) clean and (b) whi e
noise speech wi h SNR=10dB
Figu e 2: 2-D modula ion spec a o misma ch o (a) addi i e
whi e noise and (b) addi i e low-pass noise, bo h wi h
SNR=10dB
(a)
(b)
(a)
(b)
in low que encies and modula ion equencies wi h i s
maximum a he (0,0) poin . No e he di e ence along que ency
be ween bo h igu es (especially a low modula ion equencies),
while along modula ion equency hey a e simila . In bo h
igu es, when he modula ion equency inc eases, low and
middle que encies a e less a ec ed and can be used o
ecogni ion. Using FF and TF, we can emo e he dis o ed pa
o he 2-D modula ion spec um and e en be e disc imina i e
p ope ies o ea u es can be ob ained. I we emphasize wo
di e en egions in he modula ion equency dimension by
using wo di e en ime il e s, one o each o wo ea u e se s,
we can use a di e en equency il e o e e y egion in o de
o weigh di e en ly in he que ency dimension. In he
ollowing sec ions, he e ec o equency and ime il e ing on
he ecogni ion pe o mance is shown.
3. RECOGNITION TESTS
3.1 S a ic Fea u e Se
I no TF is used, we e e o he speech ep esen a ion as a s a ic
ea u e se . In he ollowing, only he e ec o he equency
il e ing is p esen ed. Fo his pu pose, we used 13 di e en
equency il e s o leng h 3 wi h sys em unc ion (z-1)(z+a),
whe e a changes om –0,2 o 1,0 wi h s ep 0,1 (no e, ha he
il e wi h a=0 is FF1 and a=1 is FF2). The ans o m esponse
o he il e is a que ency unc ion (a li e ). Changing he
pa ame e a o he il e s in he in e al <-0,2; 0,2>, he shape
o he li e s in low and middle que encies changes, while does
no change much in high que encies. When he pa ame e a
changes in he in e al <0,6; 1,0>, he li e shape in he high
que encies changes while in low and middle que encies i does
no . Since he i s and he las il e ed log FBE o each ame
con ain absolu e ene gy [2] hey can ca y much noise, so ha
hey we e no used in his ea u e se .
Figu e 3 shows he ecogni ion a es o all il e s. The clean
speech ecogni ion a e (Figu e 3(a)) inc eases when a inc eases
and FF2 gi es he bes esul s. Howe e , when he speech is
co up ed by addi i e whi e noise (Figu e 3(b)), il e s ha
a enua e low and middle que encies a e p e e able. This is due
o ac ha , al hough he low and middle que encies a e use ul
o clean speech ecogni ion, hey a e se e ally a ec ed by noise
[8].
3.2 One Time-Fil e ed Fea u e Se
In his case, ime il e ing is applied o he sequence o ea u es.
As ime il e s we used he wo di e en Slepian il e s ( he
same as hose in [4] wi h pa ame e s K=1, W=12, L=14,
deno ed as TF1 and K=2, W=12, L=14 deno ed as TF2) join
wi h equaliza ion 1-0,97z-1. TF1 p ese es he modula ion
equencies o speech oughly om 0 Hz o 3 Hz and TF2 om
2 Hz o 9 Hz.
The i s es we pe o med was wi hou equency il e ing.
F om he i s wo lines o Table 1 i seems ha he ea u es
om he TF1 egion yield mo e disc imina i e in o ma ion
(97,71% ecogni ion a e o clean speech) han hose om he
TF2 egion (95,13%). Howe e , when noisy speech is
ecognized, he TF2 ea u es gi e be e esul s and a e mo e
obus han he ea u es om he TF1 egion.
Technique Clean SNR=20dB SNR=10dB
Whi e noise
TF1, no FF 97,71 50,70 16,10
TF2, no FF 95,13 64,10 41,01
TF1, FF1 13/12 97,79 94,37 82,98
TF1, FF2 13/12 99,16 95,90 81,17
TF2, FF1 13/12 97,26 92,84 77,02
TF2, FF2 13/12 98,39 95,13 79,60
Low-pass noise
TF1, FF1 13/12 97,79 90,38 78,11
TF1, FF2 13/12 99,16 90,30 77,14
TF2, FF1 13/12 97,26 92,23 77,99
TF2, FF2 13/12 98,39 93,32 78,63
The si ua ion changes when FF is used in conjunc ion wi h TF.
Figu e 3 shows he beha io o bo h, TF1 and TF2 ea u e se s
in e ms o di e en FFs. Fo clean speech, he TF1 ea u es
gi e be e esul s o e e y FF han he TF2 ea u es (see Figu e
3(a)). In he noisy case, he equency il e ing pa ially educes
he high con en o noise in TF1 egion and he ecogni ion a es
e en ou pe o m he TF2 esul s (Figu e 3(b)). Mo eo e , a
di e en beha io o he ea u e se s om wo men ioned
modula ion equency egions can be obse ed. Fo he TF1
egion, he equency il e s which a enua e mo e he low and
Table 1: Recogni ion a es in % using TF1, TF2 wi hou
equency il e ing and wi h FF1 and FF2
90
91
92
93
94
95
96
97
98
99
100
-0,2 -0,1 0,0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1,0
F equency Fil e -> a
Recogni ion Ra e [%]
s a ic clean
TF1 clean
TF2 clean
Figu e 3: Recogni ion a e in e ms o he FF and TF used
in he pa ame e iza ion o (a) clean and (b) whi e noise
speech wi h SNR=10dB
(b)
40
45
50
55
60
65
70
75
80
-0,2 -0,1 0,0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1,0
F equency Fil e -> a
Recogni ion Ra e [%]
s a ic 10dB
TF1 10dB
TF2 10dB
(a)
middle que encies ( hose wi h a=<-0,2; 0,2>) gi e sligh ly
be e esul s han he o he s o noisy speech. A enua ing he
high que encies in his egion seems o imp o e he ecogni ion
oo ( il e s wi h a=<0,6; 1,0>). Fo he TF2 egion, an
inc easing endency in he ecogni ion a e can be obse ed
when he coe icien a inc eases and FF2 is he op imal il e .
We ied o include he i s and he las il e ed log FBE o he
ea u e ec o . Only including o he i s one imp o ed he
noisy speech ecogni ion. In gene al, he i s log FBE con ains
mo e speech ene gy and is no a ec ed by addi i e noise so
much as he las one, which con ains less speech ene gy. Table 1
shows he ecogni ion a es wi h he i s log FBE included in
he ea u e ec o .
In he addi i e low-pass noise case, he esul s om expe imen s
when only FF is used a e e y low (nea 17% o FF1 o FF2
and SNR=10dB). This is due o ac , ha he il e ed log FBEs
include he s ep o he ansi ion band o he noise spec um.
Since he s ep e ec is cons an in he ime, i can be almos
canceled by he ime il e . Resul s wi h di e en ime and
equency il e s a e in he Table 1.
3.3 Two Time-Fil e ed Fea u e Se s
The ecogni ion a e can be imp o ed using ea u es om bo h
ime- il e ed egions in wo di e en ea u e se s. We ha e
ound he s a ic ea u e se is a sou ce o e o s when used
oge he wi h ime- il e ed ea u es and we do no use i . In he
Table 2, he ecogni ion es s a e p esen ed o h ee
combina ions o ime and equency il e s. Fo clean speech, he
bes ecogni ion esul is ob ained when FF2 il e is used o
bo h ime- il e ed ea u e se s. When FF1 is used, he
ecogni ion o clean speech dec eases, bu inc eases o noisy
speech. F om Figu e 3(b) i can be obse ed (he e he ea u e
se s we e used sepa a ely), ha o he TF1 ea u e se he
equency il e s which a enua e low and middle que encies a e
p e e able and o TF2 ea u es, he FF2 is he bes il e . Using
his obse a ion, an addi ional imp o emen o noisy speech
ecogni ion was ob ained. A he end o Table 2 he bes
ecogni ion a es o low-pass noise speech a e men ioned.
Technique Clean SNR=20dB SNR=10dB
Whi e noise
FF1 13/12, TF1 &
FF1 13/12, TF2 98,15 96,18 86,68
FF2 13/12, TF1 &
FF2 13/12, TF2 99,48 97,06 84,59
FF1 13/12, TF1 &
FF2 13/12, TF2 99,12 96,74 88,01
Low-pass noise
FF1 13/12, TF1 &
FF2 13/12, TF2 99,12 94,16 81,41
3.4 Powe F equency and Time Fil e ing
In his echnique we assumed, ha in he log FBE ep esen a ion
o noisy speech he high-ene gy coe icien s a e less a ec ed by
noise han he coe icien s wi h low ene gy con en . Thus, we
use simple powe ope a ion on he log FBEs be o e hey en e o
he FF in o de o emphasize he high-ene gy coe icien s. In
gene al, he powe ope a ion can be exp essed as ()
g
nkS,log .
Table 3 shows he esul s om he same expe imen s as Table 2
bu using squa e-powe equency and ime il e ing ( 2
=
g
) o
wo ea u e se s. A clea imp o emen o noisy speech can be
ob ained while he ecogni ion a es o clean speech do no
dec ease.
Technique Clean SNR=20dB SNR=10dB
Whi e noise
FF1 13/12, TF1 &
FF1 13/12, TF2 98,43 97,22 91,51
FF2 13/12, TF1 &
FF2 13/12, TF2 99,28 98,03 90,95
FF1 13/12, TF1 &
FF2 13/12, TF2 99,16 97,63 92,31
Low-pass noise
FF1 13/12, TF1 &
FF2 13/12, TF2 99,16 96,62 89,66
4. CONCLUSIONS
So a , ime and equency il e ing ha e been s udied
sepa a ely. In his pape , we o e an in oduc ion o hei join
in es iga ion. We showed TF-FF ea u es a e obus agains
s a iona y, addi i e whi e and low-pass noises o isola ed digi
ecogni ion. A g ea ad an age o his echnique is ha i does
no dec ease clean speech ecogni ion esul s. In he u he
wo k, a 2-D il e can be designed, which will include di e en
FF in di e en TF egions in one ea u e se . Mo eo e , he
powe coe icien can be op imized. Also, we wan o ex end he
men ioned echniques o mo e di icul asks.
5. REFERENCES
1. He mansky, H., Mo gan, N. “RASTA P ocessing o
Speech”, IEEE T ans. on Speech and Audio P ocessing,
Vol.2, No. 4, 1-12, Oc obe 1994.
2. Nadeu, C., He nando, J., Go icho, M., “On he
Deco ela ion o Fil e -Bank Ene gies in Speech
Recogni ion”, P oc. Eu ospeech, 1381-84, 1995.
3. He nando, J., Nadeu, C. “Robus Speech Pa ame e s
Loca ed in he F equency Domain”, P oc. Eu ospeech,
417-20, 1997.
4. Nadeu, C., Paches-Leal, P., Juang, B. H. “Fil e ing he
Time Sequence o Spec al Pa ame e s o Speech
Recogni ion”, Speech Communica ion 22, 315-322, 1997.
5. Nadeu, C., Ma ino, J. B., He nando, J., Noguei as, A.
“F equency and Time Fil e ing o Fil e -Bank Ene gies o
HMM Speech Recogni ion”, P oc. ICSLP, 430-33, 1996.
6. A endaño, C., Vuu en, S., He mansky, H., “Da a Based
Fil e Design o RASTA-like Channel No maliza ion in
ASR”, P oc. ICSLP, 2087-90, 1996.
7. G eenbe g, S., Kingsbu y, B. E. D., “The Modula ion
Spec og am: in Pu sui o an In a ian Rep esen a ion o
Speech”, P oc. ICASSP, 1647-50, 1997.
8. He nando, J., Nadeu, C., “Linea P edic ion o he One-
Sided Au oco ela ion Sequence o Noisy Speech
Recogni ion”, IEEE T ans. on SAP, Vol. 5, N 1, 80-84,
1997.
Table 2: Recogni ion a es in % o wo ea u e se s
Table 3: Recogni ion a es in % o wo ea u e se s using squa e-
powe equency and ime il e ing