scieee Open visual document viewer

On the interaction between time and frequency filtering of speech parameters

Macho Ciena, Dusan,Nadeu Camprubí, Climent

Abstract

One of the great today's challenges in speech recognition is to ensure the robustness of the used speech representation. Usually, the recognition rate is strongly reduced when the speech is corrupted, e.g. by convolutional or additive noise, and the speech features are not designed to be robust. In this paper we study the effect of additive noise on the logarithmic filter-bank energy representation. We use time and frequency filtering techniques to emphasize the discriminative information and to reduce the mismatch between noisy and clean speech representation. A 2-D spectral representation is introduced to see the regions most affected by noise in the 2-D quefrency-modulation frequency domain and to help to design the frequency and time filter shapes. Experiments with one and two dynamic feature sets show the usefulness of the combination of time and frequency filtering for both, white and low-pass noise speech recognition. At the end the power time and frequency filtering technique is presented.

Full text

ON THE INTERACTION BETWEEN TIME AND FREQUENCY FILTERING OF SPEECH PARAMETERS FOR ROBUST SPEECH RECOGNITION Dušan Macho* and Climen Nadeu** *Dep . o Telecommunica ions, Slo ak Technical Uni e si y and Dep . o Speech Analysis and Syn hesis, Slo ak Academy o Sciences, B a isla a, Slo akia **Dep . o Signal Theo y and Communica ions, Uni e si a Poli ècnica de Ca alunya, Ba celona, Spain ABSTRACT One o he g ea oday’s challenges in speech ecogni ion is o ensu e he obus ness o he used speech ep esen a ion. Usually, he ecogni ion a e is s ongly educed when he speech is co up ed, e.g. by con olu ional o addi i e noise, and he speech ea u es a e no designed o be obus . In his pape we s udy he e ec o addi i e noise on he loga i hmic il e -bank ene gy ep esen a ion. We use ime and equency il e ing echniques o emphasize he disc imina i e in o ma ion and o educe he misma ch be ween noisy and clean speech ep esen a ion. A 2-D spec al ep esen a ion is in oduced o see he egions mos a ec ed by noise in he 2-D que ency- modula ion equency domain and o help o design he equency and ime il e shapes. Expe imen s wi h one and wo dynamic ea u e se s show he use ulness o he combina ion o ime and equency il e ing o bo h, whi e and low-pass noise speech ecogni ion. A he end he powe ime and equency il e ing echnique is p esen ed. 1. INTRODUCTION Only a pa o he in o ma ion con ained in he speech signal is used o speech ecogni ion. Mo eo e , a speech signal can be dis o ed by non-speech componen s (e.g. channel o mic ophone dis o ion, addi i e noise, e e be a ion…). I is necessa y o ex ac phone ically impo an ea u es wi h good disc imina i e p ope ies and obus ness when used in ad e se en i onmen s. Fo ecogni ion pu poses, speech is o en con e ed o a ime sequence o log il e -bank ene gies (log FBE). In his way, he conside ed speech uni is ep esen ed as a wo-dimensional (2- D) ime- equency sequence. This sequence is u he p ocessed in o de o ob ain mo e obus and disc imina i e ea u es (e.g. ans o med o mel-ceps um, RASTA il e ed [1]…). Recen ly, he au ho s in [2] showed ha a simple il e ing pe o med on he equency dimension o e e y ame (F equency Fil e ing – FF) gi es be e ecogni ion esul s o clean speech han ceps al coe icien s. The FF can be seen as a li e ing ope a ion pe o med in he spec al domain. The equency il e s in [2] we e designed o equalize he a iance o ceps al coe icien s and a simple, da abase independen , second-o de il e z-z-1 (he e deno ed as FF2) was ound as a good comp omise. In [3], he FF ea u es appea ed mo e obus han ceps al coe icien s when speech is dis o ed by addi i e whi e noise and a i s - o de il e 1-z-1 (FF1) ga e good ecogni ion esul s. The componen s o he speech ea u e ec o a y in ime, acco ding o he changes o he speech signal, desc ibing ime ajec o ies. The spec um o he ime ajec o y is called modula ion spec um. The ypical speech modula ion spec um dec eases along he modula ion equency axis [4]. Thus, he low modula ion equencies gene ally domina e he dis ance compu a ion in he classi ie (simila ly, as do he low que ency componen s) bu hey do no ca y he mos disc imina i e in o ma ion [5]. Mo eo e , when he speech is co up ed by s a iona y con olu ional noise, he 0 h modula ion equency is he mos a ec ed in he log FBE ep esen a ion. Thus, il e ing on he ime dimension (Time Fil e ing – TF) can emo e undesi able pa s o he modula ion spec um. In [5], bo h ime and equency il e ing we e p esen ed join ly, bu conside ing ha he e is no in e ac ion be ween hem. Howe e , we ecen ly obse ed some ac s ha led us o conside ha he in e ac ion exis s. Fi s ly, he no iceable be e clean speech pe o mance o FF wi h espec o ceps um ha is ob ained when only one s a ic ea u e se (wi hou TF) is used, may be educed o a sligh di e ence i dynamic ea u es a e included in he ep esen a ion. Second, FF looses i s good pe o mance o noisy speech when he noise is colo ed. In his wo k, we gain mo e insigh in o ha in e ac ion p oblem by using he 2-D modula ion spec um ep esen a ion ob ained om log FBE sequence. We obse ed, o example, ha he mean alue o ha 2-D unc ion o noisy speech shows highe alues a low indices han he co esponding unc ion o clean speech. Thus, TF and FF can imp o e he ecogni ion a e by a enua ing he mos dis o ed egions. Mo eo e , in he same way ha i can be con enien o use sligh ly di e en ime il e s in wo di e en equency bands [6], he use o di e en equency il e s in di e en modula ion equency egions can also inc ease he ecogni ion pe o mance o speech dis o ed by addi i e noise. Fo designing he il e s, we can ake ad an age o ha 2-D spec al ep esen a ion. Fo he ecogni ion es s p esen ed in his pape , we used he ollowing condi ions: single digi s om he adul po ion o he TI da abase, decima ed om 20 kHz o 8 kHz sampling a e; no p eemphasis; 30 ms long Hamming windowed ames wi h 10 ms shi ; 13-o de log FBE basic pa ame e iza ion scheme; con inuous densi y HMMs wi h 8 s a es pe digi and 3 s a es o he silence model; o noisy speech, ei he s a iona y whi e addi i e noise o low-pass addi i e noise wi h cu -o equency 1100 Hz we e added o he clean speech o ob ain SNR equal o 20 dB and 10 dB. T aining was pe o med always wi h clean speech and es ing wi h noisy speech. This wo k was ca ied ou du ing he s ay o D. Macho a UPC Ba celona and sponso ed by Spanish go e nmen and pa ially by Slo ak Academy o Sciences. 5 h In e na ional Con e ence on Spoken Language P ocessing (ICSLP 98) Sydney, Aus alia No embe 30 -- Decembe 4, 1998 ISCAA chi e h p://www.isca-speech.o g/a chi e 2. THE 2-D MODULATION SPECTRUM Fo be e analysis pu poses, we sp ead modula ion spec um ep esen a ion [4] in wo dimensions, whe e he modula ion spec um o e e y ceps al coe icien is p esen . Le log S(k,n) be he sho - ime log FBE es ima e o he speech signal wi h k deno ing he il e -bank ou pu and n he ame index. The 2-D modula ion spec um (in [7], he modula ion spec og am has been in oduced which displays he e olu ion o low modula ion equencies in ime and equency) is hen es ima ed by compu ing and a e aging unc ion () 2 , q mC o e a speech da abase. () 2 , q mC is ob ained by in e se disc e e-Fou ie ans o ming om he equency domain k o he que ency m and by he Fou ie ans o ming om he ime domain n o he modula ion equency domain q , 2 ),(),(),(),(log 2 qq mCmCnmcnkS nk FTIDFT ¾®¾¾¾®¾¾¾¾®¾ × . (2) The 2-D modula ion spec um es ima ed om clean isola ed digi s da abase is shown on Figu e 1(a). The dec easing il in bo h dimensions can be obse ed. Figu e 1(b) shows he 2-D modula ion spec um o speech dis o ed by addi i e whi e noise. The low indices in que ency and modula ion equency seem o be he mos a ec ed by noise. The misma ch be ween aining and es ing log FBE ep esen a ion is he main eason o he poo ecogni ion esul s ob ained when he speech co up ed by addi i e noise is used o es ing. We compu ed he misma ch be ween he clean and noisy speech ep esen a ion as () () 2 ,, qq mCmC cleannoisy - o all m and q ,(3) and a e aging i o e many speake s and u e ances we es ima ed he 2-D modula ion spec um o misma ch. Figu e 2 shows he 2-D modula ion spec a o misma ch o speech co up ed by addi i e whi e noise (a) and addi i e low-pass noise (b), bo h o SNR=10dB. The la ges misma ch is si ua ed Figu e 1: 2-D modula ion spec a o (a) clean and (b) whi e noise speech wi h SNR=10dB Figu e 2: 2-D modula ion spec a o misma ch o (a) addi i e whi e noise and (b) addi i e low-pass noise, bo h wi h SNR=10dB (a) (b) (a) (b) in low que encies and modula ion equencies wi h i s maximum a he (0,0) poin . No e he di e ence along que ency be ween bo h igu es (especially a low modula ion equencies), while along modula ion equency hey a e simila . In bo h igu es, when he modula ion equency inc eases, low and middle que encies a e less a ec ed and can be used o ecogni ion. Using FF and TF, we can emo e he dis o ed pa o he 2-D modula ion spec um and e en be e disc imina i e p ope ies o ea u es can be ob ained. I we emphasize wo di e en egions in he modula ion equency dimension by using wo di e en ime il e s, one o each o wo ea u e se s, we can use a di e en equency il e o e e y egion in o de o weigh di e en ly in he que ency dimension. In he ollowing sec ions, he e ec o equency and ime il e ing on he ecogni ion pe o mance is shown. 3. RECOGNITION TESTS 3.1 S a ic Fea u e Se I no TF is used, we e e o he speech ep esen a ion as a s a ic ea u e se . In he ollowing, only he e ec o he equency il e ing is p esen ed. Fo his pu pose, we used 13 di e en equency il e s o leng h 3 wi h sys em unc ion (z-1)(z+a), whe e a changes om –0,2 o 1,0 wi h s ep 0,1 (no e, ha he il e wi h a=0 is FF1 and a=1 is FF2). The ans o m esponse o he il e is a que ency unc ion (a li e ). Changing he pa ame e a o he il e s in he in e al <-0,2; 0,2>, he shape o he li e s in low and middle que encies changes, while does no change much in high que encies. When he pa ame e a changes in he in e al <0,6; 1,0>, he li e shape in he high que encies changes while in low and middle que encies i does no . Since he i s and he las il e ed log FBE o each ame con ain absolu e ene gy [2] hey can ca y much noise, so ha hey we e no used in his ea u e se . Figu e 3 shows he ecogni ion a es o all il e s. The clean speech ecogni ion a e (Figu e 3(a)) inc eases when a inc eases and FF2 gi es he bes esul s. Howe e , when he speech is co up ed by addi i e whi e noise (Figu e 3(b)), il e s ha a enua e low and middle que encies a e p e e able. This is due o ac ha , al hough he low and middle que encies a e use ul o clean speech ecogni ion, hey a e se e ally a ec ed by noise [8]. 3.2 One Time-Fil e ed Fea u e Se In his case, ime il e ing is applied o he sequence o ea u es. As ime il e s we used he wo di e en Slepian il e s ( he same as hose in [4] wi h pa ame e s K=1, W=12, L=14, deno ed as TF1 and K=2, W=12, L=14 deno ed as TF2) join wi h equaliza ion 1-0,97z-1. TF1 p ese es he modula ion equencies o speech oughly om 0 Hz o 3 Hz and TF2 om 2 Hz o 9 Hz. The i s es we pe o med was wi hou equency il e ing. F om he i s wo lines o Table 1 i seems ha he ea u es om he TF1 egion yield mo e disc imina i e in o ma ion (97,71% ecogni ion a e o clean speech) han hose om he TF2 egion (95,13%). Howe e , when noisy speech is ecognized, he TF2 ea u es gi e be e esul s and a e mo e obus han he ea u es om he TF1 egion. Technique Clean SNR=20dB SNR=10dB Whi e noise TF1, no FF 97,71 50,70 16,10 TF2, no FF 95,13 64,10 41,01 TF1, FF1 13/12 97,79 94,37 82,98 TF1, FF2 13/12 99,16 95,90 81,17 TF2, FF1 13/12 97,26 92,84 77,02 TF2, FF2 13/12 98,39 95,13 79,60 Low-pass noise TF1, FF1 13/12 97,79 90,38 78,11 TF1, FF2 13/12 99,16 90,30 77,14 TF2, FF1 13/12 97,26 92,23 77,99 TF2, FF2 13/12 98,39 93,32 78,63 The si ua ion changes when FF is used in conjunc ion wi h TF. Figu e 3 shows he beha io o bo h, TF1 and TF2 ea u e se s in e ms o di e en FFs. Fo clean speech, he TF1 ea u es gi e be e esul s o e e y FF han he TF2 ea u es (see Figu e 3(a)). In he noisy case, he equency il e ing pa ially educes he high con en o noise in TF1 egion and he ecogni ion a es e en ou pe o m he TF2 esul s (Figu e 3(b)). Mo eo e , a di e en beha io o he ea u e se s om wo men ioned modula ion equency egions can be obse ed. Fo he TF1 egion, he equency il e s which a enua e mo e he low and Table 1: Recogni ion a es in % using TF1, TF2 wi hou equency il e ing and wi h FF1 and FF2 90 91 92 93 94 95 96 97 98 99 100 -0,2 -0,1 0,0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1,0 F equency Fil e -> a Recogni ion Ra e [%] s a ic clean TF1 clean TF2 clean Figu e 3: Recogni ion a e in e ms o he FF and TF used in he pa ame e iza ion o (a) clean and (b) whi e noise speech wi h SNR=10dB (b) 40 45 50 55 60 65 70 75 80 -0,2 -0,1 0,0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1,0 F equency Fil e -> a Recogni ion Ra e [%] s a ic 10dB TF1 10dB TF2 10dB (a) middle que encies ( hose wi h a=<-0,2; 0,2>) gi e sligh ly be e esul s han he o he s o noisy speech. A enua ing he high que encies in his egion seems o imp o e he ecogni ion oo ( il e s wi h a=<0,6; 1,0>). Fo he TF2 egion, an inc easing endency in he ecogni ion a e can be obse ed when he coe icien a inc eases and FF2 is he op imal il e . We ied o include he i s and he las il e ed log FBE o he ea u e ec o . Only including o he i s one imp o ed he noisy speech ecogni ion. In gene al, he i s log FBE con ains mo e speech ene gy and is no a ec ed by addi i e noise so much as he las one, which con ains less speech ene gy. Table 1 shows he ecogni ion a es wi h he i s log FBE included in he ea u e ec o . In he addi i e low-pass noise case, he esul s om expe imen s when only FF is used a e e y low (nea 17% o FF1 o FF2 and SNR=10dB). This is due o ac , ha he il e ed log FBEs include he s ep o he ansi ion band o he noise spec um. Since he s ep e ec is cons an in he ime, i can be almos canceled by he ime il e . Resul s wi h di e en ime and equency il e s a e in he Table 1. 3.3 Two Time-Fil e ed Fea u e Se s The ecogni ion a e can be imp o ed using ea u es om bo h ime- il e ed egions in wo di e en ea u e se s. We ha e ound he s a ic ea u e se is a sou ce o e o s when used oge he wi h ime- il e ed ea u es and we do no use i . In he Table 2, he ecogni ion es s a e p esen ed o h ee combina ions o ime and equency il e s. Fo clean speech, he bes ecogni ion esul is ob ained when FF2 il e is used o bo h ime- il e ed ea u e se s. When FF1 is used, he ecogni ion o clean speech dec eases, bu inc eases o noisy speech. F om Figu e 3(b) i can be obse ed (he e he ea u e se s we e used sepa a ely), ha o he TF1 ea u e se he equency il e s which a enua e low and middle que encies a e p e e able and o TF2 ea u es, he FF2 is he bes il e . Using his obse a ion, an addi ional imp o emen o noisy speech ecogni ion was ob ained. A he end o Table 2 he bes ecogni ion a es o low-pass noise speech a e men ioned. Technique Clean SNR=20dB SNR=10dB Whi e noise FF1 13/12, TF1 & FF1 13/12, TF2 98,15 96,18 86,68 FF2 13/12, TF1 & FF2 13/12, TF2 99,48 97,06 84,59 FF1 13/12, TF1 & FF2 13/12, TF2 99,12 96,74 88,01 Low-pass noise FF1 13/12, TF1 & FF2 13/12, TF2 99,12 94,16 81,41 3.4 Powe F equency and Time Fil e ing In his echnique we assumed, ha in he log FBE ep esen a ion o noisy speech he high-ene gy coe icien s a e less a ec ed by noise han he coe icien s wi h low ene gy con en . Thus, we use simple powe ope a ion on he log FBEs be o e hey en e o he FF in o de o emphasize he high-ene gy coe icien s. In gene al, he powe ope a ion can be exp essed as () g nkS,log . Table 3 shows he esul s om he same expe imen s as Table 2 bu using squa e-powe equency and ime il e ing ( 2 = g ) o wo ea u e se s. A clea imp o emen o noisy speech can be ob ained while he ecogni ion a es o clean speech do no dec ease. Technique Clean SNR=20dB SNR=10dB Whi e noise FF1 13/12, TF1 & FF1 13/12, TF2 98,43 97,22 91,51 FF2 13/12, TF1 & FF2 13/12, TF2 99,28 98,03 90,95 FF1 13/12, TF1 & FF2 13/12, TF2 99,16 97,63 92,31 Low-pass noise FF1 13/12, TF1 & FF2 13/12, TF2 99,16 96,62 89,66 4. CONCLUSIONS So a , ime and equency il e ing ha e been s udied sepa a ely. In his pape , we o e an in oduc ion o hei join in es iga ion. We showed TF-FF ea u es a e obus agains s a iona y, addi i e whi e and low-pass noises o isola ed digi ecogni ion. A g ea ad an age o his echnique is ha i does no dec ease clean speech ecogni ion esul s. In he u he wo k, a 2-D il e can be designed, which will include di e en FF in di e en TF egions in one ea u e se . Mo eo e , he powe coe icien can be op imized. Also, we wan o ex end he men ioned echniques o mo e di icul asks. 5. REFERENCES 1. He mansky, H., Mo gan, N. “RASTA P ocessing o Speech”, IEEE T ans. on Speech and Audio P ocessing, Vol.2, No. 4, 1-12, Oc obe 1994. 2. Nadeu, C., He nando, J., Go icho, M., “On he Deco ela ion o Fil e -Bank Ene gies in Speech Recogni ion”, P oc. Eu ospeech, 1381-84, 1995. 3. He nando, J., Nadeu, C. “Robus Speech Pa ame e s Loca ed in he F equency Domain”, P oc. Eu ospeech, 417-20, 1997. 4. Nadeu, C., Paches-Leal, P., Juang, B. H. “Fil e ing he Time Sequence o Spec al Pa ame e s o Speech Recogni ion”, Speech Communica ion 22, 315-322, 1997. 5. Nadeu, C., Ma ino, J. B., He nando, J., Noguei as, A. “F equency and Time Fil e ing o Fil e -Bank Ene gies o HMM Speech Recogni ion”, P oc. ICSLP, 430-33, 1996. 6. A endaño, C., Vuu en, S., He mansky, H., “Da a Based Fil e Design o RASTA-like Channel No maliza ion in ASR”, P oc. ICSLP, 2087-90, 1996. 7. G eenbe g, S., Kingsbu y, B. E. D., “The Modula ion Spec og am: in Pu sui o an In a ian Rep esen a ion o Speech”, P oc. ICASSP, 1647-50, 1997. 8. He nando, J., Nadeu, C., “Linea P edic ion o he One- Sided Au oco ela ion Sequence o Noisy Speech Recogni ion”, IEEE T ans. on SAP, Vol. 5, N 1, 80-84, 1997. Table 2: Recogni ion a es in % o wo ea u e se s Table 3: Recogni ion a es in % o wo ea u e se s using squa e- powe equency and ime il e ing