scieee Open visual document viewer

Linear prediction of the one-sided autocorrelation sequence for noisy speech recognition

Hernando Pericás, Francisco Javier,Nadeu Camprubí, Climent

Abstract

The article presents a robust representation of speech based on AR modeling of the causal part of the autocorrelation sequence. In noisy speech recognition, this new representation achieves better results than several other related techniques.

Full text

80 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 small. Al hough inc eased aining da a may be help ul, in he speake dependen Manda in syllable ecogni ion p oblem, a limi ed da abase will p obably s ill be a no mal si ua ion o some pe iod o ime in he u u e. VII. CONCLUSION A new app oach is p oposed in his co espondence o ob ain mo e elabo a e ini ial models co e ing cha ac e is ics o di e en ones. Imp o ed s a e ansi ion opologies a e also ound o achie e be e pe o mance compa ed wi h he simple le - o- igh model wi h wo ansi ions. A h eshold decision app oach is u he de eloped o imp o e he pe o mance o BS ecogni ion o he syllables wi h he neu al one. The es esul s on e e yday Chinese show ha a o al e o a e educ ion on he o de o 20% in he op 1 a e can be ob ained when all he concep s a e p ope ly in eg a ed. Al hough he echniques he e a e p oposed specially o ecogni ion o Manda in base syllables conside ing he e ec o ones, i is ce ainly belie ed ha simila concep s a e po en ially applicable o sol e simila p oblems in speech ecogni ion in o he languages. REFERENCES [1] R. He, Ed., Guoyu bao Tzdian (Manda in Chinese Daily Dic iona y). Taipei, R.O.C.: Guoyu bao, 1976. [2] L.-S. Lee, C.-Y. Tseng, and M. Ouh-Young, “The syn hesis ules in a Chinese ex - o speech sys em,” IEEE T ans. Acous ., Speech, Signal P ocessing, pp. 1309–1320, Sep . 1989. [3] Y. R. Chao, A G amma o Spoken Chinese. Be keley, CA: Uni e si y o Cali o nia Be keley P ess, 1968. [4] W. J. Yang, J. C. Lee, Y. C. Chang, and H. C. Wang, “Hidden Ma ko model o manda in lexical one ecogni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, pp. 988–992, July 1988. [5] F.-H. Liu, Y. Lee, and L. S. Lee, “A Di ec -conca ena ion app oach o ain hidden ma ko models o ecognize he highly con using manda in syllables wi h e y limi ed aining da a,” IEEE T ans. Speech Audio P ocessing, ol. 1, no. 1, pp. 113–119, Jan. 1993. [6] L.-S. Lee e al., “Golden manda in (I)—A eal- ime manda in speech dic a ion machine o chinese language wi h e y la ge ocabula y,” IEEE T ans. Speech Audio P ocessing, ol. 1, no. 2, pp. 158–179, Ap . 1993. [7] B.-H. Juang and L. R. Rabine , “Mix u e au o eg essi e hidden Ma ko models o speech signals,” IEEE T ans. Acous ., Speech, Signal P o- cessing, ol. ASSP-33, no. 6, pp. 1404–1413, Dec. 1985. [8] X. Huang e al., “The SPHNIX-II speech ecogni ion sys em: An o e iew,” Compu . Speech Language, pp. 137–148, Feb. 1993. Linea P edic ion o he One-Sided Au oco ela ion Sequence o Noisy Speech Recogni ion Ja ie He nando and Climen Nadeu Abs ac — The aim o his co espondence is o p esen a obus ep esen a ion o speech based on AR modeling o he causal pa o he au oco ela ion sequence. In noisy speech ecogni ion, his new ep- esen a ion achie es be e esul s han se e al o he ela ed echniques. I. INTRODUCTION Linea p edic i e coding (LPC) [1] is a spec al es ima ion ech- nique widely used in speech p ocessing and, pa icula ly, in speech ecogni ion. Howe e , he con en ional LPC echnique, which is equi alen o AR modeling o he signal x ( n ) , is known o be e y sensi i e o he p esence o backg ound noise. This ac leads o poo ecogni ion a es when his echnique is used in speech ecogni ion unde noisy condi ions, e en i only a mode a e le el o con amina ion is p esen in he speech signal. Simila esul s a e ob ained wi h he well-known mel-ceps um echnique [2]. This explains why some o he main a emp s o comba he noise p oblem consis o inding no el acous ic ep esen a ions ha a e mo e esis an o noise co up ion han adi ional pa ame e iza ion echniques. Linea p edic ion o he au oco ela ion sequence has been he common app oach o se e al obus spec al es ima ion me hods o noisy signals p esen ed in he pas . Fo speech ecogni ion, Mansou and Juang [3] p oposed he sho - ime modi ied cohe ence (SMC) as a obus ep esen a ion o speech based on ha app oach. On he o he hand, Cadzow [4] in oduced he use o an o e de e mined se o Yule–Walke equa ions o obus modeling o ime se ies. Al hough Cadzow applies linea p edic ion o he signal, his me hod can also be in e p e ed as pe o ming linea p edic ion in he au oco ela ion domain. Bo h me hods ely, ei he explici ly o implici ly, on he ac ha he au oco ela ion sequence is less a ec ed by b oadband noise han he signal i sel , especially a high lag indices. In his wo k, we conside he one-sided o causal pa o he au oco ela ion sequence and i s ma hema ical p ope ies. As his sequence sha es i s poles wi h he signal x ( n ) , i p o ides a good s a ing poin o LPC modeling. In his way, he new one-sided au oco ela ion LPC (OSALPC) me hod appea s as a s aigh o wa d esul o he app oach [5]. In addi ion, i is closely ela ed o he SMC ep esen a ion and Cadzow’s me hod. All o hem can be in e p e ed as AR modeling o ei he a spec al unc ion named “en elope” o i s squa e. This in e p e a ion, which is based on he p ope ies o he one-sided au oco ela ion, p o ides mo e insigh in o he a ious me hods. In his co espondence, hei pe o mance in noisy speech ecogni ion is compa ed. The op imum model o de and ceps al li e ing ha e also been in es iga ed in noisy condi ions. The simula ion esul s show ha OSALPC ou pe o ms he o he echniques in se e e noisy condi ions and ob ains simila sco es o mode a e o high SNR. Manusc ip ecei ed Feb ua y 14, 1995; e ised No embe 9, 1995. This wo k was suppo ed by G an nos. TIC-92-0800-C05/04 and TIC-92-1026- C02/02. The associa e edi o coo dina ing he e iew o his pape and app o ing i o publica ion was D . Kuldip K. Paliwal. The au ho s a e wi h he Depa men o Signal Theo y and Communica ions, Poly echnical Uni e si y o Ca alonia, Ba celona, Spain. Publishe I em Iden i ie S 1063-6676(97)00766-9. 1063–6676/97$10.00 1997 IEEE IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 81 This co espondence is o ganized in he ollowing way. In Sec ion II, he OSALPC echnique is in oduced, and i s ela ionship wi h he con en ional LPC app oach and he o he pa ame e iza ions based on AR modeling in he au oco ela ion domain is discussed. Sec ion III epo s he applica ion o all hose pa ame e iza ion echniques o an isola ed wo d mul ispeake ecogni ion ask, using he HMM app oach, in o de o compa e hei pe o mances in he p esence o addi i e whi e noise. Finally, some conclusions a e summa ized in Sec ion IV. II. AR MODELING IN THE AUTOCORRELATION DOMAIN F om he au oco ela ion sequence R ( m ) , we de ine he one-sided (causal pa o he) au oco ela ion (OSA) sequence in he ollowing way: R + ( m )= R ( m ) m> 0 R (0) 2 m =0 0 m< 0 (1) I s Fou ie ans o m is he complex “spec um” S + ( ! )= 1 2 [ S ( ! )+ jS H ( ! )] (2) whe e S ( ! ) is he eal spec um, i.e., he Fou ie ans o m o R ( m ) , and S H ( ! ) is he Hilbe ans o m o S ( ! ) . Due o he analogy be ween S + ( ! ) in (2) and he analy ic signal used in ampli ude modula ion, a spec al “en elope” E ( ! ) [6] can be de ined as E ( ! )= j S + ( ! ) j : (3) Due o he la ge dynamic ange o speech spec a, he en elope E ( ! ) s ongly enhances he highes powe equency bands wi h espec o S ( ! ) [5]. Consequen ly, he noise componen s lying ou side he enhanced equency bands a e la gely a enua ed in E ( ! ) wi h espec o S ( ! ) , and hus, E ( ! ) is mo e obus o b oadband noise han S ( ! ) . On he o he hand, as i is well known, he OSA sequence R + ( m ) and he signal x ( n ) ha e he same poles [7]. Those wo p ope ies, i.e., obus ness o noise and pole p ese a- ion, sugges ha AR pa ame e s o he speech signal can be mo e eliably es ima ed om he OSA sequence R + ( m ) han di ec ly om he signal x ( n ) when x ( n ) is co up ed by b oadband noise. Thus, as he con en ional LPC echnique assumes an all-pole model o he speech spec um S ( ! ) , we may apply linea p edic ion o he OSA sequence, assuming an all-pole model o i s “spec um” E 2 ( ! ) . This is he basis o he one-sided au oco ela ion linea p edic i e coding (OSALPC) pa ame e iza ion echnique [5]. A s aigh o wa d algo i hm is p oposed in [5] ha calcula es he OSALPC ceps al coe icien s. I consis s o applying he (windowed) au oco ela ion me hod o linea p edic ion o an es ima ion o he OSA sequence: a) Fi s , om he speech ame o leng h N , he au oco ela ion lags un il M = N= 2 a e compu ed ( his alue o M was empi ically op imized o conside he well-known adeo be ween a iance and equency esolu ion o he spec al es ima e [8]). b) Second, he Hamming window om m =0 o M is applied on such es ima ed OSA sequence. c) Thi d, i p is he p edic ion o de , he i s p +1 au oco ela ion alues o ha OSA sequence a e compu ed om m =0 o p , using he con en ional biased es ima o , i.e., he one ha is commonly employed in speech p ocessing. d) Then, hese alues a e used as en ies o he Le inson–Du bin algo i hm o es ima e he AR pa ame e s ak , k =1 ; 111 ;p: Fig. 1. Robus ness o he OSALPC ep esen a ion o addi i e whi e noise: (a) LPC spec um and (b) OSALPC squa ed en elope o a oiced speech ame in noise ee condi ions (solid line) and SNR equal o 0 dB (do ed line). e) Finally, he ceps al coe icien s co esponding o he model a e ecu en ly compu ed om hose AR pa ame e s. The obus ness o OSALPC o addi i e whi e noise is illus a ed in Fig. 1. As can be seen in his igu e, he OSALPC squa ed en elope shows a p ominen i s o man , and i s whole cu e is mo e obus o addi i e whi e noise han ha o he LPC spec um. In his case, he con en ional biased au oco ela ion es ima o was used o compu e he OSA sequence om he signal. Fig. 1 also shows ha spu ious peaks may appea in he OSALPC squa e en elope. They a e p obably due o he ac ha he OSALPC echnique pe o ms only a pa ial decon olu ion o he speech signal [9]. In spi e o ha , OSALPC shows a be e speech ecogni ion pe o mance han con en ional LPC in se e e condi ions o addi i e whi e noise, as will be seen in he nex sec ion. The OSALPC echnique is closely ela ed o he sho - ime modi- ied cohe ence (SMC) ep esen a ion p oposed by Mansou and Juang in [3]. SMC is also based on AR modeling in he au oco ela ion domain. Howe e , whe eas in he OSALPC echnique, he en ies o he Le inson–Du bin algo i hm ( i s p alues o he au oco ela ion o he OSA sequence) a e calcula ed om he OSA sequence using he con en ional biased au oco ela ion es ima o , in he SMC ep esen- a ion, hey a e compu ed using a squa e oo spec al shape . In ac , in e ms o he abo e o mula ion, ha di e ence lies in assuming in he SMC echnique an all-pole spec al model o he en elope E ( ! ) ins ead o E 2 ( ! ) . Fu he mo e, R + (0) is se o 0 in he case o addi i e whi e noise because i is se e ely co up ed by noise. On he o he hand, he name o he SMC ep esen a ion de i es om he usage o a pa icula es ima o , which is e e ed o as cohe ence in [3], o compu e he OSA sequence om he signal. This es ima o is a mo e homogeneous measu e han he con en ional biased au oco ela ion es ima o in he sense ha e e y es ima ed alue is compu ed using he same numbe o signal samples, whe eas in he con en ional es ima o , he numbe o signal samples employed o es ima e R ( m ) dec eases along he index m . Tha p ope y does no ha e much ele ance in he es ima ion o he au oco ela ion en ies o he Le inson–Du bin algo i hm since only he i s p +1 82 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 Fig. 2. Ma ix o mula ion o OSALPC and LSMYE me hods. alues a e conside ed, and usually, p  N . Howe e , i may be impo an in he es ima ion o he OSA sequence om he speech signal since he OSA leng h conside ed in bo h OSALPC and SMC echniques is M = N= 2 and no negligible wi h espec o N . The OSALPC echnique can also be easily ela ed o he o e de- e mined se o Yule–Walke equa ions p oposed by Cadzow in [4] o seek ARMA models o ime se ies. As an AR ( p ) p ocess con amina ed by addi i e whi e noise becomes an ARMA ( p; p ) p ocess, Cadzow’s me hod can be used o es ima e he pa ame e s o his noisy AR p ocess simply by se ing he same AR and MA o de s in he so-called leas squa es modi ied Yule–Walke equa ions (LSMYWE’s) [8]. The ela ionship be ween he OSALPC and LSMYWE echniques is illus a ed by he ma ix equa ion in Fig. 2, whe e M deno es he highes au oco ela ion lag used, and e ( m ) is he e o o be minimized. The minimiza ion o he no m o he ull e o ec o e ( m ) g m =1 ; 111 ;M + p wi h espec o he AR pa ame e s a k is equi alen o he applica ion o he (windowed) au oco ela ion me hod o linea p edic ion o he sequence R ( m ) , m =1 ; 111 ;M , i.e., he OSALPC echnique. On he o he hand, he LSMYWE echnique minimizes he no m o he sub ec o e ( m ) g m = p +1 ; 111 ;M , and he e o e, i amoun s o applying he (unwindowed) co a iance me hod o linea p edic ion on he same ange o au oco ela ion lags. When M =2 p , LSMYWE a e he modi ied Yule–Walke equa ions [8] o an ARMA ( p; p ) p ocess. In bo h cases, only au oco ela ion lags co esponding o he OSA sequence a e employed. In ou compa ison, we will also conside ano he e sion o his (unwindowed) co a iance-based app oach ha will be called leas squa es Yule–Walke equa ions (LSYWE’s). Whe eas in he LSMYWE echnique he i s p edic ed au oco ela ion alue is R ( p + 1) , in he LSYWE echnique, he p edic ion begins a R (1) . Bo h LSMYWE and LSYWE me hods and hei ela ionship o OSALPC a e g aphically desc ibed in Fig. 3. As i is shown, he only di e ence be ween he a ious echniques is he ange o au oco ela ion lags conside ed in he minimiza ion o he e o . I is wo h no ing ha LSYWE conside s some nega i e au oco ela ion lags ha do no belong o he OSA sequence. In pa icula , i M is equal o p , LSYWE a e he con en ional Yule–Walke equa ions. As will be seen in he nex sec ion, in spi e o he simila i y be ween all hese echniques, he OSALPC ep esen a ion ou pe o ms he LSYWE, LSMYWE, and SMC echniques in speech ecogni ion in se e e noisy condi ions. On he o he hand, as a as he compu a ional complexi y o he algo i hms is conce ned, OSALPC and SMC echniques a e much mo e e icien han LSYWE and LSMYWE echniques because hey use he Le inson–Du bin algo i hm. Finally, i is wo h no ing ha he OSALPC echnique may be included in he ield o highe o de spec al es ima ion due o he ac ha he squa ed en elope E 2 ( ! ) is he Fou ie ans o m o he au oco ela ion o he OSA sequence, which is a pa icula ou h- o de momen o he signal. Fig. 3. In e p e a ion o he (a) OSALPC, (b) LSMYWE, and (c) LSYWE app oaches as applica ion o he au oco ela ion o co a iance me hods o linea p edic ion o an au oco ela ion sequence in di e en lag anges. III. SPEECH RECOGNITION EXPERIMENTS This sec ion epo s he applica ion o all he abo e pa ame e - iza ion echniques o ecognize isola ed wo ds in a mul ispeake ask wi h a disc e e HMM-based sys em in o de o compa e hei pe o mance and o gain some insigh in o he me i o he OSALPC ep esen a ion in he p esence o addi i e whi e noise A. Speech Da abase and Recogni ion Sys em The da abase used in ou expe imen s consis s o 10 epe i ions o he Ca alan digi s u e ed by se en male and h ee emale speake s (1000 wo ds) and eco ded in a quie oom. Fi s , he sys em was ained wi h hal o he da abase and es ed wi h he o he hal . Then, he oles o bo h hal es we e changed, and he epo ed esul s we e ob ained by a e aging hose wo esul s. The analog speech signal was i s bandpass il e ed o 100–3400 Hz by an an ialiasing il e , sampled a 8 kHz and, 12 bi s quan ized. The digi ized clean speech was manually endpoin ed o de e mine he bounda ies o each wo d. The endpoin s ob ained in his way we e used in all ou expe imen s, including hose in which noise was added o he signal. Clean speech was used o aining in all he expe imen s. Noisy speech was simula ed by adding ze o mean whi e Gaussian noise o he clean signal so ha he SNR o he esul ing signal becomes 1 (clean), 20, 10, and 0 dB. No p eemphasis was pe o med. In he pa ame e iza ion s age o he ecogni ion sys em, he signal was di ided in o ames o 30 ms a a a e o 15 ms, and each ame was cha ac e ized by i s ceps al pa ame e s ob ained ei he by he con en ional LPC me hod o by any o he echniques p esen ed in he las sec ion. Be o e en e ing he ecogni ion s age, he ceps al pa ame e s we e ec o quan ized using bo h a codebook o 64 codewo ds and he Euclidean dis ance measu e be ween li e ed ceps al ec o s. Each digi was cha ac e ized by a le - o- igh disc e e hidden Ma ko model o 10 s a es wi hou skips. T aining and es ing we e pe o med using Baum–Welch and Vi e bi algo i hms, espec i ely. B. Recogni ion Resul s Fi s o all, we ca ied ou some expe imen s wi h he abo e desc ibed speech ecogni ion sys em o op imize he model o de and IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 83 TABLE I RECOGNITION RATES OF THE CONVENTIONAL LPC TECHNIQUE FOR SEVERAL PREDICTION ORDER VALUES AND CEPSTRAL LIFTERS TABLE II RECOGNITION RATES OF THE CONVENTIONAL LPC, LSMYWE, AND LSYWE TECHNIQUES FOR p =12 AND THE SLOPE LIFTER TABLE III RECOGNITION RATES OF THE CONVENTIONAL LPC, SMC, AND OSALPC TECHNIQUES FOR p =12 AND THE SLOPE LIFTER he ype o ceps al li e in he con en ional LPC echnique. In Table I, he ecogni ion esul s o LPC model o de s p = 8, 12, and 16 and o he bandpass [10], in e se o s anda d de ia ion [11] (ISD), and slope [12] li e s a e p esen ed. The ecogni ion esul s show ha nei he he model o de no he ype o ceps al li e a e ele an o ou ask in noise- ee condi ions. Howe e , in he p esence o noise, he ecogni ion esul s a e e y sensi i e o bo h ac o s. I is also clea om Table I ha he nonsymme ical li e s—slope and ISD—ou pe o m he bandpass li e o e e y model o de . This may be due o he ac ha in he p esence o whi e noise, he lowe o de ceps al coe icien s a e mo e a ec ed han he highe o de ones in he unca ed ceps al ec o . The bes esul s o se e e noisy condi ions—10 and 0 dB o SNR—a e ob ained using slope li e and p edic ion o de p equal o 12. The con enience o his ela i ely high o de comes om he ac ha he sensi i i y o he au oco ela ion sequence o addi i e whi e noise ends o dec ease along he lag index. Model o de s ha a e oo high, howe e , yield poo ecogni ion esul s since he spec al es ima e shows spu ious peaks. Ac ually, ecogni ion a es we e calcula ed using he slope li e o a la ge ange o alues o he model o de , and he bes esul s we e hose ob ained o p = 12. In Table II, he ecogni ion a es o con en ional LPC, LSMYWE, and LSYWE app oaches a e p esen ed, using M = N= 2 and bo h op imum model o de and li e ob ained o he con en ional LPC echnique, i.e., p = 12 and he slope li e . Ob iously, hese a e no he op imum condi ions o each pa ame e iza ion echnique, bu he esul s can help o compa e hei pe o mance. As can be seen om Table II, he con en ional LPC echnique ou pe o ms no iceably he o he app oaches. Howe e , he excellen pe o mance o he LSYWE app oach in noise- ee condi ions is wo h no ing. Fig. 4. Compa ison o ecogni ion a es o he LPC, SMC, OSALPC-I and OSALPC-II echniques. Fig. 5. Block diag am o he calcula ion o he LPC, SMC, OSALPC-I and OSALPC-II ceps a. In Table III and Fig. 4, he ecogni ion a es co esponding o he con en ional LPC echnique, he SMC ep esen a ion, and he no el OSALPC app oach a e p esen ed, whe e we also use M = N= 2 , p =12 , and he slope li e . The wo e sions OSALPC-I and OSALPC-II o he OSALPC app oach co espond o he OSA es ima o s o which we e e ed in Sec ion II: OSALPC-I uses he con en ional biased au oco ela ion es ima o , and OSALPC-II like SMC uses he cohe ence es ima o (and se s R (0) o 0). Fig. 5 shows 84 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 TABLE IV RECOGNITION RATES FOR THE OSALPC-II TECHNIQUE FOR SEVERAL PREDICTION ORDER VALUES AND CEPSTRAL LIFTERS a block diag am o he calcula ion o he LPC, SMC, OSALPC-I, and OSALPC-II ceps a ha pe mi s compa ison o hei espec i e algo i hms. The OSALPC and SMC ep esen a ions clea ly ou do he con- en ional LPC echnique in se e e noisy condi ions: OSALPC-I and OSALPC-II a es a e be e han LPC ones a 10 and 0 dB, and SMC ou pe o ms LPC a 0 dB. Mo eo e , OSALPC-I and OSALPC-II ep esen a ions ou pe o m he SMC echnique in all noisy condi ions. Fo he OSALPC ep esen a ion, he use o he con en ional biased au oco ela ion es ima o o compu ing he OSA sequence ( e sion OSALPC-I) is con enien in se e e noisy condi ions, i.e., o an SNR o 10 o 0 dB. Howe e , in noise- ee condi ions, he e is a loss o ecogni ion pe o mance in he OSALPC and SMC app oaches wi h espec o he con en ional LPC echnique due o he impe ec decon olu ion o he speech signal pe o med by hose echniques. This e ec seems o be minimized by using he cohe ence es ima o o compu e he OSA sequence, as in he case o OSALPC-II and SMC. Finally, Table IV shows he ecogni ion a es co esponding o OSALPC-II o he same model o de s and ceps al li e s as in Table I. I can be no iced ha he new echnique is less sensi i e o changes in bo h he model o de and he ype o ceps al li e han he con en ional LPC app oach, p o ided ha he model o de is no oo low. IV. CONCLUSIONS In his co espondence, se e al LPC-based echniques ha wo k in he au oco ela ion domain a e p esen ed and compa ed in noisy speech ecogni ion. The OSALPC echnique, which is based on he applica ion o he (windowed) au oco ela ion me hod o linea p edic ion o he one-sided au oco ela ion sequence, yields he bes esul s among all he compa ed LPC-based echniques in se e e noisy condi ions. REFERENCES [1] F. I aku a, “Minimum p edic ion esidual p inciple applied o speech ecogni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-23, pp. 67–72, 1975. [2] S. B. Da is and P. Me mels ein, “Compa ison o pa ame ic ep e- sen a ions o monosyllabic wo d ecogni ion in con inuously spoken sen ences,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP- 28, pp. 357–366, 1980. [3] D. Mansou and B. H. Juang, “The sho - ime modi ied cohe ence ep esen a ion and i s applica ion o noisy speech ecogni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. 37, pp. 795–804, 1989. [4] J. A. Cadzow, “Spec al es ima ion: An o e de e mined a ional model equa ion app oach,” P oc. IEEE, ol. 70, pp. 907–939, 1982. [5] J. He nando and C. Nadeu, “Speech ecogni ion in noisy ca en i on- men based on OSALPC ep esen a ion and obus simila i y measu ing echniques,” in P oc. ICASSP’94, Adelaide, Ap . 1994, pp. 69–72. [6] M. A. Lagunas and M. Amengual, “Non-linea spec al es ima ion,” in P oc. ICASSP’87, Dallas, Ap . 1987, pp. 2035–2038. [7] D. P. McGinn and D. H. Johnson, “Reduc ion o all-pole pa ame e es ima ion bias by successi e au oco ela ion,” in P oc. ICASSP’83, Bos on, Ap . 1983, pp. 1088–1091. [8] S. L. Ma ple, J ., Ed., Digi al Spec al Analysis wi h Applica ions. Englewood Cli s, NJ: P en ice-Hall, 1987. [9] C. Nadeu, J. Pascual, and J. He nando, “Pi ch de e mina ion using he ceps um o he one-sided au oco ela ion sequence,” in P oc. ICASSP’91, To on o, Canada, May 1991, pp. 3677–3680. [10] B. H. Juang, L. R. Rabine , and J. G. Wilpon, “On he use o band-pass li e ing in speech ecogni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-35, pp. 947–954, 1987. [11] Y. Tohku a, “A weigh ed ceps al dis ance measu e o speech ecog- ni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-35, pp. 1414–1422, 1987. [12] B. A. Hanson and H. Waki a, “Spec al slope dis ance measu es wi h linea p edic ion analysis o wo d ecogni ion in noise,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-35, pp. 968–973, 1987. A Fas Algo i hm o Finding he Adap i e Componen Weigh ed Ceps um o Speake Recogni ion Mihailo S. Zilo ic, Ra i P. Ramachand an, and Richa d J. Mammone Abs ac — In speake ecogni ion sys ems, he adap i e componen weigh ed (ACW) ceps um has been shown o be mo e obus han he con en ional linea p edic i e (LP) ceps um. The ACW ceps um is de i ed om a pole-ze o ans e unc ion whose denomina o is he p h-o de LP polynomial A ( z ). The nume a o is a ( p 0 1 ) h-o de polynomial ha is up o now ound as ollows. The oo s o A ( z )a e compu ed, and he co esponding esidues ob ained by a pa ial ac ion expansion o 1 =A ( z ) a e se o uni y. The e o e, he nume a o is he sum o all he ( p 0 1 ) h-o de co ac o s o A ( z ). In his co espondence, we show ha he nume a o polynomial is me ely he de i a i e o he denomina o polynomial A ( z ). This g ea ly speeds up he compu a ion o he nume a o polynomial coe icien s since i in ol es a simple scaling o he denomina o polynomial coe icien s. Roo inding is comple ely elimina ed. Since he denomina o is gua an eed o be minimum phase and he nume a o can be p o en o be minimum phase, wo sepa a e ecu sions in ol ing he polynomial coe icien s es ablishes he ACW cep- s um. This new me hod, which a oids oo inding, educes he compu e ime signi ican ly and imposes negligible o e head when compa ed wi h he app oach o inding he LP ceps um. I. INTRODUCTION Speake ecogni ion is he ask o iden i ying a speake by his o he oice [1]. A common p oblem in ealizing obus speake ecogni ion sys ems is ha a misma ch in aining and es ing condi ions se iously deg ades he pe o mance [2]. One o he pu sued app oaches o Manusc ip ecei ed Feb ua y 3, 1995; e ised July 13, 1996. The associa e edi o coo dina ing he e iew o his pape and app o ing i o publica ion was D . Joseph Campbell. M. S. Zilo ic is wi h Bell Communica ions Resea ch, Red Bank, NJ USA. R. P. Ramachand an and R. J. Mammone a e wi h he CAIP Cen e , Depa men o Elec ical Enginee ing, Ru ge s Uni e si y, Pisca away, NJ 08855 USA. Publishe I em Iden i ie S 1063-6676(97)00762-1. 1063–6676/97$10.00 1997 IEEE