80 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997
small. Al hough inc eased aining da a may be help ul, in he speake
dependen Manda in syllable ecogni ion p oblem, a limi ed da abase
will p obably s ill be a no mal si ua ion o some pe iod o ime in
he u u e.
VII. CONCLUSION
A new app oach is p oposed in his co espondence o ob ain
mo e elabo a e ini ial models co e ing cha ac e is ics o di e en
ones. Imp o ed s a e ansi ion opologies a e also ound o achie e
be e pe o mance compa ed wi h he simple le - o- igh model wi h
wo ansi ions. A h eshold decision app oach is u he de eloped
o imp o e he pe o mance o BS ecogni ion o he syllables
wi h he neu al one. The es esul s on e e yday Chinese show
ha a o al e o a e educ ion on he o de o 20% in he op 1
a e can be ob ained when all he concep s a e p ope ly in eg a ed.
Al hough he echniques he e a e p oposed specially o ecogni ion
o Manda in base syllables conside ing he e ec o ones, i is
ce ainly belie ed ha simila concep s a e po en ially applicable o
sol e simila p oblems in speech ecogni ion in o he languages.
REFERENCES
[1] R. He, Ed., Guoyu bao Tzdian (Manda in Chinese Daily Dic iona y).
Taipei, R.O.C.: Guoyu bao, 1976.
[2] L.-S. Lee, C.-Y. Tseng, and M. Ouh-Young, “The syn hesis ules in
a Chinese ex - o speech sys em,” IEEE T ans. Acous ., Speech, Signal
P ocessing, pp. 1309–1320, Sep . 1989.
[3] Y. R. Chao, A G amma o Spoken Chinese. Be keley, CA: Uni e si y
o Cali o nia Be keley P ess, 1968.
[4] W. J. Yang, J. C. Lee, Y. C. Chang, and H. C. Wang, “Hidden Ma ko
model o manda in lexical one ecogni ion,” IEEE T ans. Acous .,
Speech, Signal P ocessing, pp. 988–992, July 1988.
[5] F.-H. Liu, Y. Lee, and L. S. Lee, “A Di ec -conca ena ion app oach o
ain hidden ma ko models o ecognize he highly con using manda in
syllables wi h e y limi ed aining da a,” IEEE T ans. Speech Audio
P ocessing, ol. 1, no. 1, pp. 113–119, Jan. 1993.
[6] L.-S. Lee e al., “Golden manda in (I)—A eal- ime manda in speech
dic a ion machine o chinese language wi h e y la ge ocabula y,”
IEEE T ans. Speech Audio P ocessing, ol. 1, no. 2, pp. 158–179, Ap .
1993.
[7] B.-H. Juang and L. R. Rabine , “Mix u e au o eg essi e hidden Ma ko
models o speech signals,” IEEE T ans. Acous ., Speech, Signal P o-
cessing, ol. ASSP-33, no. 6, pp. 1404–1413, Dec. 1985.
[8] X. Huang e al., “The SPHNIX-II speech ecogni ion sys em: An
o e iew,” Compu . Speech Language, pp. 137–148, Feb. 1993.
Linea P edic ion o he One-Sided Au oco ela ion
Sequence o Noisy Speech Recogni ion
Ja ie He nando and Climen Nadeu
Abs ac — The aim o his co espondence is o p esen a obus
ep esen a ion o speech based on AR modeling o he causal pa o
he au oco ela ion sequence. In noisy speech ecogni ion, his new ep-
esen a ion achie es be e esul s han se e al o he ela ed echniques.
I. INTRODUCTION
Linea p edic i e coding (LPC) [1] is a spec al es ima ion ech-
nique widely used in speech p ocessing and, pa icula ly, in speech
ecogni ion. Howe e , he con en ional LPC echnique, which is
equi alen o AR modeling o he signal
x
(
n
)
, is known o be
e y sensi i e o he p esence o backg ound noise. This ac leads
o poo ecogni ion a es when his echnique is used in speech
ecogni ion unde noisy condi ions, e en i only a mode a e le el
o con amina ion is p esen in he speech signal. Simila esul s
a e ob ained wi h he well-known mel-ceps um echnique [2]. This
explains why some o he main a emp s o comba he noise
p oblem consis o inding no el acous ic ep esen a ions ha a e
mo e esis an o noise co up ion han adi ional pa ame e iza ion
echniques.
Linea p edic ion o he au oco ela ion sequence has been he
common app oach o se e al obus spec al es ima ion me hods o
noisy signals p esen ed in he pas . Fo speech ecogni ion, Mansou
and Juang [3] p oposed he sho - ime modi ied cohe ence (SMC) as a
obus ep esen a ion o speech based on ha app oach. On he o he
hand, Cadzow [4] in oduced he use o an o e de e mined se o
Yule–Walke equa ions o obus modeling o ime se ies. Al hough
Cadzow applies linea p edic ion o he signal, his me hod can also
be in e p e ed as pe o ming linea p edic ion in he au oco ela ion
domain. Bo h me hods ely, ei he explici ly o implici ly, on he ac
ha he au oco ela ion sequence is less a ec ed by b oadband noise
han he signal i sel , especially a high lag indices.
In his wo k, we conside he one-sided o causal pa o he
au oco ela ion sequence and i s ma hema ical p ope ies. As his
sequence sha es i s poles wi h he signal
x
(
n
)
, i p o ides a good
s a ing poin o LPC modeling. In his way, he new one-sided
au oco ela ion LPC (OSALPC) me hod appea s as a s aigh o wa d
esul o he app oach [5]. In addi ion, i is closely ela ed o he
SMC ep esen a ion and Cadzow’s me hod. All o hem can be
in e p e ed as AR modeling o ei he a spec al unc ion named
“en elope” o i s squa e. This in e p e a ion, which is based on he
p ope ies o he one-sided au oco ela ion, p o ides mo e insigh in o
he a ious me hods. In his co espondence, hei pe o mance in
noisy speech ecogni ion is compa ed. The op imum model o de
and ceps al li e ing ha e also been in es iga ed in noisy condi ions.
The simula ion esul s show ha OSALPC ou pe o ms he o he
echniques in se e e noisy condi ions and ob ains simila sco es o
mode a e o high SNR.
Manusc ip ecei ed Feb ua y 14, 1995; e ised No embe 9, 1995. This
wo k was suppo ed by G an nos. TIC-92-0800-C05/04 and TIC-92-1026-
C02/02. The associa e edi o coo dina ing he e iew o his pape and
app o ing i o publica ion was D . Kuldip K. Paliwal.
The au ho s a e wi h he Depa men o Signal Theo y and Communica ions,
Poly echnical Uni e si y o Ca alonia, Ba celona, Spain.
Publishe I em Iden i ie S 1063-6676(97)00766-9.
1063–6676/97$10.00 1997 IEEE
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 81
This co espondence is o ganized in he ollowing way. In Sec ion
II, he OSALPC echnique is in oduced, and i s ela ionship wi h he
con en ional LPC app oach and he o he pa ame e iza ions based
on AR modeling in he au oco ela ion domain is discussed. Sec ion
III epo s he applica ion o all hose pa ame e iza ion echniques
o an isola ed wo d mul ispeake ecogni ion ask, using he HMM
app oach, in o de o compa e hei pe o mances in he p esence o
addi i e whi e noise. Finally, some conclusions a e summa ized in
Sec ion IV.
II. AR MODELING IN THE AUTOCORRELATION DOMAIN
F om he au oco ela ion sequence
R
(
m
)
, we de ine he one-sided
(causal pa o he) au oco ela ion (OSA) sequence in he ollowing
way:
R
+
(
m
)=
R
(
m
)
m>
0
R
(0)
2
m
=0
0
m<
0
(1)
I s Fou ie ans o m is he complex “spec um”
S
+
(
!
)=
1
2
[
S
(
!
)+
jS
H
(
!
)]
(2)
whe e
S
(
!
)
is he eal spec um, i.e., he Fou ie ans o m o
R
(
m
)
,
and
S
H
(
!
)
is he Hilbe ans o m o
S
(
!
)
.
Due o he analogy be ween
S
+
(
!
)
in (2) and he analy ic signal
used in ampli ude modula ion, a spec al “en elope”
E
(
!
)
[6] can
be de ined as
E
(
!
)=
j
S
+
(
!
)
j
:
(3)
Due o he la ge dynamic ange o speech spec a, he en elope
E
(
!
)
s ongly enhances he highes powe equency bands wi h
espec o
S
(
!
)
[5]. Consequen ly, he noise componen s lying
ou side he enhanced equency bands a e la gely a enua ed in
E
(
!
)
wi h espec o
S
(
!
)
, and hus,
E
(
!
)
is mo e obus o b oadband
noise han
S
(
!
)
. On he o he hand, as i is well known, he OSA
sequence
R
+
(
m
)
and he signal
x
(
n
)
ha e he same poles [7].
Those wo p ope ies, i.e., obus ness o noise and pole p ese a-
ion, sugges ha AR pa ame e s o he speech signal can be mo e
eliably es ima ed om he OSA sequence
R
+
(
m
)
han di ec ly om
he signal
x
(
n
)
when
x
(
n
)
is co up ed by b oadband noise. Thus,
as he con en ional LPC echnique assumes an all-pole model o he
speech spec um
S
(
!
)
, we may apply linea p edic ion o he OSA
sequence, assuming an all-pole model o i s “spec um”
E
2
(
!
)
. This
is he basis o he one-sided au oco ela ion linea p edic i e coding
(OSALPC) pa ame e iza ion echnique [5].
A s aigh o wa d algo i hm is p oposed in [5] ha calcula es he
OSALPC ceps al coe icien s. I consis s o applying he (windowed)
au oco ela ion me hod o linea p edic ion o an es ima ion o he
OSA sequence:
a) Fi s , om he speech ame o leng h
N
, he au oco ela ion
lags un il
M
=
N=
2
a e compu ed ( his alue o
M
was
empi ically op imized o conside he well-known adeo
be ween a iance and equency esolu ion o he spec al
es ima e [8]).
b) Second, he Hamming window om
m
=0
o
M
is applied
on such es ima ed OSA sequence.
c) Thi d, i
p
is he p edic ion o de , he i s
p
+1
au oco ela ion
alues o ha OSA sequence a e compu ed om
m
=0
o
p
,
using he con en ional biased es ima o , i.e., he one ha is
commonly employed in speech p ocessing.
d) Then, hese alues a e used as en ies o he Le inson–Du bin
algo i hm o es ima e he AR pa ame e s
ak
,
k
=1
;
111
;p:
Fig. 1. Robus ness o he OSALPC ep esen a ion o addi i e whi e noise:
(a) LPC spec um and (b) OSALPC squa ed en elope o a oiced speech ame
in noise ee condi ions (solid line) and SNR equal o 0 dB (do ed line).
e) Finally, he ceps al coe icien s co esponding o he model a e
ecu en ly compu ed om hose AR pa ame e s.
The obus ness o OSALPC o addi i e whi e noise is illus a ed in
Fig. 1. As can be seen in his igu e, he OSALPC squa ed en elope
shows a p ominen i s o man , and i s whole cu e is mo e obus
o addi i e whi e noise han ha o he LPC spec um. In his case, he
con en ional biased au oco ela ion es ima o was used o compu e
he OSA sequence om he signal.
Fig. 1 also shows ha spu ious peaks may appea in he OSALPC
squa e en elope. They a e p obably due o he ac ha he OSALPC
echnique pe o ms only a pa ial decon olu ion o he speech signal
[9]. In spi e o ha , OSALPC shows a be e speech ecogni ion
pe o mance han con en ional LPC in se e e condi ions o addi i e
whi e noise, as will be seen in he nex sec ion.
The OSALPC echnique is closely ela ed o he sho - ime modi-
ied cohe ence (SMC) ep esen a ion p oposed by Mansou and Juang
in [3]. SMC is also based on AR modeling in he au oco ela ion
domain. Howe e , whe eas in he OSALPC echnique, he en ies o
he Le inson–Du bin algo i hm ( i s
p
alues o he au oco ela ion
o he OSA sequence) a e calcula ed om he OSA sequence using he
con en ional biased au oco ela ion es ima o , in he SMC ep esen-
a ion, hey a e compu ed using a squa e oo spec al shape . In ac ,
in e ms o he abo e o mula ion, ha di e ence lies in assuming
in he SMC echnique an all-pole spec al model o he en elope
E
(
!
)
ins ead o
E
2
(
!
)
. Fu he mo e,
R
+
(0)
is se o 0 in he case
o addi i e whi e noise because i is se e ely co up ed by noise.
On he o he hand, he name o he SMC ep esen a ion de i es
om he usage o a pa icula es ima o , which is e e ed o as
cohe ence in [3], o compu e he OSA sequence om he signal.
This es ima o is a mo e homogeneous measu e han he con en ional
biased au oco ela ion es ima o in he sense ha e e y es ima ed
alue is compu ed using he same numbe o signal samples, whe eas
in he con en ional es ima o , he numbe o signal samples employed
o es ima e
R
(
m
)
dec eases along he index
m
. Tha p ope y does
no ha e much ele ance in he es ima ion o he au oco ela ion
en ies o he Le inson–Du bin algo i hm since only he i s
p
+1
82 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997
Fig. 2. Ma ix o mula ion o OSALPC and LSMYE me hods.
alues a e conside ed, and usually,
p
N
. Howe e , i may be
impo an in he es ima ion o he OSA sequence om he speech
signal since he OSA leng h conside ed in bo h OSALPC and SMC
echniques is
M
=
N=
2
and no negligible wi h espec o
N
.
The OSALPC echnique can also be easily ela ed o he o e de-
e mined se o Yule–Walke equa ions p oposed by Cadzow in
[4] o seek ARMA models o ime se ies. As an
AR
(
p
)
p ocess
con amina ed by addi i e whi e noise becomes an ARMA
(
p; p
)
p ocess, Cadzow’s me hod can be used o es ima e he pa ame e s
o his noisy AR p ocess simply by se ing he same AR and MA
o de s in he so-called leas squa es modi ied Yule–Walke equa ions
(LSMYWE’s) [8].
The ela ionship be ween he OSALPC and LSMYWE echniques
is illus a ed by he ma ix equa ion in Fig. 2, whe e
M
deno es
he highes au oco ela ion lag used, and
e
(
m
)
is he e o o
be minimized. The minimiza ion o he no m o he ull e o
ec o
e
(
m
)
g
m
=1
;
111
;M
+
p
wi h espec o he AR pa ame e s
a
k
is equi alen o he applica ion o he (windowed) au oco ela ion
me hod o linea p edic ion o he sequence
R
(
m
)
,
m
=1
;
111
;M
,
i.e., he OSALPC echnique. On he o he hand, he LSMYWE
echnique minimizes he no m o he sub ec o
e
(
m
)
g
m
=
p
+1
;
111
;M
,
and he e o e, i amoun s o applying he (unwindowed) co a iance
me hod o linea p edic ion on he same ange o au oco ela ion lags.
When
M
=2
p
, LSMYWE a e he modi ied Yule–Walke equa ions
[8] o an ARMA
(
p; p
)
p ocess. In bo h cases, only au oco ela ion
lags co esponding o he OSA sequence a e employed.
In ou compa ison, we will also conside ano he e sion o
his (unwindowed) co a iance-based app oach ha will be called
leas squa es Yule–Walke equa ions (LSYWE’s). Whe eas in he
LSMYWE echnique he i s p edic ed au oco ela ion alue is
R
(
p
+
1)
, in he LSYWE echnique, he p edic ion begins a
R
(1)
. Bo h
LSMYWE and LSYWE me hods and hei ela ionship o OSALPC
a e g aphically desc ibed in Fig. 3. As i is shown, he only di e ence
be ween he a ious echniques is he ange o au oco ela ion lags
conside ed in he minimiza ion o he e o . I is wo h no ing ha
LSYWE conside s some nega i e au oco ela ion lags ha do no
belong o he OSA sequence. In pa icula , i
M
is equal o
p
, LSYWE
a e he con en ional Yule–Walke equa ions.
As will be seen in he nex sec ion, in spi e o he simila i y be ween
all hese echniques, he OSALPC ep esen a ion ou pe o ms he
LSYWE, LSMYWE, and SMC echniques in speech ecogni ion in
se e e noisy condi ions. On he o he hand, as a as he compu a ional
complexi y o he algo i hms is conce ned, OSALPC and SMC
echniques a e much mo e e icien han LSYWE and LSMYWE
echniques because hey use he Le inson–Du bin algo i hm.
Finally, i is wo h no ing ha he OSALPC echnique may be
included in he ield o highe o de spec al es ima ion due o he
ac ha he squa ed en elope
E
2
(
!
)
is he Fou ie ans o m o he
au oco ela ion o he OSA sequence, which is a pa icula ou h-
o de momen o he signal.
Fig. 3. In e p e a ion o he (a) OSALPC, (b) LSMYWE, and (c) LSYWE
app oaches as applica ion o he au oco ela ion o co a iance me hods o
linea p edic ion o an au oco ela ion sequence in di e en lag anges.
III. SPEECH RECOGNITION EXPERIMENTS
This sec ion epo s he applica ion o all he abo e pa ame e -
iza ion echniques o ecognize isola ed wo ds in a mul ispeake
ask wi h a disc e e HMM-based sys em in o de o compa e hei
pe o mance and o gain some insigh in o he me i o he OSALPC
ep esen a ion in he p esence o addi i e whi e noise
A. Speech Da abase and Recogni ion Sys em
The da abase used in ou expe imen s consis s o 10 epe i ions o
he Ca alan digi s u e ed by se en male and h ee emale speake s
(1000 wo ds) and eco ded in a quie oom. Fi s , he sys em was
ained wi h hal o he da abase and es ed wi h he o he hal . Then,
he oles o bo h hal es we e changed, and he epo ed esul s we e
ob ained by a e aging hose wo esul s.
The analog speech signal was i s bandpass il e ed o 100–3400
Hz by an an ialiasing il e , sampled a 8 kHz and, 12 bi s quan ized.
The digi ized clean speech was manually endpoin ed o de e mine
he bounda ies o each wo d. The endpoin s ob ained in his way
we e used in all ou expe imen s, including hose in which noise
was added o he signal. Clean speech was used o aining in all he
expe imen s. Noisy speech was simula ed by adding ze o mean whi e
Gaussian noise o he clean signal so ha he SNR o he esul ing
signal becomes
1
(clean), 20, 10, and 0 dB. No p eemphasis was
pe o med.
In he pa ame e iza ion s age o he ecogni ion sys em, he signal
was di ided in o ames o 30 ms a a a e o 15 ms, and each ame
was cha ac e ized by i s ceps al pa ame e s ob ained ei he by he
con en ional LPC me hod o by any o he echniques p esen ed in
he las sec ion. Be o e en e ing he ecogni ion s age, he ceps al
pa ame e s we e ec o quan ized using bo h a codebook o 64
codewo ds and he Euclidean dis ance measu e be ween li e ed
ceps al ec o s. Each digi was cha ac e ized by a le - o- igh
disc e e hidden Ma ko model o 10 s a es wi hou skips. T aining and
es ing we e pe o med using Baum–Welch and Vi e bi algo i hms,
espec i ely.
B. Recogni ion Resul s
Fi s o all, we ca ied ou some expe imen s wi h he abo e
desc ibed speech ecogni ion sys em o op imize he model o de and
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 83
TABLE I
RECOGNITION RATES OF THE CONVENTIONAL LPC TECHNIQUE FOR
SEVERAL PREDICTION ORDER VALUES AND CEPSTRAL LIFTERS
TABLE II
RECOGNITION RATES OF THE CONVENTIONAL LPC, LSMYWE,
AND LSYWE TECHNIQUES FOR
p
=12
AND THE SLOPE LIFTER
TABLE III
RECOGNITION RATES OF THE CONVENTIONAL LPC, SMC, AND
OSALPC TECHNIQUES FOR
p
=12
AND THE SLOPE LIFTER
he ype o ceps al li e in he con en ional LPC echnique. In Table
I, he ecogni ion esul s o LPC model o de s
p
=
8, 12, and 16
and o he bandpass [10], in e se o s anda d de ia ion [11] (ISD),
and slope [12] li e s a e p esen ed. The ecogni ion esul s show ha
nei he he model o de no he ype o ceps al li e a e ele an o
ou ask in noise- ee condi ions. Howe e , in he p esence o noise,
he ecogni ion esul s a e e y sensi i e o bo h ac o s.
I is also clea om Table I ha he nonsymme ical li e s—slope
and ISD—ou pe o m he bandpass li e o e e y model o de . This
may be due o he ac ha in he p esence o whi e noise, he lowe
o de ceps al coe icien s a e mo e a ec ed han he highe o de
ones in he unca ed ceps al ec o .
The bes esul s o se e e noisy condi ions—10 and 0 dB o
SNR—a e ob ained using slope li e and p edic ion o de
p
equal
o 12. The con enience o his ela i ely high o de comes om he
ac ha he sensi i i y o he au oco ela ion sequence o addi i e
whi e noise ends o dec ease along he lag index. Model o de s
ha a e oo high, howe e , yield poo ecogni ion esul s since he
spec al es ima e shows spu ious peaks. Ac ually, ecogni ion a es
we e calcula ed using he slope li e o a la ge ange o alues o
he model o de , and he bes esul s we e hose ob ained o
p
=
12.
In Table II, he ecogni ion a es o con en ional LPC, LSMYWE,
and LSYWE app oaches a e p esen ed, using
M
=
N=
2
and bo h
op imum model o de and li e ob ained o he con en ional LPC
echnique, i.e.,
p
=
12 and he slope li e . Ob iously, hese a e no
he op imum condi ions o each pa ame e iza ion echnique, bu he
esul s can help o compa e hei pe o mance. As can be seen om
Table II, he con en ional LPC echnique ou pe o ms no iceably he
o he app oaches. Howe e , he excellen pe o mance o he LSYWE
app oach in noise- ee condi ions is wo h no ing.
Fig. 4. Compa ison o ecogni ion a es o he LPC, SMC, OSALPC-I and
OSALPC-II echniques.
Fig. 5. Block diag am o he calcula ion o he LPC, SMC, OSALPC-I and
OSALPC-II ceps a.
In Table III and Fig. 4, he ecogni ion a es co esponding o
he con en ional LPC echnique, he SMC ep esen a ion, and he
no el OSALPC app oach a e p esen ed, whe e we also use
M
=
N=
2
,
p
=12
, and he slope li e . The wo e sions OSALPC-I
and OSALPC-II o he OSALPC app oach co espond o he OSA
es ima o s o which we e e ed in Sec ion II: OSALPC-I uses he
con en ional biased au oco ela ion es ima o , and OSALPC-II like
SMC uses he cohe ence es ima o (and se s
R
(0)
o 0). Fig. 5 shows
84 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997
TABLE IV
RECOGNITION RATES FOR THE OSALPC-II TECHNIQUE FOR
SEVERAL PREDICTION ORDER VALUES AND CEPSTRAL LIFTERS
a block diag am o he calcula ion o he LPC, SMC, OSALPC-I,
and OSALPC-II ceps a ha pe mi s compa ison o hei espec i e
algo i hms.
The OSALPC and SMC ep esen a ions clea ly ou do he con-
en ional LPC echnique in se e e noisy condi ions: OSALPC-I and
OSALPC-II a es a e be e han LPC ones a 10 and 0 dB, and SMC
ou pe o ms LPC a 0 dB. Mo eo e , OSALPC-I and OSALPC-II
ep esen a ions ou pe o m he SMC echnique in all noisy condi ions.
Fo he OSALPC ep esen a ion, he use o he con en ional biased
au oco ela ion es ima o o compu ing he OSA sequence ( e sion
OSALPC-I) is con enien in se e e noisy condi ions, i.e., o an SNR
o 10 o 0 dB.
Howe e , in noise- ee condi ions, he e is a loss o ecogni ion
pe o mance in he OSALPC and SMC app oaches wi h espec o
he con en ional LPC echnique due o he impe ec decon olu ion o
he speech signal pe o med by hose echniques. This e ec seems
o be minimized by using he cohe ence es ima o o compu e he
OSA sequence, as in he case o OSALPC-II and SMC.
Finally, Table IV shows he ecogni ion a es co esponding o
OSALPC-II o he same model o de s and ceps al li e s as in
Table I. I can be no iced ha he new echnique is less sensi i e
o changes in bo h he model o de and he ype o ceps al li e
han he con en ional LPC app oach, p o ided ha he model o de
is no oo low.
IV. CONCLUSIONS
In his co espondence, se e al LPC-based echniques ha wo k
in he au oco ela ion domain a e p esen ed and compa ed in noisy
speech ecogni ion. The OSALPC echnique, which is based on
he applica ion o he (windowed) au oco ela ion me hod o linea
p edic ion o he one-sided au oco ela ion sequence, yields he bes
esul s among all he compa ed LPC-based echniques in se e e noisy
condi ions.
REFERENCES
[1] F. I aku a, “Minimum p edic ion esidual p inciple applied o speech
ecogni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol.
ASSP-23, pp. 67–72, 1975.
[2] S. B. Da is and P. Me mels ein, “Compa ison o pa ame ic ep e-
sen a ions o monosyllabic wo d ecogni ion in con inuously spoken
sen ences,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-
28, pp. 357–366, 1980.
[3] D. Mansou and B. H. Juang, “The sho - ime modi ied cohe ence
ep esen a ion and i s applica ion o noisy speech ecogni ion,” IEEE
T ans. Acous ., Speech, Signal P ocessing, ol. 37, pp. 795–804, 1989.
[4] J. A. Cadzow, “Spec al es ima ion: An o e de e mined a ional model
equa ion app oach,” P oc. IEEE, ol. 70, pp. 907–939, 1982.
[5] J. He nando and C. Nadeu, “Speech ecogni ion in noisy ca en i on-
men based on OSALPC ep esen a ion and obus simila i y measu ing
echniques,” in P oc. ICASSP’94, Adelaide, Ap . 1994, pp. 69–72.
[6] M. A. Lagunas and M. Amengual, “Non-linea spec al es ima ion,” in
P oc. ICASSP’87, Dallas, Ap . 1987, pp. 2035–2038.
[7] D. P. McGinn and D. H. Johnson, “Reduc ion o all-pole pa ame e
es ima ion bias by successi e au oco ela ion,” in P oc. ICASSP’83,
Bos on, Ap . 1983, pp. 1088–1091.
[8] S. L. Ma ple, J ., Ed., Digi al Spec al Analysis wi h Applica ions.
Englewood Cli s, NJ: P en ice-Hall, 1987.
[9] C. Nadeu, J. Pascual, and J. He nando, “Pi ch de e mina ion using
he ceps um o he one-sided au oco ela ion sequence,” in P oc.
ICASSP’91, To on o, Canada, May 1991, pp. 3677–3680.
[10] B. H. Juang, L. R. Rabine , and J. G. Wilpon, “On he use o band-pass
li e ing in speech ecogni ion,” IEEE T ans. Acous ., Speech, Signal
P ocessing, ol. ASSP-35, pp. 947–954, 1987.
[11] Y. Tohku a, “A weigh ed ceps al dis ance measu e o speech ecog-
ni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-35,
pp. 1414–1422, 1987.
[12] B. A. Hanson and H. Waki a, “Spec al slope dis ance measu es wi h
linea p edic ion analysis o wo d ecogni ion in noise,” IEEE T ans.
Acous ., Speech, Signal P ocessing, ol. ASSP-35, pp. 968–973, 1987.
A Fas Algo i hm o Finding he Adap i e Componen
Weigh ed Ceps um o Speake Recogni ion
Mihailo S. Zilo ic, Ra i P. Ramachand an, and Richa d J. Mammone
Abs ac — In speake ecogni ion sys ems, he adap i e componen
weigh ed (ACW) ceps um has been shown o be mo e obus han he
con en ional linea p edic i e (LP) ceps um. The ACW ceps um is
de i ed om a pole-ze o ans e unc ion whose denomina o is he
p
h-o de LP polynomial
A
(
z
). The nume a o is a (
p
0
1
) h-o de
polynomial ha is up o now ound as ollows. The oo s o
A
(
z
)a e
compu ed, and he co esponding esidues ob ained by a pa ial ac ion
expansion o
1
=A
(
z
) a e se o uni y. The e o e, he nume a o is he
sum o all he (
p
0
1
) h-o de co ac o s o
A
(
z
). In his co espondence,
we show ha he nume a o polynomial is me ely he de i a i e o he
denomina o polynomial
A
(
z
). This g ea ly speeds up he compu a ion o
he nume a o polynomial coe icien s since i in ol es a simple scaling
o he denomina o polynomial coe icien s. Roo inding is comple ely
elimina ed. Since he denomina o is gua an eed o be minimum phase
and he nume a o can be p o en o be minimum phase, wo sepa a e
ecu sions in ol ing he polynomial coe icien s es ablishes he ACW cep-
s um. This new me hod, which a oids oo inding, educes he compu e
ime signi ican ly and imposes negligible o e head when compa ed wi h
he app oach o inding he LP ceps um.
I. INTRODUCTION
Speake ecogni ion is he ask o iden i ying a speake by his o he
oice [1]. A common p oblem in ealizing obus speake ecogni ion
sys ems is ha a misma ch in aining and es ing condi ions se iously
deg ades he pe o mance [2]. One o he pu sued app oaches o
Manusc ip ecei ed Feb ua y 3, 1995; e ised July 13, 1996. The associa e
edi o coo dina ing he e iew o his pape and app o ing i o publica ion
was D . Joseph Campbell.
M. S. Zilo ic is wi h Bell Communica ions Resea ch, Red Bank, NJ USA.
R. P. Ramachand an and R. J. Mammone a e wi h he CAIP Cen e ,
Depa men o Elec ical Enginee ing, Ru ge s Uni e si y, Pisca away, NJ
08855 USA.
Publishe I em Iden i ie S 1063-6676(97)00762-1.
1063–6676/97$10.00 1997 IEEE