Full text
ENEKO AÑORGA
DEVELOPMENT OF THE FEATURE
EXTRACTOR FOR SPEECH RECOGNITION
DIPLOMA WORK
MARIBOR, OCTOBER 2009
i
FAKULTETA ZA ELEKTROTEHNIKO,
RAČUNALNIŠTVO IN INFORMATIKO
2000 Ma ibo , Sme ano a ul. 17
Diploma Wo k o Elec onic Enginee ing S uden P og am
DEVELOPMENT OF THE FEATURE
EXTRACTOR FOR SPEECH RECOGNITION
S uden : Eneko A
ñ
o ga
S uden p og am: Elec onic Enginee ing
Men o : P o . D . Riko ŠAFARIČ
Asis . P o . D . Suzana URAN
Ma ibo , Oc obe 2009
ii
ACKNOWLEDGMENTS
Thanks o P o . D . Riko ŠAFARIČ o his
assis ance and help ul ad ices in ca ying ou he
diploma wo k.
Special hanks o my amily and iends who a e in
all momen s beside me.
iii
DEVELOPMENT OF THE FEATURE
EXTRACTOR FOR SPEECH RECOGNITION
Key wo ds: oice ope a ed wheelchai , speech ecogni ion, oice ac i i y de ec ion, neu al
ne wo ks, ul asound senso ne
UDK: 004.934:681.5(043.2)
Abs ac
Wi h his diploma wo k we
ha e a emp ed
o gi e con inui y o he p e ious wo k done by
o he esea che s called, Voice Ope a ing In elligen Wheelchai – VOIC [1]. A de elopmen o
a wheelchai con olled by oice is p esen ed in his wo k and is designed o physically disabled
people, who canno con ol hei mo emen s. This wo k desc ibes basic componen s o speech
ecogni ion and wheelchai con ol sys em.
Going o he g ain, a speech ecognize sys em is comp ised o wo dis inc blocks, a Fea u e
Ex ac o and a Recognize . The p esen wo k is a ge ed a he ealiza ion o an adequa e
Fea u e Ex ac o block which uses a s anda d LPC Ceps um code , which ansla es he
incoming speech in o a ajec o y in he LPC Ceps um ea u e space, ollowed by a Sel
O ganizing Map, which classi ies he ou come o he code in o de o p oduce op imal
ajec o y ep esen a ions o wo ds in educed dimension ea u e spaces. Expe imen al esul s
indica e ha ajec o ies on such educed dimension spaces can p o ide eliable ep esen a ions
o spoken wo ds. The Recognize block is le o u u e esea che s.
The main con ibu ions o his wo k ha e been he esea ch and app oach o a new
echnology o de elopmen issues and he ealiza ion o applica ions like a oice eco de and
playe and a comple e Fea u e Ex ac o sys em.
i
RAZVOJ PREVODNIKA SIGNALA ZA
PREPOZNAVO GOVORA
Ključne besede: glaso no oden in alidski oziček, p epozna a go o a, zazna a glaso ne
ak i nos i ne onske m eže, ul az očna senzo ska m eža
UDK: 004.934:681.5(043.2)
Po ze ek
S em diplomskim delom sem poskusil nadalje a i delo aziska e z naslo om Voice Ope a ing
In elligen Wheelchai – VOIC [1]. V em diplomskem delu je udi p eds a ljen az oj glaso no
odenega in alidskega ozička, na ejenega za elesno p izade e ljudi, ki ne mo ejo nadzo o a i
s ojih gibo . To delo opisuje osno ne komponen e go o nega nadzo a in sis ema odenja
in alidskega ozička.
Sis em go o nega nadzo a u a na a a d a azlična dela; p e odnik signala in
p epozna alec. V em diplomskem delu se os edo očam na p e odnos us eznega p e odnika
signala na osno i s anda dnega LPC Ceps um kode ja, ki pos eduje p ihajajoči go o po
LPC Ceps um p os o a, emu pos opku pa sledi .i. “samoo ganizacijska ka a” (Sel
O ganizing Map), ki az s i ezul a kode ja za op imalni p ikaz besed na zmanjšanih
dimenzijah p os o a. Poskusni ezul a i kažejo, da lahko e po i na zmanjšanih dimenzijah
p os o a zago o ijo zaneslji p ikaz izgo o jenih besed. P epozna alec je lahko p edme
azisko anja š uden o udi p ihodnos i.
Gla ni namen ega diplomskega dela s a bili aziska a in poskus upo abe no e ehnologije
az ojne namene, kako udi upo aba aplikacij ko so snemalec in p ed ajalnik z oka e celo en
sis em p e odnika signala.
Table o Con en s
1
INTRODUCTION ........................................................................................................ 1
1.1
MOTIVATION ..................................................................................................... 1
1.2
OBJECTIVES AND CONTRIBUTION OF THIS WORK.................................. 3
1.3
ORGANIZATION OF THIS WORK ................................................................... 5
1.4
RESOURCES ........................................................................................................ 5
2
DEVELOPMENT ......................................................................................................... 6
2.1
BRIEF DESCRIPTION OF COLIBRI MODULE ............................................... 6
2.1.1
Ha dwa e ......................................................................................................... 6
2.1.2
So wa e ........................................................................................................ 10
2.2
DESIGN OF THE FEATURE EXTRACTOR ................................................... 11
2.2.1
Speech coding ............................................................................................... 12
2.2.1.1
Speech sampling ..................................................................................... 12
2.2.1.2
P e-emphasis il e .................................................................................. 13
2.2.1.3
Wo d Isola ion ........................................................................................ 13
2.2.1.4
Speech coding ........................................................................................ 14
2.2.1.5
Summing up ........................................................................................... 18
2.2.2
Dimensionali y educ ion using SOM ........................................................... 19
2.2.2.1
Op ional Signal Scaling .......................................................................... 22
2.2.2.2
Summing up ........................................................................................... 22
2.3
SOFTWARE DEVELOPMENT ......................................................................... 24
2.3.1
Audio Reco de and Playe ........................................................................... 24
2.3.2
Fea u e Ex ac o ........................................................................................... 27
3
CONCLUSIONS ........................................................................................................ 33
3.1
SUMMARY OF RESULTS ................................................................................ 33
3.2
DIRECTIONS FOR FUTURE RESEARCH ...................................................... 34
4
REFERENCES ........................................................................................................... 36
i
GLOSARY OF SIMBOLS
Name Desc ip ion
1/Â(z)
a
i
w
lp
(n)
H
h1
(z)
s(n)
s'(n)
(k)
w
lag
(n)
'(k)
k
i
s
0
LP syn hesis il e
LP coe icien s (a
0
= 1.0)
LP analysis window
Inpu high-pass il e
P ep ocessed/ il e ed speech signal
Windowed speech signal
Au o-co ela ion coe icien s
Co ela ion lag window
Modi ied au o-co ela ion coe icien s
Re lec ion coe icien s
Sampling equency
Bandwid h expansion
Table 1 – Glossa y o symbols
ii
GLOSSARY OF ACRONYMS
Ac onym Desc ip ion
VOIC
DSP
LPC
LP
SOM
CE
SODIMM
CAN
GPIO
BSP
ITU
ITU-T
FE
VAD
DTX
CNG
RNN
VQ
HMM
Voice Ope a ed In elligen Wheelchai
Digi al Signal P ocesso
Linea P edic ion Coding
Linea P edic ion
Sel O ganizing Maps
Compac Edi ion
Small Ou line Dual In-line Memo y Module
Con ol A ea Ne wo k
Gene al Pu pose Inpu /Ou pu
Boa d Suppo Package
In e na ional Telecommunica ion Union
Telecommunica ion S anda diza ion Sec o
Fea u e Ex ac o
Voice Ac i i y De ec ion
Discon inuous T ansmission
Com o Noise Gene a o
Recu en Neu al Ne wo k
Vec o Quan iza ion
Hidden Ma ko Model
Table 2 – Glossa y o ac onyms
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 7
Figu e 2.1 – Colib i PXA320
Module Speci ica ions:
CPU
PXA320 806MHz
Memo y
128MB DDR RAM (32Bi )
1GB NAND Flash (8Bi )
In e aces
16Bi Ex e nal BUS
Compac Flash/PCMCIA
LCD (SVGA)
Touch Sc een
Audio I/O (16Bi S e eo)
CMOS image senso
I2C
SPI
2x SD Ca d
USB Hos /De ice
100MBi E he ne
2x UART
I DA
PWM
127 GPIOs
So wa e
P e-ins alled
Windows CE 5.0/6.0
Size
67.6 x 36.7 x 5.2 mm
Tempe a u e Range
0 o +70°C
-45 o +85°C (IT e sion)
8 Eneko Año ga, Diploma Wo k
In o de o ha e a lexible de elopmen en i onmen o explo e he unc ionali y and
pe o mance o he Colib i modules he Colib i E alua ion Boa d is used (See Figu e 2.2).
Besides he use in e aces i p o ides nume ous communica ion channels as well as a
con igu able jumpe a ea o hook up he Colib i GPIOs o he desi ed unc ion. To acili a e
in e acing o he cus om ha dwa e he Colib i E alua ion Boa d p o ides he bu e ed CPU bus
on a sepa a e connec o .
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 9
Figu e 2.2 – Colib i E alua ion Boa d
Module Speci ica ions:
CPU Modules
Colib i PXA270
Colib i PXA300
Colib i PXA310
Colib i PXA320
In e aces
10/100MBi E he ne
USB Hos /De ice
USB Hos
2x PS/2
Analogue VGA
Gene ic LCD Connec o
TFT: Philips LB064V02-A1
Line-In, Line-Ou , Mic-In
I DA
2x RS232
CAN (Philips SJA1000)
SD Ca d
Compac Flash
Powe Supply:
Requi ed Inpu :
7-24VDC, 3-50W
On-boa d Con e e :
3.3V, 5V max 5A
Size:
200 x 200 mm
10 Eneko Año ga, Diploma Wo k
The ecei ed in oice om To adex, o he Colib i XScale® PXA320, plus he Colib i
E alua ion Ca ie Boa d and plus he suppo hou s a e shown in he nex Figu e 2.3:
Figu e 2.3 – In oice om To adex
2.1.2 So wa e
The module is shipped wi h a p eins alled WinCE 5.0 image wi h WinCE Co e license. O he
OS like Embedded Linux a e a ailable om he hi d-pa y.
To adex p o ides a WinCE 5.0 image and a WinCE 6.0. All WinCE images con ain he
To adex Boa d Suppo Package (BSP) which is one o he mos ad anced BSPs a ailable on he
ma ke . Besides he s anda d Windows CE unc ionali y, i includes a la ge numbe o addi ional
d i e s as well as op imized e sions o s anda d d i e s o he mos common in e aces and is
easily cus omizable by egis y se ings o adap o speci ic ha dwa e.
The Mic oso ® eMbedded Visual C++ 4.0 ool is used as desk op de elopmen en i onmen
o c ea ing he applica ions and sys em componen s o Windows® CE .NET powe ed de ices.
In conclusion, all he so wa e p esen ed in his wo k was done using he Mic oso ®
eMbedded Visual C++ 4.0 de elopmen ool and he To adex BSP ool (Figu e 2.4).
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 11
Figu e 2.4 – Mic oso ® eMbedded Visual C++ 4.0 and Windows® CE
2.2 DESIGN OF THE FEATURE EXTRACTOR
As s a ed be o e, in a speech ecogni ion p oblem he FE block has o p ocess he incoming
in o ma ion, he speech signal, so ha i s ou pu eases he wo k o he classi ica ion s age. The
app oach used in his wo k designs he FE block and di ides i in o wo consecu i e sub-blocks:
he i s is based on speech coding echniques, and he second uses a SOM o u he
op imiza ion (da a dimensionali y educ ion). The di e en blocks and sub-blocks a e shown in
he nex Figu e 2.5:
12 Eneko Año ga, Diploma Wo k
Figu e 2.5 – FE schema ic
2.2.1 Speech coding
2.2.1.1 Speech sampling
The speech was eco ded and sampled using a ela i ely inexpensi e dynamic mic ophone
and a Colib i’s audio inpu in e ace. The incoming signal was sampled a 8.000 Hz wi h 16 bi s
o esolu ion.
I migh be a gued ha a highe sampling equency, o mo e sampling p ecision, is needed
in o de o highe ecogni ion accu acy. Howe e , i a no mal digi al phone, which samples
speech a 8.000 Hz wi h a 16 bi esolu ion, is able o p ese e mos o he in o ma ion ca ied by
he signal [6], i does no seem necessa y o inc ease he sampling a e beyond 8.000 Hz o he
sampling p ecision o some hing highe han 16 bi s. Ano he eason behind hese se ings is ha
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 13
comme cial speech ecognize s ypically use compa able pa ame e alues and achie e
imp essi e esul s.
2.2.1.2 P e-emphasis il e
A e sampling he inpu signal is con enien o il e i wi h a second o de high-pass il e
wi h cu o equency a 140 Hz. The il e se es as a p ecau ion agains undesi ed low-
equency componen s.
The esul ing il e is gi en by:
( )
21
21
1
9114024
.
0
9059465
.
1
1
46363718.092724705.046363718.0
−−
−−
+
−
+−
=
z
z
zz
zH
h
(1)
2.2.1.3 Wo d Isola ion
Despi e he ac ha he sampled signal had pauses be ween he u e ances, i was s ill needed
o de e mine he s a ing and ending poin s o he wo d u e ances in o de o know exac ly he
signal ha cha ac e ized each wo d. To accomplish his, we decided o use VAD (Voice Ac i i y
De ec ion) echnique used in speech p ocessing, ins ead o using he olling a e age and he
h eshold, de e mined by he s a and end o each wo d, used in he p e ious wo ks, wi h he aim
o achie ing mo e accu acy and e iciency.
Fo ha , we based ou wo k in he Annex B om he ITU’s Recommenda ion G.729 [5],
whe e a sou ce code in C language abou he VAD is e icien ly de eloped.
VAD is a me hod which di e en ia es speech om silence o noise signal o aid in speech
p ocessing and he Annex B p o ides a high le el desc ip ion o he Voice Ac i i y De ec ion
(VAD), Discon inuous T ansmission (DTX) and Com o Noise Gene a o (CNG) algo i hms.
These algo i hms a e used o educe he ansmission a e du ing silence pe iods o speech. They
14 Eneko Año ga, Diploma Wo k
a e designed and op imized o wo k in conjunc ion wi h [ITU-T V.70]. [ITU-T V.70] manda es
he use o speech coding me hods. The algo i hms a e adap ed o ope a e wi h bo h he ull
e sion o G.729 and Annex B.
Le ’s see a gene al desc ip ion o he VAD algo i hm:
The VAD algo i hm makes a oice ac i i y decision e e y 10 ms in acco dance wi h he
ame size o he p e-p ocessed ( il e ed) signal. A se o di e ence pa ame e s is ex ac ed and
used o an ini ial decision. The pa ame e s a e he ull-band ene gy, he low-band ene gy, he
ze o-c ossing a e and a spec al measu e. The long- e m a e ages o he pa ame e s du ing non-
ac i e oice segmen s ollow he changing na u e o he backg ound noise. A se o di e en ial
pa ame e s is ob ained a each ame. These a e a di e ence measu e be ween each pa ame e
and i s espec i e long- e m a e age. The ini ial oice ac i i y decision is ob ained using a
piecewise linea decision bounda y be ween each pai o di e en ial pa ame e s. A inal oice
ac i i y decision is ob ained by smoo hing he ini ial decision.
The ou pu o he VAD module is ei he 1 o 0, indica ing he p esence o absence o oice
ac i i y espec i ely. I he VAD ou pu is 1, he G.729 speech codec is in oked o code/decode
he ac i e oice ames. Howe e , i he VAD ou pu is 0, he DTX/CNG algo i hms desc ibed
he ein a e used o code/decode he non-ac i e oice ames.
2.2.1.4 Speech coding
A e he signal was sampled, he spec um was la ened, and he u e ances we e isola ed we
ied o codi y i using he Linea P edic ion Coding (LPC) me hod [3].
In a a ie y o applica ions, i is desi able o comp ess a speech signal o e icien
ansmission o s o age. Fo example, o accommoda e many speech signals in a gi en
bandwid h o a cellula phone sys em, each digi ized speech signal is comp essed be o e
ansmission. Fo medium o low bi - a e speech code s, LPC me hod is mos widely used.
Redundancy in a speech signal is emo ed by passing he signal h ough a speech analysis il e .
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 15
The ou pu o he il e , e med he esidual e o signal, has less edundancy han he o iginal
speech signal and can be quan ized by a smalle numbe o bi s han he o iginal speech.
The sho - e m analysis and syn hesis il e s a e based on 10 h o de linea p edic ion (LP)
il e s.
The LP syn hesis il e is de ined as:
∑
=
−
+
=
10
1
ˆ
1
1
)(
ˆ1
i
i
i
za
zA
(2)
whe e â
i
, i = 1,...,10, a e he quan ized Linea P edic ion (LP) coe icien s. Sho - e m p edic ion
o linea p edic ion analysis is pe o med once pe speech ame using he au oco ela ion
me hod wi h a 30 ms (240 samples) asymme ic window. E e y 10 ms (80 samples), he
au oco ela ion coe icien s o windowed speech a e compu ed and con e ed o he LP
coe icien s using he Le inson-Du bin algo i hm. Then he LP coe icien s a e ans o med o
he LSP domain o quan iza ion and in e pola ion pu poses. The in e pola ed quan ized and
unquan ized il e s a e con e ed back o he LP il e coe icien s ( o cons uc he syn hesis and
weigh ing il e s o each sub ame).
The LP analysis window consis s o wo pa s: he i s pa is hal a Hamming window and
he second pa is a qua e o a cosine unc ion cycle. The window is gi en by:
( ) ( )
239...,002
159
2002
cos
0,...,199
399
2
cos 46.054.0
=
−
=
−
=
n
n
n
n
nw
lp
π
π
(3)
The e is a 5 ms look-ahead in he LP analysis which means ha 40 samples a e needed om
he u u e speech ame. This ansla es in o an ex a algo i hmic delay o 5 ms a he encode
s age. The LP analysis window applies o 120 samples om pas speech ames, 80 samples om
16 Eneko Año ga, Diploma Wo k
he p esen speech ame, and 40 samples om he u u e ame. The windowing p ocedu e is
illus a ed in Figu e 2.4.
Figu e 2.6 – Windowing p ocedu e in LP analysis
The di e en shading pa e ns iden i y co esponding exci a ion and LP analysis windows.
The windowed speech:
(
)
(
)
(
)
0,...,239
=
=
′
nnsnwns
lp
(4)
is used o compu e he au oco ela ion coe icien s:
( ) ( ) ( )
0,...,10
239
=−=
∑
=
kkn'sn'sk
kn
(5)
To a oid a i hme ic p oblems o low-le el inpu signals he alue o (0) has a lowe
bounda y o (0) = 1.0. A 60 Hz bandwid h expansion is applied by mul iplying he
au oco ela ion coe icien s wi h:
( )
1,...,10
2
2
1
2
0
=
π
−= k
k
expkw
s
lag
(6)
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 23
1. The incoming p essu e wa e is educed o a digi al signal h ough a sampling p ocess.
2. The s a ing and he ending poin s o he u e ance embedded in o he signal a e ob ained
using he VAD p ocess.
3. The spec um o he ex ac ed u e ance is enhanced by means o a p e-emphasis il e
which boos s he high equency componen s.
4. Se e al da a blocks a e ex ac ed om he enhanced signal.
5. The ex ac ed da a blocks a e windowed o educe leakage e ec s.
6. LPC componen s a e ex ac ed om he il e ed blocks.
7. LPC Ceps um componen s a e hen ex ac ed om he LPC ec o s.
8. The dimensionali y o he LPC Ceps um ec o s is educed using a SOM.
9. The esul ing ec o s a e scaled i he Recognize equi es i .
No hing can s ill be said abou he o e all e ec i eness o he FE block, since i depends on
he ecogni ion accu acy o he o e all sys em. As an example, i he comple e sys em achie es
low ecogni ion pe cen ages, ha can be caused by he FE block o he Recognize , bu , i i
achie es highe pe cen ages ha means ha he FE block was a leas able o p oduce da a ha
allowed hese ecogni ion accu acies. In o he wo ds, in o de o know he use ulness o he FE
block, he Recognize ou pu s mus be ob ained i s .
24 Eneko Año ga, Diploma Wo k
2.3 SOFTWARE DEVELOPMENT
2.3.1 Audio Reco de and Playe
As we ha e men ioned be o e, we buil a comple e audio Reco de and Playe . This ask was
no s ic ly necessa y bu he aim o his has been o lea n, p ac ice and imp o e he C ++
p og amming skills. The g aphic in e ace o he p og am is shown in he ollowing Figu e 2.8:
Figu e 2.8 – audioce Reco de & Playe
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 25
As we can see in he Figu e 2.8, wi h his so wa e we a e able o eco d he oice du ing
some ime, wi h di e en sample a es, di e en esolu ions and inally we can sa e i in a .wa
ile o nex p ocessing s eps like he speech coding. In addi ion, we will be able o play and
lis en o he eco ded signals. Along wi h he documen a ion (Annex 2) is included an elec onic
a achmen con aining he sou ce code in C++ used o build he so wa e and which is coming
wi h all he necessa y explana ions.
As we can see in he nex Figu e 2.9 we can see he di e ence be ween he di e en ways o
sampling. In gene al, he memo y occupied by he sound ile is p opo ional o he numbe o
samples pe second and he esolu ion o each sample. Fo he i s case he speech is sampled a
8.0 KHz and wi h 8 bi s o esolu ion (1 by e/sample), his means ha he memo y occupied o
he sound ile will be 8.000 (samples/sec) x 5 (sec) = 40.000 (samples) = 40.000 (samples) x 1
(by e/sample) = 40 Kby es.
Fo he second case he speech is sampled a 44.1 KHz and wi h 16 bi s o esolu ion (2
by es/sample), he memo y occupied o he sound ile will be 44.100 (samples/sec) x 5 (sec) =
220.500 (samples) = 220.500 (samples) x 2 (by es/sample) = 440.1 Kby es.
26 Eneko Año ga, Diploma Wo k
Figu e 2.9 – Di e en samplings o he same speech signal
I migh be a gued ha he highe he sampling equency and he highe he sampling
p ecision, he be e he ecogni ion accu acy. Howe e , i a no mal digi al phone, which samples
speech a 8.000 Hz wi h a 16 bi p ecision, is able o p ese e mos o he in o ma ion ca ied by
he signal [6], i does no seem necessa y o inc ease he sampling a e beyond 8.000 Hz o he
sampling p ecision o some hing highe han 16 bi s. Ano he eason behind hese se ings is ha
comme cial speech ecognize s ypically use compa able pa ame e alues and achie e
imp essi e esul s. So o he speech code we decided o use 8.000 Hz o sample a e and 16 bi s
o esolu ion as con igu a ion o eco d he oice.
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 27
2.3.2 Fea u e Ex ac o
A e buil he audioce Reco de and Playe we added some new unc ions o he p og am as
he pa which ep esen s he FE block. Fo ha , we based ou wo k on he ITU-T’s G.729
Recommenda ion. This Recommenda ion con ains he desc ip ion o an algo i hm o he coding
o speech signals using Linea P edic ion Coding and in which mo e p ocesses like he il e ing
o he sampled signal, he di ision in o blocks and he windowing o il e ed signal and he wo d
isola ion a e implici . The g aphic in e ace o he p og am is shown in he ollowing Figu e
2.10.
Figu e 2.10 – audioce Reco de & Playe & VAD De ec o & Speech Code
28 Eneko Año ga, Diploma Wo k
Wi h his so wa e we a e able o eco d and play oice signals and sa e hem in .wa iles.
In addi ion, we will be able o de ec and isola e he oice command om he sound ile, c ea e a
new ile wi h i and inally codi y o educe i s dimensionali y. Along wi h he documen a ion
(Annex 2) is included an elec onic a achmen con aining he sou ce code in C++ used o build
he so wa e and which is coming wi h all he necessa y explana ions.
When we p ess he bu on “VAD De ec o ...” we ha e o selec a sound ile which con ains
he eco ded oice. Then he sampled signal will go h ough all hese s eps:
(Sampling: he oice signal is sampled a 8.000 He z wi h 16 bi s o p ecision and sa ed
in a new ile called “le .wa ”.)
1. P e-emphasis il e : he sampled signal is il e ed by a second o de high-pass il e and
sa ed in a new ile called “ il e ed.wa ”.
2. Wo d Isola ion: he il e ed signal is passed h ough he VAD block o isola e he wo d
and sa ed in a ile called “ ad.wa ”.
We can ep esen he esul s o he di e en s ages o he signal using Ma lab and he ea lie
c ea ed sound iles (Figu e 2.11).
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 29
Figu e 2.11 – Di e en s ages o he signal in VAD De ec o p ocess
As we can see, he i s g aphic shows he o iginal signal (Command “ igh ”), sampled a 8.0
KHz, wi h 16 bi s o esolu ion and in mono o 1 channel. We can ealize ha he signal has a
DC o se and also ha he end o he eco ding, is a bi noisy. The second g aphic shows he
il e ed sampled signal and inally he hi d g aphic shows he isola ed wo d, a e he VAD
p ocess.
A his poin we ha e o explain ha he esul shown o he VAD p ocess was ob ained
using an algo i hm myVAD.m [10] ob aining p e y good esul s. The p oblem wi h he algo i hm
de eloped o Colib i module is ha i does no a oid he silence and he non-speech pa s om
he oice e y well as he myVAD.m algo i hm does, as is shown in he nex Figu e 2.12.
30 Eneko Año ga, Diploma Wo k
Figu e 2.12 – Di e en ways o VAD p ocess
In he abo e Figu e 2.12, he second g aphic shows he eco ded signal a e VAD p ocess
using he algo i hm de eloped o Colib i module. He e, we can see ha e en hough he sound
ile is educed in some silen pa s o he signal a e a oided; hese p ocess is no doing he
co ec wo k as he hi d g aphic does. We ied o ind he solu ion o his bu inally we did no ,
so his pa need o be imp o ed and is le o u u e esea che s as we explain in he sec ion 3.2
Di ec ions o u u e esea ch.
To con inue ou esea ch we decided o use he ile c ea ed ough Ma lab and myVAD.m
algo i hm called “ ad2.wa ”. Once we ge he isola ed wo d (“ ad2.wa ”) we can s a codi ying
he speech clicking in “Speech Code ...” bu on. The nex s eps a e he ones which he isola ed
wo d signal will ollow:
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 31
1. Blocking: he isola ed wo d is di ided in o a sequence o da a blocks o ixed leng h,
called ames and mul iplied by Hamming window o same wid h.
2. LPC analysis: o each ame, 10 LPC coe icien s a e calcula ed.
3. Ceps um analysis: he 10 LPC coe icien s a e con e ed in 10 Ceps al coe icien ones
(Figu e 2.13).
Figu e 2.13 – The 10 Ceps al coe icien s o each ame
32 Eneko Año ga, Diploma Wo k
Once we ge he 10 Ceps al coe icien s o each ame i we click in “SOM...” bu on and
selec he ile wi h he Ceps al coe icien s (“c_coe . x ”), we will ed he inpu laye o he
Kohonen SOM, wi h his inpu da a one by one. The ou pu laye will o ganize i sel o ep esen
he inpu s in wo-dimensional space.
The aining p ocedu e in ol es he ollowing s eps:
1. The neu ons a e a anged in an n-dimensional la ice. Each neu on s o es a poin in an m-
dimensional space.
2. An inpu ec o is p esen ed o he SOM. The neu ons s a o compe e un il he one ha
s o es he closes poin o he inpu ec o p e ails. Once he dynamics o he ne wo k
con e ge, all he neu ons bu he p e ailing one will be inac i e. The ou pu o he SOM
is de ined as he co-o dina es o he p e ailing neu on in he la ice.
3. A neighbou hood unc ion is cen ed on he p e ailing neu on o he la ice. The alue o
his unc ion is one a he posi ion o he ac i e neu on, and dec eases wi h he dis ance
measu ed om he posi ion o he winning neu on.
4. The poin s s o ed by all he neu ons a e mo ed owa ds he inpu ec o in an amoun
p opo ional o he neighbou hood unc ion e alua ed in he posi ion o he la ice whe e
he neu on being modi ied s ands.
5. Re u n o 2, and epea s eps 2, 3, and 4 un il he a e age e o be ween he inpu ec o s
and he winning neu ons educes o a small alue.
A e he SOM is ained, he co-o dina es o he ac i e neu on in he la ice a e used as i s
ou pu s.
Along wi h he documen a ion (Annex 2) is included an elec onic a achmen con aining he
sou ce code in C++ used o build he so wa e and which is coming wi h all he necessa y
explana ions. In his case pa o he code o SOM has been de eloped bu s ill needs o be
imp o ed and inished, so his pa is le o u u e esea che s as we explain in he sec ion 3.2
Di ec ions o u u e esea ch.
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 39
ANNEX 2:
40 Eneko Año ga, Diploma Wo k
DECLARATION:
I, Eneko Año ga, he unde signed, decla e ha I ha e made he diploma wo k by mysel . I
am awa e o he po en ial consequences in he e en o a b each o his decla a ion.