scieee Science in your language
[en] (orig)

Development of the Feature Extractor for Speech Recognition

Abstract

With this diploma work we have attempted to give continuity to the previous work done by other researchers called, Voice Operating Intelligent Wheelchair – VOIC [1]. A development of a wheelchair controlled by voice is presented in this work and is designed for physically disabled people, who cannot control their movements. This work describes basic components of speech recognition and wheelchair control system. Going to the grain, a speech recognizer system is comprised of two distinct blocks, a Feature Extractor and a Recognizer. The present work is targeted at the realization of an adequate Feature Extractor block which uses a standard LPC Cepstrum coder, which translates the incoming speech into a trajectory in the LPC Cepstrum feature space, followed by a Self Organizing Map, which classifies the outcome of the coder in order to produce optimal trajectory representations of words in reduced dimension feature spaces. Experimental results indicate that trajectories on such reduced dimension spaces can provide reliable representations of spoken words. The Recognizer block is left for future researchers. The main contributions of this work have been the research and approach of a new technology for development issues and the realization of applications like a voice recorder and player and a complete Feature Extractor system.

Read accessible full text

Development of the Feature Extractor for Speech Recognition

Author: Añorga Irigoien, Eneko
Publisher: Universitat Politècnica de Catalunya
Year: 2009
Source: https://upcommons.upc.edu/bitstream/2099.1/8751/1/Diploma%20Work.pdf
ENEKO AÑORGA
DEVELOPMENT OF THE FEATURE
EXTRACTOR FOR SPEECH RECOGNITION
DIPLOMA WORK
MARIBOR, OCTOBER 2009
i
FAKULTETA ZA ELEKTROTEHNIKO,
RAČUNALNIŠTVO IN INFORMATIKO
2000 Ma ibo , Sme ano a ul. 17
Diploma Wo k o Elec onic Enginee ing S uden P og am
DEVELOPMENT OF THE FEATURE
EXTRACTOR FOR SPEECH RECOGNITION
S uden : Eneko A
ñ
o ga
S uden p og am: Elec onic Enginee ing
Men o : P o . D . Riko ŠAFARIČ
Asis . P o . D . Suzana URAN
Ma ibo , Oc obe 2009
ii
ACKNOWLEDGMENTS
Thanks o P o . D . Riko ŠAFARIČ o his
assis ance and help ul ad ices in ca ying ou he
diploma wo k.
Special hanks o my amily and iends who a e in
all momen s beside me.
iii
DEVELOPMENT OF THE FEATURE
EXTRACTOR FOR SPEECH RECOGNITION
Key wo ds: oice ope a ed wheelchai , speech ecogni ion, oice ac i i y de ec ion, neu al
ne wo ks, ul asound senso ne
UDK: 004.934:681.5(043.2)
Abs ac
Wi h his diploma wo k we
ha e a emp ed
o gi e con inui y o he p e ious wo k done by
o he esea che s called, Voice Ope a ing In elligen Wheelchai – VOIC [1]. A de elopmen o
a wheelchai con olled by oice is p esen ed in his wo k and is designed o physically disabled
people, who canno con ol hei mo emen s. This wo k desc ibes basic componen s o speech
ecogni ion and wheelchai con ol sys em.
Going o he g ain, a speech ecognize sys em is comp ised o wo dis inc blocks, a Fea u e
Ex ac o and a Recognize . The p esen wo k is a ge ed a he ealiza ion o an adequa e
Fea u e Ex ac o block which uses a s anda d LPC Ceps um code , which ansla es he
incoming speech in o a ajec o y in he LPC Ceps um ea u e space, ollowed by a Sel
O ganizing Map, which classi ies he ou come o he code in o de o p oduce op imal
ajec o y ep esen a ions o wo ds in educed dimension ea u e spaces. Expe imen al esul s
indica e ha ajec o ies on such educed dimension spaces can p o ide eliable ep esen a ions
o spoken wo ds. The Recognize block is le o u u e esea che s.
The main con ibu ions o his wo k ha e been he esea ch and app oach o a new
echnology o de elopmen issues and he ealiza ion o applica ions like a oice eco de and
playe and a comple e Fea u e Ex ac o sys em.
i
RAZVOJ PREVODNIKA SIGNALA ZA
PREPOZNAVO GOVORA
Ključne besede: glaso no oden in alidski oziček, p epozna a go o a, zazna a glaso ne
ak i nos i ne onske m eže, ul az očna senzo ska m eža
UDK: 004.934:681.5(043.2)
Po ze ek
S em diplomskim delom sem poskusil nadalje a i delo aziska e z naslo om Voice Ope a ing
In elligen Wheelchai – VOIC [1]. V em diplomskem delu je udi p eds a ljen az oj glaso no
odenega in alidskega ozička, na ejenega za elesno p izade e ljudi, ki ne mo ejo nadzo o a i
s ojih gibo . To delo opisuje osno ne komponen e go o nega nadzo a in sis ema odenja
in alidskega ozička.
Sis em go o nega nadzo a u a na a a d a azlična dela; p e odnik signala in
p epozna alec. V em diplomskem delu se os edo očam na p e odnos us eznega p e odnika
signala na osno i s anda dnega LPC Ceps um kode ja, ki pos eduje p ihajajoči go o po
LPC Ceps um p os o a, emu pos opku pa sledi .i. “samoo ganizacijska ka a” (Sel
O ganizing Map), ki az s i ezul a kode ja za op imalni p ikaz besed na zmanjšanih
dimenzijah p os o a. Poskusni ezul a i kažejo, da lahko e po i na zmanjšanih dimenzijah
p os o a zago o ijo zaneslji p ikaz izgo o jenih besed. P epozna alec je lahko p edme
azisko anja š uden o udi p ihodnos i.
Gla ni namen ega diplomskega dela s a bili aziska a in poskus upo abe no e ehnologije
az ojne namene, kako udi upo aba aplikacij ko so snemalec in p ed ajalnik z oka e celo en
sis em p e odnika signala.

Table o Con en s
1
INTRODUCTION ........................................................................................................ 1
1.1
MOTIVATION ..................................................................................................... 1
1.2
OBJECTIVES AND CONTRIBUTION OF THIS WORK.................................. 3
1.3
ORGANIZATION OF THIS WORK ................................................................... 5
1.4
RESOURCES ........................................................................................................ 5
2
DEVELOPMENT ......................................................................................................... 6
2.1
BRIEF DESCRIPTION OF COLIBRI MODULE ............................................... 6
2.1.1
Ha dwa e ......................................................................................................... 6
2.1.2
So wa e ........................................................................................................ 10
2.2
DESIGN OF THE FEATURE EXTRACTOR ................................................... 11
2.2.1
Speech coding ............................................................................................... 12
2.2.1.1
Speech sampling ..................................................................................... 12
2.2.1.2
P e-emphasis il e .................................................................................. 13
2.2.1.3
Wo d Isola ion ........................................................................................ 13
2.2.1.4
Speech coding ........................................................................................ 14
2.2.1.5
Summing up ........................................................................................... 18
2.2.2
Dimensionali y educ ion using SOM ........................................................... 19
2.2.2.1
Op ional Signal Scaling .......................................................................... 22
2.2.2.2
Summing up ........................................................................................... 22
2.3
SOFTWARE DEVELOPMENT ......................................................................... 24
2.3.1
Audio Reco de and Playe ........................................................................... 24
2.3.2
Fea u e Ex ac o ........................................................................................... 27
3
CONCLUSIONS ........................................................................................................ 33
3.1
SUMMARY OF RESULTS ................................................................................ 33
3.2
DIRECTIONS FOR FUTURE RESEARCH ...................................................... 34
4
REFERENCES ........................................................................................................... 36
i
GLOSARY OF SIMBOLS
Name Desc ip ion
1/Â(z)
a
i
w
lp
(n)
H
h1
(z)
s(n)
s'(n)
(k)
w
lag
(n)
'(k)
k
i
s
0
LP syn hesis il e
LP coe icien s (a
0
= 1.0)
LP analysis window
Inpu high-pass il e
P ep ocessed/ il e ed speech signal
Windowed speech signal
Au o-co ela ion coe icien s
Co ela ion lag window
Modi ied au o-co ela ion coe icien s
Re lec ion coe icien s
Sampling equency
Bandwid h expansion
Table 1 – Glossa y o symbols
ii
GLOSSARY OF ACRONYMS
Ac onym Desc ip ion
VOIC
DSP
LPC
LP
SOM
CE
SODIMM
CAN
GPIO
BSP
ITU
ITU-T
FE
VAD
DTX
CNG
RNN
VQ
HMM
Voice Ope a ed In elligen Wheelchai
Digi al Signal P ocesso
Linea P edic ion Coding
Linea P edic ion
Sel O ganizing Maps
Compac Edi ion
Small Ou line Dual In-line Memo y Module
Con ol A ea Ne wo k
Gene al Pu pose Inpu /Ou pu
Boa d Suppo Package
In e na ional Telecommunica ion Union
Telecommunica ion S anda diza ion Sec o
Fea u e Ex ac o
Voice Ac i i y De ec ion
Discon inuous T ansmission
Com o Noise Gene a o
Recu en Neu al Ne wo k
Vec o Quan iza ion
Hidden Ma ko Model
Table 2 – Glossa y o ac onyms
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 7
Figu e 2.1 – Colib i PXA320
Module Speci ica ions:
CPU
PXA320 806MHz
Memo y
128MB DDR RAM (32Bi )
1GB NAND Flash (8Bi )
In e aces
16Bi Ex e nal BUS
Compac Flash/PCMCIA
LCD (SVGA)
Touch Sc een
Audio I/O (16Bi S e eo)
CMOS image senso
I2C
SPI
2x SD Ca d
USB Hos /De ice
100MBi E he ne
2x UART
I DA
PWM
127 GPIOs
So wa e
P e-ins alled
Windows CE 5.0/6.0
Size
67.6 x 36.7 x 5.2 mm
Tempe a u e Range
0 o +70°C
-45 o +85°C (IT e sion)

8 Eneko Año ga, Diploma Wo k
In o de o ha e a lexible de elopmen en i onmen o explo e he unc ionali y and
pe o mance o he Colib i modules he Colib i E alua ion Boa d is used (See Figu e 2.2).
Besides he use in e aces i p o ides nume ous communica ion channels as well as a
con igu able jumpe a ea o hook up he Colib i GPIOs o he desi ed unc ion. To acili a e
in e acing o he cus om ha dwa e he Colib i E alua ion Boa d p o ides he bu e ed CPU bus
on a sepa a e connec o .
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 9
Figu e 2.2 – Colib i E alua ion Boa d
Module Speci ica ions:
CPU Modules
Colib i PXA270
Colib i PXA300
Colib i PXA310
Colib i PXA320
In e aces
10/100MBi E he ne
USB Hos /De ice
USB Hos
2x PS/2
Analogue VGA
Gene ic LCD Connec o
TFT: Philips LB064V02-A1
Line-In, Line-Ou , Mic-In
I DA
2x RS232
CAN (Philips SJA1000)
SD Ca d
Compac Flash
Powe Supply:
Requi ed Inpu :
7-24VDC, 3-50W
On-boa d Con e e :
3.3V, 5V max 5A
Size:
200 x 200 mm
10 Eneko Año ga, Diploma Wo k
The ecei ed in oice om To adex, o he Colib i XScale® PXA320, plus he Colib i
E alua ion Ca ie Boa d and plus he suppo hou s a e shown in he nex Figu e 2.3:
Figu e 2.3 – In oice om To adex
2.1.2 So wa e
The module is shipped wi h a p eins alled WinCE 5.0 image wi h WinCE Co e license. O he
OS like Embedded Linux a e a ailable om he hi d-pa y.
To adex p o ides a WinCE 5.0 image and a WinCE 6.0. All WinCE images con ain he
To adex Boa d Suppo Package (BSP) which is one o he mos ad anced BSPs a ailable on he
ma ke . Besides he s anda d Windows CE unc ionali y, i includes a la ge numbe o addi ional
d i e s as well as op imized e sions o s anda d d i e s o he mos common in e aces and is
easily cus omizable by egis y se ings o adap o speci ic ha dwa e.
The Mic oso ® eMbedded Visual C++ 4.0 ool is used as desk op de elopmen en i onmen
o c ea ing he applica ions and sys em componen s o Windows® CE .NET powe ed de ices.
In conclusion, all he so wa e p esen ed in his wo k was done using he Mic oso ®
eMbedded Visual C++ 4.0 de elopmen ool and he To adex BSP ool (Figu e 2.4).
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 11
Figu e 2.4 – Mic oso ® eMbedded Visual C++ 4.0 and Windows® CE
2.2 DESIGN OF THE FEATURE EXTRACTOR
As s a ed be o e, in a speech ecogni ion p oblem he FE block has o p ocess he incoming
in o ma ion, he speech signal, so ha i s ou pu eases he wo k o he classi ica ion s age. The
app oach used in his wo k designs he FE block and di ides i in o wo consecu i e sub-blocks:
he i s is based on speech coding echniques, and he second uses a SOM o u he
op imiza ion (da a dimensionali y educ ion). The di e en blocks and sub-blocks a e shown in
he nex Figu e 2.5:
12 Eneko Año ga, Diploma Wo k
Figu e 2.5 – FE schema ic
2.2.1 Speech coding
2.2.1.1 Speech sampling
The speech was eco ded and sampled using a ela i ely inexpensi e dynamic mic ophone
and a Colib i’s audio inpu in e ace. The incoming signal was sampled a 8.000 Hz wi h 16 bi s
o esolu ion.
I migh be a gued ha a highe sampling equency, o mo e sampling p ecision, is needed
in o de o highe ecogni ion accu acy. Howe e , i a no mal digi al phone, which samples
speech a 8.000 Hz wi h a 16 bi esolu ion, is able o p ese e mos o he in o ma ion ca ied by
he signal [6], i does no seem necessa y o inc ease he sampling a e beyond 8.000 Hz o he
sampling p ecision o some hing highe han 16 bi s. Ano he eason behind hese se ings is ha

De elopmen o he Fea u e Ex ac o o Speech Recogni ion 13
comme cial speech ecognize s ypically use compa able pa ame e alues and achie e
imp essi e esul s.
2.2.1.2 P e-emphasis il e
A e sampling he inpu signal is con enien o il e i wi h a second o de high-pass il e
wi h cu o equency a 140 Hz. The il e se es as a p ecau ion agains undesi ed low-
equency componen s.
The esul ing il e is gi en by:
( )
21
21
1
9114024
.
0
9059465
.
1
1
46363718.092724705.046363718.0
−−
−−
+
−
+−
=
z
z
zz
zH
h
(1)
2.2.1.3 Wo d Isola ion
Despi e he ac ha he sampled signal had pauses be ween he u e ances, i was s ill needed
o de e mine he s a ing and ending poin s o he wo d u e ances in o de o know exac ly he
signal ha cha ac e ized each wo d. To accomplish his, we decided o use VAD (Voice Ac i i y
De ec ion) echnique used in speech p ocessing, ins ead o using he olling a e age and he
h eshold, de e mined by he s a and end o each wo d, used in he p e ious wo ks, wi h he aim
o achie ing mo e accu acy and e iciency.
Fo ha , we based ou wo k in he Annex B om he ITU’s Recommenda ion G.729 [5],
whe e a sou ce code in C language abou he VAD is e icien ly de eloped.
VAD is a me hod which di e en ia es speech om silence o noise signal o aid in speech
p ocessing and he Annex B p o ides a high le el desc ip ion o he Voice Ac i i y De ec ion
(VAD), Discon inuous T ansmission (DTX) and Com o Noise Gene a o (CNG) algo i hms.
These algo i hms a e used o educe he ansmission a e du ing silence pe iods o speech. They
14 Eneko Año ga, Diploma Wo k
a e designed and op imized o wo k in conjunc ion wi h [ITU-T V.70]. [ITU-T V.70] manda es
he use o speech coding me hods. The algo i hms a e adap ed o ope a e wi h bo h he ull
e sion o G.729 and Annex B.
Le ’s see a gene al desc ip ion o he VAD algo i hm:
The VAD algo i hm makes a oice ac i i y decision e e y 10 ms in acco dance wi h he
ame size o he p e-p ocessed ( il e ed) signal. A se o di e ence pa ame e s is ex ac ed and
used o an ini ial decision. The pa ame e s a e he ull-band ene gy, he low-band ene gy, he
ze o-c ossing a e and a spec al measu e. The long- e m a e ages o he pa ame e s du ing non-
ac i e oice segmen s ollow he changing na u e o he backg ound noise. A se o di e en ial
pa ame e s is ob ained a each ame. These a e a di e ence measu e be ween each pa ame e
and i s espec i e long- e m a e age. The ini ial oice ac i i y decision is ob ained using a
piecewise linea decision bounda y be ween each pai o di e en ial pa ame e s. A inal oice
ac i i y decision is ob ained by smoo hing he ini ial decision.
The ou pu o he VAD module is ei he 1 o 0, indica ing he p esence o absence o oice
ac i i y espec i ely. I he VAD ou pu is 1, he G.729 speech codec is in oked o code/decode
he ac i e oice ames. Howe e , i he VAD ou pu is 0, he DTX/CNG algo i hms desc ibed
he ein a e used o code/decode he non-ac i e oice ames.
2.2.1.4 Speech coding
A e he signal was sampled, he spec um was la ened, and he u e ances we e isola ed we
ied o codi y i using he Linea P edic ion Coding (LPC) me hod [3].
In a a ie y o applica ions, i is desi able o comp ess a speech signal o e icien
ansmission o s o age. Fo example, o accommoda e many speech signals in a gi en
bandwid h o a cellula phone sys em, each digi ized speech signal is comp essed be o e
ansmission. Fo medium o low bi - a e speech code s, LPC me hod is mos widely used.
Redundancy in a speech signal is emo ed by passing he signal h ough a speech analysis il e .
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 15
The ou pu o he il e , e med he esidual e o signal, has less edundancy han he o iginal
speech signal and can be quan ized by a smalle numbe o bi s han he o iginal speech.
The sho - e m analysis and syn hesis il e s a e based on 10 h o de linea p edic ion (LP)
il e s.
The LP syn hesis il e is de ined as:
∑
=
−
+
=
10
1
ˆ
1
1
)(
ˆ1
i
i
i
za
zA
(2)
whe e â
i
, i = 1,...,10, a e he quan ized Linea P edic ion (LP) coe icien s. Sho - e m p edic ion
o linea p edic ion analysis is pe o med once pe speech ame using he au oco ela ion
me hod wi h a 30 ms (240 samples) asymme ic window. E e y 10 ms (80 samples), he
au oco ela ion coe icien s o windowed speech a e compu ed and con e ed o he LP
coe icien s using he Le inson-Du bin algo i hm. Then he LP coe icien s a e ans o med o
he LSP domain o quan iza ion and in e pola ion pu poses. The in e pola ed quan ized and
unquan ized il e s a e con e ed back o he LP il e coe icien s ( o cons uc he syn hesis and
weigh ing il e s o each sub ame).
The LP analysis window consis s o wo pa s: he i s pa is hal a Hamming window and
he second pa is a qua e o a cosine unc ion cycle. The window is gi en by:
( ) ( )
239...,002
159
2002
cos
0,...,199
399
2
cos 46.054.0







=





−
=






−
=
n
n
n
n
nw
lp
π
π
(3)
The e is a 5 ms look-ahead in he LP analysis which means ha 40 samples a e needed om
he u u e speech ame. This ansla es in o an ex a algo i hmic delay o 5 ms a he encode
s age. The LP analysis window applies o 120 samples om pas speech ames, 80 samples om
16 Eneko Año ga, Diploma Wo k
he p esen speech ame, and 40 samples om he u u e ame. The windowing p ocedu e is
illus a ed in Figu e 2.4.
Figu e 2.6 – Windowing p ocedu e in LP analysis
The di e en shading pa e ns iden i y co esponding exci a ion and LP analysis windows.
The windowed speech:
(
)
(
)
(
)
0,...,239
=
=
′
nnsnwns
lp
(4)
is used o compu e he au oco ela ion coe icien s:
( ) ( ) ( )
0,...,10
239
=−=
∑
=
kkn'sn'sk
kn
(5)
To a oid a i hme ic p oblems o low-le el inpu signals he alue o (0) has a lowe
bounda y o (0) = 1.0. A 60 Hz bandwid h expansion is applied by mul iplying he
au oco ela ion coe icien s wi h:
( )
1,...,10
2
2
1
2
0
=















π
−= k
k
expkw
s
lag
(6)
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 23
1. The incoming p essu e wa e is educed o a digi al signal h ough a sampling p ocess.
2. The s a ing and he ending poin s o he u e ance embedded in o he signal a e ob ained
using he VAD p ocess.
3. The spec um o he ex ac ed u e ance is enhanced by means o a p e-emphasis il e
which boos s he high equency componen s.
4. Se e al da a blocks a e ex ac ed om he enhanced signal.
5. The ex ac ed da a blocks a e windowed o educe leakage e ec s.
6. LPC componen s a e ex ac ed om he il e ed blocks.
7. LPC Ceps um componen s a e hen ex ac ed om he LPC ec o s.
8. The dimensionali y o he LPC Ceps um ec o s is educed using a SOM.
9. The esul ing ec o s a e scaled i he Recognize equi es i .
No hing can s ill be said abou he o e all e ec i eness o he FE block, since i depends on
he ecogni ion accu acy o he o e all sys em. As an example, i he comple e sys em achie es
low ecogni ion pe cen ages, ha can be caused by he FE block o he Recognize , bu , i i
achie es highe pe cen ages ha means ha he FE block was a leas able o p oduce da a ha
allowed hese ecogni ion accu acies. In o he wo ds, in o de o know he use ulness o he FE
block, he Recognize ou pu s mus be ob ained i s .

24 Eneko Año ga, Diploma Wo k
2.3 SOFTWARE DEVELOPMENT
2.3.1 Audio Reco de and Playe
As we ha e men ioned be o e, we buil a comple e audio Reco de and Playe . This ask was
no s ic ly necessa y bu he aim o his has been o lea n, p ac ice and imp o e he C ++
p og amming skills. The g aphic in e ace o he p og am is shown in he ollowing Figu e 2.8:
Figu e 2.8 – audioce Reco de & Playe
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 25
As we can see in he Figu e 2.8, wi h his so wa e we a e able o eco d he oice du ing
some ime, wi h di e en sample a es, di e en esolu ions and inally we can sa e i in a .wa
ile o nex p ocessing s eps like he speech coding. In addi ion, we will be able o play and
lis en o he eco ded signals. Along wi h he documen a ion (Annex 2) is included an elec onic
a achmen con aining he sou ce code in C++ used o build he so wa e and which is coming
wi h all he necessa y explana ions.
As we can see in he nex Figu e 2.9 we can see he di e ence be ween he di e en ways o
sampling. In gene al, he memo y occupied by he sound ile is p opo ional o he numbe o
samples pe second and he esolu ion o each sample. Fo he i s case he speech is sampled a
8.0 KHz and wi h 8 bi s o esolu ion (1 by e/sample), his means ha he memo y occupied o
he sound ile will be 8.000 (samples/sec) x 5 (sec) = 40.000 (samples) = 40.000 (samples) x 1
(by e/sample) = 40 Kby es.
Fo he second case he speech is sampled a 44.1 KHz and wi h 16 bi s o esolu ion (2
by es/sample), he memo y occupied o he sound ile will be 44.100 (samples/sec) x 5 (sec) =
220.500 (samples) = 220.500 (samples) x 2 (by es/sample) = 440.1 Kby es.
26 Eneko Año ga, Diploma Wo k
Figu e 2.9 – Di e en samplings o he same speech signal
I migh be a gued ha he highe he sampling equency and he highe he sampling
p ecision, he be e he ecogni ion accu acy. Howe e , i a no mal digi al phone, which samples
speech a 8.000 Hz wi h a 16 bi p ecision, is able o p ese e mos o he in o ma ion ca ied by
he signal [6], i does no seem necessa y o inc ease he sampling a e beyond 8.000 Hz o he
sampling p ecision o some hing highe han 16 bi s. Ano he eason behind hese se ings is ha
comme cial speech ecognize s ypically use compa able pa ame e alues and achie e
imp essi e esul s. So o he speech code we decided o use 8.000 Hz o sample a e and 16 bi s
o esolu ion as con igu a ion o eco d he oice.
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 27
2.3.2 Fea u e Ex ac o
A e buil he audioce Reco de and Playe we added some new unc ions o he p og am as
he pa which ep esen s he FE block. Fo ha , we based ou wo k on he ITU-T’s G.729
Recommenda ion. This Recommenda ion con ains he desc ip ion o an algo i hm o he coding
o speech signals using Linea P edic ion Coding and in which mo e p ocesses like he il e ing
o he sampled signal, he di ision in o blocks and he windowing o il e ed signal and he wo d
isola ion a e implici . The g aphic in e ace o he p og am is shown in he ollowing Figu e
2.10.
Figu e 2.10 – audioce Reco de & Playe & VAD De ec o & Speech Code
28 Eneko Año ga, Diploma Wo k
Wi h his so wa e we a e able o eco d and play oice signals and sa e hem in .wa iles.
In addi ion, we will be able o de ec and isola e he oice command om he sound ile, c ea e a
new ile wi h i and inally codi y o educe i s dimensionali y. Along wi h he documen a ion
(Annex 2) is included an elec onic a achmen con aining he sou ce code in C++ used o build
he so wa e and which is coming wi h all he necessa y explana ions.
When we p ess he bu on “VAD De ec o ...” we ha e o selec a sound ile which con ains
he eco ded oice. Then he sampled signal will go h ough all hese s eps:
(Sampling: he oice signal is sampled a 8.000 He z wi h 16 bi s o p ecision and sa ed
in a new ile called “le .wa ”.)
1. P e-emphasis il e : he sampled signal is il e ed by a second o de high-pass il e and
sa ed in a new ile called “ il e ed.wa ”.
2. Wo d Isola ion: he il e ed signal is passed h ough he VAD block o isola e he wo d
and sa ed in a ile called “ ad.wa ”.
We can ep esen he esul s o he di e en s ages o he signal using Ma lab and he ea lie
c ea ed sound iles (Figu e 2.11).

De elopmen o he Fea u e Ex ac o o Speech Recogni ion 29
Figu e 2.11 – Di e en s ages o he signal in VAD De ec o p ocess
As we can see, he i s g aphic shows he o iginal signal (Command “ igh ”), sampled a 8.0
KHz, wi h 16 bi s o esolu ion and in mono o 1 channel. We can ealize ha he signal has a
DC o se and also ha he end o he eco ding, is a bi noisy. The second g aphic shows he
il e ed sampled signal and inally he hi d g aphic shows he isola ed wo d, a e he VAD
p ocess.
A his poin we ha e o explain ha he esul shown o he VAD p ocess was ob ained
using an algo i hm myVAD.m [10] ob aining p e y good esul s. The p oblem wi h he algo i hm
de eloped o Colib i module is ha i does no a oid he silence and he non-speech pa s om
he oice e y well as he myVAD.m algo i hm does, as is shown in he nex Figu e 2.12.
30 Eneko Año ga, Diploma Wo k
Figu e 2.12 – Di e en ways o VAD p ocess
In he abo e Figu e 2.12, he second g aphic shows he eco ded signal a e VAD p ocess
using he algo i hm de eloped o Colib i module. He e, we can see ha e en hough he sound
ile is educed in some silen pa s o he signal a e a oided; hese p ocess is no doing he
co ec wo k as he hi d g aphic does. We ied o ind he solu ion o his bu inally we did no ,
so his pa need o be imp o ed and is le o u u e esea che s as we explain in he sec ion 3.2
Di ec ions o u u e esea ch.
To con inue ou esea ch we decided o use he ile c ea ed ough Ma lab and myVAD.m
algo i hm called “ ad2.wa ”. Once we ge he isola ed wo d (“ ad2.wa ”) we can s a codi ying
he speech clicking in “Speech Code ...” bu on. The nex s eps a e he ones which he isola ed
wo d signal will ollow:
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 31
1. Blocking: he isola ed wo d is di ided in o a sequence o da a blocks o ixed leng h,
called ames and mul iplied by Hamming window o same wid h.
2. LPC analysis: o each ame, 10 LPC coe icien s a e calcula ed.
3. Ceps um analysis: he 10 LPC coe icien s a e con e ed in 10 Ceps al coe icien ones
(Figu e 2.13).
Figu e 2.13 – The 10 Ceps al coe icien s o each ame
32 Eneko Año ga, Diploma Wo k
Once we ge he 10 Ceps al coe icien s o each ame i we click in “SOM...” bu on and
selec he ile wi h he Ceps al coe icien s (“c_coe . x ”), we will ed he inpu laye o he
Kohonen SOM, wi h his inpu da a one by one. The ou pu laye will o ganize i sel o ep esen
he inpu s in wo-dimensional space.
The aining p ocedu e in ol es he ollowing s eps:
1. The neu ons a e a anged in an n-dimensional la ice. Each neu on s o es a poin in an m-
dimensional space.
2. An inpu ec o is p esen ed o he SOM. The neu ons s a o compe e un il he one ha
s o es he closes poin o he inpu ec o p e ails. Once he dynamics o he ne wo k
con e ge, all he neu ons bu he p e ailing one will be inac i e. The ou pu o he SOM
is de ined as he co-o dina es o he p e ailing neu on in he la ice.
3. A neighbou hood unc ion is cen ed on he p e ailing neu on o he la ice. The alue o
his unc ion is one a he posi ion o he ac i e neu on, and dec eases wi h he dis ance
measu ed om he posi ion o he winning neu on.
4. The poin s s o ed by all he neu ons a e mo ed owa ds he inpu ec o in an amoun
p opo ional o he neighbou hood unc ion e alua ed in he posi ion o he la ice whe e
he neu on being modi ied s ands.
5. Re u n o 2, and epea s eps 2, 3, and 4 un il he a e age e o be ween he inpu ec o s
and he winning neu ons educes o a small alue.
A e he SOM is ained, he co-o dina es o he ac i e neu on in he la ice a e used as i s
ou pu s.
Along wi h he documen a ion (Annex 2) is included an elec onic a achmen con aining he
sou ce code in C++ used o build he so wa e and which is coming wi h all he necessa y
explana ions. In his case pa o he code o SOM has been de eloped bu s ill needs o be
imp o ed and inished, so his pa is le o u u e esea che s as we explain in he sec ion 3.2
Di ec ions o u u e esea ch.
De elopmen o he Fea u e Ex ac o o Speech Recogni ion 39
ANNEX 2:

40 Eneko Año ga, Diploma Wo k
DECLARATION:
I, Eneko Año ga, he unde signed, decla e ha I ha e made he diploma wo k by mysel . I
am awa e o he po en ial consequences in he e en o a b each o his decla a ion.