scieee Open visual document viewer

Development of the Feature Extractor for Speech Recognition

Añorga Irigoien, Eneko

Abstract

With this diploma work we have attempted to give continuity to the previous work done by other researchers called, Voice Operating Intelligent Wheelchair – VOIC [1]. A development of a wheelchair controlled by voice is presented in this work and is designed for physically disabled people, who cannot control their movements. This work describes basic components of speech recognition and wheelchair control system. Going to the grain, a speech recognizer system is comprised of two distinct blocks, a Feature Extractor and a Recognizer. The present work is targeted at the realization of an adequate Feature Extractor block which uses a standard LPC Cepstrum coder, which translates the incoming speech into a trajectory in the LPC Cepstrum feature space, followed by a Self Organizing Map, which classifies the outcome of the coder in order to produce optimal trajectory representations of words in reduced dimension feature spaces. Experimental results indicate that trajectories on such reduced dimension spaces can provide reliable representations of spoken words. The Recognizer block is left for future researchers. The main contributions of this work have been the research and approach of a new technology for development issues and the realization of applications like a voice recorder and player and a complete Feature Extractor system.

Full text

ENEKO AÑORGA DEVELOPMENT OF THE FEATURE EXTRACTOR FOR SPEECH RECOGNITION DIPLOMA WORK MARIBOR, OCTOBER 2009 i FAKULTETA ZA ELEKTROTEHNIKO, RAČUNALNIŠTVO IN INFORMATIKO 2000 Ma ibo , Sme ano a ul. 17 Diploma Wo k o Elec onic Enginee ing S uden P og am DEVELOPMENT OF THE FEATURE EXTRACTOR FOR SPEECH RECOGNITION S uden : Eneko A ñ o ga S uden p og am: Elec onic Enginee ing Men o : P o . D . Riko ŠAFARIČ Asis . P o . D . Suzana URAN Ma ibo , Oc obe 2009 ii ACKNOWLEDGMENTS Thanks o P o . D . Riko ŠAFARIČ o his assis ance and help ul ad ices in ca ying ou he diploma wo k. Special hanks o my amily and iends who a e in all momen s beside me. iii DEVELOPMENT OF THE FEATURE EXTRACTOR FOR SPEECH RECOGNITION Key wo ds: oice ope a ed wheelchai , speech ecogni ion, oice ac i i y de ec ion, neu al ne wo ks, ul asound senso ne UDK: 004.934:681.5(043.2) Abs ac Wi h his diploma wo k we ha e a emp ed o gi e con inui y o he p e ious wo k done by o he esea che s called, Voice Ope a ing In elligen Wheelchai – VOIC [1]. A de elopmen o a wheelchai con olled by oice is p esen ed in his wo k and is designed o physically disabled people, who canno con ol hei mo emen s. This wo k desc ibes basic componen s o speech ecogni ion and wheelchai con ol sys em. Going o he g ain, a speech ecognize sys em is comp ised o wo dis inc blocks, a Fea u e Ex ac o and a Recognize . The p esen wo k is a ge ed a he ealiza ion o an adequa e Fea u e Ex ac o block which uses a s anda d LPC Ceps um code , which ansla es he incoming speech in o a ajec o y in he LPC Ceps um ea u e space, ollowed by a Sel O ganizing Map, which classi ies he ou come o he code in o de o p oduce op imal ajec o y ep esen a ions o wo ds in educed dimension ea u e spaces. Expe imen al esul s indica e ha ajec o ies on such educed dimension spaces can p o ide eliable ep esen a ions o spoken wo ds. The Recognize block is le o u u e esea che s. The main con ibu ions o his wo k ha e been he esea ch and app oach o a new echnology o de elopmen issues and he ealiza ion o applica ions like a oice eco de and playe and a comple e Fea u e Ex ac o sys em. i RAZVOJ PREVODNIKA SIGNALA ZA PREPOZNAVO GOVORA Ključne besede: glaso no oden in alidski oziček, p epozna a go o a, zazna a glaso ne ak i nos i ne onske m eže, ul az očna senzo ska m eža UDK: 004.934:681.5(043.2) Po ze ek S em diplomskim delom sem poskusil nadalje a i delo aziska e z naslo om Voice Ope a ing In elligen Wheelchai – VOIC [1]. V em diplomskem delu je udi p eds a ljen az oj glaso no odenega in alidskega ozička, na ejenega za elesno p izade e ljudi, ki ne mo ejo nadzo o a i s ojih gibo . To delo opisuje osno ne komponen e go o nega nadzo a in sis ema odenja in alidskega ozička. Sis em go o nega nadzo a u a na a a d a azlična dela; p e odnik signala in p epozna alec. V em diplomskem delu se os edo očam na p e odnos us eznega p e odnika signala na osno i s anda dnega LPC Ceps um kode ja, ki pos eduje p ihajajoči go o po LPC Ceps um p os o a, emu pos opku pa sledi .i. “samoo ganizacijska ka a” (Sel O ganizing Map), ki az s i ezul a kode ja za op imalni p ikaz besed na zmanjšanih dimenzijah p os o a. Poskusni ezul a i kažejo, da lahko e po i na zmanjšanih dimenzijah p os o a zago o ijo zaneslji p ikaz izgo o jenih besed. P epozna alec je lahko p edme azisko anja š uden o udi p ihodnos i. Gla ni namen ega diplomskega dela s a bili aziska a in poskus upo abe no e ehnologije az ojne namene, kako udi upo aba aplikacij ko so snemalec in p ed ajalnik z oka e celo en sis em p e odnika signala. Table o Con en s 1 INTRODUCTION ........................................................................................................ 1 1.1 MOTIVATION ..................................................................................................... 1 1.2 OBJECTIVES AND CONTRIBUTION OF THIS WORK.................................. 3 1.3 ORGANIZATION OF THIS WORK ................................................................... 5 1.4 RESOURCES ........................................................................................................ 5 2 DEVELOPMENT ......................................................................................................... 6 2.1 BRIEF DESCRIPTION OF COLIBRI MODULE ............................................... 6 2.1.1 Ha dwa e ......................................................................................................... 6 2.1.2 So wa e ........................................................................................................ 10 2.2 DESIGN OF THE FEATURE EXTRACTOR ................................................... 11 2.2.1 Speech coding ............................................................................................... 12 2.2.1.1 Speech sampling ..................................................................................... 12 2.2.1.2 P e-emphasis il e .................................................................................. 13 2.2.1.3 Wo d Isola ion ........................................................................................ 13 2.2.1.4 Speech coding ........................................................................................ 14 2.2.1.5 Summing up ........................................................................................... 18 2.2.2 Dimensionali y educ ion using SOM ........................................................... 19 2.2.2.1 Op ional Signal Scaling .......................................................................... 22 2.2.2.2 Summing up ........................................................................................... 22 2.3 SOFTWARE DEVELOPMENT ......................................................................... 24 2.3.1 Audio Reco de and Playe ........................................................................... 24 2.3.2 Fea u e Ex ac o ........................................................................................... 27 3 CONCLUSIONS ........................................................................................................ 33 3.1 SUMMARY OF RESULTS ................................................................................ 33 3.2 DIRECTIONS FOR FUTURE RESEARCH ...................................................... 34 4 REFERENCES ........................................................................................................... 36 i GLOSARY OF SIMBOLS Name Desc ip ion 1/Â(z) a i w lp (n) H h1 (z) s(n) s'(n) (k) w lag (n) '(k) k i s 0 LP syn hesis il e LP coe icien s (a 0 = 1.0) LP analysis window Inpu high-pass il e P ep ocessed/ il e ed speech signal Windowed speech signal Au o-co ela ion coe icien s Co ela ion lag window Modi ied au o-co ela ion coe icien s Re lec ion coe icien s Sampling equency Bandwid h expansion Table 1 – Glossa y o symbols ii GLOSSARY OF ACRONYMS Ac onym Desc ip ion VOIC DSP LPC LP SOM CE SODIMM CAN GPIO BSP ITU ITU-T FE VAD DTX CNG RNN VQ HMM Voice Ope a ed In elligen Wheelchai Digi al Signal P ocesso Linea P edic ion Coding Linea P edic ion Sel O ganizing Maps Compac Edi ion Small Ou line Dual In-line Memo y Module Con ol A ea Ne wo k Gene al Pu pose Inpu /Ou pu Boa d Suppo Package In e na ional Telecommunica ion Union Telecommunica ion S anda diza ion Sec o Fea u e Ex ac o Voice Ac i i y De ec ion Discon inuous T ansmission Com o Noise Gene a o Recu en Neu al Ne wo k Vec o Quan iza ion Hidden Ma ko Model Table 2 – Glossa y o ac onyms De elopmen o he Fea u e Ex ac o o Speech Recogni ion 7 Figu e 2.1 – Colib i PXA320 Module Speci ica ions: CPU PXA320 806MHz Memo y 128MB DDR RAM (32Bi ) 1GB NAND Flash (8Bi ) In e aces 16Bi Ex e nal BUS Compac Flash/PCMCIA LCD (SVGA) Touch Sc een Audio I/O (16Bi S e eo) CMOS image senso I2C SPI 2x SD Ca d USB Hos /De ice 100MBi E he ne 2x UART I DA PWM 127 GPIOs So wa e P e-ins alled Windows CE 5.0/6.0 Size 67.6 x 36.7 x 5.2 mm Tempe a u e Range 0 o +70°C -45 o +85°C (IT e sion) 8 Eneko Año ga, Diploma Wo k In o de o ha e a lexible de elopmen en i onmen o explo e he unc ionali y and pe o mance o he Colib i modules he Colib i E alua ion Boa d is used (See Figu e 2.2). Besides he use in e aces i p o ides nume ous communica ion channels as well as a con igu able jumpe a ea o hook up he Colib i GPIOs o he desi ed unc ion. To acili a e in e acing o he cus om ha dwa e he Colib i E alua ion Boa d p o ides he bu e ed CPU bus on a sepa a e connec o . De elopmen o he Fea u e Ex ac o o Speech Recogni ion 9 Figu e 2.2 – Colib i E alua ion Boa d Module Speci ica ions: CPU Modules Colib i PXA270 Colib i PXA300 Colib i PXA310 Colib i PXA320 In e aces 10/100MBi E he ne USB Hos /De ice USB Hos 2x PS/2 Analogue VGA Gene ic LCD Connec o TFT: Philips LB064V02-A1 Line-In, Line-Ou , Mic-In I DA 2x RS232 CAN (Philips SJA1000) SD Ca d Compac Flash Powe Supply: Requi ed Inpu : 7-24VDC, 3-50W On-boa d Con e e : 3.3V, 5V max 5A Size: 200 x 200 mm 10 Eneko Año ga, Diploma Wo k The ecei ed in oice om To adex, o he Colib i XScale® PXA320, plus he Colib i E alua ion Ca ie Boa d and plus he suppo hou s a e shown in he nex Figu e 2.3: Figu e 2.3 – In oice om To adex 2.1.2 So wa e The module is shipped wi h a p eins alled WinCE 5.0 image wi h WinCE Co e license. O he OS like Embedded Linux a e a ailable om he hi d-pa y. To adex p o ides a WinCE 5.0 image and a WinCE 6.0. All WinCE images con ain he To adex Boa d Suppo Package (BSP) which is one o he mos ad anced BSPs a ailable on he ma ke . Besides he s anda d Windows CE unc ionali y, i includes a la ge numbe o addi ional d i e s as well as op imized e sions o s anda d d i e s o he mos common in e aces and is easily cus omizable by egis y se ings o adap o speci ic ha dwa e. The Mic oso ® eMbedded Visual C++ 4.0 ool is used as desk op de elopmen en i onmen o c ea ing he applica ions and sys em componen s o Windows® CE .NET powe ed de ices. In conclusion, all he so wa e p esen ed in his wo k was done using he Mic oso ® eMbedded Visual C++ 4.0 de elopmen ool and he To adex BSP ool (Figu e 2.4). De elopmen o he Fea u e Ex ac o o Speech Recogni ion 11 Figu e 2.4 – Mic oso ® eMbedded Visual C++ 4.0 and Windows® CE 2.2 DESIGN OF THE FEATURE EXTRACTOR As s a ed be o e, in a speech ecogni ion p oblem he FE block has o p ocess he incoming in o ma ion, he speech signal, so ha i s ou pu eases he wo k o he classi ica ion s age. The app oach used in his wo k designs he FE block and di ides i in o wo consecu i e sub-blocks: he i s is based on speech coding echniques, and he second uses a SOM o u he op imiza ion (da a dimensionali y educ ion). The di e en blocks and sub-blocks a e shown in he nex Figu e 2.5: 12 Eneko Año ga, Diploma Wo k Figu e 2.5 – FE schema ic 2.2.1 Speech coding 2.2.1.1 Speech sampling The speech was eco ded and sampled using a ela i ely inexpensi e dynamic mic ophone and a Colib i’s audio inpu in e ace. The incoming signal was sampled a 8.000 Hz wi h 16 bi s o esolu ion. I migh be a gued ha a highe sampling equency, o mo e sampling p ecision, is needed in o de o highe ecogni ion accu acy. Howe e , i a no mal digi al phone, which samples speech a 8.000 Hz wi h a 16 bi esolu ion, is able o p ese e mos o he in o ma ion ca ied by he signal [6], i does no seem necessa y o inc ease he sampling a e beyond 8.000 Hz o he sampling p ecision o some hing highe han 16 bi s. Ano he eason behind hese se ings is ha De elopmen o he Fea u e Ex ac o o Speech Recogni ion 13 comme cial speech ecognize s ypically use compa able pa ame e alues and achie e imp essi e esul s. 2.2.1.2 P e-emphasis il e A e sampling he inpu signal is con enien o il e i wi h a second o de high-pass il e wi h cu o equency a 140 Hz. The il e se es as a p ecau ion agains undesi ed low- equency componen s. The esul ing il e is gi en by: ( ) 21 21 1 9114024 . 0 9059465 . 1 1 46363718.092724705.046363718.0 −− −− + − +− = z z zz zH h (1) 2.2.1.3 Wo d Isola ion Despi e he ac ha he sampled signal had pauses be ween he u e ances, i was s ill needed o de e mine he s a ing and ending poin s o he wo d u e ances in o de o know exac ly he signal ha cha ac e ized each wo d. To accomplish his, we decided o use VAD (Voice Ac i i y De ec ion) echnique used in speech p ocessing, ins ead o using he olling a e age and he h eshold, de e mined by he s a and end o each wo d, used in he p e ious wo ks, wi h he aim o achie ing mo e accu acy and e iciency. Fo ha , we based ou wo k in he Annex B om he ITU’s Recommenda ion G.729 [5], whe e a sou ce code in C language abou he VAD is e icien ly de eloped. VAD is a me hod which di e en ia es speech om silence o noise signal o aid in speech p ocessing and he Annex B p o ides a high le el desc ip ion o he Voice Ac i i y De ec ion (VAD), Discon inuous T ansmission (DTX) and Com o Noise Gene a o (CNG) algo i hms. These algo i hms a e used o educe he ansmission a e du ing silence pe iods o speech. They 14 Eneko Año ga, Diploma Wo k a e designed and op imized o wo k in conjunc ion wi h [ITU-T V.70]. [ITU-T V.70] manda es he use o speech coding me hods. The algo i hms a e adap ed o ope a e wi h bo h he ull e sion o G.729 and Annex B. Le ’s see a gene al desc ip ion o he VAD algo i hm: The VAD algo i hm makes a oice ac i i y decision e e y 10 ms in acco dance wi h he ame size o he p e-p ocessed ( il e ed) signal. A se o di e ence pa ame e s is ex ac ed and used o an ini ial decision. The pa ame e s a e he ull-band ene gy, he low-band ene gy, he ze o-c ossing a e and a spec al measu e. The long- e m a e ages o he pa ame e s du ing non- ac i e oice segmen s ollow he changing na u e o he backg ound noise. A se o di e en ial pa ame e s is ob ained a each ame. These a e a di e ence measu e be ween each pa ame e and i s espec i e long- e m a e age. The ini ial oice ac i i y decision is ob ained using a piecewise linea decision bounda y be ween each pai o di e en ial pa ame e s. A inal oice ac i i y decision is ob ained by smoo hing he ini ial decision. The ou pu o he VAD module is ei he 1 o 0, indica ing he p esence o absence o oice ac i i y espec i ely. I he VAD ou pu is 1, he G.729 speech codec is in oked o code/decode he ac i e oice ames. Howe e , i he VAD ou pu is 0, he DTX/CNG algo i hms desc ibed he ein a e used o code/decode he non-ac i e oice ames. 2.2.1.4 Speech coding A e he signal was sampled, he spec um was la ened, and he u e ances we e isola ed we ied o codi y i using he Linea P edic ion Coding (LPC) me hod [3]. In a a ie y o applica ions, i is desi able o comp ess a speech signal o e icien ansmission o s o age. Fo example, o accommoda e many speech signals in a gi en bandwid h o a cellula phone sys em, each digi ized speech signal is comp essed be o e ansmission. Fo medium o low bi - a e speech code s, LPC me hod is mos widely used. Redundancy in a speech signal is emo ed by passing he signal h ough a speech analysis il e . De elopmen o he Fea u e Ex ac o o Speech Recogni ion 15 The ou pu o he il e , e med he esidual e o signal, has less edundancy han he o iginal speech signal and can be quan ized by a smalle numbe o bi s han he o iginal speech. The sho - e m analysis and syn hesis il e s a e based on 10 h o de linea p edic ion (LP) il e s. The LP syn hesis il e is de ined as: ∑ = − + = 10 1 ˆ 1 1 )( ˆ1 i i i za zA (2) whe e â i , i = 1,...,10, a e he quan ized Linea P edic ion (LP) coe icien s. Sho - e m p edic ion o linea p edic ion analysis is pe o med once pe speech ame using he au oco ela ion me hod wi h a 30 ms (240 samples) asymme ic window. E e y 10 ms (80 samples), he au oco ela ion coe icien s o windowed speech a e compu ed and con e ed o he LP coe icien s using he Le inson-Du bin algo i hm. Then he LP coe icien s a e ans o med o he LSP domain o quan iza ion and in e pola ion pu poses. The in e pola ed quan ized and unquan ized il e s a e con e ed back o he LP il e coe icien s ( o cons uc he syn hesis and weigh ing il e s o each sub ame). The LP analysis window consis s o wo pa s: he i s pa is hal a Hamming window and he second pa is a qua e o a cosine unc ion cycle. The window is gi en by: ( ) ( ) 239...,002 159 2002 cos 0,...,199 399 2 cos 46.054.0        =      − =       − = n n n n nw lp π π (3) The e is a 5 ms look-ahead in he LP analysis which means ha 40 samples a e needed om he u u e speech ame. This ansla es in o an ex a algo i hmic delay o 5 ms a he encode s age. The LP analysis window applies o 120 samples om pas speech ames, 80 samples om 16 Eneko Año ga, Diploma Wo k he p esen speech ame, and 40 samples om he u u e ame. The windowing p ocedu e is illus a ed in Figu e 2.4. Figu e 2.6 – Windowing p ocedu e in LP analysis The di e en shading pa e ns iden i y co esponding exci a ion and LP analysis windows. The windowed speech: ( ) ( ) ( ) 0,...,239 = = ′ nnsnwns lp (4) is used o compu e he au oco ela ion coe icien s: ( ) ( ) ( ) 0,...,10 239 =−= ∑ = kkn'sn'sk kn (5) To a oid a i hme ic p oblems o low-le el inpu signals he alue o (0) has a lowe bounda y o (0) = 1.0. A 60 Hz bandwid h expansion is applied by mul iplying he au oco ela ion coe icien s wi h: ( ) 1,...,10 2 2 1 2 0 =                π −= k k expkw s lag (6) De elopmen o he Fea u e Ex ac o o Speech Recogni ion 23 1. The incoming p essu e wa e is educed o a digi al signal h ough a sampling p ocess. 2. The s a ing and he ending poin s o he u e ance embedded in o he signal a e ob ained using he VAD p ocess. 3. The spec um o he ex ac ed u e ance is enhanced by means o a p e-emphasis il e which boos s he high equency componen s. 4. Se e al da a blocks a e ex ac ed om he enhanced signal. 5. The ex ac ed da a blocks a e windowed o educe leakage e ec s. 6. LPC componen s a e ex ac ed om he il e ed blocks. 7. LPC Ceps um componen s a e hen ex ac ed om he LPC ec o s. 8. The dimensionali y o he LPC Ceps um ec o s is educed using a SOM. 9. The esul ing ec o s a e scaled i he Recognize equi es i . No hing can s ill be said abou he o e all e ec i eness o he FE block, since i depends on he ecogni ion accu acy o he o e all sys em. As an example, i he comple e sys em achie es low ecogni ion pe cen ages, ha can be caused by he FE block o he Recognize , bu , i i achie es highe pe cen ages ha means ha he FE block was a leas able o p oduce da a ha allowed hese ecogni ion accu acies. In o he wo ds, in o de o know he use ulness o he FE block, he Recognize ou pu s mus be ob ained i s . 24 Eneko Año ga, Diploma Wo k 2.3 SOFTWARE DEVELOPMENT 2.3.1 Audio Reco de and Playe As we ha e men ioned be o e, we buil a comple e audio Reco de and Playe . This ask was no s ic ly necessa y bu he aim o his has been o lea n, p ac ice and imp o e he C ++ p og amming skills. The g aphic in e ace o he p og am is shown in he ollowing Figu e 2.8: Figu e 2.8 – audioce Reco de & Playe De elopmen o he Fea u e Ex ac o o Speech Recogni ion 25 As we can see in he Figu e 2.8, wi h his so wa e we a e able o eco d he oice du ing some ime, wi h di e en sample a es, di e en esolu ions and inally we can sa e i in a .wa ile o nex p ocessing s eps like he speech coding. In addi ion, we will be able o play and lis en o he eco ded signals. Along wi h he documen a ion (Annex 2) is included an elec onic a achmen con aining he sou ce code in C++ used o build he so wa e and which is coming wi h all he necessa y explana ions. As we can see in he nex Figu e 2.9 we can see he di e ence be ween he di e en ways o sampling. In gene al, he memo y occupied by he sound ile is p opo ional o he numbe o samples pe second and he esolu ion o each sample. Fo he i s case he speech is sampled a 8.0 KHz and wi h 8 bi s o esolu ion (1 by e/sample), his means ha he memo y occupied o he sound ile will be 8.000 (samples/sec) x 5 (sec) = 40.000 (samples) = 40.000 (samples) x 1 (by e/sample) = 40 Kby es. Fo he second case he speech is sampled a 44.1 KHz and wi h 16 bi s o esolu ion (2 by es/sample), he memo y occupied o he sound ile will be 44.100 (samples/sec) x 5 (sec) = 220.500 (samples) = 220.500 (samples) x 2 (by es/sample) = 440.1 Kby es. 26 Eneko Año ga, Diploma Wo k Figu e 2.9 – Di e en samplings o he same speech signal I migh be a gued ha he highe he sampling equency and he highe he sampling p ecision, he be e he ecogni ion accu acy. Howe e , i a no mal digi al phone, which samples speech a 8.000 Hz wi h a 16 bi p ecision, is able o p ese e mos o he in o ma ion ca ied by he signal [6], i does no seem necessa y o inc ease he sampling a e beyond 8.000 Hz o he sampling p ecision o some hing highe han 16 bi s. Ano he eason behind hese se ings is ha comme cial speech ecognize s ypically use compa able pa ame e alues and achie e imp essi e esul s. So o he speech code we decided o use 8.000 Hz o sample a e and 16 bi s o esolu ion as con igu a ion o eco d he oice. De elopmen o he Fea u e Ex ac o o Speech Recogni ion 27 2.3.2 Fea u e Ex ac o A e buil he audioce Reco de and Playe we added some new unc ions o he p og am as he pa which ep esen s he FE block. Fo ha , we based ou wo k on he ITU-T’s G.729 Recommenda ion. This Recommenda ion con ains he desc ip ion o an algo i hm o he coding o speech signals using Linea P edic ion Coding and in which mo e p ocesses like he il e ing o he sampled signal, he di ision in o blocks and he windowing o il e ed signal and he wo d isola ion a e implici . The g aphic in e ace o he p og am is shown in he ollowing Figu e 2.10. Figu e 2.10 – audioce Reco de & Playe & VAD De ec o & Speech Code 28 Eneko Año ga, Diploma Wo k Wi h his so wa e we a e able o eco d and play oice signals and sa e hem in .wa iles. In addi ion, we will be able o de ec and isola e he oice command om he sound ile, c ea e a new ile wi h i and inally codi y o educe i s dimensionali y. Along wi h he documen a ion (Annex 2) is included an elec onic a achmen con aining he sou ce code in C++ used o build he so wa e and which is coming wi h all he necessa y explana ions. When we p ess he bu on “VAD De ec o ...” we ha e o selec a sound ile which con ains he eco ded oice. Then he sampled signal will go h ough all hese s eps: (Sampling: he oice signal is sampled a 8.000 He z wi h 16 bi s o p ecision and sa ed in a new ile called “le .wa ”.) 1. P e-emphasis il e : he sampled signal is il e ed by a second o de high-pass il e and sa ed in a new ile called “ il e ed.wa ”. 2. Wo d Isola ion: he il e ed signal is passed h ough he VAD block o isola e he wo d and sa ed in a ile called “ ad.wa ”. We can ep esen he esul s o he di e en s ages o he signal using Ma lab and he ea lie c ea ed sound iles (Figu e 2.11). De elopmen o he Fea u e Ex ac o o Speech Recogni ion 29 Figu e 2.11 – Di e en s ages o he signal in VAD De ec o p ocess As we can see, he i s g aphic shows he o iginal signal (Command “ igh ”), sampled a 8.0 KHz, wi h 16 bi s o esolu ion and in mono o 1 channel. We can ealize ha he signal has a DC o se and also ha he end o he eco ding, is a bi noisy. The second g aphic shows he il e ed sampled signal and inally he hi d g aphic shows he isola ed wo d, a e he VAD p ocess. A his poin we ha e o explain ha he esul shown o he VAD p ocess was ob ained using an algo i hm myVAD.m [10] ob aining p e y good esul s. The p oblem wi h he algo i hm de eloped o Colib i module is ha i does no a oid he silence and he non-speech pa s om he oice e y well as he myVAD.m algo i hm does, as is shown in he nex Figu e 2.12. 30 Eneko Año ga, Diploma Wo k Figu e 2.12 – Di e en ways o VAD p ocess In he abo e Figu e 2.12, he second g aphic shows he eco ded signal a e VAD p ocess using he algo i hm de eloped o Colib i module. He e, we can see ha e en hough he sound ile is educed in some silen pa s o he signal a e a oided; hese p ocess is no doing he co ec wo k as he hi d g aphic does. We ied o ind he solu ion o his bu inally we did no , so his pa need o be imp o ed and is le o u u e esea che s as we explain in he sec ion 3.2 Di ec ions o u u e esea ch. To con inue ou esea ch we decided o use he ile c ea ed ough Ma lab and myVAD.m algo i hm called “ ad2.wa ”. Once we ge he isola ed wo d (“ ad2.wa ”) we can s a codi ying he speech clicking in “Speech Code ...” bu on. The nex s eps a e he ones which he isola ed wo d signal will ollow: De elopmen o he Fea u e Ex ac o o Speech Recogni ion 31 1. Blocking: he isola ed wo d is di ided in o a sequence o da a blocks o ixed leng h, called ames and mul iplied by Hamming window o same wid h. 2. LPC analysis: o each ame, 10 LPC coe icien s a e calcula ed. 3. Ceps um analysis: he 10 LPC coe icien s a e con e ed in 10 Ceps al coe icien ones (Figu e 2.13). Figu e 2.13 – The 10 Ceps al coe icien s o each ame 32 Eneko Año ga, Diploma Wo k Once we ge he 10 Ceps al coe icien s o each ame i we click in “SOM...” bu on and selec he ile wi h he Ceps al coe icien s (“c_coe . x ”), we will ed he inpu laye o he Kohonen SOM, wi h his inpu da a one by one. The ou pu laye will o ganize i sel o ep esen he inpu s in wo-dimensional space. The aining p ocedu e in ol es he ollowing s eps: 1. The neu ons a e a anged in an n-dimensional la ice. Each neu on s o es a poin in an m- dimensional space. 2. An inpu ec o is p esen ed o he SOM. The neu ons s a o compe e un il he one ha s o es he closes poin o he inpu ec o p e ails. Once he dynamics o he ne wo k con e ge, all he neu ons bu he p e ailing one will be inac i e. The ou pu o he SOM is de ined as he co-o dina es o he p e ailing neu on in he la ice. 3. A neighbou hood unc ion is cen ed on he p e ailing neu on o he la ice. The alue o his unc ion is one a he posi ion o he ac i e neu on, and dec eases wi h he dis ance measu ed om he posi ion o he winning neu on. 4. The poin s s o ed by all he neu ons a e mo ed owa ds he inpu ec o in an amoun p opo ional o he neighbou hood unc ion e alua ed in he posi ion o he la ice whe e he neu on being modi ied s ands. 5. Re u n o 2, and epea s eps 2, 3, and 4 un il he a e age e o be ween he inpu ec o s and he winning neu ons educes o a small alue. A e he SOM is ained, he co-o dina es o he ac i e neu on in he la ice a e used as i s ou pu s. Along wi h he documen a ion (Annex 2) is included an elec onic a achmen con aining he sou ce code in C++ used o build he so wa e and which is coming wi h all he necessa y explana ions. In his case pa o he code o SOM has been de eloped bu s ill needs o be imp o ed and inished, so his pa is le o u u e esea che s as we explain in he sec ion 3.2 Di ec ions o u u e esea ch. De elopmen o he Fea u e Ex ac o o Speech Recogni ion 39 ANNEX 2: 40 Eneko Año ga, Diploma Wo k DECLARATION: I, Eneko Año ga, he unde signed, decla e ha I ha e made he diploma wo k by mysel . I am awa e o he po en ial consequences in he e en o a b each o his decla a ion.