scieee Open visual document viewer

Aprendizaje en línea en traducción automática basada en redes neuronales

Cebrián Chuliá, Luis

Abstract

[EN] High quality translations are in high demand these days. Although machine translation offers acceptable performance, it is not sufficient in some cases and human supervision is required. In order to ease the translation task of the human, machine translation systems take part in this process. When a sentence in the source language needs to be translated, it is fed to the system which outputs a hypothesis translation. The human then, corrects this hypothesis (also known as post-editing) in order to obtain a high quality translation. Being able to transfer the knowledge that a human translator exhibit when post-editing a translation to the machine translation system is a desirable feature, as it has been proven that a more accurate machine translation system helps to increase the efficiency of the post-editing process. Because the post-editing scenario requires an already trained system, online learning techniques are suited for this task. In this work, three online learning algorithms have been proposed and applied to a neural machine translation sys- tem in a post-editing scenario. They rely on the Passive-Aggressive online learn- ing approach in which the model is updated after every sample in order to fulfil a correctness criterion while remembering previously learned information. The goal is to adapt and refine an already trained system with new samples on-the- fly as the post-editing process takes place (hence, the update time must be kept under control). Moreover, these new algorithms are compared with well-stablished online learning variants of the stochastic gradient descent algorithm. Results show im- provements on the translation quality of the system after applying these algo- rithms, reducing human effort in the post-editing process.

Full text

Depa amen o de Sis emas In o má icos y Compu ación Uni e si a Poli ècnica de València Online Lea ning in Neu al Machine T ansla ion MASTER THESIS Más e en In eligencia A i icial, Reconocimien o de Fo mas e Imagen Digi al Au ho : Luis Ceb ián Chuliá Tu o : F ancisco Casacube a Nolla Ál a o Pe is Ab il Cou se 2016-2017 Resum La aducció de g an quali a es oba mol demanada en l’ac uali a . To i que la aducció au omà ica o e ix unes p es acions accep ables, en alguns casos no és su icien i és necessà ia la supe isió humana. Pe a acili a la asca de aducció de l’humà, els sis emes de aducció au omà ica p enen pa en aques p océs. Quan una no a o ació en el llengua ge o igen necessi a se aduïda, es a s’in oduïx en el sis ema, el qual ob é com a eixida una hipò esi de aducció. Lla o s, l’humà co egix aques a hipò esi ( ambé conegu com a pos -edi a ) pe a ob ind e una aducció de majo quali a . Se capaços de ans e i el coneixemen que l’ humà exhibix quan eali za la asca de pos -edició al sis ema de aducció au omà ica és una ca ac e ís ica desi jable ja que s’ha demos a que un sis ema de aducció mes p ecís ajuda a augmen a l‘e iciència del p océs de pos -edició. Pel e que el p océs de pos -edició eque ix un sis ema ja en ena , les ècni- ques d’ap enen a ge en línia són les adequades pe aques a asca. En es e eball, es p oposen es algo i mes d’ap enen a ge en línia aplica s a un aduc o au- omà ic neu onal en un escena i de pos -edició. Es os algo i mes es basen en l’ap oximació en línia Passi e-Agg essi e en la qual el model s’ac uali za desp és de cada mos a amb l’objec iu de compli un c i e i de co ecció al ma eix emps que man é in o mació p è ia ap esa. L’objec iu és adap a i e ina un sis ema ja en ena amb no es mos es al ol men e el p océs de pos -edició es du a e me (pe an , el emps d’ac uali zació ha de man eni -se con ola ). A més, es os algo i mes es compa en amb al es ben conegudes a ian s en línia de l’algo i me de descens pe g adien es ocàs ic. Els esul a s mos en una millo a en la quali a de les aduccions desp és d’aplica es os algo i mes, e- duin així l’es o ç humà en el p océs de pos -edició. Pa aules clau: Ap enen a ge en línia, T aducció au omà ica neu onal, Passi e- Agg esi e Resumen La aducción de g an calidad es á muy demandada en la ac ualidad. A pesa de que la aducción au omá ica o ece unas p es aciones acep ables, en algunos casos no es su icien e y es necesa ia la supe isión humana. Pa a acili a la a ea de aducción del humano, los sis emas de aducción au omá ica oman pa e en es e p oceso. Cuando una nue a o ación en el idioma o igen necesi a se adu- cida, es a se in oduce en el sis ema, el cual ob iene como salida una hipó esis de aducción. El humano en onces, co ige es a hipó esis ( ambién conocido como pos -edi a ) pa a ob ene una aducción de mayo calidad. Se capaz de ans e- i el conocimien o que el humano exhibe cuando ealiza la a ea de pos -edición al sis ema de aducción au omá ica es una ca ac e ís ica deseable pues o que se ha demos ado que un sis ema de aducción mas p eciso ayuda a aumen a la e iciencia del p oceso de pos -edición. Debido a que el p oceso de pos -edición equie e un sis ema ya en enado, las écnicas de ap endizaje en línea son las adecuadas pa a es a a ea. En es e aba- jo, se p oponen es algo i mos de ap endizaje en línea aplicados a un aduc o au omá ico neu onal en un escena io de pos -edición. Es os algo i mos se basan iii i en la ap oximación en línea Passi e-Agg essi e en la cual el modelo se ac ualiza después de cada mues a con el obje i o de cumpli un c i e io de co ección a la ez que man eniendo in o mación p e ia ap endida. El obje i o es adap a y e ina un sis ema ya en enado con nue as mues as al uelo mien as el p o- ceso de pos -edición se lle a a cabo (po an o, el iempo de ac ualización debe man ene se bajo con ol). Además, es os algo i mos se compa an con o as bien conocidas a ian es en línea del algo i mo de descenso po g adien e es ocás ico. Los esul ados mues- an una mejo a en la calidad de las aducciones después de aplica es os algo- i mos, educiendo así el es ue zo humano en el p oceso de pos -edición. Palab as cla e: Ap endizaje en línea, T aducción au omá ica neu onal, Passi e- Agg esi e Abs ac High quali y ansla ions a e in high demand hese days. Al hough machine ansla ion o e s accep able pe o mance, i is no su icien in some cases and human supe ision is equi ed. In o de o ease he ansla ion ask o he human, machine ansla ion sys ems ake pa in his p ocess. When a sen ence in he sou ce language needs o be ansla ed, i is ed o he sys em which ou pu s a hypo hesis ansla ion. The human hen, co ec s his hypo hesis (also known as pos -edi ing) in o de o ob ain a high quali y ansla ion. Being able o ans e he knowledge ha a human ansla o exhibi when pos -edi ing a ansla ion o he machine ansla ion sys em is a desi able ea u e, as i has been p o en ha a mo e accu a e machine ansla ion sys em helps o inc ease he e iciency o he pos -edi ing p ocess. Because he pos -edi ing scena io equi es an al eady ained sys em, online lea ning echniques a e sui ed o his ask. In his wo k, h ee online lea ning algo i hms ha e been p oposed and applied o a neu al machine ansla ion sys- em in a pos -edi ing scena io. They ely on he Passi e-Agg essi e online lea n- ing app oach in which he model is upda ed a e e e y sample in o de o ul il a co ec ness c i e ion while emembe ing p e iously lea ned in o ma ion. The goal is o adap and e ine an al eady ained sys em wi h new samples on- he- ly as he pos -edi ing p ocess akes place (hence, he upda e ime mus be kep unde con ol). Mo eo e , hese new algo i hms a e compa ed wi h well-s ablished online lea ning a ian s o he s ochas ic g adien descen algo i hm. Resul s show im- p o emen s on he ansla ion quali y o he sys em a e applying hese algo- i hms, educing human e o in he pos -edi ing p ocess. Key wo ds: Online Lea ning, Neu al Machine T ansla ion, Passi e-Agg esi e Con en s Con en s Lis o Figu es ii Lis o Tables ii 1 Mo i a ion 1 2 In oduc ion o MT 3 2.1 S a is ical machine ansla ion ...................... 3 2.1.1 Language model ......................... 4 2.1.2 T ansla ion model ......................... 5 2.1.3 Log-linea model .......................... 6 2.2 Assessmen ................................. 6 2.2.1 BLEU ................................ 7 2.2.2 TER ................................. 8 2.2.3 METEOR .............................. 8 3 Neu al machine ansla ion 11 3.1 Modelling language wi h neu al ne wo ks ............... 12 3.1.1 Con inuous wo d ep esen a ion: wo d embedding ..... 12 3.1.2 Dealing wi h sequences: ecu en neu al ne wo ks ..... 13 3.1.3 Dealing wi h con ex : bidi ec ional ecu en neu al ne wo ks 15 3.2 End- o-end ansla ion: encode -decode ................ 15 3.2.1 A en ion model .......................... 17 3.2.2 Decoding ansla ions: beam sea ch .............. 17 4 T aining neu al ne wo ks 19 4.1 The backp opaga ion algo i hm ..................... 19 4.1.1 Backp opaga ion h ough ime ................. 20 4.2 S ochas ic g adien descen ........................ 21 4.2.1 SGD wi h momen um ...................... 22 4.2.2 Adag ad .............................. 22 4.2.3 Adadel a .............................. 23 4.2.4 Adam ................................ 23 5 Online lea ning 25 5.1 Online lea ning amewo k ....................... 25 5.2 Passi e-Agg essi e online lea ning ................... 26 5.3 PA online lea ning applied o NMT ................... 27 5.3.1 Passi e-Agg essi e ia subg adien echniques ........ 28 5.3.2 Passi e-Agg essi e ia p ojec ed subg adien echniques . . 29 5.3.3 Passi e-Agg essi e ia SGD wi h egula iza ion ....... 29 6 Expe imen s and esul s 31 6.1 Expe imen al amewo k ......................... 31 i CONTENTS 6.1.1 Task desc ip ion .......................... 31 6.1.2 So wa e .............................. 32 6.1.3 Co po a .............................. 32 6.1.4 NMT sys em ............................ 33 6.2 Resul s ................................... 34 6.2.1 Hype pa ame e con igu a ion ................. 34 6.2.2 Compa ison be ween OL algo i hms .............. 35 7 Fu u e wo k and conclusions 39 7.1 Fu u e wo k ................................ 39 7.1.1 Using o he loss unc ions .................... 39 7.1.2 Inclusion o OL s a egies in IMT ................ 40 7.2 Conclusions ................................ 40 Bibliog aphy 41 Lis o Figu es 2.1 Wo d aligmen ............................... 5 3.1 LSTM .................................... 14 3.2 Bidi ec ional ecu en neu al ne wo k ................. 15 3.3 Encode -decode ............................. 16 4.1 Backp opaga ion example ........................ 21 6.1 In luence o hype pa ame e s in PA algo i hms ............ 35 6.2 E ec o uning NMT sys ems in e ms o BLEU ............ 36 6.3 E ec o uning NMT sys ems in e ms o TER ............. 37 6.4 E olu ion o NMT sys ems ........................ 38 Lis o Tables 6.1 S a is ics o he Xe ox co pus ....................... 32 6.2 S a is ics o he Emea co pus ....................... 33 6.3 S a is ics o he TED co pus ....................... 33 6.4 G id sea ch alues ............................. 34 6.5 Bes con igu a ion PA ........................... 34 6.6 Op imiza ion hype pa ame e s ..................... 36 6.7 Resul s all asks .............................. 38 ii CHAPTER 1 Mo i a ion We ha e seen in mode n his o y how he imp o emen o anspo a ion and elecommunica ions in as uc u e has had a big impac in he mo emen o peo- ple. Now we a e mo e exposed o o he cul u es, p oduc s and wo ld iews han e e be o e. As echnology imp o es, he access o any kind o in o ma ion has become some hing we can’ li e wi hou . Howe e , di e en cul u es o en mean di e en languages which implies he appea ance o obs acles when he need o communica e is manda o y. Thus, he need o a ool ha helps us o o e come his obs acle a ises. Machine T ansla ion (MT) in es iga es he use o so wa e o ansla e ex o speech om one language in o ano he . I aims a p o iding he bes possible ansla ion wi hou human assis ance. This esea ch ield was bo n in 1947 when Wa en Wea e published a memo- andum s a ing his belie s abou compu e ’s capabili y o ansla e one language in o ano he using c ip og aphy, logic and linguis ic pa e ns. In he nex yea s a lo o p ojec s eme ged in he Uni ed S a es, he So ie Union and Wes e n Eu- ope bu he p ac ical esul s we e disappoin ing (Madsen,2009;Hu chins,2005). This inally led o he publica ion o a epo om a special commi ee o med by he Uni ed S a es named ALPAC (Au oma ic Language P ocessing Ad iso y Commi ee) whe e hey s a ed ha he e was no u u e o good-quali y/cos - e ec i e MT. Du ing he 70’s he e was a new app oach o he MT ield ha elied on he p emise ha a language is based on a se o g amma ical and syn ac ic ules. This ule-based sys em needed a obus and ca e ully designed bilingual dic iona y o linguis ic in o ma ion c ea ed by human expe s. The big e o ha his app oach equi ed was a big issue and i was hen eplaced by he co pus-based sys ems. A he same ime, scien is s ocused on de eloping ools ha would acili a e he ansla ion p ocess a he han eplacing human ansla o s, leading o he de- elopmen o ansla ion memo y (TM) and o he compu e assis ed ansla ion (CAT) (Samson,2005) ools. Co pus-based sys ems ely on a pa allel co pus o sen ences in a sou ce lan- guage and i s ansla ion in a a ge language. Wi h enough numbe o samples, hese sys ems can be ained o in e he ansla ions o new sen ences. Fo his pu pose, s a is ical models a e applied o he ansla ion p ocess as hese a e good a ex ac ing ele an in o ma ion om a se o ( ansla ion) examples. 1 8 In oduc ion o MT Acco ding o he expe imen a ion made in (Papineni e al.,2002), N is se o 4, and wn=1 N. The BLEU sco e anges om 0 o 1 being 1 he closes a candida e ansla ion can be o he e e ence (iden ical in his case). This is he main measu e we will use in he expe imen a ion sec ion. TER The ansla ion e o a e (TER) (Sno e e al.,2006) is ano he measu e used o assess ansla ions. I measu es he minimum numbe o edi s needed o change a hypo hesis so ha i exac ly ma ches he e e ence, no malized by he a e age leng h o he e e ence (i we ha e mo e han one). I we ha e se e al e e ences he TER sco e will co espond o he minimum numbe o edi s needed o ma ch he closes e e ence. The possible edi s include inse ion, dele ion and subs i u ion o single wo ds as well as shi s o wo d sequences. A shi is jus a mo emen o con iguous wo ds wi hin he hypo hesis. An impo an hing o no e is ha all edi s, includ- ing shi s o any numbe o wo ds and dis ance ha e equal cos . In a pos -edi ing scena io his measu e app oxima es he human e o equi ed o co ec a ans- la ion p oduced by a MT sys em. Because we wan o educe he human e o equi ed, we wan o achie e a low TER sco e. METEOR METEOR (Bane jee and La ie,2005) was designed o add ess some o he is- sues ha BLEU in oduced, such as: he lack o ecall, he use o highe o de n- g ams o model wo d o de o he lack o explici wo d-ma ching be ween ans- la ion and e e ence. I is based on he unig am ma ching be ween he machine- p oduced ansla ion and he human-p oduced e e ence ansla ion. Gi en a pai o sen ences, METEOR gene a es an alignmen be ween he wo s ings. In his con ex , an alignmen is a ma ching be ween unig ams, such ha e e y unig am in each s ing maps o ze o o one unig am in he o he s ing. In o de o gene a e ha alignmen he p ocess is di ided in wo phases, each o which has di e en s ages. In he i s phase, based on di e en c i e ia, di e en modules p oduce uni- g am mappings be ween he wo s ings. The "exac " module p oduces ma ches be ween wo unig ams i hey a e exac ly he same. The "po e s em" module p oduces unig am ma ches i hey a e he same a e being s emmed. The "Wo d- Ne synonymy" maps wo s ings i hey a e synonyms. This modules a e o de ed by p io i y and a gi en ma ch will be chosen i s i i is possible o o m an align- men and i i also has highe p io i y han o he ma ches p oduced by lowe p io i y modules. Fo example, an exac ma ch will be always p e e ed. In he second phase, he la ges subse o unig am mappings is selec ed such ha he esul ing se cons i u es an alignmen as de ined abo e. I mo e han one subse cons i u es an alignmen , his me ic selec s he one wi h ewe mapping c osses. I bo h sen ences a e w i en one below he o he and a line is d awn 2.2 Assessmen 9 be ween he ma ching unig ams, a c oss is p oduced when hese lines in e sec wi h ano he mapping. Las ly, in o de o ob ain he METEOR sco e, an ha monic mean o p ecision and ecall is compu ed o e he unig ams and penalized i hese unig ams a e in di e en o de compa ed o he e e ence. A high sco e will mean high quali y ansla ions. CHAPTER 3 Neu al machine ansla ion S a e-o - he-a MT sys ems ha e elied on he ph ase-based app oach o a long ime. Howe e , a new neu al app oach has eme ged, being he i s echnology ha has been able o challenge o me sys ems (Luong and Manning,2015;Jean e al.,2015). Wi h he use o G aphic P ocesso Uni s (GPU), neu al machine ansla ion (NMT) has been able o cope wi h he high compu a ional cos i e- qui es o compe e wi h s a e-o - he-a ph ase-based sys ems (Ben i ogli e al., 2016). This ac eamed wi h he imp o emen s made o he encode -decode a chi- ec u e (explained la e in his chap e ) such as he a en ion model, o he use o ga ed ecu en uni s o cope wi h con ex in long sen ences, made NMT echnol- ogy o ad ance by leaps and bounds. Apa om ha , NMT demons a ed i s powe a he IWSLT12015 e alua ion campaign, whe e one o hese sys ems ou - pe o med he up- o- hen s a e-o - he-a ph ase-based sys ems on he English- Ge man ask, a pai o languages di icul due o he dispa i y in mo phology and syn ax. The nex yea , Google announced ha Google T ansla e made he leap o NMT because neu al ne wo ks inc ease bo h luency and accu acy o i s ansla ions2. Mic oso did he same wi h Mic oso T ansla o s a ing ha neu al ne wo ks be e cap u e he con ex o ull sen ences be o e ansla ing hem, p o iding much highe quali y ansla ions3. Howe e , despi e he easons hese big com- panies ha e o shi o NMT we s ill lack a solid o mal backg ound ha suppo s his echnology. Things like he ac i a ion unc ion choice, he numbe o uni s pe laye o how many laye s o use, a e hings ha need a ho ough s udy o unde s and hem. Fo now, we can only s udy NMT sys ems, by looking a he ansla ions and see wha di e ences hem om he ones p oduced by o me app oaches (Ben i ogli e al.,2016). In his chap e we will gi e an o e iew o he NMT echnology and how i can be used o pe o m he ask a hand. Fi s we will see how we can model he language wi h neu al ne wo ks. Then we will p esen he cu en neu al a chi- 1In e na ional Wo kshop on Spoken Language T ansla ion. 2Found in ansla ion: Mo e accu a e, luen sen ences in Google T ansla e. 3Mic oso T ansla o launching Neu al Ne wo k based ansla ions o all i s speech lan- guages. 11 12 Neu al machine ansla ion ec u e used o ansla e. Las ly, we will see how we pe o m he sea ch ask o ind he bes ansla ion. Modelling language wi h neu al ne wo ks Being able o ep esen language wi h neu al ne wo ks is no an easy ask. Fi s , we need o sol e he issue ha comes om he wo ds being ep esen ed as sym- bols a he han numbe s. Fo his ask, we need a con inuous ep esen a ion (a dense, eal- alued ec o ) ha bo h encapsula es ele an in o ma ion o such symbols and educes he impac o he cu se o dimensionali y. This la e e m e e s o he need o huge numbe o examples when lea ning complex unc ions (such as language). This ac ge s wo se as he numbe o a iables inc eases (i.e. he size o he ocabula y) since he model needs o disc imina e be ween a huge numbe o combina ions o such alues. Once we ha e sol ed he p e ious issue we ace ano he : we need o manage he ac ha sen ences a e no o he same leng h. No only ha , he inpu and ou pu sen ence can be o di e en leng h, so we need a way o p ocess in o ma- ion o a iable leng h. Las ly, we need a way o cap u e con ex ual in o ma ion. All hese h ee p oblems will be add essed in nex subsec ions. Con inuous wo d ep esen a ion: wo d embedding A con inuous ep esen a ion o a wo d is a ec o o ea u es which cha ac e - ize he meaning o ha wo d. To unde s and his, i a human would ha e o ex ac hese ea u es, he would choose g amma ical ea u es such as gende o plu ali y. Wi h neu al ne wo ks we le he lea ning algo i hm disco e his eal- alued ea u es ins ead. The idea is o map e e y wo d in he ocabula y o a low-dimensional con inuous- alued ec o wi h he hope ha simila wo ds ge simila ep esen a ions in he ea u e space. The main goal behind i is o allow he model o gene alize be e o sequences ha a e no seen du ing aining bu whose ea u es a e simila o hose who ha e been seen. Al hough dis ibu ed wo d ep esen a ions we e i s p oposed in he ea ly 1980’s (Hin on,1986) and 1990’s (Cas año and Casacube a,1997), i was no un- il he 2000’s ha i s app oaches using neu al ne wo ks appea ed o model lan- guage. These models we e s udied i s in e ms o eed- o wa d ne wo ks (Ben- gio e al.,2003) and la e in e ms o ecu en neu al ne wo ks (RNN) (Mikolo e al.,2010,2011) and yielded good wo d ep esen a ion ha condensed linguis ic egula i ies in he ec o ep esen a ions (Mikolo e al.,2013b). One o he i s app oaches ha success ully used neu al ne wo ks o model language was he one p oposed in Bengio e al. (2003). This wo k p ojec ed he wo ds in o a eal- alued ec o . Then, i used a eed- o wa d ne wo k ocusing on lea ning bo h he wo d ea u e ec o and he p obabili y dis ibu ion unc- ion o wo d sequences in e ms o hese ea u e ec o s by maximizing he log- likelihood o e a aining da a. The esul s ob ained wi h his app oach demon- s a ed how dis ibu ed ep esen a ion o wo ds could be eamed up wi h neu al ne wo ks o ou pe o m egula n-g am models. 3.1 Modelling language wi h neu al ne wo ks 13 All his wo k demons a ed ou s anding pe o mance in wo d-p edic ion bu also he need o a mo e compu a ionally e icien model (specially he ones in- ol ing RNN). Fo his pu pose, Mikolo e al. (2013a) p oposed a model ha could lea n wo d ep esen a ion much as e wi h highe quali y han p e ious app oaches. I showed he lea ned ela ionship be ween wo ds wi h simple al- geb a by simply adding and sub ac ing ec o s o ob aining o he wo ds. Fo example, he wo d Rome could be ob ained wi h: Pa is - F ance + I aly. Dealing wi h sequences: ecu en neu al ne wo ks When humans ead a sen ence, hey unde s and each wo d based on hei unde - s anding o p e ious wo ds. As you ead, you b ain e ains in o ma ion ha is used o unde s and wha is coming. Recu en neu al ne wo ks model his si - ua ion na u ally. This is a majo ad an age since a a ie y o p oblems depend on an a bi a y sequence o ime-dependen e en s, such as speech ecogni ion, ansla ion o language modelling. They ake in o accoun pas in o ma ion by means o a cycle be ween i s uni s. This allows hem o ha e an in e nal s a e ha holds p esen and pas in o ma- ion. In addi ion, depending on how hese cycles a e ou ed we ind di e en a chi ec u es in he li e a u e. An example o his is he Elman (Elman,1990) and Jo dan (Jo dan,1986) ne wo ks, also known as "simple ecu en ne wo ks" . The Elman a chi ec u e wo ks as ollows: gi en a sequence o ec o s x= x1...xT he ne wo k will p oduce a sequence o ou pu s y=y1...yT. A each ime s ep (Eq. 3.1) he hidden s a e a ime ,s ge s calcula ed based on bo h, he inpu a ime ,x and he hidden s a e o he p e ious ime s ep, s −1. Then, he ou pu (Eq. 3.2 ) ge s calcula ed based on s . s =σs(Wx +Us −1)(3.1) y =σy(Vs )(3.2) In hese equa ions, W,Uand Va e he weigh ma ices o he inpu , ecu en and ou pu connec ions espec i ely, σsis an ac i a ion unc ion such as sigmoid and σyis he ou pu ac i a ion unc ion, being usually, he so max unc ion (Eq. 3.18). Al e na i ely, he Jo dan a chi ec u e is simila bu he ecu en connec ion is achie ed by aking in o accoun he p e ious ou pu a he han he p e ious hidden s a e (Eq. 3.3). s =σs(Wx +Uy −1)(3.3) Howe e , one o he majo laws hese ne wo ks ha e is hei dependency on he comple e pas in o ma ion. This p e en s hem om ocusing only on ecen in o ma ion o on long- e m dependencies, as depic ed in (Bengio e al.,1994). When aining hem, hey also ea u e he so-called anishing g adien p oblem (explained in sec ion 4.1) which p e en s hem om p ope ly lea ning. To add ess his, se e al new ecu en a chi ec u es ha e been p oposed, such as he long sho e m memo y ne wo ks (Hoch ei e and Schmidhube ,1997) 14 Neu al machine ansla ion known as LSTM (Fig. 3.1), o he ga ed ecu en uni s (GRU) (Cho e al.,2014) which ha e he special ai o mi iga ing he anishing g adien p oblem. Bo h o hese a chi ec u es ea u e pa ame ized ga es ha modula e how much inpu in o ma ion is allowed o a ec he hidden s a e, how much pas in o ma ion is le h ough o he nex ime s ep and how much in o ma ion is o go en. These a e ainable pa ame e s ha allow he ne wo k o do his con enien ly. In LSTM uni s, he memo y cell m (Eq. 3.4) depends on he p e ious memo y cell m −1and he new in o ma ion ˜ m (Eq. 3.5) coming om he inpu o he uni . The amoun o in o ma ion used o upda e m om bo h m −1and ˜ m is modula ed by g (Eq. 3.6) and i (Eq. 3.7) which a e he ec o ou pu s o he o ge ga e and inpu ga e whose alues ange om 0 o 1. By doing an elemen - wise mul iplica ion  hey modula e he amoun o in o ma ion ha will be used in he upda e, being 1 comple ely e ain and 0 comple ely o ge . The ou pu o he uni s is modula ed by he ou pu ga e (Eq. 3.8) ge ing Eq. 3.9. m =g m −1+i ˜ m (3.4) ˜ m = anh(WM Ys −1+WM Xx )(3.5) g =σ(WF Ys −1+WF Xx )(3.6) i =σ(WI Ys −1+WI Xx )(3.7) o =σ(WO Ys −1+WO Xx )(3.8) s =o  anh(m )(3.9) In hese equa ions, WO Y,WI Y,WF Yand WM Ya e he ou pu ga e, inpu ga e, o - ge ga e, and memo y cell ecu en weigh ma ices and WO X,WI X,WF X,WM Xa e he ou pu ga e, inpu ga e, o ge ga e, and memo y cell inpu weigh ma ices. σ(·)and anh(·)a e he sigma and hype bolic angen ac i a ion unc ions. Inpu Inpu Fo ge Ou pu ga e ga e ga e x s −1 s −1s −1 s −1 x x x −1 s Cell ˜ s i g m −1 m o Figu e 3.1: LSTM uni 3.2 End- o-end ansla ion: encode -decode 15 Dealing wi h con ex : bidi ec ional ecu en neu al ne wo ks When p ocessing sequences, he e a e imes when u u e e en s a e use ul and p o ide mo e in o ma ion o he ne wo k in o de o be e pe o m he ask a hand. Regula RNN ha e he limi a ion ha u u e in o ma ion can no be eached om he cu en s a e. Bidi ec ional ecu en neu al ne wo ks (BRNN) (Schus e and Paliwal,1997) add ess his issue allowing u u e in o ma ion o be eachable om he cu en s a e. They do his by connec ing wo hidden laye s o opposi e di ec ion o he same ou pu laye , ha ing hen, in o ma ion o pas and u u e e en s. Wi h his a chi ec u e we call o wa d laye he one ha p ocesses he sequence in he posi i e ime di ec ion (le o igh ) and backwa d laye he one ha p o- cesses he sequence in he nega i e ime di ec ion ( igh o le ). A each ime s ep, he ou pu o bo h laye s a e combined (e.g. adding hem up) be o e apply- ing he ou pu ac i a ion unc ion. Fo wa d s a es Backwa d s a es -1 +1 Ou pu neu on g oup Hidden (s a e) neu on g oup Inpu s G oup o weigh s wi h in o ma ion low Figu e 3.2: S uc u e o a bidi ec ional ecu en neu al ne wo k shown un olded in ime o h ee ime s eps. Example bo owed om Schus e and Paliwal (1997). . End- o-end ansla ion: encode -decode We ha e seen ha when a RNN p ocesses a sequence o leng h Ti p oduces an ou pu sequence o he same leng h. In ansla ion howe e , sou ce and a ge sen ence can (and usually a e) o di e en leng hs. To o e come his, wo simila models we e p oposed by Su ske e e al. (2014) and Cho e al. (2014) which elied on he encode -decode app oach. This wo k cons i u es a common amewo k on which la e wo k aims o imp o e i in di e en ways. The encode -decode app oach consis in a wo s ep p ocess ha i s , maps he sou ce sen ence in o a ixed leng h ec o , and second, his ec o is decoded o p oduce he a ge sen ence (possibly o di e en size). I aims o di ec ly model he condi ional ansla ion p obabili y (Eq. 2.1). The inpu o he sys em is a sequence o wo ds x=x1...xJp esen in he ocabula y o he sou ce language Vxin a one-ho ep esen a ion. This means ha we ha e o each wo d xja ocabula y-sized ec o ¯xj∈N|Vx|wi h all elemen s 16 Neu al machine ansla ion se o ze o excep o he one loca ed a he posi ion o xjin Vxwhich is se o one. Each wo d is p ojec ed o a ixed-leng h eal- alued ec o : xj=Es¯ xj(3.10) whe e xj∈Rdis he embedding o wo d xj,Es∈Rd×|Vx|is he sou ce languaje p ojec ion ma ix and d he embedding size. c NULL x1x2xT y1y2yT0 y1yT0−1 Encode Decode Figu e 3.3: Ilus a ion o he encode -decode app oach. The encode is an RNN ha eads each elemen o he sequence o wo d em- beddings. A e eading each elemen in he posi i e di ec ion, he hidden s a e h jo he RNN is upda ed (Eq. 3.11). Once he end-o -sen ence symbol is eached, he hidden s a e o he RNN c(Eq. 3.12), can be seen as a ixed-leng h summa y o he whole inpu sequence. h j= (xj,h j−1)(3.11) c=h J(3.12) whe e is a non-linea unc ion. The decode is ano he RNN which is ained o gene a e he ou pu sequence by p edic ing he nex wo d yigi en he hidden s a e si, he p e ious gene a ed wo d yi−1and c. The hidden s a e o he decode a ime iis compu ed as: si= (si−1,yi−1,c)(3.13) whe e is a non-linea ac i a ion unc ion (LSTM o GRU) ha p oduce he hid- den s a e ec o si. The condi ional p obabili y o he nex wo d yiis app oxi- ma ed (Eq. 3.14) assuming ha i depends on he p e ious wo d (and all p e ious wo ds ha a e in si o some ex en ). In his case, g(·)∈R|Vy|is he so max unc- ion (Eq. 3.18) ha p oduces a ec o o p obabili ies o size |Vy|, which is he size o he ocabula y o he a ge language. ¯yi∈N|Vy|is he one-ho ep esen a ion o he wo d yi.V∈R|Vy|×Lis he weigh ma ix and ϕ(·)is he L-sized ou pu laye o a RNN. P (yi|y1...yi−1,c)≈¯y0 ig(Vϕ(si,yi−1,c)) (3.14) I is impo an o no e ha , du ing aining, we wan he decode o lea n o ou pu he e e ence sen ence. In o de o help he model, we use he eache o c- ing (Pascanu e al.,2013) echnique. This echnique consis o eeding he model 3.2 End- o-end ansla ion: encode -decode 17 wi h he co esponding wo d in he e e ence sen ence a ime s ep i−1 ins ead o he p e iously gene a ed wo d. A en ion model One o he laws ha encode -decode models exhibi is ha hey ha e o con- dense an a bi a y leng h sen ence in o a ixed-leng h ec o c. This is a bo leneck when dealing wi h long sen ences, as no iced by Cho e al. (2014), and causes he sys em o pe o m poo ly in hese si ua ions. In o de o sol e his, Bahdanau e al. (2014) p oposed he so-called a en ion model. The basic idea is o use a di e en con ex ec o depending on he cu en decoding s age. Wi h his idea, Equa ion 3.14 is ew i en as: P (yi|y1...yi−1,c)≈¯y0 ig(Vϕ(si,yi−1,ci)) (3.15) whe e siis he RNN hidden s a e a ime i: si= (si−1,yi−1,ci)(3.16) We can see ha a each ime s ep, we use a di e en con ex ec o ci. This ec o is calcula ed based on a sequence o anno a ions (h1,..., hJ)ex ac ed om he sou ce sen ence. These anno a ions ha e in o ma ion ha s ongly desc ibe he i- h wo d and su oundings. They a e calcula ed by conca ena ing he o - wa d and backwa d hidden s a es o a BRNN. The anno a ion hj= [h j;hb j]∈R2S whe e h jand hb ja e he hidden s a es o he o wa d and backwa d laye s o he BRNN and Sis he size o he hidden s a e (we assume ha hey a e o he same size). The e o e, ciis calcula ed as he weigh ed sum o he anno a ions: ci= J ∑ j=1 αijhj(3.17) whe e αij is compu ed by he so max unc ion: αij =exp(eij) ∑J k=1exp(eik)(3.18) being eij =a(si−1,hj)an alignmen model which sco es how ela ed he inpu s a posi ion jand he ou pu a posi ion ia e. The alignmen model is pa ame ized wi h a eed- o wa d neu al ne wo k which is ained join ly wi h he whole sys- em. Decoding ansla ions: beam sea ch When decoding he ansla ion, he neu al sys em p o ides, a each ime s ep, a se o p obabili ies (wi h he so max unc ion) o each wo d in he ocabula y. Because we a e maximizing equa ion 2.1 and he ou pu p obabili ies o a gi en ime s ep depend o e he p e ious chosen wo d, we canno jus choose he one 24 T aining neu al ne wo ks decaying a e age o pas squa ed g adien s and an exponen ially decaying a e age o pas g adien s m : m =β1·m −1+ (1−β1)·∇`(Θ )(4.11) =β2· −1+ (1−β2)·∇`(Θ )2(4.12) The au ho s no iced ha a ea ly s ages, he algo i hm is biased owa ds ze o. To coun e ac his hey co ec he abo e exp essions as ollows: ˆm =m 1−β 1 (4.13) ˆ = 1−β 2 (4.14) lea ing he upda e ule as: Θ +1=Θ −ρ √ˆ +eˆm (4.15) The de aul alues he au ho s p opose a e 0.9 o β1and 0.999 o β2. CHAPTER 5 Online lea ning Al hough compu a ional esou ces ha e e ol ed o he poin whe e i is ela i ely cheap o ain a Pa e n Recogni ion (PR) sys em, i is s ill a ime-consuming p o- cess ha akes days o e en weeks o comple e. As mo e and mo e da a is a ail- able his p oblem is agg a a ed, specially on sys ems ha in e ac wi h changing en i onmen s o ha equi e a quick esponse a he same ime hey lea n. In hese cases, being able o adap ou sys em wi hou aining i om sc a ch be- comes manda o y. In his chap e we will p esen he online lea ning amewo k which will se e o in oduce ou wo k, emphasizing i s applica ion o a pos -edi ing scena io in MT. Nex we will explain he Passi e-Agg essi e (PA) online lea ning algo i hm. I has he main goal o a oiding he ca as ophic in e e ence (McCloskey and Cohen,1989) p oblem ha h ea ens neu al ne wo k sys ems in which hey end o o ge p e iously lea ned in o ma ion upon lea ning new in o ma ion. Finally, we will p opose h ee new PA-based algo i hms o he ask o NMT. Online lea ning amewo k PR sys ems ha e achie ed an accep able pe o mance in e y complex asks ha in ol es s uc u ed ou pu and ambigui y such as machine ansla ion o image desc ip ion. Al hough hese sys ems can each a high pe o mance on da a simi- la o he one used o ain hem, hei pe o mance apidly d ops when he ask is sligh ly di e en . In addi ion, ge ing enough manually anno a ed da a o a spe- ci ic domain in o de o ain a whole sys em migh no be possible. The e o e, i is use ul o ain wi h gene ic ou -o -domain da a and hen une he sys em wi h domain-speci ic da a. Because a comple e e aining o he sys em migh be in easible, online lea ning echniques (Blum,1998) a e chosen o his ask. These echniques a e pa icula ly use ul in he compu e assis ed ansla ion (CAT) and in e ac i e machine ansla ion (IMT) pa adigms, whe e human ans- la o s wo k join ly wi h machine ansla o s in o de o e icien ly ob ain high quali y ansla ions. In a pos -edi ing scena io, he MT sys em ansla es a sen- ence and p opose his as a hypo hesis o a human ansla o which hen co ec s he sen ence (i needed). Once a sen ence is co ec ed wha we ha e is a new pai o sen ences ha can be used o ain he sys em. By using hese new ain- 25 26 Online lea ning ing samples o ain ou model, we p og essi ely make he pos -edi ing p ocess mo e e icien . Fi s , by making he sys em lea n om i s own e o s and second, by easing he wo k o he human ansla o as i will ha e o co ec less e o s. P io wo k in his a ea has also shown ha human ansla o s become mo e p o- duc i e as he MT quali y imp o es (Ta sumi,2009). The applica ion o such echniques has been s udied ho oughly (Ma ínez- Gómez e al.,2012;La ie,2014;Ma ínez-Gómez e al.,2011) in classical ph ase- based SMT sys ems. Resul s show ha signi ican imp o emen s can be achie ed by adap ing he sys em by means o online lea ning in e ms o he e o he human would need o co ec he hypo hesis p o ided. Howe e , he e has been no published wo k ( o he bes o ou knowledge) ela ed o he applica ion o hese echniques in NMT sys ems. This wo ks aims a p o iding new algo i hms o apply online lea ning o hese sys ems. Passi e-Agg essi e online lea ning A a ie y o online lea ning me hods ha e been p oposed in li e a u e (Lu e al., 2016). A classical OL me hod is he Pe cep on algo i hm (Rosenbla ,1958) which upda es he model by adding a misclassi ied example (pa ame ized wi h a ac- o ) o he cu en pa ame e s. F om his wo k, a lo o new online lea ning algo- i hms ha e been de eloped based on he maximum ma gin c i e ion (C amme and Singe ,2003;Gen ile,2001;Ki inen e al.,2004;Li and Long,2000) which ies o sepa a e he classes as much as possible. One impo an echnique ha alls in his ca ego y is he Passi e-Agg essi e (PA) online lea ning me hod (C amme e al.,2006). This has been p o ed as a e y success ul and popula online lea n- ing echnique o sol ing many eal-wo ld applica ions. The idea behind he PA algo i hm is e y simple. When a new sample a i es i is classi ied and a loss is ob ained, measu ing he deg ee o which he p edic ion is w ong. The idea now is o ind a new se o pa ame e s (weigh s) ha make he classi ie o co ec ly classi y he sample (agg essi eness) bu by emaining ela i ely close o he p e ious se o pa ame e s (passi eness). We will ou line he PA app oach in a bina y classi ica ion p oblem as explained in C amme e al. (2006). In a bina y classi ica ion p oblem, each sample x ∈Rdhas a unique label y ∈ {+1, −1}associa ed. We assume ha he classi ica ion unc ion is based on a ec o o weigh s Θ∈Rdwhich ake he o m o sign(Θ·x). The magni ude |Θ·x|is in e p e ed as he deg ee o con idence in he p edic ion. We e e o he e m y (Θ ·x )as he ma gin calcula ed a ime . Whene e he ma gin is posi i e, he sample has been classi ied co ec ly. Howe e , we wan he classi ie o achie e a co ec classi ica ion wi h some ma gin (in C amme e al. (2006) his ma gin is se o 1). Then we de ine he loss ` ha su e s he classi ie whene e he samples is classi ied inco ec ly: `(Θ;(x,y)) = (0 i y(Θ·x)≤1 1−y(Θ·x)o he wise (5.1) 5.3 PA online lea ning applied o NMT 27 In o de o p og essi ely lea n he weigh ec o Θ, he algo i hm mus up- da e i a e e e y sample. A ime , he new weigh ec o Θ +1is calcula ed by sol ing his cons ained op imiza ion: Θ +1=a gmin Θ∈Rd 1 2kΘ−Θ k2s. . `(Θ;(x ,y )) = 0 (5.2) In his algo i hm, i he loss is 0, he op imal solu ion is Θ +1=Θ . On he o he hand, when he loss is g ea e han 0, he algo i hm o ces Θ +1 o sa is y he cons ain `(Θ +1;(x ,y )) = 0 while being as close as possible o he p e ious weigh ec o o p ese e knowledge o p e ious samples. I we de i e Eq. 5.2, he upda e ule o he weigh ec o is: Θ +1=Θ +τ y x whe e τ =`(Θ ;(x ,y )) kx k2(5.3) Howe e , his upda e ule is oo agg essi e and i migh esul in upda ing he model in such a way ha in o de o mee he cons ain , i p oduces a lo o e o s in la e samples because he model was changed d as ically. The e o e, a gen le upda e s a egy is equi ed. To do his, he au ho s in oduce a non-nega i e slack ξ a iable in o he op imiza ion p oblem de ined in Eq. 5.2: Θ +1=a gmin Θ∈Rd 1 2kΘ−Θ k2+Cξs. . `(Θ;x ,y )) ≤ξand ξ≥0 (5.4) The pa ame e Ccon ols he in luence o he slack a iable ξ. The au ho s coin he C a iable as he agg essi eness pa ame e o he algo i hm. The upda e ule in his case is he same Θ +1=Θ +τ y x bu τ changes o: τ =min{C,`(Θ ;(x ,y )) kx k2}(5.5) PA online lea ning applied o NMT In his sec ion we p opose a new e sion o PA-based SGD. The de elopmen o his wo ks is inspi ed by he applica ion o he MIRA (a PA-based) algo i hm (C amme and Singe ,2003) used in adi ional SMT o une he weigh s o he log-linea model. Also, he wo k de elop in Ma ínez-Gómez e al. (2012) o he ask o online adap a ion gi es us a eel o wha can be achie ed when adap ing a MT sys em wi h OL. As any o he PA echnique, ou algo i hm aims o pe o m he minimum modi ica ion o he model pa ame e s while ul illing a co ec ness c i e ion. In ou case, we use he loss unc ion o he neu al model o e iciency easons (bu we could ha e used ano he loss unc ion such as BLEU). 28 Online lea ning Le h be he hypo hesis gene a ed by he NMT sys em (using he pa ame e s Θ a ime ) o he sou ce sen ence x . We conside ha Θ is inco ec i he model assigns a lowe p obabili y o he a ge e e ence sen ence y han o h : pΘ (y |x )<pΘ (h |x )(5.6) When ha happens, we wan o sea ch o a b Θsuch ha b Θis close o Θ and pΘ (y |x )>pΘ (h |x ). This is exp essed wi h he loss unc ion `: `(b Θ,x ,y ,h ) = log pb Θ(h |x )−log pb Θ(h |x )(5.7) Wi h his loss unc ion we a e measu ing how la ge is he gap be ween he p obabili y gi en o he hypo hesis and he p obabili y gi en o he e e ence. By op imizing his unc ion we educe his gap wi h he idea ha p og essi ely, he hypo hesis will be close (i.e. simila ) o he e e ence sen ence. Wi h he goal o op imizing his unc ion, h ee a ia ions o PA-based algo i hms a e de eloped. Passi e-Agg essi e ia subg adien echniques In o de o ind b Θ, he p oblem can be o mula ed as: b Θ=a gmin Θ 1 2kΘ−Θ k2+Cξs. . `(b Θ,x ,y ,h )≤ξand ξ≥0 (5.8) being C he pa ame e ha con ols he agg essi eness o he algo i hm and ξ a slack a iable (as discussed in Sec ion 5.2). F om his equa ion we ha e ξ≥ max(0, `(b Θ,x ,y ,h )) and hen, we de ine F as he unc ion o op imize: F (Θ,x ,y ,h ) = 1 2kΘ−Θ k2+Cmax(0, `(b Θ,x ,y ,h )) (5.9) aiming o ind he se o pa ame e s ha minimize his unc ion, i.e.: b Θ=a gmin Θ F (Θ,x ,y ,h )(5.10) Because we a e op imizing a unc ion (F ) wi h discon inui y poin s (because he max(·) unc ion has no de i a i e a 0) we ha e o use a subg adien me hod (Sho e al.,2003). This esul s in he de i a i e ∂ΘF (Θ,x ,y ,h )being: ∂ΘF (Θ,x ,y ,h ) =      Θ−Θ −C∇`(Θ,x ,y ,h )`(Θ,x ,y ,h )<0 Θ−Θ `(Θ,x ,y ,h )>0 [Θ−Θ ,Θ−Θ −C∇`(Θ,x ,y ,h )] `(Θ,x ,y ,h ) = 0 (5.11) 5.3 PA online lea ning applied o NMT 29 whe e, on he discon inui y poin (`(Θ,x ,y ,h ) = 0), we choose a alue in he in- e al [Θ−Θ ,Θ−Θ −C∇`(Θ,x ,y ,h )]. We assume Θ−Θ when `(Θ,x ,y ,h ) = 0. Finally, we pe o m he upda e as: Θk+1=Θk−ρ·∂ΘF (Θ,x ,y ,h )|Θk(5.12) whe e ρis he lea ning a e, ∂ΘF is he subg adien o F wi h espec o Θ. We ini ialize Θk=0=Θ and he upda e is pe o med k imes pe sample. We deno e his PA ia subg adien echnique upda e ule as PAS. Because we may wan o achie e a highe p obabili y o he e e ence sen ence bu wi h some ma gin m:pΘ(y |x ) + m>pΘ(h |x )we can ew i e he subg a- dien ∂ΘF wi h he discon inui y poin being m a he han ze o (as de ined in Eq. 5.11). Passi e-Agg essi e ia p ojec ed subg adien echniques An ex ension o he PAS me hod is he p ojec ed PA subg adien me hod (PPAS) in which he op imiza ion p oblem is e o mula ed (Boyd e al.,2003). We de ine G (Θ,x ,y ,h )as max(0, `(b Θ,x ,y ,h )). Then, Eq. 5.8 can be ew i en as: b Θ=a gmin Θ G (Θ,x ,y ,h )s. . kΘ−Θ k2≤C(5.13) In his case, ∂ΘG (Θ,x ,y ,h )is de ined as: ∂ΘG (Θ,x ,y ,h ) =      ∇`(Θ,x ,y ,h )`(Θ,x ,y ,h )<0 0`(Θ,x ,y ,h )>0 [0, ∇`(Θ,x ,y ,h )] `(Θ,x ,y ,h ) = 0 (5.14) and he upda e ule is pe o med in a wo s ep p ocess. Fi s , we calcula e he in e media e weigh upda e ¯ Θk+1as: ¯ Θk+1=Θk−ρ∂ΘG (Θ,x ,y ,h )|Θk(5.15) and second, we apply he p ojec ion ope a o , being Θk+1: Θk+1=¯ Θk+1−Θ k¯ Θk+1−Θ kC+Θ (5.16) As in he p e ious case, we ini ialize Θk=0=Θ . Passi e-Agg essi e ia SGD wi h egula iza ion We can also op imize `(Θ,x ,y ,h )by means o SGD bu wi h a egula iza ion e m, inspi ed by he PA p inciple o upda ing he model wi h new pa ame e s 30 Online lea ning ha a e close o he cu en ones. We will e e o his algo i hm as PAR. Wi h his in mind, we can de ine he minimiza ion p oblem as: b Θ=a gmin Θ (`(Θ,x ,y ,h ) + C 2kΘ−Θ k2)(5.17) Because he minimiza ion has no discon inui y poin s we don’ need o apply a subg adien me hod. The upda e ule hen, esul s in: Θk+1=Θk−ρ·(∇`(Θ,x ,y ,h )|Θk+C(Θk−Θ )) (5.18) As usual, Θk=0=Θ . We obse e ha when k=1 he upda e ule is plain SGD wi h no egula iza ion e m: Θk+1=Θk−ρ·(∇`(Θ,x ,y ,h )|Θk+ : 0 C(Θk−Θ ) ) (5.19) in ha case, we in oduce a small Gaussian noise ν o he model pa ame e s in o de o a oid eaching local minima: Θk+1=Θk−ρ·(∇`(Θ,x ,y ,h )|Θk+Cν)(5.20) CHAPTER 6 Expe imen s and esul s A e he p oposal o h ee PA-based algo i hms in Sec ion 5.3, we will p oceed o i s es ing and compa ison wi h o he well-known algo i hms ( hose desc ibed in Sec ion 4.2). In his sec ion we will desc ibe he expe imen al se up used o conduc he expe imen a ion. Fi s , we will desc ibe he ask in which we wan o e alua e ou algo i hms. Then we will desc ibe he NMT sys em used and he ools needed o wo k wi h i . A e desc ibing he co po a used, we will compa e and discuss he esul s ob ained. Expe imen al amewo k Be o e p esen ing he esul s, we will p esen he ask a hand and desc ibe he ools used o de elop he men ioned algo i hms and he co po a used o assess hem. Task desc ip ion In o de o es hese algo i hms we will simula e (o he wise i would be oo cos ly) a pos -edi ing scena io. We will simula e his scena io in h ee di e en co pus o asks (de ailed in nex subsec ions). Fi s , we will ain h ee neu al models using o each one he aining pa i ion o i s co esponding da ase . Once ained, hese models will cons i u e he baseline models o each ask. Then, we will e ine hese models wi h hei espec i e es o de elopmen se o he co pus in which hey we e ained. Fo a gi en sou ce sen ence, he sys- em p oduces i s ansla ion ( he sys em’s hypo hesis). This ansla ion is s o ed o la e e alua ion. Then we ake he pos -edi ed sen ence (in his case, he e e - ence sen ence) and upda e he sys em wi h i . This is epea ed un il he ull se o sen ences a e ansla ed. We e alua e he s o ed hypo heses and compa e hem wi h he ones p oduced by he baseline sys em. Ou hope is o see an imp o e- men in he quali y o hese ansla ion. Wi h his p ocedu e we wan o see up o wha ex en he OL algo i hm can e ine he baseline sys em. In o de o measu e he impac o each algo i hm in he pos -edi ing e o educ ion we use he TER me ic. In o de o assess he quali y o ansla ions we will use he BLEU and METEOR me ics. 31 32 Expe imen s and esul s So wa e We ha e de eloped ou wo k using he NMT-Ke as oolki 1(Pe is,2017). This (s ill in p og ess) so wa e p o ides ools o he NMT ask, such as: beam sea ch decoding, suppo o GRU and LSTM ne wo ks, unknown wo d eplacemen , use o p e ained wo d embedding ec o s an mo e. This so wa e is based on a Py hon lib a y called Ke as2(Cholle e al.,2015). Ke as o e s an abs ac ion laye on op o Theano3(Theano De elopmen Team, 2016) (a lib a y o de ining and e alua ing ma hema ical exp essions in ol ing mul i-dimensional a ays) and eases he ask o building and aining di e en neu al ne wo k a chi ec u es. I has implemen ed u ili ies o sa e and load neu al ne wo k models, moni o he aining p ocess and use p ede ined NN a chi ec- u es. All he algo i hms lis ed in Sec ion 5.3 ha e been implemen ed in Ke as and in eg a ed in he NMT-Ke as oolki . Co po a In his wo k, we es all he algo i hm in h ee di e en co pus: Xe ox, Emea and TED. We use he s anda d pa i ions and he English-F ench language pai o all expe imen s. The ex is kep ue case and i is okenized using he sc ip om Moses 4. In he nex subsec ions we will desc ibe hese co pus. Xe ox The Xe ox co pus (Ba achina e al.,2009) consis s in ansla ions o use manuals o Xe ox p in e s. Table 6.1 shows all in o ma ion ega ding he pa i ions o he English-F ench language pai . T aining De elopmen Tes Sen ences 52k 984 994 Vocabula y size English 14k 1.8k 1.7k F ench 15.5k 1.9k 1.8k Wo ds English 615k 10.9k 11.1k F ench 676k 11.7k 11.8k Table 6.1: S a is ics o he Xe ox co pus. In he able a e collec ed he numbe o sen ences and ocabula y size o each pa i ion and language. k s ands o housands 1h ps://gi hub.com/l apeab/nm -ke as/ 2h ps://ke as.io/ 3h p://deeplea ning.ne /so wa e/ heano/ 4h p://www.s a m .o g/moses/ 6.1 Expe imen al amewo k 33 Emea The Emea pa allel co pus (Tiedemann,2009) is made ou o documen s om he Eu opean Medicine Agency. Table 6.2 shows all in o ma ion ega ding he pa i- ions o he English-F ench language pai . T aining De elopmen Tes Sen ences 319k 500 1k Vocabula y size English 52k 2.8k 4.5k F ench 59.3k 2.9k 4.5k Wo ds English 4.2M 10.3k 21.4k F ench 4.8M 12.3k 25.7k Table 6.2: S a is ics o he Emea co pus. In he able a e collec ed he numbe o sen ences and ocabula y size o each pa i ion and language. k s ands o housands and M o millions TED The TED co pus (Fede ico e al.,2011) ga he s ansc ibed TED alks. Table 6.3 shows all in o ma ion ega ding he pa i ions o he English-F ench language pai . T aining De elopmen Tes Sen ences 159k 887 1.7k Vocabula y size English 46.7k 3.4k 4.9k F ench 58.2k 3.9k 4.9k Wo ds English 2.17M 21.3k 33.5k F ench 2.3M 21.8k 35.7k Table 6.3: S a is ics o he TED co pus. In he able a e collec ed he numbe o sen ences and ocabula y size o each pa i ion and language. k s ands o housands NMT sys em Fo each co pus, we ain a neu al model o e he aining pa i ion. Such model (simila o he one depic ed in Bahdanau e al. (2014)) consis s in an encode - decode LSTM ne wo k equipped wi h he a en ion mechanism. We made use o single-laye ed LSTM due o ime limi a ions. The size o he hidden s a e o each LSTM, he wo d embedding size and he a en ion mechanism laye is 512 (ou choice is based on he expe imen a ions ca ied ou in B i z e al. (2017)). The baseline sys ems ha e been ained wi h he Adadel a algo i hm wi h he de aul pa ame e s. Du ing aining, Gaussian noise (G a es,2011) and laye no maliza ion (Ba e al.,2016) a e applied as egula iza ion me hods. We also ea ly s opped he aining i he BLEU on he de elopmen se did no imp o e in 100,000 upda es. The size o he beam is 6 in all expe imen s. 40 Fu u e wo k and conclusions Inclusion o OL s a egies in IMT The IMT scena io is igh ly ela ed wi h he pos -edi ing scena io. IMT ies o collabo a e wi h he human on he ansla ion ask wi h he goal o minimizing human e o . O en, he use only needs o poin ou he posi ion in he sen ence in which i is w ongly ansla ed and he sys em will ou pu an al e na i e su - ix s a ing om ha poin . IMT is an ac i e ield and ecen wo k has s udied i s applica ion on NMT sys ems (Knowles and Koehn,2016;Pe is e al.,2017b). As in he pos -edi ing scena io, being able o apply OL echniques in o an IMT amewo k migh be bene icial in o de o de elop mo e adap i e and p oduc i e MT sys ems. Conclusions In his wo k, h ee OL algo i hms ha e been p oposed and applied o an NMT sys em. They ely on he Passi e-Agg essi e OL app oach in which he MT model is upda ed a e e e y sample in o de o ul il a co ec ness c i e ion while emembe ing p e iously lea ned in o ma ion. Because hese algo i hms ha e a high numbe o hype pa ame e s o une, an ex ensi e g id sea ch has been ca - ied ou in o de o ind he bes combina ion. In addi ion, a s udy has been made o see how he di e en pa ame e s a ec each algo i hm’s pe o mance. Resul s show ha , al hough hey depend on se e al hype pa ame e s, hose c i ical o he well beha iou o he algo i hms a e, in e e y case, he agg essi eness hy- pe pa ame e and he lea ning a e. O he pa ame e s such as he ma gin o he numbe o i e a ions pe sample o e mino imp o emen s in he o e all pe o - mance o he algo i hm. These algo i hms ha e been es ed in a pos -edi ing scena io whe e an al eady ained NMT sys em has o be e ined as soon as new samples a e a ailable. Each new sample ha a i es o he sys em is ansla ed and co ec ed (we use he e e ence sen ence ins ead). A loss unc ion is compu ed in o de o measu e he dis ance be ween bo h sen ences. Once we ha e his, we use i o adap he sys em, aiming o a oid he commi ed mis akes. In o de o e alua e how well hese new algo i hms beha e, a compa ison be ween hese algo i hms and well- known OL g adien descen a ian s has been ca ied ou . Resul s show ha OL echniques help o e ine and adap he sys em on- he- ly, inc easing he quali y o ansla ions and educing he human e o in a pos -edi ing scena io. We ha e ound ha he p oposed PA-based algo i hms o e a compe i i e pe o mance: hey imp o e he quali y o ansla ions in all asks. Ne e heless, adap i e SGD- based algo i hms pe o med gene ally be e . Finally, pa o he wo k de eloped in his hesis has been submi ed o he 2017 Con e ence on Empi ical Me hods in Na u al Language P ocessing (Pe is e al.,2017a). Mo eo e , we plan o submi an a icle o a jou nal (ye o be de e - mined). Bibliog aphy Ba, J. L., Ki os, J. R., and Hin on, G. E. (2016). Laye no maliza ion. a Xi p ep in a Xi :1607.06450. Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neu al machine ansla ion by join ly lea ning o align and ansla e. a Xi p ep in a Xi :1409.0473. Bane jee, S. and La ie, A. (2005). Me eo : An au oma ic me ic o m e alua- ion wi h imp o ed co ela ion wi h human judgmen s. In P oceedings o he acl wo kshop on in insic and ex insic e alua ion measu es o machine ansla ion and/o summa iza ion, olume 29, pages 65–72. Ba achina, S., Bende , O., Casacube a, F., Ci e a, J., Cubel, E., Khadi i, S., La- ga da, A., Ney, H., Tomás, J., Vidal, E., e al. (2009). S a is ical app oaches o compu e -assis ed ansla ion. Compu a ional Linguis ics, 35:3–28. Bengio, Y., Ducha me, R., Vincen , P., and Jau in, C. (2003). A neu al p obabilis ic language model. Jou nal o machine lea ning esea ch, 3:1137–1155. Bengio, Y., Sima d, P., and F asconi, P. (1994). Lea ning long- e m dependencies wi h g adien descen is di icul . IEEE ansac ions on neu al ne wo ks, 5(2):157– 166. Ben i ogli, L., Bisazza, A., Ce olo, M., and Fede ico, M. (2016). Neu al e - sus ph ase-based machine ansla ion quali y: a case s udy. a Xi p ep in a Xi :1608.04631. Bishop, C. M. (2006). Pa e n ecogni ion and machine lea ning. sp inge . Blum, A. (1998). On-line algo i hms in machine lea ning. In Online algo i hms, pages 306–325. Sp inge . Bo ou, L. (1991). S ochas ic g adien lea ning in neu al ne wo ks. P oceedings o Neu o-Nımes, 91. Boyd, S., Xiao, L., and Mu apcic, A. (2003). Subg adien me hods. lec u e no es o EE392o, S an o d Uni e si y, Au umn Qua e , 2004. B i z, D., Goldie, A., Luong, T., and Le, Q. (2017). Massi e explo a ion o neu al machine ansla ion a chi ec u es. a Xi p ep in a Xi :1703.03906. B own, P. F., Cocke, J., Pie a, S. A. D., Pie a, V. J. D., Jelinek, F., La e y, J. D., Me ce , R. L., and Roossin, P. S. (1990). A s a is ical app oach o machine ans- la ion. Compu a ional linguis ics, 16(2):79–85. 41 42 BIBLIOGRAPHY B own, P. F., Pie a, V. J. D., Pie a, S. A. D., and Me ce , R. L. (1993). The ma h- ema ics o s a is ical machine ansla ion: Pa ame e es ima ion. Compu a ional linguis ics, 19(2):263–311. Cas año, A. and Casacube a, F. (1997). A connec ionis app oach o machine ansla ion. In Fi h Eu opean Con e ence on Speech Communica ion and Technology. Chen, S. F. and Joshua, G. (1998). An empi ical s udy o smoo hing echniques o language modeling. Technical epo , Ha a d Compu e Science G oup. Chiang, D. (2012). Hope and ea o disc imina i e aining o s a is ical ansla- ion models. Jou nal o Machine Lea ning Resea ch, 13:1159–1187. Cho, K., Van Me iënboe , B., Gulceh e, C., Bahdanau, D., Bouga es, F., Schwenk, H., and Bengio, Y. (2014). Lea ning ph ase ep esen a ions using nn encode - decode o s a is ical machine ansla ion. a Xi p ep in a Xi :1406.1078. Cholle , F. e al. (2015). Ke as. h ps://gi hub.com/ cholle /ke as . Chow, Y.-L. and Schwa z, R. (1989). The n-bes algo i hm: An e icien p ocedu e o inding op n sen ence hypo heses. In P oceedings o he wo kshop on Speech and Na u al Language, pages 199–202. Associa ion o Compu a ional Linguis- ics. Chung, J., Gulceh e, C., Cho, K., and Bengio, Y. (2014). Empi ical e alua ion o ga ed ecu en neu al ne wo ks on sequence modeling. a Xi p ep in a Xi :1412.3555. C amme , K., Dekel, O., Keshe , J., Shale -Shwa z, S., and Singe , Y. (2006). On- line passi e-agg essi e algo i hms. Jou nal o Machine Lea ning Resea ch, 7:551– 585. C amme , K. and Singe , Y. (2003). Ul aconse a i e online algo i hms o mul- iclass p oblems. Jou nal o Machine Lea ning Resea ch, 3:951–991. Duchi, J., Hazan, E., and Singe , Y. (2011). Adap i e subg adien me hods o on- line lea ning and s ochas ic op imiza ion. Jou nal o Machine Lea ning Resea ch, 12:2121–2159. Elman, J. L. (1990). Finding s uc u e in ime. Cogni i e science, 14(2):179–211. Fede ico, M., Ben i ogli, L., Paul, M., and S üke , S. (2011). O e iew o he IWSLT e alua ion campaign. In P oceedings o he In e na ional Wo kshop on Spo- ken Language T ansla ion, pages 11–27. Gen ile, C. (2001). A new app oxima e maximal ma gin classi ica ion algo i hm. Jou nal o Machine Lea ning Resea ch, 2:213–242. Golik, P., Doe sch, P., and Ney, H. (2013). C oss-en opy s. squa ed e o ain- ing: a heo e ical and expe imen al compa ison. In P oceedings o he In e speech, pages 1756–1760. G a es, A. (2011). P ac ical a ia ional in e ence o neu al ne wo ks. In Ad ances in Neu al In o ma ion P ocessing Sys ems, pages 2348–2356. BIBLIOGRAPHY 43 Guo, J. (2013). Backp opaga ion h ough ime. Unpubl. ms., Ha bin Ins i u e o Technology. Hin on, G. E. (1986). Lea ning dis ibu ed ep esen a ions o concep s. In P o- ceedings o he eigh h annual con e ence o he cogni i e science socie y, olume 1, page 12. Amhe s , MA. Hoch ei e , S., Bengio, Y., F asconi, P., and Schmidhube , J. (2001). G adien low in ecu en ne s: he di icul y o lea ning long- e m dependencies. Hoch ei e , S. and Schmidhube , J. (1997). Long sho - e m memo y. Neu al com- pu a ion, 9(8):1735–1780. Hu chins, J. (2005). The his o y o machine ansla ion in a nu shell. Re ie ed Decembe , 20:2009. Jean, S., Fi a , O., Cho, K., Memise ic, R., and Bengio, Y. (2015). Mon eal neu al machine ansla ion sys ems o wm ’15. In P oceedings o Empi ical Me hods in Na u al Language P ocessing, pages 134–140. Jo dan, M. I. (1986). A ac o dynamics and pa allellism in a connec ionis se- quen ial machine. Kingma, D. and Ba, J. (2014). Adam: A me hod o s ochas ic op imiza ion. a Xi p ep in a Xi :1412.6980. Ki inen, J., Smola, A. J., and Williamson, R. C. (2004). Online lea ning wi h ke - nels. IEEE ansac ions on signal p ocessing, 52:2165–2176. Knese , R. and Ney, H. (1995). Imp o ed backing-o o m-g am language mod- eling. In Acous ics, Speech, and Signal P ocessing, olume 1, pages 181–184. IEEE. Knowles, R. and Koehn, P. (2016). Neu al in e ac i e ansla ion p edic ion. In P oceedings o he Con e ence o he Associa ion o Machine T ansla ion in he Ame - icas (AMTA), olume 1, pages 107–120. Koehn, P. (2004). S a is ical signi icance es s o machine ansla ion e alua ion. In P oceedings o Empi ical Me hods in Na u al Language P ocessing, pages 388– 395. Koehn, P. (2010). S a is ical machine ansla ion. La ie, M. D. C. D. A. (2014). Lea ning om pos -edi ing: Online model adap a- ion o s a is ical machine ansla ion. P oceedings o he 14 h Con e ence o he Eu opean Chap e o he Associa ion o Compu a ional Linguis ics (EACL 14), page 395. LeCun, Y., Bengio, Y., and Hin on, G. (2015). Deep lea ning. Na u e, 521:436–444. Li, Y. and Long, P. M. (2000). The elaxed online maximum ma gin algo i hm. In Ad ances in neu al in o ma ion p ocessing sys ems, pages 498–504. Lu, J., Zhao, P., and Hoi, S. C. (2016). Online passi e-agg essi e ac i e lea ning. Machine Lea ning, 103:141–183. 44 BIBLIOGRAPHY Luong, M.-T. and Manning, C. D. (2015). S an o d neu al machine ansla ion sys- ems o spoken language domains. In P oceedings o he In e na ional Wo kshop on Spoken Language T ansla ion. Madsen, M. W. (2009). The limi s o machine ansla ion ( hesis). Cen e o Lan- guage Technology, Uni . o Copenhagen, Copenhagen. Ma ínez-Gómez, P., Sanchis-T illes, G., and Casacube a, F. (2011). Online lea n- ing ia dynamic e anking o compu e assis ed ansla ion. Compu a ional Linguis ics and In elligen Tex P ocessing, pages 93–105. Ma ínez-Gómez, P., Sanchis-T illes, G., and Casacube a, F. (2012). Online adap- a ion s a egies o s a is ical machine ansla ion in pos -edi ing scena ios. Pa e n Recogni ion, 45(9):3193–3203. McCloskey, M. and Cohen, N. J. (1989). Ca as ophic in e e ence in connec ionis ne wo ks: The sequen ial lea ning p oblem. Psychology o lea ning and mo i a- ion, 24:109–165. Mikolo , T., Chen, K., Co ado, G., and Dean, J. (2013a). E icien es ima ion o wo d ep esen a ions in ec o space. a Xi p ep in a Xi :1301.3781. Mikolo , T., Ka a iá , M., Bu ge , L., Ce nock` y, J., and Khudanpu , S. (2010). Re- cu en neu al ne wo k based language model. In P oceedings o he In e speech, olume 2, page 3. Mikolo , T., Komb ink, S., Bu ge , L., ˇ Ce nock` y, J., and Khudanpu , S. (2011). Ex ensions o ecu en neu al ne wo k language model. In Acous ics, Speech and Signal P ocessing, pages 5528–5531. IEEE. Mikolo , T., Yih, W.- ., and Zweig, G. (2013b). Linguis ic egula i ies in con in- uous space wo d ep esen a ions. In P oceedings o he 2013 Con e ence o he No h Ame ican Chap e o he Associa ion o Compu a ional Linguis ics: Human Language Technologies (NAACL-HLT-2013), olume 13, pages 746–751. Och, F. J. (2003). Minimum e o a e aining in s a is ical machine ansla ion. In P oceedings o he Annual Mee ing o he Associa ion o Compu a ional Linguis ics, pages 160–167. Papineni, K., Roukos, S., Wa d, T., and Zhu, W.-J. (2002). Bleu: a me hod o au oma ic e alua ion o machine ansla ion. In P oceedings o he 40 h annual mee ing on associa ion o compu a ional linguis ics, pages 311–318. Pascanu, R., Mikolo , T., and Bengio, Y. (2013). On he di icul y o aining ecu - en neu al ne wo ks. In e na ional Con e ence on Machine Lea ning (3), 28:1310– 1318. Penning on, J., Soche , R., and Manning, C. D. (2014). Glo e: Global ec o s o wo d ep esen a ion. In P oceedings o Empi ical Me hods in Na u al Language P ocessing, olume 14, pages 1532–1543. Pe is, Á. (2017). NMT-Ke as. h ps://gi hub.com/l apeab/nm -ke as . Gi Hub eposi o y. BIBLIOGRAPHY 45 Pe is, Á., Ceb ián, L., and Casacube a, F. (2017a). Online lea ning o neu al machine ansla ion pos -edi ing. a Xi p ep in a Xi :1706.03196. Pe is, Á., Domingo, M., and Casacube a, F. (2017b). In e ac i e neu al machine ansla ion. Compu e Speech & Language, 45:201–220. Qian, N. (1999). On he momen um e m in g adien descen lea ning algo i hms. Neu al ne wo ks, 12(1):145–151. Rosenbla , F. (1958). The pe cep on: A p obabilis ic model o in o ma ion s o - age and o ganiza ion in he b ain. Psychological e iew, 65:386. Rumelha , D. E., Hin on, G. E., and Williams, R. J. (1988). Lea ning ep esen a- ions by back-p opaga ing e o s. Cogni i e modeling, 5(3):1. Samson, R. (2005). Compu e -assis ed ansla ion. T aining o he New Millennium: Pedagogies o ansla ion and in e p e ing, 60:101. Schus e , M. and Paliwal, K. K. (1997). Bidi ec ional ecu en neu al ne wo ks. IEEE T ansac ions on Signal P ocessing, 45(11):2673–2681. Shen, S., Cheng, Y., He, Z., He, W., Wu, H., Sun, M., and Liu, Y. (2016). Minimum isk aining o neu al machine ansla ion. In P oceedings o he 54 h Annual Mee ing o he Associa ion o Compu a ional Linguis ics, pages 1683–1692. Sho , N., Zhu benko, N., Likho id, A., and S e syuk, P. (2003). Algo i hms o nondi e en iable op imiza ion: De elopmen and applica ion. Cybe ne ics and Sys ems Analysis, 39:537–548. Sno e , M., Do , B., Schwa z, R., Micciulla, L., and Makhoul, J. (2006). A s udy o ansla ion edi a e wi h a ge ed human anno a ion. In P oceedings o asso- cia ion o machine ansla ion in he Ame icas, olume 200, pages 223–231. Su ske e , I., Vinyals, O., and Le, Q. V. (2014). Sequence o sequence lea ning wi h neu al ne wo ks. In Ad ances in neu al in o ma ion p ocessing sys ems, pages 3104–3112. Ta sumi, M. (2009). Co ela ion be ween au oma ic e alua ion me ic sco es, pos - edi ing speed, and some o he ac o s. P oceedings o he Twel h Machine T ans- la ion Summi (MT-Summi XII), pages 332–339. Theano De elopmen Team (2016). Theano: A Py hon amewo k o as com- pu a ion o ma hema ical exp essions. a Xi e-p in s, abs/1605.02688. Tiedemann, J. (2009). News om opus-a collec ion o mul ilingual pa allel co - po a wi h ools and in e aces. In Recen ad ances in na u al language p ocessing, olume 5, pages 237–248. Zeile , M. D. (2012). Adadel a: an adap i e lea ning a e me hod. a Xi p ep in a Xi :1212.5701. Zens, R., Och, F. J., and Ney, H. (2002). Ph ase-based s a is ical machine ansla- ion. In Annual Con e ence on A i icial In elligence, pages 18–32. Sp inge .