Depa amen o de Sis emas In o má icos y Compu ación
Uni e si a Poli ècnica de València
Online Lea ning in Neu al Machine
T ansla ion
MASTER THESIS
Más e en In eligencia A i icial, Reconocimien o de Fo mas e
Imagen Digi al
Au ho : Luis Ceb ián Chuliá
Tu o : F ancisco Casacube a Nolla
Ál a o Pe is Ab il
Cou se 2016-2017
Resum
La aducció de g an quali a es oba mol demanada en l’ac uali a . To i
que la aducció au omà ica o e ix unes p es acions accep ables, en alguns casos
no és su icien i és necessà ia la supe isió humana. Pe a acili a la asca de
aducció de l’humà, els sis emes de aducció au omà ica p enen pa en aques
p océs. Quan una no a o ació en el llengua ge o igen necessi a se aduïda,
es a s’in oduïx en el sis ema, el qual ob é com a eixida una hipò esi de aducció.
Lla o s, l’humà co egix aques a hipò esi ( ambé conegu com a pos -edi a ) pe a
ob ind e una aducció de majo quali a . Se capaços de ans e i el coneixemen
que l’ humà exhibix quan eali za la asca de pos -edició al sis ema de aducció
au omà ica és una ca ac e ís ica desi jable ja que s’ha demos a que un sis ema
de aducció mes p ecís ajuda a augmen a l‘e iciència del p océs de pos -edició.
Pel e que el p océs de pos -edició eque ix un sis ema ja en ena , les ècni-
ques d’ap enen a ge en línia són les adequades pe aques a asca. En es e eball,
es p oposen es algo i mes d’ap enen a ge en línia aplica s a un aduc o au-
omà ic neu onal en un escena i de pos -edició. Es os algo i mes es basen en
l’ap oximació en línia Passi e-Agg essi e en la qual el model s’ac uali za desp és
de cada mos a amb l’objec iu de compli un c i e i de co ecció al ma eix emps
que man é in o mació p è ia ap esa. L’objec iu és adap a i e ina un sis ema ja
en ena amb no es mos es al ol men e el p océs de pos -edició es du a e me
(pe an , el emps d’ac uali zació ha de man eni -se con ola ).
A més, es os algo i mes es compa en amb al es ben conegudes a ian s en
línia de l’algo i me de descens pe g adien es ocàs ic. Els esul a s mos en una
millo a en la quali a de les aduccions desp és d’aplica es os algo i mes, e-
duin així l’es o ç humà en el p océs de pos -edició.
Pa aules clau: Ap enen a ge en línia, T aducció au omà ica neu onal, Passi e-
Agg esi e
Resumen
La aducción de g an calidad es á muy demandada en la ac ualidad. A pesa
de que la aducción au omá ica o ece unas p es aciones acep ables, en algunos
casos no es su icien e y es necesa ia la supe isión humana. Pa a acili a la a ea
de aducción del humano, los sis emas de aducción au omá ica oman pa e en
es e p oceso. Cuando una nue a o ación en el idioma o igen necesi a se adu-
cida, es a se in oduce en el sis ema, el cual ob iene como salida una hipó esis de
aducción. El humano en onces, co ige es a hipó esis ( ambién conocido como
pos -edi a ) pa a ob ene una aducción de mayo calidad. Se capaz de ans e-
i el conocimien o que el humano exhibe cuando ealiza la a ea de pos -edición
al sis ema de aducción au omá ica es una ca ac e ís ica deseable pues o que se
ha demos ado que un sis ema de aducción mas p eciso ayuda a aumen a la
e iciencia del p oceso de pos -edición.
Debido a que el p oceso de pos -edición equie e un sis ema ya en enado, las
écnicas de ap endizaje en línea son las adecuadas pa a es a a ea. En es e aba-
jo, se p oponen es algo i mos de ap endizaje en línea aplicados a un aduc o
au omá ico neu onal en un escena io de pos -edición. Es os algo i mos se basan
iii
i
en la ap oximación en línea Passi e-Agg essi e en la cual el modelo se ac ualiza
después de cada mues a con el obje i o de cumpli un c i e io de co ección a
la ez que man eniendo in o mación p e ia ap endida. El obje i o es adap a y
e ina un sis ema ya en enado con nue as mues as al uelo mien as el p o-
ceso de pos -edición se lle a a cabo (po an o, el iempo de ac ualización debe
man ene se bajo con ol).
Además, es os algo i mos se compa an con o as bien conocidas a ian es en
línea del algo i mo de descenso po g adien e es ocás ico. Los esul ados mues-
an una mejo a en la calidad de las aducciones después de aplica es os algo-
i mos, educiendo así el es ue zo humano en el p oceso de pos -edición.
Palab as cla e: Ap endizaje en línea, T aducción au omá ica neu onal, Passi e-
Agg esi e
Abs ac
High quali y ansla ions a e in high demand hese days. Al hough machine
ansla ion o e s accep able pe o mance, i is no su icien in some cases and
human supe ision is equi ed. In o de o ease he ansla ion ask o he human,
machine ansla ion sys ems ake pa in his p ocess. When a sen ence in he
sou ce language needs o be ansla ed, i is ed o he sys em which ou pu s a
hypo hesis ansla ion. The human hen, co ec s his hypo hesis (also known as
pos -edi ing) in o de o ob ain a high quali y ansla ion. Being able o ans e
he knowledge ha a human ansla o exhibi when pos -edi ing a ansla ion o
he machine ansla ion sys em is a desi able ea u e, as i has been p o en ha a
mo e accu a e machine ansla ion sys em helps o inc ease he e iciency o he
pos -edi ing p ocess.
Because he pos -edi ing scena io equi es an al eady ained sys em, online
lea ning echniques a e sui ed o his ask. In his wo k, h ee online lea ning
algo i hms ha e been p oposed and applied o a neu al machine ansla ion sys-
em in a pos -edi ing scena io. They ely on he Passi e-Agg essi e online lea n-
ing app oach in which he model is upda ed a e e e y sample in o de o ul il
a co ec ness c i e ion while emembe ing p e iously lea ned in o ma ion. The
goal is o adap and e ine an al eady ained sys em wi h new samples on- he-
ly as he pos -edi ing p ocess akes place (hence, he upda e ime mus be kep
unde con ol).
Mo eo e , hese new algo i hms a e compa ed wi h well-s ablished online
lea ning a ian s o he s ochas ic g adien descen algo i hm. Resul s show im-
p o emen s on he ansla ion quali y o he sys em a e applying hese algo-
i hms, educing human e o in he pos -edi ing p ocess.
Key wo ds: Online Lea ning, Neu al Machine T ansla ion, Passi e-Agg esi e
Con en s
Con en s
Lis o Figu es ii
Lis o Tables ii
1 Mo i a ion 1
2 In oduc ion o MT 3
2.1 S a is ical machine ansla ion ...................... 3
2.1.1 Language model ......................... 4
2.1.2 T ansla ion model ......................... 5
2.1.3 Log-linea model .......................... 6
2.2 Assessmen ................................. 6
2.2.1 BLEU ................................ 7
2.2.2 TER ................................. 8
2.2.3 METEOR .............................. 8
3 Neu al machine ansla ion 11
3.1 Modelling language wi h neu al ne wo ks ............... 12
3.1.1 Con inuous wo d ep esen a ion: wo d embedding ..... 12
3.1.2 Dealing wi h sequences: ecu en neu al ne wo ks ..... 13
3.1.3 Dealing wi h con ex : bidi ec ional ecu en neu al ne wo ks 15
3.2 End- o-end ansla ion: encode -decode ................ 15
3.2.1 A en ion model .......................... 17
3.2.2 Decoding ansla ions: beam sea ch .............. 17
4 T aining neu al ne wo ks 19
4.1 The backp opaga ion algo i hm ..................... 19
4.1.1 Backp opaga ion h ough ime ................. 20
4.2 S ochas ic g adien descen ........................ 21
4.2.1 SGD wi h momen um ...................... 22
4.2.2 Adag ad .............................. 22
4.2.3 Adadel a .............................. 23
4.2.4 Adam ................................ 23
5 Online lea ning 25
5.1 Online lea ning amewo k ....................... 25
5.2 Passi e-Agg essi e online lea ning ................... 26
5.3 PA online lea ning applied o NMT ................... 27
5.3.1 Passi e-Agg essi e ia subg adien echniques ........ 28
5.3.2 Passi e-Agg essi e ia p ojec ed subg adien echniques . . 29
5.3.3 Passi e-Agg essi e ia SGD wi h egula iza ion ....... 29
6 Expe imen s and esul s 31
6.1 Expe imen al amewo k ......................... 31
i
CONTENTS
6.1.1 Task desc ip ion .......................... 31
6.1.2 So wa e .............................. 32
6.1.3 Co po a .............................. 32
6.1.4 NMT sys em ............................ 33
6.2 Resul s ................................... 34
6.2.1 Hype pa ame e con igu a ion ................. 34
6.2.2 Compa ison be ween OL algo i hms .............. 35
7 Fu u e wo k and conclusions 39
7.1 Fu u e wo k ................................ 39
7.1.1 Using o he loss unc ions .................... 39
7.1.2 Inclusion o OL s a egies in IMT ................ 40
7.2 Conclusions ................................ 40
Bibliog aphy 41
Lis o Figu es
2.1 Wo d aligmen ............................... 5
3.1 LSTM .................................... 14
3.2 Bidi ec ional ecu en neu al ne wo k ................. 15
3.3 Encode -decode ............................. 16
4.1 Backp opaga ion example ........................ 21
6.1 In luence o hype pa ame e s in PA algo i hms ............ 35
6.2 E ec o uning NMT sys ems in e ms o BLEU ............ 36
6.3 E ec o uning NMT sys ems in e ms o TER ............. 37
6.4 E olu ion o NMT sys ems ........................ 38
Lis o Tables
6.1 S a is ics o he Xe ox co pus ....................... 32
6.2 S a is ics o he Emea co pus ....................... 33
6.3 S a is ics o he TED co pus ....................... 33
6.4 G id sea ch alues ............................. 34
6.5 Bes con igu a ion PA ........................... 34
6.6 Op imiza ion hype pa ame e s ..................... 36
6.7 Resul s all asks .............................. 38
ii
CHAPTER 1
Mo i a ion
We ha e seen in mode n his o y how he imp o emen o anspo a ion and
elecommunica ions in as uc u e has had a big impac in he mo emen o peo-
ple. Now we a e mo e exposed o o he cul u es, p oduc s and wo ld iews han
e e be o e. As echnology imp o es, he access o any kind o in o ma ion has
become some hing we can’ li e wi hou . Howe e , di e en cul u es o en mean
di e en languages which implies he appea ance o obs acles when he need o
communica e is manda o y. Thus, he need o a ool ha helps us o o e come
his obs acle a ises. Machine T ansla ion (MT) in es iga es he use o so wa e o
ansla e ex o speech om one language in o ano he . I aims a p o iding he
bes possible ansla ion wi hou human assis ance.
This esea ch ield was bo n in 1947 when Wa en Wea e published a memo-
andum s a ing his belie s abou compu e ’s capabili y o ansla e one language
in o ano he using c ip og aphy, logic and linguis ic pa e ns. In he nex yea s
a lo o p ojec s eme ged in he Uni ed S a es, he So ie Union and Wes e n Eu-
ope bu he p ac ical esul s we e disappoin ing (Madsen,2009;Hu chins,2005).
This inally led o he publica ion o a epo om a special commi ee o med
by he Uni ed S a es named ALPAC (Au oma ic Language P ocessing Ad iso y
Commi ee) whe e hey s a ed ha he e was no u u e o good-quali y/cos -
e ec i e MT.
Du ing he 70’s he e was a new app oach o he MT ield ha elied on he
p emise ha a language is based on a se o g amma ical and syn ac ic ules. This
ule-based sys em needed a obus and ca e ully designed bilingual dic iona y o
linguis ic in o ma ion c ea ed by human expe s. The big e o ha his app oach
equi ed was a big issue and i was hen eplaced by he co pus-based sys ems.
A he same ime, scien is s ocused on de eloping ools ha would acili a e he
ansla ion p ocess a he han eplacing human ansla o s, leading o he de-
elopmen o ansla ion memo y (TM) and o he compu e assis ed ansla ion
(CAT) (Samson,2005) ools.
Co pus-based sys ems ely on a pa allel co pus o sen ences in a sou ce lan-
guage and i s ansla ion in a a ge language. Wi h enough numbe o samples,
hese sys ems can be ained o in e he ansla ions o new sen ences. Fo his
pu pose, s a is ical models a e applied o he ansla ion p ocess as hese a e good
a ex ac ing ele an in o ma ion om a se o ( ansla ion) examples.
1
8
In oduc ion o MT
Acco ding o he expe imen a ion made in (Papineni e al.,2002), N is se o 4,
and wn=1
N. The BLEU sco e anges om 0 o 1 being 1 he closes a candida e
ansla ion can be o he e e ence (iden ical in his case). This is he main measu e
we will use in he expe imen a ion sec ion.
TER
The ansla ion e o a e (TER) (Sno e e al.,2006) is ano he measu e used o
assess ansla ions. I measu es he minimum numbe o edi s needed o change
a hypo hesis so ha i exac ly ma ches he e e ence, no malized by he a e age
leng h o he e e ence (i we ha e mo e han one). I we ha e se e al e e ences
he TER sco e will co espond o he minimum numbe o edi s needed o ma ch
he closes e e ence.
The possible edi s include inse ion, dele ion and subs i u ion o single wo ds
as well as shi s o wo d sequences. A shi is jus a mo emen o con iguous
wo ds wi hin he hypo hesis. An impo an hing o no e is ha all edi s, includ-
ing shi s o any numbe o wo ds and dis ance ha e equal cos . In a pos -edi ing
scena io his measu e app oxima es he human e o equi ed o co ec a ans-
la ion p oduced by a MT sys em. Because we wan o educe he human e o
equi ed, we wan o achie e a low TER sco e.
METEOR
METEOR (Bane jee and La ie,2005) was designed o add ess some o he is-
sues ha BLEU in oduced, such as: he lack o ecall, he use o highe o de n-
g ams o model wo d o de o he lack o explici wo d-ma ching be ween ans-
la ion and e e ence. I is based on he unig am ma ching be ween he machine-
p oduced ansla ion and he human-p oduced e e ence ansla ion.
Gi en a pai o sen ences, METEOR gene a es an alignmen be ween he wo
s ings. In his con ex , an alignmen is a ma ching be ween unig ams, such ha
e e y unig am in each s ing maps o ze o o one unig am in he o he s ing.
In o de o gene a e ha alignmen he p ocess is di ided in wo phases, each o
which has di e en s ages.
In he i s phase, based on di e en c i e ia, di e en modules p oduce uni-
g am mappings be ween he wo s ings. The "exac " module p oduces ma ches
be ween wo unig ams i hey a e exac ly he same. The "po e s em" module
p oduces unig am ma ches i hey a e he same a e being s emmed. The "Wo d-
Ne synonymy" maps wo s ings i hey a e synonyms. This modules a e o de ed
by p io i y and a gi en ma ch will be chosen i s i i is possible o o m an align-
men and i i also has highe p io i y han o he ma ches p oduced by lowe
p io i y modules. Fo example, an exac ma ch will be always p e e ed.
In he second phase, he la ges subse o unig am mappings is selec ed such
ha he esul ing se cons i u es an alignmen as de ined abo e. I mo e han one
subse cons i u es an alignmen , his me ic selec s he one wi h ewe mapping
c osses. I bo h sen ences a e w i en one below he o he and a line is d awn
2.2 Assessmen
9
be ween he ma ching unig ams, a c oss is p oduced when hese lines in e sec
wi h ano he mapping.
Las ly, in o de o ob ain he METEOR sco e, an ha monic mean o p ecision
and ecall is compu ed o e he unig ams and penalized i hese unig ams a e in
di e en o de compa ed o he e e ence. A high sco e will mean high quali y
ansla ions.
CHAPTER 3
Neu al machine ansla ion
S a e-o - he-a MT sys ems ha e elied on he ph ase-based app oach o a long
ime. Howe e , a new neu al app oach has eme ged, being he i s echnology
ha has been able o challenge o me sys ems (Luong and Manning,2015;Jean
e al.,2015). Wi h he use o G aphic P ocesso Uni s (GPU), neu al machine
ansla ion (NMT) has been able o cope wi h he high compu a ional cos i e-
qui es o compe e wi h s a e-o - he-a ph ase-based sys ems (Ben i ogli e al.,
2016).
This ac eamed wi h he imp o emen s made o he encode -decode a chi-
ec u e (explained la e in his chap e ) such as he a en ion model, o he use o
ga ed ecu en uni s o cope wi h con ex in long sen ences, made NMT echnol-
ogy o ad ance by leaps and bounds. Apa om ha , NMT demons a ed i s
powe a he IWSLT12015 e alua ion campaign, whe e one o hese sys ems ou -
pe o med he up- o- hen s a e-o - he-a ph ase-based sys ems on he English-
Ge man ask, a pai o languages di icul due o he dispa i y in mo phology
and syn ax.
The nex yea , Google announced ha Google T ansla e made he leap o NMT
because neu al ne wo ks inc ease bo h luency and accu acy o i s ansla ions2.
Mic oso did he same wi h Mic oso T ansla o s a ing ha neu al ne wo ks
be e cap u e he con ex o ull sen ences be o e ansla ing hem, p o iding
much highe quali y ansla ions3. Howe e , despi e he easons hese big com-
panies ha e o shi o NMT we s ill lack a solid o mal backg ound ha suppo s
his echnology. Things like he ac i a ion unc ion choice, he numbe o uni s
pe laye o how many laye s o use, a e hings ha need a ho ough s udy o
unde s and hem. Fo now, we can only s udy NMT sys ems, by looking a he
ansla ions and see wha di e ences hem om he ones p oduced by o me
app oaches (Ben i ogli e al.,2016).
In his chap e we will gi e an o e iew o he NMT echnology and how i
can be used o pe o m he ask a hand. Fi s we will see how we can model he
language wi h neu al ne wo ks. Then we will p esen he cu en neu al a chi-
1In e na ional Wo kshop on Spoken Language T ansla ion.
2Found in ansla ion: Mo e accu a e, luen sen ences in Google T ansla e.
3Mic oso T ansla o launching Neu al Ne wo k based ansla ions o all i s speech lan-
guages.
11
12
Neu al machine ansla ion
ec u e used o ansla e. Las ly, we will see how we pe o m he sea ch ask o
ind he bes ansla ion.
Modelling language wi h neu al ne wo ks
Being able o ep esen language wi h neu al ne wo ks is no an easy ask. Fi s ,
we need o sol e he issue ha comes om he wo ds being ep esen ed as sym-
bols a he han numbe s. Fo his ask, we need a con inuous ep esen a ion
(a dense, eal- alued ec o ) ha bo h encapsula es ele an in o ma ion o such
symbols and educes he impac o he cu se o dimensionali y. This la e e m
e e s o he need o huge numbe o examples when lea ning complex unc ions
(such as language). This ac ge s wo se as he numbe o a iables inc eases (i.e.
he size o he ocabula y) since he model needs o disc imina e be ween a huge
numbe o combina ions o such alues.
Once we ha e sol ed he p e ious issue we ace ano he : we need o manage
he ac ha sen ences a e no o he same leng h. No only ha , he inpu and
ou pu sen ence can be o di e en leng h, so we need a way o p ocess in o ma-
ion o a iable leng h. Las ly, we need a way o cap u e con ex ual in o ma ion.
All hese h ee p oblems will be add essed in nex subsec ions.
Con inuous wo d ep esen a ion: wo d embedding
A con inuous ep esen a ion o a wo d is a ec o o ea u es which cha ac e -
ize he meaning o ha wo d. To unde s and his, i a human would ha e o
ex ac hese ea u es, he would choose g amma ical ea u es such as gende o
plu ali y. Wi h neu al ne wo ks we le he lea ning algo i hm disco e his eal-
alued ea u es ins ead. The idea is o map e e y wo d in he ocabula y o a
low-dimensional con inuous- alued ec o wi h he hope ha simila wo ds ge
simila ep esen a ions in he ea u e space. The main goal behind i is o allow
he model o gene alize be e o sequences ha a e no seen du ing aining bu
whose ea u es a e simila o hose who ha e been seen.
Al hough dis ibu ed wo d ep esen a ions we e i s p oposed in he ea ly
1980’s (Hin on,1986) and 1990’s (Cas año and Casacube a,1997), i was no un-
il he 2000’s ha i s app oaches using neu al ne wo ks appea ed o model lan-
guage. These models we e s udied i s in e ms o eed- o wa d ne wo ks (Ben-
gio e al.,2003) and la e in e ms o ecu en neu al ne wo ks (RNN) (Mikolo
e al.,2010,2011) and yielded good wo d ep esen a ion ha condensed linguis ic
egula i ies in he ec o ep esen a ions (Mikolo e al.,2013b).
One o he i s app oaches ha success ully used neu al ne wo ks o model
language was he one p oposed in Bengio e al. (2003). This wo k p ojec ed he
wo ds in o a eal- alued ec o . Then, i used a eed- o wa d ne wo k ocusing
on lea ning bo h he wo d ea u e ec o and he p obabili y dis ibu ion unc-
ion o wo d sequences in e ms o hese ea u e ec o s by maximizing he log-
likelihood o e a aining da a. The esul s ob ained wi h his app oach demon-
s a ed how dis ibu ed ep esen a ion o wo ds could be eamed up wi h neu al
ne wo ks o ou pe o m egula n-g am models.
3.1 Modelling language wi h neu al ne wo ks
13
All his wo k demons a ed ou s anding pe o mance in wo d-p edic ion bu
also he need o a mo e compu a ionally e icien model (specially he ones in-
ol ing RNN). Fo his pu pose, Mikolo e al. (2013a) p oposed a model ha
could lea n wo d ep esen a ion much as e wi h highe quali y han p e ious
app oaches. I showed he lea ned ela ionship be ween wo ds wi h simple al-
geb a by simply adding and sub ac ing ec o s o ob aining o he wo ds. Fo
example, he wo d Rome could be ob ained wi h: Pa is - F ance + I aly.
Dealing wi h sequences: ecu en neu al ne wo ks
When humans ead a sen ence, hey unde s and each wo d based on hei unde -
s anding o p e ious wo ds. As you ead, you b ain e ains in o ma ion ha is
used o unde s and wha is coming. Recu en neu al ne wo ks model his si -
ua ion na u ally. This is a majo ad an age since a a ie y o p oblems depend
on an a bi a y sequence o ime-dependen e en s, such as speech ecogni ion,
ansla ion o language modelling.
They ake in o accoun pas in o ma ion by means o a cycle be ween i s uni s.
This allows hem o ha e an in e nal s a e ha holds p esen and pas in o ma-
ion. In addi ion, depending on how hese cycles a e ou ed we ind di e en
a chi ec u es in he li e a u e. An example o his is he Elman (Elman,1990) and
Jo dan (Jo dan,1986) ne wo ks, also known as "simple ecu en ne wo ks" .
The Elman a chi ec u e wo ks as ollows: gi en a sequence o ec o s x=
x1...xT he ne wo k will p oduce a sequence o ou pu s y=y1...yT. A each ime
s ep (Eq. 3.1) he hidden s a e a ime ,s ge s calcula ed based on bo h, he inpu
a ime ,x and he hidden s a e o he p e ious ime s ep, s −1. Then, he ou pu
(Eq. 3.2 ) ge s calcula ed based on s .
s =σs(Wx +Us −1)(3.1)
y =σy(Vs )(3.2)
In hese equa ions, W,Uand Va e he weigh ma ices o he inpu , ecu en
and ou pu connec ions espec i ely, σsis an ac i a ion unc ion such as sigmoid
and σyis he ou pu ac i a ion unc ion, being usually, he so max unc ion (Eq.
3.18). Al e na i ely, he Jo dan a chi ec u e is simila bu he ecu en connec ion
is achie ed by aking in o accoun he p e ious ou pu a he han he p e ious
hidden s a e (Eq. 3.3).
s =σs(Wx +Uy −1)(3.3)
Howe e , one o he majo laws hese ne wo ks ha e is hei dependency on
he comple e pas in o ma ion. This p e en s hem om ocusing only on ecen
in o ma ion o on long- e m dependencies, as depic ed in (Bengio e al.,1994).
When aining hem, hey also ea u e he so-called anishing g adien p oblem
(explained in sec ion 4.1) which p e en s hem om p ope ly lea ning.
To add ess his, se e al new ecu en a chi ec u es ha e been p oposed, such
as he long sho e m memo y ne wo ks (Hoch ei e and Schmidhube ,1997)
14
Neu al machine ansla ion
known as LSTM (Fig. 3.1), o he ga ed ecu en uni s (GRU) (Cho e al.,2014)
which ha e he special ai o mi iga ing he anishing g adien p oblem. Bo h
o hese a chi ec u es ea u e pa ame ized ga es ha modula e how much inpu
in o ma ion is allowed o a ec he hidden s a e, how much pas in o ma ion is
le h ough o he nex ime s ep and how much in o ma ion is o go en. These
a e ainable pa ame e s ha allow he ne wo k o do his con enien ly.
In LSTM uni s, he memo y cell m (Eq. 3.4) depends on he p e ious memo y
cell m −1and he new in o ma ion ˜
m (Eq. 3.5) coming om he inpu o he
uni . The amoun o in o ma ion used o upda e m om bo h m −1and ˜
m is
modula ed by g (Eq. 3.6) and i (Eq. 3.7) which a e he ec o ou pu s o he
o ge ga e and inpu ga e whose alues ange om 0 o 1. By doing an elemen -
wise mul iplica ion hey modula e he amoun o in o ma ion ha will be used
in he upda e, being 1 comple ely e ain and 0 comple ely o ge . The ou pu o
he uni s is modula ed by he ou pu ga e (Eq. 3.8) ge ing Eq. 3.9.
m =g m −1+i ˜
m (3.4)
˜
m = anh(WM
Ys −1+WM
Xx )(3.5)
g =σ(WF
Ys −1+WF
Xx )(3.6)
i =σ(WI
Ys −1+WI
Xx )(3.7)
o =σ(WO
Ys −1+WO
Xx )(3.8)
s =o anh(m )(3.9)
In hese equa ions, WO
Y,WI
Y,WF
Yand WM
Ya e he ou pu ga e, inpu ga e, o -
ge ga e, and memo y cell ecu en weigh ma ices and WO
X,WI
X,WF
X,WM
Xa e
he ou pu ga e, inpu ga e, o ge ga e, and memo y cell inpu weigh ma ices.
σ(·)and anh(·)a e he sigma and hype bolic angen ac i a ion unc ions.
Inpu
Inpu
Fo ge
Ou pu
ga e
ga e
ga e
x
s −1
s −1s −1
s −1
x x
x
−1
s
Cell
˜
s
i
g
m −1
m
o
Figu e 3.1: LSTM uni
3.2 End- o-end ansla ion: encode -decode
15
Dealing wi h con ex : bidi ec ional ecu en neu al ne wo ks
When p ocessing sequences, he e a e imes when u u e e en s a e use ul and
p o ide mo e in o ma ion o he ne wo k in o de o be e pe o m he ask
a hand. Regula RNN ha e he limi a ion ha u u e in o ma ion can no be
eached om he cu en s a e.
Bidi ec ional ecu en neu al ne wo ks (BRNN) (Schus e and Paliwal,1997)
add ess his issue allowing u u e in o ma ion o be eachable om he cu en
s a e. They do his by connec ing wo hidden laye s o opposi e di ec ion o he
same ou pu laye , ha ing hen, in o ma ion o pas and u u e e en s.
Wi h his a chi ec u e we call o wa d laye he one ha p ocesses he sequence
in he posi i e ime di ec ion (le o igh ) and backwa d laye he one ha p o-
cesses he sequence in he nega i e ime di ec ion ( igh o le ). A each ime
s ep, he ou pu o bo h laye s a e combined (e.g. adding hem up) be o e apply-
ing he ou pu ac i a ion unc ion.
Fo wa d
s a es
Backwa d
s a es
-1 +1
Ou pu neu on
g oup
Hidden (s a e)
neu on g oup
Inpu s
G oup o
weigh s wi h
in o ma ion
low
Figu e 3.2: S uc u e o a bidi ec ional ecu en neu al ne wo k shown un olded in ime
o h ee ime s eps. Example bo owed om Schus e and Paliwal (1997).
.
End- o-end ansla ion: encode -decode
We ha e seen ha when a RNN p ocesses a sequence o leng h Ti p oduces an
ou pu sequence o he same leng h. In ansla ion howe e , sou ce and a ge
sen ence can (and usually a e) o di e en leng hs. To o e come his, wo simila
models we e p oposed by Su ske e e al. (2014) and Cho e al. (2014) which elied
on he encode -decode app oach. This wo k cons i u es a common amewo k
on which la e wo k aims o imp o e i in di e en ways.
The encode -decode app oach consis in a wo s ep p ocess ha i s , maps
he sou ce sen ence in o a ixed leng h ec o , and second, his ec o is decoded
o p oduce he a ge sen ence (possibly o di e en size). I aims o di ec ly
model he condi ional ansla ion p obabili y (Eq. 2.1).
The inpu o he sys em is a sequence o wo ds x=x1...xJp esen in he
ocabula y o he sou ce language Vxin a one-ho ep esen a ion. This means ha
we ha e o each wo d xja ocabula y-sized ec o ¯xj∈N|Vx|wi h all elemen s
16
Neu al machine ansla ion
se o ze o excep o he one loca ed a he posi ion o xjin Vxwhich is se o one.
Each wo d is p ojec ed o a ixed-leng h eal- alued ec o :
xj=Es¯
xj(3.10)
whe e xj∈Rdis he embedding o wo d xj,Es∈Rd×|Vx|is he sou ce languaje
p ojec ion ma ix and d he embedding size.
c
NULL
x1x2xT
y1y2yT0
y1yT0−1
Encode
Decode
Figu e 3.3: Ilus a ion o he encode -decode app oach.
The encode is an RNN ha eads each elemen o he sequence o wo d em-
beddings. A e eading each elemen in he posi i e di ec ion, he hidden s a e
h
jo he RNN is upda ed (Eq. 3.11). Once he end-o -sen ence symbol is eached,
he hidden s a e o he RNN c(Eq. 3.12), can be seen as a ixed-leng h summa y
o he whole inpu sequence.
h
j= (xj,h
j−1)(3.11)
c=h
J(3.12)
whe e is a non-linea unc ion.
The decode is ano he RNN which is ained o gene a e he ou pu sequence
by p edic ing he nex wo d yigi en he hidden s a e si, he p e ious gene a ed
wo d yi−1and c. The hidden s a e o he decode a ime iis compu ed as:
si= (si−1,yi−1,c)(3.13)
whe e is a non-linea ac i a ion unc ion (LSTM o GRU) ha p oduce he hid-
den s a e ec o si. The condi ional p obabili y o he nex wo d yiis app oxi-
ma ed (Eq. 3.14) assuming ha i depends on he p e ious wo d (and all p e ious
wo ds ha a e in si o some ex en ). In his case, g(·)∈R|Vy|is he so max unc-
ion (Eq. 3.18) ha p oduces a ec o o p obabili ies o size |Vy|, which is he size
o he ocabula y o he a ge language. ¯yi∈N|Vy|is he one-ho ep esen a ion
o he wo d yi.V∈R|Vy|×Lis he weigh ma ix and ϕ(·)is he L-sized ou pu
laye o a RNN.
P (yi|y1...yi−1,c)≈¯y0
ig(Vϕ(si,yi−1,c)) (3.14)
I is impo an o no e ha , du ing aining, we wan he decode o lea n o
ou pu he e e ence sen ence. In o de o help he model, we use he eache o c-
ing (Pascanu e al.,2013) echnique. This echnique consis o eeding he model
3.2 End- o-end ansla ion: encode -decode
17
wi h he co esponding wo d in he e e ence sen ence a ime s ep i−1 ins ead
o he p e iously gene a ed wo d.
A en ion model
One o he laws ha encode -decode models exhibi is ha hey ha e o con-
dense an a bi a y leng h sen ence in o a ixed-leng h ec o c. This is a bo leneck
when dealing wi h long sen ences, as no iced by Cho e al. (2014), and causes he
sys em o pe o m poo ly in hese si ua ions.
In o de o sol e his, Bahdanau e al. (2014) p oposed he so-called a en ion
model. The basic idea is o use a di e en con ex ec o depending on he cu en
decoding s age. Wi h his idea, Equa ion 3.14 is ew i en as:
P (yi|y1...yi−1,c)≈¯y0
ig(Vϕ(si,yi−1,ci)) (3.15)
whe e siis he RNN hidden s a e a ime i:
si= (si−1,yi−1,ci)(3.16)
We can see ha a each ime s ep, we use a di e en con ex ec o ci. This
ec o is calcula ed based on a sequence o anno a ions (h1,..., hJ)ex ac ed om
he sou ce sen ence. These anno a ions ha e in o ma ion ha s ongly desc ibe
he i- h wo d and su oundings. They a e calcula ed by conca ena ing he o -
wa d and backwa d hidden s a es o a BRNN. The anno a ion hj= [h
j;hb
j]∈R2S
whe e h
jand hb
ja e he hidden s a es o he o wa d and backwa d laye s o he
BRNN and Sis he size o he hidden s a e (we assume ha hey a e o he same
size). The e o e, ciis calcula ed as he weigh ed sum o he anno a ions:
ci=
J
∑
j=1
αijhj(3.17)
whe e αij is compu ed by he so max unc ion:
αij =exp(eij)
∑J
k=1exp(eik)(3.18)
being eij =a(si−1,hj)an alignmen model which sco es how ela ed he inpu s a
posi ion jand he ou pu a posi ion ia e. The alignmen model is pa ame ized
wi h a eed- o wa d neu al ne wo k which is ained join ly wi h he whole sys-
em.
Decoding ansla ions: beam sea ch
When decoding he ansla ion, he neu al sys em p o ides, a each ime s ep, a
se o p obabili ies (wi h he so max unc ion) o each wo d in he ocabula y.
Because we a e maximizing equa ion 2.1 and he ou pu p obabili ies o a gi en
ime s ep depend o e he p e ious chosen wo d, we canno jus choose he one
24
T aining neu al ne wo ks
decaying a e age o pas squa ed g adien s and an exponen ially decaying
a e age o pas g adien s m :
m =β1·m −1+ (1−β1)·∇`(Θ )(4.11)
=β2· −1+ (1−β2)·∇`(Θ )2(4.12)
The au ho s no iced ha a ea ly s ages, he algo i hm is biased owa ds ze o.
To coun e ac his hey co ec he abo e exp essions as ollows:
ˆm =m
1−β
1
(4.13)
ˆ =
1−β
2
(4.14)
lea ing he upda e ule as:
Θ +1=Θ −ρ
√ˆ +eˆm (4.15)
The de aul alues he au ho s p opose a e 0.9 o β1and 0.999 o β2.
CHAPTER 5
Online lea ning
Al hough compu a ional esou ces ha e e ol ed o he poin whe e i is ela i ely
cheap o ain a Pa e n Recogni ion (PR) sys em, i is s ill a ime-consuming p o-
cess ha akes days o e en weeks o comple e. As mo e and mo e da a is a ail-
able his p oblem is agg a a ed, specially on sys ems ha in e ac wi h changing
en i onmen s o ha equi e a quick esponse a he same ime hey lea n. In
hese cases, being able o adap ou sys em wi hou aining i om sc a ch be-
comes manda o y.
In his chap e we will p esen he online lea ning amewo k which will se e
o in oduce ou wo k, emphasizing i s applica ion o a pos -edi ing scena io in
MT. Nex we will explain he Passi e-Agg essi e (PA) online lea ning algo i hm.
I has he main goal o a oiding he ca as ophic in e e ence (McCloskey and
Cohen,1989) p oblem ha h ea ens neu al ne wo k sys ems in which hey end
o o ge p e iously lea ned in o ma ion upon lea ning new in o ma ion. Finally,
we will p opose h ee new PA-based algo i hms o he ask o NMT.
Online lea ning amewo k
PR sys ems ha e achie ed an accep able pe o mance in e y complex asks ha
in ol es s uc u ed ou pu and ambigui y such as machine ansla ion o image
desc ip ion. Al hough hese sys ems can each a high pe o mance on da a simi-
la o he one used o ain hem, hei pe o mance apidly d ops when he ask is
sligh ly di e en . In addi ion, ge ing enough manually anno a ed da a o a spe-
ci ic domain in o de o ain a whole sys em migh no be possible. The e o e,
i is use ul o ain wi h gene ic ou -o -domain da a and hen une he sys em
wi h domain-speci ic da a. Because a comple e e aining o he sys em migh be
in easible, online lea ning echniques (Blum,1998) a e chosen o his ask.
These echniques a e pa icula ly use ul in he compu e assis ed ansla ion
(CAT) and in e ac i e machine ansla ion (IMT) pa adigms, whe e human ans-
la o s wo k join ly wi h machine ansla o s in o de o e icien ly ob ain high
quali y ansla ions. In a pos -edi ing scena io, he MT sys em ansla es a sen-
ence and p opose his as a hypo hesis o a human ansla o which hen co ec s
he sen ence (i needed). Once a sen ence is co ec ed wha we ha e is a new
pai o sen ences ha can be used o ain he sys em. By using hese new ain-
25
26
Online lea ning
ing samples o ain ou model, we p og essi ely make he pos -edi ing p ocess
mo e e icien . Fi s , by making he sys em lea n om i s own e o s and second,
by easing he wo k o he human ansla o as i will ha e o co ec less e o s.
P io wo k in his a ea has also shown ha human ansla o s become mo e p o-
duc i e as he MT quali y imp o es (Ta sumi,2009).
The applica ion o such echniques has been s udied ho oughly (Ma ínez-
Gómez e al.,2012;La ie,2014;Ma ínez-Gómez e al.,2011) in classical ph ase-
based SMT sys ems. Resul s show ha signi ican imp o emen s can be achie ed
by adap ing he sys em by means o online lea ning in e ms o he e o he
human would need o co ec he hypo hesis p o ided. Howe e , he e has been
no published wo k ( o he bes o ou knowledge) ela ed o he applica ion o
hese echniques in NMT sys ems. This wo ks aims a p o iding new algo i hms
o apply online lea ning o hese sys ems.
Passi e-Agg essi e online lea ning
A a ie y o online lea ning me hods ha e been p oposed in li e a u e (Lu e al.,
2016). A classical OL me hod is he Pe cep on algo i hm (Rosenbla ,1958) which
upda es he model by adding a misclassi ied example (pa ame ized wi h a ac-
o ) o he cu en pa ame e s. F om his wo k, a lo o new online lea ning algo-
i hms ha e been de eloped based on he maximum ma gin c i e ion (C amme
and Singe ,2003;Gen ile,2001;Ki inen e al.,2004;Li and Long,2000) which ies
o sepa a e he classes as much as possible. One impo an echnique ha alls in
his ca ego y is he Passi e-Agg essi e (PA) online lea ning me hod (C amme
e al.,2006). This has been p o ed as a e y success ul and popula online lea n-
ing echnique o sol ing many eal-wo ld applica ions.
The idea behind he PA algo i hm is e y simple. When a new sample a i es
i is classi ied and a loss is ob ained, measu ing he deg ee o which he p edic ion
is w ong. The idea now is o ind a new se o pa ame e s (weigh s) ha make
he classi ie o co ec ly classi y he sample (agg essi eness) bu by emaining
ela i ely close o he p e ious se o pa ame e s (passi eness). We will ou line
he PA app oach in a bina y classi ica ion p oblem as explained in C amme e al.
(2006).
In a bina y classi ica ion p oblem, each sample x ∈Rdhas a unique label
y ∈ {+1, −1}associa ed. We assume ha he classi ica ion unc ion is based on
a ec o o weigh s Θ∈Rdwhich ake he o m o sign(Θ·x). The magni ude
|Θ·x|is in e p e ed as he deg ee o con idence in he p edic ion. We e e o
he e m y (Θ ·x )as he ma gin calcula ed a ime . Whene e he ma gin is
posi i e, he sample has been classi ied co ec ly. Howe e , we wan he classi ie
o achie e a co ec classi ica ion wi h some ma gin (in C amme e al. (2006) his
ma gin is se o 1). Then we de ine he loss ` ha su e s he classi ie whene e
he samples is classi ied inco ec ly:
`(Θ;(x,y)) = (0 i y(Θ·x)≤1
1−y(Θ·x)o he wise (5.1)
5.3 PA online lea ning applied o NMT
27
In o de o p og essi ely lea n he weigh ec o Θ, he algo i hm mus up-
da e i a e e e y sample. A ime , he new weigh ec o Θ +1is calcula ed by
sol ing his cons ained op imiza ion:
Θ +1=a gmin
Θ∈Rd
1
2kΘ−Θ k2s. . `(Θ;(x ,y )) = 0 (5.2)
In his algo i hm, i he loss is 0, he op imal solu ion is Θ +1=Θ . On he
o he hand, when he loss is g ea e han 0, he algo i hm o ces Θ +1 o sa is y
he cons ain `(Θ +1;(x ,y )) = 0 while being as close as possible o he p e ious
weigh ec o o p ese e knowledge o p e ious samples. I we de i e Eq. 5.2,
he upda e ule o he weigh ec o is:
Θ +1=Θ +τ y x whe e τ =`(Θ ;(x ,y ))
kx k2(5.3)
Howe e , his upda e ule is oo agg essi e and i migh esul in upda ing he
model in such a way ha in o de o mee he cons ain , i p oduces a lo o e o s
in la e samples because he model was changed d as ically. The e o e, a gen le
upda e s a egy is equi ed. To do his, he au ho s in oduce a non-nega i e slack
ξ a iable in o he op imiza ion p oblem de ined in Eq. 5.2:
Θ +1=a gmin
Θ∈Rd
1
2kΘ−Θ k2+Cξs. . `(Θ;x ,y )) ≤ξand ξ≥0 (5.4)
The pa ame e Ccon ols he in luence o he slack a iable ξ. The au ho s
coin he C a iable as he agg essi eness pa ame e o he algo i hm. The upda e
ule in his case is he same Θ +1=Θ +τ y x bu τ changes o:
τ =min{C,`(Θ ;(x ,y ))
kx k2}(5.5)
PA online lea ning applied o NMT
In his sec ion we p opose a new e sion o PA-based SGD. The de elopmen
o his wo ks is inspi ed by he applica ion o he MIRA (a PA-based) algo i hm
(C amme and Singe ,2003) used in adi ional SMT o une he weigh s o he
log-linea model. Also, he wo k de elop in Ma ínez-Gómez e al. (2012) o he
ask o online adap a ion gi es us a eel o wha can be achie ed when adap ing
a MT sys em wi h OL. As any o he PA echnique, ou algo i hm aims o pe o m
he minimum modi ica ion o he model pa ame e s while ul illing a co ec ness
c i e ion. In ou case, we use he loss unc ion o he neu al model o e iciency
easons (bu we could ha e used ano he loss unc ion such as BLEU).
28
Online lea ning
Le h be he hypo hesis gene a ed by he NMT sys em (using he pa ame e s
Θ a ime ) o he sou ce sen ence x . We conside ha Θ is inco ec i he
model assigns a lowe p obabili y o he a ge e e ence sen ence y han o h :
pΘ (y |x )<pΘ (h |x )(5.6)
When ha happens, we wan o sea ch o a b
Θsuch ha b
Θis close o Θ and
pΘ (y |x )>pΘ (h |x ). This is exp essed wi h he loss unc ion `:
`(b
Θ,x ,y ,h ) = log pb
Θ(h |x )−log pb
Θ(h |x )(5.7)
Wi h his loss unc ion we a e measu ing how la ge is he gap be ween he
p obabili y gi en o he hypo hesis and he p obabili y gi en o he e e ence. By
op imizing his unc ion we educe his gap wi h he idea ha p og essi ely, he
hypo hesis will be close (i.e. simila ) o he e e ence sen ence. Wi h he goal o
op imizing his unc ion, h ee a ia ions o PA-based algo i hms a e de eloped.
Passi e-Agg essi e ia subg adien echniques
In o de o ind b
Θ, he p oblem can be o mula ed as:
b
Θ=a gmin
Θ
1
2kΘ−Θ k2+Cξs. . `(b
Θ,x ,y ,h )≤ξand ξ≥0 (5.8)
being C he pa ame e ha con ols he agg essi eness o he algo i hm and ξ
a slack a iable (as discussed in Sec ion 5.2). F om his equa ion we ha e ξ≥
max(0, `(b
Θ,x ,y ,h )) and hen, we de ine F as he unc ion o op imize:
F (Θ,x ,y ,h ) = 1
2kΘ−Θ k2+Cmax(0, `(b
Θ,x ,y ,h )) (5.9)
aiming o ind he se o pa ame e s ha minimize his unc ion, i.e.:
b
Θ=a gmin
Θ
F (Θ,x ,y ,h )(5.10)
Because we a e op imizing a unc ion (F ) wi h discon inui y poin s (because
he max(·) unc ion has no de i a i e a 0) we ha e o use a subg adien me hod
(Sho e al.,2003). This esul s in he de i a i e ∂ΘF (Θ,x ,y ,h )being:
∂ΘF (Θ,x ,y ,h ) =
Θ−Θ −C∇`(Θ,x ,y ,h )`(Θ,x ,y ,h )<0
Θ−Θ `(Θ,x ,y ,h )>0
[Θ−Θ ,Θ−Θ −C∇`(Θ,x ,y ,h )] `(Θ,x ,y ,h ) = 0
(5.11)
5.3 PA online lea ning applied o NMT
29
whe e, on he discon inui y poin (`(Θ,x ,y ,h ) = 0), we choose a alue in he in-
e al [Θ−Θ ,Θ−Θ −C∇`(Θ,x ,y ,h )]. We assume Θ−Θ when `(Θ,x ,y ,h ) =
0. Finally, we pe o m he upda e as:
Θk+1=Θk−ρ·∂ΘF (Θ,x ,y ,h )|Θk(5.12)
whe e ρis he lea ning a e, ∂ΘF is he subg adien o F wi h espec o Θ. We
ini ialize Θk=0=Θ and he upda e is pe o med k imes pe sample. We deno e
his PA ia subg adien echnique upda e ule as PAS.
Because we may wan o achie e a highe p obabili y o he e e ence sen ence
bu wi h some ma gin m:pΘ(y |x ) + m>pΘ(h |x )we can ew i e he subg a-
dien ∂ΘF wi h he discon inui y poin being m a he han ze o (as de ined in
Eq. 5.11).
Passi e-Agg essi e ia p ojec ed subg adien echniques
An ex ension o he PAS me hod is he p ojec ed PA subg adien me hod (PPAS)
in which he op imiza ion p oblem is e o mula ed (Boyd e al.,2003). We de ine
G (Θ,x ,y ,h )as max(0, `(b
Θ,x ,y ,h )). Then, Eq. 5.8 can be ew i en as:
b
Θ=a gmin
Θ
G (Θ,x ,y ,h )s. . kΘ−Θ k2≤C(5.13)
In his case, ∂ΘG (Θ,x ,y ,h )is de ined as:
∂ΘG (Θ,x ,y ,h ) =
∇`(Θ,x ,y ,h )`(Θ,x ,y ,h )<0
0`(Θ,x ,y ,h )>0
[0, ∇`(Θ,x ,y ,h )] `(Θ,x ,y ,h ) = 0
(5.14)
and he upda e ule is pe o med in a wo s ep p ocess. Fi s , we calcula e he
in e media e weigh upda e ¯
Θk+1as:
¯
Θk+1=Θk−ρ∂ΘG (Θ,x ,y ,h )|Θk(5.15)
and second, we apply he p ojec ion ope a o , being Θk+1:
Θk+1=¯
Θk+1−Θ
k¯
Θk+1−Θ kC+Θ (5.16)
As in he p e ious case, we ini ialize Θk=0=Θ .
Passi e-Agg essi e ia SGD wi h egula iza ion
We can also op imize `(Θ,x ,y ,h )by means o SGD bu wi h a egula iza ion
e m, inspi ed by he PA p inciple o upda ing he model wi h new pa ame e s
30
Online lea ning
ha a e close o he cu en ones. We will e e o his algo i hm as PAR. Wi h his
in mind, we can de ine he minimiza ion p oblem as:
b
Θ=a gmin
Θ
(`(Θ,x ,y ,h ) + C
2kΘ−Θ k2)(5.17)
Because he minimiza ion has no discon inui y poin s we don’ need o apply
a subg adien me hod. The upda e ule hen, esul s in:
Θk+1=Θk−ρ·(∇`(Θ,x ,y ,h )|Θk+C(Θk−Θ )) (5.18)
As usual, Θk=0=Θ . We obse e ha when k=1 he upda e ule is plain
SGD wi h no egula iza ion e m:
Θk+1=Θk−ρ·(∇`(Θ,x ,y ,h )|Θk+
: 0
C(Θk−Θ ) ) (5.19)
in ha case, we in oduce a small Gaussian noise ν o he model pa ame e s in
o de o a oid eaching local minima:
Θk+1=Θk−ρ·(∇`(Θ,x ,y ,h )|Θk+Cν)(5.20)
CHAPTER 6
Expe imen s and esul s
A e he p oposal o h ee PA-based algo i hms in Sec ion 5.3, we will p oceed
o i s es ing and compa ison wi h o he well-known algo i hms ( hose desc ibed
in Sec ion 4.2). In his sec ion we will desc ibe he expe imen al se up used o
conduc he expe imen a ion. Fi s , we will desc ibe he ask in which we wan
o e alua e ou algo i hms. Then we will desc ibe he NMT sys em used and he
ools needed o wo k wi h i . A e desc ibing he co po a used, we will compa e
and discuss he esul s ob ained.
Expe imen al amewo k
Be o e p esen ing he esul s, we will p esen he ask a hand and desc ibe he
ools used o de elop he men ioned algo i hms and he co po a used o assess
hem.
Task desc ip ion
In o de o es hese algo i hms we will simula e (o he wise i would be oo
cos ly) a pos -edi ing scena io. We will simula e his scena io in h ee di e en
co pus o asks (de ailed in nex subsec ions). Fi s , we will ain h ee neu al
models using o each one he aining pa i ion o i s co esponding da ase .
Once ained, hese models will cons i u e he baseline models o each ask.
Then, we will e ine hese models wi h hei espec i e es o de elopmen
se o he co pus in which hey we e ained. Fo a gi en sou ce sen ence, he sys-
em p oduces i s ansla ion ( he sys em’s hypo hesis). This ansla ion is s o ed
o la e e alua ion. Then we ake he pos -edi ed sen ence (in his case, he e e -
ence sen ence) and upda e he sys em wi h i . This is epea ed un il he ull se o
sen ences a e ansla ed. We e alua e he s o ed hypo heses and compa e hem
wi h he ones p oduced by he baseline sys em. Ou hope is o see an imp o e-
men in he quali y o hese ansla ion. Wi h his p ocedu e we wan o see up o
wha ex en he OL algo i hm can e ine he baseline sys em.
In o de o measu e he impac o each algo i hm in he pos -edi ing e o
educ ion we use he TER me ic. In o de o assess he quali y o ansla ions we
will use he BLEU and METEOR me ics.
31
32
Expe imen s and esul s
So wa e
We ha e de eloped ou wo k using he NMT-Ke as oolki 1(Pe is,2017). This
(s ill in p og ess) so wa e p o ides ools o he NMT ask, such as: beam sea ch
decoding, suppo o GRU and LSTM ne wo ks, unknown wo d eplacemen ,
use o p e ained wo d embedding ec o s an mo e.
This so wa e is based on a Py hon lib a y called Ke as2(Cholle e al.,2015).
Ke as o e s an abs ac ion laye on op o Theano3(Theano De elopmen Team,
2016) (a lib a y o de ining and e alua ing ma hema ical exp essions in ol ing
mul i-dimensional a ays) and eases he ask o building and aining di e en
neu al ne wo k a chi ec u es. I has implemen ed u ili ies o sa e and load neu al
ne wo k models, moni o he aining p ocess and use p ede ined NN a chi ec-
u es.
All he algo i hms lis ed in Sec ion 5.3 ha e been implemen ed in Ke as and
in eg a ed in he NMT-Ke as oolki .
Co po a
In his wo k, we es all he algo i hm in h ee di e en co pus: Xe ox, Emea and
TED. We use he s anda d pa i ions and he English-F ench language pai o all
expe imen s. The ex is kep ue case and i is okenized using he sc ip om
Moses 4. In he nex subsec ions we will desc ibe hese co pus.
Xe ox
The Xe ox co pus (Ba achina e al.,2009) consis s in ansla ions o use manuals
o Xe ox p in e s. Table 6.1 shows all in o ma ion ega ding he pa i ions o
he English-F ench language pai .
T aining De elopmen Tes
Sen ences 52k 984 994
Vocabula y size English 14k 1.8k 1.7k
F ench 15.5k 1.9k 1.8k
Wo ds English 615k 10.9k 11.1k
F ench 676k 11.7k 11.8k
Table 6.1: S a is ics o he Xe ox co pus. In he able a e collec ed he numbe o sen ences
and ocabula y size o each pa i ion and language. k s ands o housands
1h ps://gi hub.com/l apeab/nm -ke as/
2h ps://ke as.io/
3h p://deeplea ning.ne /so wa e/ heano/
4h p://www.s a m .o g/moses/
6.1 Expe imen al amewo k
33
Emea
The Emea pa allel co pus (Tiedemann,2009) is made ou o documen s om he
Eu opean Medicine Agency. Table 6.2 shows all in o ma ion ega ding he pa i-
ions o he English-F ench language pai .
T aining De elopmen Tes
Sen ences 319k 500 1k
Vocabula y size English 52k 2.8k 4.5k
F ench 59.3k 2.9k 4.5k
Wo ds English 4.2M 10.3k 21.4k
F ench 4.8M 12.3k 25.7k
Table 6.2: S a is ics o he Emea co pus. In he able a e collec ed he numbe o sen ences
and ocabula y size o each pa i ion and language. k s ands o housands and M o
millions
TED
The TED co pus (Fede ico e al.,2011) ga he s ansc ibed TED alks. Table 6.3
shows all in o ma ion ega ding he pa i ions o he English-F ench language
pai .
T aining De elopmen Tes
Sen ences 159k 887 1.7k
Vocabula y size English 46.7k 3.4k 4.9k
F ench 58.2k 3.9k 4.9k
Wo ds English 2.17M 21.3k 33.5k
F ench 2.3M 21.8k 35.7k
Table 6.3: S a is ics o he TED co pus. In he able a e collec ed he numbe o sen ences
and ocabula y size o each pa i ion and language. k s ands o housands
NMT sys em
Fo each co pus, we ain a neu al model o e he aining pa i ion. Such model
(simila o he one depic ed in Bahdanau e al. (2014)) consis s in an encode -
decode LSTM ne wo k equipped wi h he a en ion mechanism. We made use
o single-laye ed LSTM due o ime limi a ions. The size o he hidden s a e o
each LSTM, he wo d embedding size and he a en ion mechanism laye is 512
(ou choice is based on he expe imen a ions ca ied ou in B i z e al. (2017)).
The baseline sys ems ha e been ained wi h he Adadel a algo i hm wi h he
de aul pa ame e s. Du ing aining, Gaussian noise (G a es,2011) and laye
no maliza ion (Ba e al.,2016) a e applied as egula iza ion me hods. We also
ea ly s opped he aining i he BLEU on he de elopmen se did no imp o e
in 100,000 upda es. The size o he beam is 6 in all expe imen s.
40
Fu u e wo k and conclusions
Inclusion o OL s a egies in IMT
The IMT scena io is igh ly ela ed wi h he pos -edi ing scena io. IMT ies o
collabo a e wi h he human on he ansla ion ask wi h he goal o minimizing
human e o . O en, he use only needs o poin ou he posi ion in he sen ence
in which i is w ongly ansla ed and he sys em will ou pu an al e na i e su -
ix s a ing om ha poin . IMT is an ac i e ield and ecen wo k has s udied
i s applica ion on NMT sys ems (Knowles and Koehn,2016;Pe is e al.,2017b).
As in he pos -edi ing scena io, being able o apply OL echniques in o an IMT
amewo k migh be bene icial in o de o de elop mo e adap i e and p oduc i e
MT sys ems.
Conclusions
In his wo k, h ee OL algo i hms ha e been p oposed and applied o an NMT
sys em. They ely on he Passi e-Agg essi e OL app oach in which he MT
model is upda ed a e e e y sample in o de o ul il a co ec ness c i e ion while
emembe ing p e iously lea ned in o ma ion. Because hese algo i hms ha e a
high numbe o hype pa ame e s o une, an ex ensi e g id sea ch has been ca -
ied ou in o de o ind he bes combina ion. In addi ion, a s udy has been made
o see how he di e en pa ame e s a ec each algo i hm’s pe o mance. Resul s
show ha , al hough hey depend on se e al hype pa ame e s, hose c i ical o
he well beha iou o he algo i hms a e, in e e y case, he agg essi eness hy-
pe pa ame e and he lea ning a e. O he pa ame e s such as he ma gin o he
numbe o i e a ions pe sample o e mino imp o emen s in he o e all pe o -
mance o he algo i hm.
These algo i hms ha e been es ed in a pos -edi ing scena io whe e an al eady
ained NMT sys em has o be e ined as soon as new samples a e a ailable. Each
new sample ha a i es o he sys em is ansla ed and co ec ed (we use he
e e ence sen ence ins ead). A loss unc ion is compu ed in o de o measu e
he dis ance be ween bo h sen ences. Once we ha e his, we use i o adap he
sys em, aiming o a oid he commi ed mis akes. In o de o e alua e how well
hese new algo i hms beha e, a compa ison be ween hese algo i hms and well-
known OL g adien descen a ian s has been ca ied ou . Resul s show ha OL
echniques help o e ine and adap he sys em on- he- ly, inc easing he quali y
o ansla ions and educing he human e o in a pos -edi ing scena io. We ha e
ound ha he p oposed PA-based algo i hms o e a compe i i e pe o mance:
hey imp o e he quali y o ansla ions in all asks. Ne e heless, adap i e SGD-
based algo i hms pe o med gene ally be e .
Finally, pa o he wo k de eloped in his hesis has been submi ed o he
2017 Con e ence on Empi ical Me hods in Na u al Language P ocessing (Pe is
e al.,2017a). Mo eo e , we plan o submi an a icle o a jou nal (ye o be de e -
mined).
Bibliog aphy
Ba, J. L., Ki os, J. R., and Hin on, G. E. (2016). Laye no maliza ion. a Xi p ep in
a Xi :1607.06450.
Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neu al machine ansla ion by
join ly lea ning o align and ansla e. a Xi p ep in a Xi :1409.0473.
Bane jee, S. and La ie, A. (2005). Me eo : An au oma ic me ic o m e alua-
ion wi h imp o ed co ela ion wi h human judgmen s. In P oceedings o he
acl wo kshop on in insic and ex insic e alua ion measu es o machine ansla ion
and/o summa iza ion, olume 29, pages 65–72.
Ba achina, S., Bende , O., Casacube a, F., Ci e a, J., Cubel, E., Khadi i, S., La-
ga da, A., Ney, H., Tomás, J., Vidal, E., e al. (2009). S a is ical app oaches o
compu e -assis ed ansla ion. Compu a ional Linguis ics, 35:3–28.
Bengio, Y., Ducha me, R., Vincen , P., and Jau in, C. (2003). A neu al p obabilis ic
language model. Jou nal o machine lea ning esea ch, 3:1137–1155.
Bengio, Y., Sima d, P., and F asconi, P. (1994). Lea ning long- e m dependencies
wi h g adien descen is di icul . IEEE ansac ions on neu al ne wo ks, 5(2):157–
166.
Ben i ogli, L., Bisazza, A., Ce olo, M., and Fede ico, M. (2016). Neu al e -
sus ph ase-based machine ansla ion quali y: a case s udy. a Xi p ep in
a Xi :1608.04631.
Bishop, C. M. (2006). Pa e n ecogni ion and machine lea ning. sp inge .
Blum, A. (1998). On-line algo i hms in machine lea ning. In Online algo i hms,
pages 306–325. Sp inge .
Bo ou, L. (1991). S ochas ic g adien lea ning in neu al ne wo ks. P oceedings o
Neu o-Nımes, 91.
Boyd, S., Xiao, L., and Mu apcic, A. (2003). Subg adien me hods. lec u e no es o
EE392o, S an o d Uni e si y, Au umn Qua e , 2004.
B i z, D., Goldie, A., Luong, T., and Le, Q. (2017). Massi e explo a ion o neu al
machine ansla ion a chi ec u es. a Xi p ep in a Xi :1703.03906.
B own, P. F., Cocke, J., Pie a, S. A. D., Pie a, V. J. D., Jelinek, F., La e y, J. D.,
Me ce , R. L., and Roossin, P. S. (1990). A s a is ical app oach o machine ans-
la ion. Compu a ional linguis ics, 16(2):79–85.
41
42
BIBLIOGRAPHY
B own, P. F., Pie a, V. J. D., Pie a, S. A. D., and Me ce , R. L. (1993). The ma h-
ema ics o s a is ical machine ansla ion: Pa ame e es ima ion. Compu a ional
linguis ics, 19(2):263–311.
Cas año, A. and Casacube a, F. (1997). A connec ionis app oach o machine
ansla ion. In Fi h Eu opean Con e ence on Speech Communica ion and Technology.
Chen, S. F. and Joshua, G. (1998). An empi ical s udy o smoo hing echniques
o language modeling. Technical epo , Ha a d Compu e Science G oup.
Chiang, D. (2012). Hope and ea o disc imina i e aining o s a is ical ansla-
ion models. Jou nal o Machine Lea ning Resea ch, 13:1159–1187.
Cho, K., Van Me iënboe , B., Gulceh e, C., Bahdanau, D., Bouga es, F., Schwenk,
H., and Bengio, Y. (2014). Lea ning ph ase ep esen a ions using nn encode -
decode o s a is ical machine ansla ion. a Xi p ep in a Xi :1406.1078.
Cholle , F. e al. (2015). Ke as.
h ps://gi hub.com/ cholle /ke as
.
Chow, Y.-L. and Schwa z, R. (1989). The n-bes algo i hm: An e icien p ocedu e
o inding op n sen ence hypo heses. In P oceedings o he wo kshop on Speech
and Na u al Language, pages 199–202. Associa ion o Compu a ional Linguis-
ics.
Chung, J., Gulceh e, C., Cho, K., and Bengio, Y. (2014). Empi ical e alua ion
o ga ed ecu en neu al ne wo ks on sequence modeling. a Xi p ep in
a Xi :1412.3555.
C amme , K., Dekel, O., Keshe , J., Shale -Shwa z, S., and Singe , Y. (2006). On-
line passi e-agg essi e algo i hms. Jou nal o Machine Lea ning Resea ch, 7:551–
585.
C amme , K. and Singe , Y. (2003). Ul aconse a i e online algo i hms o mul-
iclass p oblems. Jou nal o Machine Lea ning Resea ch, 3:951–991.
Duchi, J., Hazan, E., and Singe , Y. (2011). Adap i e subg adien me hods o on-
line lea ning and s ochas ic op imiza ion. Jou nal o Machine Lea ning Resea ch,
12:2121–2159.
Elman, J. L. (1990). Finding s uc u e in ime. Cogni i e science, 14(2):179–211.
Fede ico, M., Ben i ogli, L., Paul, M., and S üke , S. (2011). O e iew o he
IWSLT e alua ion campaign. In P oceedings o he In e na ional Wo kshop on Spo-
ken Language T ansla ion, pages 11–27.
Gen ile, C. (2001). A new app oxima e maximal ma gin classi ica ion algo i hm.
Jou nal o Machine Lea ning Resea ch, 2:213–242.
Golik, P., Doe sch, P., and Ney, H. (2013). C oss-en opy s. squa ed e o ain-
ing: a heo e ical and expe imen al compa ison. In P oceedings o he In e speech,
pages 1756–1760.
G a es, A. (2011). P ac ical a ia ional in e ence o neu al ne wo ks. In Ad ances
in Neu al In o ma ion P ocessing Sys ems, pages 2348–2356.
BIBLIOGRAPHY
43
Guo, J. (2013). Backp opaga ion h ough ime. Unpubl. ms., Ha bin Ins i u e o
Technology.
Hin on, G. E. (1986). Lea ning dis ibu ed ep esen a ions o concep s. In P o-
ceedings o he eigh h annual con e ence o he cogni i e science socie y, olume 1,
page 12. Amhe s , MA.
Hoch ei e , S., Bengio, Y., F asconi, P., and Schmidhube , J. (2001). G adien low
in ecu en ne s: he di icul y o lea ning long- e m dependencies.
Hoch ei e , S. and Schmidhube , J. (1997). Long sho - e m memo y. Neu al com-
pu a ion, 9(8):1735–1780.
Hu chins, J. (2005). The his o y o machine ansla ion in a nu shell. Re ie ed
Decembe , 20:2009.
Jean, S., Fi a , O., Cho, K., Memise ic, R., and Bengio, Y. (2015). Mon eal neu al
machine ansla ion sys ems o wm ’15. In P oceedings o Empi ical Me hods in
Na u al Language P ocessing, pages 134–140.
Jo dan, M. I. (1986). A ac o dynamics and pa allellism in a connec ionis se-
quen ial machine.
Kingma, D. and Ba, J. (2014). Adam: A me hod o s ochas ic op imiza ion. a Xi
p ep in a Xi :1412.6980.
Ki inen, J., Smola, A. J., and Williamson, R. C. (2004). Online lea ning wi h ke -
nels. IEEE ansac ions on signal p ocessing, 52:2165–2176.
Knese , R. and Ney, H. (1995). Imp o ed backing-o o m-g am language mod-
eling. In Acous ics, Speech, and Signal P ocessing, olume 1, pages 181–184. IEEE.
Knowles, R. and Koehn, P. (2016). Neu al in e ac i e ansla ion p edic ion. In
P oceedings o he Con e ence o he Associa ion o Machine T ansla ion in he Ame -
icas (AMTA), olume 1, pages 107–120.
Koehn, P. (2004). S a is ical signi icance es s o machine ansla ion e alua ion.
In P oceedings o Empi ical Me hods in Na u al Language P ocessing, pages 388–
395.
Koehn, P. (2010). S a is ical machine ansla ion.
La ie, M. D. C. D. A. (2014). Lea ning om pos -edi ing: Online model adap a-
ion o s a is ical machine ansla ion. P oceedings o he 14 h Con e ence o he
Eu opean Chap e o he Associa ion o Compu a ional Linguis ics (EACL 14), page
395.
LeCun, Y., Bengio, Y., and Hin on, G. (2015). Deep lea ning. Na u e, 521:436–444.
Li, Y. and Long, P. M. (2000). The elaxed online maximum ma gin algo i hm. In
Ad ances in neu al in o ma ion p ocessing sys ems, pages 498–504.
Lu, J., Zhao, P., and Hoi, S. C. (2016). Online passi e-agg essi e ac i e lea ning.
Machine Lea ning, 103:141–183.
44
BIBLIOGRAPHY
Luong, M.-T. and Manning, C. D. (2015). S an o d neu al machine ansla ion sys-
ems o spoken language domains. In P oceedings o he In e na ional Wo kshop
on Spoken Language T ansla ion.
Madsen, M. W. (2009). The limi s o machine ansla ion ( hesis). Cen e o Lan-
guage Technology, Uni . o Copenhagen, Copenhagen.
Ma ínez-Gómez, P., Sanchis-T illes, G., and Casacube a, F. (2011). Online lea n-
ing ia dynamic e anking o compu e assis ed ansla ion. Compu a ional
Linguis ics and In elligen Tex P ocessing, pages 93–105.
Ma ínez-Gómez, P., Sanchis-T illes, G., and Casacube a, F. (2012). Online adap-
a ion s a egies o s a is ical machine ansla ion in pos -edi ing scena ios.
Pa e n Recogni ion, 45(9):3193–3203.
McCloskey, M. and Cohen, N. J. (1989). Ca as ophic in e e ence in connec ionis
ne wo ks: The sequen ial lea ning p oblem. Psychology o lea ning and mo i a-
ion, 24:109–165.
Mikolo , T., Chen, K., Co ado, G., and Dean, J. (2013a). E icien es ima ion o
wo d ep esen a ions in ec o space. a Xi p ep in a Xi :1301.3781.
Mikolo , T., Ka a iá , M., Bu ge , L., Ce nock`
y, J., and Khudanpu , S. (2010). Re-
cu en neu al ne wo k based language model. In P oceedings o he In e speech,
olume 2, page 3.
Mikolo , T., Komb ink, S., Bu ge , L., ˇ
Ce nock`
y, J., and Khudanpu , S. (2011).
Ex ensions o ecu en neu al ne wo k language model. In Acous ics, Speech
and Signal P ocessing, pages 5528–5531. IEEE.
Mikolo , T., Yih, W.- ., and Zweig, G. (2013b). Linguis ic egula i ies in con in-
uous space wo d ep esen a ions. In P oceedings o he 2013 Con e ence o he
No h Ame ican Chap e o he Associa ion o Compu a ional Linguis ics: Human
Language Technologies (NAACL-HLT-2013), olume 13, pages 746–751.
Och, F. J. (2003). Minimum e o a e aining in s a is ical machine ansla ion. In
P oceedings o he Annual Mee ing o he Associa ion o Compu a ional Linguis ics,
pages 160–167.
Papineni, K., Roukos, S., Wa d, T., and Zhu, W.-J. (2002). Bleu: a me hod o
au oma ic e alua ion o machine ansla ion. In P oceedings o he 40 h annual
mee ing on associa ion o compu a ional linguis ics, pages 311–318.
Pascanu, R., Mikolo , T., and Bengio, Y. (2013). On he di icul y o aining ecu -
en neu al ne wo ks. In e na ional Con e ence on Machine Lea ning (3), 28:1310–
1318.
Penning on, J., Soche , R., and Manning, C. D. (2014). Glo e: Global ec o s o
wo d ep esen a ion. In P oceedings o Empi ical Me hods in Na u al Language
P ocessing, olume 14, pages 1532–1543.
Pe is, Á. (2017). NMT-Ke as.
h ps://gi hub.com/l apeab/nm -ke as
. Gi Hub
eposi o y.
BIBLIOGRAPHY
45
Pe is, Á., Ceb ián, L., and Casacube a, F. (2017a). Online lea ning o neu al
machine ansla ion pos -edi ing. a Xi p ep in a Xi :1706.03196.
Pe is, Á., Domingo, M., and Casacube a, F. (2017b). In e ac i e neu al machine
ansla ion. Compu e Speech & Language, 45:201–220.
Qian, N. (1999). On he momen um e m in g adien descen lea ning algo i hms.
Neu al ne wo ks, 12(1):145–151.
Rosenbla , F. (1958). The pe cep on: A p obabilis ic model o in o ma ion s o -
age and o ganiza ion in he b ain. Psychological e iew, 65:386.
Rumelha , D. E., Hin on, G. E., and Williams, R. J. (1988). Lea ning ep esen a-
ions by back-p opaga ing e o s. Cogni i e modeling, 5(3):1.
Samson, R. (2005). Compu e -assis ed ansla ion. T aining o he New Millennium:
Pedagogies o ansla ion and in e p e ing, 60:101.
Schus e , M. and Paliwal, K. K. (1997). Bidi ec ional ecu en neu al ne wo ks.
IEEE T ansac ions on Signal P ocessing, 45(11):2673–2681.
Shen, S., Cheng, Y., He, Z., He, W., Wu, H., Sun, M., and Liu, Y. (2016). Minimum
isk aining o neu al machine ansla ion. In P oceedings o he 54 h Annual
Mee ing o he Associa ion o Compu a ional Linguis ics, pages 1683–1692.
Sho , N., Zhu benko, N., Likho id, A., and S e syuk, P. (2003). Algo i hms o
nondi e en iable op imiza ion: De elopmen and applica ion. Cybe ne ics and
Sys ems Analysis, 39:537–548.
Sno e , M., Do , B., Schwa z, R., Micciulla, L., and Makhoul, J. (2006). A s udy
o ansla ion edi a e wi h a ge ed human anno a ion. In P oceedings o asso-
cia ion o machine ansla ion in he Ame icas, olume 200, pages 223–231.
Su ske e , I., Vinyals, O., and Le, Q. V. (2014). Sequence o sequence lea ning
wi h neu al ne wo ks. In Ad ances in neu al in o ma ion p ocessing sys ems, pages
3104–3112.
Ta sumi, M. (2009). Co ela ion be ween au oma ic e alua ion me ic sco es, pos -
edi ing speed, and some o he ac o s. P oceedings o he Twel h Machine T ans-
la ion Summi (MT-Summi XII), pages 332–339.
Theano De elopmen Team (2016). Theano: A Py hon amewo k o as com-
pu a ion o ma hema ical exp essions. a Xi e-p in s, abs/1605.02688.
Tiedemann, J. (2009). News om opus-a collec ion o mul ilingual pa allel co -
po a wi h ools and in e aces. In Recen ad ances in na u al language p ocessing,
olume 5, pages 237–248.
Zeile , M. D. (2012). Adadel a: an adap i e lea ning a e me hod. a Xi p ep in
a Xi :1212.5701.
Zens, R., Och, F. J., and Ney, H. (2002). Ph ase-based s a is ical machine ansla-
ion. In Annual Con e ence on A i icial In elligence, pages 18–32. Sp inge .