scieee Science in your language
[en] (orig)

Multilingual Email Zoning - Segmenting Multilingual Email Text Into Zones

Abstract

The segmentation of emails into functional zones (also dubbed email zoning) is a relevant preprocessing step for most NLP tasks that deal with emails. In this research, we analyze in depth the email zoning literature and develop a business case around CLEVERLY AI, a company from the Customer Service sector. We design a new email zoning classification schema and collect a multilingual corpus of emails from CLEVERLY AI clients. We develop five neural network-based email zoning systems, among those systems, we introduce OKAPI, the first multilingual email zoning model based on a language agnostic sentence encoder. Besides outperforming our other systems when tested on CLEVERLY’s emails, OKAPI shows competitive performances with current English public benchmarks and reached new state-of-the-art results for English domain adaptation tasks. Moreover, we release a new multilingual benchmark, composed of 625 emails in Portuguese, Spanish and French, and demonstrate OKAPI can effectively generalize its learnings for unseen languages.

Read accessible full text

Multilingual Email Zoning - Segmenting Multilingual Email Text Into Zones

Author: Jardim, João Bruno Morais de Sousa
Year: 2021
Source: https://run.unl.pt/bitstream/10362/119831/1/TGI0412.pdf
i
Mul ilingual Email Zoning
João B uno Mo ais de Sousa Ja dim
Segmen ing Mul ilingual Email Tex In o Zones
Disse a ion p esen ed as pa ial equi emen o ob aining
he Mas e ’s deg ee in In o ma ion Managemen
ii


NOVAIn o ma ionManagemen School
Ins i u oSupe io deEs a ís icaeGes ãodeIn o mação
Uni e sidadeNo adeLisboa

MULTILINGUALEMAILZONING
Segmen ingMul ilingualEmailIn oZones

by

JoãoB unoMo aisdeSousaJa dim





Disse a ion p esen ed as he pa ial equi emen  o  ob aining a Mas e 's deg ee in In o ma ion
Managemen ,Specializa ioninKnowledgeManagemen andBusinessIn elligence


Ad iso :P o esso aDou o aMa ianaSáCo eiaLei edeAlmeida
Co‐supe iso :Rica doCos aDiasRei



 Ma ch2021 

iii
ACKNOWLEDGEMENTS
I wan o hank Nuno Ca nei o, Rica do Rei, Ma iana Almeida, he en i e Cle e ly AI eam and all
Cle e ly anno a o s, my Mas e ’s colleagues and eache s, my lo ing pa en s and amily, E a and my
dea es iends.
This p ojec has ecei ed unding om he Eu opean Union’s Ho izon 2020 esea ch and inno a ion
p og am unde g an ag eemen No 873904.
i
ABSTRACT
The segmen a ion o emails in o unc ional zones (also dubbed email zoning) is a ele an
p ep ocessing s ep o mos NLP asks ha deal wi h emails. In his esea ch, we analyze in dep h he
email zoning li e a u e and de elop a business case a ound CLEVERLY AI, a company om he
Cus ome Se ice sec o . We design a new email zoning classi ica ion schema and collec a
mul ilingual co pus o emails om CLEVERLY AI clien s. We de elop i e neu al ne wo k-based email
zoning sys ems, among hose sys ems, we in oduce OKAPI, he i s mul ilingual email zoning model
based on a language agnos ic sen ence encode . Besides ou pe o ming ou o he sys ems when
es ed on CLEVERLY’s emails, OKAPI shows compe i i e pe o mances wi h cu en English public
benchma ks and eached new s a e-o - he-a esul s o English domain adap a ion asks. Mo eo e ,
we elease a new mul ilingual benchma k, composed o 625 emails in Po uguese, Spanish and
F ench, and demons a e OKAPI can e ec i ely gene alize i s lea nings o unseen languages.
KEYWORDS
Na u al Language P ocessing; Machine Lea ning; Email Zoning; Tex Segmen a ion; Mul ilingual;
Cus ome Se ice

i
INDEX
1. In oduc ion .................................................................................................................. 1
1.1. Mo i a ion ............................................................................................................. 1
1.1.1. Email Zoning ................................................................................................... 1
1.1.2. Cle e ly Case S udy ........................................................................................ 2
1.2. Objec i es and Me hodology ................................................................................ 3
2. Backg ound ................................................................................................................... 6
2.1. Supe ised Machine Lea ning ............................................................................... 6
2.1.1. Pe cep on ...................................................................................................... 7
2.1.2. Mul ilaye Pe cep on .................................................................................... 8
2.1.3. Backp opaga ion............................................................................................. 8
2.1.4. O e i ing and Regula iza ion ....................................................................... 9
2.1.5. Con olu ional Neu al Ne wo ks ................................................................... 10
2.1.6. Condi ional Random Fields ........................................................................... 11
2.1.7. Recu en Neu al Ne wo ks ......................................................................... 12
2.1.8. A en ion Mechanism ................................................................................... 14
2.1.9. The T ans o me ........................................................................................... 14
2.2. Tex Rep esen a ion Models ............................................................................... 15
2.2.1. Spa se Models .............................................................................................. 16
2.2.2. Dense Models ............................................................................................... 17
3. Rela ed Wo k .............................................................................................................. 22
3.1. Email Zoning ........................................................................................................ 22
3.1.1. JANGADA ......................................................................................................... 22
3.1.2. ZEBRA ............................................................................................................. 24
3.1.3. QUAGGA .......................................................................................................... 26
3.1.4. CHIPMUNK ....................................................................................................... 28
3.1.5. Email Zoning Public Co po a ........................................................................ 30
3.2. Tex Segmen a ion .............................................................................................. 32
4. Me hodology .............................................................................................................. 35
4.1. Co po a ................................................................................................................ 35
4.1.1. CLEVERLY AI Co pus ......................................................................................... 35
4.1.2. Public Co po a .............................................................................................. 39
4.1.3. New Mul ilingual Email Zoning Co pus ........................................................ 41
4.2. Models ................................................................................................................. 43
ii
4.2.1. Baseline, Wo d Embeddings + BiLSTM (W-BiLSTM) ..................................... 43
4.2.2. Wo d and Subwo d Embeddings + BiLSTM (WSw-BiLSTM) ......................... 46
4.2.3. XLM-RoBERTa Embeddings + BiLSTM (XLMR-BiLSTM) ................................. 46
4.2.4. XLM-RoBERTa Embeddings + BiLSTM + CRF (XLMR-BiLSTM-CRF o OKAPI) 48
4.3. E alua ion Me ics ............................................................................................... 48
4.3.1. Email Zoning ................................................................................................. 48
4.3.2. In e -anno a o Ag eemen ......................................................................... 50
5. Resul s and discussion ................................................................................................ 53
5.1. Expe imen al Se up ............................................................................................. 53
5.2. Cle e ly Resul s .................................................................................................... 54
5.2.1. CLEVERLY Anno a ed Co pus........................................................................... 54
5.2.2. Bes Model Analysis ..................................................................................... 56
5.2.3. Impac in CLEVERLY Pipeline ........................................................................... 58
5.3. Public Co po a Resul s ......................................................................................... 59
5.3.1. English Co po a ............................................................................................ 59
5.3.2. Mul ilingual Co pus ...................................................................................... 61
6. Conclusions ................................................................................................................. 64
7. Limi a ions and ecommenda ions o u u e wo ks ................................................. 66
8. Bibliog aphy ................................................................................................................ 68
iii
LIST OF FIGURES
Figu e 1.1 – Example email wi h wo iden i ied unc ional segmen s: au ho ed con en and
ad e isemen . ................................................................................................................... 1
Figu e 1.2 – O e iew o CLEVERLY AI icke classi ica ion pipeline wi h email zoning as a
p ep ocessing s ep. ............................................................................................................ 2
Figu e 2.1 – Un olded Recu en Neu al Ne wo k. Taken om h ps:// inyu l.com/yy79zmxo.
.......................................................................................................................................... 12
Figu e 2.2 – In e nal Rep esen a ion o an LSTM cell. Taken om
h ps:// inyu l.com/yy79zmxo ......................................................................................... 13
Figu e 2.3 – The a chi ec u e o he T ans o me Model. Taken om
h ps:// inyu l.com/y2j w3m3. ....................................................................................... 15
Figu e 2.4 – BERT-Base and BERT-La ge Encode S ack. Taken om
h ps:// inyu l.com/y3we bz7 ......................................................................................... 19
Figu e 3.1 – Example o a JANGADA labeled email message adap ed om Ca alho & Cohen
(2004). .............................................................................................................................. 23
Figu e 3.2 – Example o a labeled email message wi h bo h h ee- and nine-zone
classi ica ion adap ed om Lampe e al. (2009). .......................................................... 25
Figu e 3.3 – Example o a QUAGGA labeled email message wi h bo h wo- and i e-zone
anno a ions, adap ed om Repke & K es el (2018). ....................................................... 27
Figu e 3.4 – QUAGGA model o e iew. Taken om Repke & K es el (2018). ........................... 28
Figu e 3.5 – Example o a CHIPMUNK labeled email message, adap ed om Be endo e al.
(2020). .............................................................................................................................. 29
Figu e 3.6 – CHIPMUNK model a chi ec u e. Taken om Be endo e al. (2020). .................. 29
Figu e 3.7 – P oposed T ans o me -based segmen a ion models (Lukasik e al., 2020). ....... 33
Figu e 4.1 – Exce p om a CLEVERLY labeled email message. Pe sonal and o he sensible
ins ances we e eplaced by a de aul oken ha indica es hei con en . ...................... 36
Figu e 4.2 – Pe cen age o o al lines o email zone, o each language in CLEVERLY co pus. . 37
Figu e 4.3 – Pe cen age o he o al lines pe zone. Compa ison be ween he En on co pus
(g een/le ) and he ASF co pus (yellow/ igh ). ............................................................... 40
Figu e 4.4 – W-BiLSTM model o e iew. The model is di ided in wo pa s: 1) a sen ence-
le el wo d2 ec wo d encode ; and 2) a segmen a ion module ha uses a BiLSTM and a
so max ou pu laye o classi y each sen ence in o an email zone. Al hough he BiLSTM
ecei es he sequence o sen ences in an email, o simplici y, we illus a e he p ocess
o a single sen ence. ....................................................................................................... 44
ix
Figu e 4.5 – Sw-BiLSTM model o e iew. The model uses he same a chi ec u e as he W-
BiLSTM bu encodes sen ences a he subwo d-le el. .................................................... 45
Figu e 4.6 – WSw-BiLSTM model o e iew. The model p oduces pa allel sen ence
ep esen a ions a wo d and subwo d le el. The ep esen a ions a e conca ena ed o
each sen ence and ed in o he ou pu laye . .................................................................. 46
Figu e 4.7 – O e iew o XLM-RoBERTa embedding ex ac ion s eps. ................................... 47
Figu e 4.8 – XLMR-BILSTM is composed o wo building blocks: 1) a mul ilingual sen ence
encode (XLM-RoBERTa) o de i e sen ence embeddings; and 2) a segmen a ion
module ha uses a BiLSTM and a so max ou pu laye o classi y each sen ence in o an
email zone. ....................................................................................................................... 47
Figu e 4.9 – OKAPI model o e iew. OKAPI ollows he same a chi ec u e as XLMR-BiLSTM
model excep o he ou pu laye , in which i uses CRF o classi y each sen ence in o an
email zone. ....................................................................................................................... 48
Figu e 5.1 – Con usion ma ix o OKAPI’s email zoning esul s on CLEVERLY’s mul ilingual
co pus. On he le he ue labels and on he bo om he p edic ed labels. The da ke
he squa e, he mo e lines a e p edic ed wi h he column’s label. ................................. 57
3
company is p ep ocessing icke s by mixing a se o hand coded ules wi h he machine lea ning
package Talon
2
, in o de o emo e quo ed ex and signa u e lines. Ne e heless, his ype o
app oach may no be su icien ly dynamic o a company ha deals wi h an e e -eme ging clien
base wi h icke s di e ing in he email o mal layou and language.
This way, he CLEVERLY eam belie es ha hey will bene i om he implemen a ion o an
email zoning solu ion based on an ad anced machine lea ning model, one ha is mo e lexible o
changes in he icke s ea u es and able o deal wi h mul iple languages.
1.2. OBJECTIVES AND METHODOLOGY
Ou objec i e wi h his esea ch is o build a case s udy a ound he Cus ome Se ice company
CLEVERLY AI o show he e ec i eness o a mul ilingual email zoning sys em applied o an email
classi ica ion pipeline. This case s udy en iches his esea ch wi h a eal-li e business scena io, which
is ep esen a i e o many o he business ci cums ances. Mo eo e , we wan o con ibu e o he
email zoning li e a u e by mi iga ing he p oblems o he English-cen ic co po a, lack o a s anda d
email zoning axonomy and absence o a mul ilingual email zoning sys em. Hence, ou con ibu ions
a e he ollowing:
1. We discuss he exis ing email zoning co po a, he co esponding email zoning classi ica ion
schemas, he exis ing email zoning sys ems, and hei limi a ions.
2. We design a new email zoning classi ica ion schema o CLEVERLY and anno a e zones o
15,547 icke s in 5 languages – English, Po uguese, Spanish, F ench and I alian.
3. We c ea e he i s publicly a ailable mul ilingual anno a ed co pus o email zoning. This
co pus consis s o 625 emails in 3 languages - Po uguese, Spanish and F ench - and
encompasses 15 email zones.
4. We in oduce i e email zoning sys ems, among hem OKAPI, a mul ilingual email
segmen a ion sys em buil on op o XLM-RoBERTa (Conneau e al., 2020) ha can be easily
ex ended o 100 languages. To he bes o ou knowledge, OKAPI is he i s end- o-end
mul ilingual sys em explo ing p e- ained ans o me models (Vaswani e al., 2017) o
pe o m email zoning.
5. We use ou i e sys ems o segmen CLEVERLY’s anno a ed icke s, e alua e and compa e hei
e ec i eness in segmen ing CLEVERLY’s icke s. We also desc ibe he ime aken by ou bes
sys em (OKAPI) o p ocess a o al o 84,929 icke s and compa e i s pe o mance agains
CLEVERLY’s cu en solu ion by measu ing he pe o mance o he downs eam icke
classi ica ion model wi h each app oach.
6. We es OKAPI agains o he sys ems om he li e a u e in di e en email zoning co po a
including ou new mul ilingual co pus, showing ha besides e ec i ely gene alizing o
unseen languages, OKAPI eaches compe i i e esul s wi h cu en English benchma ks and
s a e-o - he-a pe o mances o domain adap a ion asks.
A de ailed accoun o each o he subjec s lis ed abo e is gi en in he emainde o his
hesis: Backg ound p esen s a e iew o he concep s ha a e he basis o he de elopmen o his
2
h ps://gi hub.com/mailgun/ alon

4
esea ch; Rela ed Wo k discusses p e iously implemen ed email zoning me hods and p o ides a
comp ehensi e e iew o exis ing email zoning co po a, as well as ele an wo k in he a ea o ex
segmen a ion; Me hodology de ines he zoning schema adop ed o CLEVERLY’s con ex and p esen s
a ho ough analysis o CLEVERLY’s co pus and o he public co po a, namely ou new mul ilingual
co pus. This sec ion also p esen s he email zoning sys ems (models) de eloped and he e alua ion
me ics used o measu e hei pe o mance. Resul s and discussion epo s and discusses he
expe imen s done and esul s achie ed o CLEVERLY co po a and public co po a; Conclusions
summa izes and concludes he wo k de eloped h oughou his esea ch; and Limi a ions and
ecommenda ions o u u e wo ks add esses limi a ions o ou wo k and in e es ing ideas o be
de eloped in he u u e.
Finally, we mus say ha du ing he de elopmen o his esea ch, we ha e had a pape
accep ed a EACL S uden Resea ch Wo kshop
3
aking place in conjunc ion wi h EACL 2021. The
pape is a ailable a h ps://a xi .o g/abs/2102.00461 and h ps://gi hub.com/cle e ly-
ai/mul ilingual-email-zoning (Ja dim, Rei, & Almeida, 2021).
3
h ps://si es.google.com/ iew/eacls w2021/home
5
6
2. BACKGROUND
This sec ion p esen s he undamen al machine lea ning concep s ha a e he basis o his esea ch.
We e iew supe ised machine lea ning (sec ion 2.1) and discuss how o con e na u al language
ex s in o ep esen a ions capable o being in e p e able by machine lea ning algo i hms (sec ion
2.2).
2.1. SUPERVISED MACHINE LEARNING
Supe ised machine lea ning is a ca ego y o machine lea ning algo i hms ha lea n om labeled
da a o de ine a unc ion ha co ec ly p oduces an ou pu o unlabeled da a. The algo i hms lea n
o c ea e a unc ion 𝑓(𝑥), ha will map a se o a iables 𝑥 = {𝑥1,𝑥2,...,𝑥𝑛} wi h a se o labels 𝑦 =
{𝑦1,𝑦2,...,𝑦𝑚}, by lea ning pa e ns om he supplied da a, i.e. hey unco e gene al hypo heses
om he aining da a, being hen able o apply hose hypo heses and make p edic ions on unseen
da a.
Supe ised lea ning can be di ided in o classi ica ion and eg ession p oblems. This di ision is
ela ed o he na u e o 𝑦 alues. In case 𝑦 is a se o ini e ca ego ical alues and he model seeks o
p edic he ca ego y 𝑥 belongs o, we ha e a classi ica ion p oblem. Some examples o classi ica ion
p oblems include spam de ec ion, aud de ec ion, o sen imen analysis. On he o he hand, i 𝑦 is a
con inuous a iable o eal alues, we ha e a eg ession p oblem. House p ice p edic ion and
empe a u e p edic ion a e examples o eg ession p oblems.
One o he mos simple and popula classi ica ion algo i hms is he K-Nea es Neighbo s
(KNN) (Fix & Hodges, 1989). This algo i hm elies on he assump ion ha simila ins ances belong o
he same ca ego ies. Based on he inpu ea u es, he KNN algo i hm calcula es he dis ance
be ween he new inpu and he aining ins ances using a simila i y unc ion and, inally, assigns he
inpu o he mos common class o i s 𝑘 close aining ins ances, whe e 𝑘 is a posi i e in ege ,
gene ally o small alue. This algo i hm is conside ed non-linea since i s decision bounda ies on he
ea u e space a e no linea .
On he con a y, linea me hods deal wi h linea decision bounda ies. The objec i e o linea
me hods is o model he ela ionship be ween a dependen a iable 𝑦 and he explana o y a iables
𝑥 = {𝑥1,𝑥2,...,𝑥𝑛} as a unc ion 𝑓(𝑥). A Linea Reg ession is an example o a linea me hod. This
me hod, ies o model he ela ionship be ween wo a iables by i ing a linea equa ion o he
obse ed da a:
𝑦 = 𝛽0+𝛽1𝑥(2.1)
Equa ion 2.1 desc ibes a line whe e 𝑦 is he dependen a iable 𝑥 he explana o y a iable,
𝛽0 is he alue o 𝑦 when 𝑥 =0, he in e cep , and 𝛽1 is he slope o he line.
Supe ised lea ning me hods can also be di ided in o gene a i e me hods and disc imina i e
me hods, based on he way classi ie s a e compu ed. We eso o a gene a i e model, i o compu e
a label 𝑦 based on obse a ion 𝑥, we es ima e he join dis ibu ion 𝑃(𝑥,𝑦) and om ha compu e
he p obabili y 𝑃(𝑦|𝑥), as a basis o he classi ie . This means he da a is modeled based on how i
was gene a ed. A simple example o a gene a i e me hod is he naï e Bayes model, de ined below in
equa ion 2.2:
7
𝑃(𝑦|𝑥)=𝑃(𝑥|𝑦)𝑃(𝑦)
𝑃(𝑥)(2.2)
This model assumes ha he p edic o s 𝑥 a e independen and ha e he same impo ance,
hus being called naï e. I hen calcula es he pos e io condi ional p obabili y o 𝑦 ha ing 𝑥 and
assigns 𝑦 o he class wi h he highes p obabili y.
I he p oblem o be modeled equi es he assump ion ha he inpu s a e in e dependen
be ween each o he , i.e. he inpu da a is sequen ial, we canno ollow a naï e app oach like he one
desc ibed using he naï e Bayes model. We need an app oach ha enables he modeling o a
sequence o inpu and ou pu ins ances. Hidden Ma ko Models (HMM) (Baum & Pe ie, 1966) a e
an example o gene a i e models capable o modeling he join dis ibu ion 𝑃(𝑦,𝑥) when dealing
wi h sequen ial da a.
HMM assume ha a andom obse ed a iable is no he s a e o a sys em bu simply da a
gene a ed by hidden s a es o ha sys em. Thus, he occu ence o an obse a ion is condi ioned by
he unde lying s a e. As he name e eals, HMM ely on Ma ko p ocess assump ions (Ma ko ,
1953), which s a e ha i he p esen s a e in a sequence is known, no mo e in o ma ion is needed o
p edic he immedia e u u e s a e. This way, gi en a se o hidden s a es and a se o obse ed
a iables, he model calcula es he p obabili y o a sequence o hidden s a es ( ansi ion
p obabili ies), hen, using he p esen s a e, i can calcula e wha is he obse a ion wi h he highes
p obabili y o occu ing and/o he mos likely nex s a e.
A model can be conside ed as disc imina i e i i di ec ly es ima es he condi ional
p obabili y o 𝑃(𝑦|𝑥) by de e mining a se o pa ame e s 𝜃𝑚𝑜𝑑𝑒𝑙 using aining da a consis ing o
pai s (𝑥,𝑦), i.e. inpu ec o s and hei a ge ou pu , hen using he condi ional p obabili y o make
p edic ions o 𝑦 o a new alue o 𝑥 (Be na do e al., 2007). One o he mos amous classes o
disc imina i e models is he A i icial Neu al Ne wo k. These models inhe i hei name om he
synap ic p ocesses p esen in animal b ains since hey y o ep oduce he p ocesses happening
inside biological ne ous sys ems, pa icula ly wi hin a single neu on (McCulloch & Pi s, 1943).
A i icial neu al ne wo ks a e di ec ed acyclic g aphs, in which each node ecei es in o ma ion,
ans o ms i and passes i o o he nodes.
Di e se ypes o neu al ne wo ks we e de eloped aking in o conside a ion he eme gence o
new asks and he capabili ies p o ided wi h he e e -inc easing compu a ional powe . Some o
hem like he Mul ilaye Pe cep on (Rumelha , Hin on, & Williams, 1986), Con olu ional Ne wo ks
(LeCun e al., 1989), and Recu en Neu al Ne wo ks (Elman, 1990) will be add essed la e in his
sec ion and ha e been showing s a e-o - he-a esul s in a b oad ange o NLP asks, such as
machine ansla ion and ques ion-answe ing (Zhou, Duan, Liu, & Shum, 2020).
2.1.1. Pe cep on
A Pe cep on (Rosenbla , 1958) is he mos basic o m o an a i icial neu al ne wo k, being he
building block o o he neu al ne wo k a ia ions. The pe cep on is composed o an inpu laye , wi h
a weigh ma ix 𝑊, a bias e m 𝑏, and an ac i a ion unc ion 𝑔(𝑧). The pe cep on wo ks by ecei ing
an inpu 𝑥 o 𝑛 nume ic ea u es, aking he do p oduc be ween he inpu and 𝑊, summing he
alues wi h he bias e m 𝑏, leading o a sco e 𝑧 = 𝑥𝑊+𝑏. This sco e is hen passed h ough an
ac i a ion unc ion g(z). The pe cep on is de ined as ollows:
8
𝑃𝑒𝑟𝑐𝑒𝑝𝑡𝑟𝑜𝑛(𝑥)=𝑔(𝑥𝑊+𝑏)(2.3)
Some o he mos common ac i a ion unc ions a e he Logis ic (sigmoid), he Hype bolic
Tangen ( anh), he Rec i ied Linea Uni (ReLU), and he No malized Exponen ial unc ion (so max).
These unc ions allow he Pe cep on o ou pu a nonlinea unc ion. Taking ReLU as an example, he
unc ion can o mally be de ined by he ollowing equa ion:
𝑅(𝑥)=𝑚𝑎𝑥(0,𝑥)(2.4)
This unc ion gi es an ou pu 𝑅(𝑥), whe e he alue is 0 i 𝑥 is less han 0 and 𝑧 i 𝑧 is equal
o abo e 0. Ano he popula ac i a ion unc ion is he so max. This unc ion akes i s inpu ec o
and no malizes i in o numbe s ha can be seen as a p obabili y dis ibu ion. The unc ion wo ks by
aking he exponen s o each elemen 𝑧𝑖 om he p e ious laye ou pu 𝑧 = {𝑧1,...,𝑧𝐾} and hen
no malizing each numbe by he sum o he exponen s, as shown in he equa ion below:
𝜎(𝑧)𝑖=𝑒𝑧𝑖
∑𝑒𝑧𝑗
𝐾
𝑗=1 𝑓𝑜𝑟 𝑖 =1,…,𝐾 (2.5)
The so max is ypically used as he ou pu laye o a neu al ne wo k because i enables he
ne wo k o ou pu a sco e o he inpu alue o a speci ic class. As o he ReLU, al hough i can be
used as he ou pu ac i a ion unc ion, i is mos ly used as an ac i a ion unc ion in in e media y
laye s o mo e complex neu al ne wo ks, such as he Mul ilaye Pe cep on (Rumelha e al., 1986).
2.1.2. Mul ilaye Pe cep on
While in a Pe cep on (Rosenbla , 1958) he e is only one laye whe e ac i a ion o he inpu s
occu s, in a Mul ilaye Pe cep on (MLP) (Rumelha e al., 1986) he e a e wo o mo e laye s o
ans o ma ions. MLPs consis o an inpu laye , an ou pu laye , and, in be ween hose wo, one o
mo e hidden laye s. One can look a MLPs as being composed o mo e han one Pe cep on. Fo each
inpu , he MLP applies a linea ans o ma ion by aking he do p oduc o he inpu and he se o
weigh s 𝑊𝑖 be ween he inpu laye and he hidden laye . These alues a e summed wi h he bias
e m 𝑏𝑖 and passed h ough an ac i a ion unc ion o e e y node in he hidden laye . Once he
ou pu o e e y node in he hidden laye is calcula ed, hese alues a e pushed o he nex hidden
laye o o he ou pu laye , which akes he ou pu om he las hidden laye , pe o ms simila
compu a ions o he ones done in he hidden laye s, and e u ns he ou pu alues o he ne wo k.
2.1.3. Backp opaga ion
The se o Weigh s 𝑊𝑖 and bias e ms 𝑏𝑖 used in-be ween laye s a e he MLP pa ame e s
𝜃𝑚𝑜𝑑𝑒𝑙. To ain his model and imp o e he ne wo k pe o mance o a ce ain ask, he ne wo k
should be able o lea n o adjus 𝜃𝑚𝑜𝑑𝑒𝑙. Backp opaga ion (Linnainmaa, 1976; Rumelha e al., 1986;
We bos, 1974) is a widely used algo i hm o ain neu al ne wo k models. This me hod compu es he
g adien (de i a i e) o he loss unc ion 𝐿(𝑦,𝑦; 𝜃𝑚𝑜𝑑𝑒𝑙) ega ding 𝜃𝑚𝑜𝑑𝑒𝑙 by compa ing he ou pu
alues 𝛾 wi h he co ec answe s 𝑦. The weigh s a e adjus ed o educe he loss unc ion. This
i e a i e p ocess o op imiza ion o ind he minimum o he Loss unc ion is called he G adien
Descen (GD) algo i hm (Cauchy, 1847; Cu y, 1944) and i is desc ibed in Algo i hm 1:

9
𝜃𝑚𝑜𝑑𝑒𝑙 ← 𝑎𝑛𝑦 𝑝𝑜𝑖𝑛𝑡 𝑖𝑛 𝑡ℎ𝑒 𝑝𝑎𝑟𝑎𝑚𝑒𝑡𝑒𝑟 𝑠𝑝𝑎𝑐𝑒
𝒘𝒉𝒊𝒍𝒆 𝐿(𝑦,𝑦;𝜃𝑚𝑜𝑑𝑒𝑙) > 𝜖 𝒅𝒐
𝒇𝒐𝒓 𝑤𝑖 ∈ 𝜃𝑚𝑜𝑑𝑒𝑙 𝒅𝒐
𝑤𝑖← 𝑤𝑖−𝛼 𝑑
𝑑𝑤𝑖𝐿(𝑦,𝑦;𝜃𝑚𝑜𝑑𝑒𝑙)
𝒆𝒏𝒅 𝒇𝒐𝒓
𝒆𝒏𝒅 𝒘𝒉𝒊𝒍𝒆
Algo i hm 1 – G adien Descen Op imiza ion
Algo i hm 1 has a hype -pa ame e
4
lea ning a e 𝛼, used o de ine how much 𝜃𝑚𝑜𝑑𝑒𝑙 should
be changed by each i e a ion. Besides allowing non-linea i y o he p edic ions, non-linea unc ions
also allow o he g adien s o he unc ions o be dependen on he inpu alue, while linea
unc ions ha e a cons an g adien . The numbe o 𝑦 and 𝑦 used o compu e he loss unc ion can
di e , leading o di e en GD e sions. When he algo i hm uses only one sample a a ime o
compu e he loss unc ion, we call his p ocess S ochas ic G adien Descen (SGD) (Kie e &
Wol owi z, 1952; Robbins & Mon o, 1951) . On he o he hand, i he algo i hms e alua e e e y
aining example a each s ep and only hen upda es he ne wo k pa ame e s, he p ocess is called
Ba ch G adien Descen . SGD ends o be compu a ionally as e han Ba ch GD, bu i s highe
numbe o upda es can esul in noisy g adien s.
A hi d app oach ha has been inc easingly used is he Mini-Ba ch GD, which sepa a es he
aining se in o small ba ches and upda es 𝜃𝑚𝑜𝑑𝑒𝑙 o each o hose ba ches, c ea ing a balance
be ween bo h SGD and Ba ch GD app oaches. Va ian s o he SGD ha can wo k wi h ull ba ches o
mini ba ches a e he Roo Mean Squa e P opaga ion (RMSP op) and he Adap i e Momen
Es ima ion (Adam) (Kingma & Ba, 2015).
2.1.4. O e i ing and Regula iza ion
One o he mos p ominen p oblems p esen in neu al ne wo ks wi h a conside able numbe o
hidden laye s (deep neu al ne wo ks) is he lack o con ol o e he lea ning p ocess. E en hough
he high numbe o pa ame e s and eedom o lea n a e he main cha ac e is ics ha enable neu al
ne wo ks o i complex p oblems, hey also come wi h he o e i ing d awback. O e i ing occu s
when he model closely i s he aining da a bu has di icul y gene alizing i s lea ning o unseen
da a examples. Fo ne wo ks like he MLP ha ha e complex hidden s uc u es, he abili y o ex ac
ea u es om he aining da a can lead he model o lea n i ele an ea u es ha ep esen
andomness p esen in he aining da ase . The model will hen make p edic ions based on ha
noise, which will no hold o new da a.
A me hod ha allows o check i a model is o e i ing is o di ide he aining se in o h ee
pa s – ain se , alida ion se , and es se . The pe cen age o aining ins ances ha should be held
by each se depends on he size o he whole da ase and complexi y o he model, al hough a
common di ision is a spli o 60%, 20%, and 20%, espec i ely. The model lea ns om he ain se
and uses he alida ion se o ack p og ess o each lea ning epoch o op imize i s pe o mance. As
o he es se , i is used a e he model is ained o measu e i s pe o mance and check o
o e i ing in case he e o a e on he alida ion se is much lowe han he one on he es se .
4
A hype pa ame e is a pa ame e whose alue is se be o e he lea ning p ocess begins
10
One common echnique o ace an o e i ing model is o inc ease he amoun o aining
da a. This will gi e he model he abili y o lea n om a aining se ha is a be e gene aliza ion o
unseen da a. Ne e heless, i is no always he case ha mo e da a is a ailable o ain he model.
Ano he op ion is o use ano he da a se spli ing echnique, such as k- old C oss- alida ion
(Mos elle & Tukey, 1968), which uses a aining and alida ion se spli , and o each k lea ning
i e a ion changes he ins ances ha a e in each spli , a e aging he alida ion esul s o e he
i e a ions.
Regula iza ion is ye ano he common me hod used o a oid o e i ing, and i consis s o
cons aining he complexi y o a model by adding a egula iza ion e m o he loss unc ion. An
immensely popula egula iza ion me hod used o p e en neu al ne wo ks om o e i ing is he
d opou (S i as a a, Hin on, K izhe sky, & Salakhu dino , 2014). Al hough d opou is conside ed a
egula iza ion me hod, i does no di ec ly add a e m o he loss unc ion, ins ead, e e y uni o a
neu al ne wo k (excep o he ou pu uni s) has a p obabili y 𝑝 o being igno ed du ing he lea ning
p ocess. This will esul in a smalle ne wo k compa ed o he o iginal one, called a hinned ne wo k.
Fo e e y lea ning sample, di e en nodes a e d opped, and new hinned ne wo ks a e used
o aining. This means ha , o each i e a ion, a node will ecei e di e en combina ions o inpu s.
A es ime, d opou is no used - a single ne wo k is used o aining. The ou going weigh s o each
uni a e mul iplied by 𝑝 o ensu e he ou pu o a hidden uni in es ime co esponds o he one a
aining ime.
The ad an age o his me hod is ha i allows o combine se e al neu al ne wo k models
ha sha e he same se o hype pa ame e s and can be ained wi hou addi ional compu a ion size,
inc easing gene aliza ion, and a oiding o e i ing.
2.1.5. Con olu ional Neu al Ne wo ks
Ano he ype o neu al ne wo k is he Con olu ional Neu al Ne wo k (CNN). The simples o m o a
CNN a chi ec u e was p oposed by Fukushima (1980), inspi ed by p e ious disco e ies abou he
isual co ex o mammals. A e ha , o he e inemen s modi ied CNNs, such as backp opaga ion
aining (LeCun e al., 1989). The popula i y and applica ion o CNNs comes mos ly om asks in he
a eas o image and ideo ecogni ion, al hough i s use has been inc easing ac oss NLP asks.
The ne wo k is mos ypically composed o a Con olu ional laye , a Pooling laye , and a Fully
Connec ed laye (ano he name o MLP). The Con olu ional laye , which is he building block o a
CNN, has as pa ame e s 𝜃𝑚𝑜𝑑𝑒𝑙 a se o lea nable il e s (ke nels) ha ocus on a pa ch o he inpu a
he ime and, du ing he o wa d pass o he algo i hm, ake he do p oduc be ween hei alues
and he inpu pa ch alues, in a p ocess called con olu ion, p oducing a ma ix called ac i a ion map
and educing he dimensionali y o he inpu . Applying his idea o a NLP ask wi h a one dimensional
con olu ional, we de ine an inpu sen ence composed o wo ds as 𝑥 = {𝑥𝑖,𝑥𝑖+1,...,𝑥𝑖+ℎ }.
Le 𝑥𝑖:𝑖+ℎ−1 be a window o wo ds, he Con olu ional laye il e 𝑤 is applied o he window o wo ds
gene a ing a new ea u e 𝑐𝑖, as ollows:
𝑐𝑖=𝑓(𝑤.𝑥𝑖:𝑖+ℎ−1 +𝑏) (2.6)
Whe e 𝑏 is a bias e m and 𝑓 a non-linea unc ion, such as ReLU, ha in oduces non-
linea i y o he ne wo k. This il e is applied o each combina ion o wo ds in he inpu sen ence,
11
wi h he same window size, p oducing a ec i ied ea u e map 𝑐 = [𝑐1,𝑐2,…,𝑐𝑛]. Fo a 𝑛 numbe o
il e s, a ying in window size, 𝑛 ea u e maps a e ob ained. A e compu ing he ea u e maps, hese
a e passed o a pooling laye , ypically max pooling, ha educes he size o he ep esen a ion by
aking he maximum alue 𝑐 =𝑚𝑎𝑥{𝑐} o each map. This way, he model can cap u e he mos
impo an ea u e o each il e . By educing he ea u e size, i educes he numbe o pa ame e s,
he amoun o compu a ion, and con ols o e i ing.
Finally, he ea u es 𝑐 a e passed o he Fully Connec ed (FC) Laye . Inside he FC laye , he
compu a ions a e iden ical o he ones done inside a Pe cep on o an MLP, wi h a so max ac i a ion
unc ion in he ou pu laye ha p oduces he dis ibu ion o he inpu o e classi ica ion classes.
Being 𝑝 = {𝑐𝑖,…,𝑐𝑖+𝑛} he se o ea u es passed o he FC laye , he ou pu 𝑂 is ob ained by
applying he so max 𝑓 o e he sum o he do p oduc be ween he FC laye weigh s 𝑊 and he
inpu alues 𝑝:
𝑧 = 𝑠𝑢𝑚(𝑊𝑝)
𝑂 =𝑓(𝑧) (2.7)
2.1.6. Condi ional Random Fields
MLPs and CNNs a e ypes o ne wo ks called eed o wa d neu al ne wo ks, whe e each neu on in
one laye has only di ec connec ions o he neu on in he nex laye . This means ha he
in o ma ion mo es only o wa d, om he inpu , h ough he hidden nodes, and o he ou pu
nodes. This cha ac e is ic can esul in limi a ions ega ding hei abili y o p ocess sequen ial da a
such as sequences o wo ds. In his espec , Condi ional Random Fields (CRF) (La e y, McCallum, &
Pe ei a, 2001) a e an example o disc imina i e models ha a e sui able o p ocess sequences o
da a.
CRF is used o s uc u ed p edic ion
5
, being g aphically modeled – modeled in a way ha
enables he condi ional dependence s uc u e be ween a iables o be exp essed. A popula ype o
CRF is he linea chain CRF, which conside s sequen ial dependencies be ween he model’s
p edic ions. The aim o a linea chain CRF is o calcula e he condi ional p obabili y 𝑃(𝑦|𝑥) o an
ou pu sequence 𝑦 gi en he inpu sequence 𝑥. Thus, o p edic he ou pu sequence we ex ac he
sequence wi h he highes p obabili y. The CRF o mula o calcula e 𝑝(𝑦|𝑥) can be b oken down in o
wo componen s - a weigh s and ea u es componen and no maliza ion componen :
𝑝(𝑦|𝑥,𝜆)= 1
𝑍(𝑥)𝑒𝑥𝑝∑∑𝜆𝑗𝑓𝑖(𝑥,𝑖,𝑦𝑖−1,𝑦)
𝑗
𝑛
𝑖=1 (2.8)
Whe e 𝑍(𝑥) is he no maliza ion componen ha sums all combina ions o s a e sequences
un il he o al is 1. This ans o ma ion u ns he ou pu in o a p obabili y:
𝑍(𝑥)= ∑∑∑𝜆𝑗
𝑗
𝑛
𝑖=1𝑦𝜖𝑌 𝑓𝑖(𝑥,𝑖,𝑦𝑖−1,𝑦) (2.9)
5
Supe ised machine lea ning echnique used o p edic s uc u ed objec s (sequences), ins ead o
scala disc e e o eal alues
12
F om equa ion 2.9, in he ea u es unc ion 𝑓𝑖(𝑥,𝑖,𝑦𝑖−1,𝑦), 𝑥 ep esen s he se o inpu
ec o s, 𝑖 he posi ion o da a poin s we wan o p edic , 𝑦𝑖−1 he label o he da a poin 𝑖−1 and 𝑦𝑖
he label o da a poin s 𝑖 in 𝑥. The 𝜆 pa ame e ep esen s he weigh s o he ea u e unc ion and i
is es ima ed using a maximum likelihood me hod. To ain he model, i.e op imize he pa ame e s, an
i e a i e unc ion like he G adien descen is used.
2.1.7. Recu en Neu al Ne wo ks
A ype o neu al ne wo k ha is a commonly used me hod o p ocess sequen ial da a is he
Recu en Neu al Ne wo k (RNN) (Elman, 1990). This ne wo k has ecu en connec ions be ween
laye s, inc easing i s abili y o p ocess sequences o a bi a y leng h as inpu . This abili y is enabled
using a eedback connec ion o s o e in o ma ion o e ime s eps, known as he in e nal s a e. This
way, he RNN is cons i u ed o a eed o wa d laye and a ecu en laye . I s mos simple o m, he
Vanilla RNN, can be exp essed wi h he ollowing equa ion:
ℎ𝑡 =𝑔(𝑥𝑡𝑊𝑥+𝑠𝑡−1𝑊𝑠+𝑏)
𝑠𝑡=ℎ𝑡 (2.10)
F om equa ion 2.10, he RNN wo ks by ecei ing an inpu 𝑥𝑡 oge he wi h a s a e ec o
𝑠𝑡−1, a each ime s ep, ha is linea ly ans o med by he weigh ma ices 𝑊𝑥 and 𝑊𝑠. The linea
ans o ma ions along wi h he bias e m 𝑏, simila ly o he MLP, a e passed h ough an ac i a ion
unc ion (e.g. Tanh) ha p oduces a hidden s a e ℎ𝑡. Then, he compu ed hidden s a e o he las
ime s ep is used as a s a e ec o 𝑠𝑡 o he nex inpu 𝑥𝑡+1. Figu e 2.1 depic s an un olded RNN.
Figu e 2.1 – Un olded Recu en Neu al Ne wo k. Taken om h ps:// inyu l.com/yy79zmxo.
In heo y, RNNs should be able o p ocess and ep esen in o ma ion o an en i e sen ence
wi h long- e m dependencies be ween i s cons i uen s, bu in p ac ice, his is no ue since
backp opaga ion h ough ime may no wo k. In o he wo ds, when an RNN is lea ning o s o e
in o ma ion h ough ime, he alues o he g adien s become so small ha he model s ops lea ning
o akes oo much ime o lea n, a p oblem known as anishing g adien s.
A e ined o m o he RNN, he Long Sho -Te m Memo y Ne wo ks (LSTM) (Hoch ei e &
Schmidhube , 1997), was p oposed o add ess he p oblem o he anishing g adien s. The ne wo k
a chi ec u e is designed o lea n long- e m dependencies. The LSTM holds a cell s a e ha ca ies
in o ma ion ac oss he di e en ime s eps o he sequence, ecei ing minimal upda es based on
h ee di e en ga es, o ge , inpu , and ou pu , ha con ol he in o ma ion held by he cell s a e.
Equa ion 2.11 shows he unc ions o he LSTM.
19
e icien wi h ega ds o hei pa ame e size, bu also ega ding hei abili y o ex ac complex
ea u es om he wo ds based on hei dis ibu ion (dis ibu ional hypo hesis). Howe e , hese
ep esen a ions all sho conside ing ha wo ds can ha e di e en meanings depending on he
con ex (wo d polysemy). Because hese ep esen a ions a e ixed, we a e igno ing he di e en
meanings a wo d can ha e depending on he con ex i is used, which can impac he pe o mance o
he downs eam sys em ecei ing he embeddings.
(Pe e s e al., 2018) in oduced a new ype o deep con ex ualized wo d ep esen a ion
de i ed om p e- ained bidi ec ional models - Embeddings om Language Models (ELMo). ELMo
embeddings con e each oken ( ex ual ins ance) in o a cha ac e ep esen a ion using cha ac e
embeddings. The use o cha ac e le el embeddings esembles he use o subwo ds, as cha ac e
ep esen a ions also allow, o ex ac hidden mo phological ea u es and elimina e he OOV wo ds.
Cha ac e embeddings a e hen pushed h ough a CNN laye , ha allows o pick up n-g am ea u es
using he CNN il e . The CNN o e he cha ac e s compu es wo d-le el embeddings, which a e
passed h ough a BiLSTM model. The in e nal ep esen a ions o hese models, i.e. he s a es o he
LSTM, cap u e con ex -dependen aspec s o wo d meaning, which a e op imized wi h a bidi ec ional
language model, making ELMo embeddings a unc ion o he en i e inpu sequence, ins ead o a ixed
ec o o each wo d.
BERT (De lin e al., 2019) is ano he me hod able o p oduce con ex ual embeddings, ha
elies on wo p e iously de eloped ideas – A en ion Mechanism (Bahdanau e al., 2015) and he
T ans o me (Vaswani e al., 2017). The model can be unde s ood as a ained T ans o me encode
s ack, wi h 12 encode laye s o he BERT-Base a ia ion o 24 encode laye s o BERT-La ge, each
wi h eed o wa d neu al ne wo ks wi h 768 and 1024 hidden-nodes, espec i ely (Figu e 2.4).
Figu e 2.4 – BERT-Base and BERT-La ge Encode S ack. Taken om h ps:// inyu l.com/y3we bz7
BERT ecei es subwo d ep esen a ions o inpu sequences, which a e accompanied by a
special oken [CLS], ha appea s a he beginning o he inpu sequence and ano he special oken
[SEP] ha ma ks he end o he sequence. Addi ionally, a segmen embedding, and a posi ional
embedding a e added o each subwo d o indica e i s posi ion on a documen . The inpu embeddings
a e he sum o he oken, segmen , and posi ion embeddings. Each ep esen a ion is hen passed
h ough he encode s ack.

20
BERT’s aining can be di ided in o wo asks: 1) Masked Language Modeling (MLM), in which
15% o he okens in he inpu sequence a e masked, okens a e andomly eplaced by o he okens
and he model is asked o p edic a masked oken. This lea ning ask makes BERT a bi-di ec ional
language model; 2) Nex Sen ence P edic ion, which consis s o p edic ing i a gi en sen ence B is
likely o ollow ano he sen ence A. Bo h hese asks allow BERT o gene a e a language
ep esen a ion model. The p e- ained BERT can hen be used o gene a e con ex ualized wo d
embeddings o a gi en co pus, which can be ed o a downs eam ask, o can be ine- uned o a
speci ic ask such as ques ion answe ing, language in e ence, o named en i y ecogni ion. Besides
BERT-Base and BERT-La ge, a ia ions o BERT ained wi h he di e en co pus ha e been
eleased
10
.
In hei esea ch o measu e he impac o key hype pa ame e s and aining da a size on
BERT aining, Liu e al. (2019) obse ed ha BERT was unde ained. Among o he changes, he
au ho s emo ed he nex sen ence p edic ion objec i e, dynamically changed he masking pa e n
employed o he aining da a and used a no el da ase o p e- aining. This new con igu a ion o
BERT was de ined as Robus ly op imized BERT app oach (RoBERTa) and showed s a e-o - he-a
esul s when es ed on he same asks used o es BERT and o he T ans o me based models, as
well as pe o mance imp o emen s o downs eam asks ha use RoBERTa o c ea e con ex ualized
wo d embeddings.
Recen wo ks wi h T ans o me -like models ha e shi ed hei a en ion o lea ning
mul ilingual (c oss-lingual) ep esen a ions o ex . Among hem, he Mul ilingual BERT (De lin e al.,
2019), which is ained wi h co po a om mul iple languages. La e , o he c oss-lingual T ans o me -
like models used mo e p e- aining da a o o he aining objec i es. Lample & Conneau (2019)
p oposed C oss-lingual Language Models (XLM), in oducing a new supe ised ask o lea ning c oss-
lingual ep esen a ions, he T ansla ion Language Model, in which wo ds a e masked in a sou ce
sen ence (e.g. in English) and a ge sen ence (e.g. in F ench) and he model can a end o
su ounding English wo ds o o he F ench wo ds om he a ge sen ence ( ansla ion). This
me hod encou ages he model o align he English and F ench ep esen a ions.
Mo eo e , Conneau e al. (2020) mixed RoBERTa (Liu e al., 2019) and XLM (Lample &
Conneau, 2019) app oaches, in oducing XLM-RoBERTa (XLM-R), a C oss-lingual RoBERTa
ans o me p e ained on ex om 100 languages, showing pe o mance gains in a ange o
mul ilingual ans e asks when compa ed o o he mul ilingual ans o me s. In conclusion, hese
c oss-lingual T ans o me s can lea n language models in mul iple languages boos ing he
pe o mance o models on monolingual and c oss-lingual classi ica ion and on unsupe ised and
supe ised machine ansla ion.
10
O iginal BERT eposi o y is a ailable om h ps://gi hub.com/google- esea ch/be
21
22
3. RELATED WORK
In his chap e , we add ess he email zoning ela ed wo k (sec ion 3.1) and ela ed wo k in he a ea
o ex segmen a ion (sec ion 3.2). Ch onologically, we discuss he e olu ion o he email zoning
app oaches, axonomies and he di e en models applied, while also summa izing he exis ing email
zoning co po a (sec ion 3.1.5).
3.1. EMAIL ZONING
Mos wo k ela ed o segmen ing ex in o zones was di ec ed o opic-based segmen a ion o news
ex in an unsupe ised ashion. Hea s (1997) eso ed o newspape a icles o de elop Tex Tiling,
an unsupe ised echnique o sub di ing ex s in o sub opics based on pa e ns o lexical co-
occu ence calcula ed om simila i ies be ween in e al poin s in he ex . Likewise, Bee e man,
Be ge , & La e y (1999) p oposed a s a is ical app oach o au oma ically pa i ion ex in o
segmen s based on a log-linea model ha weigh s di e en bina y ea u es o he da a, assigning o
each posi ion in he ex a p obabili y ha a bounda y be ween segmen s a ises a ha posi ion. The
au ho s elied on a news b oadcas ex such as a Wall S ee Jou nal news a chi e a ailable om he
Linguis ic Da a Conso ium
11
. E en hough hese models elied on news ex , hey opened he pa h
o ex segmen ing asks ha eso o di e en sou ces o co po a, such as email.
Chen, Hu, & Sp oa (1999) we e he pionee s in he opic o ex segmen a ion di ec ed o
email ex , bu hei wo k ocused only on he iden i ica ion o signa u e zones. Signa u e ex blocks
a e email pa s ha can usually be ound a he end o he email, con aining au oma ically inse ed
pe sonal da a. Signa u e blocks a e a majo sou ce o in o ma ion, as hey usually con ain de ails
abou he sende , such as email add ess, web add ess, elephone numbe , name, pos al add ess,
e c., and can be used o asks such as cons uc ion o a clien da abase o message e ie al. The
au ho s looked a linguis pa e ns and geome ical pa e ns inside p e iously iden i ied signa u e
zones. Linguis ic pa e ns a e ela ed o he lexical cons ain s ound in he ex , which a e common
in email and web add esses, bu less isible in o he ex passages such as names and add esses. As
o he geome ical pa e ns, hey indica e he eading sequence o a signa u e block. The
esea che s es ed hei model in 1361 signa u e blocks collec ed om emails om he Depa men
o Compu e Science a Conco dia Uni e si y
12
and om hei own pe sonal emails, claiming ha he
sys em achie es a Recall o 53% and a P ecision o 90% in he ask o signa u e iden i ica ion.
3.1.1. JANGADA
Simila ly o Chen, Hu, & Sp oa (1999), Ca alho & Cohen (2004) de eloped JANGADA
13
, a sys em ha
a emp s o iden i y signa u e blocks and quo ed ex om p e ious emails. Figu e 3.1 illus a es a
labeled email message, wi h eply (quo ed ex ) and signa u e lines iden i ied.
11
h ps://www.ldc.upenn.edu/
12
h ps://www.conco dia.ca/ginacody/compu e -science-so wa e-eng
13
h ps://www.cs.cmu.edu/~ i o /codeAndDa a.h ml
23
Zone
Line
o he
F om: [email p o ec ed]
o he
To: Vi o Ca alho
o he
Subjec : Re: Did you y o compile ja adoc ecen ly?
o he
Da e: 25 Ma 2004 12:05:51 -0500
o he
T y c s upda e –dP, his emo es iles & di ec o ies ha ha e been dele ed om c s.
o he
- W
o he
eply
On Wed, 2004-03-24 a 19:58, Vi o Ca alho w o e:
eply
> I jus checked-ou he baseline m3 code and
eply
> "An dis " is wo king ine, bu "an ja adoc" is no .
eply
> Thanks
eply
> Vi o
o he
signa u e
------------------------------------------------------------------------------------------------------------------------------------
signa u e
William W. Cohen “Would you d i e a mime
signa u e
[email p o ec ed] nu s i you played an
signa u e
h p://www.wcohen.com audio ape a ull
signa u e
Associa e Resea ch P o esso blas ?” ----
signa u e
CALD, Ca negie-Mellon Uni e si y S. W igh
Figu e 3.1 – Example o a JANGADA labeled email message adap ed om Ca alho & Cohen (2004).
The sys em s a s by classi ying i an email con ains signa u es o eply lines. Then, o he
selec ed emails, i classi ies each line using CRF (La e y e al., 2001) and sequence-awa e
pe cep ons (Collins, 2002), based on line sel - ea u es and ea u es om su ounding lines, such as
p esence o a quo e sign “>”, p esence o he name o he message sende , p esence o email
add ess pa e ns, p esence o o he ele an cha ac e s and punc ua ion pa e ns and numbe o
leading abs. JANGADA was ained and es ed wi h English emails om he 20 News-g oups
14
(Lang,
1995), and epo ed accu acy eaches 98.91%. La e , Lampe e al. (2009) es ed he pe o mance o
he JANGADA sys em in he En on email co pus
15
(Klim & Yang, 2004), epo ing ha JANGADA de ec s
less han 10% o eply and o wa d lines. None heless, Repke & K es el (2018) showed ha sligh
modi ica ions applied o JANGADA o i di e en asks lead o an accu acy simila o he one claimed
in he o iginal esea ch.
A e JANGADA, Tang, Li, Cao, & Tang (2005) p oposed an email da a cleansing sys em based
on a Suppo Vec o Machine (SVM) (Co es & Vapnik, 1995), ha aimed a il e ing he non- ex ual
noisy con en om emails based on hand-coded ea u es p esen in he email zones ex ual con en .
The au ho s collec ed email om di e en newsg oups held by Google, Yahoo o Mic oso and
ex ended p e ious email zone schemas, conside ing heade , signa u e, quo a ion, p og am code,
and able zones. Repo ed zone de ec ion 1-sco e eaches 97.76% o heade , 89.88% o signa u e,
95% o quo a ion, 81.26% o p og am code and 91.19% o pa ag aph.
14
h p://people.csail.mi .edu/j ennie/20Newsg oups/
15
h p://www.cs.cmu.edu/~en on/
24
In hei au ho p o iling and au ho iden i ica ion ask, Es i al e al. (2007) eso ed o
ec ui ed esponden s dona ed email messages
16
, pa sing each email and classi ying ex segmen s
in o i e ca ego ies: au ho ex , signa u e, ad e isemen , quo ed ex , and eply lines. The
segmen a ion o emails is a c ucial pa o hei wo k since ea u es o au ho a ibu ion and
p o iling can be ound in he au ho ex zones, which ep esen a ound 82% o he o al numbe o
wo ds in hei email co pus. They compa ed a ange o ML algo i hms oge he wi h ea u e selec ion
o classi y each line in an email, a aining imp o emen s in he end ask o au ho p o iling. The
au ho s also compa ed he pe o mance o hei model wi h JANGADA, epo ing an accu acy o 88%
wi h a h ee-zone classi ica ion (au ho ex , signa u e, and eply lines), agains 64% accu acy o
JANGADA.
3.1.2. ZEBRA
Recognizing he lack o a common syn ax and in o mal s uc u e o emails, Lampe e al. (2009)
o mally de ined he email unc ional pa s as email zones, desc ibing he di e en zones inside email
messages based on g aphic, o hog aphic, and lexical ea u es.
Alongside hei de ini ion, he au ho s e ined and ex ended Es i al e al. (2007) classi ica ion
schema, conside ing h ee zones - sende , quo ed con e sa ion, and boile pla e -, ex ensible o nine
subzones. The sende zones con ain ex w i en by he cu en email sende and can be subdi ided
in o au ho , g ee ing, and signo . The quo ed zones con ain eply zones ha hold con en quo ed
om p e ious messages in he same h ead, and o wa d zones, which consis o o wa d messages
om o he con e sa ions. Finally, he boile pla e zones con ain con en ha is eused wi hou
modi ica ion h oughou di e en messages and can be subdi ided in signa u e, ad e ising,
disclaime , and a achmen . An example labeled message is shown in Figu e 3.2.
Fu he mo e, he au ho s also p oposed ZEBRA
17
, an email zoning sys em based on a SVM.
The sys em was ained wi h 400 andom English email messages
18
om an En on email co pus
(Klim & Yang, 2004) da abase dump
19
.
Zone classi ica ion was made ollowing wo app oaches – zone agmen classi ica ion and
line classi ica ion. Fo he zone agmen classi ica ion, ZEBRA conside s email zones o be composed
o agmen s: consecu i e email lines ha a e di ided by zone bounda ies, such as whi e space-only
lines. As o he line classi ica ion me hod, i simply classi ies lines one-by-one, an app oach like he
one seen in JANGADA.
Bo h agmen s and lines a e e e ed o as ex agmen s. The ea u es used o classi y each
ex agmen a e di ided in o G aphic Fea u es, O hog aphic Fea u es, and Lexical Fea u es.
G aphic ea u es cap u e in o ma ion abou he layou o he email ex conside ing, o example, he
numbe o wo ds in he ex agmen o he a e age line leng h o a ex agmen (equal o line
leng h in Line Classi ica ion). O hog aphical ea u es cap u e he use o dis inc i e cha ac e s and
sequences o cha ac e s, such as punc ua ion, capi al le e s, and numbe . Examples o
O hog aphical ea u es a e he pe cen age o capi alized wo ds in ex agmen and whe he he
16
Co pus a ailable upon con ac wi h he au ho s
17
h p://zeb a. hough le s.o g/zoning.php
18
h p://zeb a. hough le s.o g/da a.php
19
h ps://bailando.be keley.edu/en on_email

25
ex con ains an email add ess. Las ly, Lexical Fea u es aim o encapsula e in o ma ion abou he
wo ds used, eso ing o unig ams o each wo d in he ocabula y, and big ams o cap u e sho -
ange wo d sequence in o ma ion. Lexical Fea u es also conside i he sende ’s name o ecipien ’s
name is p esen on he ex agmen o in p e ious ex agmen s.
3 Zone
9 Zone
Line
sende
g ee ing
Sa a:
sende
sende
au ho
Las week I had sen you a inal e sion o he memo andum ega ding pulp and
sende
au ho
pape ansac ions (wi h a achmen ). Howe e , his mo ning I had an
sende
au ho
compu e -gene a ed e o message ega ding ha ansmission. I am sending
sende
au ho
you
sende
au ho
he iles again jus in case you didn’ ecei e hem o couldn’ open hem
sende
au ho
sende
au ho
I you would like a ha d copy, please e-mail me back and I will send you a
sende
au ho
Se .
sende
signo
Thanks.
boile pla e
signa u e
Sco Eckas
boile pla e
signa u e
212-504-6968
boile pla e
boile pla e
a achmen
(See a ached ile: 0463515.04)(See a ached ile: 046450401)
boile pla e
boile pla e
boile pla e
disclaime
|-------------------------------------------------------------------------------------------|
boile pla e
disclaime
|NOTE: The in o ma ion in his email is con iden ial and may be|
boile pla e
disclaime
|legally p i ileged. (…)
boile pla e
disclaime
|any way om i s use.|
boile pla e
disclaime
|-------------------------------------------------------------------------------------------|
boile pla e
boile pla e
a achmen
- 046351504
boile pla e
a achmen
- 046450401
Figu e 3.2 – Example o a labeled email message wi h bo h h ee- and nine-zone classi ica ion
adap ed om Lampe e al. (2009).
ZEBRA uses he men ioned ea u es as inpu o he SVM classi ie . Repo ed accu acy eaches
highe alues wi h line classi ica ion, achie ing 91.53% a e age accu acy wi h h ee zones and
87.01% o nine zones. ZEBRA’s h ee-zone classi ica ion ou pe o ms Es i al e al. (2007) sys em and
Lampe e al. (2009) JANGADA implemen a ion. Fo a nine-zone classi ica ion, when compa ing he
pe o mance o each zone, au ho (89%), eply (91%) and o wa d (89%) achie e he highes F-
Measu e sco e, while ad e ising (34%), signa u e (60%), and disclaime (60%) ge he wo s esul s.
In hei pos e io wo k owa ds de ec ing emails con aining eques s o ac ion, Lampe e
al. (2010) aimed a classi ying emails based on he p esence o ac ion-i ems. Using ZEBRA as an email
p ep ocessing ask, hey we e able o segmen emails in o di e en unc ional zones. Then,
26
conside ing only a small numbe o zones ha had ele an pa e ns o he classi ie , hey inc eased
he accu acy o hei eques de ec ion ask om 72% o 84% accu acy.
A e he in oduc ion o ZEBRA, o almos a decade, no much in es iga ion was made on
he opic o email zoning. Talon
20
, an online lib a y o quo a ion and signa u e ex ac ion, became a
popula and easy o implemen me hod o email cleaning. The lib a y p o ides unc ions able o
ex ac quo a ions and signa u es wi hou he use o machine lea ning algo i hms, eso ing o
sophis ica ed pa e n ma ching echniques.
Fo mo e complex emails in which he pa e n sys em does no wo k, Talon also o e s a
machine lea ning solu ion eso ing o he sciki -lea n
21
lib a y o build SVM classi ie s based on
email line ea u es. The classi ie was ained wi h emails om he En on co pus and con e sa ions
om pe sonal email o he c ea o s o Talon. Addi ionally, Talon’s machine lea ning algo i hm can be
ained wi h he use ’s da ase .
3.1.3. QUAGGA
As email zoning su passed i s o iginal pu pose o signa u e iden i ica ion and ex cleansing in o
a mo e gene al ask, Repke & K es el (2018) ex ended i s u ili y o h ead econs uc ion. The au ho s
p oposed QUAGGA, a ecu en neu al ne wo k-based model ha aims o eco e con e sa ion
h eads on single messages by segmen ing and classi ying email pa s in o wo zones: heade and
body.
The heade zone consis s o blocks o me ada a au oma ically inse ed by he email p og am,
con aining in o ma ion abou he sende , ecipien , da e, and subjec o a quo ed message. The body
zone conside s he ex ha was w i en by he au ho o he cu en email. A con e sa ion h ead is
de ined as being a sequence o heade s and body blocks. This 2-zone di ision can be u he
ex ended o a 5-zone classi ica ion based on Lampe e al. (2009) ZEBRA’s app oach, by di iding he
body zone in o g ee ing, body, signo , and signa u e (see example in Figu e 3.3).
Inspi ed by he wo k on cha ac e awa e neu al language models (Kim, Je ni e, Son ag, & Rush,
2016), QUAGGA uses a CNN o encode he cha ac e s o each line using he ou pu o his ne wo k as
inpu o a bidi ec ional GRU ne wo k, ollowed by a CRF laye , which gi es he inal sco es o each
line. This p ocess is illus a ed in Figu e 3.4. The au ho s e e ha aining he CNN sepa a ely om
he GRU-CRF gi es he bes esul s.
Like Lampe e al. (2009), Repke & K es el (2018) eso ed o he En on co pus and emails
ga he ed om public mail a chi es o he Apache So wa e Founda ion
22
(ASF). The au ho s made
a ailable he anno a ed da ase o he co po a
23
, as well as he code o implemen QUAGGA
24
. Thei
anno a ed da ase conside s 800 En on messages and 500 ASF messages.
20
h ps://gi hub.com/mailgun/ alon
21
h ps://sciki -lea n.o g/s able/
22
h p://mail-a chi es.apache.o g/mod_mbox/
23
h ps://gi hub.com/HPI-In o ma ion-Sys ems/Quagga/ ee/mas e /Da ase s
24
h ps://gi hub.com/TimRepke/Quagga
27
2 Zone
5 Zone
Line
body
body
Thank you o you help.
body
signa u e
ISC Ho line
heade
03/15/2001 10:32 AM
heade
Sen by: Randi Howa d
heade
To: Je Skilling/Co p/En on@ENRON
heade
cc:
heade
Subjec : Re: My “P” Numbe
body
g ee ing
M . Skilling:
body
body
You P numbe is P00500599. Fo you con enience, you can also go o
body
body
h p://isc.en on.com/ unde Si e Highligh s and ese you passwo d o
body
body
ind you “P” numbe .
body
signo
Thanks,
body
signo
Randi Howa d
body
signa u e
ISC HOTLINE
body
body
heade
F om: Je Skilling 03/15/2001 10:01 AM
heade
To: ISC Ho line/Co p/En on@En on
heade
Subjec : My “P” Numbe
body
body
body
body
Could you please o wa d my “P” numbe . I am unable o ge in o he XMS
body
body
sys em and need his ASAP.
body
signo
Thanks o you help.
Figu e 3.3 – Example o a QUAGGA labeled email message wi h bo h wo- and i e-zone anno a ions,
adap ed om Repke & K es el (2018).
Fo he men ioned da ase s, QUAGGA esul s a e compa ed wi h Repke & K es el (2018)
implemen a ion o JANGADA
25
and ZEBRA
26
. Fo wo-zone segmen a ion, QUAGGA shows an accu acy o
98% on bo h he En on and ASF se , e ealing a i e o h ee poin s dec ease when conside ing a i e-
zone segmen a ion. Repo ed pe o mance alues o QUAGGA, a e a supe io compa ed o ZEBRA
(accu acy o 25% o wo zones and 24% o i e zones) which does no ep oduce he esul s o he
o iginal pape . The implemen a ion o JANGADA ge s close o he o iginally epo ed accu acies (88%
o wo zones and 85% o i e zones), bu i is s ill ou pe o med by QUAGGA.
JANGADA and ZEBRA a e es ed only in English emails and hei au ho s do no a i m he sys ems
a e mul ilingual. On he o he hand, due o i s cha ac e le el app oach, Repke & K es el (2018) claim
ha QUAGGA can segmen zones in emails om o he languages, such as A abic and Cy illic.
Ne e heless, no es s we e p esen ed o suppo ha claim.
25
h ps://gi hub.com/HPI-In o ma ion-Sys ems/Quagga/ ee/mas e /Compe i o s/Jangada
26
h ps://gi hub.com/HPI-In o ma ion-Sys ems/Quagga/ ee/mas e /Compe i o s/Zeb a
28
Figu e 3.4 – QUAGGA model o e iew. Taken om Repke & K es el (2018).
3.1.4. CHIPMUNK
Un il e y ecen ly, email zoning eso ed mos ly o small samples o mailing lis s o newsg oup
co pus and was limi ed o he English language. Be endo , Kha ib, Po has , & S ein (2020) we e he
i s o c awl emails a scale, eso ing o he Gmane email- o-newsg oup ga eway
27
. The au ho s
anno a ed 3,033 Gmane emails
28
om 31 languages, bu wi h 90% o he emails being in English.
Due o Gmane’s con e sa ions ichness in echnical opics, he au ho s de eloped a mo e
ine-g ained classi ica ion schema when compa ed o p e ious ela ed wo k, conside ing he
segmen a ion o blocks o code, log da a and echnical da a. Whils also p ese ing mos o he
common zones in oduced in p e ious wo ks, hey ended up wi h a o al o 15 zones: pa ag aph
(main con en , equi alen o body o au ho ed ex ), salu a ion (equi alen o g ee ing), closing
(equi alen o signo ), quo a ion ( o wa d and eply lines), quo a ion ma ke (au ho and da e o a
quo a ion), inline heade (o he de ails o quo a ion such as he subjec ), pe sonal signa u e
(equi alen o signa u e), mua signa u e (mos ly ad e ising), aw code (blocks o code), pa ch
(sou ce code di s), log da a ( ex e e ing o logging in o ma ion), echnical (a achmen s and PGP
signa u es), abula (con en in able o ma ), isual sepa a o (sequence o dashes), and sec ion
heading ( he i le o an pa ag aph segmen ). Figu e 3.5 shows an example o a Gmane labeled icke .
Thei email zoning model, dubbed CHIPMUNK, consis s o a BiGRU-CNN model as depic ed in
Figu e 3.6. 1.5 million emails we e ex ac ed om he main co pus and used o ain a as Tex
embeddings (Joulin, G a e, Bojanowski, & Mikolo , 2017) o size 100. The BiGRU ecei es he
embeddings o he cu en and 𝑛 p e ious lines, wi h he alue o 𝑛 being up o 12 lines, while
pa allelly, a embedding ma ix o size 2𝑐+1,𝑛,100, whe e 𝑐 is o size 4 and ep esen s he con ex
window o each line, and 𝑛 is he maximum oken coun pe line, is ed o a CNN o il e 4×4 ha
pe o ms 128 con olu ions and i s ollowed by a 3×3 CNN laye pe o ming he same amoun o
con olu ions. The ou pu o he second CNN is ed in o a Max Pooling laye and i is conca ena ed in
a single ec o wi h he ou pu o he BiGRU o he cu en and he p e ious lines. The conca ena ed
ec o ecei es a egula iza ion d opou o 0.25 and i is passed h ough a So max laye ha
gene a es he ou pu s.
27
h ps://news.gmane.io/
28
h ps://gi hub.com/webis-de/acl20-c awling-mailing-lis s/ ee/mas e /anno a ions
35
4. METHODOLOGY
Following he concep s and wo k add essed in sec ions 2 and 3, his sec ion desc ibes he Co po a
used in his esea ch (sec ion 4.1), de ines and explo es he sys ems (models) we de eloped o
pe o m he email zoning ask (sec ion 4.2) and desc ibes and explain he me ics used o e alua e
email zoning models on he di e en co po a and o es anno a o ag eemen (sec ion 4.3).
4.1. CORPORA
In his sec ion, we desc ibe and analyze he email zoning co po a used in he scope o his esea ch.
The collec ed co po a can be di ided based on hei pu pose as ollows:
• CLEVERLY AI co pus – a mul ilingual co pus we anno a ed con aining Cus ome Se ice emails
om 14 companies in 5 languages. Collec ed o es di e en email zoning model app oaches
and build a Cus ome Se ice case s udy.
• Email zoning public co po a – used o benchma k ou esul s in a ailable email zoning
co po a, namely Repke & K es el (2018) and Be endo e al. (2020) co po a.. These co po a
we e al eady used o benchma k o he ela ed wo k models.
• New mul ilingual email zoning co pus – collec ed and anno a ed by us o encou age u he
mul ilingual email zoning expe imen s.
Fo each co pus, we analyze he o e all s a is ics (e.g. numbe o emails, numbe o lines,
e c.) and zone dis ibu ion s a is ics (e.g. numbe o lines pe zone, pe cen age o o al lines pe zone,
e c.).
4.1.1. CLEVERLY AI Co pus
We collec ed emails om 14 di e en accoun s (CLEVERLY clien s) in he English, Po uguese, Spanish,
F ench and I alian languages. Mo eo e , by analyzing he co pus oge he wi h CLEVERLY business
needs, we p oduced an email zoning classi ica ion schema ha bes i s he company’s clien emails.
4.1.1.1. De ini ion o Zones
In iew o he commonali ies be ween zone schemas p esen ed in he ela ed wo k (sec ion
3.1), we analyzed CLEVERLY’s co pus conside ing, as well, he company business equi emen s and he
ex ual cha ac e is ics ound in he emails. This way, 8 zones we e de ined:
• The con ex de ails zone ep esen s in o ma ion ha is gene ally used o con ex ualize he
de ails wi h espec o a ce ain ansac ion being add essed in he email message, ypically
in a abula ashion.
• a achmen con ains au oma ed ex in place o a ached documen s.
• inside he g ee ing zone we can ind he e ms o add ess and ecipien names a he
beginning o a message (e.g., Good a e noon/ Dea M s.).
• he body zone con ains new con en om he cu en email sende .
• he signo zone con ains he message closing he au ho ed pa o he email (e.g., Kind
Rega ds, M .).
• signa u e conside s con en con aining con ac o o he in o ma ion ha is au oma ically
inse ed in a message. In con as o signo , signa u e con en is usually empla ed con en

36
w i en once by he email au ho and au oma ically o semi-au oma ically included in email
messages. Also con ained unde signa u e a e au oma ically gene a ed message lines ha
de ine om which de ice he email was sen o ha ad e ise some b and (e.g. Sen om
iPhone/ Secu ed by A as ).
• disclaime includes legal disclaime s, p i acy s a emen s, o o he au oma ically gene a ed
ex s ad ising he people ha in e ac wi h he email (e.g. Think o he en i onmen be o e
p in ing his email).
• inally, quo ed con e sa ion zones include bo h con en s quo ed in eply o p e ious
messages in he same con e sa ion h ead and o wa ded con en om o he con e sa ions.
Zone
Line
con ex de ails
O de de ails: NUMBER
con ex de ails
Email: [email p o ec ed]
con ex de ails
P oduc Name: PRODUCT
a achmen
[image1.png]
g ee ing
Hi.
body
Made an o de wi h you on WEEKDAY and ecei ed his email yes e day (see pic u e).
body
Jus wan o double check ha he o de was p ocessed?
body
signo
Thanks.
signo
NAME SURNAME.
signa u e
Sen om my iPhone.
signa u e
disclaime
--------------------------------------------------------------------------------------------------------------------
disclaime
Impo an : The con en s o his email and any a achmen s a e con iden ial.
disclaime
They a e in ended o he named ecipien (s) only.
disclaime
I you ha e ecei ed his email by mis ake
disclaime
please no i y he sende immedia ely and do no disclose he con en s o anyone o make a
copy he eo .
quo ed
> On WEEKDAY,
quo ed
MONTH DAYNUMBER,
quo ed
YEAR a HH:MM ACCOUNT <[email p o ec ed]> w o e:
quo ed
> Hi NAME.
(…)
(…)
Figu e 4.1 – Exce p om a CLEVERLY labeled email message. Pe sonal and o he sensible ins ances
we e eplaced by a de aul oken ha indica es hei con en .
Figu e 4.1 shows an example email con aining all 8 zones. The zone anno a ion was done by
4 CLEVERLY anno a o s eso ing o he p odigy
47
anno a ion ool, since his was he ool used by
CLEVERLY o all in e nal anno a ion p ocesses. All anno a o s we e na i e Po uguese speake s,
capable o speaking a leas 3 o he co pus languages. Zones we e de ined a he oken le el (wo d,
punc ua ion ma k, e c.) and hen mapped a he line le el. We conside lines as sequences o
cha ac e s and okens delimi ed by punc ua ion ma ks (excep o email add esses, websi e, e c.) o
by a ab “ n”. Single blank and consecu i e blank lines we e labeled wi h he same zone as he
p e ious non-blank line. A e a quo ed line, all nex lines a e conside ed as quo ed. I we wan o
47
h ps://p odi.gy/
37
ex ac mo e de ailed zone in o ma ion om quo ed, his can be done by conside ing each quo ed
segmen (sequence o quo ed lines) as a single email and passing i h ough he email zoning model.
4.1.1.2. Language and Zone Analysis
Table 4.1 cha ac e izes each language p esen in he co pus. English, wi h 8,906 icke s, ep esen s
mo e han hal o he o al numbe o icke s. Po uguese is he second language wi h mo e icke s,
wi h 4,869, oughly hal o he English icke s. Spanish (756), F ench (642) and I alian (284) icke s
summed up a e s ill less han Po uguese icke s. These numbe s e eal he p esence o an
imbalance ega ding he numbe o icke s a ailable o each language co pus.
English
Po uguese
Spanish
F ench
I alian
To al
# icke s
8906
4869
756
642
284
15547
# lines
137566
109577
15577
11231
10649
284600
# lines / icke
15.5
22.5
20.6
17.5
37.5
18.4
# zones / icke
2.7
3.0
2.8
3.0
3.1
2.8
Table 4.1 – S a is ics o CLEVERLY co pus disc imina ed by language.
As conce ning he a e age numbe o lines pe icke , I alian icke s (37.5 lines) a e, by a ,
he longe ones, while English icke s a e he sho es (15.5 lines). The a e age numbe o zones pe
icke is highe in I alian, Po uguese, and F ench - all wi h app oxima ely 3 zones pe icke on
a e age - and smalle in English icke s (2.7 zones).
Figu e 4.2 – Pe cen age o o al lines o email zone, o each language in CLEVERLY co pus.
Figu e 4.2 depic s, o each language, he pe cen age o he o al lines pe zone. Clea ly, quo ed is he
zone wi h mo e lines in e e y language, wi h be ween 40% o 80% o he lines; body is he second
mos p ominen zone in e e y language, anging be ween 30% o 40%. All o he zones ha e a simila
dis ibu ion ac oss language, wi h a achmen (less han 1%) being he zone wi h less occu ences.
0%
20%
40%
60%
80%
100%
% o o al lines
English Po uguese Spanish F ench I alian
38
We no e he I alian co pus has a disp opo iona e amoun o quo ed lines, which may explain
why i is he language wi h a highe alue o a e age lines pe icke , since quo ed zones come, on
a e age, in sequences o 60 lines (see Table 4.2).
The deg ee o dis inc ness be ween he de ined zones does no only ely on he ex ual
con en bu also on g aphical ea u es, such as zones posi ion on he icke o he numbe o
consecu i e sen ences ha cons i u e a zone. Fu he mo e, zones end o p ecede and ollow
di e en zones. These phenomena may help us unde s and he ypical cons i u ion o CLEVERLY’s
icke s.
# lines
# okens
ollows
leads
posi ion
a achmen
2.8
7.21
body
body
3º
con ex de ails
2.2
6.7
body
g ee ing
1º
g ee ing
1.2
4.54
con ex de ails
body
1º
body
4.9
10.9
g ee ing
signo
2º
signo
1.9
4.3
body
signa u e
3º
signa u e
10.0
6.1
body
disclaime
3º
disclaime
5.3
8.2
signa u e
quo ed
3º
quo ed
59.6
7.0
signa u e
-
4º
Table 4.2 – S a is ics o each email zone in CLEVERLY co pus.
In Table 4.2 we cha ac e ize each email zone om CLEVERLY co pus. In summa y, mos zones
end o ha e a simila a e age numbe o consecu i e lines, wi h some excep ions, being quo ed he
mos no able one, wi h mo e han 59 consecu i e lines on a e age. Ne e heless, quo ed zone lines
a e age numbe o okens is close o he alues ound in he o he zones, being body he zone whe e
each line has mo e okens on a e age (10.9 okens). body is he leading zone o hal o he zones
and he ollowing zone o wo o hem. Rega ding he zone’s posi ion in he icke s, g ee ing and
con ex de ails end o appea as he i s zone, body as he second, a achmen , signo , signa u e,
and disclaime as hi d and quo ed as he ou h zone.
Language
T ain
Valida ion
Tes
English
5699
1425
1782
Po uguese
3116
779
974
Spanish
483
121
152
F ench
410
103
129
I alian
182
45
57
Table 4.3 – T ain, alida ion and es spli s o each language in CLEVERLY co pus.
Finally, o e e y language, ou ain, alida ion and es spli s esemble he ones om
ela ed wo k, namely Repke & K es el (2018) and Be endo e al. (2020) co po a. This way, we spli
he co po a in o a ain and es acco ding o he ollowing a ios: 80% o ain and 20% es . Then,
we di ided he ain se in o 80% o ain and 20% o alida ion. These di isions can be ound in
Table 4.3.
39
Conside ing he amoun o a ailable icke s, hese spli s assu e ha we ha e enough
ep esen a i i y o ain, alida e and es models in each language. Ne e heless, because we a e in
he p esence o co po a wi h di e en sizes, we need o be cau ious when compa ing he esul s
om models ained in English o Po uguese and models ained in he o he h ee languages, has
hese ha e less aining da a. Likewise, he alida ion and es spli s may also lead o w ong
conclusions ac oss languages, since a smalle co pus may no ha e he same a iabili y as a la ge
one.
4.1.2. Public Co po a
In his sec ion we analyze he public co po a used in he scope o his esea ch: Repke &
K es el (2018) En on and ASF co po a and Be endo e al. (2020) Gmane and En on co po a.
4.1.2.1. QUAGGA - En on & ASF
As discussed in sec ion 3.1.3, Repke & K es el (2018) objec i e ask is h ead econs uc ion,
which can be seen as sub ask o email zoning. Th ead econs uc ion consis s o ex ac ing
consecu i e sequences o email heade s and body blocks. This ask can be done ia a 2-zone
app oach, whe e email heade s a e classi ied wi h he label heade and body blocks a e classi ied as
he label body. The 2-zone app oach can be ex ended o a 5-zone app oach, whe e he body zone
can be di ided in o g ee ing, body, signo , and signa u e (see Figu e 3.3).
En on
ASF
T ain
Tes
Val
T ain
Tes
Val
# emails
500
200
100
222
90
45
# h eads
221
106
53
147
50
29
h ead leng h
2.7
2.4
2.4
3.3
3.7
3.2
# lines
12828
2309
4457
14086
2434
4588
# lines / email
22.56
23.09
22.26
63.45
54.09
50.98
# zones / email
4.57
4.59
5.13
8.29
8.22
6.91
Table 4.4 – Repke & K es el (2018) a ailable En on and ASF co pus in numbe s conside ing 5 zones.
Table 4.4 shows he di ision o he co po a in o ain, es , and alida ion and espec i e
s a is ics. In bo h co po a, he numbe o emails wi h h eads is a ound hal o he o al numbe o
emails. The ASF co pus seems o ha e longe h eads and he email size is oughly h ee imes longe
han he size o emails in he En on co pus. This is e lec ed in he ASF co pus g ea e numbe o
zones pe email. In bo h cases, icke s a e, on a e age, longe and wi h mo e zones han CLEVERLY’s
ones.
Fu he mo e, he dis ibu ion o zone occu ence o bo h co po a shown in Figu e 4.3
e eals a simila pa e n in bo h co po a. Fo a wo-zone classi ica ion, body comp ises mo e han
80% o he lines, a pa e n ha di e s om CLEVERLY co po a, in which mos o he lines we e quo ed.
This phenomenon can be explained by he h ead econs uc ion zone axonomy, which conside s
body lines a e heade s. Fo i e zones, body is again he zone wi h mo e occu ences in bo h co pus,
making up o 76% o he lines in he En on co pus and 84% in ASF; heade is he second zone wi h
40
mo e lines o bo h co pus, while g ee ing is he leas occu ing zone in he En on co pus,
ep esen ing 3% o he lines and signa u e is he leas occu ing zone o he ASF co pus, wi h 2% o
he lines.
Figu e 4.3 – Pe cen age o he o al lines pe zone. Compa ison be ween he En on co pus
(g een/le ) and he ASF co pus (yellow/ igh ).
4.1.2.2. CHIPMUNK - Gmane & En on
Table 4.5 compiles he s a is ics o he co po a o Be endo e al. (2020).
Gmane
En on
T ain
Tes
T ain
Tes
# emails
2733
300
60
236
# zones
15
15
12
13
# languages
31
14
1
1
% english
88%
87%
100%
100%
# lines
115041
11211
907
4527
# lines / email
42.1
37.4
15.0
19.2
# zones / email
6.4
6.6
4.4
5.9
Table 4.5 – Be endo e al. (2020) a ailable Gmane and En on co pus in numbe s.
On a e age, Gmane emails a e longe and ha e mo e zones han emails om he En on
co pus, which has a s uc u e in line wi h he one obse ed in Repke & K es el (2018) En on co pus.
When i comes o he zone dis ibu ion, while he echnical na u e o he Gmane emails leads o a
la ge amoun o lines co esponding o echnical zones (e.g., log da a, pa ch), in he En on co pus
hose a e non-exis ing. Fu he mo e, in he En on co pus mos lines a e pa ag aphs and quo ed
zones ha e less occu ences han in he Gmane co pus. E en hough he co pus is composed o 31
languages, he emails a e mos ly in English, and he es se only con ains a esidual numbe o non-
76%
3% 5% 4%
12%
84%
3% 6% 2% 5%
0%
20%
40%
60%
80%
100%
body g ee ing closing signa u e heade
% o o al lines
En on ASF

41
English emails (38 emails co e ing 14 di e en languages), which is insu icien o a consis en
mul ilingual e alua ion.
F om he Gmane 3,033 anno a ed emails, 2766 a e des ined o aining, lea ing 300 o
es ing. Line dis ibu ion is he ollowing: quo a ion (50%), pa ag aph (17%), pa ch (10%), log da a
(5%), mua signa u e (4%), aw code (3%), isual sepa a o (3%), abula (2%), pe sonal signa u e
(2%), closing (2%), quo a ion ma ke (2%), salu a ion (1%), echnical (0.5%), inline heade (0.3%) and
sec ion heading (0.3%).
The En on co pus con ains a o al o 300 icke s, om which 60 a e des ined o ine- uning
he model and 236 o es ing. The line dis ibu ion o his co pus is he ollowing: pa ag aph (55%),
quo a ion (12%) and inline heade (7%), quo a ion ma ke (7%), closing (6%), pe sonal signa u e (3%),
mua signa u e (2%), salu a ion (2%), isual sepa a o (2%), sec ion heading (1%), abula (1%),
echnical (1%) and log da a (0.5%). The En on co pus does no ha e lines o aw code and pa ch.
4.1.3. New Mul ilingual Email Zoning Co pus
We sea ched he Gmane aw co pus o emails and, ollowing he zone classi ica ion schema
p oposed by Be endo e al. (2020), p oduced a o al o 625 anno a ed emails in Po uguese,
Spanish and F ench. This co pus is a ailable a h ps://gi hub.com/cle e ly-ai/mul ilingual-email-
zoning.
Po uguese
Spanish
F ench
# emails
210
200
215
# zones
15
15
15
# lines
12366
9824
6958
# lines / email
58.9
49.1
32.4
# zones / email
8.6
6.5
5.9
Table 4.6 – S a is ics o ou mul ilingual email zoning co po a. The dis ibu ion was ob ained by
a e aging he alues o bo h anno a o s.
Table 4.6 compiles a b ie desc ip ion o he email s a is ics o each o he languages. While
F ench is he language wi h mo e emails, Po uguese emails end o be longe , esul ing in a g ea e
amoun o lines and a highe a e age numbe o zones pe email. The Spanish and F ench emails
esemble he s uc u e o Be endo e al. (2020) Gmane emails. Fo he h ee languages, ou co pus
e eals a simila pa e n ega ding he numbe o lines o each zone. Simila ly o Be endo e al.
(2020) Gmane co pus, ou co pus also conside s 15 zones ha can be desc ibed in de ail as ollows:
• pa ag aph: main con en ; new con en om he cu en email sende .
• closing: con en ha closes he email (ex: Rega ds, Some Name).
• inline heade s: con ain lines om in o ma ion ha is used o con ex ualize he de ails o a
p e ious message in he h ead, such as add esses, da es, ecipien s, e c.
• log da a: con ain in o ma ion abou con igu a ions o e en s ha ha e occu ed wi hin a
so wa e applica ion like e o messages o loading o iles.
• mua signa u e: include legal disclaime s, p i acy s a emen s, o o he au oma ically
gene a ed ex s ad ising he people ha in e ac wi h he email. Also con ained unde MUA
42
signa u e a e au oma ically gene a ed message lines ha de ine om which de ice he email
was sen , mail use agen and mailing lis de ails o ad e isemen s (e.g. Sen om iPhone).
• pa ch: con ain ex wi h pieces o code used o make changes o a compu e p og am o i s
suppo ing da a in o de o upda e, ix, o imp o e i . This includes ixing secu i y
ulne abili ies and o he bugs (sou ce code di s).
• pe sonal signa u e: con en con aining con ac o o he pe sonal in o ma ion such as name,
job i le, company, phone numbe , e c., ha is au oma ically inse ed a he end o an email
message.
• quo a ion: include con en s quo ed om p e ious messages in he same con e sa ion h ead
and o wa ded con en om o he con e sa ions. Each line ypically s a s wi h a “>”.
• quo a ion ma ke : ini ia e quo a ion zones s a ing he quo a ion au ho and da e o a
quo a ion.
• aw code: con ain ex ega ding code being de eloped in any p og amming languages
(sou ce code).
• salu a ion: con ain he e ms o add ess and ecipien names a he beginning o a message.
• sec ion heading: lines ha a e heade s o o he zones, mainly o pa ag aph zone.
• abula : ex in a semi-s uc u ed abula o ma ix o ma .
• echnical: con ain lines o au oma ic messages gene a ed by Gmane ega ding ini ia ion,
ending o elimina ion o email pa s; inline a achmen s o PGP signa u es.
• isual sepa a o : lines wi h no alphanume ic cha ac e s, used o di ide he ex be ween
mul iple zones.
The dis ibu ion o zones is simila be ween he h ee languages, as de ailed in Table 4.7. The
mos common zones a e quo a ion, pa ag aph and mua signa u e, while he zones echnical, pa ch
and sec ion heading a e he zones wi h less occu ences. Mo eo e , he co pus shows a less echnical
na u e when compa ed o Be endo e al. (2020) Gmane co pus, since lines o log da a, aw code
echnical o pa ch a e less common.
When c ea ing he co pus, we wan ed o ensu e ha he anno a ion p ocess, i.e. he p ocess
o manually de ining he zones p esen in he emails, led o a co ec iden i ica ion o hose email
zones. One common way o inc ease he obus ness o he anno a ions is o ha e mo e han one
pe son anno a ing he co pus (Yan, Rosales, Fung, Sub amanian, & Dy, 2014). This way, he
anno a ion was ca ied ou by wo anno a o s. The i s anno a o was a na i e Po uguese speake
and he second anno a o a na i e Spanish speake , bo h wi h academic backg ound in F ench and
luen in he hi d language. Each email was anno a ed by bo h anno a o s using he ag og
48
anno a ion ool. We chose his op ion ins ead o he p e iously used p odigy, since hese
anno a ions we e de eloped ou side CLEVERLY’s domain and ag og was he mos e icien cloud and
on-P emises ex anno a ion ool we ound ha does no equi e a p emium accoun o ha e access
o he basic anno a ion ea u es.
48
h ps://www. ag og.ne /
43
Po uguese (%)
Spanish (%)
F ench (%)
quo a ion
52.43
59.02
46.20
pa ag aph
16.33
17.36
27.61
mua signa u e
12.04
3.84
9.04
pe sonal signa u e
3.93
4.47
2.00
isual sepa a o
2.94
2.29
2.60
quo a ion ma ke
2.72
1.54
2.10
closing
2.63
2.00
3.73
log da a
1.04
3.79
1.82
aw code
1.28
2.45
2.07
inline heade s
2.96
0.82
1.33
salu a ion
0.96
0.81
1.35
abula
0.32
0.42
0.27
echnical
0.30
1.00
0.38
pa ch
0.02
0.20
0.02
sec ion heading
0.15
0.04
0.03
Table 4.7 – Dis ibu ion, o each language, o he numbe lines pe zone in ou Mul ilingual co pus.
The dis ibu ion was ob ained by a e aging he alues o bo h anno a o s.
4.2. MODELS
In his sec ion, we p opose i e sys ems based on neu al ne wo k a chi ec u es ained in a
supe ised ashion. To ace he mul ilingual cha ac e o email co po a, we es di e en embedding
me hods wi h he objec i e o inc easing model capaci y o gene alize lea nings o unseen languages.
The gene a ed embeddings a e ed in o a BiLSTM sen ence encode . Since, email zoning can be
pos ula ed as a sequen ial ask, in which we p ocess sequences o okens o sequences o email lines
as inpu , and he classi ica ion o a zone depends on he p e ious and pos e io sequences o zones,
we based he a chi ec u e o ou sys ems in a BiLSTM ne wo k. Mo eo e , p e ious wo k in email
zoning, namely QUAGGA (Repke & K es el, 2018) and CHIPMUNK (Be endo e al., 2020), ha e used
RNN based ne wo ks such as GRU o LSTMs, also seen in s a e-o - he-a ex segmen a ion models
(Badja iya e al., 2018; Kosho ek e al., 2018; Li e al., 2018; Lukasik e al., 2020). We also conside
wo di e en app oaches o he sys em’s ou pu laye - one conside ing a so max laye and ano he
using a CRF laye , which is ypically used o sequen ial asks.
4.2.1. Baseline, Wo d Embeddings + BiLSTM (W-BiLSTM)
Ou p oposed baseline model is based on wo d-le el embeddings and a BiLSTM. We i s di ide each
email line 𝑠𝑗 in a se o wo d-le el okens, which a e ed in o a wo d2 ec
49
(Mikolo e al., 2013)
embedding gene a o o ain 𝑤𝑛
(𝑗) wo d ep esen a ions o each oken 𝑥𝑛
(𝑗), ending up wi h a
ea u es ma ix o each line 𝑠𝑗=[𝑤0,𝑤1,...,𝑤𝑛]. Each email 𝑡, composed o line-le el wo d
49
h ps://code.google.com/a chi e/p/wo d2 ec/
44
embeddings is passed h ough a BiLSTM ne wo k ha e u ns he las hidden s a e o he o wa d
and backwa d ne wo ks, esul ing in a ma ix o line ea u e ep esen a ions 𝑡 =[𝑓0,𝑓1,...,𝑓𝑗], which
goes h ough a so max laye ha weigh s each sen ence o he co esponding zone ype. The zone
wi h a highe weigh is used o classi y he sen ence. Figu e 4.4 p esen s a simpli ied o e iew o he
W-BiLSTM a chi ec u e, showing a single sen ence en e ing he BiLSTM module.
Figu e 4.4 – W-BiLSTM model o e iew. The model is di ided in wo pa s: 1) a sen ence-le el
wo d2 ec wo d encode ; and 2) a segmen a ion module ha uses a BiLSTM and a so max ou pu
laye o classi y each sen ence in o an email zone. Al hough he BiLSTM ecei es he sequence o
sen ences in an email, o simplici y, we illus a e he p ocess o a single sen ence.
This model will ha dly be able o gene alize i s lea ning o di e en languages han he ones
used in aining, since he numbe o ou o ocabula y wo ds (OOV) will be la ge. Ne e heless, i
se es as a simple baseline o compa e i s esul s o u he models ha use di e en embedding
echniques.
4.2.1.1. Subwo d Embeddings + BiLSTM (Sw-BiLSTM)
Based on he idea ha a ious wo ds a e ansla able ia subwo d uni s (Senn ich e al., 2016),
subwo ds can allow us o o e come he p oblem o OOV wo ds and o e ie e meaning o wo ds
ha sha e same subwo d uni s. This way, o es he e ec i eness o hese ep esen a ions in a
mul ilingual con ex , we use Sen encePiece
50
o di ide sen ences in o subwo d uni s. Then, simila ly
o wha we do in ou baseline, we use wo d2 ec o p oduce he subwo d embeddings o each
sen ence. Figu e 4.5 illus a es his model.
50
h ps://gi hub.com/google/sen encepiece
51
𝑘 =𝑝0−𝑝𝑒
𝑝𝑒−1 (5.8)
In equa ion 5.8, 𝑝0 ep esen s he anno a o ela i e ag eemen , he same as accu acy. 𝑝𝑒 is
he hypo he ical p obabili y o ag eemen by chance:
𝑝𝑒=1
𝑁2∑𝑛𝑘1𝑛𝑘2
𝑘(5.9)
In equa ion 5.9, 𝑛𝑘𝑖 is he numbe o imes an anno a o 𝑖 p edic ed ca ego y 𝑘 and 𝑁 is he
numbe o obse a ions. A o al ag eemen be ween he anno a o s would esul in 𝑘 =1, while a
comple e disag eemen , wi h he excep ion o he chance ag eemen gi en by 𝑝𝑒, would esul in
𝑘 =0.

52
53
5. RESULTS AND DISCUSSION
This chap e p esen s and discusses he pe o mance o he models de eloped wi hin he scope o
his esea ch. The e ec i eness o ou sys ems is e alua ed by measu ing hei capabili y o
e ec i ely segmen ing CLEVERLY’s emails (sec ion 5.2.1), e alua ing he impac o a chi ec u e
a ia ions such as di e en embedding me hods and ou pu laye s. Mo eo e , we analyze in mo e
de ail he esul s o he bes pe o ming model (sec ion 5.2.2) and unde s and he impac o
implemen ing his model in CLEVERLY’s classi ica ion pipeline (sec ion 5.2.3), by compa ing he esul s
o he downs eam classi ica ion model be o e and a e i s implemen a ion.
In sec ion 5.3, he model wi h he bes esul s on CLEVERLY co pus, OKAPI, is benchma ked in
public co po a and we compa e i s esul s wi h o he email zoning sys ems om he li e a u e
(sec ion 5.3.1). Finally, we es OKAPI in ou new mul ilingual co pus (sec ion 5.3.2).
5.1. EXPERIMENTAL SETUP
We ained and es ed each model on he CLEVERLY co pus, bo h in a monolingual and c oss-lingual
ashion, o achie e he op imal combina ion o model pa ame e s, namely he numbe o hidden
uni s used in he BiLSTM, size o a d opou laye wi h di e en alues be ween he BiLSTM and he
ou pu laye , aining op imize s o upda e he ne wo k weigh s and i s lea ning a e alues, and loss
unc ion used.
Fo he W-BiLSTM, Sw-BiLSTM and WSw-BiLSTM, we eso o he aining co pus o each
expe imen o ain and ex ac wo d2 ec embeddings o each oken. This p ocess is done
sepa a ely om he segmen a ion module and he embedding alues a e kep ixed h oughou he
model’s aining p ocess. We expe imen ed a BiLSTM wi h 16, 32, 64, 128, 256 and 512 hidden uni s
and 1 and 2 laye s. We es ed a d opou wi h alues un il 0.50, including no d opou . We also es ed
he aining op imize s Adam (Kingma & Ba, 2015) and RMSp op, and di e en alues o he
lea ning a es. In all h ee models, he bes pe o mances we e achie ed using a single BiLSTM wi h
128 hidden uni s and 1 laye , wi h a d opou o 0.25 and using he RMSp op op imize wi h a ixed
lea ning a e o 0.001 and he spa se ca ego ial c oss en opy loss unc ion. In he case o he WSw-
BiLSTM model, we use wo BiLSTMs o his kind, one o encode wo ds and ano he o subwo ds.
As o he XLMR-BiLSTM and OKAPI, we use he XLM-RoBERTa p e- ained weigh s, which a e
kep ozen du ing aining and only he segmen a ion module laye s a e upda ed. Simila ly o he
p e ious models, we expe imen ed BiLSTM wi h he same di e en numbe o hidden uni s and 1
and 2 laye s bu , in he end, ha ing a small segmen a ion module, wi h 64 hidden uni s and 1 laye ,
gene ically yielded he bes pe o mances in he alida ion spli s. We used a d opou laye o 0.25
be ween he BiLSTM and he ou pu laye , he RMSp op op imize wi h a ixed lea ning a e o 0.001
he spa se ca ego ial c oss en opy loss unc ion o he XLMR-BiLSTM and nega i e log-likelihood
loss unc ion o OKAPI.
Fo each aining language in CLEVERLY co pus, he ideal epoch numbe o ain he models
a ies. None heless, o ba ches wi h size o 32, he W-BiLSTM model is op imally ained wi h 8 o
12 epochs, he Sw-BiLSTM and he WSw-BiLSTM wi h 12 o 16 epochs. XLMR-BILSTM and OKAPI a e
op imally ained wi h a ba ch size o 16 o 5 o 10 epochs.
54
Fo he public co po a expe imen s, Okapi main ains he desc ibed pa ame e alues.
None heless he numbe o epochs used changes depending on he aining co pus (see sec ion 5.3).
5.2. CLEVERLY RESULTS
The model ha bes i s CLEVERLY’s needs should be able o co ec ly segmen icke s in o he
equi ed zones, doing so in di e en languages ha migh be unseen du ing he model’s aining
p ocess. This way, o e alua e he models’ capabili y in ul illing he e e ed needs we ain and es
he i e models p oposed in sec ion 4 in each o he i e languages p esen in CLEVERLY co pus. A e
unde s anding he mos e ec i e sys em, we do a ho ough analysis o his model’s pe o mance a
he zone le el, discuss i s icke p ocessing speed, and es i s e ec i eness as a p ep ocessing
sys em o CLEVERLY’s classi ie .
5.2.1. CLEVERLY Anno a ed Co pus
We s a by measu ing he impac o he p oposed embedding app oaches in model pe o mance,
since hese a e he i s a ia ions we in oduce in he a chi ec u e o he p oposed models (sec ion
4.2). Table 5.1 compiles hese esul s, showing he e alua ion o same language p edic ion and o
c oss-lingual ze o-sho p edic ion
52
. O e all, all i e models achie e high 1-sco e when es ed in he
language used o aining, bu gene ally show a de e io a ion in he esul s when dealing wi h
unseen languages.
When compa ing he wo d-le el embedding app oach o he W-BiLSTM model agains he
subwo d-le el app oach o he Sw-BiLSTM model, in gene al, he la e seems o achie e simila o
highe 1-sco e, excep when he models a e ained in he English language. These esul s suppo
he hypo hesis ha subwo d uni s a e especially help ul when models mus deal wi h OOV wo ds.
Ne e heless, since English icke s a e o e ep esen ed in CLEVERLY icke s and a my iad o
English exp essions a e used in emails wo ldwide, he OOV p oblem seems o be mi iga ed, as
models ained in English gene ally achie e be e pe o mances in all es scena ios compa ed o
when using o he languages o aining.
Joining bo h app oaches in he WSw-BiLSTM model seems o ha e a posi i e impac in he
cases whe e he W-BiLSTM pe o mances a e supe io o he Sw-BiLSTM ones. In he emaining
cases, he pa e n o imp o emen is no clea , o pe o mances de e io a e.
The esul s show ha he bes embedding solu ion is he XLM-Robe a p e- ained
embeddings (Conneau e al., 2020). The XLMR-BiLSTM achie es supe io 1-sco e o same language
expe imen s, while main aining i s zoning capaci y compe i i e o c oss-lingual ze o-sho email
zoning, p o ing he e ec i eness o p e- ained con ex ual mul ilingual embeddings when compa ed
o he simple neu al wo d embeddings o wo d2 ec (Mikolo e al., 2013), bo h o monolingual and
mul ilingual asks.
52
Task o es ing a model in languages no seen du ing aining.
55
Tes Language
T ain Language
Model
English
Po uguese
Spanish
F ench
I alian
English
W-BiLSTM
0.89
0.56
0.65
0.79
0.82
Sw-BiLSTM
0.88
0.57
0.73
0.76
0.78
WSw-BiLSTM
0.89
0.56
0.70
0.79
0.83
XLMR-BiLSTM
0.92
0.81
0.91
0.93
0.93
Po uguese
W-BiLSTM
0.33
0.76
0.56
0.49
0.32
Sw-BLISTM
0.41
0.71
0.61
0.56
0.49
WSw-BiLSTM
0.42
0.76
0.59
0.57
0.53
XLM-R-BiLSTM
0.82
0.93
0.88
0.92
0.93
Spanish
W-BiLSTM
0.50
0.40
0.85
0.58
0.52
Sw-BLISTM
0.51
0.42
0.82
0.60
0.72
WSw-BiLSTM
0.45
0.41
0.85
0.67
0.59
XLMR-BiLSTM
0.67
0.63
0.90
0.85
0.86
F ench
W-BiLSTM
0.35
0.35
0.56
0.87
0.51
Sw-BLISTM
0.36
0.36
0.68
0.82
0.68
WSw-BiLSTM
0.30
0.31
0.51
0.89
0.45
XLMR-BiLSTM
0.81
0.61
0.86
0.95
0.91
I alian
W-BiLSTM
0.55
0.35
0.56
0.56
0.75
Sw-BiLSTM
0.60
0.33
0.62
0.60
0.76
WSw-BiLSTM
0.60
0.38
0.66
0.63
0.77
XLMR-BiLSTM
0.77
0.54
0.85
0.87
0.94
Table 5.1 – 1-sco e compa ison o embedding me hod impac o CLEVERLY co pus. Models
we e ained and es ed o each language. Bes esul is highligh ed o each ain language.
Following he same e alua ion app oach, we also es he impac o using CRF ou pu laye
ins ead o a s anda d so max ou pu laye . These esul s a e compiled in Table 5.2.
O e all, OKAPI achie es equal o sligh ly supe io 1-sco e han XLMR-BiLSTM. E en hough
ha ing CRF as he ou pu laye does no show signi ican pe o mance imp o emen s when he
aining is done in English and Po uguese ( he wo languages wi h mo e emails), he CRF posi i e
impac becomes mo e e iden o smalle aining co pus. This pa e n is mo e pe cep ible o c oss-
lingual ze o-sho p edic ion, being Spanish and I alian he wo aining languages whe e he 1-sco e
imp o emen o OKAPI is mo e no iceable. Fo he es languages, English is he only whe e he e is a
clea pa e n o one model ou pe o ming he o he , wi h OKAPI achie ing highe 1-sco e.
In conclusion, he e alua ion esul s indica e ha using XLM-Robe a p e- ained embeddings
and ha ing CRF as he ou pu laye , gene ally leads o he bes esul s on CLEVERLY co pus. This way,
in he ollowing sec ions, we eso o OKAPI o he emaining es s ega ding CLEVERLY’s case s udy
and public co po a.
56
Tes Language
T ain Language
Model
English
Po uguese
Spanish
F ench
I alian
English
XLMR-BiLSTM
0.92
0.81
0.91
0.93
0.93
OKAPI (XLMR-BiLSTM-CRF)
0.92
0.86
0.91
0.93
0.90
Po uguese
XLMR-BiLSTM
0.82
0.93
0.88
0.92
0.93
OKAPI (XLMR-BiLSTM-CRF)
0.83
0.92
0.88
0.90
0.85
Spanish
XLMR-BiLSTM
0.67
0.63
0.90
0.85
0.86
OKAPI (XLMR-BiLSTM-CRF)
0.81
0.76
0.91
0.93
0.92
F ench
XLMR-BiLSTM
0.81
0.61
0.86
0.95
0.91
OKAPI (XLMR-BiLSTM-CRF)
0.82
0.61
0.86
0.95
0.91
I alian
XLMR-BiLSTM
0.77
0.54
0.85
0.87
0.94
OKAPI (XLMR-BiLSTM-CRF)
0.80
0.63
0.85
0.89
0.95
Table 5.2 – 1-sco e compa ison o ou pu laye impac o he model wi h highe pe o ming
embedding me hod. The models we e ained and es ed o each language. Bes esul is highligh ed
o each ain- es language combina ion.
5.2.2. Bes Model Analysis
To unde s and in mo e dep h OKAPI’s email zoning e ec i eness, we also de ail he model’s
pe o mance a he zone le el. Table 5.3 p esen s OKAPI’s p ecision, ecall and 1-sco e esul s
disc imina ed by zone. To ain he model, we compile e e y language aining se in o a single
mul ilingual co pus (compiled mul ilingual aining se ). An equi alen p ocess was ollowed o ob ain
he alida ion and es co pus o his expe imen .
OKAPI achie es an o e all p ecision o 93%, o e all ecall o 94% and o e all 1-sco e o 93%.
Hal o he zones ha e p ecision, ecall and 1-sco e alues abo e 90%. The model is be e a
classi ying zones wi h mo e ep esen a ion – quo ed (49% o he lines) and body (26% o he lines)
a e he wo zones wi h highe 1-sco e alues, wi h 98% and 94% espec i ely; while o a achmen
(only 0.3% o he lines) he model achie es a 1-sco e o only 21%.
zone
p ecision
ecall
1-sco e
suppo
all zones
0.93
0.94
0.93
47508
quo ed
0.96
0.99
0.98
20857
body
0.93
0.95
0.94
13920
g ee ing
0.93
0.93
0.93
1850
con ex de ails
0.91
0.92
0.92
1792
disclaime
0.90
0.84
0.87
3206
signo
0.89
0.81
0.84
2520
signa u e
0.87
0.80
0.83
2999
a achmen
0.65
0.13
0.21
364
Table 5.3 – OKAPI zone-le el pe o mance on CLEVERLY’s mul ilingual co po a.

57
The e is also a clea pa e n showing ha OKAPI ypically achie es ecall alues ha a e highe
han p ecision alues o zones wi h mo e ep esen a ion in he co pus, and he opposi e o zones
wi h less ep esen a ion. In o he wo ds, he model ypically p edic s mo e alse posi i es han alse
nega i es o zones wi h mo e occu ences, and less alse posi i es han alse nega i es o zones
wi h less occu ences.
To be e unde s and he p e ious pa e n, Figu e 5.1 p esen s he con usion ma ix o he
zone-le el e alua ion. The mos no able p edic ion e o OKAPI makes is o mis ake a achmen lines
o body lines, wi h 73% o he a achmen lines being classi ied as body. Con ex de ails (5%),
g ee ing (6%) and signo (10%) a e mos commonly con used wi h body, while disclaime (9%) and
signa u e (8%) a e mo e con used wi h quo ed. In bo h scena ios, zones a e being con used wi h he
su ounding zone wi h mo e ep esen a ion (see Table 4.2). When body and quo ed lines a e w ongly
p edic ed, hey end o be con used wi h each o he , wi h body being p edic ed as quo ed 2% o he
imes and quo ed p edic ed as body less han 1% o he imes.
Figu e 5.1 – Con usion ma ix o OKAPI’s email zoning esul s on CLEVERLY’s mul ilingual co pus. On
he le he ue labels and on he bo om he p edic ed labels. The da ke he squa e, he mo e lines
a e p edic ed wi h he column’s label.
Addi ionally, we expe imen wi h he hypo hesis ha only body lines should be e ained o
he downs eam classi ica ion model. Hence, we educe CLEVERLY’s zoning schema o a wo-le el
schema, in which, in addi ion o he body zone, we conside all o he zones as he zone o he . Table
58
5.4 compiles he esul s o his expe imen . OKAPI achie es an o e all accu acy o 96% and 1-sco e o
97%. In his case, o he body zone, ecall is pe haps he mos impo an measu e as i accoun s o
he numbe o alse nega i e p edic ions – lines ha ac ually belong o body bu a e p edic ed as
o he . A la ge numbe o alse nega i e p edic ions esul s in a dec ease in he amoun o use ul
aining ex p o ided o he classi ie and, consequen ly, can lead o a de e io a ion in i s
e ec i eness. Al hough only 5% o he body lines a e il e ed ou ( ecall is 95%), op imally no body
ins ances should be w ongly emo ed.
zone
p ecision
ecall
1-sco e
accu acy
suppo
o e all
0.97
0.96
0.97
0.96
13920
body
0.93
0.95
0.94
-
33588
o he
0.98
0.97
0.98
-
47508
Table 5.4 – OKAPI wo-le el classi ica ion email zoning pe o mance on CLEVERLY co pus.
5.2.3. Impac in CLEVERLY Pipeline
To measu e he impac o OKAPI when implemen ed in CLEVERLY’s icke classi ica ion pipeline, we
analyze he ime aken by OKAPI in he main s eps o i s email zoning p ocess, and compa e he
accu acy o CLEVERLY’s downs eam icke classi ica ion model when using OKAPI agains when using
CLEVERLY’s cu en solu ion. Thus, we used OKAPI ained wi h he compiled mul ilingual aining se
om he CLEVERLY co pus o zone 84,929 o he icke s mos ly in English and Po uguese, bu also in
F ench, Indonesian, La in, Kinya wanda, and o he uniden i ied languages.
In Table 5.5 we measu e he ime ake by OKAPI o p oduce he embeddings o each icke
and icke line, he segmen a ion ime pe icke ( ime aken by he segmen a ion module) and he
o e all zoning ime – sum o ime aken o p oduce he embeddings and he ime aken o segmen a
icke in o zones.
# XLM-R subwo d uni s / line
14.5
# lines / icke
20.3
ime o c ea e embedding / line
0.06s
ime o c ea e embeddings / icke
1.21s
segmen a ion ime / icke
0.01s
o e all zoning ime / icke
1.22s
Table 5.5 – P edic ed co pus s a is ics and s a is ics ega ding ime aken by OKAPI on main email
zoning s eps.
Wi h an a e age sen ence size o 14.5 XLM-R subwo d uni s and an a e age o 20.3 lines pe
icke , OKAPI akes, on a e age, 0.06 seconds o gene a e a XLM-RoBERTa line embedding and 1.21
seconds o gene a e all line embeddings in a icke . The segmen a ion p ocess is much quicke han
he embedding gene a ion, aking on a e age 0.01 seconds pe icke . These esul s indica e ha he
embedding laye , despi e i s con ibu ion o he imp o emen o email zoning pe o mance, can be a
bo leneck ega ding icke p ocessing ime. Compa ing he ime s a is ics be ween ou model and
CLEVERLY’s cu en solu ion would allow o be e comp ehend OKAPI’S imes. Ne e heless, hose
ime s a is ics a e no a ailable.
59
Fu he mo e, we compa e CLEVERLY’s classi ie accu acy using OKAPI agains he classi ie ’s
accu acy using he cu en solu ion. Table 5.6 compiles bo h accu acies. The esul s show ha he
in oduc ion o OKAPI in he classi ica ion pipeline does no ha e an impac on he classi ie ’s
pe o mance. The ac ha bo h solu ions lead o he same classi ica ion accu acy (71%), indica es
ha a e a simple il e ing om he cu en p ep ocessing me hod, he classi ie , p obably due o i s
own e ec i eness, is al eady able o dis inguish ele an con en om noisy con en in a icke .
Zoning me hod
Classi ie accu acy
cu en me hod
0.71
OKAPI
0.71
Table 5.6 – CLEVERLY’s classi ie accu acy using he cu en icke p ep ocessing me hod e sus OKAPI.
5.3. PUBLIC CORPORA RESULTS
In his sec ion, we analyze bo h mul ilingual and monolingual capabili ies o OKAPI, conside ing
di e en public email zoning co po a and zone classi ica ion schemas. To benchma k ou models in
English public co po a we eso o he co po a and zone classi ica ion schemas used by Repke &
K es el (2018) and Be endo e al. (2020), since hese a e he mos ecen wo ks a ailable and, in
bo h cases, implemen a ions o he mos impo an email zoning sys ems in he li e a u e we e
es ed. We also eso o ou new mul ilingual email zoning co pus o benchma k OKAPI’s esul s in a
public mul ilingual co pus o a c oss-lingual ze o-sho p edic ion.
5.3.1. English Co po a
Reso ing o he numbe s epo ed in he email zoning li e a u e, we compa ed OKAPI wi h exis ing
monolingual me hods using di e en English co po a and zoning schemas.
Table 5.7 compa es OKAPI wi h h ee o he zoning sys ems on he co po a anno a ed by
Repke & K es el (2018) wi h 2 and 5 zones. To ensu e compa abili y be ween he esul s, we s ic ly
ollow he ain, alida ion and es di ision done by Repke & K es el (2018) and p esen ed in sec ion
4.1.2.1. We ained OKAPI wi h a ba ch size o 16, o 9 epochs on he En on co pus aining se and 7
epochs on he ASF co pus aining se .
JANGADA (Ca alho & Cohen, 2004) achie es compa able esul s o he ones ound in i s
o iginal implemen a ion, in pa icula o a 2-zone classi ica ion (97% o he ASF co pus). ZEBRA
(Lampe e al., 2009) does no ep oduce he same pe o mance wi h accu acies below 25% o a 5-
zone classi ica ion. F om he h ee li e a u e models, QUAGGA (Repke & K es el, 2018) achie es he
highes accu acies, wi h 98% o bo h co po a wi h a 2-zone classi ica ion, and, o a 5-zone
classi ica ion, 93% and 95% accu acy, o he En on and ASF co pus, espec i ely.
OKAPI shows excellen pe o mances and, in gene al, su passes he esul s o QUAGGA. I
p oduces an almos pe ec 2-zone classi ica ion o bo h co po a, wi h 99% accu acy; o a 5-zone
classi ica ion ou model ou pe o ms QUAGGA wi h 96% accu acy in he En on co pus and achie es
he same 95% accu acy as Repke & K es el (2018) model in he ASF co pus.
60
Model
Zones
En on
ASF
JANGADA
2
0.89/0.88/0.88
0.97/0.97/0.97
ZEBRA
2
0.66/0.25/0.25
0.88/0.18/0.18
QUAGGA
2
0.98/0.98/0.98
0.98/0.98/0.98
OKAPI
2
0.99/0.99/0.99
0.99/0.99/0.99
JANGADA
5
0.82/0.85/0.85
0.92/0.90/0.91
ZEBRA
5
0.60/0.25/0.24
0.81/0.20/0.20
QUAGGA
5
0.93/0.93/0.93
0.95/0.95/0.95
OKAPI
5
0.96/0.96/0.96
0.95/0.95/0.95
Table 5.7 – OKAPI email zoning pe o mance (p ecision/ ecall/accu acy) compa ed o a ious models
om he li e a u e, o Repke & K es el (2018) co po a in a 2-zone and 5-zone schema.
Table 5.8 shows he esul s ob ained by OKAPI in Be endo e al. (2020) co po a and zoning
schema. To ain ou model, we ollow Be endo e al. (2020) o iginal ain and es di ision.
Ne e heless, since he au ho s did no p o ide a se o model alida ion, we used 10% o he ain
English Gmane co pus o alida e ou model. We choose a small spli so ha we main ain, as much as
possible, simila condi ions o he ones ollowed by Be endo e al. (2020). OKAPI was ained wi h a
ba ch size o 16, o 11 epochs on Gmane co pus aining se and 4 epochs on he En on aining se
o ine une he model and es i on he En on es se .
Gmane
En on
OKAPI
CHIPMUNK
QUAGGA
Tang
OKAPI
CHIPMUNK
QUAGGA
Tang
all zones
0.96
0.96
0.94
0.80
0.88
0.88
0.83
0.72
quo a ion
0.98
0.99
0.99
0.99
0.94
0.99
0.88
0.85
pa ch
0.98
0.95
0.95
0.46
-
-
-
-
pa ag aph
0.94
0.93
0.90
0.90
0.94
0.95
0.91
0.89
log da a
0.83
0.84
0.77
-
0.00
0.24
0.74
-
mua signa u e
0.92
0.91
0.93
0.40
0.45
0.65
0.51
0.21
pe sonal signa u e
0.86
0.77
0.85
0.73
0.85
0.78
Table 5.8 – OKAPI email zoning o e all accu acy and 6 mos common zones ecall, compa ed o
a ious models, unde he 15-le el zoning schema and co po a o Be endo e al. (2020).
OKAPI and CHIPMUNK (Be endo e al., 2020) achie e he highes pe o mance o he Gmane
co pus, wi h an all zone accu acy o 96%. Al hough OKAPI, CHIPMUNK and QUAGGA (94%) pe o m
almos equally well, OKAPI seems o achie e he mos consis en pe o mance ac oss di e en zones,
eaching he highes ecall in 3 zones – pa ch (98%), pa ag aph (98%) and pe sonal signa u e (86%).
On he o he hand, CHIPMUNK achie es he highes ecall o quo a ion (99%) and log da a (84%),
QUAGGA also achie es he highes ecall o quo a ion (99%) and mua signa u e (93%) and, while Tang
e al. (2005) p esen he lowes all zones accu acy (0.80%) i also eaches 99% ecall o quo a ion.
Rega ding he En on co pus, OKAPI eaches 88% accu acy o all zones, ma ching, once again,
CHIPMUNK’s pe o mance. QUAGGA eaches an all zones accu acy o 83% and Tang p esen s he lowes
all zones accu acy wi h 72%. E en hough ou model ma ches CHIPMUNK’s o e all accu acy, when
looking a each one o he six mos common zones, OKAPI achie es a lowe ecall. This appa en
67

68
8. BIBLIOGRAPHY
Almeida, M. S. C., Pin o, C., Figuei a, H., Mendes, P., & Ma ins, A. F. T. (2015). Aligning Opinions:
C oss-Lingual Opinion Mining wi h Dependencies. P oceedings o he 53 d Annual Mee ing o
he Associa ion o Compu a ional Linguis ics and he 7 h In e na ional Join Con e ence on
Na u al Language P ocessing, 1, 408–418. h ps://doi.o g/10.3115/ 1/P15-1040
Badja iya, P., Ku isinkel, L. J., Gup a, M., & Va ma, V. (2018). A en ion-Based Neu al Tex
Segmen a ion. Lec u e No es in Compu e Science , 10772 LNCS, 180–193.
h ps://doi.o g/10.1007/978-3-319-76941-7_14
Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neu al Machine T ansla ion by Join ly Lea ning o Align
and T ansla e. 3 d In e na ional Con e ence on Lea ning Rep esen a ions, ICLR 2015.
Baum, L. E., & Pe ie, T. (1966). S a is ical In e ence o P obabilis ic Func ions o Fini e S a e Ma ko
Chains. Ann. Ma h. S a is ., 37(6), 1554–1563. h ps://doi.o g/10.1214/aoms/1177699147
Bee e man, D., Be ge , A., & La e y, J. (1999). S a is ical Models o Tex Segmen a ion. Machine
Lea ning, 34(1), 177–210. h ps://doi.o g/10.1023/A:1007506220214
Be na do, J., Baya i, M., Be ge , J., Dawid, A., Hecke man, D., Smi h, A., … Lasse e, J. (2007).
Gene a i e o Disc imina i e? Ge ing he Bes o Bo h Wo lds. BAYESIAN STATISTICS, 8, 3–24.
Be enbu g, N., Adams, B., Hassan, A. E., & Smid , M. (2011). A Ligh weigh App oach o Unco e
Technical A i ac s in Uns uc u ed Da a. In e na ional Con e ence on P og am Comp ehension,
185–188. h ps://doi.o g/10.1109/ICPC.2011.36
Be endo , J., Kha ib, K. Al, Po has , M., & S ein, B. (2020). C awling and P ep ocessing Mailing Lis s
A Scale o Dialog Analysis. P oceedings o he 58 h Annual Mee ing o he Associa ion o
Compu a ional Linguis ics - ACL 2020, 1151–1158. h ps://doi.o g/10.18653/ 1/2020.acl-
main.108
Ca lson, L., Ma cu, D., & Oku owski, M. E. (2003). Building a Discou se-Tagged Co pus in he
F amewo k o Rhe o ical S uc u e Theo y. Cu en and New Di ec ions in Discou se and
Dialogue, 85–112. h ps://doi.o g/10.1007/978-94-010-0019-2_5
Ca alho, V. R., & Cohen, W. W. (2004). Lea ning o Ex ac Signa u e and Reply Lines om Email.
Fi s Con e ence on Email and An i-Spam (CEAS). Re ie ed om
h ps://www.cs.cmu.edu/~ i o /pape s/sigFilePape _ inal e sion.pd
Cauchy, A.-L. (1847). Me hode gene ale pou la esolu ion des sys emes d’equa ions simul anees.
25(2), 536–538. Re ie ed om h ps://www.mendeley.com/ca alogue/7a6d25c8-4c6d-33e3-
8855-
6963 320b5d/?u m_sou ce=desk op&u m_medium=1.19.4&u m_campaign=open_ca alog&us
e Documen Id=%7B575ee8d -eb07-4ee1-bd71-0a99 97c720d%7D
Chen, H., Hu, J., & Sp oa , R. W. (1999). In eg a ing geome ical and linguis ic analysis o email
signa u e block pa sing. ACM T ansac ions on In o ma ion Sys ems, 17(4), 343–366.
h ps://doi.o g/10.1145/326440.326442
Chen, M. X., Cao, Y., Tsay, J., Chen, Z., Lee, B. N., Zhang, S., … Wu, Y. (2019). Gmail sma compose:
Real- ime assis ed w i ing. P oceedings o he ACM SIGKDD In e na ional Con e ence on
Knowledge Disco e y and Da a Mining, 2287–2295. h ps://doi.o g/10.1145/3292500.3330723
69
Cho, K., an Me iënboe , B., Gulceh e, C., Bahdanau, D., Bouga es, F., Schwenk, H., & Bengio, Y.
(2014). Lea ning ph ase ep esen a ions using RNN encode -decode o s a is ical machine
ansla ion. P oceedings o he 2014 Con e ence on Empi ical Me hods in Na u al Language
P ocessing ({EMNLP}), 1724–1734. h ps://doi.o g/10.3115/ 1/D14-1179
Choi, F. Y. Y. (2000). Ad ances in domain independen linea ex segmen a ion. 6 h Applied Na u al
Language P ocessing Con e ence - ANLP 2000, 26–33. Re ie ed om
h ps://www.aclweb.o g/an hology/A00-2004/
Ch is idis, P., & Losada, Á. G. (2019). Email Based Ins i u ional Ne wo k Analysis: Applica ions and
Risks. The Social Sciences, 8(11), 306. h ps://doi.o g/10.3390/socsci8110306
Cohen, J. (1960). A Coe icien o Ag eemen o Nominal Scales. Educa ional and Psychological
Measu emen , 20(1), 37–46. h ps://doi.o g/10.1177/001316446002000104
Collins, M. (2002). Disc imina i e T aining Me hods o Hidden Ma ko Models: Theo y and
Expe imen s wi h Pe cep on Algo i hms. P oceedings o he 2002 Con e ence on Empi ical
Me hods in Na u al Language P ocessing - EMNLP 2002, 1–8.
h ps://doi.o g/10.3115/1118693.1118694
Collobe , R., & Wes on, J. (2008). A Uni ied A chi ec u e o Na u al Language P ocessing: Deep
Neu al Ne wo ks wi h Mul i ask Lea ning. P oceedings o he 25 h In e na ional Con e ence on
Machine Lea ning - ICML ’08, 160–167. h ps://doi.o g/10.1145/1390156.1390177
Conneau, A., Khandelwal, K., Goyal, N., Chaudha y, V., Wenzek, G., Guzmán, F., … S oyano , V.
(2020). Unsupe ised C oss-lingual Rep esen a ion Lea ning a Scale. P oceedings o he 58 h
Annual Mee ing o he Associa ion o Compu a ional Linguis ics, 8440–8451.
h ps://doi.o g/10.18653/ 1/2020.acl-main.747
Co es, C., & Vapnik, V. (1995). Suppo -Vec o Ne wo ks. Machine Lea ning, 20, 273–297.
h ps://doi.o g/10.1007/BF00994018
Coussemen , K., & den Poel, D. Van. (2008). Imp o ing cus ome complain managemen by
au oma ic email classi ica ion using linguis ic s yle ea u es as p edic o s. Decision Suppo
Sys ems, 44(4), 870–882. h ps://doi.o g/h ps://doi.o g/10.1016/j.dss.2007.10.010
Cu y, H. B. (1944). The me hod o s eepes descen o non-linea minimiza ion p oblems. Qua e ly
o Applied Ma hema ics, 2(3), 258–261. Re ie ed om h p://www.js o .o g/s able/43633461
De lin, J., Chang, M.-W., Lee, K., & Tou ano a, K. (2019). BERT: P e- aining o Deep Bidi ec ional
T ans o me s o Language Unde s anding. P oceedings o he 2019 Con e ence o he No h
Ame ican Chap e o he Associa ion o Compu a ional Linguis ics, 1, 4171–4186.
h ps://doi.o g/10.18653/ 1/n19-1423
Elman, J. L. (1990). Finding s uc u e in ime. Cogni i e Science, 14(2), 179–211.
h ps://doi.o g/h ps://doi.o g/10.1016/0364-0213(90)90002-E
Es i al, D., Gaus ad, T., Pham, S. B., Rad o d, W., & Hu chinson, B. (2007). Au ho p o iling o English
emails. P oceedings o he 10 h Con e ence o he Paci ic Associa ion o Compu a ional
Linguis ics, 263–272.
Fix, E., & Hodges, J. L. (1989). Disc imina o y Analysis. Nonpa ame ic Disc imina ion: Consis ency
P ope ies. In e na ional S a is ical Re iew / Re ue In e na ionale de S a is ique, 57(3), 238–
247. Re ie ed om h p://www.js o .o g/s able/1403797
70
Fukushima, K. (1980). Neocogni on: A sel -o ganizing neu al ne wo k model o a mechanism o
pa e n ecogni ion una ec ed by shi in posi ion. Biological Cybe ne ics, 36(4), 193–202.
h ps://doi.o g/10.1007/BF00344251
Gage, P. (1994). A New Algo i hm o Da a Comp ession. C Use s J., 12(2), 23–38.
Gellens, R. (2004). The Tex /Plain Fo ma and DelSp Pa ame e s. h ps://doi.o g/10.17487/RFC3676
Gla aš, G., & Somasunda an, S. (2020). Two-le el ans o me and auxilia y cohe ence modeling o
imp o ed ex segmen a ion. The Thi yFou h AAAI Con e ence on A i icial In elligence (AAAI-
20), 2306–2315. h ps://doi.o g/10.1609/aaai. 34i05.6284
G a es, A., & Schmidhube , J. (2005). F amewise phoneme classi ica ion wi h bidi ec ional LSTM and
o he neu al ne wo k a chi ec u es. Neu al Ne wo ks, 18(5), 602–610.
h ps://doi.o g/10.1016/j.neune .2005.06.042
Hea s , M. A. (1997). Tex Tiling: Segmen ing Tex in o Mul i-pa ag aph Sub opic Passages.
Compu a ional Linguis ics, 23(1), 33–64.
Hoch ei e , S., & Schmidhube , J. (1997). Long Sho -Te m Memo y. Neu al Compu a ion, 9(8), 1735–
1780. h ps://doi.o g/10.1162/neco.1997.9.8.1735
Huang, Z., Xu, W., & Yu, K. (2015). Bidi ec ional LSTM-CRF Models o Sequence Tagging. a Xi p ep.
Re ie ed om h p://a xi .o g/abs/1508.01991
Ja dim, B., Rei, R., & Almeida, M. S. C. (2021). Mul ilingual Email Zoning. EACL 2021 S uden Resea ch
W okshop. Re ie ed om h ps://a xi .o g/abs/2102.00461
Joulin, A., G a e, E., Bojanowski, P., & Mikolo , T. (2017). Bag o T icks o E icien Tex Classi ica ion.
P oceedings o he 15 h Con e ence o he Eu opean Chap e o he Associa ion o
Compu a ional Linguis ics., 2, 427–431. Re ie ed om
h ps://www.aclweb.o g/an hology/E17-2068
Kie e , J., & Wol owi z, J. (1952). S ochas ic Es ima ion o he Maximum o a Reg ession Func ion. The
Annals o Ma hema ical S a is ics, 23(3), 462–466. Re ie ed om
h p://www.js o .o g/s able/2236690
Kim, Y., Je ni e, Y., Son ag, D. A., & Rush, A. M. (2016). Cha ac e -Awa e Neu al Language Models.
P oceedings o he Thi ie h AAAI Con e ence on A i icial In elligence, 2741–2749. Re ie ed
om h p://www.aaai.o g/ocs/index.php/AAAI/AAAI16/pape / iew/12489
Kingma, D. P., & Ba, J. (2015). Adam: A Me hod o S ochas ic Op imiza ion. In Y. Bengio & Y. LeCun
(Eds.), 3 d In e na ional Con e ence on Lea ning Rep esen a ions. Re ie ed om
h p://a xi .o g/abs/1412.6980
Klim , B., & Yang, Y. (2004). The En on Co pus: A New Da ase o Email Classi ica ion Resea ch.
Lec u e No es in A i icial In elligence, 217–226. h ps://doi.o g/10.1007/978-3-540-30115-8_22
Kocayusu oglu, F., Sheng, Y., Vo, N., Wend , J., Zhao, Q., Ta a, S., & Najo k, M. (2019). RiSER: Lea ning
Be e Rep esen a ions o Richly S uc u ed Emails. The Wo ld Wide Web Con e ence, 886–895.
h ps://doi.o g/10.1145/3308558.3313720
Koehle , J., Fux, E., He zog, F. A., Lö sche , D., Wael i, K., Imobe do , R., & Budke, D. (2018). Towa ds
in elligen p ocess suppo o cus ome se ice desks: Ex ac ing p oblem desc ip ions om
noisy and mul i-lingual ex s. Lec u e No es in Business In o ma ion P ocessing, 308, 36–52.
71
h ps://doi.o g/10.1007/978-3-319-74030-0_3
Kosho ek, O., Cohen, A., Mo , N., Ro man, M., & Be an , J. (2018). Tex Segmen a ion as a Supe ised
Lea ning Task. NAACL HLT 2018 - 2018 Con e ence o he No h Ame ican Chap e o he
Associa ion o Compu a ional Linguis ics: Human Language Technologies - P oceedings o he
Con e ence, 2, 469–473. h ps://doi.o g/10.18653/ 1/n18-2075
Kudo, T. (2018). Subwo d Regula iza ion: Imp o ing Neu al Ne wo k T ansla ion Models wi h
Mul iple Subwo d Candida es. P oceedings o he 56 h Annual Mee ing o he Associa ion o
Compu a ional Linguis ics (Volume 1: Long Pape s), 66–75. h ps://doi.o g/10.18653/ 1/P18-
1007
La e y, J. D., McCallum, A., & Pe ei a, F. C. N. (2001). Condi ional Random Fields: P obabilis ic
Models o Segmen ing and Labeling Sequence Da a. P oceedings o he Eigh een h
In e na ional Con e ence on Machine Lea ning, 282–289. San F ancisco, CA, USA: Mo gan
Kau mann Publishe s Inc.
Lampe , A., Dale, R., & Pa is, C. (2009). Segmen ing email message ex in o zones. P oceedings o
he 2009 Con e ence on Empi ical Me hods in Na u al Language P ocessing, 919–928.
Lampe , A., Dale, R., & Pa is, C. (2010). De ec ing Emails Con aining Reques s o Ac ion. Human
Language Technologies: The 2010 Annual Con e ence o he No h Ame ican Chap e o he
Associa ion o Compu a ional Linguis ics, 984–992. Re ie ed om
h ps://www.aclweb.o g/an hology/N10-1142
Lample, G., & Conneau, A. (2019). C oss-lingual Language Model P e aining. A Xi . Re ie ed om
h p://a xi .o g/abs/1901.07291
Landis, J. R., & Koch, G. G. (1977). The Measu emen o Obse e Ag eemen o Ca ego ical Da a.
Biome ics, 33(1), 159–174. Re ie ed om h p://www.js o .o g/s able/2529310
Lang, K. (1995). NewsWeede : Lea ning o Fil e Ne news. P oceedings o he Twel h In e na ional
Con e ence on In e na ional Con e ence on Machine Lea ning, 331–339.
h ps://doi.o g/h ps://doi.o g/10.1016/B978-1-55860-377-6.50048-7
LeCun, Y., Bose , B., Denke , J. S., Hende son, D., Howa d, R. E., Hubba d, W., & Jackel, L. D. (1989).
Backp opaga ion Applied o Handw i en Zip Code Recogni ion. Neu al Compu a ion, 1(4), 541–
551. h ps://doi.o g/10.1162/neco.1989.1.4.541
Li, J., Sun, A., & Jo y, S. R. (2018). SegBo : A Gene ic Neu al Tex Segmen a ion Model wi h Poin e
Ne wo k. P oceedings o he Twen y-Se en h In e na ional Join Con e ence on A i icial
In elligence, 4166–4172. h ps://doi.o g/10.24963/ijcai.2018/579
Linnainmaa, S. (1976). Taylo expansion o he accumula ed ounding e o . BIT Nume ical
Ma hema ics, 16(2), 146–160. h ps://doi.o g/10.1007/BF01931367
Liu, Y., O , M., Goyal, N., Du, J., Joshi, M., Chen, D., … S oyano , V. (2019). RoBERTa: A Robus ly
Op imized BERT P e aining App oach. CoRR, abs/1907.1. Re ie ed om
h p://a xi .o g/abs/1907.11692
Lukasik, M., Dadache , B., Papineni, K., & Simões, G. (2020). Tex Segmen a ion by C oss Segmen
A en ion. P oceedings o he 2020 Con e ence on Empi ical Me hods in Na u al Language
P ocessing (EMNLP), 4707–4716. h ps://doi.o g/10.18653/ 1/2020.emnlp-main.380
Luong, T., Pham, H., & Manning, C. D. (2015). E ec i e App oaches o A en ion-based Neu al
72
Machine T ansla ion. P oceedings o he 2015 Con e ence on Empi ical Me hods in Na u al
Language P ocessing, 1412–1421. h ps://doi.o g/10.18653/ 1/D15-1166
Ma ko , A. A. (1953). The Theo y o Algo i hms. Jou nal o Symbolic Logic, 18(4), 340–341.
h ps://doi.o g/10.2307/2266585
McCulloch, W. S., & Pi s, W. (1943). A logical calculus o he ideas immanen in ne ous ac i i y. The
Bulle in o Ma hema ical Biophysics, 5(4), 115–133. h ps://doi.o g/10.1007/BF02478259
Mikolo , Tomas, Chen, K., Co ado, G., & Dean, J. (2013). E icien Es ima ion o Wo d
Rep esen a ions in Vec o Space. In Y. Bengio & Y. LeCun (Eds.), 1s In e na ional Con e ence on
Lea ning Rep esen a ions. Re ie ed om h p://a xi .o g/abs/1301.3781
Mikolo , Tomas, Su ske e , I., Chen, K., Co ado, G. S., & Dean, J. (2013). Dis ibu ed Rep esen a ions
o Wo ds and Ph ases and hei Composi ionali y. P oceedings o he 26 h In e na ional
Con e ence on Neu al In o ma ion P ocessing Sys ems, 2, 3111–3119.
Mikolo , Tomáš, Su ske e , I., Deo as, A., Le, H.-S., Komb ink, S., & Ce nocky, J. (2012). Subwo d
language modeling wi h neu al ne wo ks. In Unpublished.
Mos elle , F., & Tukey, J. W. (1968). Da a analysis, including s a is ics. Handbook o Social Psychology,
2, 80–203.
Nießen, S., & Ney, H. (2000). Imp o ing SMT Quali y wi h Mo pho-Syn ac ic Analysis. P oceedings o
he 18 h Con e ence on Compu a ional Linguis ics - Volume 2, 1081–1085.
h ps://doi.o g/10.3115/992730.992809
Penning on, J., Soche , R., & Manning, C. D. (2014). GloVe: Global Vec o s o Wo d Rep esen a ion.
P oceedings o he 2014 Con e ence on Empi ical Me hods in Na u al Language P ocessing,
1532–1543. h ps://doi.o g/10.3115/ 1/d14-1162
Pe e s, M. E., Neumann, M., Iyye , M., Ga dne , M., Cla k, C., Lee, K., & Ze lemoye , L. (2018). Deep
Con ex ualized Wo d Rep esen a ions. P oceedings o he 2018 Con e ence o he No h
Ame ican Chap e o he Associa ion o Compu a ional Linguis ics: Human Language
Technologies, 2227–2237. h ps://doi.o g/10.18653/ 1/n18-1202
P osku nia, J., Ca igh , M.-A., Ga cia-Pueyo, L., K ka, I., Wend , J. B., Kau mann, T., & Miklos, B.
(2017). Templa e Induc ion o e Uns uc u ed Email Co po a. P oceedings o he 26 h
In e na ional Con e ence on Wo ld Wide Web, 1521–1530.
h ps://doi.o g/10.1145/3038912.3052631
Qa oush, A., Kha e , I. M., & Washaha, M. (2012). Iden i ying spam e-mail based-on s a is ical heade
ea u es and sende beha io . ACM In e na ional Con e ence P oceeding Se ies, 771–778.
h ps://doi.o g/10.1145/2381716.2381863
Reime s, N., & Gu e ych, I. (2020). Sen ence-BERT: Sen ence embeddings using siamese BERT-
ne wo ks. EMNLP-IJCNLP 2019 - 2019 Con e ence on Empi ical Me hods in Na u al Language
P ocessing and 9 h In e na ional Join Con e ence on Na u al Language P ocessing, P oceedings
o he Con e ence, 3982–3992. h ps://doi.o g/10.18653/ 1/d19-1410
Repke, T., & K es el, R. (2018). B inging back s uc u e o ee ex email con e sa ions wi h ecu en
neu al ne wo ks. Ad ances in In o ma ion Re ie al, 114–126. h ps://doi.o g/10.1007/978-3-
319-76941-7_9
Robbins, H., & Mon o, S. (1951). A S ochas ic App oxima ion Me hod. The Annals o Ma hema ical

73
S a is ics, 22(3), 400–407. h ps://doi.o g/10.1214/aoms/1177729586
Rosenbla , F. (1958). The pe cep on: a p obabilis ic model o in o ma ion s o age and o ganiza ion
in he b ain. Psychological Re iew, 65 6, 386–408.
Rumelha , D. E., Hin on, G. E., & Williams, R. J. (1986). Lea ning ep esen a ions by back-p opaga ing
e o s. Na u e, 323, 533–536.
Senn ich, R., Haddow, B., & Bi ch, A. (2016). Neu al Machine T ansla ion o Ra e Wo ds wi h
Subwo d Uni s. P oceedings o he 54 h Annual Mee ing o he Associa ion o Compu a ional
Linguis ics, Volume 1: Long Pape s, 1715–1725. h ps://doi.o g/10.18653/ 1/p16-1162
Sneide s, E. (2016). Re iew o he main app oaches o au oma ed email answe ing. Ad ances in
In elligen Sys ems and Compu ing, 444, 135–144. h ps://doi.o g/10.1007/978-3-319-31232-
3_13
Søgaa d, A., Rude , S., & Vulić, I. (2018). On he limi a ions o unsupe ised bilingual dic iona y
induc ion. ACL 2018 - 56 h Annual Mee ing o he Associa ion o Compu a ional Linguis ics,
P oceedings o he Con e ence (Long Pape s), 1, 778–788. h ps://doi.o g/10.18653/ 1/p18-
1072
S i as a a, N., Hin on, G., K izhe sky, A., & Salakhu dino , R. (2014). D opou : A Simple Way o
P e en Neu al Ne wo ks om O e i ing. Jou nal o Machine Lea ning Resea ch, 15(1), 1929–
1958.
Su ske e , I., Vinyals, O., & Le, Q. V. (2014). Sequence o Sequence Lea ning wi h Neu al Ne wo ks.
P oceedings o he 27 h In e na ional Con e ence on Neu al In o ma ion P ocessing Sys ems -
Volume 2, 3104–3112. Camb idge, MA, USA: MIT P ess.
Tang, J., Li, H., Cao, Y., & Tang, Z. (2005). Email da a cleaning. P oceedings o he Ele en h ACM
SIGKDD In e na ional Con e ence on Knowledge Disco e y and Da a Mining, 489–498.
h ps://doi.o g/10.1145/1081870.1081926
Tenney, I., Das, D., & Pa lick, E. (2019). BERT edisco e s he classical NLP pipeline. ACL 2019 - 57 h
Annual Mee ing o he Associa ion o Compu a ional Linguis ics, P oceedings o he Con e ence,
4593–4601. h ps://doi.o g/10.18653/ 1/p19-1452
Vaswani, A., Shazee , N., Pa ma , N., Uszko ei , J., Jones, L., Gomez, A. N., … Polosukhin, I. (2017).
A en ion is All You Need. Ad ances in Neu al In o ma ion P ocessing Sys ems, 30, 5998–6008.
Vinyals, O., Fo una o, M., & Jai ly, N. (2015). Poin e Ne wo ks. P oceedings o he 28 h In e na ional
Con e ence on Neu al In o ma ion P ocessing Sys ems, 2, 2692–2700.
Wang, Y., Li, S., & Yang, J. (2018). Towa d as and accu a e neu al discou se segmen a ion.
P oceedings o he 2018 Con e ence on Empi ical Me hods in Na u al Language P ocessing,
EMNLP 2018, 962–967. h ps://doi.o g/10.18653/ 1/d18-1116
We bos, P. (1974). Beyond Reg ession: New Tools o P edic ion and Analysis in he Beha io al
Science. Ha a d Uni e si y.
Wol , T., Debu , L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., … Rush, A. (2020). T ans o me s:
S a e-o - he-A Na u al Language P ocessing. P oceedings o he 2020 Con e ence on Empi ical
Me hods in Na u al Language P ocessing: Sys em Demons a ions, 38–45.
h ps://doi.o g/10.18653/ 1/2020.emnlp-demos.6
74
Yan, Y., Rosales, R., Fung, G., Sub amanian, R., & Dy, J. (2014). Lea ning om mul iple anno a o s
wi h a ying expe ise. Machine Lea ning, 95(3), 291–327. h ps://doi.o g/10.1007/s10994-013-
5412-1
Zhang, X., Wei, F., & Zhou, M. (2020). Hibe : Documen le el p e- aining o hie a chical bidi ec ional
ans o me s o documen summa iza ion. ACL 2019 - 57 h Annual Mee ing o he Associa ion
o Compu a ional Linguis ics, P oceedings o he Con e ence, 5059–5069.
h ps://doi.o g/10.18653/ 1/p19-1499
Zhou, M., Duan, N., Liu, S., & Shum, H.-Y. (2020). P og ess in Neu al NLP: Modeling, Lea ning, and
Reasoning. Enginee ing, 6(3), 275–290.
h ps://doi.o g/h ps://doi.o g/10.1016/j.eng.2019.12.014
Hoch ei e , Sepp & Schmidhube , Jü gen. (1997). Long Sho - e m Memo y. Neu al compu a ion. 9.
1735-80. 10.1162/neco.1997.9.8.1735.
75
Page | i