i
Cus ome Chu n P edic ion in Insu ance:
Lau a So ia Sau ho da Pon e e Cas o
Modeling Renewal P ice Elas ici y o he Wo ke s’
Compensa ion Po olio om Ociden al Segu os
In e nship epo p esen ed as pa ial equi emen o
ob aining he Mas e ’s deg ee in Ad anced Analy ics
ii
i
Cus ome Chu n P edic ion in Insu ance:
Modeling Renewal P ice Elas ici y o he Wo ke s’ Compensa ion Po olio om
Ociden al Segu os
Lau a So ia Sau ho da Pon e e Cas o
MAA
2021
i
ii
NOVA In o ma ion Managemen School
Ins i u o Supe io de Es a ís ica e Ges ão de In o mação
Uni e sidade No a de Lisboa
CUSTOMER CHURN PREDICTION IN INSURANCE:
MODELING RENEWAL PRICE ELASTICITY OF THE WORKERS’ COMPENSATION
PORTFOLIO FROM OCIDENTAL SEGUROS
by
Lau a So ia Sau ho da Pon e e Cas o
In e nship epo p esen ed as pa ial equi emen o ob aining he Mas e ’s deg ee in Ad anced
Analy ics
Ad iso : Jo ge Mo ais Mendes
Co Ad iso : Nuno An ónio
No embe 2021
iii
ACKNOWLEDGMENTS
Fi s , I would like o s a o hank my ad iso , P o esso Jo ge Mo ais Mendes, o being a pa o his
p ojec since i s beginning, suppo ing me in all he challenges ha I ha e aced. As well, my co-ad iso ,
P o esso Nuno An ónio, o helping me go u he wi h his ema kable ision and insigh s. No only by
all he pa ience and suppo hey ha e p o ided, bu also, by all he knowledge hey ha e passed du ing
he lec u es in he i s yea o he Mas e ’s deg ee.
Mo eo e , I wan o exp ess my g a i ude o my in e nship supe iso , João Ped o Oli ei a, o us ing
me wi h his p ojec and o all he suppo , knowledge, and oppo uni ies he has gi en me o hese
mon hs. Also, And é Pousinho and Joana Neg a, o all he expe ience hey ha e sha ed wi h me in he
i s pa o my in e nship, allowing me o unde s and he insu ance ma ke and, pa icula ly, he
Wo ke s’ Compensa ion b anch.
I would also like o hank G upo Ageas Po ugal o allowing me o join hem o one yea , and e e yone
who has sha ed wi h me his pa h. Fu he mo e, o all he non-li e b anch P icing and Business
Analy ics eam who ha e made me eel so welcome and we e always willing o help in any hing I would
need.
Las ly, I would like o hank my iends and amily o belie ing in me and being a pa o my academic
jou ney and li e.
i
ABSTRACT
Cus ome chu n has been inc easing in insu ance, mainly due o echnological imp o emen s ha
allow cus ome s o explo e o he insu ance p o ide s’ o e s. Gi en his, insu ance p o ide s need o
compe e among hem, no only o ge new cus ome s bu , o main ain hei own.
This epo esul s om a p ojec de eloped du ing an in e nship in G upo Ageas Po ugal, which has
di e en insu ance b ands such as Ociden al Segu os. This p ojec 's main goal was o model, wi h a
mon hly pe iodici y, cus ome chu n o his la e ’s Wo ke s’ Compensa ion po olio o imp o e he
company’s compe i i eness and, ul ima ely, p o i .
Many o he company’s cus ome chu n happens a hei policy enewal ime, whe e he only a iable
ha he company de ains con ol o e is he p ice (p emium) a ia ion. Hence, by conside ing he
p emium a ia ion and o he ele an p edic i e a iables, he goal was o p edic he p obabili y o a
gi en cus ome o chu n, allowing he company o op imize he cu en enewal’s p icing p ocess and
maximize his b anch’s p o i .
Thus, di e en a iables ha could in luence he company’s cus ome beha io we e collec ed—one
o hose was he cus ome ’s loca ion. Gi en he high dimensionali y ha such a iable would ep esen
and he small da ase a ailable o modeling, clus e ing analysis is used o c ea e new signi ican (wi h
ewe dimensions) cus ome geog aphical a eas.
Di e en supe ised lea ning algo i hms we e hen e alua ed acco dingly o hei pe o mance in
p edic ing cus ome chu n. The p edic i e models used we e a G adien Boos ing, an Ex eme G adien
Boos ing, a Logis ic Reg ession, and a Mul ilaye Pe cep on. Gi en ha he numbe o cus ome s ha
enew hei con ac s is much supe io o he numbe o cus ome s who chu n, Syn he ic Mino i y
O e sampling Technique (SMOTE) was used o c ea e less unbalanced da ase s (wi h syn he ic
samples) and e alua e he impac on he pe o mance o one o he models.
Las ly, o gua an ee a success ul in eg a ion o he models in o he enewal’s p icing p ocess, models
we e e alua ed acco dingly o he wo business goals. Fi s , by ansla ing he obse ed e alua ion
me ics in o p o i . Secondly, by assu ing ha he cus ome ’s p ice elas ici y would be cap u ed,
assu ing a mono onic inc easing ela ionship among he policy’s p emium a ia ion and p obabili y o
chu n.
KEYWORDS
Supe ised Lea ning; Classi ica ion; Cus ome Chu n P edic ion; Non-Li e Insu ance; Renewal P ice
Elas ici y; Clus e ing; Neu al Ne wo k; Logis ic Reg ession; S ochas ic G adien Boos ing; Ex eme
G adien Boos ing
INDEX
1. In oduc ion .................................................................................................................. 1
1.1. P oblem S a emen and Objec i e ........................................................................ 1
1.2. Summa y o he P ocess ........................................................................................ 2
2. Theo e ical F amewo k ................................................................................................ 3
2.1. Li e a u e Re iew in Chu n Modeling in Insu ance ............................................... 3
2.2. Machine Lea ning .................................................................................................. 5
2.2.1. Da a Imbalance ............................................................................................... 5
2.2.2. Fea u e Selec ion ............................................................................................ 8
3. Da a............................................................................................................................. 11
4. Me hodology .............................................................................................................. 13
4.1. Business Unde s anding ...................................................................................... 13
4.2. Da a Unde s anding ............................................................................................ 14
4.3. Da a P epa a ion ................................................................................................. 14
4.3.1. Da a Selec ion ............................................................................................... 14
4.3.2. Da a Cleansing .............................................................................................. 15
4.3.3. Fea u e Enginee ing ..................................................................................... 16
4.3.4. Da a In eg a ion and Fo ma ........................................................................ 19
4.4. Modeling .............................................................................................................. 20
4.4.1. Model Selec ion ............................................................................................ 20
4.4.2. Tes Design Gene a ion ................................................................................ 21
4.4.3. Model Cons uc ion and Assessmen ........................................................... 22
4.5. E alua ion ............................................................................................................ 27
4.6. Deploymen ......................................................................................................... 29
5. Resul s and Discussion ................................................................................................ 31
5.1. Resul s ................................................................................................................. 31
5.1.1. C oss-Valida ion Resul s ............................................................................... 31
5.1.2. Model Assessmen ....................................................................................... 31
5.1.3. Model’s E alua ion ....................................................................................... 33
5.2. Discussion ............................................................................................................ 36
6. Conclusions ................................................................................................................. 38
7. Limi a ions and Recommenda ions o Fu u e Wo ks ............................................... 40
8. Bibliog aphy ................................................................................................................ 41
9. Appendix ..................................................................................................................... 47
i
9.1. Appendix A - Backwa d Elimina ion (Logis ic Reg ession Final Fea u e Lis ) ...... 47
9.2. Appendix B - Cons uc ed Models’ Cha ac e is ics ............................................. 48
9.3. Appendix C - Final Model’s Fea u e Lis .............................................................. 49
4
Model
Used in Li e a u e
Logis ic Reg ession
Y. He e al. (2020)
Va eiadis e al. (2015)
Sunda kuma & Ra i (2015)
Bolancé e al. (2016)
Spi e i & Azzopa di (2018)
Decision T ee
Va eiadis e al. (2015)
Sunda kuma & Ra i, 2015)
Bolancé e al. (2016)
Dola abadi e al. (2017)
Spi e i & Azzopa di (2018)
Sc iney e al. (2020)
Suppo Vec o Machine
Y. He e al., (2020)
Va eiadis e al. (2015)
Sunda kuma & Ra i (2015)
Bolancé e al. (2016)
Dola abadi e al. (2017)
Spi e i & Azzopa di (2018)
Sc iney e al. (2020)
Neu al Ne wo k
Y. He e al. (2020)
Va eiadis e al. (2015)
Sunda kuma & Ra i (2015)
Bolancé e al. (2016)
Dola abadi e al. (2017)
Sc iney e al. (2020)
G adien Boos ing
Y. He e al. (2020)
Naï e Bayes Classi ie
Va eiadis e al. (2015)
Dola abadi e al. (2017)
Spi e i & Azzopa di (2018)
Sc iney e al. (2020)
Random Fo es
Y. He e al. (2020)
Spi e i & Azzopa di (2018)
Ex a T ees Classi ie
Y. He e al. (2020)
Table 2.1 Models used in li e a u e o insu ance chu n modeling
5
2.2. MACHINE LEARNING
Unsupe ised Lea ning
Unsupe ised lea ning is a o m o machine lea ning ha aims o g oup da a in o segmen s based on
simila a ibu es, o na u ally occu ing ends, pa e ns, and ela ionships hidden in he da a (McCue,
2015). S anda d algo i hms used o his kind o ask a e clus e ing, anomaly de ec ion, neu al
ne wo ks, and app oaches o lea ning la en a iable models (El Bouche y & De Souza, 2020).
The main goal o clus e ing analysis is o segmen he ini e unlabeled da ase in o a ini e and disc e e
se o hidden da a s uc u es (Xu & Wunsch, 2005). The esul o his analysis is se e al g oups o he
da a named clus e s. These a e subse s o da a g ouped when, acco ding o he s udied c i e ia, ha e
simila cha ac e is ics and, when no , sepa a ed in o di e en g oups (Rokach & Maimon, 2005).
Mo eo e , wo ypes o clus e ing echniques a e pa i ional clus e ing and hie a chical clus e ing. In
he i s , g oups a e c ea ed by pa i ioning he space in o a p e-de ined numbe o subspaces. On he
o he hand, in hie a chical clus e ing, he da a objec s a e g ouped in sequence by a hie a chical
s uc u e.
Supe ised Lea ning
Ano he o m o Machine Lea ning is supe ised lea ning. Con a y o unsupe ised lea ning, i uses a
da ase ha has al eady been classi ied (labeled) as a basis o p edic ing he classi ica ion o o he
unlabeled da a (Talabis e al., 2015). The labeled da ase is a aining se composed o inpu a iables
( ea u es) and an ou pu a iable (label). The used ea u es will in luence he model’s abili y o
co ec ly classi y he p edic ed a iable, he ou pu a iable (L. Wang e al., 2021).
Supe ised Lea ning can be di ided in o classi ica ion when he label is ca ego ical and eg ession
when he label is con inuous. The e o e, he s udied p oblem is a classi ica ion ask since chu n
modeling can be ansla ed o a bina y classi ica ion p oblem: assuming 1 when he cus ome chu ns
and 0 when he cus ome does no chu n.
2.2.1. Da a Imbalance
Fo many eal supe ised lea ning p oblems in ol ing a bina y esponse a iable, da ase s p esen a
skewed dis ibu ion, ha ing one class wi h a much lowe ep esen a ion han ano he . When he
da ase shows his ype o beha io , ha ing an unde ep esen ed class, he da a is said o be
unbalanced (S. Wang & X. Yao, 2012). Resea che s ha e concluded ha his imbalance causes a
subop imal classi ica ion pe o mance (Chawla e al., 2004), since classi ie s end o gi e much highe
impo ance o he la ge classes.
H. He & Ga cia (2009) b ing ha . Usually, classi ie s, when dealing wi h an imbalanced da ase , end
o “p o ide a se e ely imbalanced deg ee o accu acy, wi h he majo i y class ha ing close o 100
pe cen accu acy and he mino i y class ha ing accu acies o 0-10 pe cen ”, such consequence can
ep esen high cos s o some indus ies so, he e o e, is i al o cons uc a model ha “will p o ide
high accu acy o he mino i y class wi hou se e ely jeopa dizing he accu acy o he majo i y class”.
In o de o deal wi h imbalanced da ase s, h ee possible app oaches can be aken: da a le el,
algo i hmic le el, and combining o ensemble me hods, o which he i s in ol es esampling o
educe he class skewness (Yap e al., 2014). Resampling is done ei he by emo ing ins ances om he
6
majo i y class (unde sampling) o adding ins ances o he mino i y class (o e sampling) by using
algo i hms such as SMOTE.
Syn he ic Mino i y O e sampling Technique
Syn he ic Mino i y O e sampling Technique (SMOTE) (Chawla e al., 2002) is one o he mos popula
and in luen ial da a p e-p ocessing algo i hms o deal wi h he da a imbalance p oblem (Ga cía e al.,
2016).
This echnique is an o e sampling app oach, said ha new ins ances om he smalle class a e
in oduced in o he da ase . Unlike basic app oaches such as andom o e sampling (ROS), which only
duplica es samples om he mino i y class, SMOTE gene a es new syn he ic samples, o e coming he
o e i ing caused by app oaches like ROS (Fe nández e al., 2018).
The i s s ep o his echnique is o de ine he amoun o o e sampling. He e, i is possible o ei he
se up his alue o app oxima e a balanced class dis ibu ion o disco e i ia a w appe p ocess
(Chawla e al., 2008). Then, based on k nea es neighbo s and linea in e pola ion ideas, he syn he ic
samples a e c ea ed. SMOTE ope a es in he ea u e space a he han in he da a space, each mino i y
class sample is conside ed along wi h i s k nea es neighbo s, and he new samples a e in oduced
along he line segmen s joining hem (conside ing any/all o he k neighbo s) (Chawla e al., 2002).
2.2.1.1. Model Calib a ion
When using such echniques o a i icially ebalancing he da ase o e en by consequence o he
da ase ’s cha ac e is ics, he aining and es se s ha e di e en dis ibu ions. This di e ence in
aining and es se s dis ibu ion iola es he basic assump ion in machine lea ning ha bo h a e
d awn om he same unde lying dis ibu ion (Pozzolo e al., 2015). By iola ing his assump ion, he
p edic ions ob ained in he es se will be biased and, he e o e, enhance he need o p obabili y
calib a ion o ob ain unbiased p edic ions.
Fu he mo e, some me hods end o bias p edic ed p obabili ies by pushing away o close o 0 and 1,
enhancing he need o calib a ion (Niculescu-Mizil & Ca uana, 2005). Do mann (2020) e en s a es
ha “i should be applied o any model ype as pa o he p edic ion p ocess, be o e p edic ing, c oss-
alida ing and making e ec plo s and maps o using p edic ions in any o he p obabilis ic
in e p e a ion”.
Two model calib a ion me hods ha can be used o co ec hese biased p obabili ies a e Pla Scaling
and Iso onic Reg ession. The i s is mo e e ec i e when he dis o ion is sigmoid-shaped, and he
la e is a mo e obus me hod ha can co ec any mono onic dis o ion bu , mo e p one o
o e i ing (Niculescu-Mizil & Ca uana, 2005).
The Pla Calib a ion calib a es p obabili ies by passing he ou pu h ough a sigmoid:
𝑃(𝑦=1|𝑓)= 1
1+𝑒(𝐴𝑓+𝐵)
( 1 )
Whe e 𝑓(𝑥) is he lea ning me hod and pa ame e s 𝐴 and 𝐵 a e es ima ed using G adien Descenden ,
such ha hey a e a solu ion o he ollowing minimiza ion unc ion:
7
𝑎𝑟𝑔𝑚𝑖𝑛𝐴,𝐵{− ∑𝑦𝑖 log(𝑝𝑖)+(1−𝑦𝑖)log(1−𝑝𝑖)
𝑖}
( 2 )
whe e,
𝑝𝑖= 1
1+𝑒(𝐴𝑓𝑖+𝐵)
( 3 )
On he o he hand, he Iso onic Calib a ion is mo e gene al gi en ha he only es ic ion is ha he
mapping unc ion is iso onic (Niculescu-Mizil & Ca uana, 2005). The basic assump ion o Iso onic
Reg ession (on which he model is based) is ha :
𝑦𝑖=𝑚(𝑓𝑖)+𝜖𝑖
( 4 )
Whe e, 𝑦𝑖 a e he ue labels, 𝑓𝑖 he model’s p edic ions, 𝑚 a mono onic inc easing (iso onic) unc ion.
The goal is o ind 𝑚, using he ue labels and model’s p edic ions as a aining se , such ha :
𝑚=𝑎𝑟𝑔𝑚𝑖𝑛𝑧∑(𝑦𝑖−𝑧(𝑓𝑖))2
( 5 )
One algo i hm ha can be used o ind a s epwise cons an solu ion o his p oblem is he pai -
adjacen iola o s (PAV) algo i hm (Aye e al., 1955).
Las ly, one way o assess how well-calib a ed he model is can be h ough a calib a ion plo . He e, “a
se o p edic ions o a bina y ou come is well calib a ed i he ou comes p edic ed o occu wi h
p obabili y p do occu abou p ac ion o he ime, o each p obabili y p ha is p edic ed” (Naeini e
al., 2015), which can be ansla ed in o a s aigh line om (0,0) o (1,1). Gi en his, in he calib a ion
plo , he x-axis ep esen s he a e age p edic ed p obabili y in each bin. The y-axis ep esen s he
obse ed ac ion o samples in he bin whose eal labels a e posi i e. The obse ed cu e is hen
compa ed wi h he s aigh line 𝑦=𝑥.
2.2.1.2. Th eshold-Mo ing Me hod
A echnique ha should be conside ed when dealing wi h class imbalance is changing he decision
h eshold (model’s con inuous ou pu cu -o ) and adap ing i o a pe o mance me ic. The main
di e ence be ween ebalancing (using echniques like SMOTE) and h eshold-based me hods is ha
he la e elies on manipula ing he con inuous ou pu o a lea ned model ins ead o elying on da a
p e-p ocessing be o e he lea ning happens (Collell e al., 2018).
P o os (2008) e en s a es ha “The bo om line is ha when s udying p oblems wi h imbalanced da a,
using he classi ie s p oduced by s anda d machine lea ning algo i hms wi hou adjus ing he ou pu
h eshold may well be a c i ical mis ake”.
The h eshold mo ing me hod uses he o iginal aining se o ain and unes o shi s he decision
h eshold by adap ing i o a pe o mance me ic. One possible app oach is o use he ROC e alua ion
p ocedu e and mo e om whe e misclassi ica ions a ain hei maximum on he posi i e class o he
poin whe e he maximum in he nega i e class is a ained, selec ing he poin whe e he cu e a ains
i s maximum (H. He & Ga cia, 2009).
8
2.2.2. Fea u e Selec ion
A well-known p oblem is he “cu se o ini e sample size”, o which he ela ionship among he
numbe o samples a ailable and he ea u es conside ed o modeling needs o be conside ed (Jain &
Chand aseka an, 1982). Each new ea u e in oduced o he model will ep esen a new dimension.
The highe he dimensionali y, he spa se he da ase becomes and, hus, lowe he ea u e space
co e age (Ve leysen & F ançois, 2005). Consequen ly, he p oblem’s complexi y apidly g ows wi h he
in oduc ion o mo e dimensions. Bellman (1966) in oduced he e m Cu se o Dimensionali y o
explain such phenomena.
The e o e, and gi en he nowadays exis ing high-dimensional da a, ea u e selec ion is one o he
essen ial echniques in da a p ep ocessing by elimina ing i ele an , edundan , o noisy ea u es
(Kalousis e al., 2007). Pe o ming such echniques allows as e algo i hms and, besides imp o ing
p edic i e powe , also imp o es comp ehensibili y (Kuma & Minz, 2014).
I is possible o b oadly classi y ea u e selec ion me hods in o il e and w appe me hods, whe e he
i s anks ea u es based on s a is ical measu es, independen ly o he lea ning algo i hm (Koha i &
John, 1997). One example o his ype o me hod is o use as impo ance measu e (sco e) he a iable’s
co ela ion wi h he a ge . On he o he hand, w appe s e alua e each candida e subse o ea u es'
impac on a pa icula lea ning algo i hm. The la e app oach usually allows achie ing be e esul s
gi en hei close in e ac ion wi h he classi ie (El Aboudi & Benhlima, 2016).
Some examples o w appe me hods a e o wa d selec ion, backwa d elimina ion, and ecu si e
ea u e elimina ion. The i s keeps adding new ea u es which imp o e he model pe o mance un il
no o he ea u e espec s he c i e ia. The second wo ks simila ly bu opposi ely, s a s wi h all
ea u es, and emo es he leas signi ican e ec ha does no mee he model’s s aying c i e ia un il
all ea u es a e signi ican . In bo h ( o wa d and backwa d elimina ion), he ea u e s ays in he model
once added/ emo ed. Las ly, ecu si e ea u e elimina ion is simila o he o wa d selec ion me hod.
In his case, e ec s a e added and emo ed in o he model such ha one o mo e backwa d
elimina ion s eps can happen a e a o wa d selec ion s ep (Bu sac e al., 2008).
Ano he amily de i ed om he wo p e ious me hod amilies ( il e and w appe s) a e embedded
me hods. These me hods combine he classi ie de elopmen wi h he sea ch o he op imal subse o
ea u es, cap u ing dependencies a a lowe compu a ional cos han w appe s bu (like w appe s) also
ha e a isk o o e i ing (Seijo-Pa do e al., 2017). Examples o embedded me hods a e Lasso and
Ridge eg ession, which ha e buil -in ea u e selec ion me hods ha employ L1 and L2 egula iza ion
( espec i ely).
2.2.2.1. Ensemble Lea ning o Fea u e Selec ion
Ensemble Lea ning is a ype o lea ning whe e mul iple models a e ained and combined o sol e he
same p oblem (Polika , 2006). Ensemble Lea ning is based on he assump ion ha combining he
solu ion o mul iple expe s is be e han using he solu ion o a single one. In his way, a se o
hypo heses is cons uc ed a he han using only one single hypo hesis o explain he da a, making i
possible o educe bias and a iance om he lea ning algo i hms (Die e ich, 2002).
Al hough usually employed o imp o e classi ica ion esul s, i is also possible o use ensemble lea ning
as a ea u e selec ion echnique. Combining mul iple ea u e selec ion me hods (ins ead o elying on
9
jus one) makes i possible o a ain mo e obus ea u e subse s, showing a g ea p omise o high-
dimensional da ase s wi h small sample sizes (Saeys e al., 2008).
The app oach bene i s om di e si y and con ol o a iance and is possible o employ in wo ways:
one is o use he same algo i hm o e ie e he ea u es’ impo ance using di e en subse s o he
da a (da a pe u ba ion / homogeneous) and, o he , is o use di e en ea u e selec ion echniques in
he same da ase ( unc ion pe u ba ion / he e ogenous) (Chiew e al., 2019).
Said his, di e en le els can be a ied and shall be chosen when employing ensemble lea ning in
ea u e selec ion, ollowing Bolón-Canedo & Alonso-Be anzos (2019) can be de ined as ollows:
• Da ase Le el: use di e en subse s o da a (o no )
• Fea u e Le el: use di e en subse s o ea u es (o no )
• Lea ne Me hod Le el: use o design di e en lea ning algo i hms (o no )
• Combina ion Le el: Use o design di e en combina ion/agg ega ion me hods
• Th eshold Le el: use o di e en h esholding me hods (in case o using anke me hods)
Bolón-Canedo & Alonso-Be anzos (2019) also e i ied ha he e ogeneous ea u e selec ion ensembles
a e mo e commonly used han homogeneous ones. Howe e , i is possible o ob ain ei he a ea u e
subse o a ea u e anking in bo h cases, depending on he ype o ea u e selec o s. Fo he la e , a
h eshold me hod needs o be de ined.
When using ea u e anke s, se e al ea u e anking algo i hms o he ensemble a e combined,
c ea ing a inal anked lis o he ea u es, gi en he ea u es’ ele ance o p edic ion. The app oach
o combining he ensemble membe s' esul s ( anks) has di e en p oposals in he li e a u e, om
simple o mo e complex solu ions (Seijo-Pa do e al., 2015). Some o he mos popula s aigh o wa d
me hods o combine such anks a ibu ed o each ea u e a e: minimum (bes ) ank, median ank,
a i hme ic mean ank, and geome ic mean (Bolón-Canedo & Alonso-Be anzos, 2019).
Gi en he h eshold me hod decision, he mos common app oach is de ining a ixed pe cen age o
he op ea u es, bu his pe cen age depends on he used da ase (Bolón-Canedo & Alonso-Be anzos,
2019). The e o e, his echnique is no op imal since i is p one o o e s a ing o unde s a ing his cu -
o alue (Chiew e al., 2019).
Pe mu a ion ea u e impo ance
Pe mu a ion ea u e impo ance (PFI) is a model-agnos ic ea u e selec ion me hod. The e o e, i is
possible o pe o m a he e ogeneous app oach by using his algo i hm as a base lea ne o he
ensemble and in oduce a iabili y by using di e en base models in he algo i hm.
This pe mu a ion ea u e impo ance measu emen was i s ly in oduced o andom o es s in 2001
(B eiman, 2001) bu can be used in any model. PFI assesses he a iable’s impo ance o he gi en
model when i s ela ionship wi h he a ge is b oken by obse ing he dec ease in he model’s sco e
when andom noise eplaces a a iable (in oduced by andomly shu ling he a iable’s alues)
(McGo e n e al., 2019). The e o e, i is possible o unde s and how impo an a gi en ea u e is o he
model’s abili y o p edic he a ge co ec ly.
10
PFI’s ope a ion me hod makes i sensible o co ela ed ea u es (S obl e al., 2008) and is, he e o e,
essen ial o add ess his issue p ima ily. Howe e , PFI has ad an ages such as obus ness no o bias
he measu es a o ing high ca dinally ea u es o e bina y ea u es.
11
3. DATA
Gi en he small dimensions o he Wo ke s’ Compensa ion po olio, he main goal in he da a
collec ion phase was o ex ac as much da a as possible. Mo eo e , o gua an ee ha he collec ed
in o ma ion uly cap u es he eali y, a 3-mon h aging is equi ed.
The e o e, he da a ex ac ion p ocess was di ided in o wo phases. A i s phase, a he p ojec ’s
beginning, in which i was possible o ex ac da a ega ding he enewals and chu ns om Janua y
2017 un il Sep embe 2020, used o he models’ cons uc ion. Then, a second phase, du ing he
Model’s Assessmen phase, o be used as a es se . This new da ase con ained da a om Oc obe
2020 un il Feb ua y 2021, including Janua y, whe e mos o he policies enew hei con ac s. Bellow,
he Da a Ex ac ion phases can be seen in Figu e 3.1.
Gi en he p oblem con ex , he goal was o iden i y olun a y chu ne s who based hei decision on
p ice a ia ion. In his way, nei he in olun a y chu ns no cancela ions ou side he de ined enewal
window (explained in sec ion 4.1) a e included in he analysis.
Mo eo e , o gua an ee ha he in o ma ion ex ac ed would cap u e he e en ha was decided o
be s udied - olun a y chu ne s ha ha e chu ned by consequence o he p emium a ia ion - some
es ic ions ha e been applied o he da a:
• No include chu ns ha a e a consequence o , o example, bank up o he company.
• No include policies om employees o he company
• No include empo a y policies
• No include he g oup o policies lagged by he company ha ecei e di e en enewal’s
p icing p ocesses om emaining cus ome s (cus ome s wi h highe o al p emiums and
policies iden i ied o go h ough a p uning p ocess)
• No include policies ha cancel o issue hei exi be o e he de ined enewal window
• No include policies o which i was no possible o ex ac hei p emium alue
• No include annui ies wi h a du a ion o less han one yea
Policies ha ha e canceled hei con ac s o acqui e a new policy in he company (policy
cannibaliza ion) we e s ill conside ed o analysis, since i is possible ha he cus ome has cancelled
due o he p ice inc ease. Usually, cus ome s acqui e a new policy o ob ain a lowe p ice main aining
he same bene i s.
Figu e 3.1 Di e en da a ex ac ion phases
12
Wi h such cons ain s applied, i was possible o ex ac a da ase wi h he policies' annui y in o ma ion
be o e he enewal da e in which, o one yea , a pa icula policy only appea s once. Rega ding he
collec ed a iables, his decision was made conside ing he in o ma ion p esen in he li e a u e and
sugges ions made by he p ojec ’s s akeholde s. The collec ed a iables and hei desc ip ion we e no
included in his epo . Howe e , i is possible o esume he a iable’s in o ma ion in o he ollowing
ca ego ies:
• Ca ego y 1: Value paid o insu ance
• Ca ego y 2: Cus ome beha io
• Ca ego y 3: Cus ome awa eness o he inc ease
• Ca ego y 4: Cus ome ’s/P oduc ’s cha ac e is ics
• Ca ego y 6: Cus ome ’s claims/cos and i s managemen by he company
• Ca ego y 6: Posi ion o Ociden al Segu os in Ma ke
• Ca ego y 7: Cus ome ’s in e ac ion channel wi h he company
• Ca ego y 8: Cus ome ’s loyal y
• Ca ego y 9: Cus ome ’s geog aphical loca ion cha ac e is ics
• Ca ego y 10: Pandemic si ua ion
• Ca ego y 11: Discoun s
Gi en ha he model needs o deli e p edic ions 60 days be o e he policy’s enewal da e (also
explained in sec ion 4.1), all he a iables collec ed had o conside his ision on ime. Mos a iables
main ain cons an along wi h he annui y. Howe e , o he s, such as he employees’ o al sala y
associa ed wi h he con ac , usually a y du ing his pe iod. The e o e, he ex ac ed a iables we e
designed o ga he he co ec ision, wi h he end o annui y s anding o he 60 days be o e i .
As men ioned be o e, he mos c ucial a iable o measu ing i s impac on chu n is he p emium
a ia ion ob ained in he enewal’s p icing p ocess. Howe e , his p emium a ia ion also depends on
he con ac ’s employees’ o al sala y, which may a y om one annui y o ano he o e en in he 60
days window. In his way, i was decided ha he p emium a ia ion should only be measu ed,
excluding he a ia ion o his a iable. The e o e, only he con ac ’s a i a ia ion was conside ed
when measu ing he p emium a ia ion.
13
4. METHODOLOGY
4.1. BUSINESS UNDERSTANDING
As men ioned be o e, his p ojec is applied o he Wo ke s’ Compensa ion po olio om Ociden al
Segu os. To gua an ee he p ojec ’s success, he i s objec i e was o unde s and his business line
and he p ojec objec i es. Fu he , his knowledge has been con e ed in o da a mining goals, and a
p ojec plan has been designed.
Wo ke s’ Compensa ion insu ance has been manda o y in Po ugal since 1993 o Thi d-Pa y
companies’ wo ke s and ex ended o sel -employed wo ke s in 1997. This ype o insu ance has only
one co e age ha co e s he isk o acciden s in he employees’ wo kplace o on hei way om/ o
home. This co e age assu es he legally needed bene i s by consequence o any o he in ol ed
employees’ acciden s, co e ing all needed expenses o ensu e he wo ke ’s o al eco e y (o
compensa ion in case o disabili y o dea h). The paid p emium o his insu ance depends no only on
he o al secu ed employees’ sala y (sum insu ed), which may a y du ing he yea bu also on he
con ac ’s a e s ipula ed by he company (which can a y yea ly in he enewal’s p icing s a egy).
The enewal’s p icing p ocess p oceeds app oxima ely wo mon hs in ad ance (gi en he cus ome ’s
enewal da e). By law, he insu ance company mus send he enewal le e in o ming he cus ome
abou he nex annui y’s p emium 30 days in ad ance. Gi en his, he ma ke de ined ha his enewal
le e should be sen 45 days in ad ance, so his is (usually) when he cus ome is awa e o he a ia ion
in he p emium (con ac ’s a e). The e o e, based on his business knowledge, i was de ined ha only
a cancella ion ha happens in a window o 45 days p io and 50 days a e he end o he policy’s
enewal da e would be classi ied as chu n. Mo eo e , he annui y’s p emium a ia ion is de e mined
60 o 45 days be o e he end o he policy’s enewal da e, so all a iables ex ac ed espec his ime
window (using he a iable’s ision 60 days be o e he policy annui y end da e).
Ociden al Segu os has a bancassu ance channel and de ains 1.5% o he Wo ke s’ Compensa ion
insu ance ma ke sha e. The company’s po olio can be segmen ed in o h ee main cus ome g oups:
• Housekeepe s – Besides housekeepe s, which is he main componen o his g oup, o he
domes ic employees a e in eg a ed, as ga dene s, o example.
• Sel -Employees – In gene al, a e small companies whe e he insu ed en i y is he wo ke
himsel .
• Thi d-Pa y - Comp ises companies om di e se dimensions (mic o/small, medium, and big)
insu ing hei employees.
Al hough mos o Ociden al Segu os’ cus ome con ac s a e om he Housekeepe s’ segmen , Thi d-
Pa y ep esen s almos 85% o he po olio’s o al annual p emium. Such beha io is explained by
he ac ha , on a e age, he o al secu ed employees’ sala y o his segmen is much highe han o
he o he segmen s.
Rega ding claims, he mos signi ican equency o claims is obse ed o he Sel -Employees’
segmen . Howe e , was he Thi d-Pa y segmen o which, a he ime, he lowes p o i abili y was
obse ed.
20
Yeo-Johnson ans o ma ion is a powe ans o ma ion amily wi h simila p ope ies as Box-Cox
ans o ma ion, bu well de ined in he whole eal line, hus app op ia e o educe skewness and
app oxima e no mali y wi hou such limi a ion (Yeo, 2000). Box-Cox ep esen s a amily o powe
ans o ma ions (like squa e oo , log, o in e se ans o ma ions - which a e conside ed a way o mee
he no mali y assump ion) ha easily ind he op imal powe ans o ma ion o a gi en a iable o
s abilize i s a iance (T. Zhang & Yang, 2017). Ne e heless, he o iginal Box-Cox (Box & Cox, 1964)
ans o ma ion is only alid o a posi i e x, con a y o Yeo-Johnson ans o ma ion.
On he o he hand, i has also been shown ha Fea u e Scaling gene ally imp o es he pe o mance
o classi ica ion algo i hms (Bollegala, 2017). Wha gene ally leads o such imp o emen is ha
ea u es’ alues, na u ally, occupy di e en anges (some anging in housands, o he s in ens) and,
ea u es wi h highe anges end o ha e a mo e decisi e ole while aining he model. Fu he mo e,
he ea u e’s ela i e alue di e ence is o en mo e in o ma i e han i s absolu e alue (Bollegala,
2017). The e o e, o he han using he da a in i s o iginal ange o m, some scaling me hods we e
a emp ed. Such ans o ma ions we e pe o med by es ima ing he scaling pa ame e s using he
aining se and, he ea u e scaling me hod is hen applied in bo h aining and es se s.
4.4. MODELING
In his phase, he used modeling echniques a e selec ed. Also, a es design is c ea ed o build he
di e en models and co ec ly assess hem, unde s anding hei alue o esol ing he p oblem a
hand.
4.4.1. Model Selec ion
Ini ially, o his p ojec , ou p edic i e models we e selec ed and cons uc ed: a Logis ic Reg ession,
a Mul ilaye Pe cep on, a G adien Boos ing model, and an Ex eme G adien Boos ing model. These
me hods a e b ie ly explained below.
Logis ic Reg ession
Simila o linea eg ession, logis ic eg ession (LR) can include one o mul iple co a ia es ( a iables)
ha , join ly wi h unknown pa ame e s es ima ed om he da a, p oduce a linea (and con inuous)
p edic o . The logi unc ion is used o gua an ee ha he ou come a iable alls in a 0-1 ange.
Fu he mo e, i is possible o add a egula iza ion e m o he logis ic eg ession, encou aging he i ed
pa ame e s o be small and helping p e en o e i ing.
Mul ilaye Pe cep on
The Mul ilaye Pe cep on (MLP) is a eed- o wa d ne wo k composed o inpu neu ons, ou pu
neu ons, and hidden laye s o neu ons be ween hem. Neu al ne wo ks a e a b anch o a i icial
in elligence inspi ed in human b ains. He e, nume ous cells called neu ons p ocess in o ma ion in
pa allel, linked oge he in a ne wo k by synapses, whe e in elligence is a gued o be encoded.
Mo eo e , in a eed- o wa d ne wo k, in o ma ion only mo es o wa d, om he inpu nodes o he
hidden nodes and inally ou pu nodes (Zell, 1994), connec ed by weigh s and ou pu signals. A
nonlinea ans e unc ion modi ies he neu on’s weigh ed inpu s, also called he ac i a ion unc ion
(Njikam & Zhao, 2016). These supe posi ions o nonlinea ans e unc ions allow he mul ilaye
pe cep on o app oxima e highly non-linea unc ions (con a y o he logis ic eg ession) and
accu a ely gene alize when p esen ed wi h new, unseen da a (Ga dne & Do ling, 1998).
21
G adien Boos ing Model
The G adien Boos ing (GB), p oposed by F iedman (2001, 2002), belongs o he amily o boos ing
me hods, a ype o ensemble lea ning. In boos ing, a new model is added o he ensemble sequence
ained based on he e o o he whole ensemble lea ned so a (Na ekin & Knoll, 2013). Hence,
misclassi ied ins ances a e emphasized by ecei ing highe weigh s in he nex i e a ion, which, o
example, usually happens o ins ances nea he decision bounda y (Mei & Rä sch, 2003). These
weigh s ep esen he impo ance ha he gi en ins ance will ha e in he ollowing base lea ne
cons uc ion. The e o e, ins ances ha he p e ious base models ha e shown di icul y unde s anding
appea mo e o en in he aining da a (Y. Zhang & Haghani, 2015). The inal model ob ained by he
boos ing algo i hm will be a linea combina ion o he se e al base-lea ne s conside ing hei
pe o mance on he da ase (G. Wang e al., 2011).
Ex eme G adien Boos ing Model
The Ex eme G adien Boos ing model (XGBoos ) is an ensemble o Classi ica ion T ee (CART) and a
mo e e icien e sion o GB. This algo i hm is e y popula in he Machine Lea ning ield, ha ing i s
impac widely ecognized in many machine lea ning and da a mining challenges (Chen & Gues in,
2016). Some o he modi ica ions done o GB a e pa allel aining (which as ens he algo i hm), ou -
o -co e compu a ion (allowing ha da a is no loaded in o memo y), and spa se da a op imiza ion (in
handling and speed up compu a ion) (Bisong, 2019). Also, besides sh inkage and subsampling (used in
GB) as egula iza ion o ms o con ol o e i ing and a ain be e esul s, he XGBoos uses a
egula ized objec i e. Mo e in o ma ion abou he XGBoos algo i hm can be ound a Chen & Gues in
(2016).
4.4.2. Tes Design Gene a ion
In o de o choose he sui able model and model’s hype pa ame e s, each cons uc ed model should
be es ed and, he e o e, a es design needs o be de ined. The goal is ha he ob ained esul s while
aining/e alua ing he model a e close o he eali y o how each model will pe o m. The e o e, bo h
he da ase ’s cha ac e is ics and he way he model would be used we e conside ed.
Fi s ly, ega ding he da a, a possible endency in chu n a e ac oss he yea s has been obse ed o
he di e en cus ome segmen s in he Da a Unde s anding phase. Also, as men ioned be o e, he
policy da a main ains p ima ily s able h ough he yea s in mos measu es. Howe e , signi ican
changes usually happen in he policies’ p emium paid (and o he a iables i depends on). Mo eo e ,
gi en ha each policy and i s da a only appea , a he mos , once in each o he yea s, he e is an
annual ( empo al) dependency on he da a.
A possible app oach o espec such empo al dependency is o use ime-spli c oss- alida ion. He e,
he model’s gene aliza ion abili y is assessed using he a e age o he pe o mance me ics in he
c ea ed da a spli s. A he same ime, he empo al dependency is espec ed by always using he las
block o da a as alida ion. Gi en he applied inc ease in p emium, he model should p edic he
p obabili y o chu n o he cus ome s expec ed o enew each mon h in he yea and he e o e deli e
mon hly p edic ions. Gi en hese model pu poses, wel e spli s we e c ea ed (one spli o each mon h
in he yea ), using he alida ion se o he spli ’s las a ailable mon h, as shown in Figu e 4.4.
22
Figu e 4.4 Rep esen a ion o he designed ime-spli c oss- alida ion
4.4.3. Model Cons uc ion and Assessmen
4.4.3.1. Fea u e Selec ion wi h Ensemble Lea ning
As shown, one possible way o pe o m ea u e selec ion is o employ ensemble lea ning, and, o his,
he app oach needs o be designed. Gi en he popula i y o he e ogeneous ea u e selec ion me hods,
i was decided ha his app oach should be used, using he same subse s o da a and ea u es ac oss
he di e en lea ne s. Fo his, many o he sugges ions p esen ed by Shah & Pe e ia ko (2021) we e
conside ed and a e desc ibed below.
Fi s ly, as o he lea ne me hod le el, many ea u e anke s we e cons uc ed using he PFI algo i hm
wi h a di e en base model o gua an ee a iabili y in he solu ions. The selec ed models we e he
bes se o base-line models (classi ie s wi hou pa ame e unning) in e ms o ecall ( ue posi i e
a e). Said ha , he selec ed models we e he ones ha appea ed o ha e a be e unde s anding o
he small class.
Then, as o combina ion le el, i has been shown ha se e al echniques a e possible o combine he
esul s and c ea e a inal ank, om simple o mo e complex me hods. Some o he mo e
s aigh o wa d me hods a e using he median and mean o he esul s as agg ega ion ules. The
median o he anks was used o his wo k since i is less sensi i e o ou lie s, being a mo e obus
me ic han he mean. Howe e , ea u es p esen ing he same median ank we e a e wa d o de ed
by he mean o he obse ed anks.
Las ly, o he h eshold le el decision, i has been shown ha he mos common app oaches a e
conside ing a ixed h eshold o he op- anked ea u es. Howe e , such a echnique comes wi h a ew
downsizes. In his way, once ha ing each o he model’s esul s, ins ead o selec ing s aigh away he
op N ea u es (gi en he di icul y o selec ing he p ope alue o N), he used s a egy was sligh ly
di e en . A e ha ing he ea u e impac o each ea u e gi en by each o he lea ne s, he o al (sum)
ea u e impac can be calcula ed ( o each o he ea u es summing all model’s e u ned impac s),
along wi h he cumula i e impac (a e o de ing by he ea u e’s o al (sum) impac ). Then, a gi en
23
a io (𝑟∈ ]0,1[ ) o he o al cumula i e impac is applied and, he numbe o ea u es ha p esen a
cumula i e impac lowe han he a io o o al cumula i e impac will be he numbe o ea u es o
selec (N) – N will be he cu ing h eshold. Wi h his numbe (N), he Top N bes - anked ea u es a e
selec ed, acco ding o he ensemble’s lea ne s’ opinions combina ion ule ea ly de ined. I is impo an
o no e ha his echnique will ne e selec a iables a ibu ed wi h ea u e impo ance o ze o o
nega i e alues.
Howe e , using his app oach, he e is s ill a alue ha needs o be selec ed – he a io (𝑟). This a io
and i s alue decision will be explained u he . As has been concluded in he pas , he op imal ea u e
size depends no only on he ea u e-label dis ibu ion bu also on he used classi ie (Hua e al., 2005).
In his way, since each model is di e en , i was decided o c ea e a mo e model-agnos ic inal a io
decision (indi idually o each o he selec ed models).
Addi ionally, an o de ed ea u e lis was c ea ed o each baseline model based on he PFI algo i hm
esul s using he gi en model only. In his case, o each o he classi ie s is possible o choose he
ea u e selec ion me hod ha p o ides be e esul s: using ensemble lea ning o ea u e selec ion o
only he w appe me hod - bo h ha ing a common app oach o selec ing he numbe o ea u es om
he anked lis .
In he end, a ecu si e me hod was implemen ed o selec he inal ea u e lis o each model. The
o al cumula i e impac a io (𝑟) will decay un il a speci ic s opping c i e ion is eached ( he a io s a s
a 1 and decays in each in e ac ion). Fo each classi ie , he bes ea u e anking me hod (compa ing
he ensemble and single model’s anked ea u e lis s), da a scale (scaling only nume ical a iables),
and a io a e selec ed conside ing he combina ion’s pe o mance in p edic ing he mino i y class
(conside ing he ecall using he designed c oss- alida ion). In his way, he numbe o selec ed
ea u es depends on he da ase i sel and he used classi ie .
Fea u e Lis C ea ion
Gi en his, he used ea u e selec ion echnique e ie es he bes combina ion o :
• Fea u e Ranking Me hod – using he ensemble anking o single model PFI’s anking
• Scaling – Scaling nume ical ea u es (wi h S anda d Scale o Robus Scale ) o wi hou scaling
• Ra io – s a ing a 1 (using all ea u es) and decaying 0.001 in each in e ac ion; The a io s ops
dec easing when h ee consecu i e loops do no imp o e he c oss- alida ed sco e (0.001
decay was decided based on he da ase ’s cha ac e is ics).
The bes combina ion is he one wi h he bes ecall sco e in c oss- alida ion. Also, i mo e han one
combina ion p esen s he same ecall, he one wi h ewe ea u es is e ie ed.
Fea u e Lis Op imiza ion
A e his, he c ea ed ea u e lis s passed h ough he ollowing app oaches desc ibed below. He e,
he inal ea u e anking is also used as an impo ance measu e.
1. T y o add possible good ea u es o each model's lis
• A e he Da a Unde s anding phase, o each o he cus ome ’s segmen s, a lis o ea u es
is c ea ed wi h ea u es ha , based on his phase, a e belie ed o be good p edic o s o
he p oblem a hand.
24
• Fo each o he ea u es ha a e no on he gi en c ea ed lis , hey a e added (indi idually,
one a he ime) o he lis o see i he ecall sco e (in c oss- alida ion) imp o es - i yes,
should be added o he lis (manually).
2. T y o emo e possible bad/unnecessa y ea u es
• A p ima y s a i ied and shu led ain/ es pa i ion is c ea ed o 70/30 ( o speed up gi en
he nume ous combina ions o his phase)
• A each i e a ion, one o he ea u es is le ou (by in e se o de o impo ance de ined
o he gi en ea u es lis in he la e phase) - i he sco e imp o es/main ains, ha speci ic
ea u e is le ou o he nex in e ac ion.
• A e all he ea u es a e es ed, he p ocess epea s. The p ocess will epea un il no
changes a e done - o he gi en i e a ion, no ea u e emo al imp o es/main ains he
ecall sco e in he es spli .
• Gi en he emo ed ea u es o he p ima y ain/ es pa i ion, each o he ea u es
emo ed is es ed o be ou in he model (indi idually, each ea u e a he ime), and i s
c oss- alida ed ecall sco e is e u ned.
• I he a e age ecall sco e in he alida ion se s imp o es/main ains, he ea u e is
excluded (manually).
3. T y o eplace nume ical a iables o i s powe ans o ma ion
• In he Da a P epa a ion phase, Yeo-Johnson Powe ans o ma ion was used o
app oxima e he a iable’s dis ibu ion o no mal
• Each i e a ion eplaces one o he ea u es by hei ans o ma ion (by impo ance
o de / anking). I he sco e imp o es, hen i s ans o ma ion o he nex in e ac ion
eplaces ha speci ic ea u e.
• A e all he ea u es a e es ed, he p ocess epea s. The p ocess will epea un il no
changes a e done - no ea u e eplacemen imp o es he a e age ecall sco e o alida ion
se s.
• The selec ed ans o med ea u es a e manually eplaced in he inal ea u e lis .
4.4.3.2. Logis ic Reg ession’s Fea u e Selec ion wi h Backwa d Elimina ion
Fo he logis ic eg ession’s ea u e selec ion, an un egula ized Logis ic eg ession was cons uc ed
using all ea u es (a e he Da a P epa a ion phase), and Backwa d Elimina ion as ea u e selec ion
me hod was pe o med. As men ioned be o e, as he name sugges s, he leas signi ican ea u es a e
i e a i ely emo ed un il no o he e ec mee s he speci ied le el o emo al (Bu sac e al., 2008).
The emo al c i e ia a e based on he Wald es o indi idual pa ame e s, using a signi icance le el o
0.05 ( he p- alue cu -o was de ined as 0.05). He e, he null hypo hesis ha he coe icien o he
independen a iable is equal o ze o is es ed e sus an al e na i e hypo hesis ha he coe icien is
nonze o, which can be w i en as: 𝐻0:𝛽𝑥=0 𝑣𝑠 𝐻1:𝛽𝑥≠0 (Fo ho e e al., 2007).
Besides he collec ed a iables, some i e a ions and polynomial e ms we e in oduced and es ed o
hei signi icance in he model. Such in e ac ions among a iables we e added whene e he e we e
suspicions ha he impac o one a iable in he ou come also depended on ano he a iable’s alue.
Polynomial e ms we e added o co e / es cases whe e hei ela ionship wi h he a ge is no linea .
25
Las ly, T- es was used o es possible (ca ego ical) a iable agg ega ions, es ing i hei model
pa ame e s we e signi ican ly di e en o no (𝐻0:𝛽𝑥𝑖− 𝛽𝑥𝑗=0 𝑣𝑠 𝐻1:𝛽𝑥𝑖− 𝛽𝑥𝑗≠0). Fo he cases
in which he null hypo heses we e no ejec ed, he agg ega ion was conside ed by summing bo h
bina y a iables since, o he es ed cases, he posi i e e en is exclusi e (ne e happens
simul aneously in bo h).
Be o e his me hod was applied, ea u es ha p esen ed low a iances we e emo ed o p e en
mul icollinea i y. Such ea u es can lead o a singula ma ix (wi h de e minan equaling o ze o), which
means ha he ma ix has no in e se, making i impossible o es ima e he eg ession pa ame e s.
The esul ing se o ea u es and co esponding summa y s a is ics can be ound in Appendix A.
4.4.3.3. Model Cons uc ion
Thi d-Pa y is he segmen o which he highes pe cen age o o al p emium and he highes chu n
a e was obse ed. The e o e, he Thi d-Pa y segmen was he i s modeled segmen and o which
he esul s will be p esen ed.
As said be o e, c oss- alida ion was pe o med o assess he models and choose he co ec
pa ame e s. Besides pa ame e uning, in which mul iple pa ame e combina ions we e es ed o ind
he pa ame e s ha enhanced he esul s, o he s eps we e aken in his phase. Fo each model
cons uc ed, he ollowing desc ibed p ocesses we e p oceeded.
Fea u e Lis and Da a Scale Choice
He e, he combina ion o ea u e lis and da a scale (S anda d Scale , Robus Scale , o no scale ) ha
gua an eed he bes esul s o each model is selec ed and used in he subsequen phases.
Bellow, an example o G adien Boos ing model (wi hou pa ame e uning) o he Thi d-Pa y’s
da ase , compa ing i s c oss- alida ed esul s using all da ase ’s ea u es and he selec ed lis , is
p esen in Table 4.2. Fo his case, i is possible o obse e ha bo h esul s a e simila , indica ing ha
he lowes numbe o ea u es should be conside ed only.
# Fea u es
Speci ici y
T ain
Speci ici y
Valida ion
AUC
T ain
AUC
Valida ion
Recall
T ain
Recall
Valida ion
84
100%
+/- 0p.p.
99.7%
+/- 0p.p.
52.2%
+/- 0p.p.
50.2%
+/- 1p.p.
4.4%
+/- 1p.p.
0.6%
+/- 1p.p.
37
100%
+/- 0p.p.
97.2%
+/- 9p.p.
52.1%
+/- 0p.p.
50.8%
+/- 2p.p.
4.2%
+/- 1p.p.
4.4%
+/- 13p.p.
Table 4.2 G adien Boos ing’s ea u e lis selec ion (wi hou pa ame e uning)
Da a P e-P ocessing using SMOTE
As has been men ioned, one o he possible app oaches o deal wi h class imbalance is a da a-le el
app oach, whe e unde sampling o o e sampling can be used. Gi en he small dimensions o he
da ase s, using an unde sampling echnique was no app op ia e and, he e o e, an o e sampling
echnique was a emp ed. Gi en SMOTE’s popula i y due o i s simplici y and obus ness, his
echnique gene a ed new eco ds in o he aining se .
26
Al hough i is shown ha non- andom sampling echniques can highly imp o e he classi ie ’s
pe o mance, ha ing a 1:1 dis ibu ion ( he posi i e a e equaling he nega i e a e) migh be
un a o able (Fo man & Cohen, 2004). The e o e, smalle a ios o he mino i y class we e a emp ed
ia a w appe p ocess so ha he co ec new a io o he aining se s’ mino i y class o he model
in ques ion could be chosen.
Below, in Table 4.3, an example o he SMOTE’s pa ame e choice ha con ols he new pe cen age o
chu n obse ed in he da ase is p esen ed. I is impo an o no e ha his esampling is only done in
he aining se , so he eali y in which he model will pe o m can be espec ed.
Chu n Ra e in
T ain
Speci ici y
T ain
Speci ici y
Valida ion
AUC
T ain
AUC
Valida ion
Recall
T ain
Recall
Valida ion
o iginal
100%
+/- 0p.p.
98.5%
+/- 2p.p.
68.8%
+/- 3p.p.
51.3%
+/- 2p.p.
37.6%
+/- 5p.p.
4.0%
+/- 5p.p.
50%
98.9%
+/- 0p.p.
23.4%
+/- 11p.p.
95.9%
+/- 0p.p.
54.4%
+/- 5p.p.
92.8%
+/- 0p.p.
85.3%
+/- 11p.p.
33%
99.5%
+/- 3p.p.
32.1%
+/- 29p.p.
92.3%
+/- 0p.p.
54.7%
+/- 0p.p.
85.1%
+/- 1p.p.
77.2%
+/- 18p.p.
23%
99.8%
+/- 0p.p.
43.5%
+/- 26p.p.
87.2%
+/- 1p.p.
54.8%
+/- 1p.p.
74.5%
+/- 1p.p.
66.1%
+/- 27p.p.
17%
99.8%
+/- 0p.p.
55.3%
+/- 31p.p.
80.8%
+/- 1p.p.
55.4%
+/- 4p.p.
61.7%
+/- 2p.p.
55.6%
+/- 30p.p.
13%
100%
+/- 0p.p.
63.3%
+/- 26p.p.
75.6%
+/- 2p.p.
54.9%
+/- 3p.p.
51.2%
+/- 3p.p.
46.5%
+/- 29p.p.
12%
100%
+/- 0p.p.
65.8%
+/- 32p.p.
74.4%
+/- 1p.p.
53.8%
+/- 4p.p.
48.8%
+/- 2p.p.
41.9%
+/- 32p.p.
11%
100%
+/- 0p.p.
67.1%
+/- 33p.p.
73.3%
+/- 2p.p.
53.0%
+/- 3p.p.
46.6%
+/- 3p.p.
38.9%
+/- 35p.p.
Table 4.3 Example o Ex eme G adien Boos ing's - SMOTE's pa ame e choice (wi hou pa ame e
uning)
Model Calib a ion & Th eshold Tuning
As shown be o e, o ob ain good models (in e ms o p edic ion p obabili ies and p obabili ies’
classi ica ion), he model’s ou pu p obabili ies should be calib a ed, and he bes decision h eshold
o classi y he model’s esul in o chu n o enewal should be de ined.
Rega ding he p obabili y calib a ion, he model p edic ed p obabili ies and he ue alues a e used
o assess each model’s calib a ion pe o mance using he calib a ion plo . Then, mul iple copies o he
model, using k- old c oss- alida ion, a e i ed and, he p obabili ies p edic ed by hese models a e
hen calib a ed on he hold-ou se s. As calib a ion me hods, Pla and Iso onic Calib a ion we e
es ed. Fo his, he new p edic ed p obabili ies a e assessed by plo ing he calib a ion cu e and
compa ing i wi h bo h, he pe ec calib a ion line, and esul s be o e calib a ion.
27
Fo he choice o he h eshold, he h eshold wi h he op imal balance be ween alse posi i e
(speci ici y) and ue posi i e a e ( ecall), his is he op imal h eshold o he ROC cu e, was loca ed.
Wi h his knowledge, he me ics in c oss- alida ion a e op imized.
In o de o accomplish his, he ue posi i e a e o ecall (TPR) and ue nega i e a e o speci ici y
(TNR) a e compu ed o he p edic ions using a se o h esholds ( his can also be used o c ea e a ROC
Cu e plo ). Then, he geome ic mean (G-mean), which o mula is p esen ed below, ep esen s he
balance be ween bo h sco es ( he highe , he be e ).
𝐺𝑀𝑒𝑎𝑛= √(𝑇𝑃𝑅∗𝑇𝑁𝑅)
( 6 )
This me ic is obse ed o he di e en a emp ed h esholds, and he h eshold ha maximizes his
me ic, his is, ha op imizes he balance be ween TPR and TNR, is selec ed. Then, he op imal
h esholds ob ained ac oss he di e en spli s o he designed c oss- alida ion a e a e aged in o a inal
op imal h eshold and a e conside ed.
4.5. EVALUATION
In he E alua ion phase, he da a mining esul s we e assessed compa a i ely wi h he business success
c i e ia. The p ocess was also e iewed, and he inal model was chosen o deploymen .
The main goal o his p ojec is o imp o e he p o i ma gin o Ociden al’s Wo ke s’ Compensa ion
b anch by e aining cus ome s ha ha e in en ions o lea e gi en he p emium a ia ion su e ed in
he eno a ion p ocess. A p o i es ima ion analysis was pe o med o ensu e ha he inal model
would allow such an inc ease in he company’s e enue.
Fo his, some supposi ions we e cons uc ed based on he b anch’s business knowledge and a e
bellow explained. The esul s o his analysis a e based on each model’s con usion ma ix esul s and
o he business me ics.
• Re enue i he model p edic s co ec ly ha he cus ome enews (T ue Nega i es - TN)
𝐺𝑎𝑖𝑛𝑇𝑁 =𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚+∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚
( 7 )
I he model p edic s co ec ly ha he cus ome will enew, he p oposed p emium a ia ion
will main ain, e aining bo h cus ome p emium and p emium a ia ion.
• Re enue i he model p edic s ha cus ome chu ns, bu cus ome enews (False Posi i es -
FP)
𝐺𝑎𝑖𝑛𝐹𝑃=𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚
( 8 )
I he model p edic s ha he cus ome will chu n, hen he e is a p emium a ia ion dec ease
(o e en emo al). Gi en ha i is di icul o es ima e he p emium a ia ion o hose cases,
he wo s -case scena io (no inc ease) was conside ed.
• Re enue i he model p edic s co ec ly ha he cus ome chu ns (T ue Posi i es - TP)
𝐺𝑎𝑖𝑛𝑇𝑃 =𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚∗𝑟𝑒𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑟𝑎𝑡𝑖𝑜
( 9 )
28
The ue posi i es a e he cus ome s ha , i no model exis ed, would be los . Howe e , hese
cus ome s can be p ese ed by co ec ly iden i ying hei in en ions and applying con ingency
measu es (in his case, he dec ease in he p emium a ia ion).
Ne e heless, hese cases a e simila o he False Posi i e (FP) cases, whe e he model p edic s
ha hese cus ome s will also chu n. Thus, he e is s ill an inc ease in p emium ha is no
accoun ed o in his es ima ion.
Besides his, i is also obse able ha i is no possible o p ese e all isk by applying his
con ingency measu e: cus ome s s ill chu n when he e is no p emium a ia ion
(∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚=0%). A possible explana ion o such beha io is a be e p emium p oposal in
he compe i ion ha he company canno ma ch.
The e o e, i was decided o apply a e en ion a io, conside ing ha only 𝑋% o a ge ed
chu ns can be e e sed h ough he applied con ingency measu e. Gi en ha , such alue
needed o be es ima ed.
I was obse ed ha , by yea , 13% o o al chu n in he Thi d-Pa y segmen (wi h a s anda d
de ia ion o 2 p.p.) happens o a 0% p emium a ia ion. The e o e, i was decided ha such
alue should be ounded up conside ing i s s anda d de ia ion and, as well, gi e a 100%
ma gin as a sa e y measu e. Thus, conside ing he e en ion a io as:
𝑟𝑒𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑟𝑎𝑡𝑖𝑜𝑇ℎ𝑖𝑟𝑑𝑃𝑎𝑟𝑡𝑦=(1−2(0.13+0.02))=0.7
( 10 )
• Re enue i he model p edic s ha cus ome enews bu cus ome Chu ns (False Nega i es
- FN)
𝐺𝑎𝑖𝑛𝐹𝑁 =0
( 11 )
Those a e he cases whe e he model canno o esee he cus ome ’s chu n, so hey a e los ,
ha ing, o no ha ing a model. The e o e, nei he he cus ome ’s p emium no p emium
a ia ion is e ained.
Ha ing es ima ed he di e en gains ha each o he con usion ma ix’s measu es, he inc ease in o al
e enue is gi en by he di e ence be ween he a e age e enue wi h he model implemen ed (To Be)
and wi h no model (As Is):
𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒 𝑅𝑒𝑣𝑒𝑛𝑒=𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝑇𝑜 𝐵𝑒− 𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝐴𝑠 𝐼𝑆
( 12 )
whe e,
𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝐴𝑠 𝐼𝑆=#𝑅𝑤𝑙𝑠∗(𝐴𝑣𝑔 𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚+ 𝐴𝑣𝑔 ∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚)
( 13 )
and,
𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝑇𝑜𝐵𝑒=
𝑻𝑵𝑹 ∗ #𝑅𝑤𝑙𝑠 ∗ 𝐺𝑎𝑖𝑛𝑇𝑁 + 𝑭𝑷𝑹 ∗ #𝑅𝑤𝑙𝑠 ∗ 𝐺𝑎𝑖𝑛𝐹𝑃
+ 𝑻𝑷𝑹 ∗ #𝐶ℎ𝑛𝑠∗ 𝐺𝑎𝑖𝑛𝑇𝑃
( 14 )
29
Fo each cus ome segmen , he #𝑅𝑤𝑙𝑠 is he numbe o e i ied enewals in a gi en yea , he #𝐶ℎ𝑛𝑠
he numbe o e i ied chu ns in a gi en yea , 𝑻𝑵𝑹 he ue nega i e a e and, 𝑭𝑷𝑹 and 𝑻𝑷𝑹 as alse
and ue posi i e a es espec i ely.
Using he a e age me ic’s esul s (𝑻𝑵𝑹,𝑭𝑷𝑹,𝑻𝑷𝑹) ob ained in each o he models and he obse ed
alues (#𝑅𝑤𝑙𝑠 , #𝐶ℎ𝑠, 𝐴𝑣𝑔 𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚, 𝐴𝑣𝑔 ∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚) in he las h ee yea s, i was
possible o es ima e, o each model, he expec ed a e age inc ease in e enue wi h hei
implemen a ion.
𝐸𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒=1
3 ∑(𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒 𝑅𝑒𝑣𝑒𝑛𝑒𝑖)
𝑖 ∈ 𝑌
( 15 )
whe e 𝑌 a e he las h ee a ailable comple e yea s.
4.6. DEPLOYMENT
Then, he model’s deploymen is planned, and i s moni o ing and main enance plan is also designed.
A e a ull e iew, he model will inally be in eg a ed in o he company’s enewal p icing s a egy and
help he company inc ease he cus ome e en ion a e in he Wo ke s’ Compensa ion b anch.
Fo his pa , i was decided o deli e wo di e en au oma ed p ocesses, desc ibed below. In bo h,
Py hon is he p ima y ool used.
Mon hly P edic ions
The i s deli e able aims o deli e mon hly p edic ions o he mon h’s po en ial enewals and is
designed as ollows.
In a pa icula mon h and yea , he po en ial enewals da a is collec ed using SAS En e p ise Guide.
Py hon accesses he inal ou pu able and ec ea es in his da ase all o he necessa y da a
p epa a ion s eps. A his s age o he p ocess, he p emium a ia ion ha each policy will su e is
unknown. The company wan s o know each cus ome 's eac ion (p obabili y o chu n) o he di e en
possible p ice inc eases. The e o e, many p ice inc eases scena ios a e c ea ed, and each policy line
will be ec ea ed as many imes as he numbe o di e en p emiums inc eases.
The ained model is hen used o p edic he chu n p obabili y o each policy and di e en p emium
a ia ion scena ios. Then, a inal able, in which he di e en policies ( ows), possible p emium
a ia ion (columns), and cus ome eac ion (p obabili y o chu n as alue), is sen o SAS En e p ise
Guide. He e, an op imiza ion p ocess (de eloped by he company) o choose he p emium a ia ion
o each policy ha maximizes he global p o i will use he cons uc ed able as inpu .
Model Pe o mance Moni o ing
The second deli e able aims o deli e a p ocess in which he model’s pe o mance can be
con inuously assessed and u he in eg a ed in o a dashboa d.
This p ocess is simila o he mon hly p edic ions p ocess desc ibed abo e. The only di e ence is ha
bo h he a ge (chu n) and he p emium a ia ion a e al eady known when i uns. The goal is o
compa e he model’s pas p edic ions wi h wha happened (i policies ha e been chu ned o no ) and
unde s and i s heal h. The e o e bo h, he da a collec ion p ocess (in SAS En e p ise Guide), da a
36
Figu e 5.4 Es ima ed a e age inc ease in e enue o he Ex eme G adien Boos ing model wi h he
mono onic cons ain and Logis ic Reg ession
5.2. DISCUSSION
As s a ed ea lie , he p ojec had wo business goals:
1. Imp o e he b anches’ p o i by co ec ly iden i ying cus ome s who in ended o chu n (and
enew) wi h he p oposed p emium a ia ion.
I is c ucial o co ec ly iden i y he cus ome s ha will enew wi h he applied a ia ion as
well. The company needs o main ain he p emium a ia ion (inc ease) applied o he policies
in mos o he cases since i signi ican ly impac s p o i .
2. The selec ed model needs o be sui able o he enewal’s p icing p ocess.
The model aims o cap u e he cus ome s’ p ice elas ici y. The e o e, a mono onic ela ionship
be ween he a ge and P emium Va ia ion needs o be assu ed.
In e ms o p edic ing chu n, he e we e al eady some signs o he complexi y o he p oblem. Looking
a he p edic i e a iables’ co ela ion wi h he a ge , he me ic a iable showing a highe co ela ion
is P emium Va ia ion (0.10 conside ing Pea son co ela ion and 0.08 o Spea man co ela ion), and he
highe co ela ed non-me ic a iable is Flag Annual Paymen (wi h a C amme ’s V o 0.08). Mo eo e ,
he da ase ’s dimensions we e small, and o he Thi d-Pa y case, only a ound 9% o ins ances we e
classi ied as chu n. Such cha ac e is ics in he da ase can easily lead o p oblems such as o e i ing.
The e o e, i is essen ial o main ain he models simple.
Conside ing he model’s cons uc ion phase, i was hen possible o obse e ha models pe o med
be e , showing ewe signs o o e i ing, when hei complexi y was educed. Fo example, using
ewe lea ne s ( o ensemble models) o ewe neu ons/laye s ( o he neu al ne wo ks). On he o he
hand, he educ ion o he da a-skewness by including syn he ic samples ( ia SMOTE) in he aining
da ase did no imp o e esul s.
The esul s we e also posi i ely impac ed by selec ing a sui able p edic ion h eshold (di e en han
he de aul alue o 0.5) allowing a sui able ade-o be ween he ue posi i e ( ecall) and ue
nega i e a es (speci ici y). Mo eo e , gi en he deployed solu ion s uc u e, he ou pu p obabili ies
appea o ha e bene i ed om calib a ion, app oxima ing hem o hei ue alues. In Figu e 5.5, i is
37
possible o obse e he di e ences in he uncalib a ed and calib a ed p edic ions made by he Ex eme
G adien Boos ing model wi h he imposed mono onic ela ionship, using calib a ion plo s. Obse ing
Figu e 5.5(a), is possible o obse e ha he uncalib a ed model unde - o ecas s, his is he ob ained
p obabili ies a e smalle han expec ed. Obse ing Figu e 5.5(b), is possible o obse e ha calib a ed
p obabili ies a e close o diagonal line, he e o e sugges ing a be e calib a ed model.
By conside ing he ob ained esul s, i is possible o obse e he ollowing:
• The cons uc ed Ex eme G adien Boos ing using SMOTE in da a p e-p ocessing has shown
he highes le el o o e i ing and he lowes AUC on he alida ion and es se s.
• The highes AUC in bo h, c oss- alida ion and a e age o es , is ob ained o he G adien
Boos ing (wi h an ad an age o 0.9 p.p. in c oss- alida ion and 1.1 p.p. in a e age o es
esul s), ollowed by he Ex eme G adien Boos ing model.
By e alua ing he i s business goal (p o i inc ease), on a e age, he bes esul s a e ob ained o he
G adien Boos ing model. This model is also he leade when indi idually compa ing he impac ha
he achie ed me ics (in c oss- alida ion and a e age o es se s) ha e on he es ima ed a e age
inc ease in e enue.
Howe e , conside ing he second business goal, only he Logis ic Reg ession model could espec he
mono onic ela ionship be ween he P emium Va ia ion and he a ge . Howe e , i was possible o
o ce his mono onic ela ionship among bo h a iables using an Ex eme G adien Boos ing. By
compa ing his new model o he i s business goal’s leade (G adien Boos ing), i is possible o
obse e a dec ease in he c oss- alida ion’s AUC (o 1.1 p.p.) and in he a e age o es ’s AUC (o 1.9
p.p.). Such educ ion also impac s he es ima ed a e age inc ease in e enue, by a ound 11K €.
Though, his new model s ill has be e esul s han he cons uc ed Logis ic Reg ession. I is only
su passed by he cons uc ed (non-mono onic) G adien Boos ing and (non-mono onic) o iginal
Ex eme G adien Boos ing.
P edic ed P obabili y
Obse ed P obabili y
P edic ed P obabili y
Obse ed P obabili y
Figu e 5.5 Calib a ion plo s o XGBoos Mono onic p edic ions (a) Uncalib a ed model
P obabili ies. (b) Calib a ed model p edic ions.
(a)
(b)
38
6. CONCLUSIONS
This epo esul s om a se en-mon h p ojec de eloped in a one-yea in e nship in G upo Ageas
Po ugal. Bellow, he main conclusions o his p ojec a e p esen ed.
G upo Ageas Po ugal is one o he la ges insu ance p o ide s in Po ugal, wi h di e en b ands such
as Ociden al Segu os, which p o ides o e s in he li e and non-li e b anches. This la e o e s a
Wo ke s’ Compensa ion insu ance, which is manda o y in Po ugal.
The Ociden al’s Wo ke s’ Compensa ion po olio can be segmen ed in o h ee main isk g oups (Thi d-
Pa y, Sel -Employees, and Housekeepe s), which ha e di e en enewal’s p icing s a egies and
p esen di e en beha io s and a ailable in o ma ion.
Nowadays, cus ome s' ease in explo ing he a ailable op ions allows hem o change he insu ance
p o ide easily. Gi en he compe i i eness o he insu ance ma ke , companies need o ake ac ion o
e ain hei cus ome s since i has a signi ican impac on hei p o i .
The enewal’s p icing p ocess occu s yea ly a he cus ome s' enewal da e, p oposing a a ia ion o
he policy’s paid p emium. The company’s goal was o educe he obse ed chu n a es consequen o
his p ocess. The e o e, he company wan ed o in eg a e a model ha could iden i y ea ly signs o
chu n acco dingly o he di e en possible p emium a ia ions ha he gi en policy could su e . So,
by cap u ing i s cus ome s’ p ice elas ici y, p emium a ia ions can be co ec ly adjus ed, and
cus ome s ha p esen a ce ain isk le el o lea ing, possibly main ained.
The CRISP-DM me hodology was applied o accomplish his p ojec , co e ing i s main s eps, and going
back and o wa d when necessa y.
The i s pa was dedica ed o unde s anding he business i sel , i s goals, and which a iables could
p o ide help ul in o ma ion abou his cus ome beha io . A ound 80 a iables (nume ical and
ca ego ical) we e ex ac ed om he company’s da abase. Howe e , as obse ed in he da a
unde s anding phase, no all a iables we e ele an , especially when sepa a ed wi hin he h ee
cus ome segmen s. This phase was ins umen al in unde s anding some o he s eps ha should be
done in he Da a P epa a ion phase, whe e da a in i s aw o m we e p epa ed o modeling. He e,
ac i i ies like edundan ea u e emo al, handling missing alues, c ea ing new a ibu es, and
o ma ing he da a in a way ha would allow (and help) he modeling phase we e ca ied on.
Mo eo e , as an al e na i e o using municipali ies o dis ic s, new geog aphical a eas we e c ea ed.
Fo his, clus e ing analysis was pe o med and, o each o he cus ome segmen s, new geog aphical
(con iguous) a eas we e c ea ed based on he municipali ies’ ele an demog aphic in o ma ion and
obse ed chu n a es. Thus, i was possible o educe eigh een new dimensions (in case dis ic s we e
used) in o only ou ele an dimensions, using in o ma ion ega ding mo e han h ee hund ed
municipali ies.
A e ha ing he da ase p epa ed, ea u es o be used by each o he models a e selec ed. The goal
was o use a sui able numbe o ea u es gi en he numbe o ows a ailable o modeling. Mo e
ea u es ep esen a highe p oblem complexi y, and o which, models should ha e mo e da a o
unde s and. The ea u es a e chosen using ensemble lea ning, combining mul iple lea ne s’ opinions
abou he ea u e’s impo ance ankings. The Top N ea u es a e selec ed, o each model, based on
39
he esul s ob ained o each se o ea u es. Fo he Logis ic eg ession, backwa d elimina ion based
on he Wald’s es was used.
The model’s pe o mance was e alua ed ac oss all he modeling phases. Time-spli c oss- alida ion
was used so he exis ing empo al dependencies on he da a could be espec ed. A he same ime,
he model’s gene aliza ion abili y can be gua an eed, app oxima ing he alida ion esul s o he ac ual
model pe o mance.
The inal model selec ion decision was based on he company’s wo business goals: inc easing he
b anches’ p o i (acco dingly wi h he model’s pe o mance) and ha e a sui able model o op imizing
he enewal’s p icing p ocess. Fo his la e , a mono onic inc easing ela ionship among P emium
Va ia ion and P obabili y o Chu n needs o be assu ed.
The inal selec ed model was an Ensemble Lea ning echnique, he Ex eme G adien Boos ing model
o which he mono onic ela ionship was o ced. This model shows in c oss- alida ion an AUC o 59.0%
(wi h a s anda d de ia ion o 3 p.p.), a speci ici y o 74.4% (wi h a s anda d de ia ion o 7 p.p.) and,
las ly, a ecall o 43.5% (wi h a s anda d de ia ion o 10 p.p.). As o he a e age esul s in he used
es se s, an AUC o 56.4% (wi h a s anda d de ia ion o 2 p.p.), a speci ici y o 72.7% (wi h a s anda d
de ia ion o 12 p.p.) and, las ly, a ecall o 40.2% (wi h a s anda d de ia ion o 12 p.p.) we e obse ed.
Fu he mo e, his model can answe he company’s needs. Fi s ly, i is sui able o he op imiza ion o
he enewal’s p icing p ocess: o a gi en enewal yea and policy, he deli e ed model allows ha i
he company decides o inc ease (o dec ease) he p emium a ia ion, hen he p obabili y o chu n
will ei he main ain he same alue o inc ease (o dec ease) in i s alue. The e o e, pe mi ing da a-
d i en decisions. Las ly, he model allows an inc ease in p o i : i was es ima ed ha his model could
con ibu e o an annual a e age inc ease o a ound 70K €.
40
7. LIMITATIONS AND RECOMMENDATIONS FOR FUTURE WORKS
I is possible o no ice ha he deli e ed solu ion has some space o imp o emen since he esul s in
e ms o pe o mance measu es could be highe . Resul s a e highly dependen on he da a used and,
he e, some imp o emen s could be made, o o he echniques a emp ed, like o example he es o
o he classi ica ion algo i hms.
Se e al issues we e iden i ied du ing he da a collec ion ask. Many inconsis encies we e de ec ed
among he di e en da a sou ces ( he company is cu en ly cons uc ing a uni ied sou ce o
in o ma ion). Al hough such p oblems we e deal wi hin he bes way possible, di e en quali y issues
ha e eme ged, impac ing he inal collec ed da ase . The e o e, he quali y in he used da ase o
modeling could no be assu ed. I was also no possible o collec da a o many o he company’s
policies, impac ing he inal da ase ’s dimensions. E en ha ing da a o ou whole yea s, a highe
amoun o da a could su ely help models in hei esul s. Addi ionally, he e a e possibly o he
a iables ha a e no a ailable o he company bu could be ele an , like ex e nal ac o s ha migh
in luence he cus ome ’s beha io .
Rega ding da a p e-p ocessing, o he echniques could be es ed as o he missing alues impu a ion
echniques, a deepe ou lie de ec ion analysis, o he o e sampling echniques, and u he
dimensionali y educ ion ( ea u e selec ion) echniques could be explo ed as well.
Las ly, as was possible o obse e, some models p esen be e esul s han o he s. The e o e, i could
be in e es ing o explo e o he models ha a e mono onic o allow a mono onic cons ain . Two
models ha would be in e es ing o es a e Ligh GBM ( om Py hon’s Ligh GBM package (Ke e al.
2017)) and la ice-based models ( om Py hon’s Tenso Flow-La ice package (Google AI Blog, 2017)).
Simila ly o he Ex eme G adien Boos ing model ( om he xgboos package (Chen & Gues in, 2016)),
hese models allow he in oduc ion o mono onic cons ain s. Howe e , due o secu i y cons ain s
ha he company needs o ensu e, i was no possible o ob ain such packages in a iable ime o his
p ojec execu ion.
41
8. BIBLIOGRAPHY
Ahn, J., Hwang, J., Kim, D., Choi, H., & Kang, S. (2020). A Su ey on Chu n Analysis in Va ious Business
Domains. IEEE Access, 8, 220816–220839. h ps://doi.o g/10.1109/ACCESS.2020.3042657
Aye , M., B unk, H. D., Ewing, G. M., Reid, W. T., & Sil e man, E. (1955). An empi ical dis ibu ion
unc ion o sampling wi h incomple e in o ma ion. The Annals o Ma hema ical S a is ics,
26(4), 641–647. h ps://doi.o g/10.1214/aoms/1177728423
Ba is a, G. E., & Mona d, M. C. (2003). An analysis o ou missing da a ea men me hods o
supe ised lea ning. Applied A i icial In elligence, 17(5–6), 519–533.
h ps://doi.o g/10.1080/713827181
Bellman, R. (1966). Dynamic p og amming. Science, 153(3731), 34–37.
h ps://doi.o g/10.1126/science.153.3731.34
Bisong, E. (2019). Building Machine Lea ning and Deep Lea ning Models on Google Cloud Pla o m. In
Ap ess, Be keley, CA. h ps://doi.o g/10.1007/978-1-4842-4470-8_29
Bolancé, C., Guillen, M., & Padilla-Ba e o, A. E. (2016). P edic ing P obabili y o Cus ome Chu n in
Insu ance. In R. León, M. Muñoz-To es, & J. Mone a (Eds.), Modeling and Simula ion in
Enginee ing, Economics and Managemen . MS 2016. Lec u e No es in Business In o ma ion
P ocessing, ol 254 (pp. 82–91). Sp inge , Cham. h ps://doi.o g/10.1007/978-3-319-40506-3_9
Bollegala, D. (2017). Dynamic ea u e scaling o online lea ning o bina y classi ie s. Knowledge-
Based Sys ems, 129, 97–105. h ps://doi.o g/10.1016/j.knosys.2017.05.010
Bolón-Canedo, V., & Alonso-Be anzos, A. (2019). Ensembles o ea u e selec ion: A e iew and
u u e ends. In o ma ion Fusion, 52, 1–12. h ps://doi.o g/10.1016/j.in us.2018.11.008
Box, G. E., & Cox, D. R. (1964). An Analysis o T ans o ma ions. Jou nal o he Royal S a is ical Socie y:
Se ies B (Me hodological), 26(2), 211–243. h ps://doi.o g/10.1111/j.2517-6161.1964. b00553.x
B eiman, L. (2001). Random Fo es s. Machine Lea ning, 45, 5–32.
h ps://doi.o g/10.1023/A:1010933404324
Bu sac, Z., Gauss, C. H., Williams, D. K., & Hosme , D. W. (2008). Pu pose ul selec ion o a iables in
logis ic eg ession. Sou ce Code o Biology and Medicine, 3(17). h ps://doi.o g/10.1186/1751-
0473-3-17
Chawla, N. V., Bowye , K. W., Hall, L. O., & Kegelmeye , W. P. (2002). SMOTE: Syn he ic Mino i y
O e -sampling Technique. Jou nal o A i icial In elligence Resea ch, 16(1), 321–357.
h ps://doi.o g/10.5555/1622407.1622416
Chawla, N. V., Cieslak, D. A., Hall, L. O., & Joshi, A. (2008). Au oma ically coun e ing imbalance and i s
empi ical ela ionship o cos . Da a Mining and Knowledge Disco e y, 17, 225–252.
h ps://doi.o g/10.1007/s10618-008-0087-0
Chawla, N. V., Japkowicz, N., & Ko cz, A. (2004). Edi o ial: special issue on lea ning om imbalanced
da a se s. SIGKDD Explo ., 6, 1-6. h ps://doi.o g/10.1145/1007730.1007733
Chen, T., & Gues in, C. (2016). XGBoos : A Scalable T ee Boos ing Sys em. P oceedings o he 22nd
Acm Sigkdd In e na ional Con e ence on Knowledge Disco e y and Da a Mining, 785–794.
h ps://doi.o g/10.1145/2939672.2939785
42
Chiew, K. L., Tan, C. L., Wong, K., Yong, K. S., & Tiong, W. K. (2019). A new hyb id ensemble ea u e
selec ion amewo k o machine lea ning-based phishing de ec ion sys em. In o ma ion
Sciences, 484, 153–166. h ps://doi.o g/10.1016/j.ins.2019.01.064
Chuang, H. C., Chen, C. C., & Li, S. T. (2020). Inco po a ing mono onic domain knowledge in suppo
ec o lea ning o da a mining eg ession p oblems. Neu al Compu ing and Applica ions,
32(15), 11791–11805. h ps://doi.o g/10.1007/s00521-019-04661-4
Collell, G., P elec, D., & Pa il, K. R. (2018). A simple plug-in bagging ensemble based on h eshold-
mo ing o classi ying bina y and mul iclass imbalanced da a. Neu ocompu ing, 275, 330–340.
h ps://doi.o g/10.1016/j.neucom.2017.08.035
De Win e , J. C., Gosling, S. D., & Po e , J. (2016). Supplemen al Ma e ial o Compa ing he Pea son
and Spea man Co ela ion Coe icien s Ac oss Dis ibu ions and Sample Sizes: A Tu o ial Using
Simula ions and Empi ical Da a. Psychological Me hods, 21(3), 273–290.
h ps://doi.o g/10.1037/me 0000079.supp
Die e ich, T. G. (2002). Ensemble Lea ning. The Handbook o B ain Theo y and Neu al Ne wo ks,
2(1), 110–125.
Dola abadi, S. H., & Keynia, F. (2017). Designing o Cus ome and Employee Chu n P edic ion Model
Based on Da a Mining Me hod and Neu al P edic o . 2017 2nd In e na ional Con e ence on
Compu e and Communica ion Sys ems (ICCCS), 74–77.
h ps://doi.o g/10.1109/CCOMS.2017.8075270
Do mann, C. F. (2020). Calib a ion o p obabili y p edic ions om machine‐lea ning and s a is ical
models. Global Ecology and Biogeog aphy, 29(4), 760–765. h ps://doi.o g/10.1111/geb.13070
El Aboudi, N., & Benhlima, L. (2016). Re iew on W appe Fea u e Selec ion App oaches. 2016
In e na ional Con e ence on Enginee ing & MIS (ICEMIS), 1–5.
h ps://doi.o g/0.1109/ICEMIS.2016.7745366
El Bouche y, K., & De Souza, R. S. (2020). Lea ning in Big Da a: In oduc ion o Machine Lea ning. In
Knowledge Disco e y in Big Da a om As onomy and Ea h Obse a ion (pp. 225–249).
Else ie . h ps://doi.o g/10.1016/b978-0-12-819154-5.00023-0
Fe nández, A., Ga cía, S., He e a, F., & Chawla, N. V. (2018). SMOTE o Lea ning om Imbalanced
Da a: P og ess and Challenges, Ma king he 15-yea Anni e sa y. Jou nal o A i icial
In elligence Resea ch, 61, 863–905. h ps://doi.o g/10.1613/jai .1.11192
Fo man, G., & Cohen, I. (2004). Lea ning om li le: Compa ison o classi ie s gi en li le aining. In J.
Boulicau , F. Esposi o, F. Gianno i, & D. Ped eschi (Eds.), Knowledge Disco e y in Da abases:
PKDD 2004. PKDD 2004. Lec u e No es in Compu e Science, ol 320 (pp. 161–172). Sp inge ,
Be lin, Heidelbe g. h ps://doi.o g/10.1007/978-3-540-30116-5_17
Fo ho e , R. N., Lee, E. S., & He nandez, M. (2007). Logis ic and P opo ional Haza ds Reg ession. In
Bios a is ics (Second Edi ion) (pp. 387–419). Academic P ess. h ps://doi.o g/10.1016/B978-0-
12-369492-8.50019-4
F iedman, J. H. (2001). G eedy unc ion app oxima ion: a g adien boos ing machine. The Annals o
S a is ics, 29(5), 1189–1232. h ps://doi.o g/10.1214/aos/1013203451
F iedman, J. H. (2002). S ochas ic g adien boos ing. Compu a ional S a is ics & Da a Analysis, 38(4),
367–378. h ps://doi.o g/10.1016/S0167-9473(01)00065-2
43
Ga cía, S., Luengo, J., & He e a, F. (2016). Tu o ial on p ac ical ips o he mos in luen ial da a
p ep ocessing algo i hms in da a mining. Knowledge-Based Sys ems, 98, 1–29.
h ps://doi.o g/10.1016/j.knosys.2015.12.006
Ga dne , M. W., & Do ling, S. R. (1998). A i icial neu al ne wo ks ( he mul ilaye pe cep on) - a
e iew o applica ions in he a mosphe ic sciences. A mosphe ic En i onmen , 32(14–15), 2627–
2636. h ps://doi.o g/10.1016/S1352-2310(97)00447-0
Gün he , C. C., T e e, I. F., Aas, K., Sandnes, G. I., & Bo gan, Ø. (2014). Modelling and p edic ing
cus ome chu n om an insu ance company. Scandina ian Ac ua ial Jou nal, 2014(1), 58–71.
h ps://doi.o g/10.1080/03461238.2011.636502
Google AI Blog (2017). Tenso low la ice: Flexibili y empowe ed by p io knowledge. Re i ed om
h ps://ai.googleblog.com/2017/10/ enso low-la ice- lexibili y.h ml
Ha is, R., Sleigh , P., & Webbe , R. (2005). In oducing Geodemog aphics. In Geodemog aphics, GIS
and neighbou hood a ge ing (pp. 16–17). John Wiley & Sons.
He, H., & Ga cia, E. A. (2009). Lea ning om Imbalanced Da a. IEEE T ansac ions on Knowledge and
Da a Enginee ing, 21(9), 1263–1284. h ps://doi.o g/10.1109/TKDE.2008.239
He, Y., Xiong, Y., & Tsai, Y. (2020). Machine Lea ning Based App oaches o P edic Cus ome Chu n
o an Insu ance Company. 2020 Sys ems and In o ma ion Enginee ing Design Symposium
(SIEDS), 1–6. h ps://doi.o g/10.1109/SIEDS49339.2020.9106691
Hua, J., Xiong, Z., Lowey, J., Suh, E., & Doughe y, E. R. (2005). Op imal numbe o ea u es as a
unc ion o sample size o a ious classi ica ion ules. Bioin o ma ics, 21(8), 1509–1515.
h ps://doi.o g/10.1093/bioin o ma ics/b i171
Inouye, D. I., Leqi, L., Kim, J. S., A agam, B., & Ra ikuma , P. (2020). Au oma ed Dependence Plo s.
P oceedings o he 36 h Con e ence on Unce ain y in A i icial In elligence (UAI), 124, 1238–
1247. h p://a xi .o g/abs/1912.01108
Jain, A. K., & Chand aseka an, B. (1982). Dimensionali y and sample size conside a ions in pa e n
ecogni ion p ac ice. In Handbook o S a is ics (Vol. 2, pp. 835–855).
h ps://doi.o g/10.1016/S0169-7161(82)02042-2
Lemaî e, G., Noguei a, F., & A idas, C. K. (2017). Imbalanced-lea n: A py hon oolbox o ackle he
cu se o imbalanced da ase s in machine lea ning. The Jou nal o Machine Lea ning
Resea ch, 18(1), 559-563.
Kalousis, A., P ados, J., & Hila io, M. (2007). S abili y o ea u e selec ion algo i hms: a s udy on high-
dimensional spaces. Knowledge and In o ma ion Sys ems, 12, 95–116.
h ps://doi.o g/10.1007/s10115-006-0040-8
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W. Ma, W., Ye, Q., Liu, T.Y (2017). Ligh gbm: A highly
e icien g adien boos ing decision ee. In Ad ances in neu al in o ma ion p ocessing sys ems,
30, 3146-3154.
Koha i, R., & John, G. H. (1997). W appe s o ea u e subse selec ion. A i icial In elligence, 97(1–2),
273–324. h ps://doi.o g/10.1016/S0004-3702(97)00043-X
Kuma , V., & Minz, S. (2014). Fea u e Selec ion: A li e a u e Re iew. The Sma Compu ing Re iew,
4(3), 211–229. h ps://doi.o g/10.6029/sma c .2014.03.007
44
Male ic, J., & Ma cus, A. (2000). Da a Cleansing: Beyond In eg i y Analysis. Iq, 200–209.
McCue, C. (2015). Iden i ica ion, Cha ac e iza ion, and Modeling. In Da a Mining and P edic i e
Analysis (Second Edi ion) (pp. 137–155). h ps://doi.o g/10.1016/B978-0-12-800229-2.00007-9
McGo e n, A., Lage quis , R., John Gagne, D., Je gensen, G. E., Elmo e, K. L., Homeye , C. R., & Smi h,
T. (2019). Making he Black Box Mo e T anspa en : Unde s anding he Physical Implica ions o
Machine Lea ning. Bulle in o he Ame ican Me eo ological Socie y, 100(11), 2175–2199.
h ps://doi.o g/10.1175/BAMS-D-18-0195.1
Mei , R., & Rä sch, G. (2003). An In oduc ion o Boos ing and Le e aging. In S. Mendelson & A. J.
Smola (Eds.), Ad anced lec u es on machine lea ning (pp. 118–183). Sp inge , Be lin,
Heidelbe g. h ps://doi.o g/10.1007/3-540-36434-X_4
Naeini, M. P., Coope , G. F., & Hausk ech , M. (2015). Ob aining Well Calib a ed P obabili ies Using
Bayesian Binning. P oceedings o he Twen y-Nin h AAAI Con e ence on A i icial In elligence,
2901–2907.
Na ekin, A., & Knoll, A. (2013). G adien boos ing machines, a u o ial. F on ie s in Neu o obo ics, 7,
21. h ps://doi.o g/10.3389/ nbo .2013.00021
Niculescu-Mizil, A., & Ca uana, R. (2005). P edic ing good p obabili ies wi h supe ised lea ning.
P oceedings o he 22nd In e na ional Con e ence on Machine Lea ning, 625–632.
h ps://doi.o g/10.1145/1102351.1102430
Njikam, A. N. S., & Zhao, H. (2016). A no el ac i a ion unc ion o mul ilaye eed- o wa d neu al
ne wo ks. Applied In elligence, 45, 75–82. h ps://doi.o g/10.1007/s10489-015-0744-0
Osbo ne, J. W. (2010). Imp o ing you da a ans o ma ions: Applying he Box-Cox ans o ma ion.
P ac ical Assessmen , Resea ch and E alua ion, 15(12). h ps://doi.o g/10.7275/qbpc-gk17
Ped egosa, F., Va oquaux, G., G am o , A., Michel, V., Thi ion, B., G isel, O., Blondel, M.,
P e enho e , P., Weiss, R., Dubou g, V., Vande plas, J., Passos, A., Cou napeau, D., B uche , M.,
Pe o , M., & Duchesnay, E. (2011). Sciki -lea n: Machine Lea ning in Py hon. Jou nal o
Machine Lea ning Resea ch, 12, 2825–2830. Re ei ed om
h ps://jml .csail.mi .edu/pape s/ olume12/ped egosa11a/ped egosa11a.pd
Pham, D. T., Dimo , S. S., & Nguyen, C. D. (2005). Selec ion o K in K-means clus e ing. P oceedings o
he Ins i u ion o Mechanical Enginee s, Pa C: Jou nal o Mechanical Enginee ing Science,
219(1), 103–119. h ps://doi.o g/10.1243/095440605X8298
Polika , R. (2006). Ensemble based sys ems in decision making. IEEE Ci cui s and Sys ems Magazine,
6(3), 21–45. h ps://doi.o g/10.1109/MCAS.2006.1688199
Pozzolo, A. D., Caelen, O., Johnson, R. A., & Bon empi, G. (2015). Calib a ing P obabili y wi h
Unde sampling o Unbalanced Classi ica ion. 2015 IEEE Symposium Se ies on Compu a ional
In elligence, 159–166. h ps://doi.o g/10.1109/SSCI.2015.33
P ice, B. (2002). Making CRM come o li e. E-Business Re iew, 25–31.
P o os , F. (2008). Machine lea ning om imbalanced da a se s 101.
Rokach, L., & Maimon, O. (2005). Clus e ing Me hods. In O. Maimon & L. Rokach (Eds.), Da a Mining
and Knowledge Disco e y Handbook (pp. 321–352). Sp inge , Bos on, MA.
h ps://doi.o g/10.1007/0-387-25465-X_15
45
Rubin, D. B. (1976). In e ence and missing da a. Biome ika, 63(3), 581–592.
h ps://doi.o g/10.1093/biome /63.3.581
Saeys, Y., Abeel, T., & Van De Pee , Y. (2008). Robus Fea u e Selec ion Using Ensemble Fea u e
Selec ion Techniques. In W. Daelemans (Ed.), Machine Lea ning and Knowledge Disco e y in
Da abases. ECML PKDD 2008. Lec u e No es in Compu e Science, ol 5212 (pp. 313–325).
Sp inge -Ve lag Be lin Heidelbe g 2008. h ps://doi.o g/10.1007/978-3-540-87481-2_21
Sasse , W. E., & Reichheld, F. F. (1990). Ze o De ec ions -Quali y Comes o Se ices. Ha a d Business
Re iew, 68(5), 105–111.
Sc iney, M., Nie, D., & Roan ee, M. (2020). P edic ing Cus ome Chu n o Insu ance Da a. In M.
Song, I. Song, G. Ko sis, A. M. Tjoa, & I. Khali (Eds.), Big Da a Analy ics and Knowledge Disco e y.
DaWaK 2020. Lec u e No es in Compu e Science, ol 12393 (pp. 256–265). Sp inge , Cham.
h ps://doi.o g/10.1007/978-3-030-59065-9_21
Seabold, S., & Pe k old, J. (2010). S a smodels: Econome ic and s a is ical modeling wi h py hon. In
P oceedings o he 9 h Py hon in Science Con e ence. (Vol. 57, pp. 61).
Seijo-Pa do, B., Bolón-Canedo, V., & Alonso-Be anzos, A. (2017). Tes ing Di e en Ensemble
Con igu a ions o Fea u e Selec ion. Neu al P ocessing Le e s, 46, 857–880.
h ps://doi.o g/10.1007/s11063-017-9619-1
Seijo-Pa do, B., Bolón-Canedo, V., Po o-Díaz, I., & Alonso-Be anzos, A. (2015). Ensemble Fea u e
Selec ion o Rankings o Fea u es. In I. Rojas, G. Joya, & A. Ca ala (Eds.), Ad ances in
Compu a ional In elligence. IWANN 2015. Lec u e No es in Compu e Science, ol 9095.
Sp inge , Cham. h ps://doi.o g/10.1007/978-3-319-19222-2_3
Shah, R., & Pe e ia ko, V. (2021). Using Fea u e Impo ance Rank Ensembling (FIRE) o Ad anced
Fea u e Selec ion. Re i ed om h ps://www.da a obo .com/blog/using- ea u e-impo ance-
ank-ensembling- i e- o -ad anced- ea u e-selec ion/
Spi e i, M., & Azzopa di, G. (2018). Cus ome Chu n P edic ion o a Mo o Insu ance Company. 2018
Thi een h In e na ional Con e ence on Digi al In o ma ion Managemen (ICDIM), 173–178.
h ps://doi.o g/10.1109/ICDIM.2018.8847066
S obl, C., Boules eix, A. L., Kneib, T., Augus in, T., & Zeileis, A. (2008). Condi ional a iable
impo ance o andom o es s. BMC Bioin o ma ics, 9, 307. h ps://doi.o g/10.1186/1471-
2105-9-307
Sunda kuma , G. G., & Ra i, V. (2015). A no el hyb id unde sampling me hod o mining unbalanced
da ase s in banking and insu ance. Enginee ing Applica ions o A i icial In elligence, 37, 368–
377. h ps://doi.o g/10.1016/j.engappai.2014.09.019
Talabis, M., McPhe son, R., Miyamo o, I., & Ma in, J. (2015). Analy ics De ined. In In o ma ion
secu i y analy ics: inding secu i y insigh s, pa e ns, and anomalies in big da a (pp. 1–12).
Syng ess. h ps://doi.o g/10.1016/b978-0-12-800207-0.00001-0
Toloşi, L., & Lengaue , T. (2011). Classi ica ion wi h co ela ed ea u es: un eliabili y o ea u e
anking and solu ions. Bioin o ma ics, 27(14), 1986–1994.
h ps://doi.o g/10.1093/bioin o ma ics/b 300
Va eiadis, T., Diaman a as, K. I., Sa igiannidis, G., & Cha zisa as, K. C. (2015). A compa ison o
machine lea ning echniques o cus ome chu n p edic ion. Simula ion Modelling P ac ice and
Theo y, 55, 1–9. h ps://doi.o g/10.1016/j.simpa .2015.03.003