scieee Science in your language
[en] (orig)

Customer Churn Prediction in Insurance: Modeling Renewal Price Elasticity of the Workers’ Compensation Portfolio from Ocidental Seguros

Abstract

Customer churn has been increasing in insurance, mainly due to technological improvements that allow customers to explore other insurance providers’ offers. Given this, insurance providers need to compete among them, not only to get new customers but, to maintain their own. This report results from a project developed during an internship in Grupo Ageas Portugal, which has different insurance brands such as Ocidental Seguros. This project's main goal was to model, with a monthly periodicity, customer churn of this latter’s Workers’ Compensation portfolio to improve the company’s competitiveness and, ultimately, profit. Many of the company’s customer churn happens at their policy renewal time, where the only variable that the company detains control over is the price (premium) variation. Hence, by considering the premium variation and other relevant predictive variables, the goal was to predict the probability of a given customer to churn, allowing the company to optimize the current renewal’s pricing process and maximize this branch’s profit. Thus, different variables that could influence the company’s customer behavior were collected—one of those was the customer’s location. Given the high dimensionality that such variable would represent and the small dataset available for modeling, clustering analysis is used to create new significant (with fewer dimensions) customer geographical areas. Different supervised learning algorithms were then evaluated accordingly to their performance in predicting customer churn. The predictive models used were a Gradient Boosting, an Extreme Gradient Boosting, a Logistic Regression, and a Multilayer Perceptron. Given that the number of customers that renew their contracts is much superior to the number of customers who churn, Synthetic Minority Oversampling Technique (SMOTE) was used to create less unbalanced datasets (with synthetic samples) and evaluate the impact on the performance of one of the models. Lastly, to guarantee a successful integration of the models into the renewal’s pricing process, models were evaluated accordingly to the two business goals. First, by translating the observed evaluation metrics into profit. Secondly, by assuring that the customer’s price elasticity would be captured, assuring a monotonic increasing relationship among the policy’s premium variation and probability of churn.

Read accessible full text

Customer Churn Prediction in Insurance: Modeling Renewal Price Elasticity of the Workers’ Compensation Portfolio from Ocidental Seguros

Author: Castro, Laura Sofia Sauthoff da Ponte e
Year: 2022
Source: https://run.unl.pt/bitstream/10362/141571/1/TAA0157.pdf
i
Cus ome Chu n P edic ion in Insu ance:
Lau a So ia Sau ho da Pon e e Cas o
Modeling Renewal P ice Elas ici y o he Wo ke s’
Compensa ion Po olio om Ociden al Segu os
In e nship epo p esen ed as pa ial equi emen o
ob aining he Mas e ’s deg ee in Ad anced Analy ics
ii
i
Cus ome Chu n P edic ion in Insu ance:
Modeling Renewal P ice Elas ici y o he Wo ke s’ Compensa ion Po olio om
Ociden al Segu os
Lau a So ia Sau ho da Pon e e Cas o
MAA
2021
i
ii
NOVA In o ma ion Managemen School
Ins i u o Supe io de Es a ís ica e Ges ão de In o mação
Uni e sidade No a de Lisboa
CUSTOMER CHURN PREDICTION IN INSURANCE:
MODELING RENEWAL PRICE ELASTICITY OF THE WORKERS’ COMPENSATION
PORTFOLIO FROM OCIDENTAL SEGUROS
by
Lau a So ia Sau ho da Pon e e Cas o
In e nship epo p esen ed as pa ial equi emen o ob aining he Mas e ’s deg ee in Ad anced
Analy ics
Ad iso : Jo ge Mo ais Mendes
Co Ad iso : Nuno An ónio
No embe 2021

iii
ACKNOWLEDGMENTS
Fi s , I would like o s a o hank my ad iso , P o esso Jo ge Mo ais Mendes, o being a pa o his
p ojec since i s beginning, suppo ing me in all he challenges ha I ha e aced. As well, my co-ad iso ,
P o esso Nuno An ónio, o helping me go u he wi h his ema kable ision and insigh s. No only by
all he pa ience and suppo hey ha e p o ided, bu also, by all he knowledge hey ha e passed du ing
he lec u es in he i s yea o he Mas e ’s deg ee.
Mo eo e , I wan o exp ess my g a i ude o my in e nship supe iso , João Ped o Oli ei a, o us ing
me wi h his p ojec and o all he suppo , knowledge, and oppo uni ies he has gi en me o hese
mon hs. Also, And é Pousinho and Joana Neg a, o all he expe ience hey ha e sha ed wi h me in he
i s pa o my in e nship, allowing me o unde s and he insu ance ma ke and, pa icula ly, he
Wo ke s’ Compensa ion b anch.
I would also like o hank G upo Ageas Po ugal o allowing me o join hem o one yea , and e e yone
who has sha ed wi h me his pa h. Fu he mo e, o all he non-li e b anch P icing and Business
Analy ics eam who ha e made me eel so welcome and we e always willing o help in any hing I would
need.
Las ly, I would like o hank my iends and amily o belie ing in me and being a pa o my academic
jou ney and li e.
i
ABSTRACT
Cus ome chu n has been inc easing in insu ance, mainly due o echnological imp o emen s ha
allow cus ome s o explo e o he insu ance p o ide s’ o e s. Gi en his, insu ance p o ide s need o
compe e among hem, no only o ge new cus ome s bu , o main ain hei own.
This epo esul s om a p ojec de eloped du ing an in e nship in G upo Ageas Po ugal, which has
di e en insu ance b ands such as Ociden al Segu os. This p ojec 's main goal was o model, wi h a
mon hly pe iodici y, cus ome chu n o his la e ’s Wo ke s’ Compensa ion po olio o imp o e he
company’s compe i i eness and, ul ima ely, p o i .
Many o he company’s cus ome chu n happens a hei policy enewal ime, whe e he only a iable
ha he company de ains con ol o e is he p ice (p emium) a ia ion. Hence, by conside ing he
p emium a ia ion and o he ele an p edic i e a iables, he goal was o p edic he p obabili y o a
gi en cus ome o chu n, allowing he company o op imize he cu en enewal’s p icing p ocess and
maximize his b anch’s p o i .
Thus, di e en a iables ha could in luence he company’s cus ome beha io we e collec ed—one
o hose was he cus ome ’s loca ion. Gi en he high dimensionali y ha such a iable would ep esen
and he small da ase a ailable o modeling, clus e ing analysis is used o c ea e new signi ican (wi h
ewe dimensions) cus ome geog aphical a eas.
Di e en supe ised lea ning algo i hms we e hen e alua ed acco dingly o hei pe o mance in
p edic ing cus ome chu n. The p edic i e models used we e a G adien Boos ing, an Ex eme G adien
Boos ing, a Logis ic Reg ession, and a Mul ilaye Pe cep on. Gi en ha he numbe o cus ome s ha
enew hei con ac s is much supe io o he numbe o cus ome s who chu n, Syn he ic Mino i y
O e sampling Technique (SMOTE) was used o c ea e less unbalanced da ase s (wi h syn he ic
samples) and e alua e he impac on he pe o mance o one o he models.
Las ly, o gua an ee a success ul in eg a ion o he models in o he enewal’s p icing p ocess, models
we e e alua ed acco dingly o he wo business goals. Fi s , by ansla ing he obse ed e alua ion
me ics in o p o i . Secondly, by assu ing ha he cus ome ’s p ice elas ici y would be cap u ed,
assu ing a mono onic inc easing ela ionship among he policy’s p emium a ia ion and p obabili y o
chu n.
KEYWORDS
Supe ised Lea ning; Classi ica ion; Cus ome Chu n P edic ion; Non-Li e Insu ance; Renewal P ice
Elas ici y; Clus e ing; Neu al Ne wo k; Logis ic Reg ession; S ochas ic G adien Boos ing; Ex eme
G adien Boos ing
INDEX
1. In oduc ion .................................................................................................................. 1
1.1. P oblem S a emen and Objec i e ........................................................................ 1
1.2. Summa y o he P ocess ........................................................................................ 2
2. Theo e ical F amewo k ................................................................................................ 3
2.1. Li e a u e Re iew in Chu n Modeling in Insu ance ............................................... 3
2.2. Machine Lea ning .................................................................................................. 5
2.2.1. Da a Imbalance ............................................................................................... 5
2.2.2. Fea u e Selec ion ............................................................................................ 8
3. Da a............................................................................................................................. 11
4. Me hodology .............................................................................................................. 13
4.1. Business Unde s anding ...................................................................................... 13
4.2. Da a Unde s anding ............................................................................................ 14
4.3. Da a P epa a ion ................................................................................................. 14
4.3.1. Da a Selec ion ............................................................................................... 14
4.3.2. Da a Cleansing .............................................................................................. 15
4.3.3. Fea u e Enginee ing ..................................................................................... 16
4.3.4. Da a In eg a ion and Fo ma ........................................................................ 19
4.4. Modeling .............................................................................................................. 20
4.4.1. Model Selec ion ............................................................................................ 20
4.4.2. Tes Design Gene a ion ................................................................................ 21
4.4.3. Model Cons uc ion and Assessmen ........................................................... 22
4.5. E alua ion ............................................................................................................ 27
4.6. Deploymen ......................................................................................................... 29
5. Resul s and Discussion ................................................................................................ 31
5.1. Resul s ................................................................................................................. 31
5.1.1. C oss-Valida ion Resul s ............................................................................... 31
5.1.2. Model Assessmen ....................................................................................... 31
5.1.3. Model’s E alua ion ....................................................................................... 33
5.2. Discussion ............................................................................................................ 36
6. Conclusions ................................................................................................................. 38
7. Limi a ions and Recommenda ions o Fu u e Wo ks ............................................... 40
8. Bibliog aphy ................................................................................................................ 41
9. Appendix ..................................................................................................................... 47
i
9.1. Appendix A - Backwa d Elimina ion (Logis ic Reg ession Final Fea u e Lis ) ...... 47
9.2. Appendix B - Cons uc ed Models’ Cha ac e is ics ............................................. 48
9.3. Appendix C - Final Model’s Fea u e Lis .............................................................. 49
4
Model
Used in Li e a u e
Logis ic Reg ession
Y. He e al. (2020)
Va eiadis e al. (2015)
Sunda kuma & Ra i (2015)
Bolancé e al. (2016)
Spi e i & Azzopa di (2018)
Decision T ee
Va eiadis e al. (2015)
Sunda kuma & Ra i, 2015)
Bolancé e al. (2016)
Dola abadi e al. (2017)
Spi e i & Azzopa di (2018)
Sc iney e al. (2020)
Suppo Vec o Machine
Y. He e al., (2020)
Va eiadis e al. (2015)
Sunda kuma & Ra i (2015)
Bolancé e al. (2016)
Dola abadi e al. (2017)
Spi e i & Azzopa di (2018)
Sc iney e al. (2020)
Neu al Ne wo k
Y. He e al. (2020)
Va eiadis e al. (2015)
Sunda kuma & Ra i (2015)
Bolancé e al. (2016)
Dola abadi e al. (2017)
Sc iney e al. (2020)
G adien Boos ing
Y. He e al. (2020)
Naï e Bayes Classi ie
Va eiadis e al. (2015)
Dola abadi e al. (2017)
Spi e i & Azzopa di (2018)
Sc iney e al. (2020)
Random Fo es
Y. He e al. (2020)
Spi e i & Azzopa di (2018)
Ex a T ees Classi ie
Y. He e al. (2020)
Table 2.1 Models used in li e a u e o insu ance chu n modeling

5
2.2. MACHINE LEARNING
Unsupe ised Lea ning
Unsupe ised lea ning is a o m o machine lea ning ha aims o g oup da a in o segmen s based on
simila a ibu es, o na u ally occu ing ends, pa e ns, and ela ionships hidden in he da a (McCue,
2015). S anda d algo i hms used o his kind o ask a e clus e ing, anomaly de ec ion, neu al
ne wo ks, and app oaches o lea ning la en a iable models (El Bouche y & De Souza, 2020).
The main goal o clus e ing analysis is o segmen he ini e unlabeled da ase in o a ini e and disc e e
se o hidden da a s uc u es (Xu & Wunsch, 2005). The esul o his analysis is se e al g oups o he
da a named clus e s. These a e subse s o da a g ouped when, acco ding o he s udied c i e ia, ha e
simila cha ac e is ics and, when no , sepa a ed in o di e en g oups (Rokach & Maimon, 2005).
Mo eo e , wo ypes o clus e ing echniques a e pa i ional clus e ing and hie a chical clus e ing. In
he i s , g oups a e c ea ed by pa i ioning he space in o a p e-de ined numbe o subspaces. On he
o he hand, in hie a chical clus e ing, he da a objec s a e g ouped in sequence by a hie a chical
s uc u e.
Supe ised Lea ning
Ano he o m o Machine Lea ning is supe ised lea ning. Con a y o unsupe ised lea ning, i uses a
da ase ha has al eady been classi ied (labeled) as a basis o p edic ing he classi ica ion o o he
unlabeled da a (Talabis e al., 2015). The labeled da ase is a aining se composed o inpu a iables
( ea u es) and an ou pu a iable (label). The used ea u es will in luence he model’s abili y o
co ec ly classi y he p edic ed a iable, he ou pu a iable (L. Wang e al., 2021).
Supe ised Lea ning can be di ided in o classi ica ion when he label is ca ego ical and eg ession
when he label is con inuous. The e o e, he s udied p oblem is a classi ica ion ask since chu n
modeling can be ansla ed o a bina y classi ica ion p oblem: assuming 1 when he cus ome chu ns
and 0 when he cus ome does no chu n.
2.2.1. Da a Imbalance
Fo many eal supe ised lea ning p oblems in ol ing a bina y esponse a iable, da ase s p esen a
skewed dis ibu ion, ha ing one class wi h a much lowe ep esen a ion han ano he . When he
da ase shows his ype o beha io , ha ing an unde ep esen ed class, he da a is said o be
unbalanced (S. Wang & X. Yao, 2012). Resea che s ha e concluded ha his imbalance causes a
subop imal classi ica ion pe o mance (Chawla e al., 2004), since classi ie s end o gi e much highe
impo ance o he la ge classes.
H. He & Ga cia (2009) b ing ha . Usually, classi ie s, when dealing wi h an imbalanced da ase , end
o “p o ide a se e ely imbalanced deg ee o accu acy, wi h he majo i y class ha ing close o 100
pe cen accu acy and he mino i y class ha ing accu acies o 0-10 pe cen ”, such consequence can
ep esen high cos s o some indus ies so, he e o e, is i al o cons uc a model ha “will p o ide
high accu acy o he mino i y class wi hou se e ely jeopa dizing he accu acy o he majo i y class”.
In o de o deal wi h imbalanced da ase s, h ee possible app oaches can be aken: da a le el,
algo i hmic le el, and combining o ensemble me hods, o which he i s in ol es esampling o
educe he class skewness (Yap e al., 2014). Resampling is done ei he by emo ing ins ances om he
6
majo i y class (unde sampling) o adding ins ances o he mino i y class (o e sampling) by using
algo i hms such as SMOTE.
Syn he ic Mino i y O e sampling Technique
Syn he ic Mino i y O e sampling Technique (SMOTE) (Chawla e al., 2002) is one o he mos popula
and in luen ial da a p e-p ocessing algo i hms o deal wi h he da a imbalance p oblem (Ga cía e al.,
2016).
This echnique is an o e sampling app oach, said ha new ins ances om he smalle class a e
in oduced in o he da ase . Unlike basic app oaches such as andom o e sampling (ROS), which only
duplica es samples om he mino i y class, SMOTE gene a es new syn he ic samples, o e coming he
o e i ing caused by app oaches like ROS (Fe nández e al., 2018).
The i s s ep o his echnique is o de ine he amoun o o e sampling. He e, i is possible o ei he
se up his alue o app oxima e a balanced class dis ibu ion o disco e i ia a w appe p ocess
(Chawla e al., 2008). Then, based on k nea es neighbo s and linea in e pola ion ideas, he syn he ic
samples a e c ea ed. SMOTE ope a es in he ea u e space a he han in he da a space, each mino i y
class sample is conside ed along wi h i s k nea es neighbo s, and he new samples a e in oduced
along he line segmen s joining hem (conside ing any/all o he k neighbo s) (Chawla e al., 2002).
2.2.1.1. Model Calib a ion
When using such echniques o a i icially ebalancing he da ase o e en by consequence o he
da ase ’s cha ac e is ics, he aining and es se s ha e di e en dis ibu ions. This di e ence in
aining and es se s dis ibu ion iola es he basic assump ion in machine lea ning ha bo h a e
d awn om he same unde lying dis ibu ion (Pozzolo e al., 2015). By iola ing his assump ion, he
p edic ions ob ained in he es se will be biased and, he e o e, enhance he need o p obabili y
calib a ion o ob ain unbiased p edic ions.
Fu he mo e, some me hods end o bias p edic ed p obabili ies by pushing away o close o 0 and 1,
enhancing he need o calib a ion (Niculescu-Mizil & Ca uana, 2005). Do mann (2020) e en s a es
ha “i should be applied o any model ype as pa o he p edic ion p ocess, be o e p edic ing, c oss-
alida ing and making e ec plo s and maps o using p edic ions in any o he p obabilis ic
in e p e a ion”.
Two model calib a ion me hods ha can be used o co ec hese biased p obabili ies a e Pla Scaling
and Iso onic Reg ession. The i s is mo e e ec i e when he dis o ion is sigmoid-shaped, and he
la e is a mo e obus me hod ha can co ec any mono onic dis o ion bu , mo e p one o
o e i ing (Niculescu-Mizil & Ca uana, 2005).
The Pla Calib a ion calib a es p obabili ies by passing he ou pu h ough a sigmoid:
𝑃(𝑦=1|𝑓)= 1
1+𝑒(𝐴𝑓+𝐵)
( 1 )
Whe e 𝑓(𝑥) is he lea ning me hod and pa ame e s 𝐴 and 𝐵 a e es ima ed using G adien Descenden ,
such ha hey a e a solu ion o he ollowing minimiza ion unc ion:
7
𝑎𝑟𝑔𝑚𝑖𝑛𝐴,𝐵{− ∑𝑦𝑖 log(𝑝𝑖)+(1−𝑦𝑖)log(1−𝑝𝑖)
𝑖}
( 2 )
whe e,
𝑝𝑖= 1
1+𝑒(𝐴𝑓𝑖+𝐵)
( 3 )
On he o he hand, he Iso onic Calib a ion is mo e gene al gi en ha he only es ic ion is ha he
mapping unc ion is iso onic (Niculescu-Mizil & Ca uana, 2005). The basic assump ion o Iso onic
Reg ession (on which he model is based) is ha :
𝑦𝑖=𝑚(𝑓𝑖)+𝜖𝑖
( 4 )
Whe e, 𝑦𝑖 a e he ue labels, 𝑓𝑖 he model’s p edic ions, 𝑚 a mono onic inc easing (iso onic) unc ion.
The goal is o ind 𝑚, using he ue labels and model’s p edic ions as a aining se , such ha :
𝑚=𝑎𝑟𝑔𝑚𝑖𝑛𝑧∑(𝑦𝑖−𝑧(𝑓𝑖))2
( 5 )
One algo i hm ha can be used o ind a s epwise cons an solu ion o his p oblem is he pai -
adjacen iola o s (PAV) algo i hm (Aye e al., 1955).
Las ly, one way o assess how well-calib a ed he model is can be h ough a calib a ion plo . He e, “a
se o p edic ions o a bina y ou come is well calib a ed i he ou comes p edic ed o occu wi h
p obabili y p do occu abou p ac ion o he ime, o each p obabili y p ha is p edic ed” (Naeini e
al., 2015), which can be ansla ed in o a s aigh line om (0,0) o (1,1). Gi en his, in he calib a ion
plo , he x-axis ep esen s he a e age p edic ed p obabili y in each bin. The y-axis ep esen s he
obse ed ac ion o samples in he bin whose eal labels a e posi i e. The obse ed cu e is hen
compa ed wi h he s aigh line 𝑦=𝑥.
2.2.1.2. Th eshold-Mo ing Me hod
A echnique ha should be conside ed when dealing wi h class imbalance is changing he decision
h eshold (model’s con inuous ou pu cu -o ) and adap ing i o a pe o mance me ic. The main
di e ence be ween ebalancing (using echniques like SMOTE) and h eshold-based me hods is ha
he la e elies on manipula ing he con inuous ou pu o a lea ned model ins ead o elying on da a
p e-p ocessing be o e he lea ning happens (Collell e al., 2018).
P o os (2008) e en s a es ha “The bo om line is ha when s udying p oblems wi h imbalanced da a,
using he classi ie s p oduced by s anda d machine lea ning algo i hms wi hou adjus ing he ou pu
h eshold may well be a c i ical mis ake”.
The h eshold mo ing me hod uses he o iginal aining se o ain and unes o shi s he decision
h eshold by adap ing i o a pe o mance me ic. One possible app oach is o use he ROC e alua ion
p ocedu e and mo e om whe e misclassi ica ions a ain hei maximum on he posi i e class o he
poin whe e he maximum in he nega i e class is a ained, selec ing he poin whe e he cu e a ains
i s maximum (H. He & Ga cia, 2009).
8
2.2.2. Fea u e Selec ion
A well-known p oblem is he “cu se o ini e sample size”, o which he ela ionship among he
numbe o samples a ailable and he ea u es conside ed o modeling needs o be conside ed (Jain &
Chand aseka an, 1982). Each new ea u e in oduced o he model will ep esen a new dimension.
The highe he dimensionali y, he spa se he da ase becomes and, hus, lowe he ea u e space
co e age (Ve leysen & F ançois, 2005). Consequen ly, he p oblem’s complexi y apidly g ows wi h he
in oduc ion o mo e dimensions. Bellman (1966) in oduced he e m Cu se o Dimensionali y o
explain such phenomena.
The e o e, and gi en he nowadays exis ing high-dimensional da a, ea u e selec ion is one o he
essen ial echniques in da a p ep ocessing by elimina ing i ele an , edundan , o noisy ea u es
(Kalousis e al., 2007). Pe o ming such echniques allows as e algo i hms and, besides imp o ing
p edic i e powe , also imp o es comp ehensibili y (Kuma & Minz, 2014).
I is possible o b oadly classi y ea u e selec ion me hods in o il e and w appe me hods, whe e he
i s anks ea u es based on s a is ical measu es, independen ly o he lea ning algo i hm (Koha i &
John, 1997). One example o his ype o me hod is o use as impo ance measu e (sco e) he a iable’s
co ela ion wi h he a ge . On he o he hand, w appe s e alua e each candida e subse o ea u es'
impac on a pa icula lea ning algo i hm. The la e app oach usually allows achie ing be e esul s
gi en hei close in e ac ion wi h he classi ie (El Aboudi & Benhlima, 2016).
Some examples o w appe me hods a e o wa d selec ion, backwa d elimina ion, and ecu si e
ea u e elimina ion. The i s keeps adding new ea u es which imp o e he model pe o mance un il
no o he ea u e espec s he c i e ia. The second wo ks simila ly bu opposi ely, s a s wi h all
ea u es, and emo es he leas signi ican e ec ha does no mee he model’s s aying c i e ia un il
all ea u es a e signi ican . In bo h ( o wa d and backwa d elimina ion), he ea u e s ays in he model
once added/ emo ed. Las ly, ecu si e ea u e elimina ion is simila o he o wa d selec ion me hod.
In his case, e ec s a e added and emo ed in o he model such ha one o mo e backwa d
elimina ion s eps can happen a e a o wa d selec ion s ep (Bu sac e al., 2008).
Ano he amily de i ed om he wo p e ious me hod amilies ( il e and w appe s) a e embedded
me hods. These me hods combine he classi ie de elopmen wi h he sea ch o he op imal subse o
ea u es, cap u ing dependencies a a lowe compu a ional cos han w appe s bu (like w appe s) also
ha e a isk o o e i ing (Seijo-Pa do e al., 2017). Examples o embedded me hods a e Lasso and
Ridge eg ession, which ha e buil -in ea u e selec ion me hods ha employ L1 and L2 egula iza ion
( espec i ely).
2.2.2.1. Ensemble Lea ning o Fea u e Selec ion
Ensemble Lea ning is a ype o lea ning whe e mul iple models a e ained and combined o sol e he
same p oblem (Polika , 2006). Ensemble Lea ning is based on he assump ion ha combining he
solu ion o mul iple expe s is be e han using he solu ion o a single one. In his way, a se o
hypo heses is cons uc ed a he han using only one single hypo hesis o explain he da a, making i
possible o educe bias and a iance om he lea ning algo i hms (Die e ich, 2002).
Al hough usually employed o imp o e classi ica ion esul s, i is also possible o use ensemble lea ning
as a ea u e selec ion echnique. Combining mul iple ea u e selec ion me hods (ins ead o elying on
9
jus one) makes i possible o a ain mo e obus ea u e subse s, showing a g ea p omise o high-
dimensional da ase s wi h small sample sizes (Saeys e al., 2008).
The app oach bene i s om di e si y and con ol o a iance and is possible o employ in wo ways:
one is o use he same algo i hm o e ie e he ea u es’ impo ance using di e en subse s o he
da a (da a pe u ba ion / homogeneous) and, o he , is o use di e en ea u e selec ion echniques in
he same da ase ( unc ion pe u ba ion / he e ogenous) (Chiew e al., 2019).
Said his, di e en le els can be a ied and shall be chosen when employing ensemble lea ning in
ea u e selec ion, ollowing Bolón-Canedo & Alonso-Be anzos (2019) can be de ined as ollows:
• Da ase Le el: use di e en subse s o da a (o no )
• Fea u e Le el: use di e en subse s o ea u es (o no )
• Lea ne Me hod Le el: use o design di e en lea ning algo i hms (o no )
• Combina ion Le el: Use o design di e en combina ion/agg ega ion me hods
• Th eshold Le el: use o di e en h esholding me hods (in case o using anke me hods)
Bolón-Canedo & Alonso-Be anzos (2019) also e i ied ha he e ogeneous ea u e selec ion ensembles
a e mo e commonly used han homogeneous ones. Howe e , i is possible o ob ain ei he a ea u e
subse o a ea u e anking in bo h cases, depending on he ype o ea u e selec o s. Fo he la e , a
h eshold me hod needs o be de ined.
When using ea u e anke s, se e al ea u e anking algo i hms o he ensemble a e combined,
c ea ing a inal anked lis o he ea u es, gi en he ea u es’ ele ance o p edic ion. The app oach
o combining he ensemble membe s' esul s ( anks) has di e en p oposals in he li e a u e, om
simple o mo e complex solu ions (Seijo-Pa do e al., 2015). Some o he mos popula s aigh o wa d
me hods o combine such anks a ibu ed o each ea u e a e: minimum (bes ) ank, median ank,
a i hme ic mean ank, and geome ic mean (Bolón-Canedo & Alonso-Be anzos, 2019).
Gi en he h eshold me hod decision, he mos common app oach is de ining a ixed pe cen age o
he op ea u es, bu his pe cen age depends on he used da ase (Bolón-Canedo & Alonso-Be anzos,
2019). The e o e, his echnique is no op imal since i is p one o o e s a ing o unde s a ing his cu -
o alue (Chiew e al., 2019).
Pe mu a ion ea u e impo ance
Pe mu a ion ea u e impo ance (PFI) is a model-agnos ic ea u e selec ion me hod. The e o e, i is
possible o pe o m a he e ogeneous app oach by using his algo i hm as a base lea ne o he
ensemble and in oduce a iabili y by using di e en base models in he algo i hm.
This pe mu a ion ea u e impo ance measu emen was i s ly in oduced o andom o es s in 2001
(B eiman, 2001) bu can be used in any model. PFI assesses he a iable’s impo ance o he gi en
model when i s ela ionship wi h he a ge is b oken by obse ing he dec ease in he model’s sco e
when andom noise eplaces a a iable (in oduced by andomly shu ling he a iable’s alues)
(McGo e n e al., 2019). The e o e, i is possible o unde s and how impo an a gi en ea u e is o he
model’s abili y o p edic he a ge co ec ly.

10
PFI’s ope a ion me hod makes i sensible o co ela ed ea u es (S obl e al., 2008) and is, he e o e,
essen ial o add ess his issue p ima ily. Howe e , PFI has ad an ages such as obus ness no o bias
he measu es a o ing high ca dinally ea u es o e bina y ea u es.
11
3. DATA
Gi en he small dimensions o he Wo ke s’ Compensa ion po olio, he main goal in he da a
collec ion phase was o ex ac as much da a as possible. Mo eo e , o gua an ee ha he collec ed
in o ma ion uly cap u es he eali y, a 3-mon h aging is equi ed.
The e o e, he da a ex ac ion p ocess was di ided in o wo phases. A i s phase, a he p ojec ’s
beginning, in which i was possible o ex ac da a ega ding he enewals and chu ns om Janua y
2017 un il Sep embe 2020, used o he models’ cons uc ion. Then, a second phase, du ing he
Model’s Assessmen phase, o be used as a es se . This new da ase con ained da a om Oc obe
2020 un il Feb ua y 2021, including Janua y, whe e mos o he policies enew hei con ac s. Bellow,
he Da a Ex ac ion phases can be seen in Figu e 3.1.
Gi en he p oblem con ex , he goal was o iden i y olun a y chu ne s who based hei decision on
p ice a ia ion. In his way, nei he in olun a y chu ns no cancela ions ou side he de ined enewal
window (explained in sec ion 4.1) a e included in he analysis.
Mo eo e , o gua an ee ha he in o ma ion ex ac ed would cap u e he e en ha was decided o
be s udied - olun a y chu ne s ha ha e chu ned by consequence o he p emium a ia ion - some
es ic ions ha e been applied o he da a:
• No include chu ns ha a e a consequence o , o example, bank up o he company.
• No include policies om employees o he company
• No include empo a y policies
• No include he g oup o policies lagged by he company ha ecei e di e en enewal’s
p icing p ocesses om emaining cus ome s (cus ome s wi h highe o al p emiums and
policies iden i ied o go h ough a p uning p ocess)
• No include policies ha cancel o issue hei exi be o e he de ined enewal window
• No include policies o which i was no possible o ex ac hei p emium alue
• No include annui ies wi h a du a ion o less han one yea
Policies ha ha e canceled hei con ac s o acqui e a new policy in he company (policy
cannibaliza ion) we e s ill conside ed o analysis, since i is possible ha he cus ome has cancelled
due o he p ice inc ease. Usually, cus ome s acqui e a new policy o ob ain a lowe p ice main aining
he same bene i s.
Figu e 3.1 Di e en da a ex ac ion phases
12
Wi h such cons ain s applied, i was possible o ex ac a da ase wi h he policies' annui y in o ma ion
be o e he enewal da e in which, o one yea , a pa icula policy only appea s once. Rega ding he
collec ed a iables, his decision was made conside ing he in o ma ion p esen in he li e a u e and
sugges ions made by he p ojec ’s s akeholde s. The collec ed a iables and hei desc ip ion we e no
included in his epo . Howe e , i is possible o esume he a iable’s in o ma ion in o he ollowing
ca ego ies:
• Ca ego y 1: Value paid o insu ance
• Ca ego y 2: Cus ome beha io
• Ca ego y 3: Cus ome awa eness o he inc ease
• Ca ego y 4: Cus ome ’s/P oduc ’s cha ac e is ics
• Ca ego y 6: Cus ome ’s claims/cos and i s managemen by he company
• Ca ego y 6: Posi ion o Ociden al Segu os in Ma ke
• Ca ego y 7: Cus ome ’s in e ac ion channel wi h he company
• Ca ego y 8: Cus ome ’s loyal y
• Ca ego y 9: Cus ome ’s geog aphical loca ion cha ac e is ics
• Ca ego y 10: Pandemic si ua ion
• Ca ego y 11: Discoun s
Gi en ha he model needs o deli e p edic ions 60 days be o e he policy’s enewal da e (also
explained in sec ion 4.1), all he a iables collec ed had o conside his ision on ime. Mos a iables
main ain cons an along wi h he annui y. Howe e , o he s, such as he employees’ o al sala y
associa ed wi h he con ac , usually a y du ing his pe iod. The e o e, he ex ac ed a iables we e
designed o ga he he co ec ision, wi h he end o annui y s anding o he 60 days be o e i .
As men ioned be o e, he mos c ucial a iable o measu ing i s impac on chu n is he p emium
a ia ion ob ained in he enewal’s p icing p ocess. Howe e , his p emium a ia ion also depends on
he con ac ’s employees’ o al sala y, which may a y om one annui y o ano he o e en in he 60
days window. In his way, i was decided ha he p emium a ia ion should only be measu ed,
excluding he a ia ion o his a iable. The e o e, only he con ac ’s a i a ia ion was conside ed
when measu ing he p emium a ia ion.
13
4. METHODOLOGY
4.1. BUSINESS UNDERSTANDING
As men ioned be o e, his p ojec is applied o he Wo ke s’ Compensa ion po olio om Ociden al
Segu os. To gua an ee he p ojec ’s success, he i s objec i e was o unde s and his business line
and he p ojec objec i es. Fu he , his knowledge has been con e ed in o da a mining goals, and a
p ojec plan has been designed.
Wo ke s’ Compensa ion insu ance has been manda o y in Po ugal since 1993 o Thi d-Pa y
companies’ wo ke s and ex ended o sel -employed wo ke s in 1997. This ype o insu ance has only
one co e age ha co e s he isk o acciden s in he employees’ wo kplace o on hei way om/ o
home. This co e age assu es he legally needed bene i s by consequence o any o he in ol ed
employees’ acciden s, co e ing all needed expenses o ensu e he wo ke ’s o al eco e y (o
compensa ion in case o disabili y o dea h). The paid p emium o his insu ance depends no only on
he o al secu ed employees’ sala y (sum insu ed), which may a y du ing he yea bu also on he
con ac ’s a e s ipula ed by he company (which can a y yea ly in he enewal’s p icing s a egy).
The enewal’s p icing p ocess p oceeds app oxima ely wo mon hs in ad ance (gi en he cus ome ’s
enewal da e). By law, he insu ance company mus send he enewal le e in o ming he cus ome
abou he nex annui y’s p emium 30 days in ad ance. Gi en his, he ma ke de ined ha his enewal
le e should be sen 45 days in ad ance, so his is (usually) when he cus ome is awa e o he a ia ion
in he p emium (con ac ’s a e). The e o e, based on his business knowledge, i was de ined ha only
a cancella ion ha happens in a window o 45 days p io and 50 days a e he end o he policy’s
enewal da e would be classi ied as chu n. Mo eo e , he annui y’s p emium a ia ion is de e mined
60 o 45 days be o e he end o he policy’s enewal da e, so all a iables ex ac ed espec his ime
window (using he a iable’s ision 60 days be o e he policy annui y end da e).
Ociden al Segu os has a bancassu ance channel and de ains 1.5% o he Wo ke s’ Compensa ion
insu ance ma ke sha e. The company’s po olio can be segmen ed in o h ee main cus ome g oups:
• Housekeepe s – Besides housekeepe s, which is he main componen o his g oup, o he
domes ic employees a e in eg a ed, as ga dene s, o example.
• Sel -Employees – In gene al, a e small companies whe e he insu ed en i y is he wo ke
himsel .
• Thi d-Pa y - Comp ises companies om di e se dimensions (mic o/small, medium, and big)
insu ing hei employees.
Al hough mos o Ociden al Segu os’ cus ome con ac s a e om he Housekeepe s’ segmen , Thi d-
Pa y ep esen s almos 85% o he po olio’s o al annual p emium. Such beha io is explained by
he ac ha , on a e age, he o al secu ed employees’ sala y o his segmen is much highe han o
he o he segmen s.
Rega ding claims, he mos signi ican equency o claims is obse ed o he Sel -Employees’
segmen . Howe e , was he Thi d-Pa y segmen o which, a he ime, he lowes p o i abili y was
obse ed.
20
Yeo-Johnson ans o ma ion is a powe ans o ma ion amily wi h simila p ope ies as Box-Cox
ans o ma ion, bu well de ined in he whole eal line, hus app op ia e o educe skewness and
app oxima e no mali y wi hou such limi a ion (Yeo, 2000). Box-Cox ep esen s a amily o powe
ans o ma ions (like squa e oo , log, o in e se ans o ma ions - which a e conside ed a way o mee
he no mali y assump ion) ha easily ind he op imal powe ans o ma ion o a gi en a iable o
s abilize i s a iance (T. Zhang & Yang, 2017). Ne e heless, he o iginal Box-Cox (Box & Cox, 1964)
ans o ma ion is only alid o a posi i e x, con a y o Yeo-Johnson ans o ma ion.
On he o he hand, i has also been shown ha Fea u e Scaling gene ally imp o es he pe o mance
o classi ica ion algo i hms (Bollegala, 2017). Wha gene ally leads o such imp o emen is ha
ea u es’ alues, na u ally, occupy di e en anges (some anging in housands, o he s in ens) and,
ea u es wi h highe anges end o ha e a mo e decisi e ole while aining he model. Fu he mo e,
he ea u e’s ela i e alue di e ence is o en mo e in o ma i e han i s absolu e alue (Bollegala,
2017). The e o e, o he han using he da a in i s o iginal ange o m, some scaling me hods we e
a emp ed. Such ans o ma ions we e pe o med by es ima ing he scaling pa ame e s using he
aining se and, he ea u e scaling me hod is hen applied in bo h aining and es se s.
4.4. MODELING
In his phase, he used modeling echniques a e selec ed. Also, a es design is c ea ed o build he
di e en models and co ec ly assess hem, unde s anding hei alue o esol ing he p oblem a
hand.
4.4.1. Model Selec ion
Ini ially, o his p ojec , ou p edic i e models we e selec ed and cons uc ed: a Logis ic Reg ession,
a Mul ilaye Pe cep on, a G adien Boos ing model, and an Ex eme G adien Boos ing model. These
me hods a e b ie ly explained below.
Logis ic Reg ession
Simila o linea eg ession, logis ic eg ession (LR) can include one o mul iple co a ia es ( a iables)
ha , join ly wi h unknown pa ame e s es ima ed om he da a, p oduce a linea (and con inuous)
p edic o . The logi unc ion is used o gua an ee ha he ou come a iable alls in a 0-1 ange.
Fu he mo e, i is possible o add a egula iza ion e m o he logis ic eg ession, encou aging he i ed
pa ame e s o be small and helping p e en o e i ing.
Mul ilaye Pe cep on
The Mul ilaye Pe cep on (MLP) is a eed- o wa d ne wo k composed o inpu neu ons, ou pu
neu ons, and hidden laye s o neu ons be ween hem. Neu al ne wo ks a e a b anch o a i icial
in elligence inspi ed in human b ains. He e, nume ous cells called neu ons p ocess in o ma ion in
pa allel, linked oge he in a ne wo k by synapses, whe e in elligence is a gued o be encoded.
Mo eo e , in a eed- o wa d ne wo k, in o ma ion only mo es o wa d, om he inpu nodes o he
hidden nodes and inally ou pu nodes (Zell, 1994), connec ed by weigh s and ou pu signals. A
nonlinea ans e unc ion modi ies he neu on’s weigh ed inpu s, also called he ac i a ion unc ion
(Njikam & Zhao, 2016). These supe posi ions o nonlinea ans e unc ions allow he mul ilaye
pe cep on o app oxima e highly non-linea unc ions (con a y o he logis ic eg ession) and
accu a ely gene alize when p esen ed wi h new, unseen da a (Ga dne & Do ling, 1998).

21
G adien Boos ing Model
The G adien Boos ing (GB), p oposed by F iedman (2001, 2002), belongs o he amily o boos ing
me hods, a ype o ensemble lea ning. In boos ing, a new model is added o he ensemble sequence
ained based on he e o o he whole ensemble lea ned so a (Na ekin & Knoll, 2013). Hence,
misclassi ied ins ances a e emphasized by ecei ing highe weigh s in he nex i e a ion, which, o
example, usually happens o ins ances nea he decision bounda y (Mei & Rä sch, 2003). These
weigh s ep esen he impo ance ha he gi en ins ance will ha e in he ollowing base lea ne
cons uc ion. The e o e, ins ances ha he p e ious base models ha e shown di icul y unde s anding
appea mo e o en in he aining da a (Y. Zhang & Haghani, 2015). The inal model ob ained by he
boos ing algo i hm will be a linea combina ion o he se e al base-lea ne s conside ing hei
pe o mance on he da ase (G. Wang e al., 2011).
Ex eme G adien Boos ing Model
The Ex eme G adien Boos ing model (XGBoos ) is an ensemble o Classi ica ion T ee (CART) and a
mo e e icien e sion o GB. This algo i hm is e y popula in he Machine Lea ning ield, ha ing i s
impac widely ecognized in many machine lea ning and da a mining challenges (Chen & Gues in,
2016). Some o he modi ica ions done o GB a e pa allel aining (which as ens he algo i hm), ou -
o -co e compu a ion (allowing ha da a is no loaded in o memo y), and spa se da a op imiza ion (in
handling and speed up compu a ion) (Bisong, 2019). Also, besides sh inkage and subsampling (used in
GB) as egula iza ion o ms o con ol o e i ing and a ain be e esul s, he XGBoos uses a
egula ized objec i e. Mo e in o ma ion abou he XGBoos algo i hm can be ound a Chen & Gues in
(2016).
4.4.2. Tes Design Gene a ion
In o de o choose he sui able model and model’s hype pa ame e s, each cons uc ed model should
be es ed and, he e o e, a es design needs o be de ined. The goal is ha he ob ained esul s while
aining/e alua ing he model a e close o he eali y o how each model will pe o m. The e o e, bo h
he da ase ’s cha ac e is ics and he way he model would be used we e conside ed.
Fi s ly, ega ding he da a, a possible endency in chu n a e ac oss he yea s has been obse ed o
he di e en cus ome segmen s in he Da a Unde s anding phase. Also, as men ioned be o e, he
policy da a main ains p ima ily s able h ough he yea s in mos measu es. Howe e , signi ican
changes usually happen in he policies’ p emium paid (and o he a iables i depends on). Mo eo e ,
gi en ha each policy and i s da a only appea , a he mos , once in each o he yea s, he e is an
annual ( empo al) dependency on he da a.
A possible app oach o espec such empo al dependency is o use ime-spli c oss- alida ion. He e,
he model’s gene aliza ion abili y is assessed using he a e age o he pe o mance me ics in he
c ea ed da a spli s. A he same ime, he empo al dependency is espec ed by always using he las
block o da a as alida ion. Gi en he applied inc ease in p emium, he model should p edic he
p obabili y o chu n o he cus ome s expec ed o enew each mon h in he yea and he e o e deli e
mon hly p edic ions. Gi en hese model pu poses, wel e spli s we e c ea ed (one spli o each mon h
in he yea ), using he alida ion se o he spli ’s las a ailable mon h, as shown in Figu e 4.4.
22
Figu e 4.4 Rep esen a ion o he designed ime-spli c oss- alida ion
4.4.3. Model Cons uc ion and Assessmen
4.4.3.1. Fea u e Selec ion wi h Ensemble Lea ning
As shown, one possible way o pe o m ea u e selec ion is o employ ensemble lea ning, and, o his,
he app oach needs o be designed. Gi en he popula i y o he e ogeneous ea u e selec ion me hods,
i was decided ha his app oach should be used, using he same subse s o da a and ea u es ac oss
he di e en lea ne s. Fo his, many o he sugges ions p esen ed by Shah & Pe e ia ko (2021) we e
conside ed and a e desc ibed below.
Fi s ly, as o he lea ne me hod le el, many ea u e anke s we e cons uc ed using he PFI algo i hm
wi h a di e en base model o gua an ee a iabili y in he solu ions. The selec ed models we e he
bes se o base-line models (classi ie s wi hou pa ame e unning) in e ms o ecall ( ue posi i e
a e). Said ha , he selec ed models we e he ones ha appea ed o ha e a be e unde s anding o
he small class.
Then, as o combina ion le el, i has been shown ha se e al echniques a e possible o combine he
esul s and c ea e a inal ank, om simple o mo e complex me hods. Some o he mo e
s aigh o wa d me hods a e using he median and mean o he esul s as agg ega ion ules. The
median o he anks was used o his wo k since i is less sensi i e o ou lie s, being a mo e obus
me ic han he mean. Howe e , ea u es p esen ing he same median ank we e a e wa d o de ed
by he mean o he obse ed anks.
Las ly, o he h eshold le el decision, i has been shown ha he mos common app oaches a e
conside ing a ixed h eshold o he op- anked ea u es. Howe e , such a echnique comes wi h a ew
downsizes. In his way, once ha ing each o he model’s esul s, ins ead o selec ing s aigh away he
op N ea u es (gi en he di icul y o selec ing he p ope alue o N), he used s a egy was sligh ly
di e en . A e ha ing he ea u e impac o each ea u e gi en by each o he lea ne s, he o al (sum)
ea u e impac can be calcula ed ( o each o he ea u es summing all model’s e u ned impac s),
along wi h he cumula i e impac (a e o de ing by he ea u e’s o al (sum) impac ). Then, a gi en
23
a io (𝑟∈ ]0,1[ ) o he o al cumula i e impac is applied and, he numbe o ea u es ha p esen a
cumula i e impac lowe han he a io o o al cumula i e impac will be he numbe o ea u es o
selec (N) – N will be he cu ing h eshold. Wi h his numbe (N), he Top N bes - anked ea u es a e
selec ed, acco ding o he ensemble’s lea ne s’ opinions combina ion ule ea ly de ined. I is impo an
o no e ha his echnique will ne e selec a iables a ibu ed wi h ea u e impo ance o ze o o
nega i e alues.
Howe e , using his app oach, he e is s ill a alue ha needs o be selec ed – he a io (𝑟). This a io
and i s alue decision will be explained u he . As has been concluded in he pas , he op imal ea u e
size depends no only on he ea u e-label dis ibu ion bu also on he used classi ie (Hua e al., 2005).
In his way, since each model is di e en , i was decided o c ea e a mo e model-agnos ic inal a io
decision (indi idually o each o he selec ed models).
Addi ionally, an o de ed ea u e lis was c ea ed o each baseline model based on he PFI algo i hm
esul s using he gi en model only. In his case, o each o he classi ie s is possible o choose he
ea u e selec ion me hod ha p o ides be e esul s: using ensemble lea ning o ea u e selec ion o
only he w appe me hod - bo h ha ing a common app oach o selec ing he numbe o ea u es om
he anked lis .
In he end, a ecu si e me hod was implemen ed o selec he inal ea u e lis o each model. The
o al cumula i e impac a io (𝑟) will decay un il a speci ic s opping c i e ion is eached ( he a io s a s
a 1 and decays in each in e ac ion). Fo each classi ie , he bes ea u e anking me hod (compa ing
he ensemble and single model’s anked ea u e lis s), da a scale (scaling only nume ical a iables),
and a io a e selec ed conside ing he combina ion’s pe o mance in p edic ing he mino i y class
(conside ing he ecall using he designed c oss- alida ion). In his way, he numbe o selec ed
ea u es depends on he da ase i sel and he used classi ie .
Fea u e Lis C ea ion
Gi en his, he used ea u e selec ion echnique e ie es he bes combina ion o :
• Fea u e Ranking Me hod – using he ensemble anking o single model PFI’s anking
• Scaling – Scaling nume ical ea u es (wi h S anda d Scale o Robus Scale ) o wi hou scaling
• Ra io – s a ing a 1 (using all ea u es) and decaying 0.001 in each in e ac ion; The a io s ops
dec easing when h ee consecu i e loops do no imp o e he c oss- alida ed sco e (0.001
decay was decided based on he da ase ’s cha ac e is ics).
The bes combina ion is he one wi h he bes ecall sco e in c oss- alida ion. Also, i mo e han one
combina ion p esen s he same ecall, he one wi h ewe ea u es is e ie ed.
Fea u e Lis Op imiza ion
A e his, he c ea ed ea u e lis s passed h ough he ollowing app oaches desc ibed below. He e,
he inal ea u e anking is also used as an impo ance measu e.
1. T y o add possible good ea u es o each model's lis
• A e he Da a Unde s anding phase, o each o he cus ome ’s segmen s, a lis o ea u es
is c ea ed wi h ea u es ha , based on his phase, a e belie ed o be good p edic o s o
he p oblem a hand.
24
• Fo each o he ea u es ha a e no on he gi en c ea ed lis , hey a e added (indi idually,
one a he ime) o he lis o see i he ecall sco e (in c oss- alida ion) imp o es - i yes,
should be added o he lis (manually).
2. T y o emo e possible bad/unnecessa y ea u es
• A p ima y s a i ied and shu led ain/ es pa i ion is c ea ed o 70/30 ( o speed up gi en
he nume ous combina ions o his phase)
• A each i e a ion, one o he ea u es is le ou (by in e se o de o impo ance de ined
o he gi en ea u es lis in he la e phase) - i he sco e imp o es/main ains, ha speci ic
ea u e is le ou o he nex in e ac ion.
• A e all he ea u es a e es ed, he p ocess epea s. The p ocess will epea un il no
changes a e done - o he gi en i e a ion, no ea u e emo al imp o es/main ains he
ecall sco e in he es spli .
• Gi en he emo ed ea u es o he p ima y ain/ es pa i ion, each o he ea u es
emo ed is es ed o be ou in he model (indi idually, each ea u e a he ime), and i s
c oss- alida ed ecall sco e is e u ned.
• I he a e age ecall sco e in he alida ion se s imp o es/main ains, he ea u e is
excluded (manually).
3. T y o eplace nume ical a iables o i s powe ans o ma ion
• In he Da a P epa a ion phase, Yeo-Johnson Powe ans o ma ion was used o
app oxima e he a iable’s dis ibu ion o no mal
• Each i e a ion eplaces one o he ea u es by hei ans o ma ion (by impo ance
o de / anking). I he sco e imp o es, hen i s ans o ma ion o he nex in e ac ion
eplaces ha speci ic ea u e.
• A e all he ea u es a e es ed, he p ocess epea s. The p ocess will epea un il no
changes a e done - no ea u e eplacemen imp o es he a e age ecall sco e o alida ion
se s.
• The selec ed ans o med ea u es a e manually eplaced in he inal ea u e lis .
4.4.3.2. Logis ic Reg ession’s Fea u e Selec ion wi h Backwa d Elimina ion
Fo he logis ic eg ession’s ea u e selec ion, an un egula ized Logis ic eg ession was cons uc ed
using all ea u es (a e he Da a P epa a ion phase), and Backwa d Elimina ion as ea u e selec ion
me hod was pe o med. As men ioned be o e, as he name sugges s, he leas signi ican ea u es a e
i e a i ely emo ed un il no o he e ec mee s he speci ied le el o emo al (Bu sac e al., 2008).
The emo al c i e ia a e based on he Wald es o indi idual pa ame e s, using a signi icance le el o
0.05 ( he p- alue cu -o was de ined as 0.05). He e, he null hypo hesis ha he coe icien o he
independen a iable is equal o ze o is es ed e sus an al e na i e hypo hesis ha he coe icien is
nonze o, which can be w i en as: 𝐻0:𝛽𝑥=0 𝑣𝑠 𝐻1:𝛽𝑥≠0 (Fo ho e e al., 2007).
Besides he collec ed a iables, some i e a ions and polynomial e ms we e in oduced and es ed o
hei signi icance in he model. Such in e ac ions among a iables we e added whene e he e we e
suspicions ha he impac o one a iable in he ou come also depended on ano he a iable’s alue.
Polynomial e ms we e added o co e / es cases whe e hei ela ionship wi h he a ge is no linea .
25
Las ly, T- es was used o es possible (ca ego ical) a iable agg ega ions, es ing i hei model
pa ame e s we e signi ican ly di e en o no (𝐻0:𝛽𝑥𝑖− 𝛽𝑥𝑗=0 𝑣𝑠 𝐻1:𝛽𝑥𝑖− 𝛽𝑥𝑗≠0). Fo he cases
in which he null hypo heses we e no ejec ed, he agg ega ion was conside ed by summing bo h
bina y a iables since, o he es ed cases, he posi i e e en is exclusi e (ne e happens
simul aneously in bo h).
Be o e his me hod was applied, ea u es ha p esen ed low a iances we e emo ed o p e en
mul icollinea i y. Such ea u es can lead o a singula ma ix (wi h de e minan equaling o ze o), which
means ha he ma ix has no in e se, making i impossible o es ima e he eg ession pa ame e s.
The esul ing se o ea u es and co esponding summa y s a is ics can be ound in Appendix A.
4.4.3.3. Model Cons uc ion
Thi d-Pa y is he segmen o which he highes pe cen age o o al p emium and he highes chu n
a e was obse ed. The e o e, he Thi d-Pa y segmen was he i s modeled segmen and o which
he esul s will be p esen ed.
As said be o e, c oss- alida ion was pe o med o assess he models and choose he co ec
pa ame e s. Besides pa ame e uning, in which mul iple pa ame e combina ions we e es ed o ind
he pa ame e s ha enhanced he esul s, o he s eps we e aken in his phase. Fo each model
cons uc ed, he ollowing desc ibed p ocesses we e p oceeded.
Fea u e Lis and Da a Scale Choice
He e, he combina ion o ea u e lis and da a scale (S anda d Scale , Robus Scale , o no scale ) ha
gua an eed he bes esul s o each model is selec ed and used in he subsequen phases.
Bellow, an example o G adien Boos ing model (wi hou pa ame e uning) o he Thi d-Pa y’s
da ase , compa ing i s c oss- alida ed esul s using all da ase ’s ea u es and he selec ed lis , is
p esen in Table 4.2. Fo his case, i is possible o obse e ha bo h esul s a e simila , indica ing ha
he lowes numbe o ea u es should be conside ed only.
# Fea u es
Speci ici y
T ain
Speci ici y
Valida ion
AUC
T ain
AUC
Valida ion
Recall
T ain
Recall
Valida ion
84
100%
+/- 0p.p.
99.7%
+/- 0p.p.
52.2%
+/- 0p.p.
50.2%
+/- 1p.p.
4.4%
+/- 1p.p.
0.6%
+/- 1p.p.
37
100%
+/- 0p.p.
97.2%
+/- 9p.p.
52.1%
+/- 0p.p.
50.8%
+/- 2p.p.
4.2%
+/- 1p.p.
4.4%
+/- 13p.p.
Table 4.2 G adien Boos ing’s ea u e lis selec ion (wi hou pa ame e uning)
Da a P e-P ocessing using SMOTE
As has been men ioned, one o he possible app oaches o deal wi h class imbalance is a da a-le el
app oach, whe e unde sampling o o e sampling can be used. Gi en he small dimensions o he
da ase s, using an unde sampling echnique was no app op ia e and, he e o e, an o e sampling
echnique was a emp ed. Gi en SMOTE’s popula i y due o i s simplici y and obus ness, his
echnique gene a ed new eco ds in o he aining se .

26
Al hough i is shown ha non- andom sampling echniques can highly imp o e he classi ie ’s
pe o mance, ha ing a 1:1 dis ibu ion ( he posi i e a e equaling he nega i e a e) migh be
un a o able (Fo man & Cohen, 2004). The e o e, smalle a ios o he mino i y class we e a emp ed
ia a w appe p ocess so ha he co ec new a io o he aining se s’ mino i y class o he model
in ques ion could be chosen.
Below, in Table 4.3, an example o he SMOTE’s pa ame e choice ha con ols he new pe cen age o
chu n obse ed in he da ase is p esen ed. I is impo an o no e ha his esampling is only done in
he aining se , so he eali y in which he model will pe o m can be espec ed.
Chu n Ra e in
T ain
Speci ici y
T ain
Speci ici y
Valida ion
AUC
T ain
AUC
Valida ion
Recall
T ain
Recall
Valida ion
o iginal
100%
+/- 0p.p.
98.5%
+/- 2p.p.
68.8%
+/- 3p.p.
51.3%
+/- 2p.p.
37.6%
+/- 5p.p.
4.0%
+/- 5p.p.
50%
98.9%
+/- 0p.p.
23.4%
+/- 11p.p.
95.9%
+/- 0p.p.
54.4%
+/- 5p.p.
92.8%
+/- 0p.p.
85.3%
+/- 11p.p.
33%
99.5%
+/- 3p.p.
32.1%
+/- 29p.p.
92.3%
+/- 0p.p.
54.7%
+/- 0p.p.
85.1%
+/- 1p.p.
77.2%
+/- 18p.p.
23%
99.8%
+/- 0p.p.
43.5%
+/- 26p.p.
87.2%
+/- 1p.p.
54.8%
+/- 1p.p.
74.5%
+/- 1p.p.
66.1%
+/- 27p.p.
17%
99.8%
+/- 0p.p.
55.3%
+/- 31p.p.
80.8%
+/- 1p.p.
55.4%
+/- 4p.p.
61.7%
+/- 2p.p.
55.6%
+/- 30p.p.
13%
100%
+/- 0p.p.
63.3%
+/- 26p.p.
75.6%
+/- 2p.p.
54.9%
+/- 3p.p.
51.2%
+/- 3p.p.
46.5%
+/- 29p.p.
12%
100%
+/- 0p.p.
65.8%
+/- 32p.p.
74.4%
+/- 1p.p.
53.8%
+/- 4p.p.
48.8%
+/- 2p.p.
41.9%
+/- 32p.p.
11%
100%
+/- 0p.p.
67.1%
+/- 33p.p.
73.3%
+/- 2p.p.
53.0%
+/- 3p.p.
46.6%
+/- 3p.p.
38.9%
+/- 35p.p.
Table 4.3 Example o Ex eme G adien Boos ing's - SMOTE's pa ame e choice (wi hou pa ame e
uning)
Model Calib a ion & Th eshold Tuning
As shown be o e, o ob ain good models (in e ms o p edic ion p obabili ies and p obabili ies’
classi ica ion), he model’s ou pu p obabili ies should be calib a ed, and he bes decision h eshold
o classi y he model’s esul in o chu n o enewal should be de ined.
Rega ding he p obabili y calib a ion, he model p edic ed p obabili ies and he ue alues a e used
o assess each model’s calib a ion pe o mance using he calib a ion plo . Then, mul iple copies o he
model, using k- old c oss- alida ion, a e i ed and, he p obabili ies p edic ed by hese models a e
hen calib a ed on he hold-ou se s. As calib a ion me hods, Pla and Iso onic Calib a ion we e
es ed. Fo his, he new p edic ed p obabili ies a e assessed by plo ing he calib a ion cu e and
compa ing i wi h bo h, he pe ec calib a ion line, and esul s be o e calib a ion.
27
Fo he choice o he h eshold, he h eshold wi h he op imal balance be ween alse posi i e
(speci ici y) and ue posi i e a e ( ecall), his is he op imal h eshold o he ROC cu e, was loca ed.
Wi h his knowledge, he me ics in c oss- alida ion a e op imized.
In o de o accomplish his, he ue posi i e a e o ecall (TPR) and ue nega i e a e o speci ici y
(TNR) a e compu ed o he p edic ions using a se o h esholds ( his can also be used o c ea e a ROC
Cu e plo ). Then, he geome ic mean (G-mean), which o mula is p esen ed below, ep esen s he
balance be ween bo h sco es ( he highe , he be e ).
𝐺𝑀𝑒𝑎𝑛= √(𝑇𝑃𝑅∗𝑇𝑁𝑅)
( 6 )
This me ic is obse ed o he di e en a emp ed h esholds, and he h eshold ha maximizes his
me ic, his is, ha op imizes he balance be ween TPR and TNR, is selec ed. Then, he op imal
h esholds ob ained ac oss he di e en spli s o he designed c oss- alida ion a e a e aged in o a inal
op imal h eshold and a e conside ed.
4.5. EVALUATION
In he E alua ion phase, he da a mining esul s we e assessed compa a i ely wi h he business success
c i e ia. The p ocess was also e iewed, and he inal model was chosen o deploymen .
The main goal o his p ojec is o imp o e he p o i ma gin o Ociden al’s Wo ke s’ Compensa ion
b anch by e aining cus ome s ha ha e in en ions o lea e gi en he p emium a ia ion su e ed in
he eno a ion p ocess. A p o i es ima ion analysis was pe o med o ensu e ha he inal model
would allow such an inc ease in he company’s e enue.
Fo his, some supposi ions we e cons uc ed based on he b anch’s business knowledge and a e
bellow explained. The esul s o his analysis a e based on each model’s con usion ma ix esul s and
o he business me ics.
• Re enue i he model p edic s co ec ly ha he cus ome enews (T ue Nega i es - TN)
𝐺𝑎𝑖𝑛𝑇𝑁 =𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚+∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚
( 7 )
I he model p edic s co ec ly ha he cus ome will enew, he p oposed p emium a ia ion
will main ain, e aining bo h cus ome p emium and p emium a ia ion.
• Re enue i he model p edic s ha cus ome chu ns, bu cus ome enews (False Posi i es -
FP)
𝐺𝑎𝑖𝑛𝐹𝑃=𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚
( 8 )
I he model p edic s ha he cus ome will chu n, hen he e is a p emium a ia ion dec ease
(o e en emo al). Gi en ha i is di icul o es ima e he p emium a ia ion o hose cases,
he wo s -case scena io (no inc ease) was conside ed.
• Re enue i he model p edic s co ec ly ha he cus ome chu ns (T ue Posi i es - TP)
𝐺𝑎𝑖𝑛𝑇𝑃 =𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚∗𝑟𝑒𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑟𝑎𝑡𝑖𝑜
( 9 )
28
The ue posi i es a e he cus ome s ha , i no model exis ed, would be los . Howe e , hese
cus ome s can be p ese ed by co ec ly iden i ying hei in en ions and applying con ingency
measu es (in his case, he dec ease in he p emium a ia ion).
Ne e heless, hese cases a e simila o he False Posi i e (FP) cases, whe e he model p edic s
ha hese cus ome s will also chu n. Thus, he e is s ill an inc ease in p emium ha is no
accoun ed o in his es ima ion.
Besides his, i is also obse able ha i is no possible o p ese e all isk by applying his
con ingency measu e: cus ome s s ill chu n when he e is no p emium a ia ion
(∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚=0%). A possible explana ion o such beha io is a be e p emium p oposal in
he compe i ion ha he company canno ma ch.
The e o e, i was decided o apply a e en ion a io, conside ing ha only 𝑋% o a ge ed
chu ns can be e e sed h ough he applied con ingency measu e. Gi en ha , such alue
needed o be es ima ed.
I was obse ed ha , by yea , 13% o o al chu n in he Thi d-Pa y segmen (wi h a s anda d
de ia ion o 2 p.p.) happens o a 0% p emium a ia ion. The e o e, i was decided ha such
alue should be ounded up conside ing i s s anda d de ia ion and, as well, gi e a 100%
ma gin as a sa e y measu e. Thus, conside ing he e en ion a io as:
𝑟𝑒𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑟𝑎𝑡𝑖𝑜𝑇ℎ𝑖𝑟𝑑𝑃𝑎𝑟𝑡𝑦=(1−2(0.13+0.02))=0.7
( 10 )
• Re enue i he model p edic s ha cus ome enews bu cus ome Chu ns (False Nega i es
- FN)
𝐺𝑎𝑖𝑛𝐹𝑁 =0
( 11 )
Those a e he cases whe e he model canno o esee he cus ome ’s chu n, so hey a e los ,
ha ing, o no ha ing a model. The e o e, nei he he cus ome ’s p emium no p emium
a ia ion is e ained.
Ha ing es ima ed he di e en gains ha each o he con usion ma ix’s measu es, he inc ease in o al
e enue is gi en by he di e ence be ween he a e age e enue wi h he model implemen ed (To Be)
and wi h no model (As Is):
𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒 𝑅𝑒𝑣𝑒𝑛𝑒=𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝑇𝑜 𝐵𝑒− 𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝐴𝑠 𝐼𝑆
( 12 )
whe e,
𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝐴𝑠 𝐼𝑆=#𝑅𝑤𝑙𝑠∗(𝐴𝑣𝑔 𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚+ 𝐴𝑣𝑔 ∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚)
( 13 )
and,
𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝑇𝑜𝐵𝑒=
𝑻𝑵𝑹 ∗ #𝑅𝑤𝑙𝑠 ∗ 𝐺𝑎𝑖𝑛𝑇𝑁 + 𝑭𝑷𝑹 ∗ #𝑅𝑤𝑙𝑠 ∗ 𝐺𝑎𝑖𝑛𝐹𝑃
+ 𝑻𝑷𝑹 ∗ #𝐶ℎ𝑛𝑠∗ 𝐺𝑎𝑖𝑛𝑇𝑃
( 14 )
29
Fo each cus ome segmen , he #𝑅𝑤𝑙𝑠 is he numbe o e i ied enewals in a gi en yea , he #𝐶ℎ𝑛𝑠
he numbe o e i ied chu ns in a gi en yea , 𝑻𝑵𝑹 he ue nega i e a e and, 𝑭𝑷𝑹 and 𝑻𝑷𝑹 as alse
and ue posi i e a es espec i ely.
Using he a e age me ic’s esul s (𝑻𝑵𝑹,𝑭𝑷𝑹,𝑻𝑷𝑹) ob ained in each o he models and he obse ed
alues (#𝑅𝑤𝑙𝑠 , #𝐶ℎ𝑠, 𝐴𝑣𝑔 𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚, 𝐴𝑣𝑔 ∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚) in he las h ee yea s, i was
possible o es ima e, o each model, he expec ed a e age inc ease in e enue wi h hei
implemen a ion.
𝐸𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒=1
3 ∑(𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒 𝑅𝑒𝑣𝑒𝑛𝑒𝑖)
𝑖 ∈ 𝑌
( 15 )
whe e 𝑌 a e he las h ee a ailable comple e yea s.
4.6. DEPLOYMENT
Then, he model’s deploymen is planned, and i s moni o ing and main enance plan is also designed.
A e a ull e iew, he model will inally be in eg a ed in o he company’s enewal p icing s a egy and
help he company inc ease he cus ome e en ion a e in he Wo ke s’ Compensa ion b anch.
Fo his pa , i was decided o deli e wo di e en au oma ed p ocesses, desc ibed below. In bo h,
Py hon is he p ima y ool used.
Mon hly P edic ions
The i s deli e able aims o deli e mon hly p edic ions o he mon h’s po en ial enewals and is
designed as ollows.
In a pa icula mon h and yea , he po en ial enewals da a is collec ed using SAS En e p ise Guide.
Py hon accesses he inal ou pu able and ec ea es in his da ase all o he necessa y da a
p epa a ion s eps. A his s age o he p ocess, he p emium a ia ion ha each policy will su e is
unknown. The company wan s o know each cus ome 's eac ion (p obabili y o chu n) o he di e en
possible p ice inc eases. The e o e, many p ice inc eases scena ios a e c ea ed, and each policy line
will be ec ea ed as many imes as he numbe o di e en p emiums inc eases.
The ained model is hen used o p edic he chu n p obabili y o each policy and di e en p emium
a ia ion scena ios. Then, a inal able, in which he di e en policies ( ows), possible p emium
a ia ion (columns), and cus ome eac ion (p obabili y o chu n as alue), is sen o SAS En e p ise
Guide. He e, an op imiza ion p ocess (de eloped by he company) o choose he p emium a ia ion
o each policy ha maximizes he global p o i will use he cons uc ed able as inpu .
Model Pe o mance Moni o ing
The second deli e able aims o deli e a p ocess in which he model’s pe o mance can be
con inuously assessed and u he in eg a ed in o a dashboa d.
This p ocess is simila o he mon hly p edic ions p ocess desc ibed abo e. The only di e ence is ha
bo h he a ge (chu n) and he p emium a ia ion a e al eady known when i uns. The goal is o
compa e he model’s pas p edic ions wi h wha happened (i policies ha e been chu ned o no ) and
unde s and i s heal h. The e o e bo h, he da a collec ion p ocess (in SAS En e p ise Guide), da a
36
Figu e 5.4 Es ima ed a e age inc ease in e enue o he Ex eme G adien Boos ing model wi h he
mono onic cons ain and Logis ic Reg ession
5.2. DISCUSSION
As s a ed ea lie , he p ojec had wo business goals:
1. Imp o e he b anches’ p o i by co ec ly iden i ying cus ome s who in ended o chu n (and
enew) wi h he p oposed p emium a ia ion.
I is c ucial o co ec ly iden i y he cus ome s ha will enew wi h he applied a ia ion as
well. The company needs o main ain he p emium a ia ion (inc ease) applied o he policies
in mos o he cases since i signi ican ly impac s p o i .
2. The selec ed model needs o be sui able o he enewal’s p icing p ocess.
The model aims o cap u e he cus ome s’ p ice elas ici y. The e o e, a mono onic ela ionship
be ween he a ge and P emium Va ia ion needs o be assu ed.
In e ms o p edic ing chu n, he e we e al eady some signs o he complexi y o he p oblem. Looking
a he p edic i e a iables’ co ela ion wi h he a ge , he me ic a iable showing a highe co ela ion
is P emium Va ia ion (0.10 conside ing Pea son co ela ion and 0.08 o Spea man co ela ion), and he
highe co ela ed non-me ic a iable is Flag Annual Paymen (wi h a C amme ’s V o 0.08). Mo eo e ,
he da ase ’s dimensions we e small, and o he Thi d-Pa y case, only a ound 9% o ins ances we e
classi ied as chu n. Such cha ac e is ics in he da ase can easily lead o p oblems such as o e i ing.
The e o e, i is essen ial o main ain he models simple.
Conside ing he model’s cons uc ion phase, i was hen possible o obse e ha models pe o med
be e , showing ewe signs o o e i ing, when hei complexi y was educed. Fo example, using
ewe lea ne s ( o ensemble models) o ewe neu ons/laye s ( o he neu al ne wo ks). On he o he
hand, he educ ion o he da a-skewness by including syn he ic samples ( ia SMOTE) in he aining
da ase did no imp o e esul s.
The esul s we e also posi i ely impac ed by selec ing a sui able p edic ion h eshold (di e en han
he de aul alue o 0.5) allowing a sui able ade-o be ween he ue posi i e ( ecall) and ue
nega i e a es (speci ici y). Mo eo e , gi en he deployed solu ion s uc u e, he ou pu p obabili ies
appea o ha e bene i ed om calib a ion, app oxima ing hem o hei ue alues. In Figu e 5.5, i is

37
possible o obse e he di e ences in he uncalib a ed and calib a ed p edic ions made by he Ex eme
G adien Boos ing model wi h he imposed mono onic ela ionship, using calib a ion plo s. Obse ing
Figu e 5.5(a), is possible o obse e ha he uncalib a ed model unde - o ecas s, his is he ob ained
p obabili ies a e smalle han expec ed. Obse ing Figu e 5.5(b), is possible o obse e ha calib a ed
p obabili ies a e close o diagonal line, he e o e sugges ing a be e calib a ed model.
By conside ing he ob ained esul s, i is possible o obse e he ollowing:
• The cons uc ed Ex eme G adien Boos ing using SMOTE in da a p e-p ocessing has shown
he highes le el o o e i ing and he lowes AUC on he alida ion and es se s.
• The highes AUC in bo h, c oss- alida ion and a e age o es , is ob ained o he G adien
Boos ing (wi h an ad an age o 0.9 p.p. in c oss- alida ion and 1.1 p.p. in a e age o es
esul s), ollowed by he Ex eme G adien Boos ing model.
By e alua ing he i s business goal (p o i inc ease), on a e age, he bes esul s a e ob ained o he
G adien Boos ing model. This model is also he leade when indi idually compa ing he impac ha
he achie ed me ics (in c oss- alida ion and a e age o es se s) ha e on he es ima ed a e age
inc ease in e enue.
Howe e , conside ing he second business goal, only he Logis ic Reg ession model could espec he
mono onic ela ionship be ween he P emium Va ia ion and he a ge . Howe e , i was possible o
o ce his mono onic ela ionship among bo h a iables using an Ex eme G adien Boos ing. By
compa ing his new model o he i s business goal’s leade (G adien Boos ing), i is possible o
obse e a dec ease in he c oss- alida ion’s AUC (o 1.1 p.p.) and in he a e age o es ’s AUC (o 1.9
p.p.). Such educ ion also impac s he es ima ed a e age inc ease in e enue, by a ound 11K €.
Though, his new model s ill has be e esul s han he cons uc ed Logis ic Reg ession. I is only
su passed by he cons uc ed (non-mono onic) G adien Boos ing and (non-mono onic) o iginal
Ex eme G adien Boos ing.
P edic ed P obabili y
Obse ed P obabili y
P edic ed P obabili y
Obse ed P obabili y
Figu e 5.5 Calib a ion plo s o XGBoos Mono onic p edic ions (a) Uncalib a ed model
P obabili ies. (b) Calib a ed model p edic ions.
(a)
(b)
38
6. CONCLUSIONS
This epo esul s om a se en-mon h p ojec de eloped in a one-yea in e nship in G upo Ageas
Po ugal. Bellow, he main conclusions o his p ojec a e p esen ed.
G upo Ageas Po ugal is one o he la ges insu ance p o ide s in Po ugal, wi h di e en b ands such
as Ociden al Segu os, which p o ides o e s in he li e and non-li e b anches. This la e o e s a
Wo ke s’ Compensa ion insu ance, which is manda o y in Po ugal.
The Ociden al’s Wo ke s’ Compensa ion po olio can be segmen ed in o h ee main isk g oups (Thi d-
Pa y, Sel -Employees, and Housekeepe s), which ha e di e en enewal’s p icing s a egies and
p esen di e en beha io s and a ailable in o ma ion.
Nowadays, cus ome s' ease in explo ing he a ailable op ions allows hem o change he insu ance
p o ide easily. Gi en he compe i i eness o he insu ance ma ke , companies need o ake ac ion o
e ain hei cus ome s since i has a signi ican impac on hei p o i .
The enewal’s p icing p ocess occu s yea ly a he cus ome s' enewal da e, p oposing a a ia ion o
he policy’s paid p emium. The company’s goal was o educe he obse ed chu n a es consequen o
his p ocess. The e o e, he company wan ed o in eg a e a model ha could iden i y ea ly signs o
chu n acco dingly o he di e en possible p emium a ia ions ha he gi en policy could su e . So,
by cap u ing i s cus ome s’ p ice elas ici y, p emium a ia ions can be co ec ly adjus ed, and
cus ome s ha p esen a ce ain isk le el o lea ing, possibly main ained.
The CRISP-DM me hodology was applied o accomplish his p ojec , co e ing i s main s eps, and going
back and o wa d when necessa y.
The i s pa was dedica ed o unde s anding he business i sel , i s goals, and which a iables could
p o ide help ul in o ma ion abou his cus ome beha io . A ound 80 a iables (nume ical and
ca ego ical) we e ex ac ed om he company’s da abase. Howe e , as obse ed in he da a
unde s anding phase, no all a iables we e ele an , especially when sepa a ed wi hin he h ee
cus ome segmen s. This phase was ins umen al in unde s anding some o he s eps ha should be
done in he Da a P epa a ion phase, whe e da a in i s aw o m we e p epa ed o modeling. He e,
ac i i ies like edundan ea u e emo al, handling missing alues, c ea ing new a ibu es, and
o ma ing he da a in a way ha would allow (and help) he modeling phase we e ca ied on.
Mo eo e , as an al e na i e o using municipali ies o dis ic s, new geog aphical a eas we e c ea ed.
Fo his, clus e ing analysis was pe o med and, o each o he cus ome segmen s, new geog aphical
(con iguous) a eas we e c ea ed based on he municipali ies’ ele an demog aphic in o ma ion and
obse ed chu n a es. Thus, i was possible o educe eigh een new dimensions (in case dis ic s we e
used) in o only ou ele an dimensions, using in o ma ion ega ding mo e han h ee hund ed
municipali ies.
A e ha ing he da ase p epa ed, ea u es o be used by each o he models a e selec ed. The goal
was o use a sui able numbe o ea u es gi en he numbe o ows a ailable o modeling. Mo e
ea u es ep esen a highe p oblem complexi y, and o which, models should ha e mo e da a o
unde s and. The ea u es a e chosen using ensemble lea ning, combining mul iple lea ne s’ opinions
abou he ea u e’s impo ance ankings. The Top N ea u es a e selec ed, o each model, based on
39
he esul s ob ained o each se o ea u es. Fo he Logis ic eg ession, backwa d elimina ion based
on he Wald’s es was used.
The model’s pe o mance was e alua ed ac oss all he modeling phases. Time-spli c oss- alida ion
was used so he exis ing empo al dependencies on he da a could be espec ed. A he same ime,
he model’s gene aliza ion abili y can be gua an eed, app oxima ing he alida ion esul s o he ac ual
model pe o mance.
The inal model selec ion decision was based on he company’s wo business goals: inc easing he
b anches’ p o i (acco dingly wi h he model’s pe o mance) and ha e a sui able model o op imizing
he enewal’s p icing p ocess. Fo his la e , a mono onic inc easing ela ionship among P emium
Va ia ion and P obabili y o Chu n needs o be assu ed.
The inal selec ed model was an Ensemble Lea ning echnique, he Ex eme G adien Boos ing model
o which he mono onic ela ionship was o ced. This model shows in c oss- alida ion an AUC o 59.0%
(wi h a s anda d de ia ion o 3 p.p.), a speci ici y o 74.4% (wi h a s anda d de ia ion o 7 p.p.) and,
las ly, a ecall o 43.5% (wi h a s anda d de ia ion o 10 p.p.). As o he a e age esul s in he used
es se s, an AUC o 56.4% (wi h a s anda d de ia ion o 2 p.p.), a speci ici y o 72.7% (wi h a s anda d
de ia ion o 12 p.p.) and, las ly, a ecall o 40.2% (wi h a s anda d de ia ion o 12 p.p.) we e obse ed.
Fu he mo e, his model can answe he company’s needs. Fi s ly, i is sui able o he op imiza ion o
he enewal’s p icing p ocess: o a gi en enewal yea and policy, he deli e ed model allows ha i
he company decides o inc ease (o dec ease) he p emium a ia ion, hen he p obabili y o chu n
will ei he main ain he same alue o inc ease (o dec ease) in i s alue. The e o e, pe mi ing da a-
d i en decisions. Las ly, he model allows an inc ease in p o i : i was es ima ed ha his model could
con ibu e o an annual a e age inc ease o a ound 70K €.
40
7. LIMITATIONS AND RECOMMENDATIONS FOR FUTURE WORKS
I is possible o no ice ha he deli e ed solu ion has some space o imp o emen since he esul s in
e ms o pe o mance measu es could be highe . Resul s a e highly dependen on he da a used and,
he e, some imp o emen s could be made, o o he echniques a emp ed, like o example he es o
o he classi ica ion algo i hms.
Se e al issues we e iden i ied du ing he da a collec ion ask. Many inconsis encies we e de ec ed
among he di e en da a sou ces ( he company is cu en ly cons uc ing a uni ied sou ce o
in o ma ion). Al hough such p oblems we e deal wi hin he bes way possible, di e en quali y issues
ha e eme ged, impac ing he inal collec ed da ase . The e o e, he quali y in he used da ase o
modeling could no be assu ed. I was also no possible o collec da a o many o he company’s
policies, impac ing he inal da ase ’s dimensions. E en ha ing da a o ou whole yea s, a highe
amoun o da a could su ely help models in hei esul s. Addi ionally, he e a e possibly o he
a iables ha a e no a ailable o he company bu could be ele an , like ex e nal ac o s ha migh
in luence he cus ome ’s beha io .
Rega ding da a p e-p ocessing, o he echniques could be es ed as o he missing alues impu a ion
echniques, a deepe ou lie de ec ion analysis, o he o e sampling echniques, and u he
dimensionali y educ ion ( ea u e selec ion) echniques could be explo ed as well.
Las ly, as was possible o obse e, some models p esen be e esul s han o he s. The e o e, i could
be in e es ing o explo e o he models ha a e mono onic o allow a mono onic cons ain . Two
models ha would be in e es ing o es a e Ligh GBM ( om Py hon’s Ligh GBM package (Ke e al.
2017)) and la ice-based models ( om Py hon’s Tenso Flow-La ice package (Google AI Blog, 2017)).
Simila ly o he Ex eme G adien Boos ing model ( om he xgboos package (Chen & Gues in, 2016)),
hese models allow he in oduc ion o mono onic cons ain s. Howe e , due o secu i y cons ain s
ha he company needs o ensu e, i was no possible o ob ain such packages in a iable ime o his
p ojec execu ion.
41
8. BIBLIOGRAPHY
Ahn, J., Hwang, J., Kim, D., Choi, H., & Kang, S. (2020). A Su ey on Chu n Analysis in Va ious Business
Domains. IEEE Access, 8, 220816–220839. h ps://doi.o g/10.1109/ACCESS.2020.3042657
Aye , M., B unk, H. D., Ewing, G. M., Reid, W. T., & Sil e man, E. (1955). An empi ical dis ibu ion
unc ion o sampling wi h incomple e in o ma ion. The Annals o Ma hema ical S a is ics,
26(4), 641–647. h ps://doi.o g/10.1214/aoms/1177728423
Ba is a, G. E., & Mona d, M. C. (2003). An analysis o ou missing da a ea men me hods o
supe ised lea ning. Applied A i icial In elligence, 17(5–6), 519–533.
h ps://doi.o g/10.1080/713827181
Bellman, R. (1966). Dynamic p og amming. Science, 153(3731), 34–37.
h ps://doi.o g/10.1126/science.153.3731.34
Bisong, E. (2019). Building Machine Lea ning and Deep Lea ning Models on Google Cloud Pla o m. In
Ap ess, Be keley, CA. h ps://doi.o g/10.1007/978-1-4842-4470-8_29
Bolancé, C., Guillen, M., & Padilla-Ba e o, A. E. (2016). P edic ing P obabili y o Cus ome Chu n in
Insu ance. In R. León, M. Muñoz-To es, & J. Mone a (Eds.), Modeling and Simula ion in
Enginee ing, Economics and Managemen . MS 2016. Lec u e No es in Business In o ma ion
P ocessing, ol 254 (pp. 82–91). Sp inge , Cham. h ps://doi.o g/10.1007/978-3-319-40506-3_9
Bollegala, D. (2017). Dynamic ea u e scaling o online lea ning o bina y classi ie s. Knowledge-
Based Sys ems, 129, 97–105. h ps://doi.o g/10.1016/j.knosys.2017.05.010
Bolón-Canedo, V., & Alonso-Be anzos, A. (2019). Ensembles o ea u e selec ion: A e iew and
u u e ends. In o ma ion Fusion, 52, 1–12. h ps://doi.o g/10.1016/j.in us.2018.11.008
Box, G. E., & Cox, D. R. (1964). An Analysis o T ans o ma ions. Jou nal o he Royal S a is ical Socie y:
Se ies B (Me hodological), 26(2), 211–243. h ps://doi.o g/10.1111/j.2517-6161.1964. b00553.x
B eiman, L. (2001). Random Fo es s. Machine Lea ning, 45, 5–32.
h ps://doi.o g/10.1023/A:1010933404324
Bu sac, Z., Gauss, C. H., Williams, D. K., & Hosme , D. W. (2008). Pu pose ul selec ion o a iables in
logis ic eg ession. Sou ce Code o Biology and Medicine, 3(17). h ps://doi.o g/10.1186/1751-
0473-3-17
Chawla, N. V., Bowye , K. W., Hall, L. O., & Kegelmeye , W. P. (2002). SMOTE: Syn he ic Mino i y
O e -sampling Technique. Jou nal o A i icial In elligence Resea ch, 16(1), 321–357.
h ps://doi.o g/10.5555/1622407.1622416
Chawla, N. V., Cieslak, D. A., Hall, L. O., & Joshi, A. (2008). Au oma ically coun e ing imbalance and i s
empi ical ela ionship o cos . Da a Mining and Knowledge Disco e y, 17, 225–252.
h ps://doi.o g/10.1007/s10618-008-0087-0
Chawla, N. V., Japkowicz, N., & Ko cz, A. (2004). Edi o ial: special issue on lea ning om imbalanced
da a se s. SIGKDD Explo ., 6, 1-6. h ps://doi.o g/10.1145/1007730.1007733
Chen, T., & Gues in, C. (2016). XGBoos : A Scalable T ee Boos ing Sys em. P oceedings o he 22nd
Acm Sigkdd In e na ional Con e ence on Knowledge Disco e y and Da a Mining, 785–794.
h ps://doi.o g/10.1145/2939672.2939785

42
Chiew, K. L., Tan, C. L., Wong, K., Yong, K. S., & Tiong, W. K. (2019). A new hyb id ensemble ea u e
selec ion amewo k o machine lea ning-based phishing de ec ion sys em. In o ma ion
Sciences, 484, 153–166. h ps://doi.o g/10.1016/j.ins.2019.01.064
Chuang, H. C., Chen, C. C., & Li, S. T. (2020). Inco po a ing mono onic domain knowledge in suppo
ec o lea ning o da a mining eg ession p oblems. Neu al Compu ing and Applica ions,
32(15), 11791–11805. h ps://doi.o g/10.1007/s00521-019-04661-4
Collell, G., P elec, D., & Pa il, K. R. (2018). A simple plug-in bagging ensemble based on h eshold-
mo ing o classi ying bina y and mul iclass imbalanced da a. Neu ocompu ing, 275, 330–340.
h ps://doi.o g/10.1016/j.neucom.2017.08.035
De Win e , J. C., Gosling, S. D., & Po e , J. (2016). Supplemen al Ma e ial o Compa ing he Pea son
and Spea man Co ela ion Coe icien s Ac oss Dis ibu ions and Sample Sizes: A Tu o ial Using
Simula ions and Empi ical Da a. Psychological Me hods, 21(3), 273–290.
h ps://doi.o g/10.1037/me 0000079.supp
Die e ich, T. G. (2002). Ensemble Lea ning. The Handbook o B ain Theo y and Neu al Ne wo ks,
2(1), 110–125.
Dola abadi, S. H., & Keynia, F. (2017). Designing o Cus ome and Employee Chu n P edic ion Model
Based on Da a Mining Me hod and Neu al P edic o . 2017 2nd In e na ional Con e ence on
Compu e and Communica ion Sys ems (ICCCS), 74–77.
h ps://doi.o g/10.1109/CCOMS.2017.8075270
Do mann, C. F. (2020). Calib a ion o p obabili y p edic ions om machine‐lea ning and s a is ical
models. Global Ecology and Biogeog aphy, 29(4), 760–765. h ps://doi.o g/10.1111/geb.13070
El Aboudi, N., & Benhlima, L. (2016). Re iew on W appe Fea u e Selec ion App oaches. 2016
In e na ional Con e ence on Enginee ing & MIS (ICEMIS), 1–5.
h ps://doi.o g/0.1109/ICEMIS.2016.7745366
El Bouche y, K., & De Souza, R. S. (2020). Lea ning in Big Da a: In oduc ion o Machine Lea ning. In
Knowledge Disco e y in Big Da a om As onomy and Ea h Obse a ion (pp. 225–249).
Else ie . h ps://doi.o g/10.1016/b978-0-12-819154-5.00023-0
Fe nández, A., Ga cía, S., He e a, F., & Chawla, N. V. (2018). SMOTE o Lea ning om Imbalanced
Da a: P og ess and Challenges, Ma king he 15-yea Anni e sa y. Jou nal o A i icial
In elligence Resea ch, 61, 863–905. h ps://doi.o g/10.1613/jai .1.11192
Fo man, G., & Cohen, I. (2004). Lea ning om li le: Compa ison o classi ie s gi en li le aining. In J.
Boulicau , F. Esposi o, F. Gianno i, & D. Ped eschi (Eds.), Knowledge Disco e y in Da abases:
PKDD 2004. PKDD 2004. Lec u e No es in Compu e Science, ol 320 (pp. 161–172). Sp inge ,
Be lin, Heidelbe g. h ps://doi.o g/10.1007/978-3-540-30116-5_17
Fo ho e , R. N., Lee, E. S., & He nandez, M. (2007). Logis ic and P opo ional Haza ds Reg ession. In
Bios a is ics (Second Edi ion) (pp. 387–419). Academic P ess. h ps://doi.o g/10.1016/B978-0-
12-369492-8.50019-4
F iedman, J. H. (2001). G eedy unc ion app oxima ion: a g adien boos ing machine. The Annals o
S a is ics, 29(5), 1189–1232. h ps://doi.o g/10.1214/aos/1013203451
F iedman, J. H. (2002). S ochas ic g adien boos ing. Compu a ional S a is ics & Da a Analysis, 38(4),
367–378. h ps://doi.o g/10.1016/S0167-9473(01)00065-2
43
Ga cía, S., Luengo, J., & He e a, F. (2016). Tu o ial on p ac ical ips o he mos in luen ial da a
p ep ocessing algo i hms in da a mining. Knowledge-Based Sys ems, 98, 1–29.
h ps://doi.o g/10.1016/j.knosys.2015.12.006
Ga dne , M. W., & Do ling, S. R. (1998). A i icial neu al ne wo ks ( he mul ilaye pe cep on) - a
e iew o applica ions in he a mosphe ic sciences. A mosphe ic En i onmen , 32(14–15), 2627–
2636. h ps://doi.o g/10.1016/S1352-2310(97)00447-0
Gün he , C. C., T e e, I. F., Aas, K., Sandnes, G. I., & Bo gan, Ø. (2014). Modelling and p edic ing
cus ome chu n om an insu ance company. Scandina ian Ac ua ial Jou nal, 2014(1), 58–71.
h ps://doi.o g/10.1080/03461238.2011.636502
Google AI Blog (2017). Tenso low la ice: Flexibili y empowe ed by p io knowledge. Re i ed om
h ps://ai.googleblog.com/2017/10/ enso low-la ice- lexibili y.h ml
Ha is, R., Sleigh , P., & Webbe , R. (2005). In oducing Geodemog aphics. In Geodemog aphics, GIS
and neighbou hood a ge ing (pp. 16–17). John Wiley & Sons.
He, H., & Ga cia, E. A. (2009). Lea ning om Imbalanced Da a. IEEE T ansac ions on Knowledge and
Da a Enginee ing, 21(9), 1263–1284. h ps://doi.o g/10.1109/TKDE.2008.239
He, Y., Xiong, Y., & Tsai, Y. (2020). Machine Lea ning Based App oaches o P edic Cus ome Chu n
o an Insu ance Company. 2020 Sys ems and In o ma ion Enginee ing Design Symposium
(SIEDS), 1–6. h ps://doi.o g/10.1109/SIEDS49339.2020.9106691
Hua, J., Xiong, Z., Lowey, J., Suh, E., & Doughe y, E. R. (2005). Op imal numbe o ea u es as a
unc ion o sample size o a ious classi ica ion ules. Bioin o ma ics, 21(8), 1509–1515.
h ps://doi.o g/10.1093/bioin o ma ics/b i171
Inouye, D. I., Leqi, L., Kim, J. S., A agam, B., & Ra ikuma , P. (2020). Au oma ed Dependence Plo s.
P oceedings o he 36 h Con e ence on Unce ain y in A i icial In elligence (UAI), 124, 1238–
1247. h p://a xi .o g/abs/1912.01108
Jain, A. K., & Chand aseka an, B. (1982). Dimensionali y and sample size conside a ions in pa e n
ecogni ion p ac ice. In Handbook o S a is ics (Vol. 2, pp. 835–855).
h ps://doi.o g/10.1016/S0169-7161(82)02042-2
Lemaî e, G., Noguei a, F., & A idas, C. K. (2017). Imbalanced-lea n: A py hon oolbox o ackle he
cu se o imbalanced da ase s in machine lea ning. The Jou nal o Machine Lea ning
Resea ch, 18(1), 559-563.
Kalousis, A., P ados, J., & Hila io, M. (2007). S abili y o ea u e selec ion algo i hms: a s udy on high-
dimensional spaces. Knowledge and In o ma ion Sys ems, 12, 95–116.
h ps://doi.o g/10.1007/s10115-006-0040-8
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W. Ma, W., Ye, Q., Liu, T.Y (2017). Ligh gbm: A highly
e icien g adien boos ing decision ee. In Ad ances in neu al in o ma ion p ocessing sys ems,
30, 3146-3154.
Koha i, R., & John, G. H. (1997). W appe s o ea u e subse selec ion. A i icial In elligence, 97(1–2),
273–324. h ps://doi.o g/10.1016/S0004-3702(97)00043-X
Kuma , V., & Minz, S. (2014). Fea u e Selec ion: A li e a u e Re iew. The Sma Compu ing Re iew,
4(3), 211–229. h ps://doi.o g/10.6029/sma c .2014.03.007
44
Male ic, J., & Ma cus, A. (2000). Da a Cleansing: Beyond In eg i y Analysis. Iq, 200–209.
McCue, C. (2015). Iden i ica ion, Cha ac e iza ion, and Modeling. In Da a Mining and P edic i e
Analysis (Second Edi ion) (pp. 137–155). h ps://doi.o g/10.1016/B978-0-12-800229-2.00007-9
McGo e n, A., Lage quis , R., John Gagne, D., Je gensen, G. E., Elmo e, K. L., Homeye , C. R., & Smi h,
T. (2019). Making he Black Box Mo e T anspa en : Unde s anding he Physical Implica ions o
Machine Lea ning. Bulle in o he Ame ican Me eo ological Socie y, 100(11), 2175–2199.
h ps://doi.o g/10.1175/BAMS-D-18-0195.1
Mei , R., & Rä sch, G. (2003). An In oduc ion o Boos ing and Le e aging. In S. Mendelson & A. J.
Smola (Eds.), Ad anced lec u es on machine lea ning (pp. 118–183). Sp inge , Be lin,
Heidelbe g. h ps://doi.o g/10.1007/3-540-36434-X_4
Naeini, M. P., Coope , G. F., & Hausk ech , M. (2015). Ob aining Well Calib a ed P obabili ies Using
Bayesian Binning. P oceedings o he Twen y-Nin h AAAI Con e ence on A i icial In elligence,
2901–2907.
Na ekin, A., & Knoll, A. (2013). G adien boos ing machines, a u o ial. F on ie s in Neu o obo ics, 7,
21. h ps://doi.o g/10.3389/ nbo .2013.00021
Niculescu-Mizil, A., & Ca uana, R. (2005). P edic ing good p obabili ies wi h supe ised lea ning.
P oceedings o he 22nd In e na ional Con e ence on Machine Lea ning, 625–632.
h ps://doi.o g/10.1145/1102351.1102430
Njikam, A. N. S., & Zhao, H. (2016). A no el ac i a ion unc ion o mul ilaye eed- o wa d neu al
ne wo ks. Applied In elligence, 45, 75–82. h ps://doi.o g/10.1007/s10489-015-0744-0
Osbo ne, J. W. (2010). Imp o ing you da a ans o ma ions: Applying he Box-Cox ans o ma ion.
P ac ical Assessmen , Resea ch and E alua ion, 15(12). h ps://doi.o g/10.7275/qbpc-gk17
Ped egosa, F., Va oquaux, G., G am o , A., Michel, V., Thi ion, B., G isel, O., Blondel, M.,
P e enho e , P., Weiss, R., Dubou g, V., Vande plas, J., Passos, A., Cou napeau, D., B uche , M.,
Pe o , M., & Duchesnay, E. (2011). Sciki -lea n: Machine Lea ning in Py hon. Jou nal o
Machine Lea ning Resea ch, 12, 2825–2830. Re ei ed om
h ps://jml .csail.mi .edu/pape s/ olume12/ped egosa11a/ped egosa11a.pd
Pham, D. T., Dimo , S. S., & Nguyen, C. D. (2005). Selec ion o K in K-means clus e ing. P oceedings o
he Ins i u ion o Mechanical Enginee s, Pa C: Jou nal o Mechanical Enginee ing Science,
219(1), 103–119. h ps://doi.o g/10.1243/095440605X8298
Polika , R. (2006). Ensemble based sys ems in decision making. IEEE Ci cui s and Sys ems Magazine,
6(3), 21–45. h ps://doi.o g/10.1109/MCAS.2006.1688199
Pozzolo, A. D., Caelen, O., Johnson, R. A., & Bon empi, G. (2015). Calib a ing P obabili y wi h
Unde sampling o Unbalanced Classi ica ion. 2015 IEEE Symposium Se ies on Compu a ional
In elligence, 159–166. h ps://doi.o g/10.1109/SSCI.2015.33
P ice, B. (2002). Making CRM come o li e. E-Business Re iew, 25–31.
P o os , F. (2008). Machine lea ning om imbalanced da a se s 101.
Rokach, L., & Maimon, O. (2005). Clus e ing Me hods. In O. Maimon & L. Rokach (Eds.), Da a Mining
and Knowledge Disco e y Handbook (pp. 321–352). Sp inge , Bos on, MA.
h ps://doi.o g/10.1007/0-387-25465-X_15
45
Rubin, D. B. (1976). In e ence and missing da a. Biome ika, 63(3), 581–592.
h ps://doi.o g/10.1093/biome /63.3.581
Saeys, Y., Abeel, T., & Van De Pee , Y. (2008). Robus Fea u e Selec ion Using Ensemble Fea u e
Selec ion Techniques. In W. Daelemans (Ed.), Machine Lea ning and Knowledge Disco e y in
Da abases. ECML PKDD 2008. Lec u e No es in Compu e Science, ol 5212 (pp. 313–325).
Sp inge -Ve lag Be lin Heidelbe g 2008. h ps://doi.o g/10.1007/978-3-540-87481-2_21
Sasse , W. E., & Reichheld, F. F. (1990). Ze o De ec ions -Quali y Comes o Se ices. Ha a d Business
Re iew, 68(5), 105–111.
Sc iney, M., Nie, D., & Roan ee, M. (2020). P edic ing Cus ome Chu n o Insu ance Da a. In M.
Song, I. Song, G. Ko sis, A. M. Tjoa, & I. Khali (Eds.), Big Da a Analy ics and Knowledge Disco e y.
DaWaK 2020. Lec u e No es in Compu e Science, ol 12393 (pp. 256–265). Sp inge , Cham.
h ps://doi.o g/10.1007/978-3-030-59065-9_21
Seabold, S., & Pe k old, J. (2010). S a smodels: Econome ic and s a is ical modeling wi h py hon. In
P oceedings o he 9 h Py hon in Science Con e ence. (Vol. 57, pp. 61).
Seijo-Pa do, B., Bolón-Canedo, V., & Alonso-Be anzos, A. (2017). Tes ing Di e en Ensemble
Con igu a ions o Fea u e Selec ion. Neu al P ocessing Le e s, 46, 857–880.
h ps://doi.o g/10.1007/s11063-017-9619-1
Seijo-Pa do, B., Bolón-Canedo, V., Po o-Díaz, I., & Alonso-Be anzos, A. (2015). Ensemble Fea u e
Selec ion o Rankings o Fea u es. In I. Rojas, G. Joya, & A. Ca ala (Eds.), Ad ances in
Compu a ional In elligence. IWANN 2015. Lec u e No es in Compu e Science, ol 9095.
Sp inge , Cham. h ps://doi.o g/10.1007/978-3-319-19222-2_3
Shah, R., & Pe e ia ko, V. (2021). Using Fea u e Impo ance Rank Ensembling (FIRE) o Ad anced
Fea u e Selec ion. Re i ed om h ps://www.da a obo .com/blog/using- ea u e-impo ance-
ank-ensembling- i e- o -ad anced- ea u e-selec ion/
Spi e i, M., & Azzopa di, G. (2018). Cus ome Chu n P edic ion o a Mo o Insu ance Company. 2018
Thi een h In e na ional Con e ence on Digi al In o ma ion Managemen (ICDIM), 173–178.
h ps://doi.o g/10.1109/ICDIM.2018.8847066
S obl, C., Boules eix, A. L., Kneib, T., Augus in, T., & Zeileis, A. (2008). Condi ional a iable
impo ance o andom o es s. BMC Bioin o ma ics, 9, 307. h ps://doi.o g/10.1186/1471-
2105-9-307
Sunda kuma , G. G., & Ra i, V. (2015). A no el hyb id unde sampling me hod o mining unbalanced
da ase s in banking and insu ance. Enginee ing Applica ions o A i icial In elligence, 37, 368–
377. h ps://doi.o g/10.1016/j.engappai.2014.09.019
Talabis, M., McPhe son, R., Miyamo o, I., & Ma in, J. (2015). Analy ics De ined. In In o ma ion
secu i y analy ics: inding secu i y insigh s, pa e ns, and anomalies in big da a (pp. 1–12).
Syng ess. h ps://doi.o g/10.1016/b978-0-12-800207-0.00001-0
Toloşi, L., & Lengaue , T. (2011). Classi ica ion wi h co ela ed ea u es: un eliabili y o ea u e
anking and solu ions. Bioin o ma ics, 27(14), 1986–1994.
h ps://doi.o g/10.1093/bioin o ma ics/b 300
Va eiadis, T., Diaman a as, K. I., Sa igiannidis, G., & Cha zisa as, K. C. (2015). A compa ison o
machine lea ning echniques o cus ome chu n p edic ion. Simula ion Modelling P ac ice and
Theo y, 55, 1–9. h ps://doi.o g/10.1016/j.simpa .2015.03.003