scieee Open visual document viewer

Customer Churn Prediction in Insurance: Modeling Renewal Price Elasticity of the Workers’ Compensation Portfolio from Ocidental Seguros

Castro, Laura Sofia Sauthoff da Ponte e

Abstract

Customer churn has been increasing in insurance, mainly due to technological improvements that allow customers to explore other insurance providers’ offers. Given this, insurance providers need to compete among them, not only to get new customers but, to maintain their own. This report results from a project developed during an internship in Grupo Ageas Portugal, which has different insurance brands such as Ocidental Seguros. This project's main goal was to model, with a monthly periodicity, customer churn of this latter’s Workers’ Compensation portfolio to improve the company’s competitiveness and, ultimately, profit. Many of the company’s customer churn happens at their policy renewal time, where the only variable that the company detains control over is the price (premium) variation. Hence, by considering the premium variation and other relevant predictive variables, the goal was to predict the probability of a given customer to churn, allowing the company to optimize the current renewal’s pricing process and maximize this branch’s profit. Thus, different variables that could influence the company’s customer behavior were collected—one of those was the customer’s location. Given the high dimensionality that such variable would represent and the small dataset available for modeling, clustering analysis is used to create new significant (with fewer dimensions) customer geographical areas. Different supervised learning algorithms were then evaluated accordingly to their performance in predicting customer churn. The predictive models used were a Gradient Boosting, an Extreme Gradient Boosting, a Logistic Regression, and a Multilayer Perceptron. Given that the number of customers that renew their contracts is much superior to the number of customers who churn, Synthetic Minority Oversampling Technique (SMOTE) was used to create less unbalanced datasets (with synthetic samples) and evaluate the impact on the performance of one of the models. Lastly, to guarantee a successful integration of the models into the renewal’s pricing process, models were evaluated accordingly to the two business goals. First, by translating the observed evaluation metrics into profit. Secondly, by assuring that the customer’s price elasticity would be captured, assuring a monotonic increasing relationship among the policy’s premium variation and probability of churn.

Full text

i Cus ome Chu n P edic ion in Insu ance: Lau a So ia Sau ho da Pon e e Cas o Modeling Renewal P ice Elas ici y o he Wo ke s’ Compensa ion Po olio om Ociden al Segu os In e nship epo p esen ed as pa ial equi emen o ob aining he Mas e ’s deg ee in Ad anced Analy ics ii i Cus ome Chu n P edic ion in Insu ance: Modeling Renewal P ice Elas ici y o he Wo ke s’ Compensa ion Po olio om Ociden al Segu os Lau a So ia Sau ho da Pon e e Cas o MAA 2021 i ii NOVA In o ma ion Managemen School Ins i u o Supe io de Es a ís ica e Ges ão de In o mação Uni e sidade No a de Lisboa CUSTOMER CHURN PREDICTION IN INSURANCE: MODELING RENEWAL PRICE ELASTICITY OF THE WORKERS’ COMPENSATION PORTFOLIO FROM OCIDENTAL SEGUROS by Lau a So ia Sau ho da Pon e e Cas o In e nship epo p esen ed as pa ial equi emen o ob aining he Mas e ’s deg ee in Ad anced Analy ics Ad iso : Jo ge Mo ais Mendes Co Ad iso : Nuno An ónio No embe 2021 iii ACKNOWLEDGMENTS Fi s , I would like o s a o hank my ad iso , P o esso Jo ge Mo ais Mendes, o being a pa o his p ojec since i s beginning, suppo ing me in all he challenges ha I ha e aced. As well, my co-ad iso , P o esso Nuno An ónio, o helping me go u he wi h his ema kable ision and insigh s. No only by all he pa ience and suppo hey ha e p o ided, bu also, by all he knowledge hey ha e passed du ing he lec u es in he i s yea o he Mas e ’s deg ee. Mo eo e , I wan o exp ess my g a i ude o my in e nship supe iso , João Ped o Oli ei a, o us ing me wi h his p ojec and o all he suppo , knowledge, and oppo uni ies he has gi en me o hese mon hs. Also, And é Pousinho and Joana Neg a, o all he expe ience hey ha e sha ed wi h me in he i s pa o my in e nship, allowing me o unde s and he insu ance ma ke and, pa icula ly, he Wo ke s’ Compensa ion b anch. I would also like o hank G upo Ageas Po ugal o allowing me o join hem o one yea , and e e yone who has sha ed wi h me his pa h. Fu he mo e, o all he non-li e b anch P icing and Business Analy ics eam who ha e made me eel so welcome and we e always willing o help in any hing I would need. Las ly, I would like o hank my iends and amily o belie ing in me and being a pa o my academic jou ney and li e. i ABSTRACT Cus ome chu n has been inc easing in insu ance, mainly due o echnological imp o emen s ha allow cus ome s o explo e o he insu ance p o ide s’ o e s. Gi en his, insu ance p o ide s need o compe e among hem, no only o ge new cus ome s bu , o main ain hei own. This epo esul s om a p ojec de eloped du ing an in e nship in G upo Ageas Po ugal, which has di e en insu ance b ands such as Ociden al Segu os. This p ojec 's main goal was o model, wi h a mon hly pe iodici y, cus ome chu n o his la e ’s Wo ke s’ Compensa ion po olio o imp o e he company’s compe i i eness and, ul ima ely, p o i . Many o he company’s cus ome chu n happens a hei policy enewal ime, whe e he only a iable ha he company de ains con ol o e is he p ice (p emium) a ia ion. Hence, by conside ing he p emium a ia ion and o he ele an p edic i e a iables, he goal was o p edic he p obabili y o a gi en cus ome o chu n, allowing he company o op imize he cu en enewal’s p icing p ocess and maximize his b anch’s p o i . Thus, di e en a iables ha could in luence he company’s cus ome beha io we e collec ed—one o hose was he cus ome ’s loca ion. Gi en he high dimensionali y ha such a iable would ep esen and he small da ase a ailable o modeling, clus e ing analysis is used o c ea e new signi ican (wi h ewe dimensions) cus ome geog aphical a eas. Di e en supe ised lea ning algo i hms we e hen e alua ed acco dingly o hei pe o mance in p edic ing cus ome chu n. The p edic i e models used we e a G adien Boos ing, an Ex eme G adien Boos ing, a Logis ic Reg ession, and a Mul ilaye Pe cep on. Gi en ha he numbe o cus ome s ha enew hei con ac s is much supe io o he numbe o cus ome s who chu n, Syn he ic Mino i y O e sampling Technique (SMOTE) was used o c ea e less unbalanced da ase s (wi h syn he ic samples) and e alua e he impac on he pe o mance o one o he models. Las ly, o gua an ee a success ul in eg a ion o he models in o he enewal’s p icing p ocess, models we e e alua ed acco dingly o he wo business goals. Fi s , by ansla ing he obse ed e alua ion me ics in o p o i . Secondly, by assu ing ha he cus ome ’s p ice elas ici y would be cap u ed, assu ing a mono onic inc easing ela ionship among he policy’s p emium a ia ion and p obabili y o chu n. KEYWORDS Supe ised Lea ning; Classi ica ion; Cus ome Chu n P edic ion; Non-Li e Insu ance; Renewal P ice Elas ici y; Clus e ing; Neu al Ne wo k; Logis ic Reg ession; S ochas ic G adien Boos ing; Ex eme G adien Boos ing INDEX 1. In oduc ion .................................................................................................................. 1 1.1. P oblem S a emen and Objec i e ........................................................................ 1 1.2. Summa y o he P ocess ........................................................................................ 2 2. Theo e ical F amewo k ................................................................................................ 3 2.1. Li e a u e Re iew in Chu n Modeling in Insu ance ............................................... 3 2.2. Machine Lea ning .................................................................................................. 5 2.2.1. Da a Imbalance ............................................................................................... 5 2.2.2. Fea u e Selec ion ............................................................................................ 8 3. Da a............................................................................................................................. 11 4. Me hodology .............................................................................................................. 13 4.1. Business Unde s anding ...................................................................................... 13 4.2. Da a Unde s anding ............................................................................................ 14 4.3. Da a P epa a ion ................................................................................................. 14 4.3.1. Da a Selec ion ............................................................................................... 14 4.3.2. Da a Cleansing .............................................................................................. 15 4.3.3. Fea u e Enginee ing ..................................................................................... 16 4.3.4. Da a In eg a ion and Fo ma ........................................................................ 19 4.4. Modeling .............................................................................................................. 20 4.4.1. Model Selec ion ............................................................................................ 20 4.4.2. Tes Design Gene a ion ................................................................................ 21 4.4.3. Model Cons uc ion and Assessmen ........................................................... 22 4.5. E alua ion ............................................................................................................ 27 4.6. Deploymen ......................................................................................................... 29 5. Resul s and Discussion ................................................................................................ 31 5.1. Resul s ................................................................................................................. 31 5.1.1. C oss-Valida ion Resul s ............................................................................... 31 5.1.2. Model Assessmen ....................................................................................... 31 5.1.3. Model’s E alua ion ....................................................................................... 33 5.2. Discussion ............................................................................................................ 36 6. Conclusions ................................................................................................................. 38 7. Limi a ions and Recommenda ions o Fu u e Wo ks ............................................... 40 8. Bibliog aphy ................................................................................................................ 41 9. Appendix ..................................................................................................................... 47 i 9.1. Appendix A - Backwa d Elimina ion (Logis ic Reg ession Final Fea u e Lis ) ...... 47 9.2. Appendix B - Cons uc ed Models’ Cha ac e is ics ............................................. 48 9.3. Appendix C - Final Model’s Fea u e Lis .............................................................. 49 4 Model Used in Li e a u e Logis ic Reg ession Y. He e al. (2020) Va eiadis e al. (2015) Sunda kuma & Ra i (2015) Bolancé e al. (2016) Spi e i & Azzopa di (2018) Decision T ee Va eiadis e al. (2015) Sunda kuma & Ra i, 2015) Bolancé e al. (2016) Dola abadi e al. (2017) Spi e i & Azzopa di (2018) Sc iney e al. (2020) Suppo Vec o Machine Y. He e al., (2020) Va eiadis e al. (2015) Sunda kuma & Ra i (2015) Bolancé e al. (2016) Dola abadi e al. (2017) Spi e i & Azzopa di (2018) Sc iney e al. (2020) Neu al Ne wo k Y. He e al. (2020) Va eiadis e al. (2015) Sunda kuma & Ra i (2015) Bolancé e al. (2016) Dola abadi e al. (2017) Sc iney e al. (2020) G adien Boos ing Y. He e al. (2020) Naï e Bayes Classi ie Va eiadis e al. (2015) Dola abadi e al. (2017) Spi e i & Azzopa di (2018) Sc iney e al. (2020) Random Fo es Y. He e al. (2020) Spi e i & Azzopa di (2018) Ex a T ees Classi ie Y. He e al. (2020) Table 2.1 Models used in li e a u e o insu ance chu n modeling 5 2.2. MACHINE LEARNING Unsupe ised Lea ning Unsupe ised lea ning is a o m o machine lea ning ha aims o g oup da a in o segmen s based on simila a ibu es, o na u ally occu ing ends, pa e ns, and ela ionships hidden in he da a (McCue, 2015). S anda d algo i hms used o his kind o ask a e clus e ing, anomaly de ec ion, neu al ne wo ks, and app oaches o lea ning la en a iable models (El Bouche y & De Souza, 2020). The main goal o clus e ing analysis is o segmen he ini e unlabeled da ase in o a ini e and disc e e se o hidden da a s uc u es (Xu & Wunsch, 2005). The esul o his analysis is se e al g oups o he da a named clus e s. These a e subse s o da a g ouped when, acco ding o he s udied c i e ia, ha e simila cha ac e is ics and, when no , sepa a ed in o di e en g oups (Rokach & Maimon, 2005). Mo eo e , wo ypes o clus e ing echniques a e pa i ional clus e ing and hie a chical clus e ing. In he i s , g oups a e c ea ed by pa i ioning he space in o a p e-de ined numbe o subspaces. On he o he hand, in hie a chical clus e ing, he da a objec s a e g ouped in sequence by a hie a chical s uc u e. Supe ised Lea ning Ano he o m o Machine Lea ning is supe ised lea ning. Con a y o unsupe ised lea ning, i uses a da ase ha has al eady been classi ied (labeled) as a basis o p edic ing he classi ica ion o o he unlabeled da a (Talabis e al., 2015). The labeled da ase is a aining se composed o inpu a iables ( ea u es) and an ou pu a iable (label). The used ea u es will in luence he model’s abili y o co ec ly classi y he p edic ed a iable, he ou pu a iable (L. Wang e al., 2021). Supe ised Lea ning can be di ided in o classi ica ion when he label is ca ego ical and eg ession when he label is con inuous. The e o e, he s udied p oblem is a classi ica ion ask since chu n modeling can be ansla ed o a bina y classi ica ion p oblem: assuming 1 when he cus ome chu ns and 0 when he cus ome does no chu n. 2.2.1. Da a Imbalance Fo many eal supe ised lea ning p oblems in ol ing a bina y esponse a iable, da ase s p esen a skewed dis ibu ion, ha ing one class wi h a much lowe ep esen a ion han ano he . When he da ase shows his ype o beha io , ha ing an unde ep esen ed class, he da a is said o be unbalanced (S. Wang & X. Yao, 2012). Resea che s ha e concluded ha his imbalance causes a subop imal classi ica ion pe o mance (Chawla e al., 2004), since classi ie s end o gi e much highe impo ance o he la ge classes. H. He & Ga cia (2009) b ing ha . Usually, classi ie s, when dealing wi h an imbalanced da ase , end o “p o ide a se e ely imbalanced deg ee o accu acy, wi h he majo i y class ha ing close o 100 pe cen accu acy and he mino i y class ha ing accu acies o 0-10 pe cen ”, such consequence can ep esen high cos s o some indus ies so, he e o e, is i al o cons uc a model ha “will p o ide high accu acy o he mino i y class wi hou se e ely jeopa dizing he accu acy o he majo i y class”. In o de o deal wi h imbalanced da ase s, h ee possible app oaches can be aken: da a le el, algo i hmic le el, and combining o ensemble me hods, o which he i s in ol es esampling o educe he class skewness (Yap e al., 2014). Resampling is done ei he by emo ing ins ances om he 6 majo i y class (unde sampling) o adding ins ances o he mino i y class (o e sampling) by using algo i hms such as SMOTE. Syn he ic Mino i y O e sampling Technique Syn he ic Mino i y O e sampling Technique (SMOTE) (Chawla e al., 2002) is one o he mos popula and in luen ial da a p e-p ocessing algo i hms o deal wi h he da a imbalance p oblem (Ga cía e al., 2016). This echnique is an o e sampling app oach, said ha new ins ances om he smalle class a e in oduced in o he da ase . Unlike basic app oaches such as andom o e sampling (ROS), which only duplica es samples om he mino i y class, SMOTE gene a es new syn he ic samples, o e coming he o e i ing caused by app oaches like ROS (Fe nández e al., 2018). The i s s ep o his echnique is o de ine he amoun o o e sampling. He e, i is possible o ei he se up his alue o app oxima e a balanced class dis ibu ion o disco e i ia a w appe p ocess (Chawla e al., 2008). Then, based on k nea es neighbo s and linea in e pola ion ideas, he syn he ic samples a e c ea ed. SMOTE ope a es in he ea u e space a he han in he da a space, each mino i y class sample is conside ed along wi h i s k nea es neighbo s, and he new samples a e in oduced along he line segmen s joining hem (conside ing any/all o he k neighbo s) (Chawla e al., 2002). 2.2.1.1. Model Calib a ion When using such echniques o a i icially ebalancing he da ase o e en by consequence o he da ase ’s cha ac e is ics, he aining and es se s ha e di e en dis ibu ions. This di e ence in aining and es se s dis ibu ion iola es he basic assump ion in machine lea ning ha bo h a e d awn om he same unde lying dis ibu ion (Pozzolo e al., 2015). By iola ing his assump ion, he p edic ions ob ained in he es se will be biased and, he e o e, enhance he need o p obabili y calib a ion o ob ain unbiased p edic ions. Fu he mo e, some me hods end o bias p edic ed p obabili ies by pushing away o close o 0 and 1, enhancing he need o calib a ion (Niculescu-Mizil & Ca uana, 2005). Do mann (2020) e en s a es ha “i should be applied o any model ype as pa o he p edic ion p ocess, be o e p edic ing, c oss- alida ing and making e ec plo s and maps o using p edic ions in any o he p obabilis ic in e p e a ion”. Two model calib a ion me hods ha can be used o co ec hese biased p obabili ies a e Pla Scaling and Iso onic Reg ession. The i s is mo e e ec i e when he dis o ion is sigmoid-shaped, and he la e is a mo e obus me hod ha can co ec any mono onic dis o ion bu , mo e p one o o e i ing (Niculescu-Mizil & Ca uana, 2005). The Pla Calib a ion calib a es p obabili ies by passing he ou pu h ough a sigmoid: 𝑃(𝑦=1|𝑓)= 1 1+𝑒(𝐴𝑓+𝐵) ( 1 ) Whe e 𝑓(𝑥) is he lea ning me hod and pa ame e s 𝐴 and 𝐵 a e es ima ed using G adien Descenden , such ha hey a e a solu ion o he ollowing minimiza ion unc ion: 7 𝑎𝑟𝑔𝑚𝑖𝑛𝐴,𝐵{− ∑𝑦𝑖 log(𝑝𝑖)+(1−𝑦𝑖)log(1−𝑝𝑖) 𝑖} ( 2 ) whe e, 𝑝𝑖= 1 1+𝑒(𝐴𝑓𝑖+𝐵) ( 3 ) On he o he hand, he Iso onic Calib a ion is mo e gene al gi en ha he only es ic ion is ha he mapping unc ion is iso onic (Niculescu-Mizil & Ca uana, 2005). The basic assump ion o Iso onic Reg ession (on which he model is based) is ha : 𝑦𝑖=𝑚(𝑓𝑖)+𝜖𝑖 ( 4 ) Whe e, 𝑦𝑖 a e he ue labels, 𝑓𝑖 he model’s p edic ions, 𝑚 a mono onic inc easing (iso onic) unc ion. The goal is o ind 𝑚, using he ue labels and model’s p edic ions as a aining se , such ha : 𝑚=𝑎𝑟𝑔𝑚𝑖𝑛𝑧∑(𝑦𝑖−𝑧(𝑓𝑖))2 ( 5 ) One algo i hm ha can be used o ind a s epwise cons an solu ion o his p oblem is he pai - adjacen iola o s (PAV) algo i hm (Aye e al., 1955). Las ly, one way o assess how well-calib a ed he model is can be h ough a calib a ion plo . He e, “a se o p edic ions o a bina y ou come is well calib a ed i he ou comes p edic ed o occu wi h p obabili y p do occu abou p ac ion o he ime, o each p obabili y p ha is p edic ed” (Naeini e al., 2015), which can be ansla ed in o a s aigh line om (0,0) o (1,1). Gi en his, in he calib a ion plo , he x-axis ep esen s he a e age p edic ed p obabili y in each bin. The y-axis ep esen s he obse ed ac ion o samples in he bin whose eal labels a e posi i e. The obse ed cu e is hen compa ed wi h he s aigh line 𝑦=𝑥. 2.2.1.2. Th eshold-Mo ing Me hod A echnique ha should be conside ed when dealing wi h class imbalance is changing he decision h eshold (model’s con inuous ou pu cu -o ) and adap ing i o a pe o mance me ic. The main di e ence be ween ebalancing (using echniques like SMOTE) and h eshold-based me hods is ha he la e elies on manipula ing he con inuous ou pu o a lea ned model ins ead o elying on da a p e-p ocessing be o e he lea ning happens (Collell e al., 2018). P o os (2008) e en s a es ha “The bo om line is ha when s udying p oblems wi h imbalanced da a, using he classi ie s p oduced by s anda d machine lea ning algo i hms wi hou adjus ing he ou pu h eshold may well be a c i ical mis ake”. The h eshold mo ing me hod uses he o iginal aining se o ain and unes o shi s he decision h eshold by adap ing i o a pe o mance me ic. One possible app oach is o use he ROC e alua ion p ocedu e and mo e om whe e misclassi ica ions a ain hei maximum on he posi i e class o he poin whe e he maximum in he nega i e class is a ained, selec ing he poin whe e he cu e a ains i s maximum (H. He & Ga cia, 2009). 8 2.2.2. Fea u e Selec ion A well-known p oblem is he “cu se o ini e sample size”, o which he ela ionship among he numbe o samples a ailable and he ea u es conside ed o modeling needs o be conside ed (Jain & Chand aseka an, 1982). Each new ea u e in oduced o he model will ep esen a new dimension. The highe he dimensionali y, he spa se he da ase becomes and, hus, lowe he ea u e space co e age (Ve leysen & F ançois, 2005). Consequen ly, he p oblem’s complexi y apidly g ows wi h he in oduc ion o mo e dimensions. Bellman (1966) in oduced he e m Cu se o Dimensionali y o explain such phenomena. The e o e, and gi en he nowadays exis ing high-dimensional da a, ea u e selec ion is one o he essen ial echniques in da a p ep ocessing by elimina ing i ele an , edundan , o noisy ea u es (Kalousis e al., 2007). Pe o ming such echniques allows as e algo i hms and, besides imp o ing p edic i e powe , also imp o es comp ehensibili y (Kuma & Minz, 2014). I is possible o b oadly classi y ea u e selec ion me hods in o il e and w appe me hods, whe e he i s anks ea u es based on s a is ical measu es, independen ly o he lea ning algo i hm (Koha i & John, 1997). One example o his ype o me hod is o use as impo ance measu e (sco e) he a iable’s co ela ion wi h he a ge . On he o he hand, w appe s e alua e each candida e subse o ea u es' impac on a pa icula lea ning algo i hm. The la e app oach usually allows achie ing be e esul s gi en hei close in e ac ion wi h he classi ie (El Aboudi & Benhlima, 2016). Some examples o w appe me hods a e o wa d selec ion, backwa d elimina ion, and ecu si e ea u e elimina ion. The i s keeps adding new ea u es which imp o e he model pe o mance un il no o he ea u e espec s he c i e ia. The second wo ks simila ly bu opposi ely, s a s wi h all ea u es, and emo es he leas signi ican e ec ha does no mee he model’s s aying c i e ia un il all ea u es a e signi ican . In bo h ( o wa d and backwa d elimina ion), he ea u e s ays in he model once added/ emo ed. Las ly, ecu si e ea u e elimina ion is simila o he o wa d selec ion me hod. In his case, e ec s a e added and emo ed in o he model such ha one o mo e backwa d elimina ion s eps can happen a e a o wa d selec ion s ep (Bu sac e al., 2008). Ano he amily de i ed om he wo p e ious me hod amilies ( il e and w appe s) a e embedded me hods. These me hods combine he classi ie de elopmen wi h he sea ch o he op imal subse o ea u es, cap u ing dependencies a a lowe compu a ional cos han w appe s bu (like w appe s) also ha e a isk o o e i ing (Seijo-Pa do e al., 2017). Examples o embedded me hods a e Lasso and Ridge eg ession, which ha e buil -in ea u e selec ion me hods ha employ L1 and L2 egula iza ion ( espec i ely). 2.2.2.1. Ensemble Lea ning o Fea u e Selec ion Ensemble Lea ning is a ype o lea ning whe e mul iple models a e ained and combined o sol e he same p oblem (Polika , 2006). Ensemble Lea ning is based on he assump ion ha combining he solu ion o mul iple expe s is be e han using he solu ion o a single one. In his way, a se o hypo heses is cons uc ed a he han using only one single hypo hesis o explain he da a, making i possible o educe bias and a iance om he lea ning algo i hms (Die e ich, 2002). Al hough usually employed o imp o e classi ica ion esul s, i is also possible o use ensemble lea ning as a ea u e selec ion echnique. Combining mul iple ea u e selec ion me hods (ins ead o elying on 9 jus one) makes i possible o a ain mo e obus ea u e subse s, showing a g ea p omise o high- dimensional da ase s wi h small sample sizes (Saeys e al., 2008). The app oach bene i s om di e si y and con ol o a iance and is possible o employ in wo ways: one is o use he same algo i hm o e ie e he ea u es’ impo ance using di e en subse s o he da a (da a pe u ba ion / homogeneous) and, o he , is o use di e en ea u e selec ion echniques in he same da ase ( unc ion pe u ba ion / he e ogenous) (Chiew e al., 2019). Said his, di e en le els can be a ied and shall be chosen when employing ensemble lea ning in ea u e selec ion, ollowing Bolón-Canedo & Alonso-Be anzos (2019) can be de ined as ollows: • Da ase Le el: use di e en subse s o da a (o no ) • Fea u e Le el: use di e en subse s o ea u es (o no ) • Lea ne Me hod Le el: use o design di e en lea ning algo i hms (o no ) • Combina ion Le el: Use o design di e en combina ion/agg ega ion me hods • Th eshold Le el: use o di e en h esholding me hods (in case o using anke me hods) Bolón-Canedo & Alonso-Be anzos (2019) also e i ied ha he e ogeneous ea u e selec ion ensembles a e mo e commonly used han homogeneous ones. Howe e , i is possible o ob ain ei he a ea u e subse o a ea u e anking in bo h cases, depending on he ype o ea u e selec o s. Fo he la e , a h eshold me hod needs o be de ined. When using ea u e anke s, se e al ea u e anking algo i hms o he ensemble a e combined, c ea ing a inal anked lis o he ea u es, gi en he ea u es’ ele ance o p edic ion. The app oach o combining he ensemble membe s' esul s ( anks) has di e en p oposals in he li e a u e, om simple o mo e complex solu ions (Seijo-Pa do e al., 2015). Some o he mos popula s aigh o wa d me hods o combine such anks a ibu ed o each ea u e a e: minimum (bes ) ank, median ank, a i hme ic mean ank, and geome ic mean (Bolón-Canedo & Alonso-Be anzos, 2019). Gi en he h eshold me hod decision, he mos common app oach is de ining a ixed pe cen age o he op ea u es, bu his pe cen age depends on he used da ase (Bolón-Canedo & Alonso-Be anzos, 2019). The e o e, his echnique is no op imal since i is p one o o e s a ing o unde s a ing his cu - o alue (Chiew e al., 2019). Pe mu a ion ea u e impo ance Pe mu a ion ea u e impo ance (PFI) is a model-agnos ic ea u e selec ion me hod. The e o e, i is possible o pe o m a he e ogeneous app oach by using his algo i hm as a base lea ne o he ensemble and in oduce a iabili y by using di e en base models in he algo i hm. This pe mu a ion ea u e impo ance measu emen was i s ly in oduced o andom o es s in 2001 (B eiman, 2001) bu can be used in any model. PFI assesses he a iable’s impo ance o he gi en model when i s ela ionship wi h he a ge is b oken by obse ing he dec ease in he model’s sco e when andom noise eplaces a a iable (in oduced by andomly shu ling he a iable’s alues) (McGo e n e al., 2019). The e o e, i is possible o unde s and how impo an a gi en ea u e is o he model’s abili y o p edic he a ge co ec ly. 10 PFI’s ope a ion me hod makes i sensible o co ela ed ea u es (S obl e al., 2008) and is, he e o e, essen ial o add ess his issue p ima ily. Howe e , PFI has ad an ages such as obus ness no o bias he measu es a o ing high ca dinally ea u es o e bina y ea u es. 11 3. DATA Gi en he small dimensions o he Wo ke s’ Compensa ion po olio, he main goal in he da a collec ion phase was o ex ac as much da a as possible. Mo eo e , o gua an ee ha he collec ed in o ma ion uly cap u es he eali y, a 3-mon h aging is equi ed. The e o e, he da a ex ac ion p ocess was di ided in o wo phases. A i s phase, a he p ojec ’s beginning, in which i was possible o ex ac da a ega ding he enewals and chu ns om Janua y 2017 un il Sep embe 2020, used o he models’ cons uc ion. Then, a second phase, du ing he Model’s Assessmen phase, o be used as a es se . This new da ase con ained da a om Oc obe 2020 un il Feb ua y 2021, including Janua y, whe e mos o he policies enew hei con ac s. Bellow, he Da a Ex ac ion phases can be seen in Figu e 3.1. Gi en he p oblem con ex , he goal was o iden i y olun a y chu ne s who based hei decision on p ice a ia ion. In his way, nei he in olun a y chu ns no cancela ions ou side he de ined enewal window (explained in sec ion 4.1) a e included in he analysis. Mo eo e , o gua an ee ha he in o ma ion ex ac ed would cap u e he e en ha was decided o be s udied - olun a y chu ne s ha ha e chu ned by consequence o he p emium a ia ion - some es ic ions ha e been applied o he da a: • No include chu ns ha a e a consequence o , o example, bank up o he company. • No include policies om employees o he company • No include empo a y policies • No include he g oup o policies lagged by he company ha ecei e di e en enewal’s p icing p ocesses om emaining cus ome s (cus ome s wi h highe o al p emiums and policies iden i ied o go h ough a p uning p ocess) • No include policies ha cancel o issue hei exi be o e he de ined enewal window • No include policies o which i was no possible o ex ac hei p emium alue • No include annui ies wi h a du a ion o less han one yea Policies ha ha e canceled hei con ac s o acqui e a new policy in he company (policy cannibaliza ion) we e s ill conside ed o analysis, since i is possible ha he cus ome has cancelled due o he p ice inc ease. Usually, cus ome s acqui e a new policy o ob ain a lowe p ice main aining he same bene i s. Figu e 3.1 Di e en da a ex ac ion phases 12 Wi h such cons ain s applied, i was possible o ex ac a da ase wi h he policies' annui y in o ma ion be o e he enewal da e in which, o one yea , a pa icula policy only appea s once. Rega ding he collec ed a iables, his decision was made conside ing he in o ma ion p esen in he li e a u e and sugges ions made by he p ojec ’s s akeholde s. The collec ed a iables and hei desc ip ion we e no included in his epo . Howe e , i is possible o esume he a iable’s in o ma ion in o he ollowing ca ego ies: • Ca ego y 1: Value paid o insu ance • Ca ego y 2: Cus ome beha io • Ca ego y 3: Cus ome awa eness o he inc ease • Ca ego y 4: Cus ome ’s/P oduc ’s cha ac e is ics • Ca ego y 6: Cus ome ’s claims/cos and i s managemen by he company • Ca ego y 6: Posi ion o Ociden al Segu os in Ma ke • Ca ego y 7: Cus ome ’s in e ac ion channel wi h he company • Ca ego y 8: Cus ome ’s loyal y • Ca ego y 9: Cus ome ’s geog aphical loca ion cha ac e is ics • Ca ego y 10: Pandemic si ua ion • Ca ego y 11: Discoun s Gi en ha he model needs o deli e p edic ions 60 days be o e he policy’s enewal da e (also explained in sec ion 4.1), all he a iables collec ed had o conside his ision on ime. Mos a iables main ain cons an along wi h he annui y. Howe e , o he s, such as he employees’ o al sala y associa ed wi h he con ac , usually a y du ing his pe iod. The e o e, he ex ac ed a iables we e designed o ga he he co ec ision, wi h he end o annui y s anding o he 60 days be o e i . As men ioned be o e, he mos c ucial a iable o measu ing i s impac on chu n is he p emium a ia ion ob ained in he enewal’s p icing p ocess. Howe e , his p emium a ia ion also depends on he con ac ’s employees’ o al sala y, which may a y om one annui y o ano he o e en in he 60 days window. In his way, i was decided ha he p emium a ia ion should only be measu ed, excluding he a ia ion o his a iable. The e o e, only he con ac ’s a i a ia ion was conside ed when measu ing he p emium a ia ion. 13 4. METHODOLOGY 4.1. BUSINESS UNDERSTANDING As men ioned be o e, his p ojec is applied o he Wo ke s’ Compensa ion po olio om Ociden al Segu os. To gua an ee he p ojec ’s success, he i s objec i e was o unde s and his business line and he p ojec objec i es. Fu he , his knowledge has been con e ed in o da a mining goals, and a p ojec plan has been designed. Wo ke s’ Compensa ion insu ance has been manda o y in Po ugal since 1993 o Thi d-Pa y companies’ wo ke s and ex ended o sel -employed wo ke s in 1997. This ype o insu ance has only one co e age ha co e s he isk o acciden s in he employees’ wo kplace o on hei way om/ o home. This co e age assu es he legally needed bene i s by consequence o any o he in ol ed employees’ acciden s, co e ing all needed expenses o ensu e he wo ke ’s o al eco e y (o compensa ion in case o disabili y o dea h). The paid p emium o his insu ance depends no only on he o al secu ed employees’ sala y (sum insu ed), which may a y du ing he yea bu also on he con ac ’s a e s ipula ed by he company (which can a y yea ly in he enewal’s p icing s a egy). The enewal’s p icing p ocess p oceeds app oxima ely wo mon hs in ad ance (gi en he cus ome ’s enewal da e). By law, he insu ance company mus send he enewal le e in o ming he cus ome abou he nex annui y’s p emium 30 days in ad ance. Gi en his, he ma ke de ined ha his enewal le e should be sen 45 days in ad ance, so his is (usually) when he cus ome is awa e o he a ia ion in he p emium (con ac ’s a e). The e o e, based on his business knowledge, i was de ined ha only a cancella ion ha happens in a window o 45 days p io and 50 days a e he end o he policy’s enewal da e would be classi ied as chu n. Mo eo e , he annui y’s p emium a ia ion is de e mined 60 o 45 days be o e he end o he policy’s enewal da e, so all a iables ex ac ed espec his ime window (using he a iable’s ision 60 days be o e he policy annui y end da e). Ociden al Segu os has a bancassu ance channel and de ains 1.5% o he Wo ke s’ Compensa ion insu ance ma ke sha e. The company’s po olio can be segmen ed in o h ee main cus ome g oups: • Housekeepe s – Besides housekeepe s, which is he main componen o his g oup, o he domes ic employees a e in eg a ed, as ga dene s, o example. • Sel -Employees – In gene al, a e small companies whe e he insu ed en i y is he wo ke himsel . • Thi d-Pa y - Comp ises companies om di e se dimensions (mic o/small, medium, and big) insu ing hei employees. Al hough mos o Ociden al Segu os’ cus ome con ac s a e om he Housekeepe s’ segmen , Thi d- Pa y ep esen s almos 85% o he po olio’s o al annual p emium. Such beha io is explained by he ac ha , on a e age, he o al secu ed employees’ sala y o his segmen is much highe han o he o he segmen s. Rega ding claims, he mos signi ican equency o claims is obse ed o he Sel -Employees’ segmen . Howe e , was he Thi d-Pa y segmen o which, a he ime, he lowes p o i abili y was obse ed. 20 Yeo-Johnson ans o ma ion is a powe ans o ma ion amily wi h simila p ope ies as Box-Cox ans o ma ion, bu well de ined in he whole eal line, hus app op ia e o educe skewness and app oxima e no mali y wi hou such limi a ion (Yeo, 2000). Box-Cox ep esen s a amily o powe ans o ma ions (like squa e oo , log, o in e se ans o ma ions - which a e conside ed a way o mee he no mali y assump ion) ha easily ind he op imal powe ans o ma ion o a gi en a iable o s abilize i s a iance (T. Zhang & Yang, 2017). Ne e heless, he o iginal Box-Cox (Box & Cox, 1964) ans o ma ion is only alid o a posi i e x, con a y o Yeo-Johnson ans o ma ion. On he o he hand, i has also been shown ha Fea u e Scaling gene ally imp o es he pe o mance o classi ica ion algo i hms (Bollegala, 2017). Wha gene ally leads o such imp o emen is ha ea u es’ alues, na u ally, occupy di e en anges (some anging in housands, o he s in ens) and, ea u es wi h highe anges end o ha e a mo e decisi e ole while aining he model. Fu he mo e, he ea u e’s ela i e alue di e ence is o en mo e in o ma i e han i s absolu e alue (Bollegala, 2017). The e o e, o he han using he da a in i s o iginal ange o m, some scaling me hods we e a emp ed. Such ans o ma ions we e pe o med by es ima ing he scaling pa ame e s using he aining se and, he ea u e scaling me hod is hen applied in bo h aining and es se s. 4.4. MODELING In his phase, he used modeling echniques a e selec ed. Also, a es design is c ea ed o build he di e en models and co ec ly assess hem, unde s anding hei alue o esol ing he p oblem a hand. 4.4.1. Model Selec ion Ini ially, o his p ojec , ou p edic i e models we e selec ed and cons uc ed: a Logis ic Reg ession, a Mul ilaye Pe cep on, a G adien Boos ing model, and an Ex eme G adien Boos ing model. These me hods a e b ie ly explained below. Logis ic Reg ession Simila o linea eg ession, logis ic eg ession (LR) can include one o mul iple co a ia es ( a iables) ha , join ly wi h unknown pa ame e s es ima ed om he da a, p oduce a linea (and con inuous) p edic o . The logi unc ion is used o gua an ee ha he ou come a iable alls in a 0-1 ange. Fu he mo e, i is possible o add a egula iza ion e m o he logis ic eg ession, encou aging he i ed pa ame e s o be small and helping p e en o e i ing. Mul ilaye Pe cep on The Mul ilaye Pe cep on (MLP) is a eed- o wa d ne wo k composed o inpu neu ons, ou pu neu ons, and hidden laye s o neu ons be ween hem. Neu al ne wo ks a e a b anch o a i icial in elligence inspi ed in human b ains. He e, nume ous cells called neu ons p ocess in o ma ion in pa allel, linked oge he in a ne wo k by synapses, whe e in elligence is a gued o be encoded. Mo eo e , in a eed- o wa d ne wo k, in o ma ion only mo es o wa d, om he inpu nodes o he hidden nodes and inally ou pu nodes (Zell, 1994), connec ed by weigh s and ou pu signals. A nonlinea ans e unc ion modi ies he neu on’s weigh ed inpu s, also called he ac i a ion unc ion (Njikam & Zhao, 2016). These supe posi ions o nonlinea ans e unc ions allow he mul ilaye pe cep on o app oxima e highly non-linea unc ions (con a y o he logis ic eg ession) and accu a ely gene alize when p esen ed wi h new, unseen da a (Ga dne & Do ling, 1998). 21 G adien Boos ing Model The G adien Boos ing (GB), p oposed by F iedman (2001, 2002), belongs o he amily o boos ing me hods, a ype o ensemble lea ning. In boos ing, a new model is added o he ensemble sequence ained based on he e o o he whole ensemble lea ned so a (Na ekin & Knoll, 2013). Hence, misclassi ied ins ances a e emphasized by ecei ing highe weigh s in he nex i e a ion, which, o example, usually happens o ins ances nea he decision bounda y (Mei & Rä sch, 2003). These weigh s ep esen he impo ance ha he gi en ins ance will ha e in he ollowing base lea ne cons uc ion. The e o e, ins ances ha he p e ious base models ha e shown di icul y unde s anding appea mo e o en in he aining da a (Y. Zhang & Haghani, 2015). The inal model ob ained by he boos ing algo i hm will be a linea combina ion o he se e al base-lea ne s conside ing hei pe o mance on he da ase (G. Wang e al., 2011). Ex eme G adien Boos ing Model The Ex eme G adien Boos ing model (XGBoos ) is an ensemble o Classi ica ion T ee (CART) and a mo e e icien e sion o GB. This algo i hm is e y popula in he Machine Lea ning ield, ha ing i s impac widely ecognized in many machine lea ning and da a mining challenges (Chen & Gues in, 2016). Some o he modi ica ions done o GB a e pa allel aining (which as ens he algo i hm), ou - o -co e compu a ion (allowing ha da a is no loaded in o memo y), and spa se da a op imiza ion (in handling and speed up compu a ion) (Bisong, 2019). Also, besides sh inkage and subsampling (used in GB) as egula iza ion o ms o con ol o e i ing and a ain be e esul s, he XGBoos uses a egula ized objec i e. Mo e in o ma ion abou he XGBoos algo i hm can be ound a Chen & Gues in (2016). 4.4.2. Tes Design Gene a ion In o de o choose he sui able model and model’s hype pa ame e s, each cons uc ed model should be es ed and, he e o e, a es design needs o be de ined. The goal is ha he ob ained esul s while aining/e alua ing he model a e close o he eali y o how each model will pe o m. The e o e, bo h he da ase ’s cha ac e is ics and he way he model would be used we e conside ed. Fi s ly, ega ding he da a, a possible endency in chu n a e ac oss he yea s has been obse ed o he di e en cus ome segmen s in he Da a Unde s anding phase. Also, as men ioned be o e, he policy da a main ains p ima ily s able h ough he yea s in mos measu es. Howe e , signi ican changes usually happen in he policies’ p emium paid (and o he a iables i depends on). Mo eo e , gi en ha each policy and i s da a only appea , a he mos , once in each o he yea s, he e is an annual ( empo al) dependency on he da a. A possible app oach o espec such empo al dependency is o use ime-spli c oss- alida ion. He e, he model’s gene aliza ion abili y is assessed using he a e age o he pe o mance me ics in he c ea ed da a spli s. A he same ime, he empo al dependency is espec ed by always using he las block o da a as alida ion. Gi en he applied inc ease in p emium, he model should p edic he p obabili y o chu n o he cus ome s expec ed o enew each mon h in he yea and he e o e deli e mon hly p edic ions. Gi en hese model pu poses, wel e spli s we e c ea ed (one spli o each mon h in he yea ), using he alida ion se o he spli ’s las a ailable mon h, as shown in Figu e 4.4. 22 Figu e 4.4 Rep esen a ion o he designed ime-spli c oss- alida ion 4.4.3. Model Cons uc ion and Assessmen 4.4.3.1. Fea u e Selec ion wi h Ensemble Lea ning As shown, one possible way o pe o m ea u e selec ion is o employ ensemble lea ning, and, o his, he app oach needs o be designed. Gi en he popula i y o he e ogeneous ea u e selec ion me hods, i was decided ha his app oach should be used, using he same subse s o da a and ea u es ac oss he di e en lea ne s. Fo his, many o he sugges ions p esen ed by Shah & Pe e ia ko (2021) we e conside ed and a e desc ibed below. Fi s ly, as o he lea ne me hod le el, many ea u e anke s we e cons uc ed using he PFI algo i hm wi h a di e en base model o gua an ee a iabili y in he solu ions. The selec ed models we e he bes se o base-line models (classi ie s wi hou pa ame e unning) in e ms o ecall ( ue posi i e a e). Said ha , he selec ed models we e he ones ha appea ed o ha e a be e unde s anding o he small class. Then, as o combina ion le el, i has been shown ha se e al echniques a e possible o combine he esul s and c ea e a inal ank, om simple o mo e complex me hods. Some o he mo e s aigh o wa d me hods a e using he median and mean o he esul s as agg ega ion ules. The median o he anks was used o his wo k since i is less sensi i e o ou lie s, being a mo e obus me ic han he mean. Howe e , ea u es p esen ing he same median ank we e a e wa d o de ed by he mean o he obse ed anks. Las ly, o he h eshold le el decision, i has been shown ha he mos common app oaches a e conside ing a ixed h eshold o he op- anked ea u es. Howe e , such a echnique comes wi h a ew downsizes. In his way, once ha ing each o he model’s esul s, ins ead o selec ing s aigh away he op N ea u es (gi en he di icul y o selec ing he p ope alue o N), he used s a egy was sligh ly di e en . A e ha ing he ea u e impac o each ea u e gi en by each o he lea ne s, he o al (sum) ea u e impac can be calcula ed ( o each o he ea u es summing all model’s e u ned impac s), along wi h he cumula i e impac (a e o de ing by he ea u e’s o al (sum) impac ). Then, a gi en 23 a io (𝑟∈ ]0,1[ ) o he o al cumula i e impac is applied and, he numbe o ea u es ha p esen a cumula i e impac lowe han he a io o o al cumula i e impac will be he numbe o ea u es o selec (N) – N will be he cu ing h eshold. Wi h his numbe (N), he Top N bes - anked ea u es a e selec ed, acco ding o he ensemble’s lea ne s’ opinions combina ion ule ea ly de ined. I is impo an o no e ha his echnique will ne e selec a iables a ibu ed wi h ea u e impo ance o ze o o nega i e alues. Howe e , using his app oach, he e is s ill a alue ha needs o be selec ed – he a io (𝑟). This a io and i s alue decision will be explained u he . As has been concluded in he pas , he op imal ea u e size depends no only on he ea u e-label dis ibu ion bu also on he used classi ie (Hua e al., 2005). In his way, since each model is di e en , i was decided o c ea e a mo e model-agnos ic inal a io decision (indi idually o each o he selec ed models). Addi ionally, an o de ed ea u e lis was c ea ed o each baseline model based on he PFI algo i hm esul s using he gi en model only. In his case, o each o he classi ie s is possible o choose he ea u e selec ion me hod ha p o ides be e esul s: using ensemble lea ning o ea u e selec ion o only he w appe me hod - bo h ha ing a common app oach o selec ing he numbe o ea u es om he anked lis . In he end, a ecu si e me hod was implemen ed o selec he inal ea u e lis o each model. The o al cumula i e impac a io (𝑟) will decay un il a speci ic s opping c i e ion is eached ( he a io s a s a 1 and decays in each in e ac ion). Fo each classi ie , he bes ea u e anking me hod (compa ing he ensemble and single model’s anked ea u e lis s), da a scale (scaling only nume ical a iables), and a io a e selec ed conside ing he combina ion’s pe o mance in p edic ing he mino i y class (conside ing he ecall using he designed c oss- alida ion). In his way, he numbe o selec ed ea u es depends on he da ase i sel and he used classi ie . Fea u e Lis C ea ion Gi en his, he used ea u e selec ion echnique e ie es he bes combina ion o : • Fea u e Ranking Me hod – using he ensemble anking o single model PFI’s anking • Scaling – Scaling nume ical ea u es (wi h S anda d Scale o Robus Scale ) o wi hou scaling • Ra io – s a ing a 1 (using all ea u es) and decaying 0.001 in each in e ac ion; The a io s ops dec easing when h ee consecu i e loops do no imp o e he c oss- alida ed sco e (0.001 decay was decided based on he da ase ’s cha ac e is ics). The bes combina ion is he one wi h he bes ecall sco e in c oss- alida ion. Also, i mo e han one combina ion p esen s he same ecall, he one wi h ewe ea u es is e ie ed. Fea u e Lis Op imiza ion A e his, he c ea ed ea u e lis s passed h ough he ollowing app oaches desc ibed below. He e, he inal ea u e anking is also used as an impo ance measu e. 1. T y o add possible good ea u es o each model's lis • A e he Da a Unde s anding phase, o each o he cus ome ’s segmen s, a lis o ea u es is c ea ed wi h ea u es ha , based on his phase, a e belie ed o be good p edic o s o he p oblem a hand. 24 • Fo each o he ea u es ha a e no on he gi en c ea ed lis , hey a e added (indi idually, one a he ime) o he lis o see i he ecall sco e (in c oss- alida ion) imp o es - i yes, should be added o he lis (manually). 2. T y o emo e possible bad/unnecessa y ea u es • A p ima y s a i ied and shu led ain/ es pa i ion is c ea ed o 70/30 ( o speed up gi en he nume ous combina ions o his phase) • A each i e a ion, one o he ea u es is le ou (by in e se o de o impo ance de ined o he gi en ea u es lis in he la e phase) - i he sco e imp o es/main ains, ha speci ic ea u e is le ou o he nex in e ac ion. • A e all he ea u es a e es ed, he p ocess epea s. The p ocess will epea un il no changes a e done - o he gi en i e a ion, no ea u e emo al imp o es/main ains he ecall sco e in he es spli . • Gi en he emo ed ea u es o he p ima y ain/ es pa i ion, each o he ea u es emo ed is es ed o be ou in he model (indi idually, each ea u e a he ime), and i s c oss- alida ed ecall sco e is e u ned. • I he a e age ecall sco e in he alida ion se s imp o es/main ains, he ea u e is excluded (manually). 3. T y o eplace nume ical a iables o i s powe ans o ma ion • In he Da a P epa a ion phase, Yeo-Johnson Powe ans o ma ion was used o app oxima e he a iable’s dis ibu ion o no mal • Each i e a ion eplaces one o he ea u es by hei ans o ma ion (by impo ance o de / anking). I he sco e imp o es, hen i s ans o ma ion o he nex in e ac ion eplaces ha speci ic ea u e. • A e all he ea u es a e es ed, he p ocess epea s. The p ocess will epea un il no changes a e done - no ea u e eplacemen imp o es he a e age ecall sco e o alida ion se s. • The selec ed ans o med ea u es a e manually eplaced in he inal ea u e lis . 4.4.3.2. Logis ic Reg ession’s Fea u e Selec ion wi h Backwa d Elimina ion Fo he logis ic eg ession’s ea u e selec ion, an un egula ized Logis ic eg ession was cons uc ed using all ea u es (a e he Da a P epa a ion phase), and Backwa d Elimina ion as ea u e selec ion me hod was pe o med. As men ioned be o e, as he name sugges s, he leas signi ican ea u es a e i e a i ely emo ed un il no o he e ec mee s he speci ied le el o emo al (Bu sac e al., 2008). The emo al c i e ia a e based on he Wald es o indi idual pa ame e s, using a signi icance le el o 0.05 ( he p- alue cu -o was de ined as 0.05). He e, he null hypo hesis ha he coe icien o he independen a iable is equal o ze o is es ed e sus an al e na i e hypo hesis ha he coe icien is nonze o, which can be w i en as: 𝐻0:𝛽𝑥=0 𝑣𝑠 𝐻1:𝛽𝑥≠0 (Fo ho e e al., 2007). Besides he collec ed a iables, some i e a ions and polynomial e ms we e in oduced and es ed o hei signi icance in he model. Such in e ac ions among a iables we e added whene e he e we e suspicions ha he impac o one a iable in he ou come also depended on ano he a iable’s alue. Polynomial e ms we e added o co e / es cases whe e hei ela ionship wi h he a ge is no linea . 25 Las ly, T- es was used o es possible (ca ego ical) a iable agg ega ions, es ing i hei model pa ame e s we e signi ican ly di e en o no (𝐻0:𝛽𝑥𝑖− 𝛽𝑥𝑗=0 𝑣𝑠 𝐻1:𝛽𝑥𝑖− 𝛽𝑥𝑗≠0). Fo he cases in which he null hypo heses we e no ejec ed, he agg ega ion was conside ed by summing bo h bina y a iables since, o he es ed cases, he posi i e e en is exclusi e (ne e happens simul aneously in bo h). Be o e his me hod was applied, ea u es ha p esen ed low a iances we e emo ed o p e en mul icollinea i y. Such ea u es can lead o a singula ma ix (wi h de e minan equaling o ze o), which means ha he ma ix has no in e se, making i impossible o es ima e he eg ession pa ame e s. The esul ing se o ea u es and co esponding summa y s a is ics can be ound in Appendix A. 4.4.3.3. Model Cons uc ion Thi d-Pa y is he segmen o which he highes pe cen age o o al p emium and he highes chu n a e was obse ed. The e o e, he Thi d-Pa y segmen was he i s modeled segmen and o which he esul s will be p esen ed. As said be o e, c oss- alida ion was pe o med o assess he models and choose he co ec pa ame e s. Besides pa ame e uning, in which mul iple pa ame e combina ions we e es ed o ind he pa ame e s ha enhanced he esul s, o he s eps we e aken in his phase. Fo each model cons uc ed, he ollowing desc ibed p ocesses we e p oceeded. Fea u e Lis and Da a Scale Choice He e, he combina ion o ea u e lis and da a scale (S anda d Scale , Robus Scale , o no scale ) ha gua an eed he bes esul s o each model is selec ed and used in he subsequen phases. Bellow, an example o G adien Boos ing model (wi hou pa ame e uning) o he Thi d-Pa y’s da ase , compa ing i s c oss- alida ed esul s using all da ase ’s ea u es and he selec ed lis , is p esen in Table 4.2. Fo his case, i is possible o obse e ha bo h esul s a e simila , indica ing ha he lowes numbe o ea u es should be conside ed only. # Fea u es Speci ici y T ain Speci ici y Valida ion AUC T ain AUC Valida ion Recall T ain Recall Valida ion 84 100% +/- 0p.p. 99.7% +/- 0p.p. 52.2% +/- 0p.p. 50.2% +/- 1p.p. 4.4% +/- 1p.p. 0.6% +/- 1p.p. 37 100% +/- 0p.p. 97.2% +/- 9p.p. 52.1% +/- 0p.p. 50.8% +/- 2p.p. 4.2% +/- 1p.p. 4.4% +/- 13p.p. Table 4.2 G adien Boos ing’s ea u e lis selec ion (wi hou pa ame e uning) Da a P e-P ocessing using SMOTE As has been men ioned, one o he possible app oaches o deal wi h class imbalance is a da a-le el app oach, whe e unde sampling o o e sampling can be used. Gi en he small dimensions o he da ase s, using an unde sampling echnique was no app op ia e and, he e o e, an o e sampling echnique was a emp ed. Gi en SMOTE’s popula i y due o i s simplici y and obus ness, his echnique gene a ed new eco ds in o he aining se . 26 Al hough i is shown ha non- andom sampling echniques can highly imp o e he classi ie ’s pe o mance, ha ing a 1:1 dis ibu ion ( he posi i e a e equaling he nega i e a e) migh be un a o able (Fo man & Cohen, 2004). The e o e, smalle a ios o he mino i y class we e a emp ed ia a w appe p ocess so ha he co ec new a io o he aining se s’ mino i y class o he model in ques ion could be chosen. Below, in Table 4.3, an example o he SMOTE’s pa ame e choice ha con ols he new pe cen age o chu n obse ed in he da ase is p esen ed. I is impo an o no e ha his esampling is only done in he aining se , so he eali y in which he model will pe o m can be espec ed. Chu n Ra e in T ain Speci ici y T ain Speci ici y Valida ion AUC T ain AUC Valida ion Recall T ain Recall Valida ion o iginal 100% +/- 0p.p. 98.5% +/- 2p.p. 68.8% +/- 3p.p. 51.3% +/- 2p.p. 37.6% +/- 5p.p. 4.0% +/- 5p.p. 50% 98.9% +/- 0p.p. 23.4% +/- 11p.p. 95.9% +/- 0p.p. 54.4% +/- 5p.p. 92.8% +/- 0p.p. 85.3% +/- 11p.p. 33% 99.5% +/- 3p.p. 32.1% +/- 29p.p. 92.3% +/- 0p.p. 54.7% +/- 0p.p. 85.1% +/- 1p.p. 77.2% +/- 18p.p. 23% 99.8% +/- 0p.p. 43.5% +/- 26p.p. 87.2% +/- 1p.p. 54.8% +/- 1p.p. 74.5% +/- 1p.p. 66.1% +/- 27p.p. 17% 99.8% +/- 0p.p. 55.3% +/- 31p.p. 80.8% +/- 1p.p. 55.4% +/- 4p.p. 61.7% +/- 2p.p. 55.6% +/- 30p.p. 13% 100% +/- 0p.p. 63.3% +/- 26p.p. 75.6% +/- 2p.p. 54.9% +/- 3p.p. 51.2% +/- 3p.p. 46.5% +/- 29p.p. 12% 100% +/- 0p.p. 65.8% +/- 32p.p. 74.4% +/- 1p.p. 53.8% +/- 4p.p. 48.8% +/- 2p.p. 41.9% +/- 32p.p. 11% 100% +/- 0p.p. 67.1% +/- 33p.p. 73.3% +/- 2p.p. 53.0% +/- 3p.p. 46.6% +/- 3p.p. 38.9% +/- 35p.p. Table 4.3 Example o Ex eme G adien Boos ing's - SMOTE's pa ame e choice (wi hou pa ame e uning) Model Calib a ion & Th eshold Tuning As shown be o e, o ob ain good models (in e ms o p edic ion p obabili ies and p obabili ies’ classi ica ion), he model’s ou pu p obabili ies should be calib a ed, and he bes decision h eshold o classi y he model’s esul in o chu n o enewal should be de ined. Rega ding he p obabili y calib a ion, he model p edic ed p obabili ies and he ue alues a e used o assess each model’s calib a ion pe o mance using he calib a ion plo . Then, mul iple copies o he model, using k- old c oss- alida ion, a e i ed and, he p obabili ies p edic ed by hese models a e hen calib a ed on he hold-ou se s. As calib a ion me hods, Pla and Iso onic Calib a ion we e es ed. Fo his, he new p edic ed p obabili ies a e assessed by plo ing he calib a ion cu e and compa ing i wi h bo h, he pe ec calib a ion line, and esul s be o e calib a ion. 27 Fo he choice o he h eshold, he h eshold wi h he op imal balance be ween alse posi i e (speci ici y) and ue posi i e a e ( ecall), his is he op imal h eshold o he ROC cu e, was loca ed. Wi h his knowledge, he me ics in c oss- alida ion a e op imized. In o de o accomplish his, he ue posi i e a e o ecall (TPR) and ue nega i e a e o speci ici y (TNR) a e compu ed o he p edic ions using a se o h esholds ( his can also be used o c ea e a ROC Cu e plo ). Then, he geome ic mean (G-mean), which o mula is p esen ed below, ep esen s he balance be ween bo h sco es ( he highe , he be e ). 𝐺𝑀𝑒𝑎𝑛= √(𝑇𝑃𝑅∗𝑇𝑁𝑅) ( 6 ) This me ic is obse ed o he di e en a emp ed h esholds, and he h eshold ha maximizes his me ic, his is, ha op imizes he balance be ween TPR and TNR, is selec ed. Then, he op imal h esholds ob ained ac oss he di e en spli s o he designed c oss- alida ion a e a e aged in o a inal op imal h eshold and a e conside ed. 4.5. EVALUATION In he E alua ion phase, he da a mining esul s we e assessed compa a i ely wi h he business success c i e ia. The p ocess was also e iewed, and he inal model was chosen o deploymen . The main goal o his p ojec is o imp o e he p o i ma gin o Ociden al’s Wo ke s’ Compensa ion b anch by e aining cus ome s ha ha e in en ions o lea e gi en he p emium a ia ion su e ed in he eno a ion p ocess. A p o i es ima ion analysis was pe o med o ensu e ha he inal model would allow such an inc ease in he company’s e enue. Fo his, some supposi ions we e cons uc ed based on he b anch’s business knowledge and a e bellow explained. The esul s o his analysis a e based on each model’s con usion ma ix esul s and o he business me ics. • Re enue i he model p edic s co ec ly ha he cus ome enews (T ue Nega i es - TN) 𝐺𝑎𝑖𝑛𝑇𝑁 =𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚+∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚 ( 7 ) I he model p edic s co ec ly ha he cus ome will enew, he p oposed p emium a ia ion will main ain, e aining bo h cus ome p emium and p emium a ia ion. • Re enue i he model p edic s ha cus ome chu ns, bu cus ome enews (False Posi i es - FP) 𝐺𝑎𝑖𝑛𝐹𝑃=𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚 ( 8 ) I he model p edic s ha he cus ome will chu n, hen he e is a p emium a ia ion dec ease (o e en emo al). Gi en ha i is di icul o es ima e he p emium a ia ion o hose cases, he wo s -case scena io (no inc ease) was conside ed. • Re enue i he model p edic s co ec ly ha he cus ome chu ns (T ue Posi i es - TP) 𝐺𝑎𝑖𝑛𝑇𝑃 =𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚∗𝑟𝑒𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑟𝑎𝑡𝑖𝑜 ( 9 ) 28 The ue posi i es a e he cus ome s ha , i no model exis ed, would be los . Howe e , hese cus ome s can be p ese ed by co ec ly iden i ying hei in en ions and applying con ingency measu es (in his case, he dec ease in he p emium a ia ion). Ne e heless, hese cases a e simila o he False Posi i e (FP) cases, whe e he model p edic s ha hese cus ome s will also chu n. Thus, he e is s ill an inc ease in p emium ha is no accoun ed o in his es ima ion. Besides his, i is also obse able ha i is no possible o p ese e all isk by applying his con ingency measu e: cus ome s s ill chu n when he e is no p emium a ia ion (∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚=0%). A possible explana ion o such beha io is a be e p emium p oposal in he compe i ion ha he company canno ma ch. The e o e, i was decided o apply a e en ion a io, conside ing ha only 𝑋% o a ge ed chu ns can be e e sed h ough he applied con ingency measu e. Gi en ha , such alue needed o be es ima ed. I was obse ed ha , by yea , 13% o o al chu n in he Thi d-Pa y segmen (wi h a s anda d de ia ion o 2 p.p.) happens o a 0% p emium a ia ion. The e o e, i was decided ha such alue should be ounded up conside ing i s s anda d de ia ion and, as well, gi e a 100% ma gin as a sa e y measu e. Thus, conside ing he e en ion a io as: 𝑟𝑒𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑟𝑎𝑡𝑖𝑜𝑇ℎ𝑖𝑟𝑑𝑃𝑎𝑟𝑡𝑦=(1−2(0.13+0.02))=0.7 ( 10 ) • Re enue i he model p edic s ha cus ome enews bu cus ome Chu ns (False Nega i es - FN) 𝐺𝑎𝑖𝑛𝐹𝑁 =0 ( 11 ) Those a e he cases whe e he model canno o esee he cus ome ’s chu n, so hey a e los , ha ing, o no ha ing a model. The e o e, nei he he cus ome ’s p emium no p emium a ia ion is e ained. Ha ing es ima ed he di e en gains ha each o he con usion ma ix’s measu es, he inc ease in o al e enue is gi en by he di e ence be ween he a e age e enue wi h he model implemen ed (To Be) and wi h no model (As Is): 𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒 𝑅𝑒𝑣𝑒𝑛𝑒=𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝑇𝑜 𝐵𝑒− 𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝐴𝑠 𝐼𝑆 ( 12 ) whe e, 𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝐴𝑠 𝐼𝑆=#𝑅𝑤𝑙𝑠∗(𝐴𝑣𝑔 𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚+ 𝐴𝑣𝑔 ∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚) ( 13 ) and, 𝐴𝑣𝑔 𝑅𝑒𝑣𝑒𝑛𝑢𝑒 𝑇𝑜𝐵𝑒= 𝑻𝑵𝑹 ∗ #𝑅𝑤𝑙𝑠 ∗ 𝐺𝑎𝑖𝑛𝑇𝑁 + 𝑭𝑷𝑹 ∗ #𝑅𝑤𝑙𝑠 ∗ 𝐺𝑎𝑖𝑛𝐹𝑃 + 𝑻𝑷𝑹 ∗ #𝐶ℎ𝑛𝑠∗ 𝐺𝑎𝑖𝑛𝑇𝑃 ( 14 ) 29 Fo each cus ome segmen , he #𝑅𝑤𝑙𝑠 is he numbe o e i ied enewals in a gi en yea , he #𝐶ℎ𝑛𝑠 he numbe o e i ied chu ns in a gi en yea , 𝑻𝑵𝑹 he ue nega i e a e and, 𝑭𝑷𝑹 and 𝑻𝑷𝑹 as alse and ue posi i e a es espec i ely. Using he a e age me ic’s esul s (𝑻𝑵𝑹,𝑭𝑷𝑹,𝑻𝑷𝑹) ob ained in each o he models and he obse ed alues (#𝑅𝑤𝑙𝑠 , #𝐶ℎ𝑠, 𝐴𝑣𝑔 𝐶𝑙𝑖𝑒𝑛𝑡′𝑠 𝑃𝑟𝑒𝑚𝑖𝑢𝑚, 𝐴𝑣𝑔 ∆ 𝑃𝑟𝑒𝑚𝑖𝑢𝑚) in he las h ee yea s, i was possible o es ima e, o each model, he expec ed a e age inc ease in e enue wi h hei implemen a ion. 𝐸𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒=1 3 ∑(𝐴𝑣𝑔 𝐼𝑛𝑐𝑟𝑒𝑎𝑠𝑒 𝑅𝑒𝑣𝑒𝑛𝑒𝑖) 𝑖 ∈ 𝑌 ( 15 ) whe e 𝑌 a e he las h ee a ailable comple e yea s. 4.6. DEPLOYMENT Then, he model’s deploymen is planned, and i s moni o ing and main enance plan is also designed. A e a ull e iew, he model will inally be in eg a ed in o he company’s enewal p icing s a egy and help he company inc ease he cus ome e en ion a e in he Wo ke s’ Compensa ion b anch. Fo his pa , i was decided o deli e wo di e en au oma ed p ocesses, desc ibed below. In bo h, Py hon is he p ima y ool used. Mon hly P edic ions The i s deli e able aims o deli e mon hly p edic ions o he mon h’s po en ial enewals and is designed as ollows. In a pa icula mon h and yea , he po en ial enewals da a is collec ed using SAS En e p ise Guide. Py hon accesses he inal ou pu able and ec ea es in his da ase all o he necessa y da a p epa a ion s eps. A his s age o he p ocess, he p emium a ia ion ha each policy will su e is unknown. The company wan s o know each cus ome 's eac ion (p obabili y o chu n) o he di e en possible p ice inc eases. The e o e, many p ice inc eases scena ios a e c ea ed, and each policy line will be ec ea ed as many imes as he numbe o di e en p emiums inc eases. The ained model is hen used o p edic he chu n p obabili y o each policy and di e en p emium a ia ion scena ios. Then, a inal able, in which he di e en policies ( ows), possible p emium a ia ion (columns), and cus ome eac ion (p obabili y o chu n as alue), is sen o SAS En e p ise Guide. He e, an op imiza ion p ocess (de eloped by he company) o choose he p emium a ia ion o each policy ha maximizes he global p o i will use he cons uc ed able as inpu . Model Pe o mance Moni o ing The second deli e able aims o deli e a p ocess in which he model’s pe o mance can be con inuously assessed and u he in eg a ed in o a dashboa d. This p ocess is simila o he mon hly p edic ions p ocess desc ibed abo e. The only di e ence is ha bo h he a ge (chu n) and he p emium a ia ion a e al eady known when i uns. The goal is o compa e he model’s pas p edic ions wi h wha happened (i policies ha e been chu ned o no ) and unde s and i s heal h. The e o e bo h, he da a collec ion p ocess (in SAS En e p ise Guide), da a 36 Figu e 5.4 Es ima ed a e age inc ease in e enue o he Ex eme G adien Boos ing model wi h he mono onic cons ain and Logis ic Reg ession 5.2. DISCUSSION As s a ed ea lie , he p ojec had wo business goals: 1. Imp o e he b anches’ p o i by co ec ly iden i ying cus ome s who in ended o chu n (and enew) wi h he p oposed p emium a ia ion. I is c ucial o co ec ly iden i y he cus ome s ha will enew wi h he applied a ia ion as well. The company needs o main ain he p emium a ia ion (inc ease) applied o he policies in mos o he cases since i signi ican ly impac s p o i . 2. The selec ed model needs o be sui able o he enewal’s p icing p ocess. The model aims o cap u e he cus ome s’ p ice elas ici y. The e o e, a mono onic ela ionship be ween he a ge and P emium Va ia ion needs o be assu ed. In e ms o p edic ing chu n, he e we e al eady some signs o he complexi y o he p oblem. Looking a he p edic i e a iables’ co ela ion wi h he a ge , he me ic a iable showing a highe co ela ion is P emium Va ia ion (0.10 conside ing Pea son co ela ion and 0.08 o Spea man co ela ion), and he highe co ela ed non-me ic a iable is Flag Annual Paymen (wi h a C amme ’s V o 0.08). Mo eo e , he da ase ’s dimensions we e small, and o he Thi d-Pa y case, only a ound 9% o ins ances we e classi ied as chu n. Such cha ac e is ics in he da ase can easily lead o p oblems such as o e i ing. The e o e, i is essen ial o main ain he models simple. Conside ing he model’s cons uc ion phase, i was hen possible o obse e ha models pe o med be e , showing ewe signs o o e i ing, when hei complexi y was educed. Fo example, using ewe lea ne s ( o ensemble models) o ewe neu ons/laye s ( o he neu al ne wo ks). On he o he hand, he educ ion o he da a-skewness by including syn he ic samples ( ia SMOTE) in he aining da ase did no imp o e esul s. The esul s we e also posi i ely impac ed by selec ing a sui able p edic ion h eshold (di e en han he de aul alue o 0.5) allowing a sui able ade-o be ween he ue posi i e ( ecall) and ue nega i e a es (speci ici y). Mo eo e , gi en he deployed solu ion s uc u e, he ou pu p obabili ies appea o ha e bene i ed om calib a ion, app oxima ing hem o hei ue alues. In Figu e 5.5, i is 37 possible o obse e he di e ences in he uncalib a ed and calib a ed p edic ions made by he Ex eme G adien Boos ing model wi h he imposed mono onic ela ionship, using calib a ion plo s. Obse ing Figu e 5.5(a), is possible o obse e ha he uncalib a ed model unde - o ecas s, his is he ob ained p obabili ies a e smalle han expec ed. Obse ing Figu e 5.5(b), is possible o obse e ha calib a ed p obabili ies a e close o diagonal line, he e o e sugges ing a be e calib a ed model. By conside ing he ob ained esul s, i is possible o obse e he ollowing: • The cons uc ed Ex eme G adien Boos ing using SMOTE in da a p e-p ocessing has shown he highes le el o o e i ing and he lowes AUC on he alida ion and es se s. • The highes AUC in bo h, c oss- alida ion and a e age o es , is ob ained o he G adien Boos ing (wi h an ad an age o 0.9 p.p. in c oss- alida ion and 1.1 p.p. in a e age o es esul s), ollowed by he Ex eme G adien Boos ing model. By e alua ing he i s business goal (p o i inc ease), on a e age, he bes esul s a e ob ained o he G adien Boos ing model. This model is also he leade when indi idually compa ing he impac ha he achie ed me ics (in c oss- alida ion and a e age o es se s) ha e on he es ima ed a e age inc ease in e enue. Howe e , conside ing he second business goal, only he Logis ic Reg ession model could espec he mono onic ela ionship be ween he P emium Va ia ion and he a ge . Howe e , i was possible o o ce his mono onic ela ionship among bo h a iables using an Ex eme G adien Boos ing. By compa ing his new model o he i s business goal’s leade (G adien Boos ing), i is possible o obse e a dec ease in he c oss- alida ion’s AUC (o 1.1 p.p.) and in he a e age o es ’s AUC (o 1.9 p.p.). Such educ ion also impac s he es ima ed a e age inc ease in e enue, by a ound 11K €. Though, his new model s ill has be e esul s han he cons uc ed Logis ic Reg ession. I is only su passed by he cons uc ed (non-mono onic) G adien Boos ing and (non-mono onic) o iginal Ex eme G adien Boos ing. P edic ed P obabili y Obse ed P obabili y P edic ed P obabili y Obse ed P obabili y Figu e 5.5 Calib a ion plo s o XGBoos Mono onic p edic ions (a) Uncalib a ed model P obabili ies. (b) Calib a ed model p edic ions. (a) (b) 38 6. CONCLUSIONS This epo esul s om a se en-mon h p ojec de eloped in a one-yea in e nship in G upo Ageas Po ugal. Bellow, he main conclusions o his p ojec a e p esen ed. G upo Ageas Po ugal is one o he la ges insu ance p o ide s in Po ugal, wi h di e en b ands such as Ociden al Segu os, which p o ides o e s in he li e and non-li e b anches. This la e o e s a Wo ke s’ Compensa ion insu ance, which is manda o y in Po ugal. The Ociden al’s Wo ke s’ Compensa ion po olio can be segmen ed in o h ee main isk g oups (Thi d- Pa y, Sel -Employees, and Housekeepe s), which ha e di e en enewal’s p icing s a egies and p esen di e en beha io s and a ailable in o ma ion. Nowadays, cus ome s' ease in explo ing he a ailable op ions allows hem o change he insu ance p o ide easily. Gi en he compe i i eness o he insu ance ma ke , companies need o ake ac ion o e ain hei cus ome s since i has a signi ican impac on hei p o i . The enewal’s p icing p ocess occu s yea ly a he cus ome s' enewal da e, p oposing a a ia ion o he policy’s paid p emium. The company’s goal was o educe he obse ed chu n a es consequen o his p ocess. The e o e, he company wan ed o in eg a e a model ha could iden i y ea ly signs o chu n acco dingly o he di e en possible p emium a ia ions ha he gi en policy could su e . So, by cap u ing i s cus ome s’ p ice elas ici y, p emium a ia ions can be co ec ly adjus ed, and cus ome s ha p esen a ce ain isk le el o lea ing, possibly main ained. The CRISP-DM me hodology was applied o accomplish his p ojec , co e ing i s main s eps, and going back and o wa d when necessa y. The i s pa was dedica ed o unde s anding he business i sel , i s goals, and which a iables could p o ide help ul in o ma ion abou his cus ome beha io . A ound 80 a iables (nume ical and ca ego ical) we e ex ac ed om he company’s da abase. Howe e , as obse ed in he da a unde s anding phase, no all a iables we e ele an , especially when sepa a ed wi hin he h ee cus ome segmen s. This phase was ins umen al in unde s anding some o he s eps ha should be done in he Da a P epa a ion phase, whe e da a in i s aw o m we e p epa ed o modeling. He e, ac i i ies like edundan ea u e emo al, handling missing alues, c ea ing new a ibu es, and o ma ing he da a in a way ha would allow (and help) he modeling phase we e ca ied on. Mo eo e , as an al e na i e o using municipali ies o dis ic s, new geog aphical a eas we e c ea ed. Fo his, clus e ing analysis was pe o med and, o each o he cus ome segmen s, new geog aphical (con iguous) a eas we e c ea ed based on he municipali ies’ ele an demog aphic in o ma ion and obse ed chu n a es. Thus, i was possible o educe eigh een new dimensions (in case dis ic s we e used) in o only ou ele an dimensions, using in o ma ion ega ding mo e han h ee hund ed municipali ies. A e ha ing he da ase p epa ed, ea u es o be used by each o he models a e selec ed. The goal was o use a sui able numbe o ea u es gi en he numbe o ows a ailable o modeling. Mo e ea u es ep esen a highe p oblem complexi y, and o which, models should ha e mo e da a o unde s and. The ea u es a e chosen using ensemble lea ning, combining mul iple lea ne s’ opinions abou he ea u e’s impo ance ankings. The Top N ea u es a e selec ed, o each model, based on 39 he esul s ob ained o each se o ea u es. Fo he Logis ic eg ession, backwa d elimina ion based on he Wald’s es was used. The model’s pe o mance was e alua ed ac oss all he modeling phases. Time-spli c oss- alida ion was used so he exis ing empo al dependencies on he da a could be espec ed. A he same ime, he model’s gene aliza ion abili y can be gua an eed, app oxima ing he alida ion esul s o he ac ual model pe o mance. The inal model selec ion decision was based on he company’s wo business goals: inc easing he b anches’ p o i (acco dingly wi h he model’s pe o mance) and ha e a sui able model o op imizing he enewal’s p icing p ocess. Fo his la e , a mono onic inc easing ela ionship among P emium Va ia ion and P obabili y o Chu n needs o be assu ed. The inal selec ed model was an Ensemble Lea ning echnique, he Ex eme G adien Boos ing model o which he mono onic ela ionship was o ced. This model shows in c oss- alida ion an AUC o 59.0% (wi h a s anda d de ia ion o 3 p.p.), a speci ici y o 74.4% (wi h a s anda d de ia ion o 7 p.p.) and, las ly, a ecall o 43.5% (wi h a s anda d de ia ion o 10 p.p.). As o he a e age esul s in he used es se s, an AUC o 56.4% (wi h a s anda d de ia ion o 2 p.p.), a speci ici y o 72.7% (wi h a s anda d de ia ion o 12 p.p.) and, las ly, a ecall o 40.2% (wi h a s anda d de ia ion o 12 p.p.) we e obse ed. Fu he mo e, his model can answe he company’s needs. Fi s ly, i is sui able o he op imiza ion o he enewal’s p icing p ocess: o a gi en enewal yea and policy, he deli e ed model allows ha i he company decides o inc ease (o dec ease) he p emium a ia ion, hen he p obabili y o chu n will ei he main ain he same alue o inc ease (o dec ease) in i s alue. The e o e, pe mi ing da a- d i en decisions. Las ly, he model allows an inc ease in p o i : i was es ima ed ha his model could con ibu e o an annual a e age inc ease o a ound 70K €. 40 7. LIMITATIONS AND RECOMMENDATIONS FOR FUTURE WORKS I is possible o no ice ha he deli e ed solu ion has some space o imp o emen since he esul s in e ms o pe o mance measu es could be highe . Resul s a e highly dependen on he da a used and, he e, some imp o emen s could be made, o o he echniques a emp ed, like o example he es o o he classi ica ion algo i hms. Se e al issues we e iden i ied du ing he da a collec ion ask. Many inconsis encies we e de ec ed among he di e en da a sou ces ( he company is cu en ly cons uc ing a uni ied sou ce o in o ma ion). Al hough such p oblems we e deal wi hin he bes way possible, di e en quali y issues ha e eme ged, impac ing he inal collec ed da ase . The e o e, he quali y in he used da ase o modeling could no be assu ed. I was also no possible o collec da a o many o he company’s policies, impac ing he inal da ase ’s dimensions. E en ha ing da a o ou whole yea s, a highe amoun o da a could su ely help models in hei esul s. Addi ionally, he e a e possibly o he a iables ha a e no a ailable o he company bu could be ele an , like ex e nal ac o s ha migh in luence he cus ome ’s beha io . Rega ding da a p e-p ocessing, o he echniques could be es ed as o he missing alues impu a ion echniques, a deepe ou lie de ec ion analysis, o he o e sampling echniques, and u he dimensionali y educ ion ( ea u e selec ion) echniques could be explo ed as well. Las ly, as was possible o obse e, some models p esen be e esul s han o he s. The e o e, i could be in e es ing o explo e o he models ha a e mono onic o allow a mono onic cons ain . Two models ha would be in e es ing o es a e Ligh GBM ( om Py hon’s Ligh GBM package (Ke e al. 2017)) and la ice-based models ( om Py hon’s Tenso Flow-La ice package (Google AI Blog, 2017)). Simila ly o he Ex eme G adien Boos ing model ( om he xgboos package (Chen & Gues in, 2016)), hese models allow he in oduc ion o mono onic cons ain s. Howe e , due o secu i y cons ain s ha he company needs o ensu e, i was no possible o ob ain such packages in a iable ime o his p ojec execu ion. 41 8. BIBLIOGRAPHY Ahn, J., Hwang, J., Kim, D., Choi, H., & Kang, S. (2020). A Su ey on Chu n Analysis in Va ious Business Domains. IEEE Access, 8, 220816–220839. h ps://doi.o g/10.1109/ACCESS.2020.3042657 Aye , M., B unk, H. D., Ewing, G. M., Reid, W. T., & Sil e man, E. (1955). An empi ical dis ibu ion unc ion o sampling wi h incomple e in o ma ion. The Annals o Ma hema ical S a is ics, 26(4), 641–647. h ps://doi.o g/10.1214/aoms/1177728423 Ba is a, G. E., & Mona d, M. C. (2003). An analysis o ou missing da a ea men me hods o supe ised lea ning. Applied A i icial In elligence, 17(5–6), 519–533. h ps://doi.o g/10.1080/713827181 Bellman, R. (1966). Dynamic p og amming. Science, 153(3731), 34–37. h ps://doi.o g/10.1126/science.153.3731.34 Bisong, E. (2019). Building Machine Lea ning and Deep Lea ning Models on Google Cloud Pla o m. In Ap ess, Be keley, CA. h ps://doi.o g/10.1007/978-1-4842-4470-8_29 Bolancé, C., Guillen, M., & Padilla-Ba e o, A. E. (2016). P edic ing P obabili y o Cus ome Chu n in Insu ance. In R. León, M. Muñoz-To es, & J. Mone a (Eds.), Modeling and Simula ion in Enginee ing, Economics and Managemen . MS 2016. Lec u e No es in Business In o ma ion P ocessing, ol 254 (pp. 82–91). Sp inge , Cham. h ps://doi.o g/10.1007/978-3-319-40506-3_9 Bollegala, D. (2017). Dynamic ea u e scaling o online lea ning o bina y classi ie s. Knowledge- Based Sys ems, 129, 97–105. h ps://doi.o g/10.1016/j.knosys.2017.05.010 Bolón-Canedo, V., & Alonso-Be anzos, A. (2019). Ensembles o ea u e selec ion: A e iew and u u e ends. In o ma ion Fusion, 52, 1–12. h ps://doi.o g/10.1016/j.in us.2018.11.008 Box, G. E., & Cox, D. R. (1964). An Analysis o T ans o ma ions. Jou nal o he Royal S a is ical Socie y: Se ies B (Me hodological), 26(2), 211–243. h ps://doi.o g/10.1111/j.2517-6161.1964. b00553.x B eiman, L. (2001). Random Fo es s. Machine Lea ning, 45, 5–32. h ps://doi.o g/10.1023/A:1010933404324 Bu sac, Z., Gauss, C. H., Williams, D. K., & Hosme , D. W. (2008). Pu pose ul selec ion o a iables in logis ic eg ession. Sou ce Code o Biology and Medicine, 3(17). h ps://doi.o g/10.1186/1751- 0473-3-17 Chawla, N. V., Bowye , K. W., Hall, L. O., & Kegelmeye , W. P. (2002). SMOTE: Syn he ic Mino i y O e -sampling Technique. Jou nal o A i icial In elligence Resea ch, 16(1), 321–357. h ps://doi.o g/10.5555/1622407.1622416 Chawla, N. V., Cieslak, D. A., Hall, L. O., & Joshi, A. (2008). Au oma ically coun e ing imbalance and i s empi ical ela ionship o cos . Da a Mining and Knowledge Disco e y, 17, 225–252. h ps://doi.o g/10.1007/s10618-008-0087-0 Chawla, N. V., Japkowicz, N., & Ko cz, A. (2004). Edi o ial: special issue on lea ning om imbalanced da a se s. SIGKDD Explo ., 6, 1-6. h ps://doi.o g/10.1145/1007730.1007733 Chen, T., & Gues in, C. (2016). XGBoos : A Scalable T ee Boos ing Sys em. P oceedings o he 22nd Acm Sigkdd In e na ional Con e ence on Knowledge Disco e y and Da a Mining, 785–794. h ps://doi.o g/10.1145/2939672.2939785 42 Chiew, K. L., Tan, C. L., Wong, K., Yong, K. S., & Tiong, W. K. (2019). A new hyb id ensemble ea u e selec ion amewo k o machine lea ning-based phishing de ec ion sys em. In o ma ion Sciences, 484, 153–166. h ps://doi.o g/10.1016/j.ins.2019.01.064 Chuang, H. C., Chen, C. C., & Li, S. T. (2020). Inco po a ing mono onic domain knowledge in suppo ec o lea ning o da a mining eg ession p oblems. Neu al Compu ing and Applica ions, 32(15), 11791–11805. h ps://doi.o g/10.1007/s00521-019-04661-4 Collell, G., P elec, D., & Pa il, K. R. (2018). A simple plug-in bagging ensemble based on h eshold- mo ing o classi ying bina y and mul iclass imbalanced da a. Neu ocompu ing, 275, 330–340. h ps://doi.o g/10.1016/j.neucom.2017.08.035 De Win e , J. C., Gosling, S. D., & Po e , J. (2016). Supplemen al Ma e ial o Compa ing he Pea son and Spea man Co ela ion Coe icien s Ac oss Dis ibu ions and Sample Sizes: A Tu o ial Using Simula ions and Empi ical Da a. Psychological Me hods, 21(3), 273–290. h ps://doi.o g/10.1037/me 0000079.supp Die e ich, T. G. (2002). Ensemble Lea ning. The Handbook o B ain Theo y and Neu al Ne wo ks, 2(1), 110–125. Dola abadi, S. H., & Keynia, F. (2017). Designing o Cus ome and Employee Chu n P edic ion Model Based on Da a Mining Me hod and Neu al P edic o . 2017 2nd In e na ional Con e ence on Compu e and Communica ion Sys ems (ICCCS), 74–77. h ps://doi.o g/10.1109/CCOMS.2017.8075270 Do mann, C. F. (2020). Calib a ion o p obabili y p edic ions om machine‐lea ning and s a is ical models. Global Ecology and Biogeog aphy, 29(4), 760–765. h ps://doi.o g/10.1111/geb.13070 El Aboudi, N., & Benhlima, L. (2016). Re iew on W appe Fea u e Selec ion App oaches. 2016 In e na ional Con e ence on Enginee ing & MIS (ICEMIS), 1–5. h ps://doi.o g/0.1109/ICEMIS.2016.7745366 El Bouche y, K., & De Souza, R. S. (2020). Lea ning in Big Da a: In oduc ion o Machine Lea ning. In Knowledge Disco e y in Big Da a om As onomy and Ea h Obse a ion (pp. 225–249). Else ie . h ps://doi.o g/10.1016/b978-0-12-819154-5.00023-0 Fe nández, A., Ga cía, S., He e a, F., & Chawla, N. V. (2018). SMOTE o Lea ning om Imbalanced Da a: P og ess and Challenges, Ma king he 15-yea Anni e sa y. Jou nal o A i icial In elligence Resea ch, 61, 863–905. h ps://doi.o g/10.1613/jai .1.11192 Fo man, G., & Cohen, I. (2004). Lea ning om li le: Compa ison o classi ie s gi en li le aining. In J. Boulicau , F. Esposi o, F. Gianno i, & D. Ped eschi (Eds.), Knowledge Disco e y in Da abases: PKDD 2004. PKDD 2004. Lec u e No es in Compu e Science, ol 320 (pp. 161–172). Sp inge , Be lin, Heidelbe g. h ps://doi.o g/10.1007/978-3-540-30116-5_17 Fo ho e , R. N., Lee, E. S., & He nandez, M. (2007). Logis ic and P opo ional Haza ds Reg ession. In Bios a is ics (Second Edi ion) (pp. 387–419). Academic P ess. h ps://doi.o g/10.1016/B978-0- 12-369492-8.50019-4 F iedman, J. H. (2001). G eedy unc ion app oxima ion: a g adien boos ing machine. The Annals o S a is ics, 29(5), 1189–1232. h ps://doi.o g/10.1214/aos/1013203451 F iedman, J. H. (2002). S ochas ic g adien boos ing. Compu a ional S a is ics & Da a Analysis, 38(4), 367–378. h ps://doi.o g/10.1016/S0167-9473(01)00065-2 43 Ga cía, S., Luengo, J., & He e a, F. (2016). Tu o ial on p ac ical ips o he mos in luen ial da a p ep ocessing algo i hms in da a mining. Knowledge-Based Sys ems, 98, 1–29. h ps://doi.o g/10.1016/j.knosys.2015.12.006 Ga dne , M. W., & Do ling, S. R. (1998). A i icial neu al ne wo ks ( he mul ilaye pe cep on) - a e iew o applica ions in he a mosphe ic sciences. A mosphe ic En i onmen , 32(14–15), 2627– 2636. h ps://doi.o g/10.1016/S1352-2310(97)00447-0 Gün he , C. C., T e e, I. F., Aas, K., Sandnes, G. I., & Bo gan, Ø. (2014). Modelling and p edic ing cus ome chu n om an insu ance company. Scandina ian Ac ua ial Jou nal, 2014(1), 58–71. h ps://doi.o g/10.1080/03461238.2011.636502 Google AI Blog (2017). Tenso low la ice: Flexibili y empowe ed by p io knowledge. Re i ed om h ps://ai.googleblog.com/2017/10/ enso low-la ice- lexibili y.h ml Ha is, R., Sleigh , P., & Webbe , R. (2005). In oducing Geodemog aphics. In Geodemog aphics, GIS and neighbou hood a ge ing (pp. 16–17). John Wiley & Sons. He, H., & Ga cia, E. A. (2009). Lea ning om Imbalanced Da a. IEEE T ansac ions on Knowledge and Da a Enginee ing, 21(9), 1263–1284. h ps://doi.o g/10.1109/TKDE.2008.239 He, Y., Xiong, Y., & Tsai, Y. (2020). Machine Lea ning Based App oaches o P edic Cus ome Chu n o an Insu ance Company. 2020 Sys ems and In o ma ion Enginee ing Design Symposium (SIEDS), 1–6. h ps://doi.o g/10.1109/SIEDS49339.2020.9106691 Hua, J., Xiong, Z., Lowey, J., Suh, E., & Doughe y, E. R. (2005). Op imal numbe o ea u es as a unc ion o sample size o a ious classi ica ion ules. Bioin o ma ics, 21(8), 1509–1515. h ps://doi.o g/10.1093/bioin o ma ics/b i171 Inouye, D. I., Leqi, L., Kim, J. S., A agam, B., & Ra ikuma , P. (2020). Au oma ed Dependence Plo s. P oceedings o he 36 h Con e ence on Unce ain y in A i icial In elligence (UAI), 124, 1238– 1247. h p://a xi .o g/abs/1912.01108 Jain, A. K., & Chand aseka an, B. (1982). Dimensionali y and sample size conside a ions in pa e n ecogni ion p ac ice. In Handbook o S a is ics (Vol. 2, pp. 835–855). h ps://doi.o g/10.1016/S0169-7161(82)02042-2 Lemaî e, G., Noguei a, F., & A idas, C. K. (2017). Imbalanced-lea n: A py hon oolbox o ackle he cu se o imbalanced da ase s in machine lea ning. The Jou nal o Machine Lea ning Resea ch, 18(1), 559-563. Kalousis, A., P ados, J., & Hila io, M. (2007). S abili y o ea u e selec ion algo i hms: a s udy on high- dimensional spaces. Knowledge and In o ma ion Sys ems, 12, 95–116. h ps://doi.o g/10.1007/s10115-006-0040-8 Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W. Ma, W., Ye, Q., Liu, T.Y (2017). Ligh gbm: A highly e icien g adien boos ing decision ee. In Ad ances in neu al in o ma ion p ocessing sys ems, 30, 3146-3154. Koha i, R., & John, G. H. (1997). W appe s o ea u e subse selec ion. A i icial In elligence, 97(1–2), 273–324. h ps://doi.o g/10.1016/S0004-3702(97)00043-X Kuma , V., & Minz, S. (2014). Fea u e Selec ion: A li e a u e Re iew. The Sma Compu ing Re iew, 4(3), 211–229. h ps://doi.o g/10.6029/sma c .2014.03.007 44 Male ic, J., & Ma cus, A. (2000). Da a Cleansing: Beyond In eg i y Analysis. Iq, 200–209. McCue, C. (2015). Iden i ica ion, Cha ac e iza ion, and Modeling. In Da a Mining and P edic i e Analysis (Second Edi ion) (pp. 137–155). h ps://doi.o g/10.1016/B978-0-12-800229-2.00007-9 McGo e n, A., Lage quis , R., John Gagne, D., Je gensen, G. E., Elmo e, K. L., Homeye , C. R., & Smi h, T. (2019). Making he Black Box Mo e T anspa en : Unde s anding he Physical Implica ions o Machine Lea ning. Bulle in o he Ame ican Me eo ological Socie y, 100(11), 2175–2199. h ps://doi.o g/10.1175/BAMS-D-18-0195.1 Mei , R., & Rä sch, G. (2003). An In oduc ion o Boos ing and Le e aging. In S. Mendelson & A. J. Smola (Eds.), Ad anced lec u es on machine lea ning (pp. 118–183). Sp inge , Be lin, Heidelbe g. h ps://doi.o g/10.1007/3-540-36434-X_4 Naeini, M. P., Coope , G. F., & Hausk ech , M. (2015). Ob aining Well Calib a ed P obabili ies Using Bayesian Binning. P oceedings o he Twen y-Nin h AAAI Con e ence on A i icial In elligence, 2901–2907. Na ekin, A., & Knoll, A. (2013). G adien boos ing machines, a u o ial. F on ie s in Neu o obo ics, 7, 21. h ps://doi.o g/10.3389/ nbo .2013.00021 Niculescu-Mizil, A., & Ca uana, R. (2005). P edic ing good p obabili ies wi h supe ised lea ning. P oceedings o he 22nd In e na ional Con e ence on Machine Lea ning, 625–632. h ps://doi.o g/10.1145/1102351.1102430 Njikam, A. N. S., & Zhao, H. (2016). A no el ac i a ion unc ion o mul ilaye eed- o wa d neu al ne wo ks. Applied In elligence, 45, 75–82. h ps://doi.o g/10.1007/s10489-015-0744-0 Osbo ne, J. W. (2010). Imp o ing you da a ans o ma ions: Applying he Box-Cox ans o ma ion. P ac ical Assessmen , Resea ch and E alua ion, 15(12). h ps://doi.o g/10.7275/qbpc-gk17 Ped egosa, F., Va oquaux, G., G am o , A., Michel, V., Thi ion, B., G isel, O., Blondel, M., P e enho e , P., Weiss, R., Dubou g, V., Vande plas, J., Passos, A., Cou napeau, D., B uche , M., Pe o , M., & Duchesnay, E. (2011). Sciki -lea n: Machine Lea ning in Py hon. Jou nal o Machine Lea ning Resea ch, 12, 2825–2830. Re ei ed om h ps://jml .csail.mi .edu/pape s/ olume12/ped egosa11a/ped egosa11a.pd Pham, D. T., Dimo , S. S., & Nguyen, C. D. (2005). Selec ion o K in K-means clus e ing. P oceedings o he Ins i u ion o Mechanical Enginee s, Pa C: Jou nal o Mechanical Enginee ing Science, 219(1), 103–119. h ps://doi.o g/10.1243/095440605X8298 Polika , R. (2006). Ensemble based sys ems in decision making. IEEE Ci cui s and Sys ems Magazine, 6(3), 21–45. h ps://doi.o g/10.1109/MCAS.2006.1688199 Pozzolo, A. D., Caelen, O., Johnson, R. A., & Bon empi, G. (2015). Calib a ing P obabili y wi h Unde sampling o Unbalanced Classi ica ion. 2015 IEEE Symposium Se ies on Compu a ional In elligence, 159–166. h ps://doi.o g/10.1109/SSCI.2015.33 P ice, B. (2002). Making CRM come o li e. E-Business Re iew, 25–31. P o os , F. (2008). Machine lea ning om imbalanced da a se s 101. Rokach, L., & Maimon, O. (2005). Clus e ing Me hods. In O. Maimon & L. Rokach (Eds.), Da a Mining and Knowledge Disco e y Handbook (pp. 321–352). Sp inge , Bos on, MA. h ps://doi.o g/10.1007/0-387-25465-X_15 45 Rubin, D. B. (1976). In e ence and missing da a. Biome ika, 63(3), 581–592. h ps://doi.o g/10.1093/biome /63.3.581 Saeys, Y., Abeel, T., & Van De Pee , Y. (2008). Robus Fea u e Selec ion Using Ensemble Fea u e Selec ion Techniques. In W. Daelemans (Ed.), Machine Lea ning and Knowledge Disco e y in Da abases. ECML PKDD 2008. Lec u e No es in Compu e Science, ol 5212 (pp. 313–325). Sp inge -Ve lag Be lin Heidelbe g 2008. h ps://doi.o g/10.1007/978-3-540-87481-2_21 Sasse , W. E., & Reichheld, F. F. (1990). Ze o De ec ions -Quali y Comes o Se ices. Ha a d Business Re iew, 68(5), 105–111. Sc iney, M., Nie, D., & Roan ee, M. (2020). P edic ing Cus ome Chu n o Insu ance Da a. In M. Song, I. Song, G. Ko sis, A. M. Tjoa, & I. Khali (Eds.), Big Da a Analy ics and Knowledge Disco e y. DaWaK 2020. Lec u e No es in Compu e Science, ol 12393 (pp. 256–265). Sp inge , Cham. h ps://doi.o g/10.1007/978-3-030-59065-9_21 Seabold, S., & Pe k old, J. (2010). S a smodels: Econome ic and s a is ical modeling wi h py hon. In P oceedings o he 9 h Py hon in Science Con e ence. (Vol. 57, pp. 61). Seijo-Pa do, B., Bolón-Canedo, V., & Alonso-Be anzos, A. (2017). Tes ing Di e en Ensemble Con igu a ions o Fea u e Selec ion. Neu al P ocessing Le e s, 46, 857–880. h ps://doi.o g/10.1007/s11063-017-9619-1 Seijo-Pa do, B., Bolón-Canedo, V., Po o-Díaz, I., & Alonso-Be anzos, A. (2015). Ensemble Fea u e Selec ion o Rankings o Fea u es. In I. Rojas, G. Joya, & A. Ca ala (Eds.), Ad ances in Compu a ional In elligence. IWANN 2015. Lec u e No es in Compu e Science, ol 9095. Sp inge , Cham. h ps://doi.o g/10.1007/978-3-319-19222-2_3 Shah, R., & Pe e ia ko, V. (2021). Using Fea u e Impo ance Rank Ensembling (FIRE) o Ad anced Fea u e Selec ion. Re i ed om h ps://www.da a obo .com/blog/using- ea u e-impo ance- ank-ensembling- i e- o -ad anced- ea u e-selec ion/ Spi e i, M., & Azzopa di, G. (2018). Cus ome Chu n P edic ion o a Mo o Insu ance Company. 2018 Thi een h In e na ional Con e ence on Digi al In o ma ion Managemen (ICDIM), 173–178. h ps://doi.o g/10.1109/ICDIM.2018.8847066 S obl, C., Boules eix, A. L., Kneib, T., Augus in, T., & Zeileis, A. (2008). Condi ional a iable impo ance o andom o es s. BMC Bioin o ma ics, 9, 307. h ps://doi.o g/10.1186/1471- 2105-9-307 Sunda kuma , G. G., & Ra i, V. (2015). A no el hyb id unde sampling me hod o mining unbalanced da ase s in banking and insu ance. Enginee ing Applica ions o A i icial In elligence, 37, 368– 377. h ps://doi.o g/10.1016/j.engappai.2014.09.019 Talabis, M., McPhe son, R., Miyamo o, I., & Ma in, J. (2015). Analy ics De ined. In In o ma ion secu i y analy ics: inding secu i y insigh s, pa e ns, and anomalies in big da a (pp. 1–12). Syng ess. h ps://doi.o g/10.1016/b978-0-12-800207-0.00001-0 Toloşi, L., & Lengaue , T. (2011). Classi ica ion wi h co ela ed ea u es: un eliabili y o ea u e anking and solu ions. Bioin o ma ics, 27(14), 1986–1994. h ps://doi.o g/10.1093/bioin o ma ics/b 300 Va eiadis, T., Diaman a as, K. I., Sa igiannidis, G., & Cha zisa as, K. C. (2015). A compa ison o machine lea ning echniques o cus ome chu n p edic ion. Simula ion Modelling P ac ice and Theo y, 55, 1–9. h ps://doi.o g/10.1016/j.simpa .2015.03.003