scieee AI-readable full text Open interactive document viewer

Application of machine learning to mortality modeling and forecasting

Levantesi, Susanna,Pizzorusso, Virginia

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Levantesi, Susanna; Pizzorusso, Virginia Article Application of machine learning to mortality modeling and forecasting Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Levantesi, Susanna; Pizzorusso, Virginia (2019) : Application of machine learning to mortality modeling and forecasting, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 7, Iss. 1, pp. 1-19, https://doi.org/10.3390/risks7010026 This Version is available at: https://hdl.handle.net/10419/257864 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ risks Article Application of Machine Learning to Mortality Modeling and Forecasting Susanna Levantesi 1,* and Virginia Pizzorusso 2 1Department of Statistics, Sapienza University of Rome, Viale Regina Elena, 295/G, 00161 Rome, Italy 2Ernst and Young Advisory, Via Meravigli, 12, 20123 Milano, Italy; virginia.pizzor[email protected] *Correspondence: [email protected]; Tel.: +39-06-4925-5303 Received: 30 November 2018; Accepted: 21 February 2019; Published: 26 February 2019   Abstract: Estimation of future mortality rates still plays a central role among life insurers in pricing their products and managing longevity risk. In the literature on mortality modeling, a wide number of stochastic models have been proposed, most of them forecasting future mortality rates by extrapolating one or more latent factors. The abundance of proposed models shows that forecasting future mortality from historical trends is non-trivial. Following the idea proposed in Deprez et al. (2017), we use machine learning algorithms, able to catch patterns that are not commonly identifiable, to calibrate a parameter (the machine learning estimator), improving the goodness of fit of standard stochastic mortality models. The machine learning estimator is then forecasted according to the Lee-Carter framework, allowing one to obtain a higher forecasting quality of the standard stochastic models. Out-of sample forecasts are provided to verify the model accuracy. Keywords: mortality; forecasting; machine learning; Lee-Carter model 1. Introduction During the 20th Century, mortality has declined at all ages, producing a steep increase in life expectancy. This decrease is mainly due to the reduction of infectious disease mortality (between 1900 and 1950), as well as cardio-circulatory diseases and cancer mortality (in the most recent decades). Knowledge of future mortality rates is an important matter for life insurance companies with the goal of achieving adequate pricing of their life products. Therefore, sophisticated techniques to forecast future mortality rates have become increasingly popular in actuarial science, in order to deal with the longevity risk. Among the stochastic mortality models proposed in the literature, the Lee-Carter model Lee and Carter (1992) is the most widely used in the world, probably for its robustness. The original model applies singular-value decomposition (SVD) to the log-force of mortality to find three latent parameters: a fixed age component and a time component capturing the mortality trend that is multiplied by an age-specific function. Then, the time component is forecasted using a random walk. More recent approaches involve non-linear regression and generalized linear models (GLM), e.g., Brouhns et al. (2002) assumed a Poisson distribution for deaths and calculated the Lee-Carter model parameters by log-likelihood maximization. In recent years, machine learning techniques have assumed an increasingly central role in many areas of research, from computer science to medicine, including actuarial science. Machine learning is an application of artificial intelligence through a series of algorithms that are optimized on data samples or previous experience. That is, given a certain model defined as a function of a group of parameters, learning consists of improving these parameters using datasets or accumulated experience (the “training data”). Even though machine learning may not explain everything, it is very useful in detecting patterns, even unknown and unidentifiable ones, as well as hidden correlations. In this way, Risks 2019,7, 26; doi:10.3390/risks7010026 www.mdpi.com/journal/risks Risks 2019,7, 26 2 of 19 it allows us to understand processes better, make predictions about the future based on historical data, and categorize sets of data automatically. We can distinguish between supervised and unsupervised learning methods. In the supervised learning methods, the goal is to establish the relations between a range of predictors (independent variables) and a determined target (dependent variable), whereas in the unsupervised learning methods, the algorithm sets patterns among a range of variables in order to group records that show similarities, without considering an output measure. While in the supervised method, the algorithm learns from the dataset the rules that are fed to the machine, in the unsupervised method, it has to identify the rules autonomously. Logistic and multiple regression, classification and regression trees, and naive Bayes are examples of supervised learning methods, while association rules and clustering are classified as unsupervised learning methods. Despite the increasing usage in different fields of research, applications of machine learning in demography are not so popular. The main reason lies in the findings often being seen as “black boxes” and considered difficult to interpret. Moreover, the algorithms are not theory driven (but quite data driven), while demographers are often interested in analyzing specific hypotheses. They are likely to be unwilling to use algorithms whose decisions cannot be rationally explained. However, we believe that machine learning techniques can be valuable as a complement to standard mortality models, rather than a substitute. In the literature related to mortality modeling, there are very few contributions on this topic. The work in Deprez et al. (2017) showed that machine learning algorithms are useful to assess the goodness of fit of the mortality estimates provided by standard stochastic mortality models (they considered Lee-Carter and Renshaw-Haberman models). They applied a regression tree boosting machine to “analyze how the modeling should be improved based on feature components of an individual, such as its age or its birth cohort. This (non-parametric) regression approach then allows us to detect the weaknesses of different mortality models” (p. 337). In addition, they investigated cause-of-death mortality. In a recent paper, the work in Hainaut (2018) used neural networks to find the latent factors of mortality and forecast them according to a random walk with drift. Finally, the work in Richman and Wüthrich (2018) extended the Lee-Carter model to multiple populations using neural networks. We investigate the ability of machine learning to improve the accuracy of some standard stochastic mortality models, both in the estimation and forecasting of mortality rates. The novelty of this paper is primarily in the mortality forecasting that takes advantage of machine learning, clearly capturing patterns that are not identifiable with a standard mortality model. Following Deprez et al. (2017), we use tree-based machine learning techniques to calibrate a parameter (the machine learning estimator) to be applied to mortality rates fitted by the standard mortality model. We analyze three famous stochastic mortality models: the Lee-Carter model Lee and Carter (1992), which is still the most frequently implemented, the Renshaw-Haberman model Renshaw and Haberman (2006), which also considers the cohort effect, and the Plat model Plat (2009), which tries to combine the parameters of the Lee-Carter model with those of the Cairns-Blake-Dowd model with the cohort effect, named “M7” (Cairns et al. (2009)). Three different kinds of supervised learning methods are considered for calibrating the machine learning estimator: decision tree, random forest, and gradient boosting, which are all tree-based. We show that the implementation of these machine learning techniques, based on features components such as age, sex, calendar year, and birth cohort, leads to a better fit of the historical data, with respect to the estimates given by the Lee-Carter, Renshaw-Haberman, and Plat models. We also apply the same logic to improve the mortality forecasts provided by the Lee-Carter model, where the machine learning estimator is extrapolated using the Lee-Carter framework. Out-of-sample tests are performed for the improved model in order to verify the quality of forecasting. The paper is organized as follows. In Section 2, we specify the model and introduce the tree-based machine learning estimators. In Section 3, we present the stochastic mortality models considered in Risks 2019,7, 26 3 of 19 the paper. In Section 4, we illustrate the usage of tree-based machine learning estimators to improve both the fitting and forecasting quality of the original mortality models. Conclusions and further research are then given in Section 5. 2. The Model We consider the following categorical variables, identifying an individual: gender ( g ), age ( a ), calendar year ( t ), and year of birth ( c ). We assign to each individual the feature x= (g , a , t , c)∈ X with X=G × A × T × C the feature space, where: G={males,f emales} , A={0, ..., ω} , T={t1, ..., tn} , C={c1, ..., cm} . Other categorical variables could be included in the feature space X, e.g., the marital status, the income, and other individual information. We assume that the number of deaths Dxmeets the following conditions: •Dxare independent in {x∈ X }; •Dx∼Pois(mx·Ex)for all {x∈ X }. where mxis the central death rate and Exare the exposures. Let us define dmdl x as the expected number of deaths estimated by a standard stochastic mortality model (such as Lee-Carter, Cairns-Blake-Dowd, etc.) and mmdl x the corresponding central death rate. Following Deprez et al. (2017), but modeling the central death death rate instead of mortality rate ( qx ), we initially set: •mx=mmdl x •Dx∼Pois(ψx·dmdl x), with ψx≡1, dmdl x=mmdl xEx The condition ψx≡ 1 means that the specified mortality model perfectly fits the crude rates. However, in the real world, a mortality model could overestimate ( ψx≤ 1) or underestimate ( ψx≥ 1) the crude rates. Therefore, we calibrate the parameter ψx , based on the feature x , according to three different machine learning techniques. We find ψx as a solution of a regression tree algorithm applied to the ratio between the death observations and the corresponding value estimated by the specified mortality model Dx dmdl x: Dx dmdl x ∼gender +age +year +cohort (1) We denote by ˆ ψmdl,ML x the machine learning estimator obtained by solving Equation (1), where mdl indicates the stochastic mortality model and ML the machine learning algorithm used to improve the mortality rates given by a certain model. The estimator ˆ ψmdl,ML x is then applied to the central death rate of the specified mortality model, mmdl x, aiming to obtain a better fit of the observed data: mmdl,ML x=ˆ ψmdl,ML x·mmdl x,∀x∈ X (2) As in Deprez et al. (2017), we measure the improvement in the mortality rates attained by the tree growing algorithm through the relative changes of central death rates: ∆mmdl,ML x=mmdl,ML x−mmdl x mmdl x =ˆ ψML x−1 (3) The work in Hainaut (2018) used neural networks to learn the logarithm of the central death rates directly from the features of the mortality data, by using age, calendar year, and gender (and region) as predictors in a neural network. We instead rely on the classical form of the Lee-Carter model that we improve ex-post using machine learning algorithms, considered complementary and not an alternative to the standard mortality modeling. To estimate ˆ ψmdl,ML x, we use the following tree-based machine learning (ML) techniques: •Decision tree Risks 2019,7, 26 4 of 19 •Random forest •Gradient boosting 2.1. Decision Trees The tree-based methods for regression and classification (Breiman et al. 1984) have become popular alternatives to linear regression. They are based on the partition of the feature space X , through a sequence of binary splits, and the set of splitting rules used to segment the predictor space can be summarized in a tree (Hastie et al. 2016). Once the entire feature space is split into a certain number of simple regions recursively, the response for a given observation can be predicted using the mean of the training observations in the region to which that observation belongs (James et al. 2017 ; Alpaydin 2010). Let (Xτ)τ∈T be the partition of X; the decision tree estimator is calculated as: ˆ ψ(x) = ∑ τ∈T ¯ ψτ 1 {x∈Xτ}(4) Decision trees (DT) algorithms have advantages over other types of regression models. As pointed out by James et al. (2017): they are easy to interpret; they can easily handle qualitative predictors without the need to create dummy variables; they can catch any kind of correlation in the data. However, they suffer from some important drawbacks: they do not always have predictive accuracy levels similar to those of traditional regression and classification models; they can lack robustness: a small modification of the data can produce a tree that strongly differs from the one initially estimated. The ML estimator is obtained using the Rpackage rpart (Therneau and Atkinson 2017). The algorithm provides the estimate of ˆ ψ(x) given by the average of the response variable values ψ(x) belonging to the same region identified by the regression tree. The values of the complexity parameter ( cp ) for the decision trees are chosen with the aim of making the number of splits considered uniform. 2.2. Random Forest The aggregation of many decision trees can improve the predictive performance of trees. Therefore, we first apply bagging (also called bootstrap aggregation) to produce a certain number, B , of decision trees from the bootstrapped training samples, in turn obtained from the bootstrap of the original training dataset. Random forest (RF) differs from bagging in the way of considering the predictors: RF algorithms account only for a random subset of the predictors at each split in the tree, as described in detail by Breiman (2001). If there is a strong predictor in the dataset, the other predictors will have more of a chance to be chosen as split candidates from the final set of predictors (James et al. 2017). The RF estimator is calculated as follows: ˆ ψ(x) = 1 B B ∑ b=1 ˆ ψ(b)(x)(5) The RF estimator is obtained by applying the algorithm from the Rpackage randomForest (Liaw 2018). Since this procedure proved to be very costly from a computational point of view, the number of trees must be carefully chosen: it should not be too large, but at the same time able to produce an adequate percentage of variance explained and a low mean of squared residuals, MSR. 2.3. Gradient Boosting Consider the loss in using a certain function to predict a variable on the training data; gradient boosting (GB) aims at minimizing the in-sample loss with respect to this function by a stage-wise adaptive learning algorithm that combines weak predictors. Let ψ(x) be the function; the gradient boosting algorithm finds an approximation ˆ ψ(x) to the function ψ(x) that minimizes the expected value of the specified differentiable loss function Risks 2019,7, 26 5 of 19 (optimization problem). At each stage i of gradient boosting (1 ≤i≤N ), we suppose that there is some imperfect models ˆ ψ(xi) , then the gradient boosting algorithm improves on ˆ ψ(xi) by constructing a new model that adds an estimator hto provide a better model: ˆ ψ(xi) = ˆ ψ(xi−1) + λihi(x)(6) where hi∈ H is a base learner function ( H is the set of arbitrary differentiable functions) and λ is a multiplier obtained by solving the optimization problem. The GB estimator is obtained using the Rpackage gbm (Ridgeway 2007). The gbm package requires choosing the number of trees ( n . trees ) and other key parameters as the number of cross-validation folds ( cv . f olds ), the depth of each tree involved in the estimate ( interaction . depth ), and the learning rate parameter ( shrinkage ). The number of trees, representing the number of GB iterations, must be accurately chosen, as a high number would reduce the error on the training set, while a low number would result in overfitting. The number of cross-validation folds to perform should be chosen according to the dataset size. In general, five-fold cross-validation, which corresponds to 20% of the data involved in testing, is considered a good choice in many cases. Finally, the interaction depth represents the highest level of variable interactions allowed or the maximum nodes for each tree. 3. Mortality Models Let us consider the generalized age period cohort (GAPC) stochastic mortality models’ family (see Villegas et al. 2015 for further details). In the GAPC models, the effects of age, calendar year, and cohort are caught by a predictor, in our framework denoted by ηx, as follows: ηx=αa+ n ∑ i=1 β(i) aκ(i) t+β(0) aγt−a,∀x= (g,a,t,c)∈ X (7) where: •αa: age-specific parameter providing the average age profile of mortality; •β(i) a·κ(i) t , ∀i : age-period terms describing the mortality trends ( κ(i) t is the time index, and β(i) a modifies the effect of κ(i) tacross ages); •β(0) a·γt−a : represents the cohort effect, where γt−a is the cohort parameter and β(0) a modifies its effect across ages (c=t−ais the year of birth). The mortality predictor is related to a link function g , so that: ηx=gEDx Ex . In this paper, we consider the log link function and assume that the numbers of deaths Dx follow a Poisson distribution. 3.1. Lee-Carter Model Under the above-described framework, the Lee-Carter (LC) model as proposed by Brouhns et al. (2002) requires a log link function to target the central death rate. In the LC model, the logarithm of the central death rate is described by: log (mx)=αa+β(1) aκ(1) t(8) with the constraints: ∑t∈T κ(1) t= 0, ∑a∈A β(1) a= 1 to avoid identifiability problems with the parameters. In order to forecast mortality with the LC model, the time index κ(1) t is modeled by an autoregressive integrated moving average (ARIMA) process. In general, a random walk with drift properly fits the data: κ(1) t=κ(1) t−1+δ+et,et∼N(0, σ2 k)(9) where δ is the drift parameter and et are the error terms, normally distributed with null mean and variance σ2 k. Risks 2019,7, 26 6 of 19 3.2. Renshaw-Haberman Model The Renshaw-Haberman model (Renshaw and Haberman (2006)) extends the LC model by including a cohort effect. The model’s predictor has the following expression, where the log link function is used to target the central death rate: log (mx)=αa+β(1) aκ(1) t+β(0) aγt−a(10) According to Haberman and Renshaw (2011) and Hunt and Villegas (2015), we set β(0) a= 1 ∀a∈ A, as the model is more stable with respect to the original version. log (mx)=αa+β(1) aκ(1) t+γt−a(11) The model is subject to the following constraints, where c=t−a : ∑t∈T κ(1) t= 0, ∑a∈A β(1) a=1, and ∑c∈C γc= 0. Parameters κ(1) t and γt−a are modeled by ARIMA processes, assuming the independence between them. 3.3. Plat Model The Plat model Plat (2009) aims to combine M7 and LC models in order to obtain a model appropriate for the entire age range and for capturing the cohort effect, thus overcoming the disadvantages of the previous models. log (mx) = αa+κ(1) t+κ(2) t(¯ a−a) + κ(3) t(¯ a−a)++γt−a(12) where (¯ a−a)+=max (¯ a−a, 0) . This model is obtained from Equation (7) by setting β(1) a= 1, β(2) a=¯ a−a , β(3) a= (¯ a−a)+ , and β(0) a= 1 and using the log link function to target the central death rate. The Plat model is subject to the following constraints: ∑tκ(1) t= 0, ∑tκ(2) t= 0, ∑tκ(3) t= 0, ∑c∈C γc= 0, ∑c∈C γcc= 0, ∑c∈C γcc2= 0. As described in Villegas et al. (2015): “the first three constraints ensure that the period indexes are centered around zero, while the last three constraints ensure that the cohort effect fluctuates around zero and has no linear or quadratic trend”. 4. Numerical Analysis 4.1. Model Fitting We fit the Lee-Carter (LC), Renshaw-Haberman (RH) and Plat models on the Italian population. Data were downloaded from the Human Mortality Database (www.mortality.org), while model fitting was performed with StMoMo package provided by Villegas et al. (2015). The following sets of gender G , ages A, years T, and cohort Cwere considered in the analysis: G={males,f emales},A={0, ..., 100},T={1915, ..., 2014}, and C={1815, ..., 2014}. The model accuracy was measured by the Bayes information criterion (BIC) and the Akaike information criterion (AIC), which are measures generally used to evaluate the goodness of fit of mortality models 1 . Log-likelihood L , AIC and BIC values are reported in Table 1, from which we observe that the RH model fits the historical data very well. It has the highest BIC and AIC values for both genders, with respect to the other models; then, in order, the LC model and the Plat model. 1 The AIC and BIC statistics are both function of the log-likelihood, L , and the number of parameters involved in the model, ν: AIC =2ν−2L, and BIC =νlog N−2L, where Nis the number of observations. Risks 2019,7, 26 7 of 19 Table 1. Log-likelihood, AIC and BIC statistics for LC, RH and Plat model. Ages 0–100 and years 1915–2014, Italian population. Gender: Males Females Model: LC RH Plat LC RH Plat ν300 499 496 300 499 496 L −643,226 −307,985 −875,549 −176,421 −137,098 −281,518 AIC (Rank) 1,287,051 (2) 616,967 (1) 1,752,089 (3) 353,442 (2) 275,193 (1) 564,028 (3) BIC (Rank) 1,289,218 (2) 620,570 (1) 1,755,671 (3) 355,608 (2) 278,796 (1) 567,610 (3) The goodness of fit is also tested by the residuals analysis. From Figure 1, we can observe that the RH provided the best fit, despite the highest number of parameters. The LC model provided a good fit especially for the old-age population, while the Plat model provided the worst performance despite the high number of parameters involved. 1920 1940 1960 1980 2000 0 20 40 60 80 100 calendar year age −3 −2 −1 0 1 2 3 (a) LC (males) 1920 1940 1960 1980 2000 0 20 40 60 80 100 calendar year age −3 −2 −1 0 1 2 3 (b) RH (males) 1920 1940 1960 1980 2000 0 20 40 60 80 100 calendar year age −3 −2 −1 0 1 2 3 (c) Plat (males) 1920 1940 1960 1980 2000 0 20 40 60 80 100 calendar year age −3 −2 −1 0 1 2 3 (d) LC (females) 1920 1940 1960 1980 2000 0 20 40 60 80 100 calendar year age −3 −2 −1 0 1 2 3 (e) RH (females) 1920 1940 1960 1980 2000 0 20 40 60 80 100 calendar year age −3 −2 −1 0 1 2 3 (f) Plat (females) Figure 1. Heat map of standardized residuals of the mortality models. Ages 0–100 and years 1915–2014, Italian population. 4.2. Model Fitting Improved by Machine Learning In the following, we specify the parameters used to calibrate the ML algorithms described in Section 2using the rpart,randomForest, and gbm packages, respectively: •ˆ ψmdl,DT xwas estimated with the rpart package by setting: cp = 0.003 (complexity parameter); •ˆ ψmdl,RF x was estimated with the randomForest package by setting: ntrees = 200 (number of trees). Since this procedure proved to be very costly from a computational point of view, we limited the number of trees to 200, in order to guarantee both an adequate percentage of variance explained by the model and a low mean of squared residuals, MSR (see Table 2); •ˆ ψmdl,GB x is estimated with the gbm package by setting: n . trees = 5000 (number of trees); cv.f olds = 5 (number of cross-validation folds); interaction . depth = 6; shrinkage = 0.001 (learning rate) according to the algorithm implementation speed. The parameter cv . f olds is used to estimate the optimal number of iterations through the function gbm.per f (see Figure 2). Risks 2019,7, 26 8 of 19 Table 2. Explained variance and MSR by the RF algorithm for the LC, RH, and Plat model. Ages 0–100 and years 1915–2014, Italian population. Model Explained Variance MSR LC 96.25% 0.0263 RH 86.48% 0.0057 Plat 91.14% 0.0058 0 1000 2000 3000 4000 5000 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Iteration Squared error loss (a) Lee-Carter 0 1000 2000 3000 4000 5000 0.020 0.025 0.030 Iteration Squared error loss (b) RH 0 1000 2000 3000 4000 5000 0.03 0.04 0.05 0.06 Iteration Squared error loss (c) Plat Figure 2. Estimates of the optimal number of boosting iterations for the LC, RH, and Plat model. Black line: Out-of-bag estimates; green line: cross-validation estimates. The level of improvement in central death rates resulting from the application of ML algorithms was measured by ∆mmdl,ML x, the relative changes described in Equation (3). Numerical results for the LC, RH, and Plat model combined with the tree-based ML algorithms are shown in Figure 3for males. Similar results were obtained for females. The white areas represent very small variations of ∆mmdl,ML x , approximately around zero. Larger white areas were observed for gradient boosting applied to the LC and RH model. In all cases, there were also significant changes that were less prominent for the RH model that best fit the historical data. Many regions were identified by diagonal splits (highlighting a cohort effect), strengthening our choice to insert the cohort parameter in the decision tree algorithms. Especially for the LC model, we point out that the relative changes were mainly concentrated in the young ages. For the Plat model, we observed small values of ∆mPL,ML x with respect to the other mortality models, with the exception of the population aged under 40 that showed quite significant changes. From these early results, DT and RT seemed to work better than the GB algorithm. Risks 2019,7, 26 15 of 19 Appendix A Appendix A.1 Plots for Time Period 1915–2014 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (a) LC-DT 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (b) RH-DT 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (c) Plat-DT 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (d) LC-RF 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (e) RH-RF 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (f) Plat-RF 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (g) LC-GB 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (h) RH-GB 0 20 40 60 80 100 1920 1940 1960 1980 2000 0 1 2 3 (i) Plat-GB Figure A1. Values of ∆mmdl,ML x. Italian female population. Ages 0–100 and years 1915–2014. Risks 2019,7, 26 16 of 19 1920 1940 1960 1980 2000 −150 −100 −50 0 50 100 κt (1) vs. t year (a) LC (males): κ(1) t 1920 1940 1960 1980 2000 −200 −100 0 100 κt (1) vs. t year (b) LC (females): κ(1) t 1920 1940 1960 1980 2000 0 10 20 30 κt (1) vs. t year (c) LC-DT (males): κ(1,ψ) t 1920 1940 1960 1980 2000 0 10 20 30 40 κt (1) vs. t year (d) LC-DT (females): κ(1,ψ) t 1920 1940 1960 1980 2000 −5 0 5 10 15 20 25 κt (1) vs. t year (e) LC-RF (males) 1920 1940 1960 1980 2000 0 20 40 60 κt (1) vs. t year (f) LC-RF (females): κ(1,ψ) t 1920 1940 1960 1980 2000 0 10 20 30 40 κt (1) vs. t year (g) LC-GB (males): κ(1,ψ) t 1920 1940 1960 1980 2000 −10 0 10 20 30 40 50 60 κt (1) vs. t year (h) LC-GB (females): κ(1,ψ) t Figure A2. κ(1) tand κ(1,ψ) t: Fitted value (1915–2000) and forecasted values (2000–2014). Risks 2019,7, 26 17 of 19 Appendix A.2 Plots for Time Period 1960–2014 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (a) LC-DT 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (b) RH-DT 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (c) Plat-DT 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (d) LC-RF 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (e) RH-RF 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (f) Plat-RF 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (g) LC-GB 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (h) RH-GB 0 20 40 60 80 100 1960 1980 2000 0 1 2 3 (i) Plat-GB Figure A3. Values of ∆mmdl,ML x. Italian female population. Ages 0–100 and years 1960–2014. Risks 2019,7, 26 18 of 19 1960 1970 1980 1990 2000 2010 −80 −60 −40 −20 0 20 κt (1) vs. t year (a) LC (males): κ(1) t 1960 1970 1980 1990 2000 2010 −100 −50 0 κt (1) vs. t year (b) LC (females): κ(1) t 1960 1970 1980 1990 2000 2010 −5 0 5 10 κt (1) vs. t year (c) LC-DT (males): κ(1,ψ) t 1960 1970 1980 1990 2000 2010 −5 0 5 10 κt (1) vs. t year (d) LC-DT (females): κ(1,ψ) t 1960 1970 1980 1990 2000 2010 −5 0 5 10 κt (1) vs. t year (e) LC-RF (males) 1960 1970 1980 1990 2000 2010 −5 0 5 10 15 κt (1) vs. t year (f) LC-RF (females): κ(1,ψ) t 1960 1970 1980 1990 2000 2010 −5 0 5 10 κt (1) vs. t year (g) LC-GB (males): κ(1,ψ) t 1960 1970 1980 1990 2000 2010 −5 0 5 10 κt (1) vs. t year (h) LC-GB (females): κ(1,ψ) t Figure A4. κ(1) tand κ(1,ψ) t: Fitted value (1960–2000) and forecasted values (2000–2014). Risks 2019,7, 26 19 of 19 References Alpaydin, Ethem. 2010. Introduction to Machine Learning, 2nd ed. Cambridge: Massachusetts Institute of Technology Press, ISBN 026201243X. Breiman, Leo, Jerome Friedman, Richard Olshen, and Charles Stone. 1984. Classification and regression trees. Boca Raton: CRC Press, ISBN 9780412048418. [CrossRef] Breiman, Leo. 2001. Random forests. Machine Learning 45: 5–32. Brouhns, Natacha, Michel Denuit, and Jeroen K. Vermunt. 2002. A Poisson log-bilinear regression approach to the construction of projected lifetables. Insurance: Mathematics and Economics 31: 373–93. [CrossRef] Cairns, Andrew J. G., David Blake, Kevin Dowd, Guy D. Coughlan, David Epstein, Alen Ong, and Igor Balevich. 2009. A quantitative comparison of stochastic mortality models using data from England and Wales and the United States. North American Actuarial Journal 13: 1–35. [CrossRef] Deprez, Philippe, Pavel V. Shevchenko, and Mario V. Wüthrich. 2017. Machine learning techniques for mortality modeling. European Actuarial Journal 7: 337–52. [CrossRef] Haberman, Steven, and Arthur Renshaw. 2011. A comparative study of parametric mortality projection models. Insurance: Mathematics and Economics 48: 35–55. [CrossRef] Hainaut, Donatien. 2018. A neural-network analyzer for mortality forecast. Astin Bulletin 48: 481–508. [CrossRef] Hastie, Jerome, Trevor Hastie, and Robert Tibshirani. 2016. The Elements of Statistical Learning, 2nd ed. Data Mining, Inference, and Prediction. New York: Springer, ISBN 0387848576. Hunt, Andrew, and Andrés M. Villegas. 2015. Robustness and convergence in the Lee-Carter model with cohorts. Insurance: Mathematics and Economics 64: 186–202. [CrossRef] James, Gareth, Daniela Witten, Trevor Hastie, and Robert Tibshirani. 2017. An Introduction to Statistical Learning: With Applications in R. New York: Springer, ISBN 1461471370. Lee, Ronald D., and Lawrence R. Carter. 1992. Modeling and forecasting US mortality. Journal of the American Statistical Association 87: 659–71. Liaw, Andy. 2018. Package Randomforest . Available online: https://cran.r-project.org/web/packages/randomForest/ randomForest.pdf (accessed on 21 May 2018). Plat, Richard. 2009. On stochastic mortality modeling. Insurance: Mathematics and Economics 45: 393–404. Renshaw, Arthur E., and Steven Haberman. 2006. A Cohort-Based Extension to the Lee-Carter Model for Mortality Reduction Factors. Insurance: Mathematics and Economics 38: 556–70. [CrossRef] Richman, Ronald, and Mario V. Wüthrich. 2018. A Neural Network Extension of the Lee-Carter Model to Multiple Populations. Rochester: SSRN. Ridgeway, Greg. 2007. Generalized Boosted Models: A Guide to the gbm Package. Available online: https: //cran.r-project.org/web/packages/gbm/gbm.pdf (accessed on 21 May 2018). Therneau, Terry M., and Elizabeth J. Atkinson. 2017. An Introduction to Recursive Partitioning Using the RPART Routines. Available online: https://cran.r-project.org/web/packages/rpart/vignettes/longintro. pdf (accessed on 21 May 2018). Villegas, Andrés M., Pietro Millossovich, and Vladimir K. Kaishev. 2015. Stmomo : An r Package for Stochastic Mortality Modelling. Available online: https://cran.r-project.org/web/packages/StMoMo/vignettes/ StMoMoVignette.pdf (accessed on 21 May 2018). c 2019 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).