A neural network approach for the mortality analysis of multiple populations: a case study on data of the Italian population
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Euthum, Maximilian; Scherer, Matthias; Ungolo, Francesco Article — Published Version A neural network approach for the mortality analysis of multiple populations: a case study on data of the Italian population European Actuarial Journal Provided in Cooperation with: Springer Nature Suggested Citation: Euthum, Maximilian; Scherer, Matthias; Ungolo, Francesco (2024) : A neural network approach for the mortality analysis of multiple populations: a case study on data of the Italian population, European Actuarial Journal, ISSN 2190-9741, Springer, Berlin, Heidelberg, Vol. 14, Iss. 2, pp. 495-524, https://doi.org/10.1007/s13385-024-00377-5 This Version is available at: https://hdl.handle.net/10419/315863 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) European Actuarial Journal (2024) 14:495–524 https://doi.org/10.1007/s13385-024-00377-5 1 3 CASE STUDY A neural network approach forthemortality analysis ofmultiple populations: acase study ondata oftheItalian population MaximilianEuthum1,2· MatthiasScherer1 · FrancescoUngolo1,3,4 Received: 27 February 2023 / Revised: 14 July 2023 / Accepted: 19 December 2023 / Published online: 6 March 2024 © The Author(s) 2024 Abstract A Neural Network (NN) approach for the modelling of mortality rates in a multipopulation framework is compared to three classical mortality models. The NN setup contains two instances of Recurrent NNs, including Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU) networks. The stochastic approaches comprise the Li and Lee model, the Common Age Effect model of Kleinow, and the model of Plat. All models are applied and compared in a large case study on decades of data of the Italian population as divided in counties. In this case study, a new index of multiple deprivation is introduced and used to classify all Italian counties based on socio-economic indicators, sourced from the local office of national statistics (ISTAT). The aforementioned models are then used to model and predict mortality rates of groups of different socio-economic characteristics, sex, and age. Keywords Case Study on Mortality· Longevity Risk· Neural Network· Multipopulation· Deprivation Index· Socio-economic characteristics· Italian data * Matthias Scherer [email protected] 1 Chair ofMathematical Finance, Technical University ofMunich, GarchingbeiMünchen, Germany 2 Munich Reinsurance Company (Munich RE), Munich, Germany 3 School ofRisk andActuarial Studies, University ofNew South Wales, Kensington, NSW, 2052, Australia 4 ARC Centre ofExcellence inPopulation Ageing Research, University ofNew South Wales, Kensington, NSW2052, Australia
496 M.Euthum et al. 1 3 1 Introduction 1.1 Mortality modelling: motivation, background, andliterature Since the seminal work of Lee and Carter [17], several stochastic models for the estimation and projection of mortality rates were developed during the last decades, see, e.g., the contributions of Brouhns etal. [2], Renshaw and Haberman [31], Cairns etal. [3], and Plat [28]. While these pioneering approaches analyzed single populations, models for multiple populations gained considerable importance in subsequent years; after it was found that the mortality profiles of multiple populations tended to converge (cf. Wilson [40]). Indeed, the multi-population paradigm often has advantages over modelling mortality rates for each population separately. Most notably, multi-population mortality models can capture common features of the mortality profile of similar populations such as neighbouring countries, populations showing similar socio-economic-, environmental-, or biological characteristics, while simultaneously reflecting population-specific features. This motivates their use for producing coherent projections of mortality rates. Recently, Neural Network (NN)-based approaches for mortality modelling were proposed as an appealing alternative to classical stochastic models. Among the first examples is the work of Richman and Wüthrich [33], who analyses the Swiss population, and Hainaut [11] who uses French, UK, and US mortality rates to compare a NN approach to the Lee–Carter model w/wo cohort effects. Another contribution is Nigri etal. [23], who use a deep learning algorithm based on a two-step Recurrent Neural Network (RNN) to enhance the forecasts obtainable under the Lee–Carter model. Indeed, there have been many recent developments in the use of NNs in the context of multi-population mortality modelling, such as Perla etal. [27], which considers the use of one-dimensional (period effect only) RNN with Long Short-Term Memory (LSTM) and of Convolutional Neural Networks (CNN) to provide direct forecasts of the mortality rates compared to the two-step approach of Nigri etal. [23]. Lindholm and Palmborg [19] consider similar models with a focus on the optimal use of data for projection. Schnürch and Korn [35] extend the RNN and CNN by proposing a two-dimensional approach involving age and period. In a slightly different fashion, Scognamiglio [36] proposes a NN architecture for the joint calibration of individual Lee–Carter models based on its classical log-normal representation as well as the Poisson Lee–Carter version of Brouhns etal. [2]. An approach that allows for coherent predictions within sub-groups of similar population is provided in Perla and Scognamiglio [26]. Finally, Wang etal. [38] develop a framework which ‘augments’ the mortality dataset to construct an image of neighborhood mortality data around the central death rate and use two CNN approaches for projecting mortality rates.
497 1 3 A neural network approach forthemortality analysis ofmultiple… 1.2 Applications of(multi‑population) mortality models A key application of multi-population mortality modelling approaches is the analysis of the mortality levels based on socio-economic characteristics. Among others, understanding mortality is relevant to policymakers in order to propose and plan sustainable state pension reforms and budgets, or to address disparities between socio-economic groups. Understanding mortality from a statistical pointof-view is also relevant for private sector players such as insurance companies and pension funds, offering mortality-linked products like annuities and pensions and designing effective solutions for longevity risk management and transfer. Considering specific populations that have already been investigated in the literature, let us mention Wen etal. [39], who compare several stochastic mortality models to fit mortality rates of small geographic areas in the UK (Lower Layer Super Output Areas), grouped in deciles of their Index of Multiple Deprivation. Cairns etal. [6] develop a multi-population mortality modelling approach for the analysis of the Danish population on the basis of the deciles of a newly created affluence index, which accounts for individual information on income and wealth. 1.3 Contributions: methodology andinvestigated population This paper investigates the use of NNs to jointly model the mortality rates of multiple populations and compares the empirical results to classical stochastic mortality models. The model for the dynamics of mortality rates draws on the work of Richman and Wüthrich [32], where they propose the use of Recurrent Neural Networks (RNN) such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU). Although computationally expensive, we focus on RNNs to exploit their suitability to deal with the time-series structure of mortality rates. Indeed, due to their recurrent connections, RNNs allow to maintain memory of past observations, since these let the information pass from past time steps to the current one. In this way, it is possible to model the information throughout the entire time series. Furthermore, RNNs can handle time series of variable length, due to their sequential processing of the data, cf. Hsu [12]. The NN approaches are then compared to well-established stochastic mortality models for multiple populations such as the Lee–Carter extension of Li and Lee [18] and the Common Age Effect model of Kleinow [16], as well as to the single-population approach represented by the Plat [28] model. In our case study, we investigate Italian mortality data that is grouped according to socio-economic characteristics. More precisely, we first propose a new deprivation index on the basis of five variables. This index allows separating the 106 Italian counties into nine groups of different socio-economic level, which are then used as populations for the empirical analysis. To the best of our knowledge, this study is the first to explicitly address the use of NNs in the context of the analysis of multiple populations on the basis of socio-economic characteristics.
498 M.Euthum et al. 1 3 The structure of the remaining paper is as follows: Sect. 2 introduces and explains the Italian mortality data used for our analysis and the creation of the new Index of Multiple Deprivation. Section3 describes how NNs are used in our study. Section4 briefly introduces the models used for comparison. Then, Sect.5 shows the results of the empirical analysis and, finally, Sect.6 concludes. 2 Data The collected mortality data comprises the ‘number of deaths’ and ‘exposure-at-risk years’ for males and females aged 50 to 95 living in Italy. It spans over the calendar years 1982–2018 and is granular to the level of the 106 Italian counties (called provinces). Our main source of data is the local office of national statistics (ISTAT).1 The data were further processed to account for splits in some counties over the last thirty years (see Euthum [10] for details). Further elaborations were needed to calculate the central exposure to risk from the population data given for January 1st of each year.2 We assume that the net migration effect does not bias our findings about mortality rates.3 Furthermore, we collected a set of twelve indicators of socio-economic characteristics for each county, which we used to create an index of multiple deprivation. For further details regarding peculiarities and adjustment of data, we refer to the supplementary material4. 2.1 Index ofmultiple deprivation For each county, we originally collected a set of twelve indicators representing different aspects of the quality of life. Out of these indicators, a subset of five was ultimately selected, reflecting a proper mix of different types of socio-economic factors. These have been aggregated to a so-called ‘Index of Multiple Deprivation’ (IMD). In this way, the 106 provinces5 from twenty regions could be classified to different groups based on their socio-economic indicators. The chosen variables are: 1. Relative poverty (in percent): The percentage of households with a consumption expenditure lower than the average per-capita, as estimated by ISTAT 6; 1 The population data was collected from the website www. istat. it. The number of deaths from 1982 to 2002 were provided by the office of statistics. 2 To approximate the central exposure for year i, populations from January 1st of year i and year i+1 were averaged. Further, some data cleaning was performed, for details we refer to Euthum [10]. 3 This assumption appears reasonable on the grounds of the relatively short length of the time series of available data, especially if we consider that censuses are conducted every ten years only, and also in consideration of the age-range we analyse, since migration (in general and so) between Italian provinces typically affects younger people. A detailed account of this observation is discussed in Cairns etal. [5] based on population data from England and Wales. Moreover, we are not aware of the existence of data that allows to track migration between Italian provinces. 4 See the corresponding Github repository to this paper. 5 ‘Sardegna’ has been kept to four provinces instead of five to keep things simpler for historic data. 6 Accessible at https:// www. istat. it/ en/ analy sisandprodu cts/ datab ases/ statb ase
499 1 3 A neural network approach forthemortality analysis ofmultiple… 2. Primary care and residential and semi-residential facilities (measured in beds per 10,000 inhabitants). This information is indicative of the expenditure in health care by the region where the county is located, which in its turn is affected by its wealth7; 3. Social services and benefits of municipalities, measured as costs in Euros per capita (2017 data). Services include, e.g., day nursery, socio-educational services for early childhood, and so on. It is assumed that higher costs correspond to more investments in social services by municipalities and hence more benefits from a sociological point of view for the population; 4. Unemployment rate (in percent for the population aged over 15); 5. Number of felonies committed by persons convicted by final judgement (per 1,000 inhabitants). From a statistical point of view, this variable is not significantly correlated with the other indicators used for building the index. It is assumed that a high relative number of felonies committed (by persons convicted by final judgement) suggests worse living conditions from a sociological point of view, but also comes along with a rather deficient economic situation. A detailed description of these variables can be found on the ‘Statbase portal’ of Istat, which provides access to a large amount of data on the Italian population,8 which is also the source of the underlying data (downloaded on 1st of March 2021 for year 2018). The variables have been chosen based on a correlation analysis between them. The selection was then validated by a ranking measure, Kendall’s tau, to check whether the rank of provinces changes significantly when omitting a variable from the selection. We found that Kendall’s tau does not significantly change when omitting or adding a further variable to the five we selected. Our index was constructed using z-scores to scale the five different variables to a comparable level and unit; the same approach is used in Osservatorio della salute [24]. This allowed to aggregate the single z-scores per province to a total value and finally permitted a ranking of the provinces. For each province, the z-score was calculated as where i indicates the respective socio-economic classifier, x j i its value, j the county, mi the mean of this classifier over all counties, and si the (unbiased) sample deviation over all counties. The interpretation is intuitive: the higher the value zj of a county, the worse is its socio-economic situation, or in other words, the deprivation in that area is higher. Implicitly, we assume that the standardized value of each variable has the same impact on the z-score (which could be easily generalized by introducing weights). z j= 5 ∑ i=1 zj i = 5 ∑ i=1 x j i −mi s i , 7 Since our index is created with the interpretation of ‘smaller values corresponding with better living conditions’, we changed the sign of this covariate and also the next one. 8 Accessible at https:// www. istat. it/ it/ datianali si-eprodo tti/ banchedati/ statb ase
500 M.Euthum et al. 1 3 Then, the aim is to rank the counties on the basis of the values of their z-score. Hence, we created nine groups of counties, each with homogeneous populations, ranging between 6 and 7 million people. This split does not account for sex and age of the Italian population. This implicitly assumes that the counties have similar distributed population by age and sex. If, by using this criterion, we found that two counties had the same z-score value,9 we allocate these in the same group. In this way, we aim at creating groups of comparable size, where the counties therein have similar socio-economic characteristics. The use of socio-economic indicators for the calendar year 2018 only, implies that the ranking and the corresponding groups do not change over time. This assumption may not reflect the socio-economic developments in Italy across the last 36 years. Nevertheless, we find this assumption reasonable in consideration of the results obtained in the preliminary analysis of the mortality and socio-economic data. Indeed, more deprived socio-economic counties are located in the south of Italy, compared to the historically wealthier areas in the north, as also demonstrated by the time series of the unemployment rate, and this evidence is further confirmed when analysing the evolution of the raw mortality rates. We purposely opted for a simple and interpretable approach to create the index, that might be refined in further studies. Our main focus is to elaborate on mortality models as described later, and this simple and intuitive approach provides very plausible results. In any case, we remark how the creation of the groups was carried out on the basis of their order, rather than their index value. To obtain a geographical impression of the index-based subdivision, the following map is reported, see Fig.1. The color-scheme is as follows: Brighter colors indicate a lower index value, reflecting better socio-economic condition based on the IMD defined above. The original map was taken from the De Agostini website10 and colored with standard graphic tools. To obtain a first impression of mortality rates, the crude period death rates m (x,t,i)= d(x,t,i) E c (x,t,i) between 1982 and 2018 are plotted in Fig. 2 for the male and female population aged 68 and 83 years old, respectively. We observe: • In general, mortality rates decrease over time and increase with age; which is consistent with the literature and the biological ageing process; • Mortality rates for females are lower than for males, as widely observed in many other national population tables; • The most deprived subpopulations (g1 in dark blue) appear to have higher mortality rates over the period analysed. The difference and ordering in subgroups is more pronounced for the female subpopulation, while for males, the difference in mortality rates is less evident for wealthier subgroups. This could be due to the chosen index or due to the underlying population. This effect may also be the consequence of a significant north–south division in terms of socio-economic well-being, while people still living longer in southern parts of the country, as 9 As per two decimal digits. 10 https:// blog. geogr afia. deasc uola. it/ artic oli/ chefinefaran noleprovi nceitali ane
501 1 3 A neural network approach forthemortality analysis ofmultiple… it is the case of regions like Sardinia and Calabria, which are famous for their exceptional longevity (see Poulain etal. [29]); • In the earlier years of the study, deprivation trends and differences are harder to detect. This may be caused by the fact that the socio-economic analysis Fig. 1 Subdivision of the Italian provinces based on the new index. We observe the tendency of northern provinces having smaller index values, i.e. better living conditions
502 M.Euthum et al. 1 3 was based on indicators of the year 2018, while different provinces evolved differently over decades; • The spike in the mortality rates for males and females aged 83 in 2003 is likely to be due to the massive heatwave of the summer in 2003 (Johnson etal. [15]). Further discussion is included in Sect.5. Throughout the analysis, we use the mortality rates of the first 33 calendar years 1982–2014 when training the models (training set), while those of 2015–2018 are used for predictions and forecasts (test set). We opted for this split in about 90%–10% due to the short length of the time-series of available data. We believe this split is a reasonable compromise between the need of sufficient data to fit the models (especially on a restricted age-range of a 50–95 years old population), whilst using the latest years to assess their predictive ability. −4.8 −4.4 −4.0 1990 2000 2010 Year Death Rate (Log) Group g1 g2 g3 g4 g5 g6 g7 g8 g9 Age 68, Female −4.0 −3.5 −3.0 1990 2000 2010 Year Death Rate (Log) Group g1 g2 g3 g4 g5 g6 g7 g8 g9 Age 68, Male −2.8 −2.4 −2.0 1990 2000 2010 Year Death Rate (Log) Group g1 g2 g3 g4 g5 g6 g7 g8 g9 Age 83, Female −2.50 −2.25 −2.00 −1.75 1990 2000 2010 Year Death Rate (Log) Group g1 g2 g3 g4 g5 g6 g7 g8 g9 Age 83, Male Fig. 2 Crude death rates in log-scale for ages 68 resp. 83, female and male population. Group 1 (g1) is the worst group socio-economically speaking
509 1 3 A neural network approach forthemortality analysis ofmultiple… 4.1 Li andLee (LL) model This model extends the single-population Lee and Carter [17] model to the analysis of multiple populations. The Li and Lee [18] model introduces a set of parameters which are common to the set of analysed groups, as well as group-specific parameters capturing the unexplained variance. Model 4.1 (Li and Lee model) For age (in years) x, time period t, and group i, the Li & Lee model describes the logarithm of the ‘central mortality rate’ m(x,t,i) as Here, m(x, t, i) can be approximated as the ratio between the death counts, denoted as D(x,t,i), and the central exposure at risk Ec(x,t,i) . 𝛼(x,i) , 𝛽(x,i) , and 𝜅(t,i) are group-specific parameters. 𝛼(x,i) indicates the average over time of the log mortality rate. The common function K(t) explains the evolution of the mortality over time for all groups, and B(x) is a global age modulating parameter, indicating how rates change by age for changes in the time factor K(t). 𝛽(x,i) and 𝜅(t,i) have the same role as B(x) and K(t), but act on a group-specific level. 4.2 Common age effect (CAE) model The CAE model of Kleinow [16] assumes that age effects are common to all populations, following the assumption that age effects may be very similar in countries sharing a similar socio-economic structure. Model 4.2 (Kleinow model) For age x, time period t, and group i, the Kleinow model of order p assumes the following model for the logarithm of the central death rate The order p follows from the allowance of further age–time interaction parameters. In our analysis we set p=2 . 4.3 Plat model The third model we use for comparison was initially proposed by Plat [28]. However, in this work we use a simplified version, which includes two period-specific factors, without accounting for the cohort effect. Model 4.3 (Plat model) For age x, time period t, and group i, the Plat model (without cohort effects) models the logarithm of the central death rates as (4.1) log ( m(x,t,i) ) =𝛼(x,i)+B(x)K(t)+𝛽(x,i)𝜅(t,i),t min ≤t≤t min +T− 1. (4.2) log ( m(x,t,i) ) =𝛼(x,i)+𝛽 (1) (x)𝜅 (1) (t,i)+…+𝛽 (p) (x)𝜅 (p) (t,i) .
510 M.Euthum et al. 1 3 where x denotes the average age of the observed age range. The first stochastic component 𝜅(1) represents changes in the level of mortality for all ages, while 𝜅(2) allows for changes in mortality to vary between ages. 4.3.1 Forecasting For the models of Li & Lee, Kleinow, and Plat, we performed forecasts via a classical time series approach. In fact, time dependent 𝜅 -processes are modelled as a stochastic time series to predict mortality rates through ARIMA (Auto Regressive Integrated Moving Average) models. These can be readily implemented in R, e.g. by using the function auto.arima from the package forecast (Hyndman and Khandakar [14]). 5 Empirical results All models are fitted based on data spanning from 1982 to 2014. However, fitted values are just compared for the years 1992 to 2014, since the NN-based approach does not deliver values for the first ten years, see Richman and Wüthrich [33]. In what follows, then i∈{1, …,9},x∈{50, …, 95} , and t∈{1992, …, 2014} . Furthermore, we graphically inspect the models by using the standardized residuals, as defined in Wen etal. [39], see also Table4. 5.1 In‑sample fit We first compare the three competing stochastic mortality models, that we briefly introduced in Sect.4, in terms of their explanation ratio. Then we compare these to the rates obtained by using the NN models of this paper. Table1 shows the explanation ratios for the models of Li & Lee, Kleinow, and Plat, defined for group i and model M as where 𝛼 c(x,i) ∶= 1 T ∑ t log d(x,t,i) Ec(x,t,i ) is the average log crude death rate over time. The explanation ratio is useful for analysing how much information about the crude death rates d(x,t,i) E c (x,t,i) is explained by the respective model. We observe how the Li & Lee model performs best in terms of explanation ratios for eight out of nine females subgroups, and the same is noted for the (4.3) log ( m(x,t,i) ) =𝛼(x,i)+𝜅 (1) (t,i)+𝜅 (2) (t,i)(x−x) , (5.1) R M i=1− ∑ x,t � log d(x,t,i) Ec(x,t,i)−log (m(x,t,i)) �2 ∑ x,t� log d(x,t,i) Ec(x,t,i)−𝛼c(x,i) � 2 ,
511 1 3 A neural network approach forthemortality analysis ofmultiple… Kleinow model when looking at the male population. The Plat model, analysed at the level of a single population, still yield comparable explanation ratios, despite being lower compared to the Li & Lee and the CAE model of Kleinow. Furthermore, the Plat model has a lower number of parameters, which may in part explain these results. Our results are further confirmed by the analysis of their Akaike Information Criterion (AIC) (Akaike [1]), shown in Table2, which indicates the relative quality of a statistical model with a penalty term for the number of parameters. The AIC is calculated as where 𝓁( 𝜃) indicates the log-likelihood value at (optimal) parameter 𝜃 and p denotes the number of parameters of the respective model (see Table9). Tables3 shows the mean squared error for model M and subpopulation i, calculated as follows: where m(x,t,i)M denotes the fitted mortality rate derived from model M. Table 4 shows the MSE for the RNN models analysed in this paper. (5.2) AIC =− 2 ⋅𝓁( 𝜃)+ 2 ⋅ p, (5.3) MSE M i= 1 n⋅T ∑ x,t ( m(x,t,i)− m(x,t,i)M ) 2 , Table 1 Explanation ratios for different models The highest ratios for a specific sex are emphasized in bold Deprivation Female Male Combined Group RLL i RCAE i RPlat i R LL i RCAE i RPlat i RLL i RCAE i RPlat i 10.862 0.853 0.848 0.782 0.823 0.811 0.881 0.867 0.874 20.868 0.856 0.838 0.822 0.810 0.804 0.877 0.854 0.870 30.862 0.857 0.822 0.842 0.846 0.826 0.893 0.879 0.875 40.859 0.848 0.802 0.855 0.856 0.836 0.892 0.881 0.872 50.831 0.831 0.774 0.868 0.872 0.842 0.886 0.879 0.852 60.876 0.866 0.826 0.896 0.903 0.882 0.917 0.910 0.901 7 0.876 0.882 0.842 0.888 0.912 0.894 0.929 0.924 0.917 80.866 0.857 0.826 0.880 0.890 0.864 0.910 0.902 0.895 90.866 0.861 0.831 0.885 0.898 0.866 0.914 0.906 0.909 Average 0.863 0.857 0.823 0.857 0.868 0.847 0.900 0.889 0.885 Table 2 AIC values for different models Lowest values for specific sex in bold Model Female Male Both sexes Li and Lee -66,185,533 -67,260,680 -134,742,870 Kleinow -66,188,446 -67,256,443 -134,750,703 Plat -66,194,573 -67,261,519 -134,749,793
512 M.Euthum et al. 1 3 Table 3 Mean squared errors for different models and groups Lowest values for specific sex in bold. Scale: 10 − 4 Deprivation Female Male Combined Group LL CAE Plat LL CAE Plat LL CAE Plat 11.3327 1.4275 1.3743 3.6259 3.0236 2.7788 1.4856 1.6027 1.4515 20.9692 1.2805 1.2509 2.0578 1.8407 1.8716 1.1946 1.2721 1.1912 3 0.7436 0.7396 0.8416 2.4992 1.9319 2.0930 0.9395 0.9666 0.9221 40.6853 0.8357 0.8987 2.8050 2.1437 2.3239 0.8530 0.9327 0.9533 50.8433 0.9518 1.0570 2.4233 2.1329 1.8905 0.8928 1.0593 1.0390 60.5896 0.6464 0.7379 2.6853 2.3437 2.1192 0.7630 0.8736 0.7749 70.6317 0.6708 0.7788 3.7611 3.6867 2.8752 0.7681 0.8645 0.7284 80.7108 0.7783 0.7892 3.4809 2.4448 2.6673 0.8677 1.0031 0.8662 90.6903 0.7570 0.8233 4.4645 2.8037 3.3212 0.8259 0.9085 0.7876 Sum 7.1964 8.0876 8.5518 27.8031 22.3517 21.9406 8.5902 9.4831 8.7140 Table 4 Mean squared errors for different Recurrent network models and groups Lowest values for specific sex and network type in bold. Scale: 10 − 4 Deprivation Female Male Combined Group ELSTM3 i ELSTM2 i ELSTM1 i ELSTM3 i ELSTM2 i ELSTM1 i ELSTM3 i ELSTM2 i ELSTM1 i 1 0.6477 0.6242 0.6495 1.1999 1.1560 1.2223 0.5309 0.5071 0.5362 2 0.5710 0.5521 0.5907 0.6545 0.6475 0.6853 0.4347 0.4295 0.4485 3 0.4006 0.3876 0.4139 0.8020 0.7556 0.8101 0.3449 0.3294 0.3451 4 0.3114 0.2992 0.3196 0.6739 0.6364 0.6783 0.2379 0.2245 0.2413 5 0.4202 0.4039 0.4270 0.8119 0.7816 0.8481 0.3581 0.3511 0.3709 6 0.3363 0.3257 0.3411 0.7937 0.7621 0.8044 0.3046 0.2900 0.3096 7 0.3213 0.3123 0.3215 0.9426 0.8860 0.9429 0.2753 0.2599 0.2771 8 0.2922 0.2781 0.3000 0.8838 0.8420 0.8981 0.2760 0.2591 0.2801 9 0.2979 0.2827 0.3007 0.7111 0.6564 0.7067 0.2772 0.2580 0.2766 Sum 3.5986 3.4658 3.6639 7.4734 7.1236 7.5963 3.0395 2.9086 3.0853 EGRU3 i EGRU2 i EGRU1 i EGRU3 i EGRU2 i EGRU1 i EGRU3 i EGRU2 i EGRU1 i 10.4052 0.4219 0.5124 0.5439 0.5636 0.8325 0.3051 0.3396 0.3979 20.3478 0.3599 0.4389 0.3925 0.4212 0.5785 0.2744 0.2982 0.3658 30.2756 0.2951 0.3586 0.4108 0.4216 0.5702 0.2111 0.2458 0.2905 40.2633 0.2686 0.2928 0.3824 0.3978 0.5464 0.1955 0.2099 0.2358 50.3323 0.3449 0.3901 0.4831 0.5321 0.6863 0.2736 0.3038 0.3401 60.2736 0.2774 0.3059 0.4036 0.4290 0.5686 0.2219 0.2499 0.2778 7 0.2577 0.2562 0.2752 0.4324 0.4461 0.6428 0.2101 0.2247 0.2450 80.2291 0.2327 0.2633 0.4676 0.4806 0.6703 0.1826 0.2038 0.2341 90.2387 0.2398 0.2612 0.4237 0.4296 0.5115 0.1967 0.2160 0.2408 Sum 2.6232 2.6966 3.0984 3.9399 4.1217 5.6069 2.0710 2.2918 2.6278
513 1 3 A neural network approach forthemortality analysis ofmultiple… From the results in Table3, we observe that it is not possible to draw a decisive conclusion about the best performing model, since for different groups and sex there is a mixed evidence about the best performing model. When the data for both males and females are combined, then we observe that the Plat model has a better performance in the two extremes of the deprivation group, or in other words, shows a better performance for the least and most deprived county groups. The two RNN models seem to sensibly improve the in-sample fit compared to the three competing stochastic models. In more detail, the two-layer LSTM outperforms the other two LSTM models over all socio-economic groups for males, females, and combined. On the other hand, the GRU models with three layers turn out to perform even better. For female and both sexes combined mortality rates, the most deprived groups show higher in-sample errors compared to less deprived ones. A similar evidence can be observed for the LL, CAE, and Plat models. This may be indicative of the issues of mortality models in general when fitting more deprived subpopulations via a multi-population approach. This can also be the effect of the created index in Sect.2 and of the underlying data. Nevertheless, we note that the difference between females and males is smaller for the RNN approach compared to the competing models. In conclusion, the multi-population RNN models outperform the well established stochastic mortality models when analysing county-based socio-economic subgroups of the Italian population, based on the mean squared error. The mean squared errors for males are higher compared to females and both sexes combined. For a deeper inspection of these results we restrict the MSE analysis to a shorter age range for males and females, in a way such that they both have the same life expectancy. These ranges were chosen based on remaining life expectancy of Italian males at age 50 in 2017 (32.08 years) and Italian females at age 54 in year 2017 (approx. 32 years). Therefore, based on the Italian life tables from the Human MortalityDatabase [13], we restrict the analysis to the male population aged 50 to 82 and to the females aged 54 to 86. Table 5 Mean squared errors for specifically tailored age ranges, females (54–86), and males (50–82), SMM models Scale: 10 − 4 Deprivation Female Male Group LL CAE Plat LL CAE Plat 1 0.0996 0.0982 0.1025 0.1113 0.1117 0.1147 20.0855 0.0904 0.0925 0.1004 0.0969 0.1026 30.0795 0.0851 0.0828 0.1006 0.0985 0.1035 40.0859 0.0875 0.0886 0.0899 0.0869 0.0922 5 0.0915 0.0884 0.0885 0.1052 0.1036 0.1059 60.0770 0.0792 0.0800 0.1178 0.1169 0.1189 70.0790 0.0806 0.0825 0.1081 0.1067 0.1067 80.0804 0.0867 0.0853 0.1098 0.1059 0.1090 90.0782 0.0793 0.0802 0.1072 0.1097 0.1115 Sum 0.7566 0.7755 0.7830 0.9503 0.9367 0.9650
514 M.Euthum et al. 1 3 Again, we recalculated the deprivation group-specific MSE for males and females for the restricted age-ranges, still based on the model fitted to the original dataset (males and females aged 50 to 95). The results are shown in Table5 (LL, CAE, and Plat models) and Table6 (for RNN). 5.2 Mean squared errors forspecific ages For the selected age ranges, we observe that for the three competing stochastic mortality models, the mean squared error is less than 0.1 of the mean squared error for all ages, even if these reduced age ranges cover 72% of the modelled years. Again, we observe that for the female population the MSE is larger for the most deprived socio-economic groups for both the competing models as well as for the RNN-based ones. A similar evidence is obtained for the males at a lesser extent. For the models fitted using female subpopulation data, we still face more difficulties when considering the most deprived groups compared to less deprived ones, especially for SMM models. It is also evident, that all models within one model class (SMM or RNN) have similar mean squared errors. This means, the models perform very well for these ages through all selected models. Let us also mention the fact that mean squared errors in the RNN case are by far smaller than in the SMM case. Furthermore, it is interesting to observe an alignment of female and male rate fitting (as it was aimed through the analysis of different age ranges): the SMM mean squared error for the male population were on average approximately 3.02 times compared to females. When restricting the age range observation, this factor reduces to about 1.23, which means that for our data both male and female rates are almost fitted equally well on average for an age range with similar remaining life expectation. For the three selected RNN models from above, the factors are 1.52 and 1.01 respectively, that is the alignment of female and male rates fitting ability is almost perfect in the RNN case. Table 6 Mean squared errors for specifically tailored age ranges, females (54–86), and males (50–82), selected RNN models Scale: 10 − 4 Deprivation Female Male Group GRU3 GRU2 LSTM1 GRU3 GRU2 LSTM1 1 0.0322 0.0318 0.0327 0.0287 0.0289 0.0293 2 0.0186 0.0185 0.0195 0.0193 0.0199 0.0189 3 0.0209 0.0211 0.0211 0.0173 0.0178 0.0178 4 0.0163 0.0164 0.0167 0.0159 0.0165 0.0154 5 0.0169 0.0168 0.0171 0.0171 0.0173 0.0180 6 0.0153 0.0151 0.0153 0.0184 0.0187 0.0181 7 0.0188 0.0185 0.0195 0.0208 0.0207 0.0199 8 0.0122 0.0122 0.0124 0.0157 0.0164 0.0162 9 0.0174 0.0173 0.0173 0.0167 0.0173 0.0172 Sum 0.1689 0.1677 0.1715 0.1700 0.1735 0.1707
515 1 3 A neural network approach forthemortality analysis ofmultiple… Concluding, it seems that the oldest people (of age 87–95) drive most of the mean squared error. In addition, one should be aware of looking at different age ranges for females and males based on the evidence that life expectancy significantly differs for these two groups. The point in time for older ages, where mortality is harder to fit due to more volatility in death and exposure data, starts earlier for males which could explain a higher mean squared error when looking at all ages. Another possible reason for larger errors in the male population through all models may come from the difference in population sizes of the underlying deprivation groups. These are higher in the female case, since in the observed age ranges females outlive males. A larger sample size could give the statistical model more stability when estimating parameters. However, the difference is not very large15. 5.2.1 Standardized residuals We perform a graphical analysis of these results by investigating the standardized residuals. When a model fits the data reasonably well, standardized residuals should not exhibit any pattern based on years or ages. To quote Cairns etal. [4], “if the model fits the data well, then the standardized residuals should be independent of 50 60 70 80 90 1980 1990 2000 2010 Year Age −5 0 5 10 Res Female − LiLee − Group 7 50 60 70 80 90 1980 1990 2000 2010 Year Age −5 0 5 10 Res Male − Kleinow − Group 3 Fig. 5 Selected residual heatmap plots per group ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● −2.5 0.0 2.5 5.0 50 60 70 80 90 Age Residuals 1990 2000 2010 Year Female − LiLee − Group 8 ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● −0.3 −0.2 −0.1 0.0 0.1 0.2 50 60 70 80 90 Age Residuals 1995 2000 2005 2010 Year Male − LSTM3 − Group 6 Fig. 6 Selected residual plots function of age, per group, both sexes 15 Of the population in Italy aged 50–95 on January 1st 2021, 46.31% were males (Istat Statbase).
516 M.Euthum et al. 1 3 each other, meaning that the heat plot should exhibit a high degree of randomness, with no discernible patterns”. Figure5 plots the heatmaps of the standardized residuals for the fitted rates under the Li and Lee model for group 7 and for the group 3 of the CAE model. The plots for the other groups under the three SMMs are similar. These are calculated for model M as where m(x,t,i)M denotes the fitted mortality rate under model M. The two plots show a diagonal line, which is indicative of a cohort effect for those individuals born around 1918, which corresponds to the end of the first world war and the Spanish flu pandemic. We discuss these points in more detail in Appendix A.2. For all other ages and years, the standardized residuals appear to lack any specific pattern (Figs.6, 7). For RNN models, the standardized residuals shown in Fig.8 (for the other groups we have similar evidences) seem to be spread homogeneously across the ages and years. There are no visible cohort effects in residuals as in earlier models and residuals seem to be much smaller than in the previous models. These observations suggest that RNN models used in this work are able to capture cohort effects in the data and yield a better in-sample fit for observed years. In any case, we observe that for ages 70+/80+ in 2003, residuals are large and positive, meaning that the observed mortality is higher than expected. Conversely, for 2004 mortality rates seem to be over-estimated by the models. A possible explanation for this under-estimation in 2003 could be the massive European heat wave in summer 2003, see Johnson etal. [15], the hottest summer in Europe for centuries. ISTAT reported over 18,000 deaths in that summer compared to the year before (+ 11.6%). As reported in Mattone [21], 91.8% of these were aged 75 and older. This could explain in part the large number of negative residuals at older ages in 2004. These extraordinary circumstances may explain why the models could not fully detect such a sudden increase in the number (5.4) Z (x,t,i)M=d(x,t,i)−E c (x,t,i)⋅m(x,t,i) M √ Ec(x,t,i)⋅m(x,t,i)M , −4 −2 0 2 4 6 1990 2000 2010 Year Residuals 50 60 70 80 90 Age Male − LiLee − Group 9 −0.3 −0.2 −0.1 0.0 0.1 0.2 0.3 1995 2000 2005 2010 2015 Year Residuals 50 60 70 80 90 Age Female − GRU2 − Group 5 Fig. 7 Selected residual plots function of year, per group, both sexes
517 1 3 A neural network approach forthemortality analysis ofmultiple… of deaths. A similar result might be observed in 2020 or later following the COVID19 pandemic. We further compare the residuals as function of the age (Fig.6), and as a function of the calendar year (Fig.7), where groups and sex have been chosen randomly and 50 60 70 80 90 1995 2000 2005 2010 2015 Year Age −0.2 0.0 0.2 Res Female − LSTM1 − Group 1 50 60 70 80 90 1995 2000 2005 2010 2015 Year Age −0.2 −0.1 0.0 0.1 0.2 Res Female − GRU3 − Group 8 50 60 70 80 90 1995 2000 2005 2010 2015 Year Age −0.2 −0.1 0.0 0.1 0.2 Res Male − GRU1 − Group 6 50 60 70 80 90 1995 2000 2005 2010 2015 Year Age −0.2 −0.1 0.0 0.1 0.2 Res Male − LSTM2 − Group 3 Fig. 8 Randomly selected residual heatmap plots per group Table 7 Out-of-sample mean squared errors for different models and groups Lowest values for specific sex in bold. Scale: 10 − 4 Deprivation Female Male Both Sexes Group ELL i EKL i EPlat i ELL i EKL i EPlat i ELL i EKL i EPlat i 1 1.0038 0.9955 2.1453 7.7344 4.2412 4.8872 1.1059 1.0278 2.5211 20.6560 0.7621 1.4115 2.3270 0.7531 1.4893 0.7326 5.1101 1.1578 30.7433 1.1019 1.9106 3.2611 2.1388 1.7976 1.0087 9.9613 1.4781 40.5366 0.6045 16.3794 9.7053 1.8354 16.1711 8.8701 0.8926 13.5179 5 1.0720 0.9237 2.3226 2.8685 2.9456 3.2342 0.8967 1.1394 2.0517 6 1.4249 1.3634 2.7108 3.5877 4.0032 2.2568 1.3117 1.5394 2.0675 70.9823 1.9311 2.0970 2.8236 6.2054 2.3788 1.0659 0.9724 1.7426 8 1.2298 1.1849 2.3117 3.9409 3.5925 1.6364 1.0616 1.4791 1.7348 91.3697 2.2630 2.1508 5.4839 2.3254 1.3812 1.2547 1.1114 1.5911 Sum 9.0183 11.1301 33.4397 41.7324 28.0406 35.2326 17.3079 23.2334 27.8627
518 M.Euthum et al. 1 3 where we exclude the data for the cohorts born in 1916, 1917, and 1920, based on observations in Appendix A.2. Patterns are not observable if not for yellow points which indicate age 90+. In general, residuals for the Li and Lee model (the same holds for the other models of type SMM) spread on much higher levels than those of RNN models, by a factor larger than ten. However, both model types suggest evenly distributed residuals with exception for the oldest years which seem harder to model. Presumably, this latter observation is driven by the smaller exposure for older ages. 5.2.2 Out‑of‑sample measures We analyse the out-of-sample performance of the implemented models by using a forecast period of four years, with n=46 and T=4 . All figures have been obtained on the basis of the procedure described in Sect. 3 for RNN and Sect. 4 for the SMMs. Results are shown in Tables7 and8. Table 8 Out-of-sample mean squared errors for different RNNs and groups Lowest values for specific sex and network type in bold. Scale: 10 − 4 Deprivation Female Male Both sexes Group ELSTM3 i ELSTM2 i ELSTM1 i ELSTM3 i ELSTM2 i ELSTM1 i ELSTM3 i E LSTM2 i ELSTM1 i 1 1.0916 0.9873 1.0225 2.0643 2.1807 2.0877 1.3548 1.3108 1.2765 2 0.6805 0.6004 0.6265 1.0510 1.1554 1.1369 0.7402 0.7240 0.7030 3 0.7193 0.6268 0.6695 0.9951 1.0019 0.9764 0.6845 0.6732 0.6642 4 0.7020 0.6073 0.6596 1.1020 1.1364 1.0489 0.7590 0.7397 0.6967 5 0.7966 0.6839 0.7631 1.1623 1.2082 1.0256 0.8945 0.8979 0.8581 6 1.2054 1.0858 1.1674 1.3938 1.5168 1.4414 1.1915 1.1963 1.1657 7 0.8209 0.7249 0.8058 1.1760 1.2602 1.1435 0.7999 0.8020 0.7914 8 1.0526 0.9342 0.9979 1.2587 1.2837 1.1041 1.0766 1.0483 0.9947 9 1.0315 0.8979 0.9617 1.2335 1.3099 1.2287 1.0079 0.9829 0.9457 Sum 8.1005 7.1485 7.6741 11.4367 12.0532 11.1932 8.5090 8.3750 8.0959 EGRU3 i EGRU2 i EGRU1 i EGRU3 i EGRU2 i EGRU1 i EGRU3 i E GRU2 i EGRU1 i 1 0.9718 0.8977 1.1457 2.1442 2.2496 2.2774 1.1988 1.1511 1.2861 2 0.5922 0.5266 0.7010 1.1787 1.1653 1.1986 0.6606 0.6301 0.7223 3 0.6240 0.5631 0.7322 1.0778 1.0036 1.0708 0.6095 0.5845 0.6802 4 0.5709 0.5095 0.6293 1.0881 1.1330 1.1358 0.6224 0.5863 0.6684 5 0.6083 0.5181 0.7028 1.3581 1.4528 1.2105 0.6910 0.6191 0.7283 6 1.0879 1.0308 1.1805 1.5759 1.6477 1.6592 1.1063 1.0329 1.1135 7 0.6631 0.5915 0.6908 1.1824 1.2642 1.2480 0.6614 0.6132 0.6786 8 0.8777 0.7958 0.9119 1.3260 1.4557 1.3781 0.9272 0.8538 0.9174 9 0.8758 0.8304 0.8947 1.3269 1.4161 1.3662 0.9129 0.8486 0.8722 Sum 6.8717 6.2636 7.5890 12.2580 12.7880 12.5448 7.3902 6.9196 7.6669