Neural networks for the joint development of individual payments and claim incurred
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Delong, Łukasz; Wüthrich, Mario V. Article Neural networks for the joint development of individual payments and claim incurred Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Delong, Łukasz; Wüthrich, Mario V. (2020) : Neural networks for the joint development of individual payments and claim incurred, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 8, Iss. 2, pp. 1-34, https://doi.org/10.3390/risks8020033 This Version is available at: https://hdl.handle.net/10419/257988 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
risks Article Neural Networks for the Joint Development of Individual Payments and Claim Incurred Łukasz Delong 1,* and Mario V. Wüthrich 2 1Institute of Econometrics, SGH Warsaw School of Economics, Niepodległo´sci 162, 02-554 Warsaw, Poland 2Department of Mathematics, ETH Zurich, RiskLab, Rämistrasse 101, 8092 Zurich, Switzerland; [email protected] *Correspondence: [email protected].pl Received: 27 February 2020; Accepted: 2 April 2020; Published: 7 April 2020 Abstract: The goal of this paper is to develop regression models and postulate distributions which can be used in practice to describe the joint development process of individual claim payments and claim incurred. We apply neural networks to estimate our regression models. As regressors we use the whole claim history of incremental payments and claim incurred, as well as any relevant feature information which is available to describe individual claims and their development characteristics. Our models are calibrated and tested on a real data set, and the results are benchmarked with the Chain-Ladder method. Our analysis focuses on the development of the so-called Reported But Not Settled (RBNS) claims. We show benefits of using deep neural network and the whole claim history in our prediction problem. Keywords: neural networks; individual claims; reported but not settled claims; claims simulations 1. Introduction Stochastic models for individual claims reserving were introduced roughly 30 years ago in the work of (Arjas 1989) and (Norberg 1993 1999). These papers introduce the stochastic framework of marked point processes for individual claims modeling. However, they provide little guidance on statistical aspects and on the application of these models to practical problems. Surprisingly little progress has been made in using these approaches for individual claims reserving. The proposed approaches are based on (semi-) parametric models, see (Larsen 2007;Taylor et al. 2008; Zhao et al. 2009 ; Jessen et al. 2011;Pigeon et al. 2013;Antonio and Plat 2014). Unfortunately, most of these proposals turn out to have too restrictive assumptions and they are not sufficiently practical in real data applications. Individual claims reserving has only started to become popular with the emergence of machine learning methods in insurance: some papers are based on the application of regression trees and gradient boosting techniques, see (Wüthrich 2018;Lopez et al. 2019;De Felice and Moriconi 2019; Duval and Pigeon 2019 ;Baudry and Robert 2019). Probably the most popular machine learning method, neural networks, has mainly been used on aggregate data, see (Gabrielli 2020;Kuo 2019). This seems surprising, because neural networks have gained a lot of attention in recent years due to their excellent performance on individual cases. A good introduction to neural networks and their application to insurance pricing can be found in (Goodfellow et al. 2016;Ferrario et al. 2018; Denuit et al. 2019 ). Neural networks served at developing an individual claim history simulation machine, see (Gabrielli and Wüthrich 2018), but these authors have not followed up on their method for the purpose of claims reserving. Moreover, (Gabrielli and Wüthrich 2018) only model claim payments but not claim incurred because of lack of the corresponding information. Clearly, claim incurred, together with past information about the claim development process, is an important information for Risks 2020,8, 33; doi:10.3390/risks8020033 www.mdpi.com/journal/risks
Risks 2020,8, 33 2 of 34 the prediction of future payments and should not be excluded a priori from regression models used for claim prediction. The goal of this paper is two fold. Firstly, we would like to jointly model the development of individual claim payments and claim incurred, and secondly, we would like to explore neural networks for this task because they seem to be particularly suited for this problem. Our analysis focuses on the development of the so-called Reported But Not Settled (RBNS) claims. One can expect that the development process of individual RBNS claims could be characterized with regression models since claims reported with different features and claim histories should generate different cash flows in time and amount. Consequently, regression models for individual claims should improve reserving methods and provide more detailed information about claim developments and ultimate losses in portfolios (e.g., across segments or claim types). In this paper, we develop regression models and postulate distributions which can be used in practice to describe the joint development process of individual claim payments and claim incurred. We apply neural networks to estimate our regression models. As regressors we use the whole claim history of incremental payments and claim incurred, as well as any feature information which is available to describe individual claims and their development characteristics. Due to a large number of regressors we would like to use in individual claim prediction, neural networks seem to be the most appropriate choice for estimating the regression functions. The fit of our regression models and distributions is tested on a real data set. We see five contributions of this paper. The first contribution are the regression models themselves. We point out that regression models which describe the development process of claims from individual policies are fundamentally different from regression models which have been used for aggregate data. Hence, it is not possible to simply move the well-known chain-ladder type models to individual claims reserving. The second contribution is that we allow for non-Markovian dynamics of the claim development process and we test the Markovian assumption on a real data set. The standard approach in claims reserving is to only use the most recent information about the claim developments, hence, to assume a Markovian structure of the claim development process. It turns out that, in our practical example, the Markovian assumption is too strong and it is beneficial, in terms of prediction accuracy, to use the whole claim history in our regression models. Thanks to the application of neural networks, the claim history can be efficiently used for predictions of future claims. The third contribution is to jointly model incremental payments and claim incurred. From the paper by (Quarg and Mack 2004) , it is suggested that both payments and claim incurred should be used for estimating the reserve for the outstanding claim liabilities. Intuitively, the linear regression structures proposed by (Quarg and Mack 2004) in their Munich chain-ladder model is not fully appropriate. Using the tools from neural networks, we can fit a better relation between future payments and claim incurred and past payments and claim incurred. To the best of our knowledge, there is only one paper by (Kuo 2019) which uses neural networks with simultaneous consideration of payments and claim incurred. However, (Kuo 2019) applies neural networks to aggregated payments and case reserves observed across all accident years and development years and multiple insurance companies. Consequently, claims from individual policies and models suited for individual claims reserving cannot be used. In (Kuo 2019), recurrent neural networks with a mean-squared error loss function are calibrated to the whole available history of aggregated payments and case reserves. However, the non-Markovian assumption of the claim development process is not discussed in any aspect and is not validated in the paper. Moreover, there is no clear distinction in (Kuo 2019) between random variables and their predictions, the latter typically being estimated expected values. This distinction is important, if these random variables are used as regressors in later periods. The fourth contribution of this paper is that we present a self-contained estimation procedure for our regression models and neural networks. We recommend to use the Combined Actuarial Neural Network (CANN) approach of (Schelldorfer and Wüthrich 2019) —we suggest to use generalized linear models (GLMs), generalized additive models (GAMs) or regression trees as initial predictions, or as the initial models, from which
Risks 2020,8, 33 3 of 34 we start training the neural networks. We use multinomial/binomial cross-entropy and Gamma loss functions for calibrations of our neural networks. These loss functions are also related to the distributions of the variables which we would like to use for predictions in the claim development process. The goal in the simulation of outstanding claims is to model not only the mean response but also the distribution of the response. This might be a real challenge in practice since simple distributions, which are common in regression problems and focus on the mean response, may not fit real data sufficiently well. We aim at correctly modeling at least the first two moments of the distribution and we focus on modeling the mean as well as the dispersion of a response with a continuous distribution. To model extreme events, we separate large claims from attritional claims and use Pareto distributions for claims above a high threshold. Finally, to the best of our knowledge this is the first paper in insurance data science where outstanding liabilities are predicted with neural networks and compared with classical chain-ladder estimates. This comparison on a real data set is the fifth contribution of this paper. The remainder of this paper is organized as follows. In Section 2we introduce all claim development models. In Section 3we discuss the estimation procedure. The numerical example on real data set is investigated in Section 4, and in Section 5we benchmark our results with the chain-ladder predictions. Finally, in Section 6we conclude. All calculations were done in Keras, which is an open-source API to TensorFlow. 2. Models for Individual Claims Development Let i∈ { 1, 2, . . .} denote the accident period of the occurrence date of an insurance claim. The accident period can be measured in days, weeks, months, quarters or years. Let j∈ { 0, 1, 2, . . .} denote the reporting delay after the claim occurrence date, measured in the same time units as the accident period i . Consequently, i+j denotes the reporting period (in days, weeks, months, quarters or years) of a particular claim. After a claim is reported, it develops over its settlement time. Let k∈ { 0, 1, 2, . . .} measure the development periods of a reported claim, initialized to the respective reporting date i+j. We investigate individual claims from individual insurance policies. Let Pi,j k denote the incremental payment in development period k for a claim of accident period i reported with delay j . The payment Pi,j k is made in calendar period i+j+k . Let Ii,j k denote the claim incurred at the end of development period k for a claim from accident period i reported with delay j . The claim incurred Ii,j k is observed at the end of calendar period i+j+k . We introduce the case reserve for an individual claim, which is a part of the claim incurred and corresponds to an individual claim amount estimate made by a claims adjuster. The case reserve at the end of development period k for a claim from accident period i reported with delay jis denoted by Ri,j k, and it is given by Ri,j k=Ii,j k− k ∑ l=0 Pi,j l,k=0, 1, 2, . . . . The variables Pi,j k , Ii,j k , Ri,j k should also be indexed with unique individual claim identifier (claim number), for notational convenience this index is omitted for the moment. At the reporting date of a claim we observe the first payment and the first evaluation of the total claim incurred, i.e., we have information (Pi,j 0 , Ii,j 0) . Please note that Pi,j 0 can also be zero if there no payment has been made in the initial development period k= 0. Having observed the values (Pi,j 0 , Ii,j 0) means that the time point i+j of the claim, reporting was fully observed and we aim at modeling the development (Pi,j k , Ii,j k) for all later time points k= 1, 2, . . . . Our goal is to study the development process of RBNS claims, i.e., claims for which we observed the initial state (Pi,j 0 , Ii,j 0) . In this paper we aim at modeling the two-dimensional process (Pi,j k , Ii,j k)k=1,2,... , conditionally given
Risks 2020,8, 33 4 of 34 the initial value (Pi,j 0 , Ii,j 0) . Since we are interested in individual claims reserving, the two-dimensional process (Pi,j k,Ii,j k)k=1,2,... is modeled for each single reported claim. Let us define a filtration (Ck)k=0,1,... which describes the history of payments and claim incurred on an individual claim. The individual claim filtration is defined by Ci,j k=σnPi,j s,Ii,j s; 0 ≤s≤ko,k=0, 1, 2, . . . , and it describes the history of payments and claim incurred for a claim from accident year i and reported with delay j . Moreover, to each individual claim we associate a vector of (static or dynamic) features, which we denote by zi,j k . E.g., the vector zi,j k may include any feature of the insurance policy such as the line of business involved, the age of the injured person, the claim type, the accident period, the reporting delay etc. The information included in the filtration Ci,j k−1 and the features zi,j k−1 are used as regressors (explanatory variables) in our regression problems to predict (Pi,j k , Ii,j k) in the next development period k. 2.1. Model 1: Occurrence of Payments and Changes in Claim Incurred In the first step we model the indicators which indicate whether the policy generates a new payment and/or the value of the claim incurred is changed in the next development period. The severities of the incremental payments and the change in claim incurred, if needed, are modeled in the next subsections. We define the indicator process (Ii,j k,Pi,j k)k=1,2,... as follows: Ii,j k=1{Ii,j k−Ii,j k−16=0}and Pi,j k=1{Pi,j k6=0}, (1) and we introduce a stochastic process (Yi,j k)k=1,2,... , which takes values in the set { 0, 1, 2, 3 } , by setting: Yi,j k=2Ii,j k+Pi,j k= 0 if Pi,j k=0 and Ii,j k=0, 1 if Pi,j k=1 and Ii,j k=0, 2 if Pi,j k=0 and Ii,j k=1, 3 if Pi,j k=1 and Ii,j k=1. (2) To model occurrences of payments and changes in claim incurred, we have to model the conditional probabilities for the sequence of random variables (Yi,j k)k=1,2,... . We should model the conditional probabilities since we expect the categorical distribution of Yi,j k in development period k to depend on the individual claim history Ci,j k−1and the individual claim feature zi,j k−1. We use a multinomial logistic regression to model these categorical conditional probabilities as follows: log PYi,j k=yCi,j k−1,zi,j k−1 PYi,j k=0Ci,j k−1,zi,j k−1 =fyCi,j k−1,zi,j k−1,y=1, 2, 3, k≥1, (3) where we use three appropriate regression functions f= (f1 , f2 , f3) . The regression model (3) is called Model 1. A presently well-understood approach is to use a generalized linear model (GLM) with a linear additive function fy , or a generalized additive model (GAM) with a non-linear additive function fy , and selected predictor variables from the set (Ci,j k−1 , zi,j k−1) . Alternatively, one can use regression trees where the best predictors from the set (Ci,j k−1 , zi,j k−1) are chosen as a part of the calibration and fy is estimated as a piecewise constant function on subspaces defined by the predictors; in the context
Risks 2020,8, 33 5 of 34 of individual claims reserving this was considered in (Wüthrich 2018;Lopez et al. 2019;De Felice and Moriconi 2019;Duval and Pigeon 2019). A more advanced approach, gaining popularity among actuaries, is to use a neural network with a non-linear non-additive function fy and including all predictor variables from the set (Ci,j k−1 , zi,j k−1) , this was considered in (Kuo 2019;Gabrielli 2020), however still on aggregated claims. We aim at fitting neural networks to individual claims. Let us recall the categorical cross-entropy loss function for a sample of size n of independent random variables v`,1 , v`,2 , v`,3 , v`,4`=1,...,n with categorical distribution (4 categories) and estimated probabilities ˆ p`,1,ˆ p`,2,ˆ p`,3,ˆ p`,4`=1,...,n, it is given by Dcat =−1 n n ∑ `=1 4 ∑ k=1 v`,klog(ˆ p`,k). (4) The categorical cross-entropy loss function is closely related to the deviance loss function of the multinomial model. The deviance is equal to 2 ∑n `=1∑4 k=1v`,klog(v`,k/ˆ p`,k) . We remark that minimizing the cross-entropy is equivalent to minimizing the deviance, which is performed when we fit a categorical GLM/GAM or a categorical regression tree model. If Yi,j k= 0, then we immediately know the values of the process (Pi,j k , Ii,j k) in the next development period k , because there is not any change in claim incurred and payments are equal to zero. Therefore, we only need to further model the cases Yi,j k∈ {1, 2, 3}. 2.2. Models 2 and 3: Claim Severities of Incremental Payments If Yi,j k= 1 or Yi,j k= 3, we have to model a non-zero payment in development period k for the claim under consideration. In practice, we observe both positive and negative incremental payments and we can expect that the behavior of positive and negative payments differs. By negative payments, we mean salvages and subrogations. Consequently, we can use a spliced distribution to model non-zero payments. We introduce the following sequences of random variables: Pi,j,(+) k=Pi,j kYi,j k∈ {1, 3},Pi,j k>0, k=1, 2, . . . , Pi,j,(−) k=−Pi,j kYi,j k∈ {1, 3},Pi,j k<0, k=1, 2, . . . , (5) where the first sequence models positive incremental payments, given their occurrence, and the second sequence models negative recovery payments, again given their occurrence. We have to model the conditional probabilities of a positive and a negative payment together with the conditional distributions of the payments (Pi,j,(+) k)k=1,2,... and (Pi,j,(−) k)k=1,2,... , respectively. As discussed in the previous subsection, the probabilities and the distributions of Pi,j,(+) k and Pi,j,(−) k , respectively, in development period k should depend on the available information included in Ci,j k−1 , zi,j k−1 . The probabilities and the distributions may also depend on the change in claim incurred indicator Ii,j k , since the events {Yi,j k= 1 } or {Yi,j k= 3 } include two different scenarios for the change in claim incurred. The conditional probabilities of a positive or a negative payment in development period k , respectively, given the occurrence of a non-zero payment, can be modeled with a binomial distribution. We use a binomial logistic regression model, and we set log PPi,j k>0Yi,j k∈ {1, 3},Ii,j k,Ci,j k−1,zi,j k−1 PPi,j k<0Yi,j k∈ {1, 3},Ii,j k,Ci,j k−1,zi,j k−1 =fIi,j k,Ci,j k−1,zi,j k−1,k≥1. (6) The regression model (6) is called Model 2. The cross-entropy loss function for a sample of random variables with binomial distribution is calculated analogously to (4).
Risks 2020,8, 33 6 of 34 As far as distributions of Pi,j,(+) k , Pi,j,(−) k are concerned, we can expect that the payments should have a skewed distribution. It is common in actuarial practice to model claim sizes with Gamma distributions and fit Gamma regression models to claims severities. For this reason, we start with assuming that Pi,j,(+) kand Pi,j,(−) khave Gamma distributions where we postulate the relationships: log EhPi,j,(+) kIi,j k,Ci,j k−1,zi,j k−1i=fIi,j k,Ci,j k−1,zi,j k−1,k≥1, (7) for an appropriate regression function f, and Var hPi,j,(+) kIi,j k,Ci,j k−1,zi,j k−1i=ψ·EhPi,j,(+) kIi,j k,Ci,j k−1,zi,j k−1i2,k≥1, (8) where ψ> 0 is called a dispersion coefficient. It was observed in many data sets, including insurance data sets, that the dispersion coefficient may not be constant. In such cases, it should be modeled with a separate regression function. A classical approach to model dispersion coefficients in GLMs/GAMs is to use a second Gamma regression model with a response directly modeling the dispersion coefficient. Consequently, we assume that Pi,j,(+) k and Pi,j,(−) k have Gamma distributions where we postulate the relationships: log EhPi,j,(+) kIi,j k,Ci,j k−1,zi,j k−1i =fIi,j k,Ci,j k−1,zi,j k−1,k≥1, (9) Var hPi,j,(+) kIi,j k,Ci,j k−1,zi,j k−1i=eφIi,j k,Ci,j k−1,zi,j k−1 ·EhPi,j,(+) kIi,j k,Ci,j k−1,zi,j k−1i2,k≥1, (10) for another regression function φ . Let Resi,j k denote the unscaled Pearson residual from the regression model (7) and (8) . We assume that the squared residual Resi,j k 2 follows a Gamma regression with moments: log EhResi,j k 2Ii,j k,Ci,j k−1,zi,j k−1i=φIi,j k,Ci,j k−1,zi,j k−1,k≥1, (11) Var hResi,j k 2Ii,j k,Ci,j k−1,zi,j k−1i=ϑ·EhResi,j k 2Ii,j k,Ci,j k−1,zi,j k−1i2,k≥1, (12) where ϑ> 0 denotes a dispersion coefficient for the Pearson residuals. Similar assumptions are made for negative payments. The regression model (9) – (12) is called Model 3_positive, or Model 3_negative if negative payments are modeled. We recall that the unscaled deviance loss function for a sample of n independent random variables (v`)`=1,...,n with Gamma distribution with moments (9) and (10) and estimated expected values (ˆ µ`)`=1,...,nis given by Dunscaled Gamma =−1 n n ∑ `=1log v` ˆ µ`−v`−ˆ µ` ˆ µ`. (13) We omit the constant 2 and use the average deviance loss function compared to the traditional definition of the Gamma deviance in GLMs/GAMs/trees. The scaled deviance loss function, which takes into account the dispersion coefficient, is given by Dscaled Gamma =−1 n n ∑ `=1 log v` ˆ µ`−v`−ˆ µ` ˆ µ` eˆ φ` , (14) where ˆ φ`denotes the prediction from (11) and (12) for observation `=1, . . . , n.
Risks 2020,8, 33 7 of 34 If we combine the binomial distribution for a positive or a negative payment with the Gamma distributions for the severities of positive or negative payments, we receive the distribution of the incremental payment Pi,j k|Yi,j k∈ {1, 3}in development period k. We remark that if Yi,j k= 1, our modeling process is complete for period k and we can derive the values of the process (Pi,j k+1 , Ii,j k+1) in the next development period k+ 1. If Yi,j k= 3, we have to model the change in claim incurred at the end of the development period k which is obviously related to the payment made in development period k . If Yi,j k= 2, the payment Pi,j k= 0 is zero but we need to consider a change in claim incurred. The next two subsections deal with changes in claim incurred. 2.3. Model 4: Closing Times We consider the two remaining events {Yi,j k= 2 } and {Yi,j k= 3 } related to changes in claim incurred. We should differentiate between two cases: (a) the value of the claim incurred changes and the case reserve is different from zero; (b) the value of the claim incurred changes and the resulting case reserve is equal to zero. The latter is interpreted that the claim is closed at the end of period k , and we do not expect any further claim developments on that claim, unless such a claim is re-opened. Re-openings are assumed to only happen with a small probability (which still needs to be modeled). For claims that have zero case reserves at the end of period k we exactly know the change in claim incurred (i.e., in this case the severity of change in claim incurred does not need to be modeled). Therefore, it is reasonable to include an indicator for zero case reserves directly into all regression functions discussed in this section. Please note that this indicator is already included indirectly since the regression functions are assumed to depend on the whole history of paid and incurred claims from which the value of the case reserve can be derived. However, a direct inclusion will allow the neural network to learn the relevant structure more quickly. We introduce the following sequences of random variables: Ri,j k=Ri,j kYi,j k∈ {2, 3},k=1, 2, . . . . (15) The event {Ri,j k= 0 } is interpreted as claim closing in development period k , given there is a change in claim incurred in the development period k . Of course, a claim can be re-opened in the future and this is modeled through probabilities (3) . The conditional probability for the event {Ri,j k= 0 } should depend on (Ci,j k−1 , zi,j k−1) , as well as on Pi,j k since the events {Yi,j k= 2 } and {Yi,j k= 3 } include two different scenarios for payment occurrences. Finally, if a payment is made, then the value of the payment Pi,j kcould also have an impact on the probability of the event {Ri,j k=0}. As Model 4, we use the binomial logistic regression model described by log PRi,j k=0Pi,j k,Pi,j k,Ci,j k−1,zi,j k−1 PRi,j k6=0Pi,j k,Pi,j k,Ci,j k−1,zi,j k−1 =fPi,j k,Pi,j k,Ci,j k−1,zi,j k−1,k≥1. (16) 2.4. Model 5: Severities for Claim Incurred for Open Claims We are left with changes in claim incurred in the non trivial scenario a) where the case reserve at the end of the development period is not equal to zero. We introduce the sequence of random variables: Ii,j,(open) k=Ii,j k|Yi,j k∈ {2, 3},Ri,j k6=0, k=1, 2, . . . . (17) As for payments, we assume that Ii,j,(open) khas a Gamma distribution with moments: log EhIi,j,(open) kPi,j k,Pi,j k,Ci,j k−1,zi,j k−1i=fPi,j k,Pi,j k,Ci,j k−1,zi,j k−1,k≥1, (18)
Risks 2020,8, 33 8 of 34 for a suitable regression function f, and choosing another suitable regression function φ Var hIi,j,(open) kPi,j k,Pi,j k,Ci,j k−1,zi,j k−1i=eφPi,j k,Pi,j k,Ci,j k−1,zi,j k−1 ·EhIi,j,(open) kPi,j k,Pi,j k,Ci,j k−1,zi,j k−1i2,k≥1. (19) Let Resi,j k denote the unscaled Pearson residual from regression model (18) and (19) . We assume that the squared residual Resi,j k 2follows a Gamma regression with moments: log EhResi,j k 2Pi,j k,Pi,j k,Ci,j k−1,zi,j k−1i=φPi,j k,Pi,j k,Ci,j k−1,zi,j k−1,k≥1, (20) Var hResi,j k 2Ii,j k,Ci,j k−1,zi,j k−1i=ϑ·EhResi,j k 2Pi,j k,Pi,j k,Ci,j k−1,zi,j k−1i2. (21) The set of predictors in (18) – (21) can be deduced in a similar way as in the previous sections. The regression model (18) – (21) is called Model 5. Negative claim incurred are assigned to data errors, henceforth, they are not modeled. Combining the results from Sections 2.2–2.4, we can model claim incurred Ii,j k|Yi,j k∈ { 2, 3 } in development period k and we can derive the value of the process (Pi,j k , Ii,j k) in the next development period k . The modeling process for the next development period is complete for all scenarios of potential claims developments. 2.5. Large Incremental Payments and Claim Incurred in Models 3 and 5 Typically, general insurance claims are skewed and heavy tailed. A traditional approach in actuarial pricing and reserving is to separate large claims from attritional claims. Large claims are usually identified by using techniques from Extreme Value Theory (EVT), and claims above a high threshold are modeled with Pareto distributions. In our case, we should model large incremental payments and large changes in claim incurred in each development period k= 1, 2, . . . . By a large claim we understand a large incremental payment in Models 3 or a large change in claim incurred in Model 5. Let dk denote a threshold and Pi,j,(+) k denote a positive incremental payment in development period k. We postulate a Pareto tail PPi,j,(+) k−dk>xPi,j,(+) k>dk,Ii,j k,Ci,j k−1,zi,j k−1 =λIi,j k,Ci,j k−1,zi,j k−1 λIi,j k,Ci,j k−1,zi,j k−1+xγIi,j k,Ci,j k−1,zi,j k−1,x>0, k≥1, (22) where γ is called tail index. A similar model can be used for large negative incremental payments and large claim incurred. In the most general claim development model, regression models should be built for the probability of a large claim and the severity of the large claim (for λ and γ ) since claims with particular features (e.g., bodily claims) are likely to have higher propensity to generate large claims in consecutive development periods. We refrain from doing this and choose a simpler approach. In each development period k , we set a fixed probability for the occurrence of a large claim, from which we deduce the threshold dk , and estimate constant parameters (λk , γk) of the large claim distribution (22) using EVT. Next, based on descriptive statistics and knowledge about business, we determine key features in zi,j k−1 which have high propensity to generate large claims and allocate large claims in simulations to claims with these features in the first place. Details are presented in Section 4.2, below. Consequently, we use a spliced distribution for the response in Models 3 and 5. We use a binomial distribution to decide whether the claim is attritional or large. If the claim is attritional then we use the double Gamma regression model and the Gamma distribution to predict the response, otherwise we
Risks 2020,8, 33 15 of 34 • We removed the claims for which we identified inconsistencies in accident dates, reporting dates and settlement months. In a company, one should apply data cleaning to such claims by doing further investigation with involved persons. • We investigated incremental payments and case reserves. To simplify the modeling process for the purpose of the paper, we only model positive payments and resign from modeling negative payments (salvages and subrogations). Consequently, we do not fit Model 2 and Model 3_negative in this paper. For simplicity, negative payments and recoveries were set to zero in the data set. • We ended up with a cleaned data set consisting of 1, 331, 856 individual claims. We removed less than 0.1% of the claims from the date set. The change in the aggregate payments historically observed was 1.2%, mainly caused by removal of recoveries. •The claims incurred were calculated for all claims. For computational reasons, we fit the neural networks on quarterly data. The unit of the accident period, the reporting period and the development period is 3 months. Consequently, Pi,j k denotes the sum of incremental monthly payments made in the k -th quarter starting from the reporting date, and Ii,j k denotes the value of the claim incurred at the end of the k -th quarter starting from the reporting date (in quarterly units) for a specific claim. We deal with k= 1, 2, . . . , 55, since we have claims from 14 accident years in the data set. Table 1presents the number of observations which we can be used for fitting our regression models in consecutive development periods k≥ 1. It is clear that the number of available observations decreases with the development period. This means that we are not able to estimate separate regression models for all later development periods k= 1, 2, . . . , 55. We decide to fit separate neural networks for k= 1, . . . , 16, and one neural network for all development periods k= 17, . . . , 55, for each Model 1, 3_positive, 4, 5. The neural network for the last development periods is denoted as a neural network for development period k= 17. The empirical probability that the incremental payment is zero and the claim incurred do not change (Case 0) in Model 1 in period k= 16 is 98.78%, and increases above 99% for later periods. We use 99% as a threshold probability in Model 1 for fitting separate neural networks in each development period. Table 1. Number of observations available for estimating the regression models for development periods k≥1. Development Model 1 Model 3_Positive Model 4 Model 5 k Case 0 Case 1 Case 2 Case 3 Payments Case 0 Case 1 Claims Incurred 1 321,618 87,926 123,745 762,394 850,320 82,430 803,709 82,430 2 1,094,980 10,386 46,495 114,158 124,544 39,053 121,600 39,053 4 1,148,925 1573 28,366 24,109 25,682 23,650 28,825 23,650 8 1,062,232 360 18,056 6589 6949 16,801 7844 16,801 12 948,516 157 16,619 3485 3642 15,563 4541 15,563 14 893,918 139 13,311 2102 2241 10,954 4459 10,954 16 842,072 86 9155 1162 1248 7511 2806 7511 20 732,314 40 3328 559 599 1990 1897 1990 30 433,175 8 137 89 97 113 113 113 40 170,629 0 20 9 9 20 9 20 45 88,120 0 10 2 2 8 4 8 50 35,571 0 4 1 1 1 4 1 In the next subsections we present the estimation results of our regression models for the selected development periods k= 2, 8, 16, 17. We work with neural networks with 2 hidden layers. The rule of thumb is that the number of neurons in the first layer should be between the number of regressors and three times the number of regressors. The largest number of regressors is in Models 4 and 5, where we have 20 regressors plus regressors related to the claim history in the number of 2 times the development period considered. We investigate two choices for the number of neurons in the first layer: q1= ( 20 + 2 ·development_period)· 2 and q1= ( 20 + 2 ·development_period)· 1.5. In the second layer we choose, respectively, q2=10 and q2=5.
Risks 2020,8, 33 16 of 34 The data set is split to a training set and a validation set. We use a random split where 90% of the observations are allocated to the training set. If we fit a categorical regression model (Model 1 or Model 4), then we use a stratified sampling and the proportions of the classes in the training set are the same as in the whole data set. In training the neural networks, we use a batch size of 10, 000 observations, 500 epochs for the categorical regressions, 1000 epochs for the Gamma regressions, and drop-out probabilities from 1% to 30%. Our experiments show that we need fewer epochs to train the categorical regressions than the Gamma regressions. We use a low learning rate at 0.001. We also test higher learning rates and we do not observe improvements in calibrations. 4.2. Predictions in Development Period k =2 In development period k= 2 we have many observations available for fitting each regression model, see Table 1. Moreover, the number of predictors, which we want to use in neural networks NN1 , is small due to the short history of the claim development. Hence, we expect that we can fit deep neural networks in development period k= 2 with a large number of trainable parameters and a low drop-out probability. We choose drop-out probability at 1%. First, we analyze large claims, i.e., large incremental payments and large changes in claim incurred. We use Hill plots and we try to find the most stable region in the Hill’s estimates of the tail index, see Figure A1. We choose the probability of a large claim at 1% for incremental payments and claim incurred. In Figure A2 we study the proportions of claim features in the whole data set available for estimation of the model under investigation and in the data set consisting of large claims. We can conclude that the proportion of claims from Segment 1, bodily injury claims and claims from abroad increase in the subset of large claims. This conclusion is better pronounced in later development periods in Figures A6,A10 and A14. Consequently, in simulations of ultimate payments, we assume that 75% of large incremental claims are allocated in the first place to bodily injury claims and claims from abroad, and 75% of large changes in claim incurred are allocated in the first place to Segment 1, bodily injury claims and claims from abroad. Remaining large claims are allocated to random claims. The proportion 75% is an arbitrary number. Next, we fit Models 1, 3_positive, 4, 5. We should decide which model should be chosen as an initial model M0 from which we start training our neural networks. It turns out that GAMs and regression trees with claim incurred in the last development period and cumulative payments do not reduce significantly the loss function on the training set compared to a homogeneous model for Models 1 and 4. Hence, we decide to initiate the training of the neural networks for Models 1 and 4 starting from empirical estimates. For Models 3_positive and 5 we use regression trees with claim incurred in the last period and cumulative payments as initial models M0 since we observe a significant decrease in the loss function on the training set compared to a homogeneous model. We initiate the training of the neural networks for Models 3_positive and 5 starting from the estimates produced by the regression tree. We could start with GAMs as M0 and reduce the loss function at the initial training step of neural networks even more. Yet, regression trees are faster for predictions than GAMs. The disadvantage is that we have to train Gamma neural networks for more epochs. The cross-entropy and deviance loss functions on the validation sets are presented in Table 2 and Figure A3. As a benchmark, we also provide the loss functions in the GAMs where we use claim incurred in the last development period and cumulative payments as the only regressors. The predictions from M0 and GAM can be improved with neural networks NN0 . We can conclude that not only the cumulative payments and the claim incurred in the last development period are useful for predictions in the next development period, but we should also use other claim features and interactions between the predictors in claims reserving. Moreover, we can observe that the use of the whole claim history also improves the predictions. Of course, for k= 2, the use of the whole claim history means to use the two historical values of the incremental payments and the claim incurred observed in k= 0 and k= 1. Hence, the regression models are not significantly enhanced by additional information about the claim development process. The models’ enhancement is much bigger for later
Risks 2020,8, 33 17 of 34 development periods where, as we show in the next subsections, we should also use the whole claim history, see Tables 3and 4and Figures A7 and A11. It is clear that the predictions are improved in all Models 1, 3_positive, 4, 5 when we add incremental payments and claim incurred (observed in all past development periods) as regressors in the regression functions, i.e., when we switch from neural network NN0 to neural network NN1 . Consequently, the Markovian assumption should be rejected for the claim development processes in our data set. Finally, deeper neural networks are preferred in k= 2 since we have a large number of observations at our disposal and there is no danger that we overfit the models (by appropriately early stopping and drop-out probabilities). Table 2. Minimal cross-entropy and deviance loss functions on validation sets observed during the training of the neural networks in k=2. DGAM DNN0DNN11−DNN0 DGAM 1−DNN1 DGAM Model 1: q= (48, 10)0.3865 0.3243 0.3206 16.11% 17.06% Model 1: q= (36, 5)0.3865 0.3249 0.3229 15.95% 16.46% Model 3_positive: q= (48, 10)0.5309 0.4842 0.4762 8.79% 10.30% Model 3_positive: q= (36, 5)0.5309 0.4880 0.4823 8.08% 9.14% Model 4: q= (48, 10)0.4900 0.1895 0.1888 61.32% 61.47% Model 4: q= (36, 5)0.4900 0.2125 0.2117 56.63% 56.79% Model 5: q= (48, 5)0.1714 0.1263 0.1250 26.34% 27.10% Model 5: q= (36, 5)0.1714 0.1311 0.1304 23.54% 23.94% Table 3. Minimal cross-entropy and deviance loss functions on validation sets observed during the training of the neural networks in k=8. DGAM DNN0DNN11−DNN0 DGAM 1−DNN1 DGAM Model 1: q= (72, 10)0.1086 0.0499 0.0487 54.03% 55.17% Model 1: q= (54, 5)0.1086 0.0503 0.0488 53.73% 55.03% Model 3_positive: q= (72, 10)0.6190 0.5763 0.5593 6.90% 9.64% Model 3_positive: q= (54, 5)0.6190 0.5707 0.5657 7.80% 8.61% Model 4: q= (72, 10)0.5533 0.1941 0.1742 64.92% 68.52% Model 4: q= (54, 5)0.5533 0.1939 0.1792 64.96% 67.61% Model 5: q= (72, 10)0.1152 0.0349 0.0390 69.71% 66.17% Model 5: q= (54, 5)0.1152 0.0431 0.0424 62.61% 63.19% We now validate the distributional assumption that the incremental payments and the claim incurred are from Gamma distributions with the moments, respectively, given by (9) – (10) , (18) – (19) . As in GLMs/GAMs, we use Pearson and deviance residuals. If the response v` follows Gamma distribution with true expected value µ` and dispersion coefficient ψ` , then the Pearson residual, after adding 1, i.e., Res`+ 1 =v`−µ` µ`+ 1 has a Gamma distribution with expected value equal to 1 and dispersion ψ` . We can use a version of a QQ normal plot to verify the Gamma distributional assumption. If Res` denotes an observation from the Gamma distribution with true expected value equal to 1 and dispersion ψ` , in our case Res` denotes the Pearson residual from the fitted Gamma regression model, then FΓ(1,ψ`)(Res`) has a uniform distribution, where FΓ(1,ψ) denotes the cumulative Gamma distribution function for parameters 1 and ψ . Consequently, Φ−1FΓ(1,ψ`)(Res`) has a normal distribution, where Φ−1 denotes the inverse of the standard normal distribution function. If the Gamma distribution is the correct distribution for the response, then we should observe a straight line in the QQ-plot for the Pearson residuals from the fitted model. Let us recall that the dispersion coefficients are estimated for individual claims with the second Gamma regression model. We also investigate the so-called Tukey-Anscombe plot where we plot the deviance residuals from the fitted Gamma regression model, scaled with the dispersion coefficients, against the log-predictions
Risks 2020,8, 33 18 of 34 of the response in the fitted model. If the mean and variance assumptions for the distribution of the response are correct, then the deviance residuals should be fluctuate randomly around the horizontal line through zero and the deviation in the residuals should be constant. The QQ normal plots and Tukey-Anscombe plots for the Gamma regressions fitted with NN1 are presented in Figure A4. We observe that the left tail is underestimated for incremental payments and both right and left tail are underestimated for claim incurred. If we look at later development periods, see Figures A8,A12,A16, we can conclude that the Gamma distribution can be accepted for incremental payments. The choice of the Gamma distribution for claim incurred can be questioned and an improvement would be desirable. In practice, it is difficult to get a perfect fit to a Gamma distribution (with constant or varying dispersion). We should at least aim at correctly modeling the mean response and the dispersion with two neural networks. If we investigate deviance residuals, then we can conclude that the first two moments of the response are correctly modeled for incremental payments and claim incurred in all development periods. This property should be sufficient for projecting the aggregate ultimate payments in a large portfolio, estimating the best estimate and not-extreme quantiles of the aggregate ultimate payments. We would like to point out that varying dispersion improves the deviance residuals and the second neural network for the dispersion coefficient in Model 3_positive and 5 is indeed required. 4.3. Predictions in Development Period k =8 We present the same tables and plots as in the previous section, see Table 3and Figures A5–A8. The conclusions are the same. We choose the probability of a large claim at 1%. We could differentiate the probabilities of large claims in development periods, but, if possible, we try to choose the same probability for large claims in all development periods. The drop-out probability is still set at 1%. We can deduce that the cumulative payments and the claim incurred in the last development period are not sufficient, except in Model 5, for precise predictions of the claim development processes. The Markovian assumption has to be rejected for the claim development processes. The fit of a Gamma distribution to incremental payments and claim incurred is acceptable. We prefer larger neural networks. Yet, the need to apply the early stopping rule, in order not to overfit Model 3_positive, can be observed in Figure A7. 4.4. Predictions in Development Period k =16 In development period k= 16 (4 years), the set of predictors which we want to use in neural network NN1 is rather large. We want to use 51 or 52 regressors in total (where 32 regressors are related to incremental payments and claim incurred observed in the past development periods). At the same time, the number of observations is relatively small for Models 3_positive. Since the number of neurons in the first hidden layer should be at least equal to the number of regressors, the number of trainable parameters in Models 3_positive quickly approaches, and exceeds, the number of observations when we increase the number of neurons in NN1 , which we would do to improve the fit of the neural network. Consequently, the loss function may, already after a few iterations, increase on a validation set and the fitting algorithm may produce a poor prediction model. To prevent over-fitting, we can apply regularization techniques (early stopping may not be sufficient). We decide to increase the drop-out probability to get stable calibration results for Model 3_positive. We choose 30% in k= 16. In fact, we decide to increase the drop-out probability from 1% to 30% for Model 3_positive starting from the development period k=11. We also modify the probability of a large claim in Model 3_positive. The reason is again a low number of observed payments. We choose 2% in development period k= 16, and 1.5% in k= 15. Hill plots, Pareto quantile plots and proportions of claim features in the two data sets of all claims and large claims are presented in Figures A9 and A10. The results of fitting neural networks for Models 1, 3_positive, 4, 5 are presented in Table 4and Figure A11. The conclusions are the same as in the previous sections. The larger neural network is
Risks 2020,8, 33 19 of 34 also preferred for Model 3_positive, since it is regularized with a large drop-out probability. However, when we re-run the fitting algorithm then it turns out that in many runs the smaller network is preferred, see Figure A17. Hence, we decide to switch to smaller neural networks for Model 3_positive starting from development period k= 11. For Models 1, 4 an 5 we use larger neural networks in all development periods. Table 4. Minimal cross-entropy and deviance loss functions on validation sets observed during the training of the neural networks in k=16. DGAM DNN0DNN11−DNN0 DGAM 1−DNN1 DGAM Model 1: q= (104, 10)0.0653 0.0132 0.0130 79.84% 80.10% Model 1: q= (78, 5)0.0653 0.0136 0.0132 79.15% 79.85% Model 3_positive: q= (104, 10)0.5425 0.4988 0.4781 8.05% 11.87% Model 3_positive: q= (78, 5)0.5425 0.4972 0.4943 8.35% 8.88% Model 4: q= (104, 10)0.5350 0.2540 0.2417 52.51% 54.82% Model 4: q= (78, 5)0.5350 0.2569 0.2568 51.98% 51.99% Model 5: q= (104, 10)0.0347 0.0297 0.0285 14.41% 17.96% Model 5: q= (78, 5)0.0347 0.0319 0.0305 8.23% 12.22% The residuals are presented in Figure A12. The fit of Gamma distributions is acceptable. Yet, the estimated Gamma distribution now overestimates both tails of the claim incurred. 4.5. Predictions beyond Development Period k =16 All claims with development periods bigger than 16 quarters are grouped into one data set and one neural network with development period information as a regressor is fitted. Due to this grouping, we cannot fit neural networks NN1 since claims in different development periods have different lengths of claim history. Hence, we fit neural network NN0 . This is justified since for latter development periods we expect that only the most recent claim history should have the greatest power in the prediction of claim development. Moreover, the results from the previous sections show that neural networks NN1are only slightly better than NN0in predictions on the validation sets. Hill plots, Pareto quantile plots and proportions of claim features in two data sets of all claims and large claims are presented in Figures A13 and A14. We choose the probability of a large incremental payment at 3%, and the probability of a large claim incurred at 0.5%. The probability of a large payment is higher than in the previous models. This agrees with intuition as heavier claims occur at later development periods. The results of training the neural networks are presented in Table 5and Figure A15. It is clear that NN0 has better predictive power than M0 and GAM, and we prefer larger neural networks. Residuals are presented in Figure A16. We conclude that the fit is acceptable. Finally, in Figure A18 we present the estimated tail indices of Pareto distributions for large claims for all development periods. Moreover, the mean values of the dispersion coefficients estimated for individual claims in Models 3_positive and 5 with the second Gamma regression model, which coincide with the constant dispersion coefficients in the first Gamma regression model, are presented in Figure A18.
Risks 2020,8, 33 20 of 34 Table 5. Minimal cross-entropy and deviance loss functions on validation sets observed during the training of the neural networks in k=17. DGAM DNN01−DNN0 DGAM Model 1: q= (48, 10)0.0176 0.0048 73.013% Model 1: q= (36, 5)0.0176 0.0048 73.012% Model 3_positive: q= (48, 10)0.6044 0.5555 8.08% Model 3_positive: q= (36, 5)0.6044 0.5656 6.42% Model 4: q= (48, 10)0.6431 0.3880 39.67% Model 4: q= (36, 5)0.6431 0.4506 29.94% Model 5: q= (48, 10)0.0712 0.0300 57.84% Model 5: q= (36, 5)0.0712 0.0344 51.77% 5. Projection to Ultimate Payments for RBNS Claims In the first step, we use Chain-Ladder techniques to estimate the ultimate payments for RBNS claims. In Table 6we present the Chain-Ladder estimates of the development factors calculated on quarterly data for aggregate payments and the numbers of reported claims. We use all observations in the period 2005–2018. We can conclude that we deal with an insurance portfolio with long-tailed development patterns, where late development factors for payments are still bigger than 1. Table 6. Chain-Ladder (CL) estimates of development factors for aggregate payments and the number of reported claims. Development 1 2 3 5 10 20 30 40 50 CL factor: payments 2.124071 1.167995 1.083916 1.035170 1.015100 1.005782 1.003623 1.002830 1.001223 CL factor: no. of claims 1.121514 1.024738 1.011397 1.003963 1.001042 1.000209 1.000317 1.000355 1.000311 We calculate the Chain-Ladder predictions for the aggregate ultimate payments and the ultimate number of reported claims per each accident quarter. Next, we calculate the average ultimate payment per accident quarter by diving these two predictions by each other. To differentiate the ultimate payment by the reporting delay of the claim, we calculate the mean values of the claims incurred available at the end of December 2018 per reporting delay and we regress the mean value of the claims incurred versus the reporting delay. We use a GAM regression model to obtain the scaling factors for ultimate payments for claims reported with reporting delays compared to the ultimate payment for a claim reported without a delay. Using the average ultimate payments per accident quarter and reporting delay and the number of claims to be reported in the future periods, we can calculate the IBNYR (Incurred But Not Yet Reported) and RBNS reserves. The aggregate ultimate payments for RBNS claims per accident quarter are calculated as the predicted aggregate ultimate payments for the portfolio minus the predicted number of claims to be reported in the future periods times the average ultimate payment taking into account the reporting delay of the claims (IBNYR). The RBNS reserve is calculated by subtracting the aggregate payments up to December 2018 from the projected aggregate ultimate payments for the RBNS claims. We remark that this is still a rough estimate of the RBNS reserve to receive numbers that are comparable to our individual claim reserving method. The accident years from 2009 till 2018 cover 98.98% of the total RBNS reserve of all accident years 2005–2018, so we only focus on these years. The results are presented in Table 7. Once Models 1, 3_positive, 4 and 5 are calibrated, we can use them to simulate incremental payments in consecutive development periods (quarters) and generate ultimate payments for RBNS claims. Since we deal with a very large portfolio of claims, we do not have to perform many simulations to receive stable results, unless we are interested in a very high quantile of aggregate ultimate payments. We use 50 simulations, which is sufficient for the results we present below. We consider the claims from the years 2009–2018 but we model their development over the full 55 development periods.
Risks 2020,8, 33 21 of 34 In Figure 1we present histograms of the aggregate ultimate payments for the whole insurance portfolio, as well as for segments and different claim features (property versus bodily injury and Poland versus abroad). Such results can help the insurer to improve claims reserving as the insurer can now set claims reserves for policies grouped by claim features. Even at the portfolio level, we expect that the mean prediction from our models should be more precise than the Chain-Ladder estimate derived from aggregate data since our models are calibrated based on detail information about the claim development process, they can cope with seasonality and with changes in portfolio mix. In our case, the proportion of body claims in the portfolio decreases from 10% in 2005 to 5.5% in 2018, and the proportion of the claims caused abroad decreases from 5.7% to 2.6%. Since body claims and claims from abroad are expected to be the most severe, the Chain-Ladder estimates might be overestimated for the recent accident periods. Apart from calculating the best estimate of liabilities, the histograms in Figure 1can also help the insurer to understand which claim features generate the most uncertainty of future payments. Such results can be used to calculate risk adjustments in IFRS 17 or define a safety buffer in claims reserves. We remark that there is a peak in the right tail of the histogram of the aggregate ultimate payments. We conclude that the portfolio is exposed to large claims, which mainly occur in late development periods from body claims and claims from abroad. Aggregate ultimate claims Frequency 900 902 904 906 908 0 4 8 Segment 1 Frequency 482 486 0 4 8 Segment 2 Frequency 340 344 348 0 10 Segment 3 Frequency 34.5 35.5 0 4 8 Segment 4 Frequency 41.0 42.0 0 6 12 Segment 5 Frequency 0.210 0.220 0 6 12 Property Frequency 668 670 672 0 5 15 Body Frequency 230 234 238 0 4 8 Poland Frequency 804 808 812 0 4 8 Abroad Frequency 96 98 100 102 0 5 15 Figure 1. Histograms of the aggregate ultimate payments from the RBNS claims (in MM). The dashed lines indicate the mean value from the simulations. In Figure 2we find boxplots of the aggregated ultimate payments per accident year. Deviations in the simulated payments are small which justifies a small number of simulations, see also 25th and 75th quantiles in Table 7. In Figure 2we compare the ultimate payments for the RBNS claims simulated with our model with the results estimated with the Chain-Ladder technique, and in Table 7 we compare the RBNS reserves. We can observe that our model produces lower numbers for the years
Risks 2020,8, 33 22 of 34 2009–2018 than the Chain-Ladder estimates. This observation may be explained with the change of the proportion of body claims and claims from abroad in the insurance portfolio, which we discuss above. Moreover, the Chain-Ladder development factors for old accident years and early development periods tend to be systematically higher than the development factors for recent accident years, see Figure A19 where dark blue points cumulate in the top left corner of the triangle, especially for the first development period for the old accident years. Consequently, the Chain-Ladder development factors tend to overestimate the future payments. 2009 2011 2013 2015 2017 70 80 90 100 110 Aggregate ultimate claims Accident year Figure 2. Boxplots of the aggregate ultimate payments from the RBNS claims (in MM). The dots indicate the Chain-Ladder estimates. Table 7. Simulations results and Chain-Ladder (CL) estimates (in MM). RBNS Reserve Accident Year Mean 25th Quantile 75th Quantile Chain-Ladder Case Reserve 2009 1.26 1.13 1.39 1.35 0.65 2010 1.51 1.32 1.68 1.97 1.56 2011 2.29 2.03 2.39 2.66 3.14 2012 3.65 3.26 3.78 3.85 3.63 2013 5.10 4.69 5.18 5.11 5.07 2014 5.99 5.43 6.35 6.61 7.00 2015 8.27 7.57 8.36 9.36 6.90 2016 11.91 11.48 12.27 13.43 10.81 2017 17.20 16.67 17.80 19.62 11.65 2018 35.09 34.64 35.59 39.77 28.02 All 92.27 88.21 94.78 103.74 78.43 We can also compare the company’s case reserves with the RBNS reserves estimated with our model, see Table 7. The comparison is not obvious since case reserves depend on company’s claims reserving policy. For some old accident years the case reserves are above the RBNS reserves. This might be caused by the fact that for old claims, which are not closed, the case reserve is not revaluated and is kept unchanged until the full settlement is made. For recent accident years the case reserves are below the RBNS reserves, as expected. We estimate the expected ultimate payment for an individual claim, see Figure 3. It agrees with our intuition that the expected ultimate payment for a property claim is lower than for bodily injury one. It also agrees with our intuition that the expected ultimate payment for a claim caused in Poland is lower than for a claim caused abroad. We can also confirm that the expected ultimate payment for an individual claim increases with the reporting delay of the claim.
Risks 2020,8, 33 23 of 34 2010 2012 2014 2016 2018 650 800 950 Expected ultimate claim Accident year 0 20 40 60 80 100 6789 Log expected ultimate claim Reporting delay in months Property Body 0 1000 2000 12345 Expected ultimate claim Segment 0 400 800 Poland Abroad 0 1000 2500 Figure 3. Expected aggregate ultimate payments from an individual claim. Finally, we check stability of our calibration procedure. We re-calibrate all neural networks and we re-simulate the developments of all reported claims. When refitting the neural networks, we also change the split of the data of training and validation sets and, this time, we use 80% of observations as training set and the remaining 20% as test set. The results are presented in Table 8. Comparing Tables 7and 8 , we can conclude that the calibration procedure is very stable on the aggregate level (see RBNS reserves of 89.80 vs. 92.27). Bigger changes are only observed on older accident years where RBNS are very low and, thus, only marginally contribute to the overall RBNS reserves. This uncertainty on older claims is, of course, caused by the fact that payments for these claims are very rare because most of the claims were already settled. Table 8. Simulations results and Chain-Ladder (CL) estimates (in MM) - a new re-calibration of NNs and a new simulation run. RBNS Reserve Accident Year Mean 25th Quantile 75th Quantile Chain-Ladder Case Reserve 2009 0.94 0.79 1.08 1.35 0.65 2010 1.31 1.18 1.36 1.97 1.56 2011 1.94 1.74 2.09 2.66 3.14 2012 3.11 2.94 3.32 3.85 3.63 2013 4.51 4.29 4.71 5.11 5.07 2014 5.49 5.17 5.78 6.61 7.00 2015 7.78 7.43 8.02 9.36 6.90 2016 11.40 10.84 11.79 13.43 10.81 2017 17.16 16.83 17.44 19.62 11.65 2018 36.17 35.64 36.68 39.77 28.02 All 89.80 86.87 92.28 103.74 78.43
Risks 2020,8, 33 24 of 34 6. Conclusions We developed regression models and postulated distributions which can be used in practice to describe the joint development process of individual claim payments and claim incurred. We applied neural networks to estimate our regression models. The models can improve claims reserving techniques for RBNS claims, we showed benefits of using deep neural network and the whole claim history, and provide additional information about the risk factors which trigger the future payments. Our regression models estimated with individual claims can be used at any level of granularity, from the portfolio level to the policy level. Consequently, the RBNS reserve, which is traditionally calculated at the portfolio level, can be now directly allocated to sub-portfolios/segments/units of accounts in our modelling framework. Finally, since the projected ultimate payments can be split among claims features, the result can also be used to improve pricing and estimate the ultimate loss. Author Contributions: Conceptualization and methodology, Ł.D. and M.V.W.; writing—original draft preparation, Ł.D.; writing—review and editing, M.V.W. All authors read and agreed to the published version of the manuscript. Funding: The research of Ł.Delong was funded with grant NCN 2018/31/B/HS4/02150. Acknowledgments: We would like to thank Marcin Szatkowski for helping in analyzing the insurance portfolio. Conflicts of Interest: The authors declare no conflict of interest. Appendix A. Figures 0.0 1.0 2.0 −7 −2 Model 3_positive Log observation Log quantile 0.0 0.2 0.4 0.6 0.8 1.0 1.5 3.5 Model 3_positive Theta Hill estimate 012345 −6 −2 Model 5 Log observation Log quantile 0.0 0.2 0.4 0.6 0.8 1.0 0.8 1.8 Model 5 Theta Hill estimate Figure A1. Pareto quantile plots and altHill plots for incremental payments and changes in claim incurred in k=2.
Risks 2020,8, 33 31 of 34 1234 Model 3_positive Segment 0.0 0.3 0.6 Property Body Model 3_positive Claim type 0.0 0.4 0.8 Poland Abroad Model 3_positive Claim origin 0.0 0.4 0.8 12345 Model 5 Segment 0.0 0.3 0.6 Property Body Model 5 Claim type 0.0 0.4 0.8 Poland Abroad Model 5 Claim origin 0.0 0.4 0.8 All claims Large claims Figure A14. Proportions of claim features in the whole data set and in the data set of large claims in k=17. 0 100 200 300 400 500 0.004 0.014 Model 1 Epoch Loss 0 200 400 600 800 1000 0.50 0.65 Model 3_positive Epoch Loss 0 100 200 300 400 500 0.4 0.6 Model 4 Epoch Loss 0 200 400 600 800 1000 0.04 0.10 Model 5 Epoch Loss GAM NN_0 q(36,5) NN_0 q(48,10) Figure A15. Cross-entropy and deviance loss functions on validation sets observed during the training of the neural networks in k=17.
Risks 2020,8, 33 32 of 34 −4 −2 0 2 4 −4 2 Model 3_positive Theoretical Pearson residual Sample Pearson residual −3 −2 −1 0 1 −3 1 Model 3_positive Log prediction Sample deviance residual −4 −2 0 2 4 −4 2 Model 5 Theoretical Pearson residual Sample Pearson residual −4 −2 0 2 4 −3 1 Model 5 Log prediction Sample deviance residual Figure A16. QQ normal plots and Tukey-Anscombe plots in Models 3_positive and 5 fitted with neural networks NN1in k=17 0 200 400 600 800 1000 0.40 0.55 Model 3_positive Epoch Loss 0 200 400 600 800 1000 0.45 0.55 Model 3_positive Epoch Loss 0 200 400 600 800 1000 0.45 0.55 Model 3_positive Epoch Loss 0 200 400 600 800 1000 0.45 0.55 Model 3_positive Epoch Loss GAM NN_0 q(78,5) NN_0 q(104,10) NN_1 q(78,5) NN_1 q(104,10) Figure A17. Deviance loss functions on validation sets observed during the training of the neural networks in k=16.
Risks 2020,8, 33 33 of 34 5 10 15 0.0 1.0 2.0 Dispersion coefficient Development period 5 10 15 1.5 2.5 Tail index Development period Model 3_positive Model 5 Figure A18. Estimates of tail indices and constant dispersion coefficients in Models 3_positive and 5. Figure A19. Residuals from Mack model fitted to development factors in a run-off triangle. References Antonio, Katrien, and Richard Plat. 2004. Micro-level stochastic loss reserving for general insurance. Scandinavian Actuarial Journal 7: 649–69. Arjas, Elja. 1989. The claims reserving problem in non-life insurance: Some structural ideas. ASTIN Bulletin 19: 139–52. [CrossRef] Baudry, Maximilien, and Christian Y. Robert. 2019. A machine learning approach for individual claims reserving in insurance. Applied Stochastic Models in Business and Industry 35: 1127–55. [CrossRef] Cybenko, George. 1989. Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals, and Systems 2: 303–14. [CrossRef] De Felice, Massimo, and Franco Moriconi. 2019. Claim watching and individual claims reserving using classification and regression trees. Risks 7: 102. [CrossRef] Denuit, Michel, Donatien Hainaut, and Julien Trufin. 2019. Effective Statistical Learning Methods for Actuaries III: Neural Networks and Extensions. New York: Springer. Duval, Francis, and Mathieu Pigeon. 2019. Individual loss reserving using a gradient boosting-based approach. Risks 7: 79. Ferrario, Andrea, Alexander Noll, and Mario V. Wüthrich. 2018. Insights from inside neural networks. SSRN Electronic Journal 3226852. [CrossRef] Gabrielli, Andrea. 2020. A neural network boosted double overdispersed Poisson claims reserving model. ASTIN Bulletin 50: 25–60. [CrossRef] Gabrielli, Andrea, and Mario V. Wüthrich. 2018. An individual claims history simulation machine. Risks 6: 29. [CrossRef]
Risks 2020,8, 33 34 of 34 Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. Cambridge: MIT Press. Grohs, Philipp, Dmytro Perekrestenko, Dennis Elbrächter, and Helmut Bölcskei. 2019. Deep neural network approximation theory. IEEE Transactions on Information Theory, invited paper. Hornik, Kurt, Maxwell Stinchcombe, and Halbert White. 1989. Multilayer feedforward networks are universal approximators. Neural Networks 2: 359–66. [CrossRef] Jessen, Anders H., Thomas Mikosch, and Gennady Samorodnitsky. 2011. Prediction of outstanding payments in a Poisson cluster model. Scandinavian Actuarial Journal 3: 214–37. [CrossRef] Kuo, Kevin. 2019. DeepTriangle: A deep learning approach to loss reserving. Risks 7: 97. [CrossRef] Larsen, Christian R. 2007. An individual claims reserving model. ASTIN Bulletin 37: 113–32. [CrossRef] Lopez, Olivier, Xavier Milhaud, and Pierre-Emmanuel Thérond. 2019. A tree-based algorithm adapted to microlevel reserving and long development claims. ASTIN Bulletin 49: 741–62. [CrossRef] Norberg, Ragnar. 1993. Prediction of outstanding liabilities in non-lifeinsurance. ASTIN Bulletin 23: 95–115. [CrossRef] Norberg, Ragnar. 1999. Prediction of outstanding liabilities II. Model variations and extensions. ASTIN Bulletin 29: 5–25. [CrossRef] Pigeon, Mathieu, Katrien Antonio, and Michel Denuit. 2013. Individual loss reserving with the multivariate skew normal framework. ASTIN Bulletin 43: 399–428. [CrossRef] Quarg, Gerhard, and Thomas Mack. 2004. Munich chain ladder. Blätter DGVFM XXVI: 597–630. Schelldorfer, Jürg, and Mario V. Wüthrich. 2019. Nesting classical actuarial models into neural networks. SSRN Electronic Journal 3320525. [CrossRef] Taylor, Greg, Gráinne McGuire, and James Sullivan. 2008. Individual claim loss reserving conditioned by case estimates. Annals of Actuarial Science 3: 215–56. [CrossRef] Wüthrich, Mario V. 2018. Machine learning in individual claims reserving. Scandinavian Actuarial Journal 6: 465–80. [CrossRef] Wüthrich, Mario V. 2019. From generalized linear models to neural networks, and back. SSRN Electronic Journal 3491790. [CrossRef] Wüthrich, Mario V. 2020. Bias regularization in neural networks for generalized insurance pricing. European Actuarial Journal, to be published. Zhao, Xiao B., Xian Zhou, and Jing L. Wang. 2009. Semiparametric model for prediction of individual claim loss reserving. Insurance: Mathematics and Economics 45: 1–8. [CrossRef] c 2020 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).