scieee AI-readable full text Open interactive document viewer

Modelling unobserved heterogeneity in claim counts using finite mixture models

Bermúdez, Lluís,Karlis, Dimitris,Morillo, Isabel

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Bermúdez, Lluís; Karlis, Dimitris; Morillo, Isabel Article Modelling unobserved heterogeneity in claim counts using finite mixture models Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Bermúdez, Lluís; Karlis, Dimitris; Morillo, Isabel (2020) : Modelling unobserved heterogeneity in claim counts using finite mixture models, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 8, Iss. 1, pp. 1-13, https://doi.org/10.3390/risks8010010 This Version is available at: https://hdl.handle.net/10419/257965 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ risks Article Modelling Unobserved Heterogeneity in Claim Counts Using Finite Mixture Models Lluís Bermúdez 1,* , Dimitris Karlis 2and Isabel Morillo 1 1Departament de Matemàtica Econòmica, Financera i Actuarial, Universitat de Barcelona, Diagonal 690, 08034 Barcelona, Spain; [email protected] 2Department of Statistics, Athens University of Economics and Business, 10434 Athens, Greece; [email protected] *Correspondence: [email protected]; Tel.: +34-93-403-4853; Fax: +34-93-403-4892 Received: 20 December 2019; Accepted: 22 January 2020; Published: 29 January 2020 Abstract: When modelling insurance claim count data, the actuary often observes overdispersion and an excess of zeros that may be caused by unobserved heterogeneity. A common approach to accounting for overdispersion is to consider models with some overdispersed distribution as opposed to Poisson models. Zero-inflated, hurdle and compound frequency models are typically applied to insurance data to account for such a feature of the data. However, a natural way to deal with unobserved heterogeneity is to consider mixtures of a simpler models. In this paper, we consider k -finite mixtures of some typical regression models. This approach has interesting features: first, it allows for overdispersion and the zero-inflated model represents a special case, and second, it allows for an elegant interpretation based on the typical clustering application of finite mixture models. k -finite mixture models are applied to a car insurance claim dataset in order to analyse whether the problem of unobserved heterogeneity requires a richer structure for risk classification. Our results show that the data consist of two subpopulations for which the regression structure is different. Keywords: zero-inflation; overdispersion; automobile insurance; risk classification; risk selection JEL Classification: C51 1. Introduction and Aims In insurance datasets, for the purposes of modelling claim counts, there is a problem of unobserved heterogeneity caused by differences in driving habits and behaviour among policyholders that cannot be observed or measured by the actuary (for example, driving ability, driving aggressiveness or the degree of obeying traffic regulations). This often leads to overdispersion and a relatively large number of zeros, which cannot be fully remedied by Poisson regression models. Many attempts have been made in the actuarial literature to account for such features of the data (for example, compound frequency models, also known as mixture models, and their zero-inflated or hurdle versions). This paper aims to explore whether the problem of unobserved heterogeneity requires a richer structure, though the use of finite mixtures of regression models, than the previous models have. In a competitive market, insurance companies need to use a pricing structure that ensures that the exact weight of each risk is fairly distributed within the portfolio. If an insurance company does not achieve at least the same success with respect to this goal as their competitors, the policyholders with lower risk will be tempted to move to another company that offers better rates for them. Such an adverse selection process would lead the less unsuccessful company to lose its financial equilibrium, with insufficient income from premiums to pay for the claims reported by the remaining policyholders with higher risk. Risks 2020,8, 10; doi:10.3390/risks8010010 www.mdpi.com/journal/risks Risks 2020,8, 10 2 of 13 In most developed countries, the car insurance market is a highly competitive market. Therefore, to avoid such an adverse selection process, a particularly complex pricing structure is designed by actuaries. A thorough review of the modelling of claim counts for car insurance can be found in (Denuit et al. 2007). In general terms, to handle this problem, the actuary segments the portfolio into homogeneous classes so that all the insured parties belonging to a particular class pay the same premium. This procedure is referred to as risk classification, tariff segmentation or a priori ratemaking. In short, the classification or segmentation of risks involves establishing different classes of risk according to the nature of claims and probability of their occurrence. To this end, factors are determined to classify each risk, and its influence on the observed number of claims is estimated. To achieve this, risk analysis based on generalized linear models (GLMs) is widely accepted. Focusing on claim frequency, a regression component is included in the claim count distribution to take individual characteristics into account. A very common GLM used for these purposes is the Poisson regression model and its generalisations. Introduced by Dionne and Vanasse (1989) in the context of car insurance, the model can be applied if a series of classification variables, referred to as a priori variables, plus the number of claims for each individual policy are known. However, the Poisson regression model is usually rejected because of the presence of overdispersion and an excess of zeros. This rejection may be interpreted as a sign that the portfolio is still heterogeneous: not all factors influencing risk can be identified, measured and introduced into the a priori modelling. This phenomena is known as the problem of unobserved heterogeneity. In parallel, another way to account for unobserved heterogeneity is to consider that the claims record for each insured party reveals the differences in driving habits and behaviour among policyholders that cannot be observed or measured via the a priori variables. Therefore, the idea of considering individual differences in policies within the same a priori class by using an a posteriori mechanism has emerged, i.e., tailoring an individual premium based on the claims record for each insured party. This concept has received the name of a posteriori ratemaking, experience rating or the bonus-malus system (see Denuit et al. (2007)). One way to deal with overdispersion is to consider compound frequency models (mixture models) with some overdispersed distribution. This is best achieved by moving from the simple Poisson model to the negative binomial model (Dionne and Vanasse (1992)) or to the Poisson-inverse Gaussian model (Dean et al. (1989)). To account for the excess of zeros, some generalizations of the Poisson model have been considered. Lambert (1992) introduced the zero-inflated Poisson regression model and, since then, there has been a considerable increase in the number of applications of zero-inflated regression models based on several different distributions. A comprehensive discussion of these applications can be found in Winkelmann (2008). Similarly, hurdle models are also widely applied to insurance claim count data. A common assumption in all these models is that all policyholders behave in the same way with regard to a priori variables, and thus they all have the same regression structure. In this paper, we examine whether this assumption is realistic. The models proposed in this paper account for unobserved heterogeneity by choosing a finite number of subpopulations. To account for overdispersion and an excess of zeros, we consider a k -finite mixture of Poisson and negative binomial regression models. As Park and Lord (2009) show for vehicle crash data analysis, a finite mixture of Poisson or negative binomial regression models is especially useful where count data are drawn from heterogeneous populations. For modelling claim counts, the idea behind this is that the data consist of subpopulations of policyholders, “caused" by the unobserved heterogeneity, for which the regression structure, used to account for the observed or a priori variables, is different. These models allow each component in the discrete mixture to have its own score, i.e., for there to be different behaviour for each group of policyholders, whereas classical claim frequency models use a single score. To sum up, this paper aims to explore whether resolving the problem of unobserved heterogeneity requires a richer structure than that which is present in typical compound frequency models and their zero-inflated or hurdle versions. By applying finite mixtures of regression models, we will examine Risks 2020,8, 10 3 of 13 whether unobserved risk factors that are not considered in the a priori tariff, such as a driver’s reflexes, aggressiveness, or knowledge of the Highway Code, establish the existence of subpopulations of policyholders with different a priori behaviour. To achieve this goal, the proposed models are fitted to a set of car insurance claims data to compare their goodness of fit with the traditional claim frequency models and to assess if we need to account for this extra heterogeneity. Finally, we discuss whether the proposed models help to search for better alternatives to account for unobserved heterogeneity. In the next section, the models and computational details used are defined. In Section 3, we summarize the database obtained from a Spanish insurance company and the results from fitting the models to it. Finally, we offer some concluding remarks in Section 4. 2. Finite Mixture of Regression Models The central idea for a finite mixture of regression models is that we assume that the entire population can be split into k subpopulations (also called clusters, components or segments). Assuming a discrete-valued response yifor the i-th individual, we then assume that P(yi) = P(Yi=yi) = k ∑ j=1 πjP(yi|θij),θij >0, yi=0, 1, . . . , where 0 <πj< 1 with k ∑ j=1 πj= 1 are the mixing proportions indicating the probability that a randomly selected observation belongs to the j -th subpopulation and P(y|θ) is some discrete distribution indexed by some parameter vector θ . In our case presented below, P(y|·) will be assumed to belong to one of the Poisson or negative binomial families. Note that we assume that for each individual we have a set of parameters θij that depend on each component and they may depend on some covariate information for the i-th individual. We further assume that the mean of the j -th component can be modelled by a vector of covariates containing information on the i -th individual, denoted by xi . In the general setting, this covariate vector that characterises the i -th individual can be different for different components, and therefore we should use also a subscript j . As, in our model, we use the same covariates for all components, we drop this second index. Assuming, without loss of generality, that θ= (µ , φ) , where µ is the mean of the distribution (this can be easily obtained with a reparameterisation) and φ some parameter related to overdispersion (set equal to 1 for the Poisson distribution), we further assume that: log µij =x0 iβj where now βjis a component-specific vector of coefficients. Note that the above formulation can be seen in the context of a GLM. However, we prefer to describe the model in a more general setting since some of the families we may use instead of Poisson and negative binomial models do not belong to the exponential family or therefore to the general GLM setting. The above generic formulation can be expanded by allowing additional covariates to the rest of the parameters for each component, as well as to the vector of mixing proportions. A well-known model of this type is the finite mixture of Poisson regressions in Wang et al. (1996) (see also Grun and Leisch 2007,2008). Finite mixtures of regression models have been widely used in different settings, see Hennig (2000) for a thorough discussion. This type of modelling has some interesting features: first, the zero-inflated model is a special case; second, it allows for overdispersion; and third, it allows for a neat interpretation based on the typical clustering application of finite mixture models. Risks 2020,8, 10 4 of 13 It is useful to show that if we denote by µj and σ2 j the mean and the variance of the j -th component, then the mean µand the variance σ2of the mixture are given by µ= k ∑ j=1 πjµj, and σ2= k ∑ j=1 πj(µ2 j+σ2 j)−µ2(1) These formulas will be useful later on for our calculations. 2.1. Finite Mixture of Poisson Regressions The case of a finite mixture of Poisson regressions is by far the best known and most commonly applied in practice. It dates back to Wang et al. (1996) and assumes that P(yi|µij) = k ∑ j=1 πj exp(−µij)µyi ij yi!,yi=0, 1, . . . with µij =exp(xi0βj) . The zero-inflated Poisson regression is a special case. The model allows for overdispersion with respect to the simple Poisson regression model. For more details see Grun and Leisch (2007); Wang et al. (1996). 2.2. Finite Mixture of Negative Binomial Regressions For the negative binomial model, we assume P(yi|µij,φj) = Γ(φj+yi) Γ(φj)yi! µij φj+µij !yi φj φj+µij !φj ,φj>0, yi=0, 1, . . . and µij =exp(xi0βj) , i.e., the probability function of a negative binomial with mean µij and variance µij +µ2 ij φj. Note that we assume a separate overdispersion parameter φj for each component. Such a model has been fitted by Byung-Jung et al. (2014) and Zou et al. (2013). With respect to the finite mixture of Poisson regressions, the model has an extra overdispersion parameter and therefore allows for more flexible distributions in terms of components. It is evident that the negative binomial model also contains the simple Poisson model as a special case (φj→∞). 2.3. Other Models Although in this paper we focus on the two families of models introduced above, there are other models that fit into this context for which we do not present results. They relate to Poisson-inverse Gaussian regression models (Dean et al. (1989)) and finite mixtures of them; some nonparametric random effects Poisson regression models (see Aitkin (1999)), i.e., the model assumes some random effect on the intercept of the Poisson regression and thus actually fits a finite mixture of Poisson regression model where the estimated coefficients (apart from the intercept) are the same for all components; and hurdle-type models (a hurdle model is a modified count model in which the two processes generating the zeros and the positives are not constrained to be the same, see Mullahy (1986)) . As mentioned above, we do not formulate zero-inflated models as we treat them as special cases of finite mixture models. 2.4. Estimation via EM Algorithm Under the umbrella of a finite mixture, estimation for this particular family of models is rather simple. We follow the standard approach of combining the observed data (Yi , Xi) with unobserved Risks 2020,8, 10 5 of 13 latent vectors Zi= (Zi1 , . . . , Zik) with Zij = 1 if the i -th observation belongs to the j -th component, and 0 otherwise. As is typical, the EM algorithm consists of estimating the Z ’s by their conditional expectation and then fitting a standard regression model to the response, Y , using a weighted likelihood, based on the weights derived during the E-step. A formal description of a generic algorithm is given in what follows. E-step: Using the current estimates, ˆ πjand ˆ θij,i=1, . . . , nand j=1, . . . , k, calculate wij =E(Zij) = ˆ πjP(yi|ˆ θij) k ∑ j=1 ˆ πjP(yi|ˆ θij) ,i=1, . . . , n,j=1, . . . , k(2) and then M-step: M1 Update the mixing proportions using ˆ πj= n ∑ i=1 wij n,j=1, . . . , k M2 Update the regression coefficients and the component-specific parameters by fitting a single regression model for the j -th component with response yi , covariates xi using a weighted likelihood approach with weights wij. It is clear that the M-step is not in a closed form. Also, note that actually we fit k models with the same data but different weights. This can be run in parallel to speed up the process. All the pros and cons of the EM algorithm for finite mixtures apply. Also, standard procedures for finite mixtures are applicable, such as for example model selection. We will discuss some computational details later. Finally, we need to emphasize the issue of identifiability. Conditions for identifiability for such finite mixtures of regression models are given in Hennig (2000). For such a finite mixture of regression models for count data, problems may occur if the covariates are categorical and they can have a small number of different combinations. In our case, we have seven binary variables leading to 2 7 combinations. Not all of them appear in the data but we still have quite a large number of distinct combinations for the model matrix. In general, it is hard to show that identifiability exists since the conditions are hard to evaluate. We believe that in our case no particular problem exists. From a practical point of view we have worked with several initial values to examine whether we became trapped with different solutions. This did not happen, adding to our belief that our model is identifiable. 2.5. Computational Details An important aspect for the successful application of the EM algorithm is that appropriate initial values need to be selected, as otherwise one may be trapped in local rather than global maxima. We selected our initial values as follows. We started by fitting a simple Poisson regression model. This also gave sufficient initial values for the simple negative binomial regression. Initial values for the overdispersion parameter in these two models were set equal to the observed overdispersion (as proposed in Breslow (1984)). From here on we describe the approach for each model. Therefore, when we refer to “model”, we imply either the mixture of Poisson models or the mixture of negative binomial models. Initial values for k= 2 were selected by perturbing the simple ( k= 1) regression model. Specifically, we fitted a single regression and keeping the fitted values, we split them into two components with mixing probabilities of 0.5 each, and means equal to 1.2 and 0.8 of the fitted values. Then, to fit a model with k+ 1 components, we used the solution with k components and a new component at the centre (that of a single one-component regression), with mixing probability 0.05. The other mixing Risks 2020,8, 10 6 of 13 probabilities were rescaled to sum to 1. Extensive simulation has shown that this approach works well to locate the maximum. Other approaches can be found in Papastamoulis et al. (2016). All our computations were made in R . We used our own code, while some of the models can be fitted using the gamlss , VGAM and flexmix packages in R . However, we found some convergence problems and less flexibility while using the standard packages. Convergence was detected when the relative change between two successive iterations was smaller than 10 −8 . For fitting the separate regression models we used the standard GLM approach (IRLS algorithm) for Poisson and negative binomial regression. 3. Data and Results 3.1. Data Description The original database is a random sample of the car portfolio of a major insurance company operating in Spain in 1996. Only cars categorized as being for private use were considered. The data contain information from 80,994 policyholders. Seven exogenous variables plus the annual number of accidents recorded were considered here. For each policy, the information at the beginning of the period and the total number of claims from policyholders were reported. The definition and some descriptive statistics of the variables are presented in Table 1. This dataset has previously been used in Pinquet et al. (2001), Brouhns et al. (2003), Bolancé et al. (2003,2008), Boucher et al. (2007,2009), Boucher and Denuit (2008), Bermúdez (2009) and Bermúdez and Karlis (2012). The meaning of the variables that refer to the Spanish market should also be clarified. The variable ZON distinguishes between driving zones of greatest risk (Madrid, Catalonia and central northern Spain) and the rest. Regarding the type of coverage provided by the of policies (variable COV), the classification adopted here responds to the most common types of car insurance policy available on the Spanish market. The simplest policy only includes third-party liability. This simplest type of policy makes up the baseline group, while variable COV equals 1 denotes policies which, apart from the guarantees contained in the simplest policies, also include comprehensive and collision coverage. Table 1. Dependent and explanatory variables used in the models. Variable Definition Mean St. dev. N total number of claims reported by policyholders 0.1833 0.5873 (0: 71,087; 1: 6,744; 2: 2,067; 3: 690; 4: 248; 5: 95; 6: 34; >6: 29) GEN equals 1 for women and 0 for men 0.1600 0.3666 URB equals 1 when driving in urban area, 0 otherwise 0.6690 0.4706 ZON equals 1 when driving in Madrid, Catalonia or northern Spain, 0 otherwise 0.4326 0.4954 LIC equals 1 if the driving license is 4 or more years old, 0 otherwise 0.9766 0.1511 LOY equals 1 if the client is in the company for more than 5 years, 0 otherwise 0.1441 0.3512 COV equals 1 if includes comprehensive and collision coverage, 0 otherwise 0.5087 0.4999 POW equals 1 if horsepower is greater than or equal to 5500cc, 0 otherwise 0.8058 0.3955 3.2. Fitted Models We fitted models of increasing complexity to this dataset, starting from a simple Poisson regression model. We used AIC and BIC to select the best among a series of candidate models. All models were run in R . Table 2compares the fitted models for Poisson and negative binomial distributions, resulting in the best fit being obtained with a 2-finite mixture of negative binomial regression models (2FMNB). Finite mixture models with k> 2 were also fitted, but no improvement in terms of AIC or BIC was achieved. This result gives rise to the conclusion that this portfolio is comprised of two groups of policyholders. Risks 2020,8, 10 7 of 13 Table 2. Information criteria for selecting the best model for the data. Model Log-Likelihood Parameters AIC BIC Poisson −42,585.08 8 85,186.15 85,260.57 Negative binomial −38,453.13 9 76,924.27 77,007.98 Zero-inflated Poisson −38,836.59 9 77,691.19 77,774.91 Zero-inflated negative binomial −38,453.13 10 76,926.27 77,019.28 2-Finite Poisson mixture −38,449.61 17 76,933.21 77,091.36 2-Finite negative binomial mixture −38,347.81 19 76,733.62 76,910.36 As expected, a large improvement is obtained by moving from a simple Poisson model to a compound frequency model with some overdispersed distributions. The best fit is achieved by the negative binomial model. Zero-inflated models, while providing an improvement on the basic Poisson model, allowing for overdispersion in this case, were not helpful for the negative binomial model. It seems that the problem is not extra zeros but the existence of another group of policyholders. Therefore, assuming that we have two distinct subpopulations, we may move towards a finite mixture model. In this case, using a 2-finite mixture of regression models, a large improvement was obtained by moving from one component Poisson to a 2-finite mixture of Poisson regression models. Note that this improvement is better than that obtained with the zero-inflated Poisson model. However, as the best fit is obtained by the 2FMNB, it seems that there is still some extra overdispersion which needs to be modelled appropriately, assuming within each component an overdispersed distribution like a negative binomial. Figure 1shows boxplots for the fitted mean values per component for both mixture models. We can observe that the group separation is not the same for the two models. Different models can have similar likelihoods but very different properties and potential. Comparing Poisson and negative binomial mixtures we see that they model different aspects and therefore, as they are close in likelihood terms, can focus on separate things. The first component for the Poisson model is more concentrated towards 0. The opposite is true for the second component. Clearly, 2FMNB fits the data better than the 2-finite mixture of Poisson regression models. Figure 1. Boxplots of the fitted means for each of the two components for both models. Figure 1shows the distinct characteristics of the two assumed distributions. For the case of the Poisson distribution, as the variance is determined from the mean, we see that the two components are further away as an attempt to model the excess of variance. Recall that the total variance is in fact the Risks 2020,8, 10 8 of 13 sum of the between variance (how much the components differ) and the within variance (inside each component). In contrast, in the negative binomial case, the two components are closer since the extra overdispersion parameter regulates the variability. This is also a warning that use of the Poisson model can lead to an erroneous inference of the mean for each component. From Figure 1, we can also observe that the group separation is characterised by a low mean for the first component and a high mean with higher variance for the second. One may assume that this group separation is revealed by driving characteristics, such as driving ability, aggressiveness or degree of obeying traffic regulations, that are the source of the unobserved heterogeneity. In this case, we can consider those policyholders who belong to the first component to be “good” drivers, whereas policyholders in the second component can be considered “bad” drivers. Table 3summaries the results for the 2FMNB and the case with k= 1, i.e., no mixture. For the 2FMNB, we report the estimated regression coefficients for each component and p-values for testing the hypothesis that the variable is statistically significant. For the simple negative binomial model, we report the coefficient and standard p-values based on the Wald test. Table 3. The fitted models for both the negative binomial and the 2FMNB. The p-value for the 2FMNB refers to that of LRT when the variables is removed from both components, whereas for the simple negative binomial it refers to the Wald test. 2FMNB Negative Binomial 1st comp. 2nd comp. p-Value Estimate p-Value Intercept −6.1420 −1.1364 <0.0001 Intercept −2.4144 <0.0001 GEN 0.2633 0.0086 0.0124 GEN 0.0774 0.0103 URB 0.3407 −0.0762 0.0017 URB 0.0165 0.4870 ZON 0.3745 0.0564 <0.0001 ZON 0.1324 <0.0001 LIC 0.2413 −0.2423 0.0448 LIC −0.1610 0.0230 LOY 0.3707 0.1289 <0.0001 LOY 0.2019 <0.0001 COV 3.1438 0.6373 <0.0001 COV 1.0024 <0.0001 POW 0.2502 0.1148 <0.0001 POW 0.1440 <0.0001 φ0.2321 0.6051 φ0.2527 π0.6686 0.3314 For the 2FMNB model, to assess the significance of the variables, we calculated a Likelihood Ratio Test (LRT) statistic. Note that this, as a variable selection problem, is not standard, as each covariate appears in both components. Also note that even since standard errors can be derived through the Hessian, such a procedure can be very unstable. Also bootstrap based standard errors can be very time-consuming. Therefore, to see the importance of the covariates, we maximised the log-likelihood with and without each variables and we obtained the LRT, compared with a χ2 distribution with 2 degrees of freedom. Also, note that covariates that perhaps were not significant for a simple model (no mixture) can be significant in the mixture model, as the two components can allow for separate effects, which are lost when combining to one model. Comparing the models from Table 3, we can see some interesting points. First, coefficient estimates of the negative binomial model are in some way a linear combination of coefficient estimates of each component of the 2FMNB. Second, in the negative binomial model, only URB is not significant at a level of 95%, whereas in the 2FMNB all covariates are significant at that level: the URB variable that is deemed significant for the mixture is not significant in the simple model. The reason is that they have opposite signs in the mixture, and therefore when we estimate one coefficient for the simple case, the effect is cancelled out: we estimate some average effect which is close to zero and has large variance. This implies that the mixture can more clearly reflect the importance of the variables and the existence of two groups of policyholders that behave in different ways with regard to these a priori