Using trajectory analysis to test and illustrate microsimulation outcomes
Full text
Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 3 of 17 ReSeaRch aRtIcle *For correspondence: janne. salonen@ etk. fi http:// creativecommons. org/ licenses/ by/ 4. 0/ This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited. Author Keywords: finite mixtures, trajectory analysis, groupbased modeling, dynamic microsimulation © 2019, Salonen et al; DOI: https:// doi. org/ 10. 34196/ ijm. 00198 Using trajectory analysis to test and illustrate microsimulation outcomes Janne Salonen1*, Heikki Tikanmäki1, Tapio Nummi2 1Finnish centre for Pensions, helsinki, Finland; 2University of tampere, tampere, Finland Abstract We propose a new datadriven way of testing and visualizing dynamic microsimulation outcome data. The proposed statistical methodology is based on trajectory analysis (Nagin, 1999), which can be used to identify several subpopulations from a population measured longitudinally. We briefly introduce the statistical basis of trajectory analysis and discuss its use in the context of microsimulation. Finally, we report our results from the Finnish microsimulation model ELSI (Tikanmäki etal., 2014; Tikanmäki etal., 2015) to illustrate the possibilities and benefits of this technique. Trajectory analysis is available in many statistical software packages (e.g., SAS, R, Stata and Mplus). We conclude that trajectory analysis is a useful tool for investigating microsimulation outcomes. JEL classification: C63, C10, H55 DOI: https:// doi. org/ 10. 34196/ ijm. 00198 1. Introduction This paper has two goals. First, we present a statistical technique called groupbased trajectory analysis and demonstrate its usefulness in a microsimulation context. We use trajectory analysis to identify unknown groups or subpopulations that can yield valuable information that is not necessarily easily accessible by other means. Second, we use three examples to show how this technique, among others, can be used to validate models and to reveal possible misspecification of the microsimulation model. Trajectory analysis has recently gained much popularity in a number of fields including psychology, criminology (Nagin and Odgers, 2010; Nagin, 2016; van der Geest etal., 2016), sociology (Hynes and Clarkberg, 2005; Don and Mickelson, 2014), education (Kokko etal., 2008), marketing (Mani and Nandkumar, 2016) and health sciences (Nummi et al., 2014; Nummi et al., 2017a). It has increasingly been used in studies on labor market attachment (Peutere etal., 2015; Nummi etal., 2017b). To our knowledge, trajectory analysis has very seldom been used in a microsimulation context. Dynamic microsimulation results are often reported using some ex ante classification (e.g., education, age, or labor market state). The trajectory approach offers key advantages over ex ante classification, which stem from the a priori use of an economic taxonomy. The basic advantage of trajectory analysis is that it can reveal latent patterns in longitudinal data that might otherwise remain hidden, hence complementing ex ante classifications. Furthermore, the use of a formal statistical methodology has the capacity to distinguish chance variation across individuals from real differences caused by latent subgroups (Nagin and Odgers, 2010). In this paper, we propose two kind of trajectory analysis implementations. First, the data can be stratified by various background factors (e.g., by labor market state) and thus the latent subgroups are to be found within the strata. This would yield information on, for example, how common the earnings trajectories (by population state) are in the stratas investigated. Second, another way would be to classify the trajectory groups by some background factor after trajectory analysis. This would yield information on how the ex ante classifier (e.g., labor market state) is divided in trajectory groups.
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 4 of 17 Population heterogeneity is an important topic in dynamic microsimulation. Credible heterogeneity in a simulated population is a desired property of any microsimulation model. Trajectory analysis is a powerful method for demonstrating this heterogeneity. It allows for the analysis of any microsimulation outcome with interesting population heterogeneity. For instance, in the Finnish ELSI model we could analyze the individual population state, pension contributions, working lives or incomes. The technique is especially suited to cohortbased analysis, but it can also handle several cohorts simultaneously. The approach we propose has a number of advantages. First, trajectory analysis has the potential to reveal interesting subgroups of individuals. Second, the technique of trajectory analysis with normal distribution (like other members of the exponential family) is a wellestablished method, with software packages readily available for microsimulation practitioners (e.g. Jones etal., 2001; Leisch, 2004; Grun and Leisch, 2007; Haughton etal., 2009; Muthén and Muthén, 2010). Third, trajectory analysis is a flexible method that can be applied to both crosssectional data (single period measurements), and more importantly, to longitudinal data (multiple periods of measurements). The most common metric for indexing time is age or year. Another possible metric would be time before or since a lifeevent. Descriptive analysis is often inadequate for the purposes of empirical research. Trajectory analysis can provide a more rigorous examination with timedependent covariates or risk factors affecting trajectory group membership (Nagin, 2005; Jones and Nagin, 2007). In addition to single outcome analysis, trajectory analysis also allows for the simultaneous analysis of multiple outcomes Nagin etal., 2016). Indeed, there is now a growing body of research that uses multiple trajectory analysis (e.g. Hsu, 2015; Nummi etal., 2017b). Nummi etal. (2017b) provide an example of the simultaneous modeling of employment, education, unemployment and parental leaves using a multivariate trajectory model. These abovementioned extensions also provide interesting possibilities for microsimulation data analysis. The method has also its disadvantages. First, trajectory analysis can only be performed with discretetime data. Most dynamic microsimulation models are defined in discretetime, with timepoint intervals of one year (Zaidi and Rake, 2001; Li and O’Donoghue, 2013; Li etal., 2014). This leads to a second disadvantage, which can be described as state vs. event microsimulation. Trajectory analysis is easiest to implement with the calendar year as the time interval, that is, using state microsimulation models. Trajectory analysis is often applied to normally distributed data, but it is also applicable to discrete distributions such as binomial, Poisson, multinomial, etc., making this method a useful tool for the exploratory analysis of data sets. In this paper, we show three examples of trajectory analysis in microsimulation contexts. These examples cover labor market simulation topics such as wage earnings, education and pensions. Trajectory analysis of these outcomes illustrates the technique with different data distributions. 2. Trajectory analysis in microsimulation 2.1. Trajectory analysis Trajectory analysis is, in essence, the application of finite mixture modeling to longitudinal data. It can be used for modeling the unobserved heterogeneity of individuals measured longitudinally (e.g. Nagin, 1999; Nagin, 2005). In what follows we describe the statistical background for our application of trajectory analysis. The statistical foundation of trajectory analysis is on finite mixture modeling. Wellestablished in the field of statistics, the mixture modeling method concerns modeling a statistical distribution by a mixture (or weighted sum) of distributions (see Titterington etal., 1985; McLachlan and Peel, 2000). Böhning etal. (2007) give examples of the use of finite mixture modeling in various statistical applications. Two basic modeling approaches are called growth mixture modeling and groupbased trajectory modeling. They share the same analytical objective of measuring and explaining differences across population members in their developmental course. The difference between approaches lies in the way they model individuallevel heterogenity in developmental trajectories.
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 5 of 17 Growth mixture models depict the average trend of outcome and individualspecific variation around the average trend with random effects using the same parameters of change (Nagin and Odgers, 2010). Our specific aim is to identify individuals or objects in microsimulation data with the same kind of unknown developmental profiles (trajectories or subgroups). The microsimulation population is then splitted into several subpopulations. Let yi=(y i 1,y i 2, ... ,y iT )′ represent the sequence of measurements on an individual i over T periods and let f i (yi|X i ) denote the marginal probability distribution of yi with possible timedependent covariates Xi . It is assumed that f i (yi|X i ) follows a mixture of K densities fi ( yi|Xi ) = K ∑ k=1 πkfik ( yi|Xi ) , K ∑ k=1 πk=1with πk> 0, (1) where πk is the probability of belonging to the subgroup k and f ik (yi|X i ) is the density for the kth subgroup. Trajectory analysis can handle discrete or continuous data. The simplest choice is to use the Bernoulli distribution {0, 1} for the mixture components f ik (yi|X i ) : but other members of the exponential family are often applied as well. It is assumed that given kth subgroup measurements are independent. For Bernoulli mixtures we can write fik ( yi|Xi ) = ∏T t=1 p y it itk ( 1 − p 1−y it itk ), (2) where the probability pitk is a function of covariates Xi . For modeling the conditional distribution of pitk we use the logistic regression model. For the ith individual, we can then use the equation pitk =exp ( x ′ iβk ) 1+exp ( x′ iβk ), (3) where x ′ i(t) is the tth row of Xi and βk is the parameter vector for the kth subgroup. For the analysis of continuous data, one alternative is the multivariate normal distribution fik ( yi|Xi ) = ( 2π ) − pi 2|Σik |− pi 2exp {− 1 2( yi − µik )′ Σ− 1 ik ( yi − µik )}, (4) where µik is a function of covariates Xi with parameters βk and Σik =σ2 kI a covariance matrix within kth subgroup. Thus, the measurements are assumed to be independent within subgroup k, with the variance σ2 k . One advantage of this assumption is that it considerably simplifies the likelihood function and thus yields a computationally lighter and more stable analysis. For modeling the trajectory mean in time t, simple linear models are usually applied, e.g. lowdegree polynomials. For our three examples we used cubic polynomial model x ′ βk=β0k+β1kt+β2kt2+β3kt3 (5) to model the development within the subgroup k in time (age). The mean model can also include other timedependent covariates or timestable covariates (riskfactors). For parameter estimation, Maximum Likelihood (ML) estimates can be calculated by maximizing the loglikelihood N ∑ i=1 f i over unknown parameters βik and σk,k= 1, ...,K (see e.g. Nagin, 1999; Jones etal., 2001; Jones and Nagin, 2007). In most software packages, the method used for ML estimation is the EM (Expectation and Maximization) algorithm (see Dempster etal., 1977; McLachlan and Peel, 2000). The algorithm in an iterative technique involving two steps. E step finds the expected log likelihood under current parameter estimates, the subsequent M step maximizes the expected log likelihood function. These steps are iterated until the estimates converge. When applied to trajectory analysis, the E step calculates the posterior probability for subgroup membership wik =πk f k (y i |X i ,∅ k ) f ( y i| X i,∅k) (6) under all parameter estimates ∅k . The estimated subgroup probabilities are then
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 6 of 17 ˆπ k=1 N N ∑ i=1 ˆ wik (7) Once the model parameters have been estimated, the posterior probability estimates provide a way to assign each individual to a specific subgroup or trajectory. Individuals can be assigned to the specific subgroup in which their posterior probability is highest. From the equations, we see that the assignment of individuals into groups takes into account regression parameter estimates, the number of groups, probability distribution, membership probability estimates and time span of longitudinal dataset. These are major topics in trajectory analysis and need careful consideration from microsimulation practitioner. Nagin (2005) discusses the selection of the number of groups (p. 78–87) and related statistical information criteria (p. 63–76) and the question of conditional independence (p. 26–27). The question of groups as real entities is discussed in Nagin and Odgers (2010) and Nagin (2016). Nagin (1999) discusses the use of group membership probabilities in the calculations, as well as the links between group membership probabilities and other timedependent covariates besides age (or time). The nature of trajectory groups in growth mixture modeling and groupbased trajectory analysis is discussed in Nagin, 2005, pp. 54–56), Nagin (2016) and Nathalie etal. (2017). Don and Mickelson (2014) discuss the good practices of trajectory analysis, especially in terms of model selection. Note that the subgroups revealed by trajectory analysis are not fixed constructs. They are just approximations of a more complex reality. Each individual belong to a specific group with a certain probability. It would be good if we could investigate the fit of the assumed model using the identified trajectory groups. However, since the groups are based on maximum posterior probability, this would probably introduce some correlation to withingroup residuals, as the actual group could also contain individuals from other groups with a small probability or weight. The correct approach from a statistical point of view, however, would be based on the investigation, using alternative covariance structures (for example, likelihoodratio type test statistics or some other information criterion). However, this kind of testing is impossible in Nagin’s basic model. This topic has been discussed in Nagin and Odgers (2010) and Nagin (2016). As with any statistical analysis of dynamic microsimulation outcome data, a change in the period of observation would naturally change the results of the model. The trajectory groups could also be affected to a certain extent. However, we believe that in our examples the basic main grouping structure would remain quite similar. 2.2. Finnish ELSI microsimulation model ELSI is a longitudinal microsimulation model (Tikanmäki etal., 2014; Dekkers and Van den Bosch, 2016) that is used to assess the development of the statutory pensions in Finland. The dynamic ageing model has been developed at the Finnish Centre for Pensions, a statutory cooperation body providing research and expertise services related to the Finnish pension system. ELSI model has been designed to assess the future earningsrelated pensions and the national pensions (Tikanmäki etal., 2017). It can also be used to analyze changes in the pension system and in the underlying demographic or macroeconomic conditions. One of its uses has been to assess the distributional effects of the pension reform of 2017 (Tikanmäki etal., 2015). The Finnish pension system is mainly based on pension rights accrued on the basis of the individual’s lifetime earnings. The only exception is the survivor’s pension, which accounts for no more than 6% of total pension expenditure. The ELSI model is therefore based on individuallevel information and calculations of pensions received in one’s own right. The model comprises both pension recipients and those still working. The model simulates each individual’s working life prior to retirement. The base population consists of all adults aged 18 or over covered by the social insurance system in Finland. Most of the material is drawn from administrative records maintained by the Finnish Centre for Pensions and the Social Insurance Institution of Finland. Information on educational level from Statistics Finland is also added to the model. The ELSI model provides the opportunity to run large populations. For the present analysis we used a random sample of 32% of the base population, which in numerical terms translates into 1.5 million individuals in the baseline year of 2012. The population is then simulated until 2085. Deceased people remain in the model and a new cohort as well as new immigrants enter the model each year. Consequently the population increases over the course of the simulation.
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 7 of 17 The ELSI model has a modular structure (Figure1). Each colored box represents a module of the model, while white boxes represent external sources of information. There are no feedback loops from later to earlier modules. The simulation starts from the population module, followed by the earnings module, pension modules, and the taxation module, and finally brings together the results. The population module has several functions. It simulates population and labor market transitions as well as educational changes. The population module is based on transition probabilities that are estimated from historical data for 2010–2014. The module also uses Statistics Finland’s official population forecast to replicate general trends in the sample population. Transition probabilities, with oneyear time steps, are by and large deterministic, that is, based on exogenous information. In the population module, we simulate a new population or labor market state for each individual based on transition probabilities. There are 22 states in the model, the most important of which are: active (employed), inactive, unemployed, various pension states for different pension benefits, and deceased. The labor market transitions are of Markovian type, which means that the transition probabilities are based on current state rather than former history. However, it is possible to add memory to a Markov process by extending the state space. For instance, in model ELSI we have three different active states. One for those employed first consecutive year, another for those employed second year and third one for the rest. Hence, unemployment risks may be higher for those who do not yet have an established labor market position. Education is also simulated in the population module. Postbasic education dynamics is based on age and genderspecific transition probabilities. Changes in education level are possible at any age, although they are not very common after age 35. Figure 1. Structure of the ELSI microsimulation model
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 8 of 17 The transition probabilities are updated each simulation year using the population level information produced by the semiaggregated LTP model (Tikanmäki etal., 2017). There is thus a simple alignment of microsimulation outcomes to macro level aggregates. Wage earnings are simulated in the earnings module. Wages are simulated for each individual annually based on the labor market state and an underlying earningsequation, which is a timeseries model with a stochastic component. The earnings equation takes also into account gender, age and level of education. Wages are simulated for active workers and those in partial retirement. The earnings module is described in Tikanmäki etal. (2014). The earningsrelated pension module calculates pension amounts based on the simulated time of retirement (retirement age) and lifecourse earnings. The pension calculation takes no shortcuts but is as detailed as possible given the data in use. The national pension and guarantee pension amounts are calculated as a residual of the earningsrelated pension, since they are income tested. The calculation of national pensions is described in Sihvonen (2015). The taxation module finalizes the substantive simulation. The previous modules have produced gross wages and pensions. In the taxation module, the current (2016) rules for income taxation are applied to both simulated wages and pension earnings. The calculation of income taxes is described in Sihvonen (2015). After the simulation run, the results are analyzed in the results module, which calculates aggregate results over the course of the simulation, based on individuallevel outcomes. Measures of the distribution (mean, percentage points, Gini coefficient) of pensions can be produced by ex ante classifiers such as gender, level of education and year of birth. Aggregate measures on the duration of working life and partition of lifecourse into active and passive stages are also calculated in the result module. The module collects individuallevel output data containing information on labor market state, wage earnings, residence, education level, pension earnings, pension benefit, working life, pension accrual, etc. Therefore the material is also available for the statistical analysis illustrated in this paper. The proposed statistical technique could be used with many other outcomes as well. In the following analysis we use a longitudinal individuallevel data set that covers the period from 2008–2085 with a 25% subsample of the original output data. The analysis is based on three outcomes: wage earnings, education level and pension earnings. The wage and education trajectories are presented for the male cohort born in 1995 and pension trajectories for the male cohort born in 1960. With trajectory analysis, it would be possible to analyze several cohorts at the same time so long as the data is longitudinal. Similar longitudinal data is produced in many other dynamic microsimulation models, too. For example, the Swedish SESIM model and the Norwegian MOSART model are in this respect similar to ELSI, providing rich individual or household level labor market output data (Fredriksen and Stølen, 2007; Klevmarken and Lindgren, 2008; Flood etal., 2012). 2.3. Trajectory model selection Choosing the number of mixture components K is an important stage of finite mixture modeling. The trajectory method requires the researcher to specify the assumed number of subgroups in the data. The optimal number of subgroups K can be assessed using information criteria, which are widely used in this context.[1] One commonly used criterion to assess the model fit is the Bayesian information criterion (BIC). The model with largest BIC is preferred (SAS implementation). One may also calculate the groupspecific average posterior probability over individuals to measure the fit. If this values exceeds the minimum threshold (at least 0.7) the fit is considered satisfactory. Nagin (2005, p. 63–77) provides an overview of model selection in trajectory analysis context. Ultimately, meaningful realworld interpretations of subgroups requires not just information criteria, but also theoretical evaluation and good judgment. 1. SAS package PROC TRAJ counts BIC and AIC routinely.There are other information criteria, like crossvalidation available, if the number of subgroups is critical for the microsimulation practitioner (see Nielsen etal., 2014).
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 9 of 17 In trajectory analysis missing data is considered missing completely at random according to the taxonomy layed by Little and Rubin (1987). Software packages handle missing data in such a way that all available data is used in the estimation. In other words, all individuals or objects with some missing longitudinal data values are included in the analysis. From Table1 we can see the BIC values of our example analyses on three outcomevariables (earnings, education and pension). The BIC score indicates that, a fivegroup solution fits the wage earnings and pension variables, while for education BIC increases with more groups. However, increasing the number of education groups further would yield some infinitesimal groups. Therefore our choice is the fivegroup solution for the wage earnings and pension outcomes, and the sixgroup solution for the education outcome variable. In this study the computations were carried out using SAS software package with accompanying PROC TRAJ application, an easy to use PC SAS procedure for analyzing Nagin’s model (e.g. Jones etal., 2001; Jones and Nagin, 2007; Andruff etal., 2009).[2 Appendix A.1 gives an example of the programming code for wage earnings trajectories (Example 1, k=5). The trajectory plots (Figures2–4) present conditional means of time points calculated over the simulation period. Relative sizes ˆπk of the subgroups are also presented in the figures. These plots are the main tool for interpreting the results obtained. The estimated model is summarized in Appendix (A.2), which includes groupspecific parameter estimates. The SAS procedure also produces a confidence interval (95%) and modelpredicted values by subgroup if necessary. 2. SAS and STATA (see Jones and Nagin, 2013) trajectory analysis package can be downloaded from the website: https://www.andrew.cmu.edu/user/bjones/. R package for trajectory analysis is Flexmix (see Leisch, 2004). Table 1. BIC scores and number of subgroups k Outcome 3 4 5 6 Earnings (N=3,325) –50288.7 –48178.8 –45974.7 –46375.9 Education (N=13,173) –549399.6 –544342.9 –537412.2 –535026.8 Pension (N=10,761) –834881.4 –811092.8 –786210.9 786234.1 Notes: BIC = log(L) –.5 * log(n) k, where L = log likelihood, n = sample size and k = number of parameters. Figure 2. Earnings trajectories Figure 3. Earnings trajectories, k=5
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 10 of 17 It must be emphasized that the trajectories in the following examples are group means, which indicates that there is a range of individual tracks around the trajectory curves. Software packages routinely output the group estimates and confidence intervals to visualize withingroup population heterogeneity. For the sake of readability, the model fit curves and confidence intervals are not presented in this paper. In the examples we use cubic age model (see Equation 5). The choice depends on the assumed development of the data over time. One advantage in PROC TRAJ is that the age or time model can be groupspecific, which means that for example one group can be fitted with a second order polynomial model, whereas another group can be fitted with a cubic model etc. 3. Examples 3.1. Earnings trajectories: binary outcome For this example, we have chosen a young cohort to illustrate a simulated lifecourse. The cohort, born in 1995, enters labor markets at age 17 (year 2012), and subsequent lifecourse including individual wage earnings is simulated in the ELSI model’s population and earnings modules. The example shows a trajectory analysis with binary outcome data. Individual yearly wage is dichotomized (yes = 1/no = 0), so wage earnings could also be interpreted as employment. In terms of content, wage earnings could be analyzed without transformations using continuous data (euros). The BIC score (Table1) indicates that a fivegroup solution yields the best fit for the earnings outcome, but we also present the solutions for three and fourgroups solutions. Figure2 shows that the relative sizes ˆπk vary somewhat with an increasing number of groups. Nevertheless all solutions have essentially the same major subgroups. In terms of content, we can see that there is a large group (k = 3: 49%, k = 4: 33% and k = 5: 29%) with strong labor market integration, or a strong earnings profile, until retirement age. Other groups show an early declining trend in employment. Labor market integration is weak in one group (k = 3: 12%, k = 4: 10% and k = 5: 8%), especially at older age. Unemployment and permanent disability draw individuals out of employment prior to oldage retirement. Retirement age for oldage pension in this cohort is 67 years and 9 months. Partial oldage pension can be drawn three years before retirement age. 3.1.1. elSI model recalibration This trajectory analysis experiment revealed a slight misspecification in the ELSI model’s earnings module (see spikes at age 65 in all trajectory solutions). The spikes are visible in meanbased trajectories, but they would not be seen in fitted means models. This goes to show how useful groupbased mean curves can be in revealing model misspecification. In the earnings module the mechanism for working after retirement did not work as intended. During the course of simulation, many of those who had stopped working earlier suddenly decided Figure 4. Education trajectories
Research article Methodology, Dynamic microsimulation Salonen etal. International Journal of Microsimulation 2019; 12(2) DOI: https:// doi. org/ 10. 34196/ ijm. 00198 11 of 17 to return to work after drawing their pension (Figure2). Of course, such behavior is not plausible on a larger scale. The Finnish pension system allows fulltime employment even after retirement on an oldage pension. However, wage earnings for oldage pensioners are typically quite low. The misspecification therefore did not have a major impact on the main results and was not observed in comparison with the LTP macro model. In the ELSI model earnings are calibrated with the LTP model. In practice, this is done by comparing projected average earnings each year by gender, population state and age. Any deviations imply changes in the ELSI model’s parameters. In the LTP model, working after retirement is not modeled in detail, and the relevant point of reference is provided by the corresponding statistical figures. In the ELSI model, the share of new retirees working after retirement is the same as in the observed statistics. Trajectory analysis showed that the pool of retirees working after retirement also included individuals who had not worked in the year preceding retirement, which was not intended. Following the trajectory analysis of simulated earnings, the ELSI model was recalibrated appropriately. There was no change in the aggregate figures, namely wage sum and number of people working after retirement. Trajectory analysis is not yet routinely part of ELSI model testing, but it could be incorporated as part of its validation procedures. 3.1.1.1. Composition of trajectory groups Checking the composition of the groups is a good way to validate also the trajectory analysis results. It is good practice to crosstabulate the trajectory groups by the available background factors of the microsimulation. We have added an example which shows the labor market state or population state of earnings groups (k = 5) at age 60 in the simulated lifecourse (Figure3). Depending on the available information of the microsimulation model, also other background factor could be crosstabulated. The labor market states (Table2) confirm the earnings trajectory results. The high and stable wage earnings (employment) group G4 (29%) consists of people who are mainly (89.9%) in active labor market states. The weak attachment group G1 (8%) consists of individuals who had a disability or Table 2. Composition of earnings trajectories at age 60, per cent G1 (8%) G2 (11%) G3 (17%) G4 (29%) G5 (35%) Total Active first year 0.8 2.3 2.3 0.5 3.4 2.1 Active second year 0.4 0.6 0.7 0.3 0.9 0.6 Active (working) 2.3 51.1 1.2 89.9 54.2 50.6 Unemployed or in education 1.5 9.7 6.2 1.6 10.4 6.5 Sickness benefits . 1.1 0.4 . 2.5 1.1 Partial disability pension and active . 1.4 0.2 2.3 0.4 1 Partial disability pension and nonactive 0.8 1.1 0.9 . 0.7 0.6 Full disability pension 23.4 2.8 26.9 0.5 2.9 7.9 Full disability pension second year 1.1 1.4 3.4 . 2 1.6 Oldage pension . . 0.5 0.1 0.7 0.4 Only national disability pension 15.5 . . . . 1.2 Out of labor markets I 18.9 1.4 17.2 0.3 1.8 5.3 Out of labor markets II 14.3 . 2.5 . 0.8 1.9 Inactive 5.7 18.8 30.8 1.5 15.5 13.8 Deceased 15.5 8.2 6.9 2.8 3.8 5.5 Total 100 100 100 100 100 100