Full text
A Monte Carlo method to estimate the confidence intervals for the concentration index using aggregated population register data Sonja Lumme •Reijo Sund •Alastair H. Leyland •Ilmo Keskima ¨ki Received: 2 April 2014 / Revised: 19 January 2015 / Accepted: 5 February 2015 / Published online: 18 February 2015 The Author(s) 2015. This article is published with open access at Springerlink.com Abstract In this paper, we introduce several statistical methods to evaluate the uncertainty in the concentration index (C) for measuring socioeconomic equality in health and health care using aggregated total population register data. The Cis a widely used index when measuring socioeconomic inequality, but previous studies have mainly focused on developing statistical inference for sampled data from population surveys. While data from large population-based or national registers provide complete coverage, registration comprises several sources of error. We simulate confidence intervals for the Cwith different Monte Carlo approaches, which take into account the nature of the population data. As an empirical example, we have an extensive dataset from the Finnish cause-of-death register on mortality amenable to health care interventions between 1996 and 2008. Amenable mortality has been often used as a tool to capture the effectiveness of health care. Thus, inequality in amenable mortality provides evidence on weaknesses in health care performance between socioeconomic groups. Our study shows using several approaches with different parametric assumptions that previously introduced methods to S. Lumme (&)I. Keskima ¨ki The Social and Health Systems Research Unit, The Department of Health and Social Care Systems, The National Institute for Health and Welfare (THL), P.O. Box 30, 00271 Helsinki, Finland e-mail: [email protected] I. Keskima ¨ki e-mail: [email protected] R. Sund Department of Social Research, Faculty of Social Sciences, Centre for Quantitative Methods, University of Helsinki, P.O. Box 33, 00014 Helsinki, Finland e-mail: [email protected] A. H. Leyland MRC/CSO Social and Public Health Sciences Unit, University of Glasgow, 200 Renfield Street, Glasgow G2 3QB, Scotland, UK e-mail: [email protected] I. Keskima ¨ki School of Health Sciences, University of Tampere, Tampere 33014, Finland 123 Health Serv Outcomes Res Method (2015) 15:82–98 DOI 10.1007/s10742-015-0137-1
estimate the uncertainty of the Cfor sampled data are too conservative for aggregated population register data. Consequently, we recommend that inequality indices based on the register data should be presented together with an approximation of the uncertainty and suggest using a simulation approach we propose. The approach can also be adapted to other measures of equality in health. Keywords Monte Carlo simulation Health and health care register data Equality Concentration index Confidence interval Amenable mortality 1 Background A major health policy goal in many countries is to reduce disparities in health and in access to and the quality of health care. Measuring these disparities is a challenge. In order to obtain extensive and precise knowledge of inequalities, comprehensive methods to study equality are necessary. Most studies of equality in health and health care, and most methodological papers, have focused on survey data. Register-based data provide another possible source of data for equality studies. So far, good-quality individual-level administrative data including information on socioeconomic status have been available only in a few countries, such as the Nordic countries, but the importance of such data is likely to increase as better information systems increasingly become available and changes in data privacy regulations will enable broader utilization of the individual-level data in other countries. Register-based data are typically secondary data, i.e., they have not been collected for the purposes of specific studies (Sund 2003). Another, possibly even more important, difference is that register-based data often contain total populations instead of samples. In other words, it may be invalid to use statistical methods that assume sampling variation is the source of uncertainty when measuring the phenomenon of interest. For example, when estimating the uncertainty of the measured indicator from sample data, only sampling error is traditionally taken into account. Other possible sources of uncertainty are ignored. When using total population data, such sampling error does not exist. It is, however, obvious that other sources of error exist, since many events such as deaths are assumed to be stochastic and consequently produce a natural variability in vital statistics (Brillinger 1986). In addition, people in the register at one particular time could be seen as a sample of a super-population, and recorded events on these people can be considered to be one of a series of possible results that could have occurred under the same circumstances (Curtin and Klein 1995). It is, however, a complicated task to assess the uncertainty in the indicator of interest using population-based data (Sørensen et al. 1996). There are multiple sources of errors that can affect the uncertainty, and the sources evidently vary between situations (Sund 2003). The quality of the data is the main influence on uncertainty. Errors in the data may have originated at the stage of registration due to varying practices and accuracy in the processes. Merging different databases, data handling (for example through aggregation of the data), or processing errors may form challenges. Deterioration of the data quality is possible also later in the analysis phase as a result of mistakes in variable coding or programming errors. In addition to data quality, other sources may introduce uncertainty into the indicator such as the definition of the variables or coding practices. What if the variable is recorded correctly, but does not describe the phenomenon under examination Health Serv Outcomes Res Method (2015) 15:82–98 83 123
for all individuals properly? The uncertainty is commonly quantified using a confidence interval which provides a means of assessing and reporting the uncertainty and is intuitively straightforward to interpret. The concentration index (C) is a widely used indicator for the quantification of socioeconomic equality in health and in the use of health care (e.g., van Doorslaer et al. 1997; Wagstaff 2000; Vikum et al. 2012). The Cgives comprehensive summary information about the whole distribution of the studied outcome in a single value, which is a particular advantage when making comparisons in time or between genders, areas, countries, or hospitals. In addition, it has the benefit that the level of the inequality can be visualized with the concentration curve. This paper is about the methodology of the concentration index when measuring socioeconomic equality in health and health care using aggregated register data. The measured health or health care variable can be, for example, deaths, hospitalizations, or certain procedures. In this context, the term ‘‘aggregated data’’ is taken to mean data that are originally individual level and are later grouped by income. Assessment of the registerbased estimates involves the above-mentioned uncertainties. Thus, we introduce several techniques to evaluate uncertainty by calculating confidence intervals for the Cusing empirical data, as the majority of the previous studies using and developing the methodology of the Chave focused on survey data or hypothetical data, which require different methods (for example, see Kakwani et al. 1997; Waters 2000; Burstro ¨m et al. 2005; van Ourti 2004; Wagstaff 2005; van Doorslaer et al. 2006; Chen and Roy 2009; Clarke and van Ourti 2010; Konings et al. 2010; Chen et al. 2012). In many of these papers, the uncertainty in the Chas not been assessed, but Kakwani et al. (1997), van Ourti (2004), Chen and Roy (2009), Konings et al. (2010), and Chen et al. (2012) used improved methods to estimate the uncertainty. The uncertainty in the indicator is an essential question, particularly when making comparisons of equality. Our approaches to estimate the uncertainty in the Care based on several Monte Carlo simulations. Simulation has previously been shown to be an effective method for the Cas well as other inequality indices using survey data (Chen et al. 2012; Mills and Zandvakili 1997; Sergeant and Firth 2006; Modarres and Gastwirth 2006; van Ourti and Clarke 2011). We demonstrate the results of this study empirically using an extensive Finnish aggregated register dataset on amenable mortality. We compare our results to the commonly used standard regression method and the improved method developed by Kakwani et al. (1997). 2 Methods The concentration index (C) can be used to measure the degree of socioeconomic inequality in health and health care across the distribution of the whole study population (Wagstaff et al. 1989). The index is based on the Gini coefficient, which is used to assess inequality in income or wealth. The Cis based on the concentration curve L(s), which is a tool to visualize the degree of inequality. When using aggregated data, L(s) plots the cumulative proportion of the health outcome variable against the cumulative proportion of the population (s), ranked by socioeconomic group (SEG) from the least to the most advantaged. The Cis defined as twice the area between the diagonal and L(s). In a case of complete equality, the Cgets a value of 0. Negative values indicate a disproportionate concentration of the health outcome among those classed as disadvantaged and vice versa. The Cis restricted to values between -1 and 1 when the health variable is not binary (Wagstaff 2005). 84 Health Serv Outcomes Res Method (2015) 15:82–98 123
For aggregated data—in which the groups comprise SEGs and the socioeconomic indicator is measured on an ordinal scale—the quantitative measure of inequality can be estimated as C¼2 yX G g¼1 ygfgRg1ð1Þ where y g is the health outcome (such as the annual mortality rate) of the gth SEG, and yis the mean of the y g across SEGs weighted by the population share f g . The R g is the relative rank of the SEG, defined as R g =P c=1 g-1 f c ?0.5 f g and indicates the cumulative proportion of the population up to the midpoint of each group interval. In this study, y g denotes the directly age-standardized amenable mortality rate (per 100,000 person-years) of the gth income group: yg¼PI i¼1 dig pigwi;where iis the age group, d ig is the number of deaths and p ig is the population size in the ith age group of the gth SEG, w i is the weight of the age group according to the standard population. The sum of the standard population is 100,000, i.e., P i=1 I w i =100,000. The Chas also been estimated using a weighted least squares method (WLS) (Lerman and Yitzhaki 1984; Wagstaff et al. 1991). The use of aggregated data necessitates the use of weights. The slope parameter b 1 of the regression model has a computational equivalence with the Cand is obtained from the WLS model 2r2 R yg yffiffiffiffiffi pg p¼b0ffiffiffiffiffi pg pþb1Rgffiffiffiffiffi pg pþeg;ð2Þ where p g is the population size in the gth SEG. The size of the weight indicates the power of the information contained in the associated observation. Thus, if SEGs are of equal size, the weights do not have any effect since each group has equal influence on the final estimate. The variance r R 2 is the weighted variance of the rank R g , defined as r R 2 = P g=1 G f g (R g -0.5) 2 . This convenient regression method gives the C, irrespective of whether the principal model assumptions apply, because the regression method is an artificial technique to calculate the C(Kakwani et al. 1997). Thus, the possible serial correlation resulting from the ranked nature of the independent variable (rank is ordered and cumulative) does not affect the estimated regression coefficient. In addition, it is important to note that a linear relationship between the dependent variable and the rank is not necessary due to the artificial nature of this estimation. Both models (1) and (2) can be applied to population or sample data to calculate the C. The standard error of the regression slope in the WLS model (2) describes the variability of the estimate around the unknown slope parameter b 1 . However, in order to construct a confidence interval for the b 1 using formula (2), the key WLS regression assumptions should not be violated. The independence of errors may be violated due to the above-mentioned serial correlation causing either underor overestimated standard errors. The error terms, e.g., in the regression, either have to be normally distributed and independent, or the number of observations in the regression has to be sufficiently large. Usually, when studying equality, the number of SEGs is relatively low (5–20); thus, the latter assumption is unlikely to be met. Due to its simplicity, using this regression method in statistical packages appears to be a conventional means of obtaining confidence intervals also for the C(denoted as REG in this study). Kakwani et al. (1997) developed estimators of the standard error of the C, which take into account the serial correlation in the data applicable to data drawn from a sample. We denote this technique as KWV in this study. Health Serv Outcomes Res Method (2015) 15:82–98 85 123
In this study, we introduce five different Monte Carlo simulation techniques to estimate the confidence interval for the Cusing income as a socioeconomic indicator. These simulation techniques differ from each other in distributional assumptions and in the phase of the simulation process; one technique simulates the outcome variable of the regression method (2), three of them simulate observed events (d ig ), and one simulates observed ageadjusted rates (y g ). As one of the techniques applies the regression method (2), the rest of the simulation techniques apply either the regression or the formula method (1). Our techniques can be applied to the datasets where socioeconomic variable is grouped by proportions, for example income quintiles. The SEGs must be defined by the proportions of the person years (ordered by income). The number of person years in each income group can be assumed to be rather stable when using register data due to large datasets. Due to fixed proportions, the ranking is fixed in all our simulation techniques when estimating the uncertainty (the possible miscoding of the income record and health variable) and this allows modeling variation only in the health outcome variable. Consequently, even though the proportions in each income group are fixed, the uncertainty involved in recording income information is incorporated in our method. The biasing effect of correlation between the rank and the outcome variable is avoided in all our approaches because the standard error is not estimated from the regression model; we simulate the original data to estimate the uncertainty of the C. In addition, the advantage of our approaches is that they aim to model the assumed error of the data and the concentration curve and not the error of the slope parameter of the fitted regression line. One of these simulation techniques was developed in our recent study (Lumme et al. 2012) in which we made a simple assumption of uncertainty around the dependent variable 2r R 2 (y g /y g y) in Eq. (2). We denote this technique as MC in this study, and it can be applied only to the regression method (2) to estimate the C. We accounted for that uncertainty by assuming 2r2 R yg yNðlg;r2 gÞ, where the mean l g is the observed value of 2r2 R yg yfrom the dataset and the variance r g 2 is the observed value of 2r2 R yg yffiffiffiffi ng p 2 ;with n g being the number of events (such as amenable deaths) in the gth income group. We then re-estimated the Cby replicating the regression estimation 10,000 times to account for the uncertainty. The lower and upper limits of the 95 % confidence interval of the Cwere obtained as the 2.5 and 97.5 percentiles of the distribution of the simulated slopes. The median of the distribution of the slopes is equal to the Ccalculated from the observed dataset, determined by setting the observed values of the dependent variable in Eq. (2) as the expected values in the simulations. The property that the median of the distribution of the slopes is equal to the Cresults straight from the normal distribution assumption, since the median and the mean are equal by definition. This produces symmetrical confidence intervals around the observed C. If there is a reason to assume larger errors (for example due to consistent miscoding of some variable), the variance r g 2 can be enlarged by reducing the factor of the denominator n g . This would indicate fewer events in an income group and thus assume more variation. Changing the size of the disturbance, however, would not change the median of the distributions (i.e., the C), but naturally it would enlarge the confidence intervals of the C. The second model (denoted as MC rate) assumes that the age-adjusted rates follow normal distributions ygNðyg; y2 g ngÞ;, and both methods (1) and (2) can be used to estimate the confidence intervals. In the next approach, we made more assumptions and developed the MC technique further to model the uncertainty in more detail. A requirement for the independence of the 86 Health Serv Outcomes Res Method (2015) 15:82–98 123
error terms is not needed because this method does not use errors estimated from a regression. It repeats the estimation of the Cby allowing some variability in the health outcome (events) by income groups and also in the total number of events. Both methods (1) and (2) can be used to assess the C. We denote this technique as BIN. The variability is approximated from the observed data with the following assumptions and steps such as: 1. The observed p ig (the population size), the denominator of the rate, is held fixed in the simulation. This is based on the assumption that the information on age and personyears in the registers is perfect. 2. The second assumption concerns the events that are treated as being random. The number of events is allowed to vary due to fact that there is some error in the coding of the events (such as causes of deaths). The observed total number of events in age group iis the sum over income groups P g=1 G d ig =D i . 3. The number of events is allowed to vary between income groups within the age group. This is permitted because the income information presumably does not exactly measure the person’s real wealth. It might not describe the real wealth level of a person since all assets are not recorded in the administrative registers. Income is obtained from multiple administrative sources and, in addition, may vary considerably even over a short period. Now, the number of events in group ig is simulated assuming to follow a binomial distribution d ig *B(D i ,q ig ). The denominator D i is the same for all income groups within the same age group. Probabilities q ig (with the constraints that P g=1 G q ig =1 and 0 Bq ig B1) are estimated from the observed data qig ¼dig Di: 4. Simulation step (3) is repeated Ntimes; thus, Nis the number of simulated datasets. 5. Next, Nsets of age-adjusted rates are calculated using the simulated number of events, the observed person-years at risk from the original dataset, and the weights from the original standard population. 6. Now, Nvalues of the Care calculated using methods (1)or(2) from the simulated data yielding a distribution of the C. The 2.5 and 97.5 percentiles of this distribution comprise the 95 % confidence intervals for the C. Binomial distribution may not be symmetric, but with large nand not too extreme q ig , it is in practice quite often very symmetric. Thus, the median of the distribution of the slopes is likely very close to the Ccalculated from the observed data. To test the robustness of the above-mentioned simulation techniques to estimate the confidence interval for the C, we performed more analyses using different assumptions. The fourth model (denoted POIS) is a simulation approach equivalent to BIN except the number of events in the step (3) follows a Poisson distribution d ig *Pois(k ig ) where the parameter k ig is the observed number of events in a group ig. Poisson distribution is asymmetrical when the mean is small. However, as the mean becomes large, the distribution becomes more and more symmetric and approaches normal distribution. In the fifth simulation method (denoted MN), the total number of events within age group D i is held fixed. The number of deaths is, however, allowed to vary between income groups within each age group. The number of events is simulated from a multinomial distribution with parameters D i and q, and mean E{X ig }=D i q ig with the constraint that P g=1 G X ig =D i . The probabilities q i ={q i1 ,…,q iG } (with constraints P g=1 G q ig =1 and 0\q ig B1) are estimated from the observed data qig ¼dig Di: Table 1presents all five methods and related modelling assumptions. Health Serv Outcomes Res Method (2015) 15:82–98 87 123
3 Empirical example As an empirical example, we used Finnish register data on deaths amenable to health care interventions. Monitoring socioeconomic inequalities in mortality amenable to health care interventions—which is used to measure health system performance based on certain premature deaths that should not occur if health care works effectively and is timely— provides useful information on changes in differentials in health service utilization and effectiveness (Schwarz and Pamuk 2008). These amenable deaths are an indication of potential weaknesses in health care that can then undergo more in-depth investigation (Nolte and McKee 2004). Table 1 Assumptions of the simulation methods Method Variable simulated Modelling assumptions Population assumptions Estimation of the C MC 2r2 R yg y2r2 R yg yNlg;r2 g l g is the observed value of 2r2 R yg y The regression method (2) r g 2 is the observed value of 2r2 g yg yffiffiffiffi ng p n g is the number of events in the income group MC rate y g ygNy g; y2 g ng y g is the observed rate of the event The arithmetic method (1)or the regression method (2) BIN d ig d ig *B(D i ,q ig )D t is the observed number of events in age group i:P g=1 G d ig =D i Method (1)or (2) The population size in group ig is held fixed q ig =d ig /D i The number of events is allowed to vary between income groups within the age group POIS d ig d ig *Pois(k ig )k ig is the observed number of events in group ig Method (1)or (2) MN d ig d ig follows a multinomial distribution with parameters D, and qand mean E{X ig }=D i q ig with the constraint that P g=1 G X ig =D i The probabilities qi¼ qi1;...;qig (with constraints P g=1 G q ig =1 and 0\q ig B1) are estimated from the observed data: q ig =d ig /D i Method (1)or (2) The number of events within age group D i is held fixed The number of deaths is allowed to vary between income groups within each age group 88 Health Serv Outcomes Res Method (2015) 15:82–98 123
Table 2 List of causes of death considered amenable to health care and corresponding ICD-10 codes Place of intervention Cause of death Age ICD-10 Primary health care Timing of intervention Primary prevention Intestinal infections 1–14 A00–09 Diphtheria, Tetanus, Poliomyelitis, and Varicella 1–74 A35–36, A80, B01 Whooping cough 1–14 A37 Measles 1–14 B05 Rubella 1–74 B06 Scarlatina 1–74 A38 Meningococcus 1–74 A39 Erysipelas 1–74 A46 Legionellosis 1–74 A48.1 Malaria 1–74 B50–54 Streptococcal pharyngitis 1–74 J02.0 Cellulitis 1–74 L03 Early detection and treatment Tuberculosis 1–74 A15–19, B90 Malignant neoplasm of colon and rectum 1–74 C18–21 Melanoma of skin 1–74 C43 Malignant neoplasm of skin 1–74 C44 Malignant neoplasm of breast 1–74 C50 Malignant neoplasm of cervix uteri 1–74 C53 Malignant neoplasm of cervix uteri and body of uterus 1–44 C54–55 Malignant neoplasm of bladder 1–74 C67 Benign tumors 1–74 D10–36 Hypertensive disease 1–74 I10–13.I15 Cerebrovascular disease 1–74 I60–69 Improved treatment and medical care Diseases of the thyroid 1–74 E00–07 Diabetes mellitus 1–49 E10–14 Epilepsy 1–74 G40–41 All respiratory diseases (excl. pneumonia/ influenza) 1–14 J00–09, J20–99 Asthma 15–49 J45–46 COPD 15–49 J40–44 Specialized health care Septicemia 1–74 A40–41 Malignant neoplasm of testis 1–74 C62 Hodgkin’s disease 1–74 C81 Leukemia 1–44 C91–95 Rheumatic and other valvular heart disease 1–74 I01–09 Influenza 1–74 J09–11 Pneumonia 1–74 J12–18 Peptic ulcer 1–74 K25–28 Health Serv Outcomes Res Method (2015) 15:82–98 89 123
Our dataset comprised all resident Finnish citizens aged 1–74 in 1996–2008. For this population, we received yearly information on deaths from an amenable cause including individual demographic and socioeconomic variables such as income, age, and gender. Due to data protection regulations, all variables were categorized. By means of unique identification codes, the information on mortality came from the cause-of-death register and the demographic variables came from the annual individual-level employment statistics database. Both registers are compiled and maintained by Statistics Finland. As an indicator of socioeconomic status to study equality, we had disposable family net income, adjusted for family size on the OECD equivalence scale (OECD 1982) and categorized into 20 income groups according to the Finnish income distribution, separately for each year. The income record applied was for the year before death. Age was grouped from 1 to 4 years and then in 5 year age bands. The selection of causes of death considered amenable to health care focuses on conditions for which effective clinical interventions exist in people \75 years old and in this study was an adaptation of classifications by Page et al. (2006), Nolte and McKee (2008), and McCallum et al. (2013) (Table 2). Causes of death (as a main cause) were coded according to the 10th Revision of the International Classification of the Diseases (ICD). When age-standardized values were required, we calculated annual amenable mortality rates (per 100,000 person-years) for 20 income groups in 1996–2008. The ageand income-specific rates, i.e., the number of amenable deaths as a proportion of person-years in the follow-up of the corresponding population, were directly age-standardized to the European standard population (Waterhouse et al. 1976). We repeated the simulation approaches 10,000 times in our analyses; thus, Nwas 10,000, which we found to be a sufficient number of runs, since adding more runs did not change the lengths of the confidence intervals. The computer time with 10,000 repetitions for one simulation was negligible using a standard computer system for all methods. We used SAS (SAS Institute Inc., Cary, NC, USA) version 9.2 to analyze the data. 4 Results 4.1 Overview of data In Finland in 1996, according to our definition, the total number of deaths considered amenable to health care interventions was 4087, of which 52 % occurred among men. The number of amenable deaths decreased evenly during the follow-up (pvalue for linear Table 2 continued Place of intervention Cause of death Age ICD-10 Appendicitis 1–74 K35–38 Abdominal hernia 1–74 K40–46 Cholelithiasis and cholecystitis 1–74 K80–81 Nephritis, nephrosis, and nephropathy 1–74 N00–09,N17–19, N25–27 Obstructive uropathy and prostatic hyperplasia 1–74 N13,N20–21, N35, N40 Maternal death All O00–99 Congenital cardiovascular anomalies 1–74 Q20–28 90 Health Serv Outcomes Res Method (2015) 15:82–98 123
Konings, P., Harper, S., Lynch, J., Hosseinpoor, A.R., Berkvens, D., Lorant, V., Geckova, A., Speybroeck, N.: Analysis of socioeconomic health inequalities using the concentration index. Int. J. Public Health 55, 71–74 (2010). doi:10.1007/s00038-009-0078-y Kunst AE: Cross-national comparisons of socio-economic differences in mortality. PhD thesis. Erasmus University, Rotterdam, the Netherlands (1997) Lahti, R.A., Penttila ¨, A.: The validity of death certificates: Routine validation of death certification and its effects on mortality statistics. Forensic Sci. Int. 115(1–2), 15–32 (2001) Lerman, R.I., Yitzhaki, S.: A note on the calculation and interpretation of the Gini index. Econ. Lett. 15, 363–368 (1984) Lumme, S., Sund, R., Leyland, A.H., Keskima ¨ki, I.: Socioeconomic equality in amenable mortality in Finland 1992–2008. Soc. Sci. Med. 75(5), 905–913 (2012). doi:10.1016/j.socscimed.2012.04.007 Manderbacka, K., Arffman, M., Lyytika ¨inen, O., Sajantila, A., Keskima ¨ki, I.: What really happened with pneumonia mortality in Finland in 2000–2008? A cohort study. Epidemiol. Infect. 141(4), 800–804 (2013). doi:10.1017/S0950268812001562 McCallum, A., Manderbacka, K., Arffman, M., Leyland, A.H., Keskima ¨ki, I.: Socioeconomic differences in mortality amenable to health care among Finnish adults 1992–2003: 12 year follow up using individual level linked population register data. BMC Health Serv. Res. (2013). doi:10.1186/1472-6963-13-3 Mills, J., Zandvakili, S.: Statistical inference via bootstrapping for measures of inequality. J. Appl. Econom. 12, 133–150 (1997) Modarres, R., Gastwirth, J.L.: A cautionary note on estimating the standard error of the gini index of inequality. Oxford Bull. Econ. Stat. 68, 385–390 (2006) Nolte, E., McKee, M.: Does Healthcare Save Lives? Avoidable Mortality Revisited. The Nuffield Trust, London (2004) Nolte, E., McKee, M.: Measuring the health of nations: Updating an earlier analysis. Health Aff. 27(1), 58–71 (2008). doi:10.1377/hlthaff.27.1.58 O’Donnell, O., van Doorslaer, E., Wagstaff, A., Lindelow, M.: Analyzing health equality using household survey data. The World Bank, Washington D.C (2008) Organisation for Economic Co-operation and Development (OECD).: The OECD list of social indicators. OECD Publications and Information Center, Paris (1982) Page, A., Tobias, M., Glover, J., Wright, C., Hetzel, D., Fisher, E.: Australian and New Zealand Atlas of Avoidable Mortality. PHIDU, University of Adelaide, Adelaide (2006) Schwarz, F., Pamuk, E.R.: Causes of death responsible for the widening gap in mortality among educational groups in Austria between 1981 and 1991. Wien. Klin. Wochenschr. 120, 547–557 (2008). doi:10. 1007/s00508-008-1009-2 Sergeant, J.C., Firth, D.: Relative index of inequality: definition, estimation, and inference. Biostatistics 7(2), 213–224 (2006). doi:10.1093/biostatistics/kxj002 Sørensen, H.T., Sabroe, S., Olsen, J.: A framework for evaluation of secondary data sources for epidemiological research. Int. J. Epidemiol. 25(2), 435–442 (1996). doi:10.1093/ije/25.2.435 Sund, R.: Utilisation of administrative registers using scientific knowledge discovery. Intell. Data Anal. 7(6), 501–519 (2003) Sund, R.: Quality of the Finnish hospital discharge register: A systematic review. Scand. J. Public Health 40(6), 505–515 (2012). doi:10.1177/1403494812456637 van Doorslaer, E., Wagstaff, A., Bleichrodt, H., Calonge, S., Gerdtham, U.G., Gerfin, M., Geurts, J., Gross, L., Ha ¨kkinen, U., Leu, R.E., O’Donell, O., Propper, C., Puffer, F., Rodrı ´guez, M., Sundberg, G., Winkelhake, O.: Income-related inequalities in health: Some international comparisons. J. Health Econ. 16(1), 93–112 (1997). doi:10.1016/S0167-6296(96)00532-2 van Doorslaer, E., Masseria, C., Koolman, X., OECD Health Equality Research Group: Inequalities in access to medical care by income in developed countries. Can. Med. Assoc. J. 174, 177–183 (2006) van Ourti, T.: Measuring horizontal inequality in Belgian health care using a Gaussian random effects two part count data model. Health Econ. 13(7), 705–724 (2004). doi:10.1002/hec.920 van Ourti, T., Clarke, P.: A simple correction to remove the bias of the gini coefficient due to grouping. Rev. Econ. Stat. 93(3), 982–994 (2011) Vikum, E., Krokstad, S., Westin, S.: Socioeconomic inequalities in health care utilisation in Norway: the population-based HUNT3 survey. Int. J. Equal. Health 11, 48 (2012). doi:10.1186/1475-9276-11-48 Wagstaff, A., van Doorslaer, E., Paci, P.: Equality in the finance and delivery of health care: some tentative cross-country comparisons. Oxford Rev. Econ. Policy 5(1), 89–112 (1989). doi:10.1093/oxrep/5.1.89 Wagstaff, A., Paci, P., van Doorslaer, E.: On the measurement of inequalities in health. Soc. Sci. Med. 33(5), 545–557 (1991) Wagstaff, A.: Socioeconomic inequalities in child mortality: Comparisons across nine developing countries. Bull. World Health Organ. 78(1), 19–29 (2000) Health Serv Outcomes Res Method (2015) 15:82–98 97 123
Wagstaff, A., van Doorslaer, E.: Measuring and testing for inequality in the delivery of health care. J. Hum. Resour. 35(4), 716–733 (2000) Wagstaff, A.: The bounds of the concentration index when the variable of interest is binary, with an application to immunization inequality. Health Econ. 14, 429–432 (2005) Waterhouse, J., Muir, C.S., Correa, P., Powell, J. (Eds.): Cancer Incidence in Five Continents, Vol. III. Lyon: International Agency for Research on Cancer, IARC Scientific Publications No. 15 (1976) Waters, H.R.: Measuring equality in access to health care. Soc. Sci. Med. 51(4), 599–612 (2000). doi:10. 1016/S0277-9536(00)00003-4 98 Health Serv Outcomes Res Method (2015) 15:82–98 123