Robust Estimation of Wage Dispersion with Censored Data: An Application to Occupational Earnings Risk and Risk Attitudes
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Pollmann, Daniel; Dohmen, Thomas; Palm, Franz Article — Published Version Robust Estimation of Wage Dispersion with Censored Data: An Application to Occupational Earnings Risk and Risk Attitudes De Economist Provided in Cooperation with: Springer Nature Suggested Citation: Pollmann, Daniel; Dohmen, Thomas; Palm, Franz (2020) : Robust Estimation of Wage Dispersion with Censored Data: An Application to Occupational Earnings Risk and Risk Attitudes, De Economist, ISSN 1572-9982, Springer US, New York, NY, Vol. 168, Iss. 4, pp. 519-540, https://doi.org/10.1007/s10645-020-09374-x This Version is available at: https://hdl.handle.net/10419/288476 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) De Economist (2020) 168:519–540 https://doi.org/10.1007/s10645-020-09374-x 1 3 Robust Estimation ofWage Dispersion withCensored Data: AnApplication toOccupational Earnings Risk andRisk Attitudes DanielPollmann1· ThomasDohmen2 · FranzPalm3 Published online: 17 September 2020 © The Author(s) 2020 Abstract We present a semiparametric method to estimate group-level dispersion, which is particularly effective in the presence of censored data. We apply this procedure to obtain measures of occupation-specific wage dispersion using top-coded administrative wage data from the German IAB Employment Sample. We then relate these robust measures of earnings risk to the risk attitudes of individuals working in these occupations. We find that willingness to take risk is positively correlated with the wage dispersion of an individual’s occupation. Keywords Dispersion estimation· Earnings risk· Censoring· Quantile regression· Occupational choice· Sorting· Risk preferences· SOEP· IABS JEL Classification C14· C21· C24· J24· J31· D01· D81 1 Introduction Important economic issues often center on the shape of distributions. Examples include questions relating to income inequality, the shape of wage offer distributions, or the riskiness of returns to financial assets. In various settings, empirical labor economists have been interested in measures of wage dispersion. More than often, such measures have to be estimated from censored data. For example, the March Current Population Survey (CPS), which contains survey responses on weekly earnings top-coded for anonymization purposes, has been used in several studies. Researchers have frequently dealt with this problem by multiplying top-coded * Thomas Dohmen [email protected] 1 QuantCo, Inc., Boston, USA 2 Rheinische Friedrich-Wilhelms-Universität Bonn, Bonn, Germany 3 Maastricht University, Maastricht, TheNetherlands
520 D.Pollmann et al. 1 3 earnings by a factor of 1.3 to 1.5 (e.g., Katz and Murphy 1992; Juhn etal. 1993). Other studies have relied on distributional assumptions to impute censored earnings in their data (e.g., Dustmann etal. 2009). Closely related, moments can typically be recovered if the shape of the distribution and the censoring rule are known. In many settings, however, the shape of the wage distribution is unknown and possibly itself of interest, and estimation methods that require parametric assumptions typically yield inconsistent estimates when these are violated.1 More advanced semiparametric methods have been used for social security earnings records matched to the CPS, which suffer from a much higher degree of censoring due to a legal contribution limit (e.g., Chay and Honoré 1998; Hu 2002). We present a measure of group-level dispersion that can be straightforwardly obtained from quantile regression (QR). Our method does not require parametric assumptions on the error terms and is as such consistent under heteroskedasticity and non-normality even for censored data. In addition, by using this simple-to-com- pute method, which is based on group coefficient estimates at different quantiles rather than residuals, we can avoid dealing with censored residuals. Our semiparametric approach allows to estimate differential patterns of dispersion across occupations. We are thus able to adequately characterize the entire conditional wage distribution while explicitly incorporating the dispersion effect of covariates. We then demonstrate the usefulness of the estimation procedure in an application in which we relate the estimated occupation-specific wage dispersion in the German labor market as a measure of occupation-specific earnings risk to the risk attitudes of individuals working in these occupations. In order to estimate the occupationspecific cross-sectional earnings risk, we rely on German administrative wage data from the IAB Employment Sample (IABS) that contains wage information censored at the statutory limit for social security contributions. The IABS offers great sample size, such that we are able to work with more precise occupation definitions than previous studies and reduce the effect of aggregation on variation. We then match the estimated wage dispersion measure of occupations to individuals in the German Socio-Economic Panel Study (SOEP) working in these occupations. The SOEP provides us with survey information on risk attitudes and other individual and household characteristics. Consistent with previous studies (e.g., Bonin etal. 2007; Fouarge etal. 2014) that have assessed the relation between occupational earnings risk and risk preferences, we find evidence of a statistically significant correlation between our measure of occupational earnings risk and the risk attitudes of individuals working in a particular occupation: Those who state to be more willing to take risks are more likely to work in occupations with higher cross-sectional wage dispersion. 1 For example, the Tobin–Amemiya maximum likelihood estimator (Tobin 1958; Amemiya 1973) and the two-step Heckit approach (Heckman 1976, 1979) are inconsistent under deviations from homoskedasticity (e.g., Maddala and Nelson 1975; Hurd 1979; Arabmazar and Schmidt 1981; Brown and Moffitt 1983; Donald 1995) and normality (e.g., Arabmazar and Schmidt 1982; Goldberger 1983; Paarsch 1984). The simulation study of Vijverberg (1987) for the case of non-normality shows that the estimated error variance is often seriously biased, which may trouble our dispersion analysis.
521 1 3 Robust Estimation ofWage Dispersion withCensored Data: An… Our empirical application is related to a large literature that investigates the relationships between risk preferences and occupational choice. Early studies (e.g., Bellante and Link 1981) have assessed how risk preferences affect the choice between private sector and public sector employment. The typical finding in this strand of the literature is that higher levels of individual risk aversion significantly increase the probability of working in the public sector (see, e.g., Guiso and Paiella 2005; Fuchs-Schündeln and Schündeln 2005; Dohmen and Falk 2010 for evidence based on Italian and German data). A second class of studies has focused on the relationship between risk preferences and the probability of self-employment, which is considered to be more risky than dependent employment. Using data for different countries and employing different measures of risk attitudes, these studies consistently find that a higher propensity to take risks increases the probability of being self-employed (see, e.g., Cramer etal. 2002; Guiso and Paiella 2005; Ekelund etal. 2005; Caliendo etal. 2009; Dohmen etal. 2011; Beauchamp etal. 2017 for evidence from the Netherlands, Italy, Sweden, Germany, Germany, and Sweden respectively). Most closely related to our empirical application are studies that have related proxies of risk aversion or direct measures of risk attitudes to occupational earnings risk. Saks and Shore (2005), for example, use data from the National Postsecondary Student Aid Survey in the U.S. and find that, as expected under decreasing absolute risk aversion utility, individuals with higher parental wealth more frequently choose college majors leading into occupations with greater conditional earnings variation (see also King 1974), as estimated on U.S. data from the Panel Study of Income Dynamics (PSID) and the Baccalaureate & Beyond survey.2 Bonin et al. (2007) and Fouarge etal. (2014) use direct measures of risk attitudes and relate them to an explicit statistic for the riskiness of occupations, the occupation-specific standard deviation of the residuals from a Mincer wage regression, in the German and Dutch labor markets respectively. They find a significant positive relationship between this cross-sectional earnings risk measure and individuals’ stated willingness to take risks. While Bonin etal. (2007) carry out all estimation on data from the SOEP, Fouarge etal. (2014) compute the occupation-specific cross-sectional earnings risk based on administrative wage data from Statistics Netherlands (CBS) and relate it to the self-reported risk attitudes of respondents to the ROA School Leavers Survey, which is based on the SOEP questions on willingness to take risks. Schulhofer- Wohl (2011) uses responses to the question on risky jobs in the Health and Retirement Study to relate them to the amount of income risk experienced by individuals, estimated based on data from matched social security earnings records. Schulhofer- Wohl classifies individuals into a low and a high risk tolerance group and finds that the latter carry significantly more of both aggregate and idiosyncratic risk. 2 There is a related literature on the relationship between risk preferences and educational choice (e.g., Belzil and Hansen 2004; Belzil and Leonardi 2007; Chen 2008; Shaw 1996). Theoretical predictions about the relationship between risk preferences and educational choice are less clear cut as education may be considered a risky investment (Levhari and Weiss 1974), but also shield against unemployment (Mincer 1991; Nickell and Bell 1996).
522 D.Pollmann et al. 1 3 Our main contribution to this strand of the literature is the introduction of a robust measure of occupation-specific earnings risk that does not rely on parametric assumptions for the error terms and yields consistent estimates for occupation-specific wage dispersion even if homoskedasticity and normality assumptions are violated. Moreover, the earnings risk measure we propose in the paper can even be estimated in the presence of censored wage information. This can be of great advantage in empirical work as administrative wage data are often top-coded. Importantly, Monte Carlo simulations show that our method for the estimation of wage dispersion is particularly effective compared to conventional approaches. In our application, we find that individuals with greater stated willingness to take risks work in occupations with higher cross-sectional wage dispersion. After estimating risk profiles of occupations on the IABS data, we match them to individuals in the SOEP working in these occupations. The SOEP provides us with survey information on risk preferences and other individual and household characteristics. The IABS on the other hand offers great sample size, such that we are able to work with more precise occupation definitions than previous studies and reduce the effect of aggregation on variation. The econometric approach we propose can be utilized in any setting where the researcher or analyst is interested in the dispersion of an outcome variable that is censored. An additional example is a demand planner who needs to balance the cost of being under- or oversupplied and therefore has to model the distribution of demand. Historical demand may only be observed censored because the number of units that can be sold is limited by the amount of inventory. The organization of the paper is as follows. In Sect.2, we briefly discuss QR and present our method for estimating dispersion in more detail. In addition, we describe a particularly useful estimation algorithm for censored data, the 3-step censored quantile regression (CQR) estimator by Chernozhukov and Hong (2002), which we use in our application on risk preferences and occupational sorting in Sect.3. Section4 concludes. 2 Estimation ofGroup‑Level Dispersion Our method for the estimation of dispersion is not based on residuals, but rather on the difference of coefficient estimates at particular quantiles. As such, it is in the spirit of the heteroskedasticity test of Koenker and Bassett (1982), which carries out a Wald test on the differences of coefficient estimates at different quantiles. Specifically, we first estimate the model by (C)QR at different quantiles, such as the 10th, 25th, 50th, 75th, and 90th percentile, including dummy variables for the groups which are to be compared. In our application, for instance, we include dummies for all occupations. We then consider the differences of the coefficient estimates for these dummies at two particular quantiles, such as the 10th and 90th percentile (“10–90 spread”), and compare their values across occupations. Our approach is not only computationally simple, but it also controls for the dispersion effect of
523 1 3 Robust Estimation ofWage Dispersion withCensored Data: An… covariates, and thereby filters out the (possibly) heteroskedastic effect of, for example, education and tenure in our application. To introduce notation and build intuition, we briefly summarize quantile regression in Sect.2.1 before introducing our dispersion measure in Sect.2.2. Section2.3 discusses a particularly simple estimator for censored quantile regression used in our application. 2.1 Quantile Regression Quantile regression (QR), introduced by Koenker and Bassett (1978) as a generalization of median regression, allows us to parsimoniously describe the entire conditional wage distribution by estimating conditional quantile functions (CQF) Q𝜏(Yi|Xi) .3 We denote the conditional 𝜏 -quantile of Y given X as For a linear quantile model, q𝜏(Xi)=X� i𝛽(𝜏) , and we can write Yi as While specifying a parametric model of the conditional quantiles, we are agnostic about the error distribution in our semiparametric framework. All we rely on is a conditional quantile restriction, stipulating that the conditional quantile of the error is equal to a constant. We assume that Xi always includes a constant term or full set of dummies, which affords us the following normalization: Estimation typically proceeds by characterizing conditional quantiles as the solution to a particular expected loss minimization problem, in the context of which it is useful to define the “check” (or weighted absolute loss) function. For 𝜏∈(0, 1) , In the case of our linear quantile model, q𝜏( X i)= X � i𝛽(𝜏) , We can define the QR estimator as its sample equivalent and the optimal predictor minimizing the realized loss: (1) Q 𝜏(Yi | Xi)≡q𝜏(Xi)≡F −1 Y i| X i (𝜏)=inf r∈ℝ{ r∶FYi | Xi(r)>𝜏 }. (2) Yi =X � i 𝛽(𝜏)+u i (𝜏) . (3) Q𝜏(ui(𝜏)|Xi)=0. (4) 𝜌𝜏(u) ≡ 𝜏𝟏(u ≥ 0)u+(1−𝜏)𝟏(u<0)(−u) (5) =[𝜏−𝟏(u<0)]u. (6) 𝛽 (𝜏)=arg min b 𝔼 [ 𝜌𝜏 ( Yi−X � ib )| Xi ]. 3 An excellent non-technical introduction with illustrative examples and an overview of applications can be found in Koenker and Hallock (2001). Buchinsky (1998) summarizes a range of points relevant to the empirical researcher.
524 D.Pollmann et al. 1 3 Asymptotic normality and consistency of the QR estimator can be shown (Bassett and Koenker 1978). 2.2 Differences ofQuantile Coefficients To construct our dispersion statistics, we estimate linear quantile models with occupation dummies at different quantiles. For consistency, we require that the number of observations per occupation group grows large. To make the treatment of the occupation dummies more explicit, we write Xi =( X � i , X � i ) � , where Xi includes all regressors but the occupation dummies, which are stacked in a separate vector Xi s.t. Xij = 1 if individual i works in occupation j and 0 otherwise: To build intuition for how a statistic like 𝜂 j(𝜏1)− 𝜂 j(𝜏2) measures relative dispersion in occupation j, we consider the illustrative yet likely simplistic special case of a location-scale model. Suppose Yi is dependent on Xi and Xi in mean and through a re-scaling of variances, where 𝜖i∣Xi iid ∼F𝜖( ⋅ ) for some distribution function F𝜖 s.t. 𝔼[𝜖i]=0 : In this model, the dummies in Xi have a location effect through 𝛼 and a scale effect through 𝜁 . For the regression specification (8), 𝛽 (𝜏) p �����→ 𝛼 +F−1 𝜖 (𝜏) 𝜁 and 𝜂 (𝜏) p �����→ 𝛼 +F−1 𝜖 (𝜏) 𝜁 , respectively. Therefore, for any occupation j and two different quantiles 𝜏1 and 𝜏2 , 𝜂 j (𝜏 1 )− 𝜂 j (𝜏 2 ) p �����→ 𝜁 j [F−1 𝜖 (𝜏 1 )−F−1 𝜖 (𝜏 2)] , and we can consistently estimate each occupation’s scale effect up to a multiplying constant. Chamberlain (1994, p. 186) starts by discussing comparable normal-location models but considers them inadequate for characterizing the conditional wage distribution. In particular, they imply constant covariate slopes across the quantiles, which are at odds with the quantile patterns of industry wage effects he finds. Moreover, and closest to our application, Chamberlain presents differential patterns across industries and relates these to industry-specific residual dispersion. Close inspection of the industry coefficients in Chamberlain (1994, Table5.4) at different quantiles reveals that even the location-scale model may be too restrictive. A statistical test of the location-scale hypothesis can be based on a Khmaladze transformation, but is only available for uncensored data. Applying a human capital model including occupation dummies to self-reported earnings in the SOEP, the Khmaladze test (Koenker and Xiao 2002) rejects the location-scale hypothesis at the 1% level. The parametric heteroskedasticity of the location-scale model implies the same relative dispersion pattern across occupations regardless of what quantiles we use for our difference metric, and it thereby rules out differential tail behavior in occupations. (7) 𝛽 (𝜏)=arg min b N ∑ i=1 𝜌𝜏(Yi−X� ib) . (8) q𝜏 (X i )= X� i 𝛽(𝜏)+ X i � 𝜂 (𝜏) . (9) Yi = X � i 𝛼 + X � i 𝛼 + ( X � i 𝜁+ X � i 𝜁 ) 𝜖 i.
525 1 3 Robust Estimation ofWage Dispersion withCensored Data: An… Our dispersion statistic does not require a parametric location-scale model. For example, it can accommodate heterogeneous shock distributions across occupations. In this case, estimates at different quantiles will result in different estimates of relative dispersion across occupations. More generally, (8) measures the effect of our set of occupation groups on all different quantiles while controlling for the dispersion effect of the additional covariates. This gives us a measure of the conditional dispersion effect of each occupation. While the linear specification of the conditional quantiles may appear restrictive, a linear quantile model is frequently only intended as a reduced-form approximation, such as for the minimum distance (MD) estimators in Buchinsky (1994,p. 409) and Chamberlain (1994,p. 181).4 To evaluate the practical merits of the approach, we carry out a set of Monte Carlo simulations (“Monte Carlo Evidence on Estimation of Dispersion” section in the “Appendix”). We find that compared to conventional approaches, our method is particularly effective for the estimation of dispersion in the presence of interaction effects in variance. For a moderate degree of censoring of 10%, very similar to that in our application, the 10–90 spread obtained from CQR does just as well as when we leave the data uncensored. 2.3 Censored Quantile Regression A particular feature of conditional quantiles not shared by conditional expectations is equivariance to monotone transformations. For any non-decreasing function g( ⋅ p) , As a result, QR is particularly suited for censoring problems. In addition, it does not require the restrictive assumptions of parametric censored estimators. In the case of top-coding, we observe Yi=min(Ci,Y∗ i) , where Y∗ i is the latent true value of the process of interest and Ci is some observed upper limit, of which we assume Y∗ i to be independent conditional on Xi . Since for any Ci∈ℝ , min(Ci, ⋅ p) is a non-decreasing function, we have (Powell 1986): The censored quantile regression (CQR) estimator follows trivially as the minimizing argument of the Powell objective function (Powell 1986): (10) Q𝜏[g(Yi)|Xi]=g[Q𝜏(Yi|Xi)]. (11) Q𝜏( Y ∗ i| X i) =X � i 𝛽(𝜏)⇒Q 𝜏 (Y i| X i )=min [ C i ,X � i 𝛽(𝜏) ]. 4 Formally, Chamberlain (1994,p. 181) recognized that the QR estimator provides a linear approximation to the CQF, albeit of a less “transparent” nature than in the OLS and MD case. Angrist etal. (2006) show that QR minimizes a weighted mean-squared error loss function for specification error, implicitly providing a weighted MD approximation to the true nonlinear CQF. Applying the framework to wage regressions with a focus on the education variable, they find QR to provide a useful approximation to the conditional wage distribution.
526 D.Pollmann et al. 1 3 As in the uncensored case, the QR-based estimator is consistent under general nonnormal distributions and heteroskedasticity (Powell 1984, 1986). Buchinsky (1994) gives a well-known application to changes in the US wage structure. For estimation, he presents his iterative linear programming algorithm (ILPA), which iteratively performs QR on observations with predictions in the uncensored region, based on the previous iteration. Convergence is achieved if two subsequent sets of observations are the same; while this need not occur, convergence guarantees local optimality. Another alternative, the BRCENS algorithm, is proposed in Fitzenberger (1997). Unfortunately, both algorithms have less than reliable convergence properties with respect to the Powell estimator (12), particularly in large samples and for high dimensionality, as in our application. Chernozhukov and Hong (2002) present the three-step CQR method, which avoids a great deal of problems by selecting a more “benign” sample based on an initial regression of the probability of censoring, and subsequently works with standard QR.5,6 Step 1 Let 𝜂i=𝟏(Yi≠Ci) ; that is, 𝜂i is an indicator of non-censoring (with censoring point Ci ). We estimate a parametric (e.g., probit or logit) model for the probability of non-censoring: Here, Xi is a vector of suitable transformations of ( X � i ,C i)� . In general, model (13) will be misspecified and any corresponding estimators such as MLE will therefore be inconsistent for the true propensity score h(Xi,Ci) . However, it is only used as an auxiliary regression to select an initial sample J0 with propensity score h(Xi,Ci)>𝜏 , necessary for consistent estimation of quantile 𝜏 .7 To ensure this, we do not base our selection on the condition that p( X� 𝛾 )>𝜏 , but rather that p( X � i 𝛾 )>𝜏+ k , where k is a trimming constant strictly between 0 and 1−𝜏 . Since we do not necessarily have to select the largest subset J0 , there is some freedom in choosing k. For this, we write J0 as a function of k, J0 (k)={i∶p( X � i 𝛾 )>𝜏+k } . The approach taken here, following Chernozhukov and Hong, is to choose the trimming constant k such that (12) 𝛽 CQR(𝜏)=arg min b N ∑ i=1 𝜌𝜏[Yi−min (Ci,X� ib)] . (13) � Pr (𝜂 i =1∣X i ,C i )=p ( X� i 𝛾 ). (14) #J0(k)∕#J0(0)=90%. 5 Applications include Melly (2005) on wage inequality, Kowalski (2009) on medical expenditure, and Schmillen and Möller (2012) on lifetime unemployment. 6 The estimators of Buchinsky and Hahn (1998) and Khan and Powell (2001) similarly carry out a firststage selection, but are impractical for high dimensionality and large data sets. 7 Note the deviation from Chernozhukov and Hong, who select a sample with propensity score h(Xi,Ci) > 1−𝜏 . This is an important difference between left- and right-censoring.
533 1 3 Robust Estimation ofWage Dispersion withCensored Data: An… Table 3 IABS risk profiles OLS estimates. Robust standard errors of coefficient estimates, allowing for clustering at the IABS KldB 88 occupation level, in parentheses; ***/**/*Indicate significance at 1%/5%/10% level. Dependent variable are the median differences of occupation dummy estimates from QR, on IABS KldB 88 level. “General Risk Attitude” is the response to the 2004 general risk question, “General Risk Attitude (av.)” is the corresponding average over the 2004, 2006, 2008, and 2009 responses Dependent variable: median differences of occupation effects 10–50 25–50 50–75 50–90 (1) (2) (3) (4) (5) (6) (7) (8) General Risk Attitude 0.002* (0.001) 0.000 (0.000) 0.001** (0.000) 0.002* (0.001) General Risk Attitude (av.) 0.003** (0.001) 0.001* (0.001) 0.002*** (0.001) 0.003** (0.001) Experience 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) −0.000 (0.000) −0.000 (0.000) −0.000 (0.000) −0.000 (0.000) Tenure −0.000 (0.000) −0.000 (0.000) 0.000 (0.000) 0.000 (0.000) 0.000** (0.000) 0.000** (0.000) 0.000* (0.000) 0.000* (0.000) Years of Education 0.005* (0.003) 0.005* (0.003) 0.002* (0.001) 0.002* (0.001) 0.003* (0.001) 0.003* (0.001) 0.005** (0.002) 0.005** (0.002) Married and living together −0.009** (0.004) −0.009** (0.004) −0.003 (0.002) −0.004 (0.002) −0.005* (0.003) −0.005* (0.003) −0.008** (0.003) −0.008** (0.003) Body height 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) 0.000 (0.000) Public Sector Employment −0.000 (0.016) −0.000 (0.016) −0.009 (0.007) −0.009 (0.007) −0.001 (0.009) −0.001 (0.009) −0.016 (0.012) −0.016 (0.012) Median wage (occ.) −0.012 (0.048) −0.012 (0.048) −0.009 (0.022) −0.009 (0.022) −0.002 (0.022) −0.002 (0.022) −0.110*** (0.032) −0.110*** (0.032) Constant −0.114 (0.209) −0.123 (0.208) −0.046 (0.096) −0.049 (0.095) −0.109 (0.084) −0.112 (0.083) 0.278** (0.122) 0.273** (0.121) Observations 2740 2747 2740 2747 2740 2747 2740 2747 R-squared 0.025 0.027 0.020 0.022 0.032 0.034 0.133 0.135
534 D.Pollmann et al. 1 3 preferences are rather stable is accumulating. Sahm (2008), for example, shows that risk preferences change only gradually with age but are rank-order stable. Changes in macroeconomic conditions have an impact on measured risk tolerance, but changes in income, wealth or other major events that reduce expected lifetime wealth, such as job displacement or a deterioration in health, do not affect individuals’ willingness to take risk. Dohmen et al. (2007) analyze the stability of responses to the general risk question in the SOEP. For two subject pools, one a subset of the SOEP, the other a separate one, they find a test–retest correlation of 0.62 and 0.60, respectively, over a 6-week horizon. It is plausible to assume that risk preferences do not change dramatically over such a short time period so that the variation in answers in the test–retest samples can be attributed to measurement error. The correlation between the 2004 and 2006 waves of the SOEP, in comparison, is 0.50, which is not too far below the 6-week benchmark; this suggests that risk attitudes constitute an inherent and stable trait. Beauchamp et al. (2017) support this interpretation, as they find very similar results for Swedish data using the same risk measure as is used in the SOEP. In our setting, a sorting interpretation also requires that the ranking of occupations with respect to their occupational earnings risk has remained stable. Otherwise, the risk profile estimated on the 2004 cross section may not have been relevant at the time when individuals chose their occupation. In an extreme case, risk attitudes might not have been related to differences in occupation-spe- cific earnings risk when individuals sorted into an occupation. Instead, a wage setting mechanism in which preferences of incumbents shape the occupational earnings risk might be a potential channel through which a correlation between risk preferences and wage dispersion can arise. To address the question whether there have been considerable changes in occupation-specific wage dispersion, we estimate wage dispersion measures for the years 1979, 1984, 1989, 1994, and 1999, and compute Pearson correlation coefficients with occupations as crosssectional unit (Table 4). The correlation coefficient is decreasing in the time span considered, but remains positive and high; it is larger than 0.65 for any pair of years. This suggests that the relative wage dispersion of occupations has been rather stable in West Germany in the period from 1979 to 2004, and that the risk profiles we estimate from a cross section for 2004 are quite close to those relevant at the point of labor market entry for most of the individuals in the SOEP. Table 4 Temporal stability of risk profiles Pearson correlation coefficients of 10–90 spread per occupation across years 1979 1984 1989 1994 1999 2004 1979 1.000 1984 0.818 1.000 1989 0.768 0.902 1.000 1994 0.781 0.854 0.874 1.000 1999 0.654 0.778 0.810 0.899 1.000 2004 0.655 0.694 0.753 0.896 0.922 1.000
535 1 3 Robust Estimation ofWage Dispersion withCensored Data: An… Finally, we cannot rule out that the correlation between risk attitudes and wage dispersion is driven by cognitive abilities rather than risk preferences: There is evidence for a negative relationship between risk aversion and cognitive abilities (e.g., Dohmen etal. 2010), and at the same time, dispersion may be particularly attractive for high-ability individuals. 4 Conclusion We discuss a particular method to estimate group-level wage dispersion, which is based on semiparametric methods. Specifically, we estimate a human capital model, including dummy variables for each of the groups of interest, at a number of different quantiles; we then take the differences of the dummy coefficients at different quantiles as a measure of dispersion within each group. The method is particularly useful when working with data which is either censored or top-coded, such as administrative data and some survey data, since it is more robust to deviations from homoskedasticity and distributional assumptions than parametric estimators. In addition, it controls for the dispersion effect of covariates, and allows us to estimate the entire conditional wage distribution and its differences across groups. In an application which connects a large German administrative data set, the IAB Employment Sample (IABS), which is subject to censoring due to a legislative contribution limit, and a household survey, we find that individuals with greater willingness to take risks work in occupations with higher cross-sectional wage dispersion. Acknowledgements We thank Denis de Crombrugghe and an anonymous referee for valuable comments. Thomas Dohmen gratefully acknowledges funding from the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) through CRC TR 224 (Project A01) and Germany’s Excellence Strategy - EXC 2126/1- 390838866. Funding Open Access funding provided by Projekt DEAL. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creat iveco mmons .org/licen ses/by/4.0/. Appendix: Monte Carlo Evidence onEstimation ofDispersion In this section, we review the performance of the estimation method described in Sect.2.2 for both censored and uncensored data and compare it to a residual-based method. For uncensored and censored data, we use the difference between the coefficient estimates of the group dummies at the 90th and 10th percentile from (C)QR.
536 D.Pollmann et al. 1 3 For uncensored data only, we estimate a conventional OLS regression including group dummies and compute the standard deviation of residuals per group. After each of 1000 simulations, we compute the correlation of an occupation-specific scale 𝜎j and the three statistics. The models investigated are stylized versions of the wage distribution setting in our empirical analysis; specifically, we first consider a model with only group-specific scale, and then turn to location-scale models in which a regressor has a heteroskedastic effect. The censored data is derived directly from the uncensored data through right-censoring at the 90th percentile such that in each case, 10% of the data are censored, which is intended to resemble the degree of censoring in the IABS data used in our application. Group‑Specific Scale Model The DGP has the following linear representation: for i∈{1, …,N} and j∈{1, …,M} , cj iid ∼N(0, 𝜔2 ) , 𝜖i iid ∼N(0, 1 ) . In our simulation, we set N=1000 , X i iid ∼U(0, 1 ) , 𝛼=2 , 𝜔=0.2 , and 𝜎j=0.1 +0.2Uj , where Uj iid ∼U(0, 1 ) . Individuals i are randomly assigned to one of M=10 groups j(i) according to a uniform distribution. Table5 shows a very similar performance for all three statistics, with a correlation close to unity in each case. Linear Location‑Scale Model Leaving all else the same, (16) Yi =𝛼 X i +c j(i) +𝜎 j(i) 𝜖 i (17) Yi =𝛼 X i +c j(i) + ( 𝛿X i +𝜎 j(i)) 𝜖 i Table 5 Group-specific scale model Correlation of risk measures with true dispersion Mean Std. Dev. Min. Max. 10–90 spread 0.975 0.021 0.794 0.999 10–90 spread (cens.) 0.968 0.026 0.729 0.999 Resid. std. dev. 0.984 0.013 0.881 0.999 Table 6 Linear location-scale model Correlation of risk measures with true dispersion Mean Std. Dev. Min. Max. 10–90 spread 0.899 0.072 0.228 0.993 10–90 spread (cens.) 0.891 0.076 0.198 0.994 Resid. std. dev. 0.852 0.107 0.159 0.988
537 1 3 Robust Estimation ofWage Dispersion withCensored Data: An… with 𝛿=0.5 . Hence, the independent variable X now exerts a heteroskedastic effect. Already, the 10–90 spread does slightly better (Table6) for the uncensored data. Notably, it also works just as well when only censored data is available. Nonlinear Location‑Scale Model We adapt the DGP such that the scale effect of the independent variable X is now negatively related to the occupation variance: For 𝛿=1 , As reported in Table7, the statistics based on QR are a lot more robust in this case, since the scale effect of X at the different quantiles is explicitly controlled for. The discrepancy will likely be even larger for more irregular distributions. Also, our method for dispersion estimation works equally well for censored data. References Amemiya, T. (1973). Regression analysis when the dependent variable is truncated normal. Econometrica, 41(6), 997–1016. Angrist, J., Chernozhukov, V., & Fernández-Val, I. (2006). Quantile regression under misspecification, with an application to the U.S. wage structure. Econometrica, 74(2), 539–563. Arabmazar, A., & Schmidt, P. (1981). Further evidence on the robustness of the Tobit estimator to heteroskedasticity. Journal of Econometrics, 17(2), 253–258. Arabmazar, A., & Schmidt, P. (1982). An investigation of the robustness of the Tobit estimator to nonnormality. Econometrica, 50(4), 1055–1063. Bassett, G, Jr., & Koenker, R. (1978). Asymptotic theory of least absolute error regression. Journal of the American Statistical Association, 73(363), 618–622. Beauchamp, J., Cesarini, D., & Johannesson, M. (2017). The psychometric properties of measures of economic risk preferences. Journal of Risk and Uncertainty, 54(3), 203–237. Bellante, D., & Link, A. N. (1981). Are public sector workers more risk averse than private sector workers? Industrial and Labor Relations Review, 34(3), 408–412. Belzil, C., & Hansen, J. (2004). Earnings dispersion, risk aversion and education. In S. W. Polachek (Ed.), Accounting for worker well-being, volume 23 of research in labor economics (pp. 335–358). Bingley: Emerald Group Publishing Limited. (18) Yi =𝛼 X i +c j(i) + [( 1−𝛿𝜎 j(i)) X i +𝜎 j(i)] 𝜖 i. (19) Yi =𝛼 X i +c j(i) + [ X i +(1− X i )𝜎 j(i)] 𝜖 i. Table 7 Nonlinear locationscale model Correlation of risk measures with true dispersion Mean Std. Dev. Min. Max. 10–90 spread 0.819 0.112 0.217 0.991 10–90 spread (cens.) 0.837 0.105 0.262 0.992 Resid. std. dev. 0.682 0.191 −0.347 0.973
538 D.Pollmann et al. 1 3 Belzil, C., & Leonardi, M. (2007). Can risk aversion explain schooling attainments? Evidence from Italy. Labour Economics, 14(6), 957–970. Bonin, H., Dohmen, T., Falk, A., Huffman, D., & Sunde, U. (2007). Cross-sectional earnings risk and occupational sorting: The role of risk attitudes. Labour Economics, 14(6), 926–937. Brown, C., & Moffitt, R. (1983). The effect of ignoring heteroscedasticity on estimates of the Tobit model. NBER Technical Working Papers 0027, National Bureau of Economic Research, Inc. Buchinsky, M. (1994). Changes in the U.S. wage structure 1963–1987: Application of quantile regression. Econometrica, 62(2), 405–458. Buchinsky, M. (1998). Recent advances in quantile regression models: A practical guideline for empirical research. Journal of Human Resources, 33(1), 88–126. Buchinsky, M., & Hahn, J. (1998). An alternative estimator for the censored quantile regression model. Econometrica, 66(3), 653–671. Caliendo, M., Fossen, F., & Kritikos, A. (2009). Risk attitudes of nascent entrepreneurs: New evidence from an experimentally validated survey. Small Business Economics, 32(2), 153–167. Chamberlain, G. (1994). Quantile regression, censoring, and the structuring of wages. In C. A. Sims (Ed.), Advances in econometrics: Sixth World Congress (chapter5) (Vol.1, pp. 171–209). Cambridge: Cambridge University Press. Chay, K. Y., & Honoré, B. E. (1998). Estimation of semiparametric censored regression models: An application to changes in black-white earnings inequality during the 1960s. Journal of Human Resources, 33(1), 4–38. Chen, S. H. (2008). Estimating the variance of wages in the presence of selection and unobserved heterogeneity. The Review of Economics and Statistics, 90(2), 275–289. Chernozhukov, V., Fernández-Val, I., Han, S., & Kowalski, A. E. (2011). Stata command to implement CQIV. Retrieved Jan 30, 2012 from https ://sites .lsa.umich .edu/amand a-kowal ski/stata -comma nds/. Chernozhukov, V., & Hong, H. (2002). Three-step censored quantile regression and extramarital affairs. Journal of the American Statistical Association, 97(459), 872–882. Cramer, J. S., Hartog, J., Jonker, N., & van Praag, C. M. (2002). Low risk aversion encourages the choice for entrepreneurship: An empirical test of a truism. Journal of Economic Behavior & Organization, 48(1), 29–36. Dohmen, T., & Falk, A. (2010). You get what you pay for: Incentives and selection in the education system. Economic Journal, 120(546), 256–271. Dohmen, T., Falk, A., Huffman, D., & Sunde, U. (2010). Are risk aversion and impatience related to cognitive ability? American Economic Review, 100(3), 1238–1260. Dohmen, T., Falk, A., Huffman, D., Sunde, U., Schupp, J., & Wagner, G. G. (2007). The measurement and stability of risk attitudes. New York: Mimeo. Dohmen, T., Falk, A., Huffman, D., Sunde, U., Schupp, J., & Wagner, G. G. (2011). Individual risk attitudes: Measurement, determinants and behavioral consequences. Journal of the European Economic Association, 9(3), 522–550. Donald, S. G. (1995). Two-step estimation of heteroskedastic sample selection models. Journal of Econometrics, 65(2), 347–380. Drews, N. (2008). Das Regionalfile der IAB-Beschäftigtenstichprobe 1975–2004: Handbuch-Version 1.0.3. FDZ Datenreport 02/2008(DE), Institut für Arbeitsmarkt- und Berufsforschung (IAB), Nuremberg, Germany. Retrieved Sept 15, 2020 from http://doku.iab.de/fdz/repor te/2008/DR_02-08. pdf. Dustmann, C., Ludsteck, J., & Schönberg, U. (2009). Revisiting the German wage structure. Quarterly Journal of Economics, 124(2), 843–881. Ekelund, J., Johansson, E., Jarvelin, M.-R., & Lichtermann, D. (2005). Self-employment and risk aversion—Evidence from psychological test data. Labour Economics, 12(5), 649–659. Fitzenberger, B. (1997). Computational aspects of censored quantile regression. Lecture Notes-Mono- graph Series, 31, 171–186. Fitzenberger, B., Osikominu, A., & Völter, R. (2006). Imputation rules to improve the education variable in the IAB employment subsample. Schmollers Jahrbuch, 126(3), 405–436. Fouarge, D., Kriechel, B., & Dohmen, T. (2014). Occupational sorting of school graduates: The role of economic preferences. Journal of Economic Behavior & Organization, 106, 335–351. Fuchs-Schündeln, N., & Schündeln, M. (2005). Precautionary savings and self-selection: Evidence from the German reunification experiment. Quarterly Journal of Economics, 120(3), 1085–1120.
539 1 3 Robust Estimation ofWage Dispersion withCensored Data: An… Goldberger, A. (1983). Abnormal selection bias. In S. Karlin, T. Amemiya, & L. Goodman (Eds.), Studies in econometrics, time series, and multivariate statistics (pp. 67–85). New York, NY: Academic Press. Guiso, L., & Paiella, M. (2005). The role of risk aversion in predicting individual behavior. Temi di discussione (Economic working papers) 546, Bank of Italy, Economic Research Department. Heckman, J. J. (1976). The common structure of statistical models of truncation, sample selection and limited dependent variables and a simple estimator for such models. In Annals of economic and social measurement, volume5 of NBER chapters (pp. 475–492). Cambridge: National Bureau of Economic Research, Inc. Heckman, J. J. (1979). Sample selection bias as a specification error. Econometrica, 47(1), 153–161. Hu, L. (2002). Estimation of a censored dynamic panel data model. Econometrica, 70(6), 2499–2517. Hurd, M. (1979). Estimation in truncated samples when there is heteroscedasticity. Journal of Econometrics, 11(2–3), 247–258. Juhn, C., Murphy, K. M., & Pierce, B. (1993). Wage inequality and the rise in returns to skill. Journal of Political Economy, 101(3), 410–442. Katz, L. F., & Murphy, K. M. (1992). Changes in relative wages, 1963–1987: Supply and demand factors. The Quarterly Journal of Economics, 107(1), 35–78. Khan, S., & Powell, J. L. (2001). Two-step estimation of semiparametric censored regression models. Journal of Econometrics, 103(1–2), 73–110. King, A. G. (1974). Occupational choice, risk aversion, and wealth. Industrial and Labor Relations Review, 27(4), 586–596. Koenker, R., & Bassett, G, Jr. (1978). Regression quantiles. Econometrica, 46(1), 33–50. Koenker, R., & Bassett, G, Jr. (1982). Robust tests for heteroscedasticity based on regression quantiles. Econometrica, 50(1), 43–61. Koenker, R., & Hallock, K. F. (2001). Quantile regression. Journal of Economic Perspectives, 15(4), 143–156. Koenker, R., & Xiao, Z. (2002). Inference on the quantile regression process. Econometrica, 70(4), 1583–1612. Kowalski, A. E. (2009). Censored quantile instrumental variable estimates of the price elasticity of expenditure on medical care. NBER Working Papers 15085, National Bureau of Economic Research, Inc. Levhari, D., & Weiss, Y. (1974). The effect of risk on the investment in human capital. American Economic Review, 64(6), 950–963. Maddala, G. S., & Nelson, F. D. (1975). Specification errors in limited dependent variable models. NBER Working Papers 0096, National Bureau of Economic Research, Inc. Melly, B. (2005). Decomposition of differences in distribution using quantile regression. Labour Economics, 12(4), 577–590. Mincer, J. A. (1991). Education and unemployment. NBER Working Papers 3838, National Bureau of Economic Research, Inc. Nickell, S., & Bell, B. (1996). Changes in the distribution of wages and unemployment in OECD countries. The American Economic Review, 86(2), 302–308. Papers and proceedings of the hundredth and eighth annual meeting of the American Economic Association, San Francisco, CA, January 5–7, 1996. Paarsch, H. J. (1984). A Monte Carlo comparison of estimators for censored regression models. Journal of Econometrics, 24(1–2), 197–213. Powell, J. L. (1984). Least absolute deviations estimation for the censored regression model. Journal of Econometrics, 25(3), 303–325. Powell, J. L. (1986). Censored regression quantiles. Journal of Econometrics, 32(1), 143–155. Sahm, C. R. (2008). How much does risk tolerance change? Finance and Economics Discussion Series 2007-66, Board of Governors of the Federal Reserve System (U.S.). Saks, R. E., & Shore, S. H. (2005). Risk and career choice. The B.E. Journal of Economic Analysis & Policy, 5(1), 1–43. Schmillen, A., & Möller, J. (2012). Distribution and determinants of lifetime unemployment. Labour Economics, 19(1), 33–47. Schulhofer-Wohl, S. (2011). Heterogeneity and tests of risk sharing. Journal of Political Economy, 119(5), 925–958. Shaw, K. L. (1996). An empirical analysis of risk aversion and income growth. Journal of Labor Economics, 14(4), 626–653.
540 D.Pollmann et al. 1 3 Skeels, C. L., & Vella, F. (1999). A Monte Carlo investigation of the sampling behavior of conditional moment tests in Tobit and Probit models. Journal of Econometrics, 92(2), 275–294. Tobin, J. (1958). Estimation of relationships for limited dependent variables. Econometrica, 26(1), 24–36. Vijverberg, W. P. M. (1987). Non-normality as distributional misspecification in single-equation limited dependent variable models. Oxford Bulletin of Economics and Statistics, 49(4), 417–430. Wagner, G. G., Frick, J. R., & Schupp, J. (2007). The German Socio-Economic Panel Study (SOEP)— Scope evolution and enhancements. Schmollers Jahrbuch, 127(1), 139–169. Wooldridge, J. M. (2010). Econometric analysis of cross section and panel data. Number 0262232588 in MIT Press Books (2nd edn.). The MIT Press. Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.