Propensity score in the tails and returns to education in Italy
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Furno, Marilena; Caracciolo, Francesco Article Propensity score in the tails and returns to education in Italy Economies Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Furno, Marilena; Caracciolo, Francesco (2025) : Propensity score in the tails and returns to education in Italy, Economies, ISSN 2227-7099, MDPI, Basel, Vol. 13, Iss. 2, pp. 1-28, https://doi.org/10.3390/economies13020050 This Version is available at: https://hdl.handle.net/10419/329330 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Academic Editor: Ralf Fendel Received: 26 November 2024 Revised: 7 January 2025 Accepted: 21 January 2025 Published: 13 February 2025 Citation: Furno, M., & Caracciolo, F. (2025). Propensity Score in the Tails and Returns to Education in Italy. Economies,13(2), 50. https://doi.org/ 10.3390/economies13020050 Copyright: © 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/ licenses/by/4.0/). Article Propensity Score in the Tails and Returns to Education in Italy Marilena Furno * and Francesco Caracciolo Department of Agricultural Sciences, Universitàdegli Studi di Napoli Federico II, 80055 Napoli, Italy; [email protected] *Correspondence: [email protected] Abstract: The propensity score defining the probability of completing a given degree of education—to balance covariates—and the Mincer equation is here estimated at various degrees of higher education. The novelty is in implementing propensity score and regression estimators together in a double-robust approach in order to ensure against misspecification. The model is analyzed not only at the average but also in the tails of both components to gain a detailed analysis of the tail behavior and robustness. Analyzing survey data from the 2010 and 2020 waves, we find a negative impact of southern regions and gender on education. This impact becomes milder at the mean and is not significant in the right tail. The mixing of propensity score and quantile regression shows the irrelevance of education at low wages and, in a few cases, decreasing premia as school years increase. The private sector rewards lower premiums to young workers, and these distributions are more dispersed, i.e., show higher inequality. In the women’s subset, there is a marked pay gap, even wider for those working in the private sector. Keywords: double–robust; propensity score; quantile regression JEL Classification: C20; E24 1. Introduction In this paper, we focus on returns on education in Italy in the last decade, analyzing the most recent wave of the Banca d’Italia Survey of Household Income and Wealth, SHIW. Returns may differ due to educational attainment but also due to gender and to regional economic divide—with northern regions having better opportunities and draining skilled workers from the other regions. Many methods have been implemented to measure returns on education, and this analysis proposes another approach. Here, the propensity score and the regression model are coupled in a double-robust (DR) approach, and we implement it not only at the center/average but also in the tails of both components. The DR estimator (Robins et al., 1994;Lunceford & Davidian,2004) is generally implemented to compute average treatment effects. It combines propensity score (Rosenbaum & Rubin,1983) and OLS-fitted values to guarantee against misspecification. In this study, we implement DR to evaluate returns on education both on average and in the tails to gain a more detailed analysis of the tail behavior together with more robust results. In DR, the propensity score estimator (PS) computes the probability of each observation being treated conditionally on the observed covariates. Weighting the observations in each group by the inverse probability of being in that group defines the potential outcome. The estimated treatment effect is given by the comparison of the potential outcomes of treatment and control, and it is generally implemented on average. It aims to control for Economies 2025,13, 50 https://doi.org/10.3390/economies13020050
Economies 2025,13, 50 2 of 28 the potential bias occurring when treated individuals systematically differ from those that are untreated by balancing the covariates of the two groups. 1 The other component of the DR approach is OLS regression, which relates individuals’ earnings to their degree of education, controlling for several other factors such as age, gender, region of residence, and field of education, and which provides the fitted values of the outcome in each group as computed at the conditional mean. The combination of these two estimators is very helpful since pooling PS and OLS together yields a consistent estimator if at least one of the two is correctly specified (Neugebauer & van der Laan,2005). The comparison of potential outcome and OLS-fitted values for treated and untreated individuals computes the average treatment effect. However, the impact of treatment on the outcome distribution is not necessarily constant, as in the case of heterogeneity. Therefore, we focus on the treated/untreated difference, not only at the mean but also in the tails, at various locations. Caracciolo and Furno (2017) proposed to analyze a binary treatment effect at the quantiles by introducing a quantile regression estimator in place of the OLS results. The impact on the tails may differ from the treatment estimated on the average: additional schooling may grant higher earnings in the upper tail, for more qualified jobs, compared to its impact at average or lower incomes. It may even be the case that the effect of treatment is opposite in the tails, negative at the lower and positive at the upper quantiles, thus balancing on average. Replacing the OLS-estimated outcomes in each group with quantile regression estimates allows for computing the treatment effect in the tails of the outcome distribution while also granting greater robustness. Indeed, quantile regressions are less affected by anomalous values than OLS. Next, Furno and Caracciolo (2020) extended this approach to the multivariate setting, in case of more than one treatment option, such as differing lengths in training programs/schooling, different policy interventions, or diverse drugs/doses in clinical trials. In these two works, the focus is on moving the estimated regression away from the mean, providing fitted values of the outcome at the quantiles, while keeping the PS estimates constant at the conditional mean. This approach excludes by assumption heterogeneity in the probability of treatment. However, the probability may change as well. For instance, the probability of highly educated workers being employed is lower in the left tail for less-qualified jobs and higher in the right tail. The systematic discrepancy between treated and untreated individuals is not necessarily constant, and a non-constant balancing probability is called for. In what follows, we examine the odds of being treated in the tails and compute the probability of being in one group or another in the tails, moving both propensity score and regression estimates away from the mean. We couple propensity scores in the tails and quantile regression estimates. This provides a tail estimator of the treatment effect with respect to both components of the double-robust, the propensity score, and the regression model. The first example considers only the propensity score and shows the presence of changing coefficients across different locations. We compute a logit model to define the probability of being employed using SHIW data from the 2010 and 2020 waves of the Banca d’Italia. The impact of the explanatory variables on the probability of employment does change across the selected locations. In our findings, education is positive at and above the median, the negative impact of southern regions does not disappear in the right tail, and gender is positive at the median and becomes non-significant at the upper quartile. In a second analysis, the double-robust approach is implemented to compute returns to education at the mean and in the tails, using SHIW data from the years 2010 and 2020. The 2024 OECD (OECD,2024) country note states that education levels in Italy have grown
Economies 2025,13, 50 3 of 28 slower than the OECD average. Italy remains one of the 12 OECD members where university degrees are not the most common education qualification for the 25–34 age bracket. This delay has long been known despite tertiary education warranting better employment, but in Italy, its reward is lower than elsewhere. Our estimates show the following: 1. There is little or no relevance of education at low wages. 2. Rewards in the private sector were higher than in the public sector for post-university education in 2020, with a more dispersed distribution—the private sector is characterized by greater wage inequality. 3. The educational premiums for younger generations are limited in both waves, except for the post-university premiums in 2020. This can be interpreted as an excess supply of educated workers or else as seniority, with older cohorts like baby boomers shifting uncertainty onto the younger generations; both interpretations agree with the OECD statement of low rewards to tertiary education in Italy. 4. The women’s subset results are characterized by a gender pay gap, which is wider in the private sector. 2. Previous Results in Analyzing SHIW Data In this section, while recognizing the existence and relevance of various data sets provided by the Italian Statistical Agency (ISTAT), the Italian Research Institute on Labor (ISFOL), the European Union, and the Italian Social Security Institute (INPS), we focus on works analyzing SHIW data for the sake of comparability. Several empirical studies analyze SHIW data to estimate the returns to schooling and the gender wage gap in Italy. Their results do not provide unequivocal measures of the returns to education and the gender wage gap. Cannari and D’Alessio (1998) use instrumental variables for the 1993 wave and, by choosing family background variables as instruments, obtain an estimate of the returns to education close to 7%. Colussi (1997) uses the same sample and similar instrumental variables and provides an estimate of 6.6%, which is not significantly different. Flabbi (1999) analyzes the 1991 SHIW wave to compute the returns to schooling for women and men separately. He finds that the estimated coefficients obtained using instrumental variables are higher for men, 0.62 compared to 0.56 for women, while the OLS returns are 0.22 for women and 0.17 for men. Brunello and Miniaci (1999) use the 1993 and 1995 waves and select family background as instrumental variables. Their OLS estimate of male education returns is 4.8%, while the instrumental variable result is 5.7%. Conversely, Brunello et al. (2001), analyzing the 1984, 1989, and 1995 waves, find higher returns for women with both OLS and instrumental variable estimators, thus reversing the sign of the wage gap. Giustinelli (2004), using SHIW data from 1990 to 2000 also finds higher returns for women. However, these results are at odds with the general findings on the gender wage gap in the literature. In 2019, the European Commission ranked Italy among the European countries with the lowest gender pay gap, at less than 8%, a small but not negligible figure. Ciccone et al. (2006) analyzing the 1987, 1995, 1998, and 2000 waves find increasing returns at higher educational levels, with further improvements in the southern regions. Zizza (2013), analyzing data from 1995 to 2008, estimates a raw gender wage gap of around 6%, which increases to 11–12% in an extended version of the model. Unlike previous studies, our approach uses a double-robust method to compute returns to education, not only at the mean but also in the tails. This method uncovers subtler insights, such as the fading relevance of education at lower wage levels and the different rewards of post-university education between the private and public sectors in 2020, highlighting a pronounced gender pay gap, particularly in the private sector.
Economies 2025,13, 50 4 of 28 3. The Double Robust Approach In what follows, we describe the DR approach selected to measure returns to education in Italy. It merges propensity score and quantile regression. Central to this method is weighting observations by their probability of being educated/treated, calculated through a logistic model. The propensity score approach is further refined by implementing expectiles, which move the logistic regression away from the conditional mean, adding details on the tail behavior of the probability distribution. Meanwhile, in the regression model, we compute the unconditional distributions of more and less-educated individuals. Their difference, computed at various quantiles, provides a richer understanding of the returns to education at several points of the wage distribution. Consider the outcome variable Y i , which assumes the value Y i0 in the control group, if the treatment variable Z i is 0, and Y i1 when the treatment variable takes a unit value ( Zi= 1 for treatment). To measure the treatment effect, the standard double-robust approach combines the propensity score and regression approaches and is defined as follows: n−1∑n i=1(ZiYi P(Zi=1|X)−Zi−P(Zi=1|X) P(Zi=1|X)ˆ Yi1)−∑n i=1((1−Zi)Yi 1−P(Zi=1|X)+Zi−P(Zi=1|X) 1−P(Zi=1|X)ˆ Yi0)(1) The terms ˆ Yi1 and ˆ Yi0 represent the fitted values of the OLS regression computed in the treated and the control groups, respectively. The propensity score weights each observation by the probability of being treated. The probability weights, P(Z i = 1|X), with 0 < P(Zi= 1|X)<1 , are a function of the unobservable latent variable Z i , which is approximated by a set of observed covariates X. Consequently, the probability of treatment is estimated by assuming that P(Z i = 1|X) follows a parametric model, namely logistic regression: P(Zi=1|X) = exp(Xβ) {1+exp(Xβ)}(2) Equation (2) computes the probability of being highly educated as a function of X, which, in our model includes cohort, region of residence, gender, and field of study.2This probability provides the weights used to compute the potential outcome. The outcome Y i is weighted by the inverse of this probability, yielding the potential outcome for the treated, ZiYi P(Zi=1|X)=Yi P(Zi=1|X) , when Z i = 1, and the potential outcome for the untreated (1−Zi)Yi 1−P(Zi=1|X)=Yi 1−P(Zi=1|X)when Zi= 0. Equation (1) calculates the average difference between the treated and control groups. In the treatment group, when Z i = 1, it compares weighted observed and fitted values of the regression, respectively, Y i and ˆ Yi1 with weights equal to 1 P(Zi=1|X) and P(Zi=0|X) P(Zi=1|X) , respectively: n−1∑(Yi P(Zi=1|X)−1−P(Zi=1|X) P(Zi=1|X)ˆ Yi1)−∑ˆ Yi0for Zi=1 (3) In the control group, when Z i = 0, Y i is weighted by 1 P(Zi=0|X) and ˆ Yi0 is by P(Zi=1|X) P(Zi=0|X) , yielding the following: n−1∑ˆ Yi1 −∑(Yi 1−P(Zi=1|X)−P(Zi=1|X) 1−P(Zi=1|X)ˆ Yi0)for Zi=0 (4) The next step is to move beyond the average difference between the treated and control groups. To achieve this, we analyze the propensity score estimator and the regression at various locations, deviating from the conditional mean. We consider expectiles ( Newey & Powell,1987) to introduce a shifting weight that adjusts the estimated logistic
Economies 2025,13, 50 5 of 28 regression beyond the conditional mean. Equation (2) is modified to include an asymmetric weighting system, which shifts the equation up or down toward the tails: P(wiZi=1|X) = wi exp(Xβ) {1+exp(Xβ)}(5) where the asymmetric weight is defined as wi=(θi f u >0 1−θelsewhere and urepresents the error term. For instance, to compute the θ = 25th expectile, wi assigns weights 0.75 to observations below the regression, pulling the estimated equation toward the lower tail, while assigning a weight of 0.25 to observations above it. Regarding the regression component, the terms ˆ Yi1 and ˆ Yi0 in Equations (3) and (4) represent the fitted values of the regression model Y ij =X α + e ij computed separately for each group, with j= 0.1. In our analysis, Y represents earnings and Xdenotes the matrix of explanatory variables. The fitted values ˆ Yi0 and ˆ Yi1 are replaced with the unconditional distributions of the fitted values within each group, now denoted as ˆ ∼ Yi1 and ˆ ∼ Yi0 . While in the OLS framework, conditional and unconditional effects coincide, their interpretation differs when analyzing the tails, i.e., in the quantile framework (Frolich & Melly,2010).3 To estimate the unconditional distributions, we implement the approach proposed by Melly (2006). First, the conditional distributions of the dependent variables are estimated through quantile regression (Koenker,2005) at multiple quantiles (e.g., k= 100), separately for the treated and control groups, using the following objective function: ∑Y>Xα θ|Y−Xα|+∑Y<Xα(1−θ)|Y−Xα|(6) where θrepresents the selected quantile. The analysis within each group yields two sets of estimated coefficients, ˆα1(θ) for the treated group and ˆα0(θ) for the control. The corresponding fitted values, ˆ Yi0(θ)|X0= X0ˆα0(θ)and ˆ Yi1(θ)|X1=X1ˆα1(θ) , represent the outcome distribution at a given quantile, which is conditional on the covariates Xwithin each group. Estimating kquantile regressions within each group results in kvalues of ˆ ∼ α0(θ)and of ˆ ∼ α1(θ) . Additionally, the covariates are bootstrapped within each group as well, yielding ksamples of ∼ X0and ∼ X1 . By bootstrapping both the covariate coefficients, it is possible to estimate the unconditional distributions of the dependent variable for treated and untreated groups as follows: ˆ ∼ Yi0 =∼ X0ˆ ∼ α0(θ) and ˆ ∼ Yi1 =∼ X1ˆ ∼ α1(θ) . These terms replace the OLS-fitted values ˆ Yi0 and ˆ Yi1 in Equation (1), resulting in the following: Q∑n i=1(ZiYi P(wiZi=1|X)−Zi−P(wiZi=1|X) P(wiZi=1|X) ˆ ∼ Yi1)− ∑n i=1((1−wiZi)Yi 1−P(wiZi=1|X)+Zi−P(wiZi=1|X) 1−P(wiZi=1|X) ˆ ∼ Yi0) (7) Equation (7) enables the computation of the double-robust difference between the treated and control groups at any quantile of interest Q. This formulation allows for quantile-level comparison and ensures robustness against both model misspecification and extreme values.4 4. The Data and the Model The SHIW data are analyzed over the past decade, focusing on the 2010 and 2020 waves. 5 The sample consists of employees aged 20 to 65. Education is measured by the number of years required to complete a degree.
Economies 2025,13, 50 6 of 28 Table 1reports the summary statistics of the model variables. The education section shows a substantial increase in the number of workers with university and post-university degrees over time. This growth may raise concerns about job mismatch and overeducation, where highly educated workers are employed in lower-skilled jobs due to an oversupply of highly educated labor exceeding market demand. The other sections analyze gender, where the percentage of working women slightly decreases in the last wave; age, where the average age of young workers slightly increases over time—i.e., it takes a little longer to secure the first job; region of residence, where the percentage of southern workers decreases in the 2020 wave, possibly due to migration toward the wealthier northern regions; real annual wages, with average values increasing over time while the doubled standard deviations in 2020 signals higher inequality; and number of hours worked, with an almost stable mean but increased dispersion over time, possibly due to greater flexibility in job contracts.6 Table 1. Descriptive statistics. 2010, n = 13,733 2020, n = 10,876 Education Women Men Women Men At most junior high 4004 4324 2000 2278 High school 1727 2013 1469 1886 University 802 738 1443 1448 Post-university 46 79 146 206 Gender 2010 2020 Women 6579 47.9% 5058 47% Men 7154 52.1% 5818 53% 100% 100% Young workers 2010 2020 Average age 25.7 (s d = 3.01) 25.9 (s d = 3.06) Region of residence South 2010 2020 0 9039 65.8% 7213 66.4% 1 4694 34.2% 3663 33.7% 100% 100% Real annual wages net of taxes Sample mean 2010 2020 6828.80 (s d = 9997.2) 8599.50 (s d = 18475.2) Number of worked hours Sample mean 2010 2020 36.947 (s d = 9.45) 36.345 (s d = 10.53) Note: standard deviations in parenthesis. 4.1. Employment As an example of changing probability, Table 2considers the probability of being employed in the 2010 and 2020 waves as a function of age and age square—reflecting a non-linear impact of work experience, education, gender (a dummy variable equal to one for men), and region of residence. This table presents the results of the logistic model, computed at the median and in the tails. The expectile weights allow the logistic regression
Economies 2025,13, 50 7 of 28 to move away from the conditional mean, as described in Equation (5), while at the mean, the weights are set to one. Table 2. Probability of unemployment, expectile logistic regressions. 0.25 0.50 0.75 Year 2010, n = 5599 Coef. Std. Err. z Coef. Std. Err. z Coef. Std. Err. z Age −0.0196 0.005 −3.46 0.1129 0.011 9.61 0.0229 0.008 2.92 Age square 0.0005 0.0001 5.98 −0.0006 0.0001 −3.86 0.0003 0.0001 2.81 Education −0.0679 0.008 −8.74 0.0944 0.010 9.28 0.0202 0.009 2.23 South −1.149 0.056 −20.4 −1.284 0.068 −18.9 −1.133 0.067 −16.8 Men −0.1381 0.057 −2.40 0.3156 0.068 4.60 0.0978 0.066 1.48 Year 2020, n = 3957 Coef. Std. Err. z Coef. Std. Err. z Coef. Std. Err. z Age −0.0645 0.007 −8.63 0.1390 0.018 7.68 0.0076 0.009 0.79 Age square 0.0011 0.0001 10.6 −0.0009 0.0002 −4.08 0.0004 0.0001 2.96 Education −0.0259 0.096 −2.71 0.1798 0.012 14.8 0.0829 0.011 7.57 South −1.194 0.068 −17.4 −1.249 0.082 −15.1 −1.048 0.081 −12.8 Men −0.0936 0.068 −1.37 0.3115 0.084 3.71 −0.0907 0.083 −1.09 Note: the non-significant estimated coefficients are in italics. The impact of the explanatory variables on the probability of being employed changes, demonstrating the importance of accounting for heterogeneous probabilities. In homogenous settings, linear regressions at different locations would produce parallel lines that shift up or down according to the weights defining the specific location, with only the intercept changing from one location to another. However, differing estimated slopes and non-parallel regression lines indicate heterogeneity, as the regression line shifts and tilts to maintain the proportion of positive and negative residuals defined by the weights. In this example, age and education increase the probability of being employed at and above θ = 0.50, while at θ = 0.25, they have a negative impact in both waves. The impact of education varies significantly across θ , following an inverse u-shaped pattern—negative at θ = 0.25, positive and significant at θ = 0.50, and still positive but smaller at θ = 0.75. This suggests a declining probability of employment for highly educated workers, possibly due to an excess supply. The negative impact of residing in southern regions remains almost stable across θ. Gender is not statistically significant in the upper tail but it is positive and significant at θ = 0.50, and negative—though not significant in 2020—at θ = 0.25. Figure 1presents the box plots of probabilities at θ = 0.25, 0.50, and 0.75. These distributions differ substantially in terms of dispersion, particularly as measured by the interquartile range: IR(0.25) = 0.075, IR(0.50) = 0.148, and IR(0.75) = 0.231. Smaller probabilities, and thus larger potential output values, are required to balance employed/unemployed covariates in the upper tail.
Economies 2025,13, 50 8 of 28 Economies 2025, 13, x FOR PEER REVIEW 7 of 27 slopes and non-parallel regression lines indicate heterogeneity, as the regression line shifts and tilts to maintain the proportion of positive and negative residuals defined by the weights. In this example, age and education increase the probability of being employed at and above = 0.50, while at = 0.25, they have a negative impact in both waves. The impact of education varies significantly across , following an inverse u-shaped pattern—negative at θ = 0.25, positive and significant at = 0.50, and still positive but smaller at = 0.75. This suggests a declining probability of employment for highly educated workers, possibly due to an excess supply. The negative impact of residing in southern regions remains almost stable across Gender is not statistically significant in the upper tail but it is positive and significant at = 0.50, and negative—though not significant in 2020—at = 0.25. Figure 1 presents the box plots of probabilities at = 0.25, 0.50, and 0.75. These distributions differ substantially in terms of dispersion, particularly as measured by the interquartile range: IR(0.25) = 0.075, IR(0.50) = 0.148, and IR(0.75) = 0.231. Smaller probabilities, and thus larger potential output values, are required to balance employed/unemployed covariates in the upper tail. Figure 1. Box plot showing the probability of being employed at = 0.25, 0.50, 0.75, in 2020. 4.2. Returns to Education Next, we consider returns to education as computed by the double-robust method at = 0.25, 0.50, and 0.75. In the PS component, the probability of completing a higher level of education—namely high school, university, and post-university degrees—is estimated through logit and weighted logit models, with weights shifting the model towards the tails. This probability is a function of the cohort, defined as the year of the wave minus age, which is introduced to account for changes in the educational system over time that may have affected workers in different cohorts differently,7 gender, and region of residence, to measure the negative impact of the lagging southern economy—as documented in the literature and throughout Table 3. Fields of education also play a significant role, as the probability of achieving a degree depends on the field of the previously completed degree (Ballarino & Bratti, 2009). Figure 1. Box plot showing the probability of being employed at θ= 0.25, 0.50, 0.75, in 2020. 4.2. Returns to Education Next, we consider returns to education as computed by the double-robust method at θ = 0.25, 0.50, and 0.75. In the PS component, the probability of completing a higher level of education—namely high school, university, and post-university degrees—is estimated through logit and weighted logit models, with weights shifting the model towards the tails. This probability is a function of the cohort, defined as the year of the wave minus age, which is introduced to account for changes in the educational system over time that may have affected workers in different cohorts differently, 7 gender, and region of residence, to measure the negative impact of the lagging southern economy—as documented in the literature and throughout Table 3. Table 3. Probability of completing a degree, expectile logistic regression. High school Year 2010, n = 13480 θ= 0.25 θ= 0.50 θ= 0.75 Coef. z Coef. z Coef. z Cohort −0.0007 −32.4 −0.018 −20.8 0.0008 37.6 South −0.151 −4.08 −0.177 −4.62 −0.106 −3.26 Men 0.013 0.37 −0.048 −1.28 −0.070 −2.21 Year 2020, n = 8086 Coef. z Coef. z Coef. z Cohort −0.0009 −0.87 −0.026 −25.8 0.0008 27.6 South −0.083 1.69 −0.220 −4.52 −0.216 −4.39 Men −0.626 −13.5 0.058 1.24 0.022 0.47
Economies 2025,13, 50 15 of 28 Table 6. Cont. Year 2020 Coef. Std. Err. t Coef. Std. Err. t Coef. Std. Err. t High school −0.545 0.079 −6.84 −0.027 0.039 −0.70 0.396 0.097 4.07 University −0.230 0.043 −5.28 0.221 0.040 5.47 0.754 0.061 12.29 Post-university −0.260 0.071 −3.64 0.326 0.037 8.66 0.740 0.040 18.25 Note: the non-significant estimated coefficients are in italics. Economies 2025, 13, x FOR PEER REVIEW 14 of 27 Figure 3. Unconditional wage distributions vary across different education levels, with dispersion increasing as education level rises, especially for those with post-university degrees. In 2010, the upper quantile for post-university degree holders was significantly lower compared to those with university degrees. By 2020, wage premiums for post-university degree holders improved. Figure 3. Unconditional wage distributions vary across different education levels, with dispersion increasing as education level rises, especially for those with post-university degrees. In 2010, the upper quantile for post-university degree holders was significantly lower compared to those with university degrees. By 2020, wage premiums for post-university degree holders improved. In the 2020 wave, however, the university and post-university educational premiums increase in the private sector. The post-university wage distribution rises significantly above the university plot and surpasses its counterpart in the graph depicting the entire sample. The greater dispersion of wage distributions in the private sector leads to increased inequality, but it also offers greater opportunities for highly educated workers in 2020. The post-university plots confirm the compressed wage distributions for higher education in the public sector in 2020, though this pattern was not observed in 2010. Figure 4 illustrates the public sector, showing higher premiums and lower dispersion compared to the private sector.
Economies 2025,13, 50 16 of 28 Economies 2025, 13, x FOR PEER REVIEW 14 of 27 Figure 3. Unconditional wage distributions vary across different education levels, with dispersion increasing as education level rises, especially for those with post-university degrees. In 2010, the upper quantile for post-university degree holders was significantly lower compared to those with university degrees. By 2020, wage premiums for post-university degree holders improved. Economies 2025, 13, x FOR PEER REVIEW 15 of 27 Figure 4. Unconditional wage distributions in the public sector across different education levels increased in 2020 compared to 2010. In contrast to the private sector, as shown in Figure 3, these distributions are less dispersed and are centered around higher medians. 4.2.2. Young Workers The next section of Table 6 focuses on young workers aged 20−30, a group characterized by a relatively high unemployment rate. 13 Previous studies found reduced educational premiums for higher education within this subset. A possible explanation suggests an excess supply of highly educated workers, resulting in overqualification for available jobs (Ballarino & Scherer, 2013). The reduced premiums may be attributed to skill-biased technical change (Naticchioni et al., 2010), where overeducated workers accept lowerqualified jobs due to an oversupply of skills, displacing workers with high school or lower degrees. This phenomenon, referred to as “unskilled bias”, suggests that the demand for high-skilled workers has increased less than their supply, thereby reducing the educational premiums. As stated by (Naticchioni et al., 2010), “the increase in relative supply of education has exerted a negative impact on relative wages, while technical change−proxy for the demand for workers−is not statistically different from zero”. This could explain the low premiums at = 0.25. An alternative explanation considers generational shifts, where baby boomers transfer uncertainty to younger generations, who, in turn, face diminishing opportunities for career advancement and wage growth (Barbieri, 2011).14 Compared to other OECD countries, Biagetti and Scicchitano (2011) find that returns to education in Italy are lower. Our results show that, at the lower end of the wage distribution, the impact of education is generally negative, whereas it becomes positive and increases at and above the median. Thus, young highly educated workers receive minimal educational premiums in lower-qualified jobs. In Table 1, the percentage of university and post-university degrees in 2020 more than doubles the values of 2010, supporting the excess supply/unskilled bias hypothesis as an explanation for lower premiums among young workers. However, when they secure better jobs, their educational premiums improve, particularly at the 75th percentile for university and post-university degrees. Figure 5 displays the unconditional wage distributions for this subset, with post-university degree holders exhibiting the highest dispersion in both waves. The distributions are highly skewed, with a longer left tail in 2010 and a longer right tail in 2020. The bottom section of Table 6 reports the results for young workers in the private sector. In general, educational premiums are lower for young workers compared to previous results, with the post-university degree in 2010 providing no significant premium. The graphs in Figure 6 compare unconditional wage distributions for the full data set (left), the private sector (middle), and young workers (right). The substantial dispersion in post-university degree distributions is evident, with increased dispersion/inequality among young workers. The graphs also show that young workers’ wage distributions are generally positioned lower than their counterparts in the full sample. Figure 4. Unconditional wage distributions in the public sector across different education levels increased in 2020 compared to 2010. In contrast to the private sector, as shown in Figure 3, these distributions are less dispersed and are centered around higher medians. 4.2.2. Young Workers The next section of Table 6focuses on young workers aged 20 − 30, a group characterized by a relatively high unemployment rate. 13 Previous studies found reduced educational premiums for higher education within this subset. A possible explanation suggests an excess supply of highly educated workers, resulting in overqualification for available jobs (Ballarino & Scherer,2013). The reduced premiums may be attributed to skill-biased technical change (Naticchioni et al.,2010), where overeducated workers accept lower-qualified jobs due to an oversupply of skills, displacing workers with high school or lower degrees. This phenomenon, referred to as “unskilled bias”, suggests that the demand for high-skilled workers has increased less than their supply, thereby reducing the educational premiums. As stated by (Naticchioni et al.,2010), “the increase in relative supply of education has exerted a negative impact on relative wages,while technical change − proxy for the demand for workers − is not statistically different from zero”. This could explain the low premiums at θ = 0.25. An alternative explanation considers generational shifts, where baby boomers transfer uncertainty to younger generations, who, in turn, face diminishing opportunities for career advancement and wage growth (Barbieri, 2011). 14 Compared to other OECD countries, Biagetti and Scicchitano (2011) find that returns to education in Italy are lower. Our results show that, at the lower end of the wage distribution, the impact of education is generally negative, whereas it becomes positive and
Economies 2025,13, 50 17 of 28 increases at and above the median. Thus, young highly educated workers receive minimal educational premiums in lower-qualified jobs. In Table 1, the percentage of university and post-university degrees in 2020 more than doubles the values of 2010, supporting the excess supply/unskilled bias hypothesis as an explanation for lower premiums among young workers. However, when they secure better jobs, their educational premiums improve, particularly at the 75th percentile for university and post-university degrees. Figure 5displays the unconditional wage distributions for this subset, with post-university degree holders exhibiting the highest dispersion in both waves. The distributions are highly skewed, with a longer left tail in 2010 and a longer right tail in 2020. Economies 2025, 13, x FOR PEER REVIEW 16 of 27 Young workers in the private sector receive smaller educational premiums, as highlighted by the comparison of the third and fourth sections of the table and illustrated in Figure 7 for the 2020 wave. In the private sector, the dispersion of post-university degree premiums significantly decreases, suggesting that the greater rewards offered by the private sector do not necessarily benefit younger generations. Figure 5. Unconditional wage distributions at different degrees of education for young workers. The post-university distribution is quite dispersed and shows a longer left tail in 2010, while in 2020, there is a longer right tail. There are a few outliers at the upper end of the, at most, junior high plot. Figure 5. Unconditional wage distributions at different degrees of education for young workers. The post-university distribution is quite dispersed and shows a longer left tail in 2010, while in 2020, there is a longer right tail. There are a few outliers at the upper end of the, at most, junior high plot. The bottom section of Table 6reports the results for young workers in the private sector. In general, educational premiums are lower for young workers compared to previous results, with the post-university degree in 2010 providing no significant premium. The graphs in Figure 6compare unconditional wage distributions for the full data set (left), the private sector (middle), and young workers (right). The substantial dispersion in postuniversity degree distributions is evident, with increased dispersion/inequality among young workers. The graphs also show that young workers’ wage distributions are generally positioned lower than their counterparts in the full sample.
Economies 2025,13, 50 18 of 28 Young workers in the private sector receive smaller educational premiums, as highlighted by the comparison of the third and fourth sections of the table and illustrated in Figure 7for the 2020 wave. In the private sector, the dispersion of post-university degree premiums significantly decreases, suggesting that the greater rewards offered by the private sector do not necessarily benefit younger generations. 4.2.3. Gender Gap Finally, we examine the subset of women to evaluate the presence of a gender pay gap. The gender gap is not only associated with equal pay for women, but also with differences in the process of selection for employment. Women’s participation rates are lower and are concentrated in high-wage positions (Picchio & Mussida,2011), facing barriers such as the glass ceiling. Depalo and Giordano (2010) found evidence of a gender pay gap in Italy. The results in Table 7indicate that, in 2010, women’s educational premiums were generally lower compared to the full sample results, showing the existence of a gender gap in educational premiums. The bottom section of Table 7compares, at the median, the full sample results from Table 4(middle column) with those for the women’s subset and the women in the private sector. The table highlights in bold the subset estimates that significantly differ from the full sample results. Examining the women’s subset and considering a confidence interval of ± 2 ˆ σ , it is evident that almost all results significantly differ from the full sample results in both waves. The only exception is the field of the humanities, which offers the same rewards to both men and women. Economies 2025, 13, x FOR PEER REVIEW 17 of 27 Figure 6. Cont.
Economies 2025,13, 50 19 of 28 Economies 2025, 13, x FOR PEER REVIEW 18 of 27 Figure 6. Unconditional wage distributions are presented across different categories. The top left graphs show distributions for the entire data set, the top right graphs represent the private sector, and the bottom graphs focus on young workers aged 20–30. In 2010, the post-university wage distribution in the private sector was more dispersed compared to the entire data set, with even greater dispersion observed in the post-university premiums for young workers. By 2020, the private sector exhibited greater dispersion in university wage premiums relative to the entire data set. Post-university premiums for young workers showed an even wider dispersion, while other education levels were less dispersed and had lower premiums. Figure 6. Unconditional wage distributions are presented across different categories. The top left graphs show distributions for the entire data set, the top right graphs represent the private sector, and the bottom graphs focus on young workers aged 20–30. In 2010, the post-university wage distribution in the private sector was more dispersed compared to the entire data set, with even greater dispersion observed in the post-university premiums for young workers. By 2020, the private sector exhibited greater dispersion in university wage premiums relative to the entire data set. Post-university premiums for young workers showed an even wider dispersion, while other education levels were less dispersed and had lower premiums. In both waves, all estimates for the women’s subset are in bold, except for those related to graduates in the humanities in 2010 and postgraduates in the same field in 2020. Women across all degrees and fields—aside from the humanities—receive lower rewards. In the private sector subset, the statistically significant differences indicate reduced returns compared to both the full sample and the women’s subset, with the sole exception of the post-university degrees in social sciences in 2020. Figure 8presents the box plot of the unconditional distributions of educational premiums. In 2020, the premiums exhibit greater dispersion at lower educational levels, and the post-university premium at the 75th percentile is lower than in 2010. This effect is even more pronounced in the private sector, as shown in the middle section of the table.
Economies 2025,13, 50 20 of 28 Economies 2025, 13, x FOR PEER REVIEW 19 of 27 Figure 7. Unconditional wage distributions are shown across different education levels: the top left graph represents the full sample, the top right graph shows young workers, and the bottom graph focuses on young workers in the private sector. In the private sector, the dispersion of post-university wage premiums significantly decreases, and the higher rewards seen in the private sector in Figure 2 do not benefit the younger generation as much. 4.2.3. Gender Gap Finally, we examine the subset of women to evaluate the presence of a gender pay gap. The gender gap is not only associated with equal pay for women, but also with differences in the process of selection for employment. Women’s participation rates are lower and are concentrated in high-wage positions (Picchio & Mussida, 2011), facing barriers such as the glass ceiling. Depalo and Giordano (2010) found evidence of a gender pay gap in Italy. The results in Table 7 indicate that, in 2010, women’s educational premiums were generally lower compared to the full sample results, showing the existence of a gender gap in educational premiums. The bottom section of Table 7 compares, at the median, the full sample results from Table 4 (middle column) with those for the women’s subset and the women in the private sector. The table highlights in bold the subset estimates that significantly differ from the full sample results. Examining the women’s subset and considering a confidence interval of ±2𝛔, it is evident that almost all results significantly differ from the full sample results in both waves. The only exception is the field of the humanities, which offers the same rewards to both men and women. Figure 7. Unconditional wage distributions are shown across different education levels: the top left graph represents the full sample, the top right graph shows young workers, and the bottom graph focuses on young workers in the private sector. In the private sector, the dispersion of post-university wage premiums significantly decreases, and the higher rewards seen in the private sector in Figure 2 do not benefit the younger generation as much. Table 7. (a) Women’s subset. (b) Women in the private sector. (c) Comparison of double-robust results: full sample versus women’s subsets at the median. (a) 0.25 0.50 0.75 Year 2010 Coef. Std. Err. t Coef. Std. Err. t Coef. Std. Err. t High school −0.317 0.028 − 11.30 0.226 0.019 11.87 11.87 0.023 30.68 University 0.141 0.024 5.74 0.550 0.012 42.99 42.99 0.016 55.77 Post-university 0.480 0.012 37.15 0.780 0.011 67.18 67.18 0.015 73.95
Economies 2025,13, 50 21 of 28 Table 7. Cont. (a) 0.25 0.50 0.75 Year 2020 Coef. Std. Err. t Coef. Std. Err. t Coef. Std. Err. t High school −0.371 0.035 − 10.60 0.101 0.027 3.65 0.640 0.039 16.05 University −0.259 0.051 −5.02 0.538 0.027 19.85 1.08 0.024 44.12 Post-university −0.177 0.152 −1.16 0.628 0.042 14.94 1.01 0.034 29.38 (b) 0.25 0.50 0.75 Year 2010 Coef. Std. Err. t Coef. Std. Err. t Coef. Std. Err. t High school −0.302 0.042 −7.21 0.244 0.023 10.32 0.704 0.028 24.67 University 0.160 0.016 9.57 0.467 0.016 28.18 0.870 0.020 42.50 Post-university −0.142 0.024 5.95 0.286 0.019 14.56 0.598 0.013 43.02 Year 2020 Coef. Std. Err. t Coef. Std. Err. t Coef. Std. Err. t High school −0.460 0.055 −8.33 0.100 0.033 2.98 0.657 0.058 11.21 University −0.055 0.056 −0.98 0.593 0.027 21.91 1.04 0.033 31.46 Post-university 0.262 0.021 12.19 0.605 0.019 30.81 0.928 0.020 45.79 (c) Year 2010 Full sample Women subset Women in private sector Coef. Std. Err. t Coef. Std. Err. t Coef. Std. Err. t University humanities 0.543 0.008 68.74 0.529 0.012 42.49 0.305 0.016 18.59 University science 0.560 0.008 71.01 0.524 0.012 43.43 0.389 0.016 24.23 University social sc. 0.589 0.008 73.85 0.544 0.012 44.31 0.387 0.016 23.88 Post-university humanities 0.701 0.007 92.07 0.662 0.011 57.49 0.929 0.132 7.00 Post-university science 0.764 0.008 90.39 0.672 0.039 17.20 0.286 0.019 14.81 Post-university social sc. 0.692 0.008 85.92 0.743 0.011 65.32 0.867 0.132 6.55 Year 2020 Full sample Women subset Women in private sector Coef. Std. Err. t Coef. Std. Err. t Coef. Std. Err. t University humanities 0.558 0.012 45.49 0.461 0.020 22.50 0.574 0.047 12.03 University science 0.625 0.013 47.87 0.495 0.022 22.28 0.581 0.025 22.88 University social sc. 0.596 0.013 46.14 0.499 0.021 23.35 0.501 0.024 20.47 Post-university humanities 0.501 0.009 50.23 0.507 0.013 36.56 0.100 0.018 5.49 Post-university science 0.711 0.011 61.69 0.659 0.015 43.12 0.496 0.020 24.47 Post-university social sc. 0.769 0.011 70.65 0.714 0.013 52.69 0.817 0.023 34.12 Note: The non-significant estimated coefficients are in italics. In bold are the estimated coefficients in each women’s subset that significantly differ from the analogous estimates computed in the entire sample.
Economies 2025,13, 50 22 of 28 Economies 2025, 13, x FOR PEER REVIEW 21 of 27 University science 0.625 0.013 47.87 0.495 0.022 22.28 0.581 0.025 22.88 University social sc. 0.596 0.013 46.14 0.499 0.021 23.35 0.501 0.024 20.47 Post-university humanities 0.501 0.009 50.23 0.507 0.013 36.56 0.100 0.018 5.49 Post-university science 0.711 0.011 61.69 0.659 0.015 43.12 0.496 0.020 24.47 Post-university social sc. 0.769 0.011 70.65 0.714 0.013 52.69 0.817 0.023 34.12 Note: The non-significant estimated coefficients are in italics. In bold are the estimated coefficients in each women’s subset that significantly differ from the analogous estimates computed in the entire sample. Figure 8 presents the box plot of the unconditional distributions of educational premiums. In 2020, the premiums exhibit greater dispersion at lower educational levels, and the post-university premium at the 75th percentile is lower than in 2010. This effect is even more pronounced in the private sector, as shown in the middle section of the table. Figure 9 compares the unconditional distributions across the full sample (left), the women’s subset (middle), and women in the private sector (right). The graphs illustrate that, in both waves, post-university premiums for women in the private sector are lower than university-level premiums. Figure 8. Unconditional wage distributions for women across different education levels show notable changes. In 2020, wage premiums for lower education levels are more dispersed, while the top premiums for post-university degrees are lower compared to 2010. Additionally, there are a few outliers at the lower end of the 2020 post-university distribution. Figure 9compares the unconditional distributions across the full sample (left), the women’s subset (middle), and women in the private sector (right). The graphs illustrate that, in both waves, post-university premiums for women in the private sector are lower than university-level premiums. In summary, education provides wage premiums in each wave. While higher degrees yield greater premiums in the private sector, young workers and women employed in the private sector generally receive lower educational premiums.
Economies 2025,13, 50 23 of 28 Economies 2025, 13, x FOR PEER REVIEW 22 of 27 Figure 8. Unconditional wage distributions for women across different education levels show notable changes. In 2020, wage premiums for lower education levels are more dispersed, while the top premiums for post-university degrees are lower compared to 2010. Additionally, there are a few outliers at the lower end of the 2020 post-university distribution. Figure 9. Cont.
Economies 2025,13, 50 24 of 28 Economies 2025, 13, x FOR PEER REVIEW 23 of 27 Figure 9. A comparison of unconditional wage distributions is presented: to the top left, for the entire sample; to the top right, for the women’s subset; and to the bottom, for women working in the private sector. In summary, education provides wage premiums in each wave. While higher degrees yield greater premiums in the private sector, young workers and women employed in the private sector generally receive lower educational premiums. 5. Conclusions The earning profiles across different levels of education are analyzed using the double-robust approach at the mean and in the tails. The propensity score is implemented to compute changing probabilities at different locations. These results are combined with quantile regression-based unconditional distributions to analyze returns to education, focusing on the 2010 and 2020 waves of Banca d’Italia SHIW data. The double-robust approach is applied not only at the center but also in the tails of both components, the propensity score and the regression model, providing a deeper understanding of the behavior and enhancing robustness. In the propensity score, cohort, gender, field of studies, and region of residence significantly influence the likelihood of attaining higher educational degrees. The negative Figure 9. A comparison of unconditional wage distributions is presented: to the top left, for the entire sample; to the top right, for the women’s subset; and to the bottom, for women working in the private sector. 5. Conclusions The earning profiles across different levels of education are analyzed using the doublerobust approach at the mean and in the tails. The propensity score is implemented to compute changing probabilities at different locations. These results are combined with quantile regression-based unconditional distributions to analyze returns to education, focusing on the 2010 and 2020 waves of Banca d’Italia SHIW data. The double-robust approach is applied not only at the center but also in the tails of both components, the propensity score and the regression model, providing a deeper understanding of the behavior and enhancing robustness. In the propensity score, cohort, gender, field of studies, and region of residence significantly influence the likelihood of attaining higher educational degrees. The negative impact of southern regions diminishes at the mean and becomes non-significant at the top quartile. The double-robust results reveal sectoral differences in educational premiums, with the