Modelling the occupational and educational choices of young people in Poland using Bayesian multinomial logit models
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Grzenda, Wioletta Article Modelling the occupational and educational choices of young people in Poland using Bayesian multinomial logit models Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Grzenda, Wioletta (2021) : Modelling the occupational and educational choices of young people in Poland using Bayesian multinomial logit models, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, Vol. 22, Iss. 3, pp. 175-191, https://doi.org/10.21307/stattrans-2021-033 This Version is available at: https://hdl.handle.net/10419/266277 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/
STATISTICS IN TRANSITION new series, September 2021 Vol. 22, No. 3, pp. 175–191, DOI 10.21307/stattrans-2021-033 Received – 24.11.2019; accepted – 07.06.2021 Modelling the occupational and educational choices of young people in Poland using Bayesian multinomial logit models Wioletta Grzenda1 ABSTRACT Binomial logit models are commonly used in the analysis of the situation of respondents on the labour market. Consequently, in most cases researchers consider two states: of being unemployed and employed or economically inactive and active. This paper focuses on the situation of young people aged 18 to 29 on the labour market in Poland. A major part of the people who comprise the studied group are still in education or combine education with work. Therefore, the participants of the research were divided into the following groups: the employed and not learning, those combining education with work, the unemployed, learners/students only, and those economically inactive and not at school. The model allowing an analysis which includes both the most common division into working and nonworking persons as well as the division proposed in this study is a nested logit model. This model has a hierarchical structure and is a special case of a multinomial logit model. In this paper, all models were estimated within the Bayesian approach. The findings show that continuing education by young people may result from their problems with finding a job; moreover, combining work with education is not the preferred form of professional activity. In addition, the study examines the inequalities observed on the Polish labour market. Key words: young people, labour market, education, multinomial logit model, Bayesian approach. 1. Introduction In socio-economic research, models for the dichotomous dependent variables are very popular (Cramer, 2003; Allison, 2009). Unfortunately, with their use, only two states or two events for a given unit can be analysed. In the case of issues related to the labour market, division into economically active and economically inactive, as well as employed and unemployed persons is usually made. In the case where the examined feature has more than two levels, a better solution than combining selected categories is to use models for discrete outcome variables that can take more than two possible 1 SGH Warsaw School of Economics, Collegium of Economic Analysis, Institute of Statistics and Demography, Poland. E-mail: [email protected]. ORCID: https://orcid.org/0000-0002-2226-4563.
176 W. Grzenda: Modelling the occupational and educational… values. Among this group of models, two main classes are distinguished: models for ordinal response variables and models for dependent variables with unordered categories. The second group includes: the MultiNomial Logit Model (MNLM), the Conditional Logit Model (CLM), the Mixed Logit Model (MLM) and the Nested Logit Model (NLM) (Cameron and Trivedi, 2005). The choice of a model depends primarily on whether the independent variables included in the model vary across alternatives or they are the same across alternatives (Cameron and Trivedi, 2005). The standard multinomial logit model can be used when the model takes into account only the features of the individuals studied without taking into account the features of the selected categories. If this assumption is not met, the conditional logit model is used unless both types of features are considered. In the latter case the mixed logit model is used. In addition, according to Stanisz (2016), in order for the standard multinomial model to be used, the categories of the responding variable should be independent and distinguishable for the decision maker. Both the multinomial logit model and the conditional logit model have some limitations regarding the assumption of independence from irrelevant (unrelated) alternatives (IIA). The model in which this assumption can be slightly weakened, and also can take into account the hierarchy of alternatives, is the nested logit model considered in this paper. This model is not widely used due to the problems related to the estimation of its parameters. To avoid these problems, the Bayesian approach and Markov Chain Monte Carlo methods (MCMC) were used in this work (Robert and Casella, 2004). The purpose of this study is to analyse the occupational and educational choices of young people aged 18 to 29 in Poland. In most studies on this issue, the division of young people into those who have already completed education and those who continue their education, e.g. at a higher level (de Dios Jiménez and Salas-Velasco, 2000), economically active and economically inactive (MRPiPS, 2018) or unemployed and employed (Gallie and Paugam, 2000; Grzenda, 2012; Bieszk-Stolorz and Markowicz, 2013) are considered. The binary divisions presented above can be further detailed. For example, among the economically inactive there are both those who are unwilling to take up employment despite their abilities and young people who remain in the education system and have not started their careers yet. In addition, it is worth considering in the research that young people sometimes combine education with work. Therefore, in this study, the respondents were divided into employed, combining education with work, learners only, and unemployed or persons economically inactive but not being learners. The methodological approach proposed in this work makes it possible to consider in the analysis both a more general division into working and nonworking persons, as well as a more detailed division taking into account education of youth. Information on educational and economic activity of young people in Poland was obtained from the Labour Force Survey (LFS).
STATISTICS IN TRANSITION new series, September 2021 177 The subject addressed in this study is very important because according to many reports (CSO, 2016a; CSO, 2016b; MRPiPS, 2018) the situation of young people on the labour market in Poland is the worst compared to other age groups. In addition, economists are concerned about the growing phenomenon of NEET (not in employment, education or training) (Chłoń-Domińczak and Strawiński, 2013), which affects young people who are neither in education nor working. The consequences of this phenomenon apply to the entire economy as well as to individuals who lose their competence over time. Youth unemployment has also a social dimension, lack of employment negatively affects family and fertility decisions, and, as a result, the demographic situation of the country. Therefore, the identification of factors determining the educational and professional decisions of young people may help identify solutions that may improve the situation of these people on the labour market in Poland. 2. Multinomial models Models for unordered categorical dependent variable are also considered as discrete choice models and are most often used in marketing research (Anderson, De Palma and Thisse, 1992). In the case of the binomial logit model, it can be assumed that a given unit has two variants to choose from. Suppose now that the i-th unit 𝑖1,…,𝑛 has to select not two but J unordered categories. These categories are mutually exclusive and constitute a whole set of possible selection options for the units under consideration. In the case where the independent variables do not differ for the alternatives considered, a standard multinomial logit model (MNLM) is considered. For this model, the probability of observing the choice by the i-th unit 𝑖1,…,𝑛 of j-th category 𝑗1,…,𝐽 is given by the formula: 𝑝𝑒𝑥𝑝𝐱′𝛃 ∑𝑒𝑥𝑝𝐱′𝛃 ,𝑖1,…,𝑛,𝑗1,…,𝐽, where 𝐱 denotes the vector of independent variables and 𝛃 is the vector of parameters. The sum of these probabilities for all categories 𝑗1,…,𝐽 is 1. If the independent variables differ for the alternatives considered, the standard multinomial model cannot be used; the conditional logit model (CLM) is considered then. In the case of this model, the probability of observing the selection of the j-th category 𝑗1,…,𝐽 by the i-th unit 𝑖1,…,𝑛 is given by the formula: 𝑝𝑒𝑥𝑝𝐱′𝛃 ∑𝑒𝑥𝑝𝐱′𝛃 ,𝑖1,…,𝑛,𝑗1,…,𝐽. The combination of both considered models is the mixed logit model (MLM) (Cameron and Trivedi, 2005).
178 W. Grzenda: Modelling the occupational and educational… The presented models can also be considered more generally in the context of the additive random utility models (ARUM) and discrete choice theory. In this approach, each unit assigns to each category j certain utility 𝑈, 𝑗1,…,𝐽 and selects the one with the highest utility. Let 𝑈𝐱′𝛃𝜀,𝑖1,…,𝑛,𝑗1,…,𝐽, denote the utility function. By making different assumptions about the random component of utility, different multinomial logit models can be obtained. In the standard multinomial logit model, the random components 𝜀 𝑗1,…,𝐽 are independent and identically Gumbel distributed (have the type I extreme-value distribution), with the density function given by the formula: 𝑓𝜀𝑒𝑒𝑥𝑝𝑒,𝑗1,…,𝐽. According to assumptions made in (McFadden, 1974), to be able to use a standard logit multinomial model, the categories analysed must meet the assumption of independence from irrelevant alternatives (IIA). This assumption also applies to the conditional logit model. However, it is often not fulfilled. By eliminating or adding one alternative, the quotient of the probability of the categories considered so far often changes. Unfortunately, there are no tests that conclusively determine whether IIA assumption is met. Cheng and Long (2007) have shown that two existing tests by Hausman and McFadden (1984) and Small and Hsiao (1985) can be unreliable. Then the solution may be to use another model, namely the nested logit model (Train, 2009). The nested logit model has a hierarchical structure. The set of all possible alternatives is divided into the so-called nests so that the assumption of independence from irrelevant alternatives (IIA) is met only in each nest, but it does not have to be met between the nests. Therefore, in the nested logit model, all random components 𝜀 𝑗1,…,𝐽 do not have to be independent. In addition, instead of the Gumbel distribution, the generalized extreme-value distribution (GEV) is assumed for these components. Let K denote the number of disjoint subsets (nests) 𝑆,𝑆,…,𝑆, into which the possible alternatives have been divided. Then, the cumulative distribution function for the random components vector 𝛆𝜀,𝜀,…,𝜀, is given by the formula: 𝐹𝛆𝑒𝑥𝑝𝑒𝑥𝑝𝜀 𝜆 ⁄ ∈ . Within each of the nests, random components 𝜀 𝑗1,…,𝐽 are correlated. The 𝜆 parameter is a function of the correlation coefficient between possible alternatives in the k-th nest and is used to measure the correlation between the categories in the nest. The value of 1 for the 𝜆 parameter means no correlation in the
STATISTICS IN TRANSITION new series, September 2021 179 k-th nest, therefore if the value of this parameter for all nests is 1, then the nested logit model can be replaced with a standard multinomial logit model. With the previously introduced notation, the choice probability for alternative 𝑗∈𝑆 by i-th 𝑖1,…,𝑛 unit for the nested logit model is given by the formula: 𝑃𝑦1𝑒𝑥𝑝𝐱′𝛃𝜆 ⁄∑𝑒𝑥𝑝𝐱′𝛃𝜆 ⁄ ∈ ∑∑𝑒𝑥𝑝𝐱′𝛃𝜆 ⁄ ∈ . Then, the likelihood function is in the form: 𝑝𝐲|𝛃,𝛌𝑃𝑦1 , where 𝛌𝜆,…,𝜆. In this article, the Bayesian approach was used to estimate the parameters of the nested logit model (Lahiri and Gao, 2002; Rossi, Allenby and McCulloch, 2005). This approach requires a prior distribution for the vector of coefficient parameters 𝛃 and the parameter vector 𝛌. For the parameter vector 𝛃, depending on the prior information, the most common are flat priors or the normal prior distributions. For the components of the 𝛌 parameter vector and for 𝑎0, the following prior distribution was used in this paper: 𝑝𝜆𝑎𝜆exp𝜆for 𝜆0, 0 for 𝜆0. Examples of other prior distributions for the parameter vector 𝛌 can be found in Lahiri and Gao (2002). This could be, for example, a beta or gamma distribution. Using the notation applied for the nested logit model, the formula for the posterior distribution has the form: 𝑝𝛃,𝛌|𝐲∝𝑝𝐲|𝛃,𝛌𝑝𝛃𝑝𝛌. In this paper, Markov Chain Monte Carlo (MCMC) methods were used to determine the marginal posterior distributions, in particular the methods used were the Metropolis algorithm (Gelman, et al., 2000) and the Gamerman algorithm (Gamerman, 1997). 3. Reference data To analyse the situation of young people on the labour market in Poland, data from the Labour Force Survey (LFS) were used. The LFS is a quarterly panel survey with a rotational sample selection scheme. In this study, the research sample comprised units that were surveyed for two consecutive quarters in 2015. These are people from the samples numbered 63-65 and 67-69. This selection of the sample enabled, inter alia,
180 W. Grzenda: Modelling the occupational and educational… verification of the answers given. In the first stage of the analysis, in accordance with the adopted research objective, people aged 18 to 29 were selected from the entire data set, thus separating a sample of 16,144 respondents. Then, the respondents were divided into five categories due to their situation on the labour market: 1. only learners/students, 2. employed but not being learners, 3. combining education with work, 4. unemployed persons but economically active, 5. economically inactive people but not being learners. Learners were selected based on question No. 90 (In the last 4 weeks, including as a last week the week of the survey, were you a student?). Then, they were divided into people who had a job and those who did not. Having a job as described in this article means doing professional work in accordance with survey question 12 or having a job but temporarily not doing related work, as identified based on the answer to question 13. (Questions: 12. Did you perform work for at least 1 hour, which provided earnings or income in the week under study from Monday to Sunday, or assist in a family business for free? 13. Did you have a job in the week under study, but did not perform it temporarily?). Then, from among persons who did not have a job and did not learn, economically active and economically inactive people were distinguished. Economically active persons mean those who were looking for a job and were ready to take up a job in accordance with survey questions 71 and 79. (71. In the last 4 weeks including the week of the survey, did you look for a job? 79. Could you take a job in the 2 weeks following the week of the survey?). Table 1. A set of potential explanatory variables Variable Description Categories Percent age_group Age group at the time of the survey 1 = from 18 to 19 years old 2 = from 20 to 24 years old 3 = from 25 to 29 years old 17.68 40.96 41.35 sex Sex 0 = woman 1 = man 49.27 50.73 education Level of education 1 = higher 2 = post-secondary and secondary professional 3 = secondary general 4 = basic vocational 5 = primary school 22.76 22.01 23.19 12.38 19.66 marital_status Marital status 0 = unmarried, a widower, a widow, separated or divorced 1 = married 78.81 21.19
STATISTICS IN TRANSITION new series, September 2021 181 Table 1. A set of potential explanatory variables (cont.) Variable Description Categories Percent child The presence of a child under 15 years in the household 0 = no 1 = yes 77.87 22.13 place_residence Class of place of residence during the survey 0 = village 1 = town 47.63 52.37 region Region of Poland 1 = Central Łódzkie, Mazowieckie) 2 = Southwest (Dolnośląskie, Opolskie) 3 = South (Małopolskie, Śląskie) 4 = Northwest (Wielkopolskie, Zachodniopomorskie, Lubuskie) 5 = North (Kujawsko-Pomorskie, Warmińsko-Mazurskie, Pomorskie) 6 = East (Lubelskie,Podkarpackie, Świętokrzyskie, Podlaskie) 15.03 11.92 14.33 15.04 18.14 25.55 The employed only persons were the largest part of the entire group – 43.99%. Learners constituted 30.22%, among them were both economically inactive and unemployed people. Introducing a more detailed breakdown of learners would mean introducing more values of a dependent variable, and when interpreted against one reference level, it could give hardly clear results. In addition, substantive considerations also had an impact on this division. Namely, this group includes, for example, part-time students who did not enter full-time studies and often have difficulties in determining whether they are not working because they cannot find a job, or because a lot of their time is consumed by studying or it can be an obstacle that they have to attend weekend classes starting on Fridays. The next subgroup includes persons economically inactive but not being learners. The high percentage of economically inactive and not learning persons is worrying as the share of this group is 10.83%. The share of unemployed was 8.46%. Considering the general population in the period under study, it is worth emphasizing that among all the unemployed people aged 18 to 29 accounted for as much as 37.39% (CSO, 2016a). The smallest percentage share was obtained for working and studying people – 6.5%. Based on the presented breakdown, the dependent variable was constructed. To do this, the last two groups, i.e. groups 4 and 5 were combined into one group: the unemployed and the inactive but not learning. In this way, a group of people unemployed and persons economically inactive but not being learners, which, consider phenomenon of NEET, was then selected as a reference group in the paper. One of the
182 W. Grzenda: Modelling the occupational and educational… research objectives was to analyse the impact of individual characteristics of the respondents on their situation on the labour market. Therefore, a set of potential explanatory variables included in this study was developed, which is presented in Table 1. 4. The model estimation In the first stage of the analysis, the nested logit model was estimated in the Bayesian approach. Due to the primary division of the surveyed respondents into working and non-working persons, the two-nest model was chosen. The first nest contains both learning and unemployed, and persons economically inactive but not being learners, while the second one employed and people who combine work and education. Taking into account the large sample size, all considered models were estimated using normal non-informative prior distributions. For the parameter vector 𝛃, the normal prior distributions with mean equal to 0 and variance equal to 100 were adopted in all models. The formula for the prior distribution for the lambda parameter has been presented in Section 2. In this paper, the Metropolis algorithm (Gelman, et al., 2000) or the Gamerman algorithm (Gamerman, 1997) have been used for sampling from multidimensional distributions, depending on the model under consideration. The results for the nested logit model are presented in Table 2. The assessment of convergence of generated chains was made using the Geweke test. Based on the results obtained for both models at the significance level of α = 0.05, the null hypothesis that the obtained chains for the considered parameters of these models are convergent cannot be rejected (Table 2). Two nests were included in the model and none of them was degenerated, therefore posterior values for two lambda parameters were determined. These parameters are used to measure the correlation between alternatives in each nest. The lambda values obtained are less than 1, therefore the nested logit model is a better model for analysing the situation of young people on the labour market in Poland in the examined period, compared to the standard multinomial logit model, because it takes into account the correlation in the considered nests. Based on the results contained in Table 2, it can be concluded that if the option of non-working and not in education is not considered, then the second option, i.e. employed but not learning is the most important for the respondents, while the second most important one is only learning, in both cases compared to the option of nonworking and not in education. On the other hand, the option of combining education with professional work definitely loses its significance, also compared to the reference option.
STATISTICS IN TRANSITION new series, September 2021 189 This study has been prepared as part of the project granted by the National Science Centre, Poland, entitled "The modelling of parallel family and occupational careers with Bayesian methods" (2015/17/B/HS4/02064). References Allison, P. D., (2009). Logistic Regression Using the SAS®. Theory and Application. 8th ed. Cary, NC: SAS Institute Inc. Anderson, S. P., De Palma, A. and Thisse, J. F., (1992). Discrete choice theory of product differentiation. Massachusetts: The MIT Press. Becker, G. S., (1991). A Treatise on the Family. Cambridge: Harvard University Press. Becker, G. S., (2010). The Economics of Discrimination. USA: University of Chicago Press. Bieszk-Stolorz, B., Markowicz, I., (2013). Men’s and Women’s Economic Activity in Poland. Acta Universitatis Lodziensis. Folia Oeconomica, 285, pp. 221–227. Brooks, R., (2006). Learning and work in the lives of young adults. International Journal of Lifelong Education, 25(3), pp. 271–289. Cameron, A. C., Trivedi, P. K., (2005). Microeconometrics: methods and applications. Cambridge: Cambridge University Press. Castellano, R., Rocca, A., (2017). The dynamic of the gender gap in the European labour market in the years of economic crisis. Quality & Quantity, 51(3), pp. 1337–1357. Cheng, S., Long, J. S., (2007). Testing for IIA in the multinomial logit model. Sociological Methods & Research, 35(4), pp. 583–600. Chłoń-Domińczak, A. and Strawiński, P., (2013). Wchodzenie osób młodych na rynek pracy w Polsce. In Proceedings of the 9th Congress of Polish Economists, pp. 28–29. Cramer, J. S., 2003. Logit Models from Economics and Other Fields. Cambridge: Cambridge University Press. Central Statistical Office of Poland, (2016a). Aktywność ekonomiczna ludności Polski w latach 2013 – 2015. Warszawa: CSO. Central Statistical Office of Poland, (2016b). Monitoring Rynku Pracy, Kwartalna informacja o rynku pracy. Warszawa: CSO.
190 W. Grzenda: Modelling the occupational and educational… Dale, K., (2009). Household skills and low wages. Journal of Population Economics, 22(4), pp. 1025–1038. Davidescu, A. A. M., Roman, M., Strat, V. A. and Mosora, M., (2019). Regional sustainability, individual expectations and work motivation: A multilevel analysis. Sustainability, 11(12), 3331. de Dios Jiménez, J., Salas-Velasco, M., (2000). Modeling educational choices. A binomial logit model applied to the demand for higher education. Higher Education, 40(3), pp. 293–311. Gallie, D., Paugam, S. eds., (2000). Welfare regimes and the experience of unemployment in Europe. Oxford: OUP Oxford. Gamerman, D., (1997). Sampling from the posterior distribution in generalized linear mixed models. Statistics and Computing, 7(1), pp. 57–68. Gelman, A., Carlin, J. B., Stern, H. S. and Rubin, D. B., (2000). Bayesian Data Analysis. London: Chapman & Hall/CRC. Grzenda, W., (2012). Badanie determinant pozostawania bez pracy osób młodych z wykorzystaniem semiparametrycznego modelu Coxa. Przegląd Statystyczny, 59(1), pp. 123–139. Grzenda, W., (2019). Modelowanie karier zawodowej i rodzinnej z wykorzystaniem podejścia bayesowskiego. Warszawa: Wydawnictwo Naukowe PWN. Hausman, J., McFadden, D., (1984). Specification tests for the multinomial logit model. Econometrica, 52, pp. 1219–1240. Hsiao, C., Small, K., (1985). Multinomial logit specification tests. International economic review, 26(3), pp. 619–627. Jasiński, M., Bożykowski, M., Chłoń-Domińczak, A., Zając, T. and Żółtak, M., (2017). Who gets a job after graduation? Factors affecting the early career employment chances of higher education graduates in Poland. Edukacja Quarterly, 143(4). Kołaczek, B., (2005). Podstawowe uwarunkowania społeczne dostępu młodzieży do kształcenia. Polityka Społeczna, 1. Lahiri, K., Gao, J., (2002). Bayesian Analysis of Nested Logit Model by Markov Chain Monte Carlo. Journal of Econometrics, 11, pp. 103–133. McFadden, D., (1978). Modelling the Choice of Residential Location. In: A. Karlqvist, L. Lundqvist, F. Snickars, and J. Weibull, eds. Spatial Interaction Theory and Planning Models. Amsterdam: North-Holland, pp. 75–96.
STATISTICS IN TRANSITION new series, September 2021 191 Michaud, P. C., Tatsiramos, K., (2011). Fertility and female employment dynamics in Europe: the effect of using alternative econometric modeling assumptions. Journal of Applied Econometrics, 26(4), pp. 641–668. Ministerstwo Rodziny, Pracy i Polityki Społecznej, Departament Rynku Pracy, (2018). Sytuacja na rynku pracy osób młodych w 2017 roku, Warszawa: MRPiPS. Robert, C. P., Casella, G., (2004). Monte Carlo Statistical Methods. 2nd ed. New York: Springer. Rossi, P. E., Allenby, G. M. and McCulloch, R., (2005). Bayesian Statistics and Marketing. Chichester. UK: John Wiley & Sons. Stanisz, A., (2016). Modele regresji logistycznej: zastosowania w medycynie, naukach przyrodniczych i społecznych, Kraków: Wydawnictwo StatSoft Polska. Train, K. E., (2009). Discrete Choice Methods with Simulation. 2nd ed. Cambridge: Cambridge University Press.