scieee AI-readable full text Open interactive document viewer

How Robust are Country Rankings in Educational Mobility?

Engzell, Per

Full text

How Robust are Country Rankings in Educational Mobility? Ely Strömberg, Per Engzell Document type Post-print: This manuscript has passed peer review and been accepted for publication by a journal. It may include final edits by the author(s) but no editing or formatting by the publisher. Funding information This research was funded by the European Research Council, Grant Agreement No. 101165962 Markets and Mobility: How Employers Structure Economic Opportunity. Suggested citation Strömberg, Ely and Per Engzell. (2025). How Robust are Country Rankings in Educational Mobility? Sociological Science, forthcoming. Date of record October 16, 2025. Terms of use This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International license: https://creativecommons.org/licenses/by-nc-nd/4.0/ How Robust are Country Rankings in Educational Mobility? Ely Strömberg University of Amsterdam Per Engzell University College London October 16, 2025 Abstract We investigate the impact of analytical choices on country comparisons in intergenerational educational mobility using a multiverse approach. A literature survey gives rise to 2,880 plausible ways of measuring educational mobility, which we apply to European Social Survey (ESS) data from 16 countries. While some countries consistently appear at the top or bottom of the mobility rankings, most show substantial variation. Beyond its methodological contribution, we report two substantive findings. First, some countries often characterized as low-mobility emerge as matching or surpassing the egalitarian Nordic countries, reinforcing the view that wider mobility differences cannot be attributed solely to the education system but must be sought elsewhere, like the labor market. Second, the choice of parameter—such as regression coefficients, correlations, or categorical measures—is the single most influential factor that shifts country rankings. As different parameters carry distinct theoretical meanings, researchers should treat parameter choice not merely as a robustness check but as an opportunity to test and refine competing theories. Intergenerational educational mobility is a measure of the degree to which children’s educational attainment correlates with that of their parents. In a society with perfect educational mobility, children’s level of education is completely independent of that of their parents. The opposite of mobility is persistence, where the attainment of parents plays a significant role in the level of education reached by their children. Educational mobility is widely viewed as a desirable goal, being an indicator of equality of opportunity (Breen and Jonsson, 2005). Hence, a large literature attempts to understand its variation across place and time in an effort to identify institutional factors that promote or hinder that goal. Like in all social science, a challenge in this field is that small and seemingly inconsequential choices of data and specification can exert a large influence on results. This problem of researcher degrees of freedom is not new. In a now classic review on how parents influence their children’s attainments, Haveman and Wolfe (1995, p. 1855) noted how comparability was hampered by 1 substantial variation in estimation methods . . . distressingly small overlap in the explanatory variables included in the models . . . large variation in specification of the variables designed to indicate the same phenomenon (determinant) across the studies; . . . inconsistency among the studies in reporting effects for entire samples as opposed to specific subsamples of the population . . . and substantial variation across the studies in the specification of the outcome variable of interest. To achieve comparability, researchers often pare down the model to the simplest possible comparison: a single parameter of association between the same educational outcome for parent and child. Even so, choices abound. Should we measure education as years of schooling, credentials attained, or something else? If years of schooling, are we asking about the actual number of years spent studying or the minimum needed to reach a given level? If credentials, how do we reconcile the often large differences between education systems in different countries? Are we interested in the father’s education, mother’s education, or both? If both, how do we best combine information about the two? What parameter do we use to estimate the association? Even if we manage to fully harmonize data in a given comparison, there might always be a different set of choices that would have yielded a different ranking. Against this background, it becomes necessary to investigate to what degree analytical choices influence country comparisons in educational mobility. We approach this question through the method of multiverse analysis (Engzell and Mood, 2023; Young and Cumberworth, 2025). This method starts from the premise that many of the alternatives a researcher faces are equally plausible, and can all be justified post hoc. Pressured by theoretical expectations, there is a temptation to veer toward combinations of choices that yield a clean or easily interpretable result. A multiverse analysis removes this pressure by laying bare the full range of possible conclusions that emerge from a given dataset and question. Instead of trying to reach a single preferred specification, it records all alternative decisions and performs the analysis of interest under all possible combinations thereof. Much debate on cross-national differences in educational mobility has centered on the idea of a “Nordic exceptionalism.” Early comparative studies helped establish the view that the Nordic countries stand out for their egalitarian school systems and relatively high levels of educational mobility (Erikson and Jonsson, 1996; Shavit and Blossfeld, 1993). More recent work, however, has complicated this picture. While the Nordic countries consistently place high in studies of income mobility (Blanden, 2013; Bratberg et al., 2017; Engzell and Mood, 2023), evidence of a Nordic advantage in educational or occupational mobility is far less clear (Gregg et al., 2017; Karlson and Birkelund, 2024; Landersø and Heckman, 2017). This discrepancy, where strong income mobility is combined with middling educational or occupational mobility, has been described as a mobility paradox (Breen et al., 2016; Karlson, 2021), raising broader questions about how institutions shape intergenerational outcomes. We focus on the case of Europe, where one high-quality dataset, the European Social Survey (ESS), allows us to measure education in the same way across 16 societies. In constructing our multiverse, we carry out a systematic literature search 2 to derive relevant dimensions of variation. This lets us identify 2,880 plausible ways of measuring educational mobility, and examine how results vary across them. There is wide variation in estimated levels of mobility, as well as country rankings, across specifications. Our findings regarding a possible Nordic exceptionalism remain ambiguous: while Nordic countries tend to place in the upper half among the countries examined, so too do several countries more commonly characterized as low-mobility, such as the UK or Germany. A more consistent divide in our data is therefore between Western European countries that are more mobile, and their less mobile Southern and Central European counterparts. We also go further and test the stability of rankings in relation to various model components, asking which ones exert the greatest influence on results. It turns out that different parameters of association—regression coefficients, correlations, log odds ratios, and rank correlations—are the most important contributor and can dramatically shift the rank of individual countries. We argue that different parameters imply different constructs and speak to different questions, which suggests the need for theory in interpreting results. Other model components, such as gender differences, do not significantly affect mobility rankings despite being stressed in previous research. Strikingly, individual model components do not contribute much in isolation, implying that much variation stems from secondand higher-order interactions between model components which makes variation challenging to predict and interpret. These results call into question a simplistic view that conflates different dimensions of mobility (education, occupation, income) and treats them as primarily determined by educational systems. Even when focusing on educational mobility alone, model variation emerges as an irreducible source of uncertainty in country rankings. Crucially, this is not merely a statistical nuisance: rankings can shift substantially depending on the choice of association parameter, each of which carries distinct substantive meanings. We should not expect rankings to align across different dimensions of mobility, nor assume that country positions are stable across alternative operationalizations of one dimension. At a minimum, researchers should consider multiple parameters and, if favoring one, justify their choice theoretically. More ambitiously, variation across measures and models should be treated as informative in itself: an opportunity to explain their limited overlap and ultimately challenge or reformulate existing theory. Previous literature The question of robustness in educational mobility research was recently brought to the fore by a heated exchange comparing Denmark and the US. In both scholarly literature and public debate, Denmark is often portrayed as a model egalitarian welfare state, with substantial public investment aimed at equalizing opportunities from early childhood. The US, by contrast, is typically seen as a society of entrenched inequality, where life outcomes are largely determined at birth. A study by Landersø and Heckman (2017) questioned this conventional wisdom by claiming that the two countries are similar in their level of educational mobility. At stake are 3 important policy implications. If education is shown to drive economic mobility, it can be viewed as a successful policy lever for promoting equality of opportunity. If not, it suggests that achieving economic mobility may require other means, such as labor market reform or redistributive policies. Soon followed a rebuke and a subsequent exchange involving several other researchers (Andrade and Thomsen, 2018, 2021; Karlson, 2021; Thomsen et al., 2025), each presenting alternative analyses leading to different conclusions. But without a principled way of exploring the model universe, exchanges like this risk generating more heat than light. Each side can provide ever new analyses supporting their preferred conclusion, and defend the assumptions behind them. A multiverse analysis offers a principled way out: rather than making arbitrary analytical decisions, researchers identify all reasonable choices and evaluate them within a common framework. This approach not only makes explicit the extent of model dependence but also places individual results in the context of the broader distribution of possible outcomes. In doing so, it shifts the debate from defending single specifications to understanding patterns of variation across specifications and what this implies for theories of intergenerational mobility. While the US does not appear in our sample of countries, its closest European counterpart is arguably the UK. Both represent “liberal” welfare regimes (EspingAndersen, 1990), and combine high levels of social stratification with a formally open education system, in which the earliest formal branching occurs relatively late, in England with GCSEs around age 16 (Schneider, 2008a). Most children attend statefunded schools designed to accommodate students of all ability levels. Nonetheless, selective grammar schools continue to operate in some areas (Burgess et al., 2018), admitting students on the basis of an entrance exam at age 11, while fee-paying independent schools provide an alternative pathway for the wealthy (Henderson et al., 2020). Private schools, despite educating only a small minority, remain overrepresented among social elites (Reeves et al., 2017). At the same time, the view of the UK as a low-mobility regime stems less from comparative evidence on education than from cultural narratives, findings on income mobility, and portrayals of elite institutions. A more ideal-typical case of a rigid education system may therefore be Germany, which represents a “conservative” welfare regime (Esping-Andersen, 1990), with strong differentiation between academic and vocational tracks and where crucial educational decisions are made as early as age 10 (Schneider, 2008b; Van de Werfhorst and Mijs, 2010). In different ways, then, the UK and Germany provide strategic comparisons for testing the notion of Nordic exceptionalism: the UK as a culturally salient “class society” and the closest European analogue to the US, and Germany as a clearer case of low educational mobility, epitomized by a strong vocational–academic divide and rigid tracking that channels opportunities at an early age. Various studies have produced country rankings in educational mobility, covering 17 countries in Chevalier et al. (2009), 19 in Pfeffer (2008), 20 in Liu and Ding (2020), 42 in Hertz et al. (2008), and 185 in Narayan et al. (2018). A consistent finding is that mobility tends to be higher in more economically equal societies, such as the Nordic welfare states, and lower in less equal contexts, including parts of Southern and Central Europe and Latin America. At the level of individual 4 countries, however, results are less stable. Nordic countries usually appear among the more mobile, but Sweden ranks only mid-distribution in Hertz et al. (2008), while Norway is among the least mobile in Pfeffer (2008). Germany also shifts dramatically across studies, classified as highly immobile in Pfeffer (2008) yet among the most mobile in Liu and Ding (2020). The UK is often ranked somewhere around the middle (Chevalier et al., 2009; Liu and Ding, 2020), but its position, too, appears sensitive to methodological choices.1 A crucial but understudied question is how to construct the model space. If reproducing a study, one can combine the original model with those of published comments and rejoinders (Muñoz and Young, 2018), or investigate which models authors have used in their previous research in the same field (Steegen et al., 2016). One can also use own experience to map a set of reasonable analytical choices (Simonsohn et al., 2020b; Young and Holsteen, 2017). Other multiverse studies have taken a many-analysts approach, letting several research teams devise a way to answer the same question and then mapping the factorial combinations of all choices involved (e.g., Schweinsberg et al., 2021). Most of the examples above replicate one previous finding. Here our aim is not to replicate a given study, but to address a whole research field. To arrive at a plausible set of specifications, we therefore perform a systematic literature search to assess the range of model variation in the field. While our analysis focuses on model uncertainty, we do not wish to downplay other sources of variation. Estimation bias (e.g., due to missing data or measurement error) can be important, particularly when different data sources are used (Engzell and Jonsson, 2015). Sampling variation also matters, as each estimate is accompanied by a confidence interval that may depend on the chosen model and metric (Mogstad et al., 2024). We abstract from sampling error in our main analysis, as we believe there is value in demonstrating the influence of model variation independent of other forms of uncertainty. Supplementary analyses reported in a separate section below show that while sampling variation adds further uncertainty to country rankings, it is small compared to model variation and does not alter the overall patterns or conclusions of our study. Theoretical considerations Why, then, do different methods and models produce divergent results? A common view is to treat such variation as analogous to sampling error: mere statistical noise that weakens confidence in results. Yet this perspective is both unrealistic and limiting. Variation is not only inevitable but often theoretically meaningful: models do not simply estimate the same quantity with error, but rather capture different facets of the phenomenon. From this perspective, variation becomes a resource rather than a flaw. As Engzell and Mood (2023) argue, a theory-informed multiverse analysis can use this heterogeneity to clarify why results diverge and what those divergences reveal about the mechanisms of stratification. In studying educational mobility, one model may highlight returns to human capital (Becker and Tomes, 1979), another its signaling value (Arrow, 1973; Spence, 5 1973), and a third the positional nature of schooling (Hirsch, 1976; Shavit and Park, 2016). Looking beyond education, findings of a mobility “paradox” become less paradoxical once we recognize that education, occupation, and income capture distinct phenomena, each shaped by different processes. Occupational attainment reflects not only education but also school-to-work linkages, local industrial structures, the role of informal networks, and many other factors that vary across societies and over time (Bernardi and Ballarino, 2016). Income mobility, in turn, is additionally conditioned by wage structures within and across occupations, and features of the tax and transfer system (Breen et al., 2016; Landersø and Heckman, 2017). Although we stop short of a full theoretical exposition of our results, we use the remainder of this section to draw some relevant distinctions. Human capital theory views education as a quantitative asset of the individual, just like other forms of capital such as wealth. Asking respondents about the number of years they spent studying will produce a continuous measure of education that is easy to understand and use. In theory, the measure is comparable across countries, with one year of schooling carrying the same meaning (a full year spent studying) regardless of national education system. However, this assumes an equal increase of skills and knowledge for each additional year of schooling and that all types of skills hold similar value (Braun and Müller, 1997; Schneider, 2009), assumptions which can be questioned. An alternative way of measuring years of schooling is therefore to assign theoretical years based on the typical time necessary to reach a given qualification. In signaling theory, the education acquired is not so much a good in itself, as a signal of the underlying skills that allowed a person to reach a given level of education. To employers, the most readily observed signal is usally not years of schooling but an attained degree. In some systems, all students follow a straight path through programmes of increasing complexity, whereas in others there is a horizontal differentiation with multiple different tracks ranging from theoretical to vocational. Measuring these national levels or programmes can result in categories that are often quite heterogenous and country specific. To harmonize these on a common scale, international standards exist. One such standard is the International Standard Classification of Education (ISCED) which we will use below. Recent research argues that as a result of massive expansion of higher education, schooling increasingly serves as a positional good (Bol, 2015; Bukodi and Goldthorpe, 2016). Employers may judge applicants not on whether they have met some minimum qualification, but on their relative standing amongst applicants. The main gain from educational investment is then neither skills nor signals, but the positioning of oneself above one’s peers. This points to percentile ranking which situates people within the education distribution of their own age and cohort. Each of these constructs maps onto different ways of measuring education, which in turn imply different parameters, as we turn to next. 6 Measuring educational mobility Grounded in the above theoretical motivations, we estimate five different parameters of association: regression and correlation coefficients in years of schooling, log odds ratios between educational levels, unidiff coefficients that aggregate log odds ratios across several categories, and correlations in educational ranks. Starting with human capital theory, years of schooling are typically used to estimate a regression or correlation coefficient between parent and child. The regression coefficient is the slope βin a linear equation with child education on the left-hand side and the parent’s education on the right. Its interpretation is the additional years of child schooling associated with a one-year increment in parent schooling: Yt=α+βYt−1+ε. (1) In Equation 1, Ytrepresents the child’s years of schooling, Yt−1represents the parent’s years of schooling, αis a constant and εis a well-behaved error term. The regression coefficient βis given by: β=Cov(Yt, Yt−1) V ar(Yt−1)=Corr(Yt, Yt−1)pV ar(Yt) pV ar(Yt−1).(2) As Equation 2 makes clear, the regression coefficient depends on the relative dispersion of years of schooling in both generations. For example, if the variance in children’s years of schooling is greater than that among parents, a given correlation will imply a higher regression coefficient β. The process of educational expansion typically increases dispersion at first, as new opportunities open up, and then decreases it, as educational attainment reaches a new plateau at higher levels. To avoid conflating changes of composition with the dependence structure between parent and child, one may want to factor out the relative dispersion in both generations (Black and Devereux, 2011, p. 1504). One way to do so is by the intergenerational correlation in years of schooling: r=Corr(Yt, Yt−1) = Cov(Yt, Yt−1) pV ar(Yt)pV ar(Yt−1).(3) Treating years of schooling as continuous assumes that the relationship between parent’s and child’s education is monotonic and linear (Blanden, 2013, p. 44). This can be questioned, especially in complex education systems with multiple paths. Sociologists have long advocated for conceptualizing education as a series of discrete transitions (Breen and Jonsson, 2000; Mare, 1980). An alternative way is to estimate probabilities to reach a certain educational level given a level of parental education and without assuming linearity of effects between the different levels. One such measure is the log odds ratio, which is a function of conditional associations between levels of parent and child education. The log odds can be represented as: ln ORij,i′j′= ln pij/pi′j pij′/pi′j′= ln pijpi′j′ pij′pi′j,(4) 7 contrasting, for example, the probability pfor a university educated parent’s child jof completing university education ias opposed to a lower level i′, relative to the corresponding probability for a child of parent who has less than university education j′. A categorical model takes degrees as the outcome and therefore corresponds closest to the signalling view of education. But credentials may also reflect relevant skills or relative standing, as per the human capital or positional good theories.2 One issue with categorical measures is that, as soon we deal with a mobility table composed of more than two categories, the number of possible comparisons grows large. This yields several disparate coefficients instead of a single summary measure. To address this, researchers often employ log-multiplicative models known as “unidiff” or uniform difference models (Erikson and Goldthorpe, 1992; Xie, 1992). Such models impose a proportionality constraint on how log-odds ratios vary across tables (e.g., countries). Specifically, it assumes that the pattern of relative mobility captured by the log-odds ratios is structurally invariant, differing only by a scalar multiplier that varies between contexts and reflects the overall level of mobility. In the simplest form, this includes a term for excess cases along the diagonal, reflecting immobility between parent and child categories, while allowing the relative mobility level to vary across countries. A positional view of education instead implies that the rank-order correlation should be the preferred measure. This parameter is equivalent to the correlation coefficient rabove, with the distinction that each variable, instead of measured as years of schooling, is transformed to percentile ranks: ρ=Corr(Rt, Rt−1) = Cov(Rt, Rt−1) pV ar(Rt)pV ar(Rt−1),(5) where R(Y)∈ {1,2,...,100}represents the rank transform. The distribution of percentiles is uniform, so this transformation also equalizes the dispersion of education in both generations. Hence, the distinction between regression and correlation coefficients that arises with years of schooling does not matter here. Mapping the literature In addition to parameter selection, many other choices need to be made. To map this space of variation, we conduct a literature search. We searched Scopus and Web of Science, looking for peer-reviewed articles published in the last twenty years. Specifically, we used the following search terms for Scopus: (TITLE ("EDUCATIONAL FLUIDITY") OR TITLE ("EDUCATIONAL MOBILITY") OR TITLE ("EDUCATIONAL TRANSMISSION") OR TITLE ("INTERGENERATIONAL TRANSMISSION OF EDUCATION") AND PUBYEAR > 2002) AND (LIMIT TO (DOCTYPE, "AR")) And for Web of Science: 8 Correlations Figure 2, left panel, shows the dispersion of correlations by country. The correlation models work the same way as those for regression coefficients, but both respondent and parental education are standardized to factor in educational expansion between parental and respondent generations. Distributions still overlap to a high degree, and values range from 0.16 for Denmark to 0.67 for Hungary. Compared with its low minimum estimate, Denmark has a mean estimate at the higher 0.31 (Appendix Table A4). which is the smallest mean value followed by Germany and the UK, a clear change from the regression models. If we consider the distribution of coefficients as a measure of the importance of model choice, correlation results are comparable to regression coefficient models with an average IQR at 0.08 and an average range at 0.28, meaning model choice can change comparisons and perception of country mobility levels. Finland has a remarkably high range of 0.37, with values quite evenly spread. Some countries such as Sweden have a greater range than in the regression coefficient models, while others such as Denmark have a narrower one. Figure 2, right panel, shows a box plot of the ranking of countries by correlation. Compared with the regression models, there is less grouping in the more mobile half of the distribution. Denmark now seems to be the most mobile with a median ranking of 2, and Sweden has dropped from first to the middle of the ranking distribution. Germany has moved up to second highest median rank, from rank 9– 10 in regression models, appearing more mobile when using correlations. Belgium, Poland, Hungary, and Slovenia are still in the bottom for median ranks, while Spain appears more mobile. Only Finland is ranked both highest and lowest. 8 out of 16 countries can be ranked as first by choosing specific models. Sweden and Finland, which were grouped in the top with regression models, are now in the middle of the ranking distribution, which suggests that the distribution is more dispersed in the parent generation, and has subsequently contracted with mass higher education in the child generation.9 Log odds ratios Figure 3, left panel, shows the distributions of estimates from categorical models with log odds for tertiary education by parental tertiary education. As with the previous models, estimates overlap to a high degree, especially in the top half. However, ranges are visibly smaller for many countries, with highly dense distributions for Belgium and Poland. For Germany, the Netherlands, and Switzerland, distributions are instead non-normally distributed, something which might imply a sensitivity to certain analytical choices. Estimates range from 0.72 (Sweden) to 3.79 (Switzerland) with much overlap in the top half (Appendix Table A5). Estonia has the smallest range at 0.68 and Switzerland the greatest at 2.69, with the average at 1.30. The average IQR is at 0.29 so the impression of a country’s mobility level might not change much over models, but seeing the high overlap, it can still change country comparisons. Figure 3, right panel, shows a box plot of the ranking of countries over categorical models. The overlap in the top half of the left panel is here visible as clear 15 Poland Hungary Ireland Spain France Belgium Switzerland Slovenia Netherlands Norway United Kingdom Estonia Denmark Finland Germany Sweden 1234 Coefficient Poland Hungary Ireland Spain France Belgium Switzerland Slovenia Netherlands Norway United Kingdom Estonia Denmark Finland Germany Sweden 1 3 5 7 9 11 13 15 Rank Figure 3: Log odds ratios, distribution of coefficients and country rankings. Note: Countries are ordered by the median country rank. 508 models per country. overlap in rankings of the top 7 countries. Sweden now appears the most mobile, trailed by Germany, with a median rank of 2 and 3, respectively. We can note that Sweden’s position when calculating log odds ratios is similar to that from the regression coefficient models, while Germany’s position instead is similar to that of the correlation models. In the bottom we find Poland and Hungary. Slovenia has however risen towards the middle of the ranking distribution, with a median rank of nine and outliers at 1 and 2. As we can see in the left figure, Slovenia and Switzerland show much dispersion in log odds ratios, while other countries such as Belgium and Poland show very small dispersion. Switzerland is ranked as both 2 and 16 in different models, but otherwise ranges of ranks are smaller than in the previous models, and there is a visible order to the country rankings of the bottom half, even though overlap is still high. Unidiff coefficients Figure 4, left panel, shows the distributions of estimates from unidiff models, standardized to have mean 0 and standard deviation 1 across each specification.10 As with the previous models, estimates overlap, especially in the middle. However, there is a clear order of countries, with some having very tight distributions, especially Belgium and Ireland. The UK, Sweden, the Netherlands, and Germany show highly nonnormal distributions, implying sensitivity to analytical choices. Estimates range from −2.38 for Sweden to 2.32 for Hungary with much overlap in the top half (Appendix Table A6). Hungary has the smallest range at 1.26 and Sweden the greatest at 3.03, with the average at 2.05. The average IQR is at 0.47 with the highest being Sweden at 0.83, and the lowest France and Hungary at 0.31. Figure 4, right panel, shows a box plot of the ranking of countries over categorical models. The UK now appears the most mobile, trailed by Finland and Sweden, both 16 Poland Hungary Slovenia Switzerland Ireland Germany Norway Belgium Estonia France Netherlands Spain Denmark Sweden Finland United Kingdom −2 0 2 Coefficient Poland Hungary Slovenia Switzerland Ireland Germany Norway Belgium Estonia France Netherlands Spain Denmark Sweden Finland United Kingdom 1 3 5 7 9 11 13 15 Rank Figure 4: Unidiff, distribution of coefficients and country rankings. Note: Countries are ordered by the median country rank. 482 models per country. with a median rank of 3. Hungary and Poland are clearly placed in the bottom, followed by Switzerland and Slovenia that however both have long tails with many outliers. Germany, who was ranked second in the log odds models, is now ranked 11, but with a range from 1 to 14. Sweden, ranked third, and Slovenia, ranked 14, have the same range, pointing to all three countries being sensitive to analytical choices in unidiff analyses. Overall, many countries show a wide range when taking outliers into account, but with otherwise quite dense distributions. Therefore, there appears to be a clearer ranking than for other parameters, implying that unidiff analyses might be less sensitive to the analytical choices in our multiverse. Although we do not examine how unidiff results depend on the granularity of categories, supplementary analyses in Karlson (2021) find that country rankings are not overly sensitive to this analytical choice either, consistent with our conclusions. Rank correlations The last model type is positional educational attainment, where regression coefficients are calculated based on relative educational attainment instead of absolute. Figure 5, left panel, shows a ridge plot of regression coefficient values from the positional models. Looking at the top we find the UK with a mean of 0.28, coming out as the most mobile (Appendix Table A7). The other countries in the top half overlap to such a degree that it is hard to distinguish which is more mobile, however the distribution of Finland stands out with an exceptionally flat distribution with a range of 0.34 and the lowest estimate at 0.17. Hungary has the highest value at 0.64 and is again found in the bottom together with Poland and Slovenia. The average range is 0.26 and the average IQR is 0.08, meaning that using relative education does not remove the effect of analytical choices. Figure 5, right panel, shows country rankings under positional models. The 17 Hungary Poland Slovenia Belgium Estonia Ireland Spain Sweden Switzerland Finland Netherlands Denmark Norway Germany France United Kingdom 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Coefficient Hungary Poland Slovenia Belgium Estonia Ireland Spain Sweden Switzerland Finland Netherlands Denmark Norway Germany France United Kingdom 1 3 5 7 9 11 13 15 Rank Figure 5: Rank correlations, distribution of coefficients and country rankings. Note: Countries are ordered by the median country rank. 528 models per country. pattern diverges sharply from absolute measures: the UK holds the highest median rank, followed by France, an unexpected result given that France has consistently ranked 7th or lower using absolute education. Sweden and Estonia drop to the bottom half (9th and 12th, respectively), while Slovenia, Poland, and Hungary remain at the bottom. Spain also ranks low (11th), Switzerland sits in the middle (8th), and three countries span the full range from first to 15th or 16th. These shifts highlight how positional measures can drastically reshape country rankings, supporting Bol’s (2015) argument that relative education differs from absolute education and Bukodi and Goldthorpe’s (2016) finding that the effects of expansion depend on whether education is treated as positional. Stability of rankings So far, we have treated the five parameters separately under the assumption that they tap different constructs, and should not be compared on the same scale. Conflating the different parameters leads to even greater uncertainty in rankings, as shown in Appendix Figure A6. In this section, we formalize this notion and also test how variation in rankings due to parameter choice compares to that stemming from other model components. To quantify stability, we use the coefficient of determination (R2) in a leastsquares regression with rank (1,2,...K) as the outcome and Kcountry indicators as regressors. Intuitively, this yields a measure that equals 0 if rank distributions for countries are completely overlapping, and 1 if they are completely separated, that is, if the rank of a country does not vary across specifications. We refer to this measure as rank stability. To further understand the proportion of variation in rankings due to different model components, we calculate the same measure using country indicators interacted with a given component, such as the parameter of 18 0.55 0.56 0.56 0.58 0.59 0.73 0.54 Split by: Foreign born incl? Which parent? Education variable Gender Age restriction Parameter Baseline 0.00 0.25 0.50 0.75 1.00 Rank stability Figure 6: Rank stability. Note: The figure shows rank stability at baseline and split by model components. association or the respondent’s gender. Intuitively, this tells us what the aggregate rank stability is within different values of each component, for example within a given parameter or within the groups of men and women. We first confirm that the parameter of association is indeed the most important contributor to variation in rank across countries, in Figure 6. Pooling all parameters of association together (regression coefficients, correlations, log odds, unidiff, rank correlations), the rank stability coefficient amounts to 0.54. In other words, 54% of variation in ranks occurs between countries, and the rest across different specifications within countries. Adding the parameter of association to the rank regression, thereby moving this factor from the withinto the between-variation, pushes the rank stability up to 0.73. In other words, agreement about ranks is markedly higher within individual parameters than it is taken across the multiverse as a whole. No other single model component makes a comparable contribution to rank stability. Accounting for any other component in the model pushes the rank stability coefficient no higher than 0.60 (Figure 6). In supplemental materials, we compare rankings using different parameters: how consistent they are with each other, and which parameters yield rankings that deviate more or less from the overall pattern. Appendix Table A8 shows the correlation matrix of rankings between parameters and Appendix Table A10 shows the factor loadings for different parameters on a shared, underlying dimension. Correlations and rank correlations yield rankings closest to the common underlying factor, with factor loadings of 0.88 and 0.85, respectively. For remaining parameters (regression coefficients, log odds ratios, unidiff coefficients) the factor loadings are weaker, and all in the range 0.65–0.73. These measures produce rankings that are more distinct from the overall pattern across all rankings. Rank stability varies considerably between parameters, as the values labeled “baseline” in Appendix Figure A7 show. It is lowest for regression coefficients, correlations, and rank correlations, all at 0.67–0.68. It is higher for categorical 19 measures such as the log odds ratio (0.80) and unidiff coefficient (0.84). In other words, countries display more similar levels of mobility when using linear measures of association, and appear more distinct when using categorical measures.11 Appendix Figure A7 also shows how rank stability improves further when splitting the analysis by any of the other model components. None of the components contribute to rank stability by more than 16%, the largest improvement coming from age restriction in Appendix Figure A7, top right panel (0.78/0.67 = 1.16). One implication of this is that most model variation depends on higher-order interactions and is therefore hard to predict. Notably, gender differences are not a major contributor to country variation in mobility rankings. This implies that rankings are similar whether we study men, women, or both. At the face of it, this runs counter to the findings of Engzell and Mood (2023) where gender emerges as the main organizing frame for differences in income mobility trends. They argue that gender dynamics are “a force powerful enough to be the main driver of trends and differences across countries, and cannot be ignored” (p. 619). However, this discrepancy is unsurprising in light of the fact that we study education. As Engzell and Mood (2023) show, their findings are driven by gender differences in the economic returns to education, rather than differences in the importance of family background for education. Gender may still matter for mobility in other outcomes such as occupation or earnings, or in education for older cohorts where women had yet to reach parity with men. Incorporating sampling error So far, we have abstracted from the issue of sampling error. Model and sampling variability can be addressed jointly, for example through bootstrapping or similar approaches that generate empirical distributions incorporating both sources of uncertainty (Ibarra et al., 2025; Simonsohn et al., 2020a). We believe there are three issues at stake. First, is model variation large or small compared to the sampling variation covered in most extant research? Second, are our estimates of model variation confounded by sampling variation? This is a subtler point: because the underlying data change from one specification to the next, we would expect some degree of uncertainty simply from this aspect of the analysis. Third, how much do results shift when we account for both sources of variability together? We take each question in turn. To assess the size of the two sources of variation, we compute two statistics at the parameter–country level: the mean standard error within a country (sampling variation) and the standard deviation of point estimates across specifications (its empirical analogue for model variation). This comparison is limited to the four parameter types that produce parametric standard errors by default: regression coefficients, correlations, log odds ratios, and rank correlations. Results are summarized in Appendix Figure A8. Across the board, sampling uncertainty is smaller than model uncertainty, often by a wide margin. For regression coefficients, correlations, and rank correlations, the standard deviation of point estimates is about four times the mean standard error; for log odds ratios, the ratio is roughly two-toone. These findings reinforce our view that model uncertainty, often overlooked, is 20 a dominant source of variation that researchers ignore at their peril. To probe whether our results are confounded by sampling error, we illustrate how estimates and ranks vary purely as a function of sampling. We arbitrarily select one specification for each parameter (regression, correlation, odds ratio, unidiff, and rank correlation).12 For each specification within that parameter, we then draw 250 bootstrap samples per country using ESS weights, and report the results in the same format as our main figures. As shown in Appendix Figures A9–A13, parameter estimates and rankings fluctuate much less when variation is induced solely by sampling uncertainty, demonstrating that sampling error alone does not explain model variation. Finally, we assess how results change when sampling and model variation are considered together. Because this involves a larger number of models, instead of computationally intensive bootstrapping, we use parametric standard errors to simulate plausible values. Specifically, for each specification we draw 250 values from a normal distribution centered on the point estimate, with a standard deviation equal to its reported standard error. Appendix Figures A14–A17 present the results. The distributions resemble those obtained when ignoring sampling variability, because the additional noise from sampling uncertainty largely overlaps with the variation already introduced by model choice. For country rankings, we observe an increase in outliers, but the overall ordering and broad conclusions remain the same. Thus, while sampling error adds further uncertainty, it does not alter our substantive findings. Discussion: Nordic exceptionalism in a new light? How well do our results align with stylized facts in the field? There are at least two ways to look at this. First, we can ask whether the variation we document is artificially suppressed in published studies. Conventional publications often present a single model specification—chosen for theoretical or pragmatic reasons—as if it were definitive. This practice may obscure the extent of disagreement across plausible alternatives, giving the impression of greater certainty than actually exists. Second, we can ask whether, despite this variation, the broad ordering of countries in our results is nevertheless consistent with what other studies have found. If our results diverge sharply from published findings, this may in turn suggest that what is perceived as common knowledge may partly reflect selective reporting of specification choices. To first examine the range of variation, we turn to the case of Denmark that we have used to motivate our analysis. Figure 7 shows a density curve of our regression coefficient estimates for Denmark (containing the same information as the ridge plot for Denmark in Figure 1). Overlaid are point estimates from Andrade and Thomsen (2018) who find regression coefficients of 0.37 looking at fathers and 0.42 looking at both parents, Hertz et al. (2008) who puts the regression coefficient at 0.49, and from Liu and Ding (2020) who found estimates of 0.20 for mothers only, and 0.23 for fathers only. Our distribution has two clusters around 0.25 and 0.37, which corresponds quite well to paternal estimates by Liu and Ding (2020) and Andrade 21 .37 A&T 2018 Fathers .42 A&T 2018 Parents .49 Hertz et al 2008 .20 L&D 2020 Mothers .23 L&D 2020 Fathers 0 1 2 3 4 5 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Coefficient size Density Figure 7: Kernel density plot of the distribution of regression coefficient values for Denmark across all models. Estimates of regression coefficients for Denmark from previous research estimates overlaid as dashed lines. and Thomsen (2018). The average point estimate in our analyses lies around 0.32, and while that does not correspond to any previous estimate, they all fall within the multiverse distribution, which would point to the analysis capturing the analytical choices of previous estimates quite well. The estimate by Hertz et al. (2008) falls at the right end of the distribution.13 Should this variation be seen as large or small? Previously, Karlson (2021, p. 349) has argued that a difference between US and Danish regression coefficients of 0.04 (as found by Andrade and Thomsen 2018) is not to be considered substantial, comparing it with differences in the magnitude of 0.27 in income elasticity. We can compare these differences with the overall dispersion of estimates for one country in our study. For example, the interquartile range of regression coefficients for Denmark amounts to 0.13 (the average IQR across all countries is 0.09). This is certainly greater than the difference between Denmark and the US of 0.04, and points to limited robustness of educational mobility point estimates. One could therefore argue that researchers should not draw too big conclusions from educational mobility comparisons when differences are of this magnitude, not because they are smaller than income elasticity differences, but because the difference could easily stem from model construction. What does the ranking of countries in our study say about a possible Nordic exceptionalism? If we combine all rankings, as shown in Appendix Figure A6, the Nordic countries are ranked as second to sixth most mobile, with Denmark, Sweden, and Finland having a median ranking of 4 across all models. On average, then, our results confirm the stylized fact that the Nordics tend to cluster in the upper half of the mobility distribution.14 Whether occupying the top half of the distribution 22 qualifies as “exceptional” is a separate matter that depends on whether exceptionalism is understood as being consistently at the top, or simply as performing better than most other countries. At the same time, we present robust results that qualify some common assumptions about how institutions shape intergenerational mobility. The UK and Germany are often classified as low-mobility regimes, due to a selective private sector in the UK and early educational tracking in Germany. Yet in our analyses, both countries appear among the more mobile, matching or even surpassing the levels observed in the Nordic countries. The most striking case is the UK, which across all specifications emerges as the single most mobile country, with a median ranking of 3. This is unexpected: no previous ranking has placed the UK at the very top, and it has alternated between the top and bottom halves depending on the study. While our results place it first overall, the wide range of 13 indicates that very high and very low rankings are both possible depending on analytical choices. Germany also performs better than often assumed, consistently appearing in the top half of the distribution, with a median rank of 5 across all models (Appendix Figure A6), despite having been placed near the bottom in some earlier work (Chevalier et al., 2009; Pfeffer, 2008). Have the UK and Germany been mischaracterized in earlier literature? A closer reading suggests that some of these patterns have been hiding in plain sight. For example, Gregg et al. (2017) show that differences in income mobility across Sweden, the UK, and the US cannot be explained by disparities in educational attainment, which are strikingly similar across the three. Likewise, Pensiero and Barone (2024) demonstrate that the intergenerational persistence of education has declined in the UK, placing it on par with the Nordics. In the German case, the apparent contradiction between relatively high measured mobility and the expectation of rigidity under early tracking has long drawn attention, with scholars such as Schneider (2009, 2010) highlighting the mismatch and proposing revisions of the ISCED classification to better capture theoretical intuitions. These examples suggest that the anomalies are not entirely absent from previous research; rather, they are acknowledged but often fail to shift the prevailing consensus. The issue is not so much selective reporting as the subjective updating of what is taken to be common knowledge in the field. Here lies one of the advantages of a multiverse approach: by presenting the full range of plausible results, it becomes harder to dismiss unexpected findings as statistical outliers. Instead, we are confronted with genuine theoretical puzzles that demand explanation. Conclusion This study has shown that estimates of intergenerational educational mobility, and the country rankings derived from them, are highly sensitive to analytical choices. Using a multiverse of 2,880 specifications applied to ESS data from 16 European countries, we find that while some patterns are consistent—Nordic, UK, and German cases often placing in the upper half, and Southern and Central European countries in the lower half—rankings for many countries vary widely, sometimes 23 spanning both top and bottom positions. Most importantly, the choice of association parameter is the single factor with the greatest impact on results, underscoring that different measures capture distinct theoretical constructs rather than interchangeable versions of the same phenomenon. These findings have two implications. Substantively, they challenge the idea of a stable “Nordic exceptionalism” in educational mobility and point instead to a broader West–East/South divide in Europe, with economic mobility shaped by institutions beyond schooling alone. Methodologically, they caution against presenting single estimates as definitive. Analytical variation is not just noise but reflects theoretical disagreements about how education operates: as human capital, positional good, or signal. Future research should therefore (1) treat parameter choice as a substantive theoretical decision, (2) consider multiple specifications rather than relying on a single preferred model, and (3) use variation across models as an opportunity to refine theory rather than dismissing it as uncertainty. In this way, multiverse analysis moves us beyond robustness checks toward a more transparent and theoretically grounded understanding of mobility processes. We have emphasized the need for theory in adjudicating between rankings. Does this imply that if choices of measurement are systematically guided by a single theoretical orientation, results will become more robust? We are skeptical. Theory can help explain why rankings diverge, but it cannot resolve the fact that different frameworks capture distinct mechanisms, all of which are relevant. The obvious answer to whether education functions as human capital, signal, or positional good is: all of the above. Thus, while theory-guided choices may reduce “within-framework” variation, they cannot eliminate disagreement across frameworks. We therefore see the value of multiverse analysis not in collapsing toward a single “true” ranking, but in making explicit how results hinge on theory, clarifying which conclusions are robust across orientations and where genuine theoretical debate remains. This also highlights the limitations of a common practice in sociological research: selecting one mobility measure, often justified on theoretical grounds, and assessing robustness only by varying sample definitions or control variables. Our results suggest that this strategy risks obscuring how different measures embody distinct assumptions and capture different facets of mobility. Future work should treat the choice of measure not as a purely technical decision but as a substantive one, explicitly tied to the dimension of mobility under study. Multiverse analysis offers a way forward: rather than asking whether results are robust to small perturbations of a preferred model, researchers can examine how findings vary across theoretically informed alternatives. This shift encourages more cautious interpretation of single estimates while enriching theoretical dialogue by showing which findings hold consistently and which depend on particular modeling choices. Some variation can be understood as reflecting real theoretical puzzles, pointing to opportunities for deeper investigation. Nevertheless, considerable variation in rankings is likely to remain unexplained. Should these results then lead us to rethink the literature? On the one hand, the sheer range of variation is hard to dismiss. For countries that place in the middle range, it becomes possible to pick specifications to support almost any conclusion. Even for countries that place closer to one extreme of the distribution, there is room for considerable disagreement. On 24 Schweinsberg, Martin, Michael Feldman, Nicola Staub, Olmo R van den Akker, Robbie CM van Aert, Marcel ALM Van Assen, Yang Liu, Tim Althoff, Jeffrey Heer, Alex Kale, et al. 2021. “Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis.” Organizational Behavior and Human Decision Processes 165:228– 249. Shavit, Yossi and Hans-Peter Blossfeld (eds.). 1993. Persistent Inequality: Changing Educational Attainment in Thirteen Countries. Boulder, CO: Westview Press. Shavit, Yossi and Hyunjoon Park. 2016. “Introduction to the Special Issue: Education as a Positional Good.” Research in Social Stratification and Mobility 43:1–3. Simonsohn, Uri, Joseph P. Simmons, and Leif D. Nelson. 2020a. “Specification Curve Analysis.” Nature Human Behaviour 4:1208–14. Simonsohn, Uri, Joseph P. Simmons, and Leif D. Nelson. 2020b. “Supplement: Specification Curve Analysis.” Nature Human Behaviour 4:1208–14. Spence, Michael. 1973. “Job Market Signaling.” Quarterly Journal of Economics 87:355–374. Steegen, Sara, Francis Tuerlinckx, Andrew Gelman, and Wolf Vanpaemel. 2016. “Increasing Transparency Through a Multiverse Analysis.” Perspectives on Psychological Science 11:702–12. Thaning, Max and Martin Hällsten. 2020. “The end of dominance? Evaluating measures of socio-economic background in stratification research.” European Sociological Review 36:533–547. Thomsen, Jens-Peter, Stefan Andrade, Florian Hertel, Øyvind Wiborg, and Max Thaning. 2025. “Intergenerational Educational Mobility in Scandinavia and the U.S.” Mimeo. Van de Werfhorst, Herman G and Jonathan JB Mijs. 2010. “Achievement inequality and the institutional structure of educational systems: A comparative perspective.” Annual Review of Sociology 36:407–428. Xie, Yu. 1992. “The log-multiplicative layer effect model for comparing mobility tables.” American Sociological Review pp. 380–395. Young, Cristobal and Erin Cumberworth. 2025. Multiverse Analysis: Computational Methods for Robust Results. Cambridge University Press. Young, Cristobal and Katherine Holsteen. 2017. “Model Uncertainty and Robustness: A Computational Framework for Multimodel Analysis.” Sociological Methods & Research 46:3–40. 31 Table 1: Variation in analytical choices from article search. Dimension Alternatives Sample selection Age ranges and respective countries in which they were used: 24–65 (GR), 25–35 (2x Europe & CA), 25–65 (Europe), 25–69 (IS), 25–70 (2x CA), 25–74 (CA), 25+ (Europe), 27–46 (AT) 27–47 (SE), 28–55 (FR), 29–33 (DK & US), 30+ (DK). 4 articles stated that they excluded foreign-born respondents. 3 articles looked only at sons, 4 at only daughters, 18 looked at both separately, 8 looked at both separately and pooled. Child education 9 articles used actual years of schooling, 7 used years of schooling calculated from highest degree attained, 3 used dichotomous measures of a yes/no nature, 15 used a categorical measure of highest degree attained, 2 used highest degree calculated from years of schooling. Parent education 6 articles used actual years, 6 recoded level to theoretical years, 2 recoded ISCED to theoretical years, 3 used dichotomous measures, 11 used levels of education, 5 recoded actual years to levels, 1 used ISCED. Which parent? 10 articles used dominant parental education, 7 used only fathers’ education, 3 used only mother’s education, 4 used an average measure, 6 used both in the same analysis, 8 used both in separate analyses. Parameter 11 articles presented regression coefficients, 7 presented regression coefficients and correlations, 4 presented odds ratios, 2 presented risk ratios, 9 presented transitions matrices, 5 presented predicted probabilities, and 5 articles used other types of measures not common in the literature. 32 Table 2: Variation in analytical choices considered in this study. Dimension Alternatives Sample selection Studied countries were: Belgium, Denmark, Estonia, Finland, France, Germany, Hungary, Ireland, the Netherlands, Norway, Poland, Slovenia, Spain, Sweden, Switzerland, and the United Kingdom. Age ranges were: 24–65, 25–35, 25–65, 25–69, 25– 70, 25–74, 25+, 27–46, 27–47, 28–55, 29–33, 30+. Exclusion or inclusion of foreign-born respondents was varied. Analyses run on men and women separately. Child education Measured with: Actual years of schooling top coded at 25 or 30 years. Theoretical years of schooling calculated from ISCED levels. Relative actual years of schooling based on top-coded absolute measures. Relative theoretical years of schooling calculated from ISCED-levels. Primary/secondary/tertiary calculated from ISCED-levels and ES-ISCED levels. Parent education Measured with: Theoretical years of schooling calculated from ISCED-levels. Primary/secondary/tertiary calculated from ISCED-levels and ES-ISCED levels. Relative theoretical years calculated from ISCED-levels. Which parent? For continuous measures only father, only mother, average, and dominant parental attainment. For categorical measures: only father, only mother, and dominant parental attainment. Parameter Regression coefficients, correlations, log odds, unidiff, rank correlations. 33 Table 3: Descriptive statistics of recoded educational variables. New variable Based on Range Mean N Rounds Actual years of schooling (max 25) eduyrs 0–25 12.8 294,945 1–10 Actual years of schooling (max 30) eduyrs 0–30 12.9 295,722 1–10 Theoretical years of schooling edulvla & edulvlb 0–24.25 12.7 286,418 1–10 Father’s theoretical years of schooling edulvlfa & edulvlfb 0–24.25 10.5 243,235 1–10 Mother’s theoretical years of schooling edulvlma & edulvlmb 0–24.25 9.9 252,697 1–10 Average parental relative theoretical years of schooling Theoretical years father & mother 0–24.25 10.1 262,925 1–10 Dominant parental theoretical years of schooling Theoretical years father & mother 0–24.25 11.1 262,925 1–10 Educational level (ISCED) edulvla & edulvlb Prim/Sec/Tert 295,152 1–10 Educational level (ESISCED) eisced Prim/Sec/Tert 298,445 1–10 Father’s educational level (ISCED) edulvlfa & edulvlfb Prim/Sec/Tert 239,717 1–10 Mother’s educational level (ISCED) edulvlma & edulvlmb Prim/Sec/Tert 246,734 1–10 Dominant parental educational level (ISCED) edulvl3f & edulvl3m Prim/Sec/Tert 256,628 1–10 Father’s educational level (ES-ISCED) eiscedf Prim/Sec/Tert 165,070 4–10 Mother’s educational level (ES-ISCED) eiscedm Prim/Sec/Tert 170,988 4–10 Dominant parental educational level (ES-ISCED) eisced3f & eisced3m Prim/Sec/Tert 175,221 4–10 Relative actual years of schooling (max 25) Actual years of schooling (max 25) 1–100 46.0 294,945 1–10 Relative actual years of schooling (max 30) Actual years of schooling (max 30) 1–100 46.0 295,722 1–10 Relative theoretical years of schooling Theoretical years of schooling 1–100 44.2 286,418 1–10 Father’s relative theoretical years of schooling Father’s theoretical years of schooling 1–100 40.8 243,235 1–10 Mother’s relative theoretical years of schooling Mother’s theoretical years of schooling 1–100 40.1 252,697 1–10 Average parental relative theoretical years of schooling Average parental theoretical years of schooling 1–100 42.2 262,925 1–10 Dominant parental relative theoretical years of schooling Dominant parental theoretical years of schooling 1–100 41.8 262,925 1–10 34 Supplemental Material for: How Robust are Country Rankings in Educational Mobility? Contents 1 Descriptive statistics and coding 2 2 Separate results by gender 7 3 Stability and overlap across parameters 12 4 Sampling vs model uncertainty 15 4.1 Single-model results with bootstrap . . . . . . . . . . . . . . . . . . 16 4.2 Multiverse results with pseudo-bootstrap . . . . . . . . . . . . . . . 21 5 Result of article search 23 1 1 Descriptive statistics and coding 2 Table A1: Descriptive statistics on ESS data, unweighted. Round BE CH DE DK EE ES FI FR HU IE NL NO PL SE SI UK 1 N 1899 2040 2919 1506 1729 2000 1503 1685 2046 2364 2036 2110 1999 1519 2052 Response rate 59.2 33.5 55.7 67.6 - 53.2 73.2 43.1 69.9 64.5 67.9 65.0 73.2 69.5 70.5 55.5 Average Age 44.8 47.6 47.3 46.4 48.6 45.6 47.3 46.1 45.7 48.1 45.8 42.9 46.3 44.4 48.6 2 N 1778 2141 2870 1487 1989 1663 2022 1806 1498 2286 1881 1760 1716 1948 1442 1897 Response rate 61.2 48.6 51.0 64.2 79.1 54.9 70.7 43.6 65.9 62.5 64.3 66.2 73.7 65.4 70.2 50.6 Average Age 45.2 48.1 46.8 46.9 47.2 45.1 47.3 48.9 46.6 48.0 49.4 45.5 42.1 46.9 45.4 47.9 3 N 1798 1804 2916 1505 1517 1876 1896 1986 1518 1800 1889 1750 1721 1927 1476 2394 Response rate 61.0 51.5 54.5 50.8 65.0 65.9 64.4 46.0 66.1 56.8 59.8 65.5 70.2 65.9 65.1 54.6 Average Age 46.1 49.9 48.0 49.6 47.5 45.9 48.4 48.1 51.1 46.3 48.9 45.6 43.7 46.9 46.4 49.5 4 N 1760 1819 2751 1610 1661 2576 2195 2073 1544 1764 1778 1549 1619 1830 1286 2352 Response rate 58.9 49.9 48.0 53.9 57.4 66.8 68.4 49.4 61.3 51.6 49.8 60.4 71.2 62.2 59.1 55.8 Average Age 46.5 48.6 49.0 49.3 47.8 46.8 48.0 48.7 47.8 47.6 49.3 45.8 44.6 47.6 46.6 49.1 5 N 1704 1506 3031 1576 1793 1885 1878 1728 1561 2576 1829 1548 1751 1497 1403 2422 Response rate 53.4 53.3 30.5 55.4 56.2 68.5 59.5 47.1 49.2 65.2 60.0 58.0 70.3 51.0 64.4 56.3 Average Age 46.8 47.8 47.6 48.5 48.7 45.9 48.8 49.4 47.6 46.1 50.4 46.4 44.4 48.6 47.4 49.9 6 N 1869 1493 2958 1650 2380 1889 2197 1968 2014 2628 1845 1624 1898 1847 1257 2286 Response rate 58.7 51.7 33.8 49.1 67.8 70.3 67.3 52.1 64.5 67.9 55.1 54.9 74.9 52.4 57.7 53.1 Average Age 47.3 47.4 48.7 48.7 49.4 47.6 50.0 51.8 47.1 47.3 51.2 46.0 46.1 47.8 48.3 51.8 7 N 1769 1532 3045 1502 2051 1925 2087 1917 1698 2390 1919 1436 1615 1791 1224 2264 Response rate 57.0 52.7 31.4 51.9 59.9 67.9 62.7 50.9 52.7 60.7 58.6 53.9 65.8 50.1 52.3 43.6 Average Age 47.0 47.4 50.0 48.1 50.3 48.5 51.3 49.9 49.9 49.4 50.7 46.8 47.3 49.7 49.6 52.2 8 N 1766 1525 2852 2019 1958 1925 2070 1614 2757 1681 1545 1694 1551 1307 1959 Response rate 56.8 52.2 30.6 - 68.4 67.7 57.7 52.4 42.7 64.5 49.7 52.8 69.6 43.0 55.9 42.8 Average Age 47.0 47.8 48.6 49.6 49.6 50.0 52.4 50.8 50.2 51.2 47.0 47.2 51.6 49.1 51.4 9 N 1767 1542 2358 1572 1904 1668 1755 2010 1661 2216 1673 1406 1500 1539 1318 2204 Response rate 57.6 51.8 27.6 48.8 62.7 53.8 51.8 48.1 40.7 62.0 49.6 43.3 60.4 39.0 64.1 41.0 Average Age 47.9 47.5 49.6 49.8 50.7 48.5 51.0 52.4 51.0 52.2 48.6 47.1 47.6 52.5 49.4 52.4 10 N 1341 1523 8725 1542 2283 1577 1977 1849 1770 1470 1411 2065 2287 1252 1149 Response rate 39.19 49.5 37.0 - 47.2 35.5 41.1 39.6 40.4 36.31 35.7 37.9 39.2 37.9 54.7 20.88 Average Age 49.0 49.6 50.3 51.6 49.4 52.6 49.5 50.5 53.5 48.6 47.3 49.0 52.1 49.4 55.7 All Average N 1745 1693 3443 1551 1873 1945 1953 1904 1664 2223 1833 1607 1769 1822 1348 2098 Average resp. rate 56.3 49.5 40.0 55.2 62.6 60.5 61.7 47.2 55.3 59.2 55.1 55.8 66.9 53.6 61.4 47.4 Total N 17451 16925 34425 12408 16856 19452 19532 19038 16642 22233 18329 16065 17689 18216 13484 20979 3 Table A2: Coding schemes for categorical aggregations. edulvl3 edulvla edulvlb Primary 1 113, 129 Secondary 2, 3, 4 212, 213, 221, 222, 223, 229, 311, 312, 313, 321, 322, 323, 412, 413, 421, 422, 423 Tertiary 5 510, 520, 610, 620, 710, 720, 800 eisced3 eisced Primary 1 Secondary 2, 3, 4, 5 Tertiary 6, 7 Table A3: Regression coefficients, descriptive statistics by country. Country Median Mean Min Max IQR Range Sweden 0.275 0.273 0.130 0.366 0.059 0.236 Norway 0.311 0.319 0.174 0.507 0.095 0.333 Denmark 0.323 0.317 0.156 0.494 0.133 0.338 Finland 0.329 0.326 0.157 0.475 0.106 0.318 Ireland 0.334 0.344 0.245 0.511 0.052 0.267 France 0.357 0.353 0.194 0.497 0.097 0.304 Estonia 0.369 0.366 0.237 0.495 0.085 0.258 United Kingdom 0.371 0.370 0.209 0.506 0.069 0.297 Germany 0.384 0.376 0.218 0.528 0.083 0.311 Switzerland 0.386 0.383 0.246 0.491 0.059 0.245 Netherlands 0.387 0.391 0.241 0.559 0.100 0.318 Belgium 0.451 0.450 0.298 0.576 0.072 0.278 Spain 0.461 0.456 0.313 0.608 0.083 0.296 Slovenia 0.462 0.463 0.244 0.620 0.108 0.376 Hungary 0.466 0.466 0.302 0.643 0.101 0.341 Poland 0.512 0.501 0.406 0.583 0.056 0.177 Mean 0.386 0.385 0.236 0.529 0.085 0.293 4 Table A4: Correlations, descriptive statistics by country. Country Median Mean Min Max IQR Range Denmark 0.310 0.307 0.155 0.427 0.062 0.271 Germany 0.327 0.327 0.212 0.441 0.049 0.229 United Kingdom 0.333 0.331 0.175 0.463 0.067 0.288 Norway 0.333 0.341 0.181 0.481 0.080 0.300 Switzerland 0.377 0.366 0.231 0.475 0.080 0.245 Sweden 0.380 0.382 0.196 0.505 0.076 0.309 Netherlands 0.386 0.388 0.250 0.525 0.066 0.275 Finland 0.396 0.393 0.195 0.569 0.132 0.374 France 0.401 0.392 0.218 0.526 0.101 0.308 Spain 0.413 0.415 0.290 0.592 0.077 0.302 Ireland 0.423 0.432 0.312 0.605 0.059 0.292 Estonia 0.433 0.432 0.299 0.557 0.083 0.258 Belgium 0.455 0.452 0.310 0.577 0.074 0.268 Slovenia 0.465 0.465 0.272 0.568 0.068 0.296 Hungary 0.507 0.517 0.406 0.666 0.092 0.259 Poland 0.515 0.515 0.443 0.582 0.037 0.139 Mean 0.403 0.403 0.259 0.535 0.075 0.276 Table A5: Log odds ratios, descriptive statistics by country. Country Median Mean Min Max IQR Range Sweden 1.381 1.394 0.719 2.038 0.261 1.319 Germany 1.403 1.387 0.923 1.896 0.358 0.972 Finland 1.453 1.421 0.826 2.066 0.303 1.241 Denmark 1.463 1.447 0.506 1.998 0.208 1.491 Estonia 1.492 1.506 1.202 1.881 0.183 0.680 United Kingdom 1.540 1.518 0.826 2.079 0.272 1.252 Norway 1.553 1.543 1.058 2.063 0.258 1.005 Netherlands 1.680 1.719 1.136 2.309 0.312 1.173 Slovenia 1.756 1.774 0.782 2.524 0.398 1.742 Switzerland 1.834 1.920 1.099 3.792 0.537 2.693 Belgium 1.942 1.950 1.597 2.366 0.167 0.769 France 2.123 2.140 1.626 2.910 0.308 1.284 Spain 2.147 2.147 1.440 2.691 0.286 1.251 Ireland 2.215 2.232 1.462 2.946 0.436 1.484 Hungary 2.385 2.400 1.699 2.936 0.230 1.237 Poland 2.519 2.538 1.867 3.106 0.159 1.238 Mean 1.806 1.815 1.173 2.475 0.292 1.302 5 Table A6: Unidiff, descriptive statistics by country. Country Median Mean Min Max IQR Range United Kingdom -1.481 -1.548 -3.129 -0.683 0.574 2.446 Sweden -1.054 -0.979 -2.383 0.648 0.825 3.032 Finland -1.030 -1.066 -2.039 -0.024 0.467 2.015 Denmark -0.784 -0.785 -2.503 0.281 0.459 2.784 Spain -0.549 -0.550 -1.539 0.275 0.384 1.814 France -0.316 -0.302 -0.964 0.903 0.314 1.867 Netherlands -0.274 -0.316 -1.225 0.508 0.569 1.733 Estonia -0.083 -0.056 -0.836 0.908 0.440 1.744 Belgium -0.003 0.030 -0.634 1.014 0.339 1.648 Norway 0.146 0.130 -0.798 0.772 0.494 1.570 Ireland 0.216 0.203 -0.487 0.936 0.410 1.423 Germany 0.233 0.189 -1.335 1.043 0.625 2.378 Slovenia 0.882 0.825 -1.400 1.496 0.448 2.896 Switzerland 0.888 0.844 -0.642 2.130 0.417 2.771 Poland 1.675 1.705 0.974 2.313 0.391 1.339 Hungary 1.702 1.678 1.059 2.317 0.314 1.258 Mean 0.011 0.000 -1.118 0.927 0.467 2.045 Table A7: Rank correlations, descriptive statistics by country. Country Median Mean Min Max IQR Range United Kingdom 0.281 0.283 0.173 0.405 0.050 0.232 France 0.324 0.321 0.177 0.450 0.083 0.273 Norway 0.340 0.340 0.184 0.476 0.076 0.291 Germany 0.344 0.343 0.240 0.442 0.045 0.202 Denmark 0.347 0.337 0.166 0.444 0.073 0.278 Netherlands 0.349 0.347 0.210 0.471 0.072 0.261 Finland 0.353 0.354 0.169 0.504 0.115 0.335 Switzerland 0.372 0.367 0.191 0.489 0.078 0.298 Sweden 0.372 0.372 0.220 0.486 0.073 0.266 Spain 0.402 0.397 0.260 0.564 0.104 0.304 Ireland 0.410 0.407 0.267 0.527 0.077 0.260 Estonia 0.418 0.420 0.280 0.536 0.089 0.256 Belgium 0.420 0.418 0.330 0.517 0.049 0.187 Slovenia 0.427 0.431 0.248 0.579 0.097 0.331 Hungary 0.500 0.505 0.402 0.635 0.072 0.233 Poland 0.507 0.498 0.400 0.615 0.065 0.215 Mean 0.385 0.384 0.245 0.509 0.076 0.264 6 Table A8: Correlation between country rankings based on different parameters. Regression Correlation Log odds Unidiff Rank corr Regression 1.000 Correlation 0.599 1.000 Log odds 0.531 0.643 1.000 Unidiff 0.505 0.470 0.566 1.000 Rank corr 0.570 0.838 0.541 0.538 1.000 Table A9: Factor analysis: Eigenvalues and proportion of variance. Factor Eigenvalue Difference Proportion Cumulative 1 2.93631 2.73209 1.0144 1.0144 2 0.20422 0.20971 0.0705 1.0849 3 -0.00549 0.05728 -0.0019 1.0830 4 -0.06277 0.11476 -0.0217 1.0613 5 -0.17753 -0.0613 1.0000 Table A10: Factor loadings (pattern matrix) and unique variances. Parameter Factor 1 Factor 2 Uniqueness Regression 0.6920 0.0994 0.5113 Correlation 0.8782 -0.2187 0.1810 Log odds 0.7298 0.1936 0.4299 Unidiff 0.6535 0.2515 0.5097 Rank corr 0.8524 -0.2140 0.2276 13 (a) Regression coefficients 0.68 0.70 0.72 0.77 0.77 0.68 Split by: Foreign born incl? Which parent? Gender Education variable Age restriction Baseline 0.00 0.25 0.50 0.75 1.00 Rank stability (b) Correlations 0.68 0.69 0.73 0.74 0.78 0.67 Split by: Foreign born incl? Which parent? Education variable Gender Age restriction Baseline 0.00 0.25 0.50 0.75 1.00 Rank stability (c) Log odds ratios 0.80 0.81 0.82 0.82 0.84 0.80 Split by: Foreign born incl? Education variable Age restriction Which parent? Gender Baseline 0.00 0.25 0.50 0.75 1.00 Rank stability (d) Unidiff coefficients 0.84 0.84 0.85 0.86 0.87 0.84 Split by: Foreign born incl? Education variable Which parent? Age restriction Gender Baseline 0.00 0.25 0.50 0.75 1.00 Rank stability (e) Rank correlations 0.68 0.69 0.73 0.75 0.77 0.67 Split by: Foreign born incl? Which parent? Education variable Gender Age restriction Baseline 0.00 0.25 0.50 0.75 1.00 Rank stability Figure A7: Rank stability by parameter of association. Note: The figure shows rank stability at baseline and split by model components, by parameter of association. 14 4 Sampling vs model uncertainty (a) Regression coefficients BE CH DE DK EE ES FI FR UK HU IE NL NO PL SE SI .2 .3 .4 .5 .6 Mean estimate ± mean SE .2 .3 .4 .5 .6 Mean estimate ± SD of estimates Sampling uncertainty (mean SE) Model uncertainty (SD of estimates) (b) Correlations BE CH DE DK EE ES FI FR UK HU IE NL NO PL SE SI .2 .3 .4 .5 .6 Mean estimate ± mean SE .2 .3 .4 .5 .6 Mean estimate ± SD of estimates Sampling uncertainty (mean SE) Model uncertainty (SD of estimates) (c) Log odds ratios BE CH DE DK EE ES FI FR UK HU IE NL NO PL SE SI 11.5 22.5 3 Mean estimate ± mean SE 11.5 22.5 3 Mean estimate ± SD of estimates Sampling uncertainty (mean SE) Model uncertainty (SD of estimates) (d) Rank correlations BE CH DE DK EE ES FI FR UK HU IE NL NO PL SE SI .2 .3 .4 .5 .6 Mean estimate ± mean SE .2 .3 .4 .5 .6 Mean estimate ± SD of estimates Sampling uncertainty (mean SE) Model uncertainty (SD of estimates) Figure A8: Sampling and model uncertainty. Note: The figure plots average point estimates across specifications for each country, by parameter of association. Vertical error bars represent sampling uncertainty (mean of standard errors across specifications). Horizontal error bars represent model uncertainty (standard deviation of point estimates across specifications). 15 4.1 Single-model results with bootstrap BE ES PL IE GB FR CH HU SI FI DE NL SE NO EE DK 0.2 0.3 0.4 0.5 0.6 Coefficient BE ES PL IE GB FR CH HU SI FI DE NL SE NO EE DK 1 3 5 7 9 11 13 15 Rank BE ES SI PL FR GB HU CH IE DE NO NL FI DK EE SE 0.2 0.3 0.4 0.5 0.6 Coefficient BE ES SI PL FR GB HU CH IE DE NO NL FI DK EE SE 1 3 5 7 9 11 13 15 Rank Figure A9: Regression slopes, 1 specification with 250 bootstrap iterations per country. Top: men, bottom: women. 16 BE HU PL IE ES FR SI CH FI SE GB DK NL NO EE DE 0.2 0.3 0.4 0.5 0.6 Coefficient BE HU PL IE ES FR SI CH FI SE GB DK NL NO EE DE 1 3 5 7 9 11 13 15 Rank BE HU PL SI FR ES IE CH NO DK NL FI GB SE DE EE 0.3 0.4 0.5 0.6 Coefficient BE HU PL SI FR ES IE CH NO DK NL FI GB SE DE EE 1 3 5 7 9 11 13 15 Rank Figure A10: Correlations, 1 specification with 250 bootstrap iterations per country. Top: men, bottom: women. 17 PL HU IE FR ES CH BE SE GB SI NO DK DE FI NL EE 1.0 1.5 2.0 2.5 3.0 Coefficient PL HU IE FR ES CH BE SE GB SI NO DK DE FI NL EE 1 3 5 7 9 11 13 15 Rank HU PL FR IE CH SI ES BE NL GB DE NO DK FI SE EE 1.5 2.0 2.5 3.0 Coefficient HU PL FR IE CH SI ES BE NL GB DE NO DK FI SE EE 1 3 5 7 9 11 13 15 Rank Figure A11: Log odds ratios, 1 specification with 250 bootstrap iterations per country. Top: men, bottom: women. 18 HU PL CH SI DE IE NO BE FR DK SE NL ES EE FI GB −2 −1 0 1 2 3 HU PL CH SI DE IE NO BE FR DK SE NL ES EE FI GB 1 3 5 7 9 11 13 15 Rank HU SI PL CH DE IE NO NL BE DK FR FI ES EE SE GB −2 −1 0 1 2 HU SI PL CH DE IE NO NL BE DK FR FI ES EE SE GB 1 3 5 7 9 11 13 15 Rank Figure A12: Unidiff, 1 specification with 250 bootstrap iterations per country. Top: men, bottom: women. 19 HU IE BE ES PL SE CH SI FR FI DE NL NO EE GB DK 0.3 0.4 0.5 Coefficient HU IE BE ES PL SE CH SI FR FI DE NL NO EE GB DK 1 3 5 7 9 11 13 15 Rank HU PL ES SI BE IE CH FR DK NO NL DE SE FI EE GB 0.3 0.4 0.5 0.6 Coefficient HU PL ES SI BE IE CH FR DK NO NL DE SE FI EE GB 1 3 5 7 9 11 13 15 Rank Figure A13: Rank correlations, 1 specification with 250 bootstrap iterations per country. Top: men, bottom: women. 20 4.2 Multiverse results with pseudo-bootstrap Poland Spain Slovenia Belgium Hungary Netherlands Switzerland Germany United Kingdom France Estonia Finland Ireland Denmark Norway Sweden 0.0 0.2 0.4 0.6 Coefficient Poland Spain Slovenia Belgium Hungary Netherlands Switzerland Germany United Kingdom France Estonia Finland Ireland Denmark Norway Sweden 1 3 5 7 9 11 13 15 Rank Figure A14: Regression coefficients, pseudo-bootstrapped, 250 rankings per country-specification. Poland Hungary Slovenia Belgium Ireland Estonia Spain Finland France Netherlands Sweden Switzerland Norway United Kingdom Germany Denmark 0.0 0.2 0.4 0.6 0.8 Coefficient Poland Hungary Slovenia Belgium Ireland Estonia Spain Finland France Netherlands Sweden Switzerland Norway United Kingdom Germany Denmark 1 3 5 7 9 11 13 15 Rank Figure A15: Correlations, pseudo-bootstrapped, 250 rankings per countryspecification. 21 Poland Hungary Ireland Spain France Belgium Switzerland Slovenia Netherlands Norway United Kingdom Estonia Denmark Finland Sweden Germany 0 2 4 6 Coefficient Poland Hungary Ireland Spain France Belgium Switzerland Slovenia Netherlands Norway United Kingdom Estonia Denmark Finland Sweden Germany 1 3 5 7 9 11 13 15 Rank Figure A16: Log odds ratios, pseudo-bootstrapped, 250 rankings per countryspecification. Hungary Poland Slovenia Belgium Estonia Ireland Spain Sweden Switzerland Finland Netherlands Norway Germany Denmark France United Kingdom 0.2 0.4 0.6 Coefficient Hungary Poland Slovenia Belgium Estonia Ireland Spain Sweden Switzerland Finland Netherlands Norway Germany Denmark France United Kingdom 1 3 5 7 9 11 13 15 Rank Figure A17: Rank correlations, pseudo-bootstrapped, 250 rankings per countryspecification. 22