Just cheap talk? Investigating fairness preferences in hypothetical scenarios
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Hufe, Paul; Weishaar, Daniel Article — Published Version Just cheap talk? Investigating fairness preferences in hypothetical scenarios The Journal of Economic Inequality Provided in Cooperation with: Springer Nature Suggested Citation: Hufe, Paul; Weishaar, Daniel (2025) : Just cheap talk? Investigating fairness preferences in hypothetical scenarios, The Journal of Economic Inequality, ISSN 1573-8701, Springer US, New York, NY, Vol. 23, Iss. 3, pp. 881-907, https://doi.org/10.1007/s10888-025-09700-w This Version is available at: https://hdl.handle.net/10419/330533 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
https://doi.org/10.1007/s10888-025-09700-w RESEARCH Just cheap talk? Investigating fairness preferences in hypothetical scenarios Paul Hufe1 ·Daniel Weishaar2 Received: 27 June 2025 / Accepted: 8 July 2025 © The Author(s) 2025 Abstract The measurement of preferences often relies on surveys in which individuals evaluate hypothetical scenarios. This paper proposes and validates a novel factorial survey tool to measure fairness preferences. We examine whether a non-incentivized survey captures the same distributional preferences as an impartial spectator design, where choices may apply to a real person. In contrast to prior studies, our design involves high stakes, with respondents determining a real person’s monthly earnings, ranging from $500 - $5,700. We find that the non-incentivized survey module yields nearly identical results compared to the incentivized experiment and recovers fairness preferences that are stable over time. Furthermore, we show that most respondents adopt intermediate fairness positions, with fewer exhibiting strictly egalitarian or libertarian preferences. In sum, these findings suggest that high-stake incentives do not significantly impact the measurement of fairness preferences and that non-incentivized survey questions covering realistic scenarios offer valuable insights into the nature of these preferences. Keywords Fairness preferences ·Survey experiment ·Vignette studies JEL Classification C90 ·D63 ·I39 This paper benefited strongly from discussions with Ingvild Almås. It is part of a larger research project in which we measure fairness preferences and beliefs about inequality around the world (see Riksbankens Jubileumsfond (P22-0564) and UKRI Future Leaders Fellowship (MR/X033333/1): “(Un)Fair inequality in the labor market: a global study”). We received useful feedback from seminar audiences at LMU Munich. This research was funded by the British Academy (TDA21/210082) and by Deutsche Forschungsgemeinschaft through CRC TRR 190 (No. 280092119) and under Germany’s Excellence Strategy – EXC 2126/1-390838866. Hufe gratefully acknowledges financial support from UKRI (MR/X033333/1). Weishaar gratefully acknowledges financial support from the Joachim Herz Foundation and the Fritz Thyssen Foundation. The questionnaire and core analysis were pre-registered via the Open Science Framework (OSF), No. DV3KP. We obtained ethical approval from the Institutional Review Boards at the University of Bristol, LMU Munich, and NHH Norwegian School of Economics. All remaining errors are our own. The online version contains supplementary material available at https://doi.org/10.1007/s10888-025-09700-w. BPaul Hufe [email protected] Daniel Weishaar [email protected] 1University of Bristol, Bristol, England 2University of Cologne, Cologne, Germany 123 The Journal of Economic Inequality (2025) 23:881–907 / Published online: 16 September 2025
P. Hufe, D. Weishaar 1 Introduction There is expanding literature in economics and other social sciences that investigates which inequalities are seen as unfair by people (Alesina et al. 2018; Almås et al. 2020; Andre 2025; Cappelen et al. 2007; Jasso and Webster 1999; Konow 2000). In these papers, fairness preferences are typically elicited using incentivized experiments or non-incentivized surveys. Researchers considering the choice between these two research designs face a trade-off. On the one hand, experiments combine stylized representations of real-world situations with payout-relevant choices of respondents. On the other hand, survey questions often mirror real-world contexts more closely; however, respondents’ answers have no consequences in the real world. Therefore, survey-based methods are often considered unreliable predictors of actual behavior. This raises the question of whether researchers can employ non-incentivized surveys to analyze fairness preferences or whether such answers must be considered “cheap talk.” In this paper, we address this question by using a representative sample of the US adult population to test whether answers to hypothetical questions align with those from an incentivized experiment. Our survey tool integrates core functionalities of impartial spectator experiments (Almås et al. 2020; Almås et al. 2024a; Andre 2025; Cappelen et al. 2013; Konow 2000; Konow et al. 2020) with the methodological advantages of factorial surveys (Auspurg et al. 2017; Gaertner and Schwettmann 2007; Jasso and Webster 1999; Konow 1996).1The questions in our survey tool show respondents pairs of hypothetical persons that are described in terms of observable characteristics, i.e., their gender, age, educational attainment, parental background, working hours, and labor market earnings. Respondents then distribute earnings between these two persons based on the given information. In impartial spectator designs, choices are not payoff-relevant for the respondents themselves. Yet, answers in a hypothetical task might still differ from their incentivized analogs. First, hypothetical distribution tasks are prone to different biases, including experimenter demand effects and social desirability biases (Stantcheva 2023). For instance, in the absence of real-world consequences, respondents may tend towards more equal income allocations to comply with perceived expectations from the researcher. Such demand effects are less of a concern when real money is at stake (Haaland et al. 2023). Second, even without such systematic biases, hypothetical distribution tasks may be prone to measurement error, e.g., if non-incentivized respondents are less attentive when completing the tasks. Our survey tool differs from prior literature on fairness preferences which has largely focused on the extent to which individuals reward rather abstract concepts such as “luck,” “productivity,” “hard work,” and “talent” (e.g., Cappelen et al. 2010; Mollerstrom et al. 2015). In contrast to these studies, our survey tool allows us to elicit fairness preferences that can be directly mapped to observable labor market inequalities, e.g., gender gaps (Blau et al. 2017), returns to hours (Kuhn and Lozano 2008), education premia (Harmon et al. 2003), and intergenerational persistence (Roemer and Trannoy 2016). However, the focus on these real-world inequalities also makes it prohibitively costly to elicit the relevant preferences in an experimental design where respondents’ choices are consequential for the earnings of actual persons. Therefore, it is important to test whether hypothetical distribution tasks deliver credible results that align with the “gold standard” of incentivized experiments. 1For detailed reviews on experimental and survey-based evidence on fairness preferences, see Almås et al. (2023,2024b) and Gaertner and Schokkaert (2012). 123 882
Just cheap talk? Investigating fairness preferences To address this question, we collected data from 1,602 adults in the United States between October and November 2022. The sampling was designed to be representative of various demographic characteristics such as gender, age, education, employment status, and region of residence. The survey modules, as well as core analyses, were pre-registered via the Open Science Framework (OSF), No. DV3KP. We validate our survey tool along three dimensions. First, we test whether the distributional choices of respondents are different if they are payoff-relevant. For this purpose, we run an experiment with a between-subject design. All respondents answer a survey where they face a selection of tasks from our survey tool. Respondents in the treatment group are informed that one of the persons shown to them is a real person and that the decision made by a randomly chosen respondent will determine the monthly earnings of this individual. Thus, in contrast to the control group, they know that each choice may have substantial financial consequences for a real person. This design allows us to test whether fairness preferences in our hypothetical tool are consistent with the “gold standard” of an incentivized experiment (Bauer et al. 2020; Enke et al. 2022;Falketal.2023). Second, we test whether the distributional choices are stable over time. For this purpose, we employ a within-subject design and run an obfuscated followup one week after the baseline survey.2In particular, we invite respondents to another survey, where they again face a selection of tasks from our survey tool. Some tasks are repeated from the baseline wave, allowing us to calculate intertemporal correlations. This design enables us to test the stability of fairness preferences in our hypothetical tool and gives crucial information on measurement error in the elicited preference data (Gillen et al. 2019;Stantcheva2023). Lastly, next to the methodological validation of the survey design, we conduct a suggestive substantive analysis. Specifically, we describe the nature of fairness preferences identified through our survey and the heterogeneity of fairness views within the US population. This analysis has several caveats since the survey design was premised on methodological validation. Nevertheless, the substantive analysis provides an important cross-check on whether our hypothetical survey tool recovers preferences that are consistent with previous studies on fairness preferences in the US (Almås et al. 2020; Fisman et al. 2023;Konowetal.2020). Our results can be summarized as follows. First, the distributional choices of respondents are not affected by making them relevant to the earnings of real persons. The point estimates for the “real-person treatment” are small and insignificant at conventional levels of statistical significance. This conclusion remains unaffected when considering treatment effects on the distribution of choices and treatment effects within various population subgroups. Second, the distributional choices of respondents are relatively stable over time. The average (intra-respondent) intertemporal correlation of distributional choices is 0.56, which lies in the range of test-retest correlations of other preference survey modules (Enke et al. 2022). Furthermore, the intertemporal correlation is slightly higher in the incentivized group, suggesting that incentives have a small positive effect on reducing measurement error in the elicited preferences. Third, we find that the nature of the recovered fairness preferences is broadly consistent with previous studies on the US. Inequality acceptance ranges between Gini coefficients of 0.30 and 0.53 (e.g., Almås et al. 2020), the majority of respondents adopt intermediate fairness positions that are influenced by discretionary variables such as education and working hours (e.g., Konow 2000), and the distributional choices of different population subgroups are consistent with self-serving biases (Costa-Font and Cowell 2014). 2This choice of the time lag is consistent with several recent articles that validate survey-based measurement tools using time-lags between one and two weeks (Bauer et al. 2020;Enkeetal.2022; Falk et al. 2023; Fallucchi et al. 2020). 123 883
P. Hufe, D. Weishaar In summary, the results from our validation suggest that the proposed hypothetical survey tool recovers fairness preferences that are consistent with incentivized choices, stable, and reasonable in light of the existing literature. This paper contributes to two strands of the literature. First, we contribute to the literature on fairness preferences. There is a large literature in economics and other social sciences trying to understand the nature and anatomy of fairness preferences in different population groups (Almås et al. 2020; Andre 2025; Cappelen et al. 2007; Gaertner and Schokkaert 2012; Jasso and Webster 1999; Konow 2000; Starmans et al. 2017). In this paper, we validate a vignette-based survey tool that allows researchers to investigate fairness preferences in a flexible and cost-efficient way. Therefore, this study provides a crucial step to strengthen the methodological toolkit for investigating fairness preferences in applied research. Furthermore, there is a growing literature investigating the upstream determinants and downstream consequences of fairness preferences (Adriaans 2023; Alesina et al. 2018; Andersen et al. 2023;Fehretal.2024). These studies often rely on survey-based measures of fairness preferences. Our results provide encouraging news for such research designs as the consistency of hypothetical and incentivized choices suggests that survey-based measures are not systematically biased compared to their incentivized analogs. Second, we contribute to a growing methodological literature that validates survey-based measurement tools in various domains, including risk, time, competition, and social preferences (Bauer et al. 2020;Enkeetal.2022; Falk et al. 2023; Fallucchi et al. 2020). To the best of our knowledge, our study is the first validation of an impartial spectator task under high monetary stakes. The remainder of the paper is organized as follows. Section 2describes the survey tool and provides information on the data collection. In Section 3, we present results for the effects of the “real-person treatment.” Section 4describes the stability of fairness preferences. Section 5 provides a suggestive comparison of the recovered fairness preferences relative to existing literature. Section 6concludes the paper. 2 Survey tool and data collection Survey structure Figure 1provides an overview of the survey used in our analysis. The survey is structured into two waves, each consisting of multiple modules. In the first module of the baseline wave, we elicit the demographic characteristics of respondents. The second and third modules measure inequality perceptions and fairness preferences. The final module Baseline Wave Demographics Fairness Preferences Labor Markets Real Person Hypothetical Real Person Hypothetical 5 Questions 5 Questions 7 Questions 7 Questions Real Person Hypothetical 9 Questions 9 Questions Follow-Up Wave Inequality Perceptions Financial Inc. No Financial Inc. Financial Inc. No Financial Inc. 6 Questions 6 Questions 8 Questions 8 Questions Financial Inc. No Financial Inc. 10 Questions 10 Questions Inequality Perceptions Fairness Preferences 1week Fig. 1 Survey Structure. Note: This figure visualizes the structure of the survey with two waves (baseline, follow-up). Each wave consists of multiple modules. The modules on inequality perceptions are blurred out since they are not covered in this paper. The main treatment group (control group) is highlighted in red (gray) 123 884
Just cheap talk? Investigating fairness preferences Fig. 2 Fairness Preference Elicitation, Exemplary Question. Note: This figure provides an example of a question screen in the fairness preference module. Each question shows the characteristics of two persons in six dimensions (earnings, gender, age, education, parental education, working hours) in a table format. Each of the two persons has been allocated a random person identifier from 1 to 9999. The ordering of characteristics in the table is randomized at the respondent-question level. A slider allows respondents to select their preferred distribution of earnings between the two persons. The chosen allocation is also shown numerically above the slider contains additional questions about the labor market. The two modules of the follow-up wave mirror the perceptions and preference modules of the baseline wave. In this paper, we focus exclusively on fairness preferences.3 Preference module The module on fairness preferences consists of multiple questions that follow the design of factorial survey experiments. Factorial survey experiments are wellestablished tools in the social sciences to assess preferences and beliefs (Auspurg et al. 2015,2017;Fismanetal.2020; Jasso and Webster 1997,1999; Wiswall and Zafar 2018). In such experiments, respondents evaluate multiple hypothetical scenarios that vary at random in pre-defined characteristics. The random variation of characteristics has two main advantages. First, the design can replicate the complexities of the real world. In particular, respondents are forced to make trade-offs and weigh the importance of different real-world attributes against each other when making their choices. Second, the simultaneous variation of characteristics mitigates experimenter demand effects and social desirability biases—concerns that are particularly relevant in the domain of fairness preferences (Auspurg and Hinz 2014). In our survey module, respondents receive information on the observational characteristics of two persons—see Fig. 2for an example. We describe these persons in terms of six characteristics, i.e., their gender, age, own education, parental education, working hours, and labor market earnings. The choice of these characteristics was guided by the following considerations. First, we selected the 3The perceptions modules are designed to assess respondents’ perceptions of inequality in the labor market. All treatments in this module are independent of the treatments in the preference module, allowing us to analyze these data in isolation. 123 885
P. Hufe, D. Weishaar dimensionality of the vignettes by acknowledging the trade-off between precision and complexity in the design of factorial experiments. On the one hand, a higher number of vignette characteristics has the benefit that more confounding factors can be held constant. On the other hand, increasing the number of characteristics also puts a larger cognitive burden on participants when making distributional choices. With six dimensions, we opted for an intermediate level of complexity (Auspurg and Hinz 2014). Second, we then selected characteristics that are both relevant for understanding earnings inequality and widely available in standard household survey data (Bick et al. 2022; Goldin 2014;Lemieux2006; Magnac and Roux 2021; Mazumder 2005). We are aware that the resulting selection of characteristics is somewhat discretionary. Therefore, we also validate our selection ex-post by asking respondents about characteristics they consider particularly important when making distributional choices. Appendix Figure A1 shows that five out of the six selected characteristics are chosen by at least ten percent of respondents. Furthermore, many of our chosen characteristics correlate strongly with other characteristics that participants find important, e.g., age (our selection) and work experience (respondents selection). Each characteristic can take multiple expressions. For instance, the characteristic of education can take three values, i.e., High School Dropout,High School,orUniversity.The potential expressions for each characteristic are shown in Table 1. The order of characteristics is randomized at the respondent-question level to ensure that results are not driven by order effects (Day et al. 2012). Combining all potential expressions yields a set of 720 profiles (2×2×3×3×4×5) and a set of 258,840 (720×719 2) unique unordered profile pairs. In the following, we will refer to each unique profile as a vignette person and each unique profile pair as a vignette.Weshow respondents a selection of vignettes determined via a random draw from the full set. Based on the presented information, respondents can use a slider to adjust the initial earnings and to implement their preferred earnings distribution in the vignettes. All respondents answered the same selection of vignettes, with some variation in the number of vignettes per respondent— see also our discussion on the “length treatment” below. Treatment 1: “Real-person treatment” To assess whether hypothetical questions in factorial surveys recover fairness preferences consistent with incentivized experiments, we randomize respondents into two groups. Respondents in the control group complete a series of hypothetical distribution tasks. Respondents in the treatment group complete the same Table 1 Fairness Preference Elicitation, Characteristics and Expressions Characteristic Number Displayed Values for Expression Gender 2 Male / Female Age 2 (26 , 35 , 40) / (50, 55, 59) Education 3 High School Drop-Out / High School / University Parental Education 3 High School Drop-Out / High School / University Working Hours 4 (2, 5) / (20, 27, 31) / (39, 40) / (48, 51, 59) Earnings 5 $1,100 / $2,700 / $4,100 / $5,900 / $11,400 Note: This table shows the characteristics displayed in questions of the fairness preference module (column 1), the number of coarse expressions within each characteristic (column 2), and the displayed values for each expression (column 3). We employ a second randomization for age and working hours to display exact values instead of ranges. For each of the two age range groups (25-44, 45-65) and each of the four working hours range groups (1-9, 10-34, 35-44, More than 44), we draw from integers in the respective range 123 886
Just cheap talk? Investigating fairness preferences series of tasks. However, they are informed that one of the vignette persons is a real person. Furthermore, they are informed that the choice of one respondent will be selected to determine the monthly earnings of this person. Respondents know that the real person is between 25 and 65 years old, is a resident of the United States, and works in a job to earn money. Importantly, this information does not allow respondents to distinguish the real person from any other vignette person. They also know that the total earnings of the real person consist of two parts: (i) a fixed payment of $500 and (ii) a flexible payment that can be changed by the respondent. The research team hired a real person in August 2022—see Appendix Table A1 for the characteristics of the real person as displayed in their vignette. The hired person was informed that the fixed-term contract would have a duration of one month and that the exact amount of their earnings would be determined by another individual; however, they did not know the exact process of how this happens. The person only knew their total earnings would be $500 or above. To determine the potential earnings of the real person, we proceeded in two steps. First, we allocated the real person a monthly earnings value by randomly drawing from the set of potential earnings displayed in Table 1. Second, we randomly matched the real person with another (hypothetical) vignette person. This two-step procedure fixed the volume of earnings in the vignette of the real person at $5,200. Therefore, including the fixed payment of $500, the upper bound of potential earnings for the real person was $5,700. This upper bound would be realized if the decisive respondent allocated all the vignette earnings to the real person. Importantly, the vignette with the real person was presented alongside all other vignettes, and the identity of the real person was concealed from respondents. Consequently, respondents also faced situations in which the potential earnings implications were even higher. The average earnings volume in the displayed vignettes—and therefore the average upper bound of potential payments to the real person from the respondents’ perspective— was $11,420 in the main vignettes of the preference module. Furthermore, we ensured the salience of the real-world consequences through a training task. Specifically, we trained respondents on an example vignette and highlighted the potential earnings consequences for the real person after respondents had made their distributional choices. Our incentivization is based on high stakes with a low probability of implementation. Alternatively, we could have used an incentive scheme based on lower stakes but a higher probability of implementation. Due to our focus on real-world characteristics, the former path appears more natural since it allows us to focus on payouts that reflect realistic monthly wages drawn from the US working population. However, we contend that a generalization of our results to alternative incentive schemes is an interesting path for future research. Treatment 2: Length treatment To assess the sensitivity of our conclusions to the length of the survey, we vary the number of vignettes in the survey module. All respondents made at least five distributional choices; however, 1/3 of respondents received 2 or 4 additional vignettes, respectively. For the main validation, we focus on the first five questions answered by all respondents and use the variation from the length treatments in robustness analyses. The assignment to the length treatments was independent of the allocation to the “real-person treatment.” Therefore, we obtain six groups of approximately equal size that vary in their exposure to the “real-person treatment” and the survey length (Fig. 1). Baseline wave We administered the baseline wave of the survey to 1,602 adult citizens of the United States. Data collection took place between October and November 2022. Respondents were contacted through the survey provider Dynata and received a participation payment 123 887
P. Hufe, D. Weishaar depending on the expected survey length.4The mean (median) completion time was 24 (19) minutes for the baseline wave (Appendix Figure A2). Respondents were targeted to match the population along five dimensions (gender, age, education, employment status, and region of residence). In Panel A of Table 2, we compare our sample to the American Community Survey (ACS) regarding the targeted characteristics. In general, we match the data well. Our sample has a slight underrepresentation of people with low education. In addition, we received over-proportional (under-proportional) responses from mid-western states (southern states). Panel B of Table 2further shows that our sample is also broadly representative regarding other observable characteristics like ethnicity and income. We take various steps to ensure the quality of survey answers. First, we included an attention check at the beginning of the preference module. Respondents who failed this attention check were screened out directly and were not part of the sample. Second, we asked a training question after explaining the tasks in the preference module. Around 72% of respondents passed this question on the first try. In robustness checks, we show that our results are not sensitive to excluding respondents who did not pass the training question on their first attempt. Follow-up wave We invited respondents to a follow-up wave one week after they completed the baseline wave. Among others, the follow-up wave consists of a fairness preference module with six questions. As in the baseline wave, respondents faced a selection of distribution tasks based on the vignettes from our survey module. Three of the questions are repetitions from the baseline wave, which were presented to the respondents in an obfuscated way. This feature allows us to assess the stability of fairness preferences over time. The follow-up obtained a response rate of around 44%, and around 90% of respondents answered within two weeks. The resulting sample is slightly older but otherwise broadly comparable to our baseline sample regarding observable demographics (Appendix Table A2). The mean (median) completion for the follow-up wave was 14 (9) minutes (Appendix Figure A2). 3 Effects of the “real-person treatment” In this section, we investigate whether the potential for real-world implementation affected the distributional choices of respondents. First, we present methodological checks on the randomization and the anonymity of the real person. Second, we present the treatment effects on distributional choices. Third, we investigate potential heterogeneities by population subgroups. Fourth, we present robustness analyses. All analyses in this section are pre-registered unless noted otherwise. Balancing To give our estimates a causal interpretation, the treatment assignment must be uncorrelated with any respondent characteristics that may predict their distributional choices. Therefore, we test the balance of respondents’ socio-demographic characteristics between the treatment and the control group. In particular, we regress the treatment status Treat iof respondent ion Kpre-specified individual characteristics denoted by xk i: xk i=αk+βkTreati+εk i.(1) 4For the baseline wave, respondents were able to earn between $0.20 and $1.50. For the shorter follow-up wave, the payment varied between $0.10 and $1.20. The varying participation payment was used by Dynata to obtain responses from demographic segments of the population that are more difficult to reach. 123 888
Just cheap talk? Investigating fairness preferences Table 6 Allocations, “Real-Person Treatment”, Alternative Samples Question NPoint Estimate Control Mean Model p-value Resample p-value Romano-Wolf p-value Panel A: Directly Passed Training Question Q1 1154 168.87 1308.25 0.237 0.244 0.686 Q2 1154 31.76 -2574.27 0.774 0.777 0.949 Q3 1154 112.27 3320.52 0.813 0.813 0.949 Q4 1154 -147.78 -4508.78 0.613 0.629 0.935 Q5 1154 -290.95 6884.88 0.361 0.339 0.774 Panel B: Exclude Response Time Outliers Q1 924 162.99 1394.95 0.317 0.321 0.766 Q2 924 109.22 -2699.89 0.364 0.375 0.766 Q3 924 -270.11 4030.09 0.606 0.607 0.854 Q4 924 -30.23 -4623.23 0.925 0.933 0.933 Q5 924 -477.60 7159.91 0.169 0.177 0.572 Note: This table presents results of the regression analysis outlined in Eq. 2for different restricted samples. Panel A displays results for the sample of respondents that passed the training question of the fairness preference module at the first try. Panel B presents results for a sample that excludes respondents with high (above p90) and low (below p10) response times of the fairness preference module. We present point estimates for the coefficient of interest βj, the mean of the control group, and heteroskedasticity robust uncorrected analytical p-values (model p-values), uncorrected bootstrapped p-values (resample p-values), as well as pvalues adjusted for multiple hypothesis testing using 1000 bootstrap replications. See Appendix Tables B6-B7 for the corresponding balancing and Kolmogorov-Smirnov tests Source: Own calculation based on survey responses Results in Table 8show that there are virtually no differences between treatment and control groups, suggesting that the effort of respondents does not decrease when facing hypothetical instead of incentivized scenarios.8 In this section, we have shown that there are no systematic differences in distributional choices between hypothetical and incentivized scenarios. This conclusion holds for average allocations, distributions of allocations, and within various demographic subgroups. Furthermore, this conclusion is robust to various sensitivity checks, such as excluding low-quality responses. In summary, the results suggest that hypothetical vignettes capture the same fairness preferences as their incentivized analogs. 4 Stability of fairness preferences In this section, we investigate whether the survey module captures genuine fairness preferences that are stable over time. In particular, we use the longitudinal variation between the baseline wave and the follow-up wave of the survey. The follow-up wave consists of a fairness 8To limit the influence of extreme outliers, we focus on respondents whose response time for a particular question is below the 99th percentile of the question-specific response time distribution. In Appendix Table 25, we repeat the exercise without this restriction. Treatment effects increase due to single outliers in the treatment and control groups. However, none of the differences is statistically significant, and our general conclusion remains unaffected. 123 895
P. Hufe, D. Weishaar Table 7 Allocations, “Real-Person Treatment”, by Survey Length Subgroups Rejected Hypotheses at 5% (10%) Module Length NModel p-value Resample p-value RW p-value 5 Questions 534 0 (1) 0 (1) 0 (0) 7 Questions 536 0 (0) 0 (0) 0 (0) 9 Questions 532 1 (1) 1 (1) 0 (0) Note: This table presents results of the regression analysis outlined in Eq. 2for a subsample of respondents based on the survey module length, i.e, whether respondents answered five, seven, or nine questions in the preference module. We present the number of respondents, the number of rejected hypotheses according to heteroskedasticity robust model p-values, resample p-values, and p-values adjusted for multiple hypothesis testing. In total, we test five, seven, and nine hypotheses for each subsample depending on the number of questions. Detailed information on regression results are shown in Appendix Tables B22 - B24 Source: Own calculations based on survey responses preference module with six questions, three of which are repetitions from the baseline survey. To avoid respondents anchoring their responses on their answers in the baseline survey, we obfuscate the repeated questions by mixing them in random order with novel questions that have not been shown to respondents previously. We present results in three steps. First, we present intertemporal correlations based on the pooled follow-up sample. Second, we investigate whether these intertemporal correlations vary by treatment status in the baseline survey. Third, we present robustness analyses. We registered the follow-up survey in our pre-analysis plan. However, since the survey provider expressed considerable uncertainty about the likely response rates, we did not pre-specify the associated analyses presented in this section. Intertemporal correlations We estimate intertemporal correlations through ordinary leastsquares using the following model: yij,t=α+σyij,t−1+εij,(3) where yij is again the difference in allocations to vignette persons A and B by respondent iin vignette jin the baseline wave (t−1) and the follow-up wave (t), respectively. In all estimations, we standardize yij,tand yij,t−1on the estimation samples such that they Table 8 Preferences, Response Time (Min.), “Real-Person Treatment” Question NPoint Estimate Control Mean Model p-value Resample p-value Romano-Wolf p-value Q1 1585 0.00 1.00 0.892 0.904 0.989 Q2 1585 0.00 0.46 0.936 0.940 0.989 Q3 1585 0.00 0.50 0.837 0.839 0.989 Q4 1585 -0.01 0.37 0.501 0.555 0.936 Q5 1585 0.01 0.36 0.729 0.743 0.981 Note: This table presents results of the regression analysis outlined in Eq. 2using the response time in minutes as the dependent variable. For every question, we focus on response times below the 99th percentile of the question-specific response time distribution. We present point estimates for the coefficient of interest βj,the mean of the control group, and heteroskedasticity robust uncorrected analytical p-values (model p-values), uncorrected bootstrapped p-values (resample p-values), as well as p-values adjusted for multiple hypothesis testing using 1000 bootstrap replications Source: Own calculation based on survey responses 123 896
Just cheap talk? Investigating fairness preferences Fig. 4 Stability of Fairness Preferences, Correlation. Note: This figure displays the intertemporal correlation between the baseline and follow-up wave (Fig. 4a) and the cumulative distribution function of withinrespondent correlations (Fig. 4b). In Fig. 4a, variables are standardized on the full sample, and the line indicates the line of best fit from a linear regression. In Fig. 4b, variables are standardized at the individual level, and the solid (dashed) line indicates the mean (median) correlation across respondents. Source: Own calculation based on survey responses have a mean of zero and a standard deviation of one. As a result, estimates of σcan be interpreted as intertemporal correlation coefficients. Figure 4a plots the raw standardized data of yij,tagainst yij,t−1, with the fitted line indicating the point estimate of σ. The intertemporal correlation is estimated at 0.61, suggesting sizable stability of distributional choices over time. Figure 4b visualizes the cumulative distribution of intertemporal correlations at the individual level. For each respondent, estimates of σare based on the three repeated questions from baseline and follow-up. More than 80% of respondents display a positive correlation, and more than 65% have a correlation of 0.50 or higher. The mean (median) correlation across respondents is around 0.56 (0.91). These high intra-respondent correlations reaffirm our conclusion that the recovered distributional choices are fairly stable over time for most respondents in our sample. Effect of “real-person treatment” The estimated intertemporal correlations in Eq. 3may be attenuated by measurement error in yij,t−1, i.e., the distributional choices in the baseline wave. Therefore, we can use estimates of σto assess whether incentivized survey questions increase the signal-to-noise ratio in the recovered fairness preferences. If measurement error in yij,t−1was less pronounced in incentivized scenarios, σwould be significantly higher in the treatment group than in the control group. Such a finding would suggest that incentivized survey modules yield less noisy estimates of fairness preferences.9 In Fig. 5a, we replicate Fig. 4a by splitting our sample into the treatment and control groups from the baseline wave. Estimates of σare slightly higher in the treatment (0.66) than in the control group (0.55). The difference of 0.10 is statistically significant at the five percent level (p-value=0.02). In Fig. 5b, we show that the difference in stability is less pronounced 9In the follow-up wave, all questions were hypothetical. Since we use yij,tas outcomes in equation (2), the associated (classical) measurement error will not bias our estimates of σ. 123 897
P. Hufe, D. Weishaar Fig. 5 Stability of Fairness Preferences, “Real-Person Treatment”. Note: This figure displays the intertemporal correlation between the baseline and follow-up wave (Fig. 5a) and the cumulative distribution function of within-respondent correlations (Fig. 5b) separately for treatment and control groups. In Fig. 5a, variables are standardized at the group level, and solid lines indicate lines of best fit from a linear regression. In Fig. 5b, variables are standardized at the individual level, and solid lines indicate mean correlations across respondents. Source: Own calculation based on survey responses when considering correlations at the individual level. The average intra-individual correlation is still slightly higher in the treatment (0.57) than in the control group (0.54). However, the difference of 0.03 is not statistically significant at conventional levels of statistical significance (p-value=0.49). These patterns suggest that incentivized survey modules may yield slightly less noisy estimates of fairness preferences. However, these gains are relatively moderate and may be quickly outweighed by the benefits of an unincentivized survey, e.g., lower cost, the potential to target broader population samples, etc. Robustness We again implement a series of robustness checks to analyze the sensitivity of the previous findings. These robustness checks are summarized in Table 9. First, we check whether intertemporal correlations change when excluding low-quality answers from inattentive respondents who do not pass the training question on the first try. Intertemporal correlations increase slightly but remain very close to our full sample estimate. In an alternative test, we exclude respondents in the tails of the response time distribution. This sample restriction has virtually no effect on the estimated intertemporal correlations. Second, we check whether the estimated intertemporal correlations are especially driven by individuals who always leave the slider close to its original position. In particular, we exclude observations where respondents leave the vignette slider within a two-sided five percentage point band around the initial earnings distribution in the baseline and the followup. Indeed, there seems a slight drop in intertemporal correlations when excluding these respondents. However, we also emphasize that the implemented test is likely too stringent. On the one hand, we exclude respondents who leave the slider unaltered in bad faith. On the other hand, we also exclude respondents with genuine libertarian preferences. Therefore, we interpret the still substantial intertemporal correlation as a positive signal that we can recover stable preferences in areas further away from initial income positions. 123 898
Just cheap talk? Investigating fairness preferences Table 9 Stability of Fairness Preferences, Correlation, Robustness Full Sample Restricted Sample Training Question Response Time Not at Status Quo Aggregate 0.605 (2127) 0.644 (1710) 0.609 (1707) 0.541 (1206) Q1 0.375 (424) 0.389 (338) 0.391 (338) 0.413 (295) Q2 0.328 (426) 0.335 (346) 0.298 (340) 0.436 (199) Q3 0.407 (425) 0.390 (334) 0.399 (337) 0.356 (326) Q4 0.332 (427) 0.337 (341) 0.312 (352) 0.317 (193) Q5 0.382 (425) 0.377 (351) 0.335 (340) 0.497 (193) Note: This table displays intertemporal correlations between the baseline and follow-up waves for the full sample and three restricted samples. The first restriced sample focuses on respondents that passed the training question of the fairness preference module in the baseline wave at the first try. The second restricted sample excludes respondents with high (above p90) and low (below p10) response times of the fairness preference module in the baseline wave. The third restricted sample excludes respondents whose allocated shares are at most 5 percentage points away from the status quo distribution of earnings. Variables are standardized on the sample used in the corresponding regression. Sample sizes are shown in parenthesis Source: Own calculations based on survey responses Third, all previous conclusions hold when calculating intertemporal correlations at the level of individual questions. In our previous discussion, we especially focused on intertemporal correlations at the individual level. This is the appropriate level of analysis since factorial survey designs mostly use intra-respondent variation across multiple vignettes to identify the relevant preferences (Wiswall and Zafar 2018). However, depending on the design, researchers may want to infer preferences from fewer vignettes per individual than in our setting. In Table 9, we, therefore, assess the extreme case where preferences would be identified based on a single question only. In this case, the preference signal is more noisy, translating into lower intertemporal correlations. Nonetheless, even in the extreme case of using only one vignette, the correlations are still substantial, ranging from 0.33 to 0.41 in the full sample. In this section, we have shown that the distributional choices are relatively stable over time. This conclusion is robust to various sensitivity checks, among others, excluding low-quality responses. The presence of incentives slightly decreases the noise in elicited preferences. This decrease in noise, however, is fairly moderate and may be quickly outweighed by the potential benefits of running an unincentivized survey. In summary, the results suggest that hypothetical distribution tasks can yield high-quality data on stable fairness preferences. 5 Nature of fairness preferences In this section, we accompany the main methodological validation of the previous sections by giving some suggestive insights into the nature of elicited fairness preferences. To be sure, this analysis comes with caveats. The primary purpose of this paper is to assess the measurement of fairness preferences in hypothetical settings as compared to the “gold standard” of incentivized experiments. Therefore, we made several methodological choices that prevent a full substantive analysis of the recovered preferences. For example, to maximize statistical power to detect differences between the treatment and control group, we show all respondents the same randomly selected subset of vignettes. Consequently, vignette characteristics are 123 899
P. Hufe, D. Weishaar not equally represented, and correlations may exist among them. These features may affect respondents’ willingness to tolerate inequality and how they incorporate different vignette characteristics into their choices. Therefore, we view the following analysis as a suggestive test of whether the recovered preferences are broadly consistent with findings from the existing literature. With these caveats in mind, we will present the results of this section as follows. First, we analyze the level of inequality implemented by respondents. Second, we will analyze the prevalence of different fairness types in our sample. Lastly, we show the sensitivity of fairness preferences to different characteristics of the evaluated vignette persons. In all analyses, we will focus on unincentivized scenarios from the control group. However, our conclusions remain unaffected when focusing on the incentivized sample—see Appendix Figures C4, C6 and Appendix Table C26 for replications of the main exhibits of this section based on the treatment group. The analyses of this section are exploratory. Therefore, they have not been registered in our pre-analysis plan. Implemented inequality Figure 6compares implemented inequality by respondents to the initial inequality separately for each vignette. The implemented Gini coefficients show substantial variation across vignettes (0.30–0.53). For comparison, Almås et al. (2020) use a representative sample of American respondents to show that they would implement a Gini coefficient of 0.35 (0.54) if the income-generating process were purely based on luck (merit). This suggests the range of implemented inequality across different scenarios in our setting is plausible. Fig. 6 Gini Coefficient. Note: This figure compares implemented Gini coefficients (gray dots) to initial Gini coefficients (green crosses) in each vignette. Gray bars indicate 95 percent confidence intervals. Gini coefficients are calculated at the respondent level as |x−y| x+ywhere x(y) is the amount allocated to vignette person A (B). We focus exclusively on respondents in the control group, i.e., those respondents who faced hypothetical scenarios. Appendix Figure C4 replicates the analysis for respondents in the treatment group. Source: Own calculations based on survey responses 123 900
Just cheap talk? Investigating fairness preferences In Appendix Figure C5, we furthermore illustrate how inequality acceptance varies across respondents with different socio-demographic characteristics. Respondents who are more inequality-accepting tend to be older, more educated, and work longer hours. Those with a lower inequality tolerance tend to be female and non-white. Again, these patterns are broadly consistent with existing literature. For instance, the findings of Almås et al. (2020) indicate that women and individuals with lower educational attainment are less accepting of inequality compared to men and those with higher education. The authors interpret these patterns in the light of potential self-serving biases in fairness preferences. It is reassuring that our survey replicates these patterns as well. Fairness types Experimental literature has focused on estimating the prevalence of different fairness types that can be mapped to fairness principles in the philosophical literature—see Almås et al. (2024b) for a recent overview. On one end of the spectrum is the egalitarian position. Egalitarians consider all inequalities unfair, regardless of how these inequalities come about. Therefore, the egalitarian position prescribes an equal income distribution in any distributive situation. At the opposite end of the spectrum is the libertarian position. Libertarians consider all inequalities fair regardless of how these inequalities come about (Nozick 1974). Therefore, the libertarian position prescribes a distribution of income that corresponds to the initial distribution in any distributive situation. Between these two extreme positions, there are several intermediate positions, such as the responsibility-sensitive positions proposed by Arneson (1989); Cohen (1989); Dworkin (1981a,b). These intermediate positions advocate for distinguishing between different sources of inequality, such as discretionary choices, ability, preferences, or circumstantial factors. We estimate the prevalence of the egalitarian position by calculating the share of respondents who implement equal splits in all vignettes. Similarly, we estimate the prevalence of the libertarian position by calculating the share of respondents who accept initial inequality in all vignettes. When calculating these shares, we allow for “trembling hand” mistakes (Choi et al. 2007). For our baseline estimates, we use two-sided five percentage point bands around the egalitarian and libertarian answers to a vignette and allow respondents to be outside of the corresponding band for at most one vignette without repercussions on their classification as egalitarians or libertarians. We estimate the prevalence of the intermediate position as the remaining share of respondents who are not classified as egalitarians or libertarians. Table 10 shows the results, where the highlighted areas represent our baseline estimates. Around two percent of respondents are classified as egalitarian, whereas around nine percent Table 10 Preference Types Share Egalitarians (%) Share Libertarians (%) Max. abs. difference (pp) 2 5 10 2 5 10 Allow for 0 inconsistent answers 0.00 0.51 1.73 5.68 6.41 7.65 Allow for 1 inconsistent answers 0.00 1.46 3.75 7.89 9.47 13.84 Allow for 2 inconsistent answers 0.89 2.71 7.69 10.75 13.90 22.22 Note: This table presents shares of egalitarians (libertarians) according to consistent choices in all questions of the baseline wave. We also vary the leniency of the classification by allowing for 0, 1, 2 answers that are inconsistent with egalitarian (libertarian) choices. In the baseline (highlighted estimates), we allow for a deviation of +/-5 percentage points and inconsistent choices in one question only. We focus exclusively on respondents in the control group, i.e., those respondents who faced hypothetical scenarios. Appendix Table C26 replicates the analysis for respondents in the treatment group Source: Own calculations based on survey responses 123 901
P. Hufe, D. Weishaar are classified as libertarians. The remaining 89% percent of respondents adopt intermediate positions. Therefore, most respondents adopt fairness positions that vary with the characteristics of the respective vignette. We note that this conclusion does not vary with the leniency with which we accept “trembling hand” mistakes. Even in the most lenient specifications where we allow for two-sided ten percentage point bands and two inconsistent answers, the share of respondents adopting intermediate positions is still 70%. Furthermore, we note that conditional on the adopted rule for “trembling hand” mistakes, the presented estimates for the prevalence of egalitarian and libertarian positions should be interpreted as upper bounds. We only presented respondents with a limited selection of five to nine vignettes. Therefore, in additional questions, the number of divergences from the egalitarian and libertarian positions can stay constant at best but not decrease. The estimated shares of egalitarians and libertarians in the US are smaller than the corresponding shares estimated in Almås et al. (2020). Their estimates classify 15% and 29% of the US population as egalitarians and libertarians, respectively. This difference may be rationalized by variations in how different preference types are identified or by the increased richness of the distributional scenarios in our setting.10 Since the vignettes provide multidimensional information on the earnings-relevant characteristics of the recipients, respondents can express positions that deviate from the polar cases of egalitarian/libertarian fairness preferences in more nuanced ways. Importance of vignette person characteristics In the last step, we analyze the impact of particular earnings-relevant characteristics on the fairness preferences of respondents. To this end, we transform our data as follows. We create a data set where each row represents one person mfrom vignette j. Then, we replicate these data for each respondent iwho made a distributional choice for vignette jand include the corresponding income allocations yim(j) as the outcome variable of interest. Stacking these data, we obtain a panel data set with multiple observations for each vignette person m(j)and each respondent i. We then estimate the following model via ordinary least-squares: ln yim(j)=β1genderm(j)+β2agem(j)+β3educm(j) +β4educparm(j)+β5hoursm(j)+β6ln earnm(j) +θ[earnA(j)+earnB(j)]+im(j). (4) The right-hand side variables in the first two lines of Eq. 4represent the six vignette characteristics considered in our fairness preference module. The associated coefficients β1−β6capture the linear effect of personal characteristics on fair earnings conditional on the remaining vignette characteristics. We control non-parametrically for the total sum of vignette earnings earnA(j)+earnB(j)by including corresponding fixed effects. Thereby, we account for the fact that vignette persons who are paired with high-earning persons in their vignette may mechanically receive higher income allocations. Standard errors are clustered at the respondent level. The interpretation of the estimated coefficients comes with two important caveats. First, we cannot control how respondents associate the displayed vignette characteristics with other unobserved characteristics, such as productivity or job performance. Therefore, β1−β6are composite parameters that capture the reward for a certain characteristic, and the reward for unobserved factors correlated with it while holding all other observed characteristics 10 Almås et al. (2020) identify egalitarians as respondents who distribute resources equally in a situation where initial inequality is driven by productivity. They identify libertarians as respondents who do not redistribute at all when initial inequality is purely based on luck. 123 902
Just cheap talk? Investigating fairness preferences (a) Regression Analysis (b) Text Analysis Fig. 7 Fairness Preferences, Importance of Profile Characteristics. Note: This figure displays how respondents take the vignette characteristics into account for their distributional choices. Figure 7a displays the point estimates and 95 percent confidence intervals from Eq. 4. Figure 7b displays a word cloud from a text analysis of the open-ended question in the fairness preference module. At the end of the module, people were asked about how they came up with their distributional choices and to describe their reasoning in their own words. Based on the text corpus from the open-ended answers, we used natural language processing techniques to rank the frequency of specific terms. We focus exclusively on respondents in the control group, i.e., those respondents who faced hypothetical scenarios. Appendix Table C27 replicates the analysis for different transformations of the outcome variable, and under the inclusion of respondent fixed effects. Appendix Figure C6 replicates the analysis for respondents in the treatment group. Source: Own calculation based on survey responses constant. Second, we disregard non-linear effects across the ordinal vignette characteristics. The methodological choices mentioned at the beginning of this section limit the available variation in our data and prevent us from relaxing this stringent functional form assumption. Figure 7a displays the coefficients from regression Eq. 4along with the corresponding 95% confidence intervals. On the one hand, the results indicate that fair labor market earnings increase with vignette characteristics relevant to labor market performance and that are (partially) under the control of individuals. For example, there are strong positive effects of education and weekly working hours on fair income allocations. This result suggests that fairness preferences in the US are at least partially consistent with normative theories that emphasize the role of discretionary choices in determining fair income shares, (e.g., Konow 2000). Similar conclusions can be drawn for age and initial earnings, if interpreted as proxy indicators for relevant labor market experience and on-the-job performance, respectively.11 On the other hand, the results show that non-discretionary vignette characteristics are not rewarded in fair income allocations. For example, conditional on the other vignette characteristics, the point estimate of gender cannot be distinguished from zero, suggesting that US residents perceive adjusted gender gaps in labor market earnings as unfair. Furthermore, the 11 By controlling for initial labor market earnings, we ensure that we measure the desired reward of all other characteristics, holding initial labor market earnings constant. For instance, a positive coefficient on education implies that conditional on initial market earnings, higher-educated people are considered more deserving than lower-educated people. In a full roll-out of our factorial survey, initial labor market earnings would be uncorrelated with all other vignette characteristics and we could replicate our analyses excluding initial labor market earnings. However, due to the methodological focus of our data collection, we operate with a limited number of vignettes where the characteristics of interest are correlated with each other. To account for such correlations, we include initial labor market earnings in all our estimations. 123 903
P. Hufe, D. Weishaar respondents in our sample assign higher (lower) earnings to individuals with lower (higher) parental education. This finding could be rationalized by the fact that respondents are willing to compensate people for a disadvantaged socio-economic family background. Note, however, that our analysis only presents relative effects of vignette characteristics. Therefore, we cannot distinguish whether respondents compensate individuals for a disadvantaged background (low parental education) or penalize individuals coming from an advantaged family background (high parental education). Given the abovementioned caveats, we interpret these findings as suggestive. However, in Appendix Table C27, we show that all previous conclusions are robust to alternative specifications. In particular, we use preferred income shares or preferred income ratios as outcome variables of interest. While the magnitude of coefficients changes, the direction of effects remains unaltered. Also, our results do not change qualitatively when controlling for respondent fixed effects, i.e., when we only use within-respondent variation to estimate preferred rewards. In addition, while our baseline analysis only uses responses from respondents with hypothetical tasks, the measured rewards are similar in the real-person treatment group (Appendix Figure C6). To substantiate the quantitative evidence, we also use natural language processing techniques to present findings from a text analysis (Ferrario and Stantcheva 2022). At the end of the baseline survey, we asked respondents how they made their distributional choices and allowed them to describe their reasoning in an open-text field. Figure 7b visualizes the text analysis in a word cloud, highlighting the frequency of observed terms. The word cloud shows that respondents put a strong emphasis on working hours, earnings, and the education of the vignette persons when making their distributional choices. The emphasis on these characteristics, therefore, echoes the results from our quantitative analysis. In this section, we have shown that our survey module recovers fairness preferences that are broadly consistent with the existing literature. This conclusion holds for the degree of inequality acceptance, the prevalence of fairness types, and the characteristics determining the extent of fair income allocations. Since the data collection was designed for methodological validation, we urge readers to treat these substantive results cautiously. However, the results point to the ability of our survey module to uncover nuanced fairness positions and to describe fairness preferences in societies more broadly. 6 Conclusion This study validates a novel survey tool designed to measure fairness preferences using realistic yet hypothetical scenarios. We conduct this validation using a two-wave survey covering a representative sample of the US population. Our results demonstrate that fairness preferences are not influenced by the prospect of real-world implementation, even when monetary stakes are high. This conclusion holds true for both the general population and across various demographic subgroups. Moreover, comparing individual responses across the two waves reveals that fairness preferences are stable over time, regardless of whether they originate from hypothetical or incentivized scenarios. We furthermore provide suggestive evidence that the elicited preferences are consistent with established findings on fairness preferences in the US. Therefore, our validation provides compelling evidence that fairness preferences from hypothetical surveys are not “just cheap talk.” Instead, they can yield credible insights into the nature and anatomy of these preferences. We emphasize that these conclusions are context123 904
