Strategic attribute learning
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Benkert, Jean-Michel; Matyskova, Ludmila; Starkov, Egor Working Paper Strategic attribute learning Discussion Papers, No. 24-11 Provided in Cooperation with: Department of Economics, University of Bern Suggested Citation: Benkert, Jean-Michel; Matyskova, Ludmila; Starkov, Egor (2024) : Strategic attribute learning, Discussion Papers, No. 24-11, University of Bern, Department of Economics, Bern This Version is available at: https://hdl.handle.net/10419/312880 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Faculty of Business, Economics and Social Sciences Department of Economics Strategic Attribute Learning Jean-Michel Benkert, Ludmila Matyskova, Egor Starkov 24-11 December, 2024 Schanzeneckstrasse 1 CH-3012 Bern, Switzerland http://www.vwi.unibe.ch DISCUSSION PAPERS
Strategic Attribute Learning * Jean-Michel Benkert, Ludmila Matyskov´a, and Egor Starkov December 13, 2024 Abstract A researcher allocates a budget of informative tests across multiple unknown attributes to influence a decision-maker. We derive the researcher’s equilibrium learning strategy by solving an auxiliary single-player problem. The attribute weights in this problem depend on how much the researcher and the decision-maker disagree. If the researcher expects an excessive response to new information, she forgoes learning altogether. In an organizational context, we show that a manager favors more diverse analysts as the hierarchical distance grows. In another application, we show how an appropriately opposed advisor can constrain a discriminatory politician, and identify the welfare-inequality Pareto frontier of researchers. Keywords: Attributes, Information acquisition, Gaussian distribution, Strategic learning JEL classification: D72, D81, D83 * For valuable comments, we thank: Arjada Bardhi, Nathan Hancart, Toomas Hinnosaar, Alessandro Ispano, Igor Letina, Antoine Loeper, Marc M¨oller, Nick Netzer, Francisco Poggi, Johannes Schneider, Armin Schmutzler, Jakub Steiner, Peter Norman Sørensen, Dezs¨o Szalay, and the audiences at the Universities of Alicante, Bern, Copenhagen, Madrid Carlos III, Mannheim, Zurich, and CERGE-EI, as well as the Nordic Theory Meetings, the Annual Meeting of European Association of Young Economists, the Lisbon Meetings in Game Theory and Applications, the Barcelona Summer Forum, EEA-ESEM, and the Annual Meeting of the VfS. Matyskov´a gratefully acknowledges funding from the Generalitat Valenciana (Prometeo/2021/073) and from the grant PID2022-142356NB-I00 financed by MICIU/AEI /10.13039/501100011033 and FEDER, UE. Benkert: Department of Economics, University of Bern. Matyskov´a: Department of Economics, University of Alicante. Starkov: Department of Economics, University of Copenhagen. Email: [email protected], ludmila.matyskov[email protected], egor.stark[email protected]. 1
1 Introduction Decision-making under uncertainty often involves multiple dimensions of potentially unequal importance. For example, when designing a new product, a company may consider the product’s uncertain reception across various market segments, some of which may be more important than others. Similarly, a politician’s optimal policy may need to address the diverse and uncertain needs of distinct social groups, with some groups potentially carrying more weight than others do. In both examples, decision-makers must learn about the various dimensions, or attributes, influencing their decisions but often lack the resources for a thorough investigation. Consequently, a specialized researcher is often in charge of the learning task. However, the researcher may prioritize attributes differently from the decision-maker and strategically procure information about the unknown attributes to influence the final decision. In this paper, we explore questions pertaining to such strategic concerns in learning about complex decisions. First, what kinds of biases may arise when multiple attributes are at play, and how do they interact? For instance, in an organizational application, we explore two orthogonal biases—hierarchical distance and diversity—and analyze when a manager (decision-maker) prefers diverse analysts (researchers) as a function of hierarchical distance. Second, how can biases mitigate socially undesirable decisions? For instance, in the context of political discrimination, we show how an appropriately opposed advisor (researcher) can constrain a discriminatory politician whose decision affects well-being of various social groups (attributes). More broadly, how does preference misalignment in a multiattribute environment affect learning, when researchers shape not only the extent of learning but also its direction? To address these questions, we develop a framework that captures the strategic interaction between a researcher, who learns about an unknown state of the world, and a decision-maker, who acts based on the resulting information. Crucially, we assume that the state of the world consists of multiple attributes and capture the player’s preference misalignment by allowing for differing weights on the attributes. As such, our main contribution is a tractable model for studying how preference misalignment affects multi-attribute learning and decision-making. 2
In Section 2, we formally introduce our framework. In our model, the attributes are independently and normally distributed and jointly determine the state of the world and thus the players’ preferred decisions. The state can be imperfectly learned by allocating a given budget of tests across different attributes. The more tests are allocated to an attribute, the more informative a signal about it becomes. Players aim to minimize the quadratic loss between the final decision and their bliss point, which is a weighted sum of attributes.1 We begin by analyzing a helpful benchmark in Section 3: a situation in which a single player controls both the decision and the learning, where the optimal learning strategy is the primary focus of our analysis. Theorem 1 shows that—absent strategic motives— the player chooses the test allocation which achieves the highest reduction in residual uncertainty regarding her optimal decision. In particular, optimal learning reflects both the attribute’s relevance to the agent and its prior uncertainty. Put differently, the more important and more uncertain an attribute is, the more tests are allocated to it. Moreover, the budget size determines how many attributes the agent learns about. In Section 4, we return to the primary framework, in which a researcher controls the learning and a decision-maker makes the decision. Our main result (Theorem 2) shows that the researcher’s equilibrium test allocation then coincides with the single-player solution with auxiliary weights, which are different from the actual weights of either player. We show that these auxiliary weights decrease as the misalignment between players grows. If the researcher values an attribute less (or more) than the decision-maker does, he views the decision-maker’s reaction to any information about that attribute as excessive (or insufficient). When the misalignment results in overreaction, it can lead the researcher to abstain from learning about the attribute entirely. Notably, when all of the researcher’s weights are sufficiently low relative to the decision-maker’s weights, the researcher entirely forgoes learning—contrasting with the single-player case, where learning generically takes place. The fact that the researcher’s equilibrium test allocation coincides with the single-player solution with auxiliary weights clarifies how the strategic motive influences the researcher’s learning, reducing its impact to a shift in attribute weights. Thus, by comparing these auxiliary weights to the researcher’s original weights, we can pinpoint how misalignment drives adjustments in the researcher’s learn1Our model is a static variant of the framework in Liang, Mu, and Syrgkanis (2022), who study a single agent’s dynamic learning under attribute correlation. 3
ing strategy—insight that would be harder to achieve if the strategic solution deviated fundamentally from the single-player case. We explore two economical applications, each with a tailored parametrization of player misalignment. These applications demonstrate the versatility of our analytical tools in studying misalignment and yield new insights within each specific context. In our first application, we examine how biases across layers of hierarchy influence learning in organizations. Here, an analyst (researcher) gathers information, and a manager (decision-maker) acts on it. While the manager trivially prefers an analyst whose preferences match hers, we examine the preferred analyst when full alignment is unfeasible. We parametrize misalignment by decomposing the analyst’s preferences into sensitivity (how closely the players align in absolute terms) and distortion (how much the players disagree in relative terms). Proposition 2 shows that the manager may prefer analysts with high distortion when the analysts have low sensitivity. An analyst with low sensitivity views the manager’s reaction to any information as excessive and thus may completely forgo learning. In such cases, the manager then prefers an analyst who is very distorted toward some attribute, as then the manager’s reaction to information about that attribute is no longer viewed as excessive, resulting in some learning. As any learning is better for the manager than no learning, she prefers a strongly distorted analyst over one with little or no distortion. Assuming that the more hierarchical the organization, the bigger the difference in the sensitivity between the managers and the analysts, our results suggest balancing diversity and uniformity (reflected in distortion) in organizations as a function of hierarchical distance. In our second application, we explore the issue of discrimination. A politician (decisionmaker) chooses a policy to meet the uncertain needs of two a priori identical social groups (attributes) but favors one group over the other. With an unchecked politician—who controls both decision-making and learning—such favoritism results in inequality and reduces utilitarian welfare compared to the situation without favoritism. We examine how delegating learning to an advisor (researcher) with different preferences can mitigate these negative effects. Proposition 4 demonstrates that there exists a welfareinequality Pareto frontier of potential advisors that strictly dominates unchecked discrimination. At one end of this frontier is an impartial advisor, who—given the politician’s favoritism—maximizes welfare, but leaves the groups with unequal outcomes (Lemma 3). At the other end of the frontier is an advisor who is suitably partial towards 4
the disadvantaged group. Such an advisor learns precisely enough about this group to counteract the politician’s favoritism, and thus eliminates inequality when appointed (Lemma 4). However, it comes at a cost to welfare, as the politician then relies less on the advisor’s information, and information is thus “wasted.” Notably, the level of advisor’s partiality required to eliminate inequality does not increase with the politician’s favoritism but exhibits a non-monotone relationship. In Section 7, we outline two extensions developed in detail in the online appendix. The first examines multiple researchers in the context of media markets, showing that while competition between media outlets (researchers) leads to polarization in equilibrium, this outcome is actually beneficial for a voter (decision-maker). The second extension addresses uncertainty about the decision-maker’s preferences. We study this within a dual-self model, where a single agent may be either sophisticated or na¨ıve about potential changes in her preferences between the learning and the decision-making stages, finding that na¨ıvete can, in fact, be advantageous. We conclude in Section 8 by discussing our results and prospective avenues for future research. Related literature. Our work is closely related to Bardhi (2024), who examines a strategic multi-attribute learning problem where a project’s payoff is composed of correlated attributes, weighted differently by a decision-maker and a researcher. The researcher selects which attributes to sample (and perfectly learn), given that he can sample a limited number. Bardhi (2024) primarily focuses on the role of correlation between attributes. In contrast, we aim to provide a tractable model to study preference misalignment. First, we focus on independent attributes to separate misalignment effects from correlation. Second, we consider noisy learning, where the choice involves not only which attributes to learn about but also how much to learn about each. Our analysis demonstrates that with as few as two attributes, noisy learning effectively captures the complexity of bias interactions.2 Kirneva (2023), building on Tamura (2018), considers a related multidimensional setting, where the decision-maker can acquire additional costly information on top of what 2More broadly, our framework is related to the Gaussian-sampling literature such as Bardhi and Bobkova (2023), Callander (2011), and Carnehl and Schneider (2024). 5
is provided by the researcher. The need to influence the decision-maker’s learning strategy (and not only the decision) then shapes the researcher’s learning strategy. The researcher may then provide partial information about the attribute on which the players’ preferences are misaligned in order to divert the decision-maker’s learning away from it.3In contrast, we do not allow for independent learning by the decision-maker, creating a distinct strategic environment and shutting down the “attention diversion” channel. Our framework builds on Liang et al. (2022), adapting their dynamic model of multiple attributes to a static form. Furthermore, while they study a single agent’s dynamic learning under attribute correlation, we focus on strategic learning and abstract from correlation to highlight the role of preference misalignment between players. Notably, there is a key parallel in our findings. Both papers show that a “greedy” learning strategy, which achieves the highest reduction in residual uncertainty and is optimal in a single-player model without correlation, remains optimal in the presence of correlation (Liang et al., 2022) and in the presence of strategic motives (our paper), respectively. Our first application, on diversity in organization, connects to the literature on delegated expertise, where a decision-maker delegates learning about the state of interest to a biased researcher.4We explore the effects of different biases between the researcher and the decision-maker, which is related to studies by Ball and Gao (2024), Che and Kartik (2009), and Ilinov et al. (2022). These papers show that delegating learning to a researcher with some preference misalignment can be optimal in single-dimensional settings, as misalignment encourages the acquisition of more costly information. In contrast, we explore a multidimensional setting, where the researcher decides which attributes to learn about and to what extent. Our findings in Section 6.1 reveal that higher misalignment on a certain type of bias can be beneficial, as it may prevent the breakdown of learning caused by misalignment on another, orthogonal, type of bias. Our second application, on discrimination in policymaking, relates to Fosgerau et al. (2023) and Echenique and Li (2023), who explore how decision-makers’ strategic learn3Similar insights—namely, that when the receiver can acquire information, the sender’s choices are primarily driven by a desire to influence the learning process rather than the final decision—have also been obtained by Matveenko and Starkov (2023). 4Pioneered by Demski and Sappington (1987), this literature has seen renewed interest, for instance, see Deimen and Szalay (2019) and Lindbeck and Weibull (2020). 6
ing choices reinforce discrimination.5They show that employers’ discriminatory beliefs shape job candidates’ incentives to invest in their skills, potentially leading to a self-sustaining discriminatory equilibrium. In contrast, we focus on countering discriminatory tendencies by separating learning from decision-making and delegating it to independent advisors. Further, our results in Section 6.2 complement Liang et al. (2024). They study the fairness-accuracy Pareto frontier in the context of an interaction between an egalitarian researcher and a utilitarian decision-maker, where the researcher may coarsen or ban information about certain attributes. In contrast, we examine the welfare-inequality frontier in a setting where the decision-maker is inherently unfair but can be paired with a researcher holding different preferences, who can flexibly allocate informative tests across different attributes. 2 Model 2.1 Setup Players. There are two players, a decision-maker (D, she) and a researcher (R, he). First, the researcher chooses the learning strategy to examine payoff-relevant attributes of the state of the world by choosing how to allocate a budget of tests T. Second, the decision-maker observes the results of these tests and makes a decision. Attributes. There is an unknown state of the world ˜ θ= (˜ θ1,...,˜ θK) consisting of K≥2 attributes.6All attributes ˜ θkfor k= 1, . . . , K are jointly normally distributed with commonly known prior means µ0 k∈Rand prior variances Σ0 k>0; that is, ˜ θk∼ Nµ0 k,Σ0 k.Moreover, the attributes are independent: ˜ θl⊥˜ θjfor l=j. Actions. The decision-maker chooses a decision d∈R. The researcher, given an exogenous budget of tests T > 0, chooses a test allocation τ∈ T := {τ∈RK +: τ1+. . . +τK≤T}, where τkis the amount of tests allocated to learn about attribute 5See Onuchic (2024) for a recent overview of theories of discrimination. 6We denote a random variable and its realization by ˜xand x, respectively. 7
Figure 1. The researcher’s solution: Auxiliary single-player weights 02αR k ˆαk αD k αR k αR k Panel (a) 0 ˆαk αR k αD k/2αD k αD k 45◦ Panel (b) Notes: The figures depict ˆαkas a function of the decision-maker’s weight, αD k, in panel (a), and as a function of the researcher’s weight, αR k, in panel (b), keeping the weight of the other player constant. how misalignment drives adjustments in the researcher’s learning strategy—insight that would be harder to achieve if the strategic solution deviated fundamentally from the single-player case. From the researcher’s perspective, the ex ante marginal value of learning about attribute ˜ θkis proportional to the parameter λk, which—unlike in the single-player case—can be negative. Theorem 2 explains the link between the auxiliary weights ˆαkand the parameters λk. Note that λk(and thus ˆαk) is decreasing in the misalignment between the players’ weights ∆k. When ∆k>0, the researcher views the decision-maker’s reaction to information about ˜ θkas either excessive (αD k> αR k) or insufficient (αD k< αR k), thereby reducing the researcher’s incentives to learn about that attribute. If the reaction is too strong (αD k≥2αR k) or too weak (αD k= 0), the researcher avoids learning about that attribute entirely. If λk≤0 for all k, the researcher abstains from learning, as any information would lead to undesirable overreaction (or no reaction) by the decisionmaker, making the status quo decision the researcher’s preferred outcome. When players disagree on the weight of an attribute, the auxiliary weight differs from the weight of either player. Figure 1 illustrates how ˆαkdepends on αR kand αD k, keeping the weight of the other player fixed. Panel (a) shows that ˆαkalways satisfies ˆαk≤ αR k: the misalignment effectively reduces the importance the researcher assigns to the attribute. Nevertheless, the extent and direction of any distortion in the equilibrium test allocation—compared to the researcher’s non-strategic optimum—then depend on how the ratios of the auxiliary weights differ from the ratios of the researcher’s weights. 14
5 Equivalent payoff specifications This section introduces alternative payoff structures that result in the same equilibrium test allocation as the baseline model. The objective of presenting these frameworks is to offer alternatives that may be better tailored to particular economic applications, thus demonstrating the adaptability of our model and the analytical tools detailed in Sections 3 and 4. Baseline model. Introduced in Section 2, the baseline model features a decisionmaker taking a single decision, d∈R. Given a decision dand a realized state θ= (θ1, . . . , θK), the utility of player i=R, D is ui(d, θ) = −(d−bi(θ))2=− d−X k αi kθk!2 .(15) Framework A. In this framework, the decision-maker also takes a single decision, d∈R. Given a decision dand a realized state θ= (θ1, . . . , θK), the utility of player i=R, D is ui A(d, θ) = −X k αi k(d−θk)2,(16) where Pkαi k= 1 for both players. Framework B. This framework features the decision-maker simultaneously taking K distinct decisions d1, . . . , dK∈R. Given decisions d= (d1, . . . , dK) and a realized state θ= (θ1, . . . , θK), the utility of player i=R, D is ui B(d, θ) = −X k (dk−αi kθk)2.(17) In all three models: (i) players aim to minimize a loss given by a quadratic distance, and (ii) the weights αidetermine the relative importance of different attributes. Beyond these similarities, the frameworks differ in how the players aggregate losses across attributes and decisions. Despite these differences, the following proposition (proved in 15
Online Appendix C.1) shows that Frameworks A and B are equivalent to the baseline model in terms of the equilibrium test allocation. Proposition 1. Given the decision-maker’s weights αD, the researcher’s weights αR, the test budget T, and the prior distribution of the state ˜ θ, the researcher’s equilibrium test allocations in the baseline model, Framework A, and Framework B are identical. The displayed flexibility allows us to study a rich set of economic problems. For instance, we use Framework A to study discrimination (Section 6.2) and media polarization (Online Appendix B.1). To illustrate an application of Framework B, consider a portfolio choice problem, where an investor makes decisions dkon how much to invest in each asset k= 1, . . . , K. Here, ˜ θkrepresents the uncertain future return on asset k. The investor evaluates the future returns objectively (αD k= 1 for all k), while an advisor may have an incentive to steer the investor towards some assets and away from others: αR k= 1 for some k. 6 Applications In this section, we illustrate how our framework can be used to study preference misalignment in two applications, each with a tailored parametrization of the misalignment. First, we adopt an organizational setting to showcase how biases across hierarchical layers impact learning in organizations. Next, we analyze a policy-making model with a discriminatory politician and explore how introducing independent researchers to inform policy decisions can mitigate negative effects of discrimination on welfare and inequality. In the applications, we adopt the simplest specification, a two-attribute case, which is already sufficiently rich to demonstrate the tension between the players arising in the multi-attribute context. Throughout, we assume that both weights of the decision-maker are strictly positive, αD k>0 for k= 1,2. 6.1 Diversity in organizations In this application, we consider a stylized model of an organization, comprising a manager (decision-maker) and an analyst (researcher). This setup reflects an organizational 16
Figure 2. The bias decomposition of an analyst αR 2 αR 1 β γ γ > 0 αD Notes: The dashed line represents different αRvalues where γis fixed at a positive value and various values of βare considered. hierarchy where the manager holds authority over the analyst and takes decisions based on the analyst’s information. Given a diverse pool of analysts with varying weights, the manager selects one to provide information. The key question is which analyst the manager will choose. First, it is immediate that the manager’s ideal choice would be an analyst whose weights match hers. However, our focus is on identifying the type of analyst the manager would prefer when, due to an inherent hierarchical structure, perfect alignment is unachievable. To analyze this question, we begin by establishing a suitable parametrization of the misalignment between the players’ weights. We fix the manager’s weight vector, αD= (αD 1, αD 2), and define an orthogonal vector ¯αD:= (−αD 2, αD 1). Next, we express the analyst’s weight vector αRas a linear combination of the manager’s weight vector and its orthogonal counterpart as αR(β, γ) := βαD+γ¯αD(18) for some β∈R+and γ∈Γ(β) := h−βαD 2 αD 1 , β αD 1 αD 2i.13 Here, βrepresents the absolute bias, referred to as sensitivity (to new information), and γrepresents the relative bias, referred to as distortion. We say that the analyst becomes more sensitive when βincreases.14 13The constraint γ∈Γ(β) follows from the requirement that αR k≥0 for both k= 1,2. As noted in Footnote 8, this restriction is without loss and only for ease of exposition. 14Note that when β= 1, the analyst has the same sensitivity as the manager. 17
Figure 3. Equilibrium test allocation Panel (a) αR 2 αR 1 αD 1 2 αD 2 2 β no testing for β≤1/2 manager’s first best for β > 1/2 αD Panel (b) αR 2 αR 1 αD 1 2 αD 2 2 αD L1 L2L12 L0 γ τ∗ 1↓, τ∗ 2↑ γ Notes: Panel (a) depicts the undistorted analyst’s equilibrium test allocation: no testing if β≤1/2, and manager’s first best if β > 1/2. Panel (b) details the equilibrium test allocation for different (β, γ) pairs: the analyst learns (i) only about attribute ˜ θ1if αR∈L1; (ii) only about attribute ˜ θ2if αR∈L2; (iii) about both attributes if αR∈L12; (iv) about no attribute if αR∈L0. The black thick lines captures analysts with increasing distortion towards attribute ˜ θ2from an undistorted analyst (dashed line) for fixed sensitivity at β < 1/2 (line closer to the origin) and β > 1/2 (line further away from the origin). We further say the analyst becomes more distorted when γ≥0 increases or γ≤0 decreases. Specifically, when γ > 0 (γ < 0), the analyst exhibits bias towards attribute ˜ θ2(attribute ˜ θ1). When γ= 0, we call the analyst undistorted. Figure 2 visualizes this bias decomposition.15 In the context of our organizational application, we posit that analysts who are separated by more layers of hierarchy from the manager are less sensitive. First, we examine how the manager’s payoff changes with the analyst’s distortion γ. Proposition 2 below shows that the manager is worse off as the analyst becomes more distorted, provided the analyst is sufficiently sensitive (high β). However, the converse is true when the analyst is too insensitive (low β). Proposition 2. Fix the analyst’s level of sensitivity β∈R+. (i) If β≤1/2 (the analyst is too insensitive), then the manager’s expected equilibrium 15Alternative decompositions of the analyst’s weights in terms of “absolute” and “relative” bias exist. For instance, one could present both αRand αDin polar coordinates and let βbe the difference in distances from the origin and γbe the difference in angles. The insights from our Propositions 2 and 3 would extend to such alternative representations. 18
payoff weakly increases as the analyst becomes more distorted. (ii) If β > 1/2 (the analyst is sufficiently sensitive), then the manager’s expected equilibrium payoff weakly decreases as the analyst becomes more distorted. The manager and an undistorted analyst share the same weight ratios, so the analyst’s equilibrium learning strategy aligns with the manager’s first-best solution, provided the analyst’s sensitivity is high enough to motivate learning (β > 1/2). In this case, if the analyst’s distortion increases, his equilibrium learning strategy deviates further from the manager’s first best, making the manager worse off. Conversely, when β≤1/2, an undistorted analyst abstains from learning, perceiving the manager’s reaction to any information as excessive. Then, if the analyst’s distortion increases, his preferences align more closely with the manager’s on one particular attribute, potentially prompting him to learn about it. While the added distortion does increase misalignment on the other attribute, it does not negatively impact learning since the initial misalignment was already sufficient to deter the analyst from learning about it. Overall, since any acquired information is preferable to none, the manager ultimately benefits. The results, which are illustrated in Figure 3, highlight the importance of both diversity and uniformity within organizations. Our findings suggest a strategic approach to organizational structure if employees lower down the hierarchy are less engaged and thus less sensitive to new information than the managers higher up are. In organizations with many layers of hierarchy and thus potentially significant differences in engagement, it is beneficial to have substantial diversity at the lower levels, as indicated by relatively high degrees of distortion |γ|compared to the managers. Conversely, in smaller, less hierarchical organizations, a more uniform workforce with minimal distortion is preferable as the variation in engagement levels among managers and analysts is less problematic in these settings. The above underscores the nuanced role of distortion in shaping the analyst’s alignment with the manager’s objectives. However, when we shift focus from distortion to sensitivity, the effects become more straightforward. Proposition 3. Fix the analyst’s distortion γ∈R. Then, the manager’s expected equilibrium payoff is weakly increasing in the analyst’s sensitivity β. The manager always prefers a more sensitive analyst, as this leads to an equilibrium 19
test allocation that aligns more closely with her first-best solution. In other words, the manager seeks highly engaged employees who are sensitive to new information relevant to the organization. To prove this result, we show that for any distortion γ, the equilibrium test allocation τ∗increasingly aligns with the manager’s first best as βincreases, effectively neutralizing the impact of distortion. 6.2 Discrimination, welfare and inequality In this application, we explore a scenario where a politician decides on a policy affecting two social groups and favors one group. We investigate how appointing an advisor with different preferences, who strategically curates the information provided to the politician, can mitigate the negative impact of this type of discrimination on utilitarian welfare and inequality. Suppose there are two social groups k= 1,2. Each group khas an unknown optimal policy ˜ θk∼ N(0,1), and so the two groups are a priori identical. A politician (decisionmaker) decides on a common policy, d∈R. The utility of group kis given by uk(d, θk) = −(d−θk)2. A budget T= 1 of tests is available to inquire about the optimal policies of both groups. An advisor (researcher) chooses a test allocation τ= (τ1, τ2)∈ T. The learning process and notation are the same as in the baseline model. The politician and the advisor are utilitarian with particular weights. The politician cares about both groups but favors group k= 1, her payoff being uD(d, θ1, θ2;δ) = 1 + δ 2 |{z} αD 1(δ):= u1(d, θ1) + 1−δ 2 |{z} αD 2(δ):= u2(d, θ2) for a given level of discrimination δ∈(0,1). On the other hand, the advisor (possibly) favors group k= 2, his payoff being uR(d, θ1, θ2;p) = 1−p 2 |{z} αR 1(p):= u1(d, θ1) + 1 + p 2 |{z} αR 2(p):= u2(d, θ2) for a given level of partiality p∈[0,1). When p= 0, indicating the advisor cares equally about both groups, we call the advisor impartial. 20
Our focus in this application is on welfare W(p, δ), defined as the sum of ex ante equilibrium expected payoffs of both groups (and is hence utilitarian with equal weights), and inequality I(p, δ), defined as their difference (and is hence an egalitarian measure): W(p, δ) := Ehu1˜ d∗(τ∗),˜ θ1i+Ehu2˜ d∗(τ∗),˜ θ2i,(19) I(p, δ) := Ehu1˜ d∗(τ∗),˜ θ1i−Ehu2˜ d∗(τ∗),˜ θ2i.(20) In the expressions above, the right-hand side depends on pand δvia the advisor’s choice of test allocation τ∗and the politician’s policy choice ˜ d∗(τ∗), which can be more fully described as ˜ d∗(τ∗(p, δ), δ). When I(p, δ) = 0, we say there is equality in equilibrium.16 As a benchmark, we consider the case of unchecked discrimination, when the politician controls both learning and decision-making. Here, welfare and inequality are defined analogously to equations (19) and (20), with the politician selecting the optimal test allocation instead of an advisor doing so. In this context, increased discrimination negatively impacts both welfare and inequality: welfare declines while inequality rises. We aim to understand how appointing an advisor can mitigate these adverse effects. We analyze a scenario where welfare and inequality are influenced solely through the advisor’s role in the learning stage, while the politician retains full control over decisionmaking. First, we ask which advisor maximizes welfare. Since both welfare and the preferences of the impartial advisor are utilitarian with equal weights, it immediately follows that appointing an impartial advisor maximizes welfare, regardless of the politician’s level of discrimination. In contrast, appointing a partial advisor or allowing unchecked discrimination yields lower welfare. The next lemma describes the (welfare-maximizing) learning strategy of the impartial advisor. Lemma 3. For every level of discrimination δ, welfare is maximized by appointing the impartial advisor (p= 0). Moreover, the impartial advisor’s equilibrium test allocation τ∗(0, δ) = (1/2,1/2) is independent of δ. As the level of discrimination δincreases, two effects emerge. From the advisor’s perspective, information about group 1 is increasingly overemphasized in the politician’s 16This model fits Framework A in Section 5. By Proposition 1, Theorem 2 can be used to find the equilibrium test allocation. 21
decision, while information about group 2 is increasingly underused. Each effect independently reduces the advisor’s incentive to learn about the respective group. However, for the impartial advisor, these effects are equal in magnitude and cancel each other out. This leads him to be unresponsive to changes in δand always choose an equal test allocation. Consequently, as δincreases, inequality increases under an impartial advisor. The above raises the question of whether a partial advisor can counterbalance the politician’s discrimination and restore equality. If so, does the required level of partiality p increase with the level of discrimination δ? The next lemma addresses these questions. Lemma 4. For every politician’s level of discrimination δ, there exists a unique advisor’s level of partiality ˆp(δ)>0 that ensures equality in equilibrium: I(ˆp(δ), δ) = 0. The function ˆp(δ) is continuous and non-monotone: there exists a unique δ∈(0,1) such that ˆp(δ) is strictly increasing for δ < δ and strictly decreasing for δ > δ. As noted earlier, increases in δreduce the advisor’s incentive to learn about both groups, as information about group 1 is increasingly overemphasized in the politician’s decision, while information about group 2 is increasingly underused. For an advisor with partiality p > 0, the former effect dominates the latter, prompting him to learn more about group 2 as δincreases. Moreover, the disparity between the two effects is greater for advisors with higher partiality p, as they prioritize group 2 over group 1 more. As such, advisors with higher partiality adjust their learning strategies more sharply in response to changes in δthan do those with lower partiality. On the other hand, as the level of discrimination δincreases, restoring equality requires more wasted information: the advisor must learn less about group 1 and more about group 2, even though the politician becomes less responsive to information about group 2. Hence, the politician’s decision increasingly relies on her prior beliefs—where there is no disagreement between groups—rather than the advisor’s information. As the amount of “effectively” utilized information diminishes with higher δ, the adjustments required in the learning strategy to restore equality become progressively smaller. In summary, two effects are at play: advisors with higher partiality prespond more strongly to changes in δ, while the degree of adjustment in tests to achieve equality diminishes with higher δ. These dynamics then lead to the non-monotonicity result in Lemma 4. Additionally, since restoring equality involves wasting more information as δ 22
increases, it results in a welfare loss. Indeed, welfare W(ˆp(δ), δ) decreases with δ. As we have seen, welfare is maximized with an impartial advisor, but then inequality increases with δ. Conversely, with partial advisors who restore equality, welfare then decreases with δ. Hence, welfare maximization and inequality minimization are misaligned objectives, as captured in the following proposition. Proposition 4. For any level of discrimination δ, a welfare-inequality Pareto frontier is formed by advisors having partiality levels p∈[0,ˆp(δ)], where both welfare and inequality strictly decrease in p. Moreover, welfare with the equality-restoring advisor ˆp(δ) is strictly higher than under unchecked discrimination. Proposition 4 illustrates a trade-off between welfare and inequality. Reducing inequality requires wasting more information, achieved by advisors with higher levels of partiality, while increasing welfare requires using information more efficiently, achieved by less partial advisors. This dynamic gives rise to the Pareto frontier formed by advisors ranging from impartial to those restoring equality. Moreover, any advisor along this frontier unambiguously improves both welfare and inequality compared to unchecked discrimination. Notably, even an advisor who restores equality results in higher welfare than does unchecked discrimination. This is because welfare is concave in test allocation, and the learning strategy of an unchecked politician is heavily skewed towards learning about the needs of group 1, whereas an “equalizing” advisor learns more evenly about both groups (even though his learning is skewed towards group 2). 7 Extensions In this section, we outline two extensions, presented in more detail in Online Appendix B. The first examines media polarization by introducing two researchers into the model, representing media outlets competing to influence a voter (decision-maker). Each outlet and the voter seek implementation of their preferred policy mix on two policy issues. In the model, first, the voter allocates her attention between the two outlets. Second, the outlets simultaneously decide how much coverage to devote to the two policy issues. The combined attention and coverage provide the voter with signals about the two policy issues. Finally, the voter casts her ballot. We provide conditions under which media 23
and (iii) cov(˜ θk,˜ θj) = 0 for all k=jsince the attributes are independent. A.2.2 Proof of Theorem 2 Recall that ψD(τ) = σ2,D 0−ˆσ2,D(τ) is the variance of the decision-maker’s decision, where ˆσ2,D(τ) is given by (8). Then, from Lemma 2 and dropping terms that are constant in τ, the researcher’s value is given by Vst(τ) = 2 cov ˜ d∗(τ),˜ bR+ ˜σ2,D(τ) = 2 X k αD kαR kΣ0 kτkˆ Σk(τk) + X kαD k2ˆ Σk(τk) =X k2αD kαR kΣ0 kτk+αD k2Σ0 k 1 + τkΣ0 k subject to the non-negativity and the budget constraints. The partial derivatives of Vst(τ) with respect to τk≥0 are ∂V st(τ) ∂τk =λkΣ0 k 1 + τkΣ0 k2 =λk1 Σ0 k +τk−2 (30) where we denote λk:= αD k2αR k−αD kfor every k. Note that when λk≤0, the function Vst(τ) is (weakly) decreasing in τk. Then the solution dictates τ∗ k= 0 for such k.21 Hence, the maximizers of function Vst(τ) are the same as the maximizers of an adjusted function Vst,adj(τ) := Vst(τ) where we set αD k=αR k= 0 for all kwith λk≤0 (and hence Vst,adj(τ) is independent of τkfor such k). For every k, denote ˆαk=pmax{0, λk}and, w.l.o.g., let us relabel the attributes such that ˆα1Σ0 1≥. . . ≥ˆαKΣ0 K. By noting that (i) the function Vst,adj(τ) is concave, (ii) for every k, the partial derivative ∂V st,adj (τ) ∂τkis the same as the partial derivative ∂V (τ) ∂τk of a single-player objective function (21) in which we set αk= ˆαkfor all k, and (iii) the constraints are the same for both (the adjusted strategic and the single-player) maximization problems, we obtain the statement. 21When λk≤0 for all kand there exists at least one jsuch that λj= 0, we use the assumption that in case of indifference the researcher abstains from learning, thus yielding τ∗= (0,0,...,0). Otherwise, setting τ∗ k= 0 whenever λk≤0 is uniquely optimal: (i) either there exists at least one jwith λj>0 (and thus the function Vst(τ) is strictly increasing in τj), or (ii) λk<0 for all k(and thus the function Vst(τ) is strictly decreasing in every τk). 30
A.3 Proofs for Section 6.1: Diversity in organizations We first state four lemmas (proved in Online Appendix C.2) that we use to prove both Propositions 2 and 3. Lemma 5 and Lemma 6 characterize the equilibrium learning strategy τ∗(β, γ) as a function of γwhen β≤1/2 and β > 1/2, respectively. Lemma 7 establishes that the decision-maker’s interim expected payoff is strictly increasing and concave in (τ1, τ2). Lemma 8 derives the researcher’s equilibrium test allocation for γ= 0. Lemma 5. Let αD 1≥αD 2>0 and assume the analyst’s weight vector αRis given by decomposition (18). For any β∈[0,1/2] there exist unique γ1(β), γ2(β)∈Γ(β) such that γ1(β)≤0≤γ2(β) and: τ∗(β, γ) = (0, T) if γ > γ2(β), (0,0) if γ∈[γ1(β), γ2(β)], (T, 0) if γ < γ1(β). (31) Further, there exist β1, β2∈Rwith 0 < β2≤β1<1/2 such that γ1(β)∈int (Γ(β)) if and only if β∈(β1,1/2], and γ2(β)∈int (Γ(β)) if and only if β∈(β2,1/2]. Lemma 6. Fix the manager’s weight vector αD= (αD 1, αD 2) with αD k>0 for both k= 1,2. Let the analyst’s weight vector αR(β, γ) be given by the decomposition (18). Fix β > 1/2. Then there exist γI(β), γII(β)∈int(Γ(β)) with γI(β)< γII(β) such that the analyst’s equilibrium test allocation is τ∗(β, γ) = (T, 0) if γ≤γI(β) (˜τ1(β, γ),˜τ2(β, γ)) if γ∈(γI(β), γII(β)) (0, T) if γ≥γII(β) (32) where ˜τ1(β, γ) := ˆα1(β, γ)Σ0 1−ˆα2(β, γ)Σ0 2 Σ0 1Σ0 2(ˆα1(β, γ) + ˆα2(β, γ)) +ˆα1(β, γ) ˆα1(β, γ) + ˆα2(β, γ)T ˜τ2(β, γ) :=T−˜τ1(β, γ) (33) 31
and ˆα1(β, γ) := rmax n(2β−1) αD 12−2γαD 1αD 2,0o ˆα2(β, γ) := rmax n(2β−1) αD 22+ 2γαD 1αD 2,0o(34) Moreover, ˜τ1(β, γ) (˜τ2(β, γ)) is strictly decreasing (strictly increasing) in γfor γ∈ (γI(β), γII(β)). Lemma 7. The decision-maker’s interim expected equilibrium payoff is a strictly increasing and strictly concave function of (τ1, τ2). Lemma 8. Fix the test budget T > 0 and the decision-maker’s weight vector αD∈R2 ++. Suppose γ= 0 (the agent is not distorted). (i) If β≤1/2 (the agent is too insensitive), then the agent optimally chooses no testing: τ∗(β, 0) = (0,0). (ii) If β > 1/2 (the agent is sufficiently sensitive), then the agent optimally chooses the decision-maker’s most preferred test allocation: τ∗(β, 0) = τ∗(1,0). A.3.1 Proof of Proposition 2 Suppose the premise of Proposition 2 holds. Proposition 2(i) is a direct implication of Lemma 5. The manager’s interim expected payoff is strictly increasing in τ1and τ2 by Lemma 7. Hence, the manager always strictly prefers any positive test allocation τ= (0,0) to abstaining from learning. Now take β > 1/2. Lemma 8 shows that (given the constraint τ∈ T) the manager’s interim expected payoff VD(τ) is maximized when the agent’s relative bias is γ= 0. Furthermore, since VD(τ) is strictly increasing in τ1and τ2, the constraint τ1+τ2≤Tmust be binding at the manager’s most preferred test allocation. Furthermore, Lemma 6 shows when β > 1/2, the agent’s equilibrium test allocation τ∗(β, γ) is such that τ∗ 1(β, γ) and τ∗ 2(β, γ) are weakly decreasing and weakly increasing in γ, respectively, and τ∗ 1(β, γ) + τ∗ 2(β, γ) = T. The claim of Proposition 2(ii) then follows from the strict concavity of VD(τ) (shown in Lemma 7). 32
A.3.2 Proof of Proposition 3 Suppose γ= 0. Then, Lemma 8 shows that the researcher’s learning strategy is to not learn for β≤1/2 and then to implement the decision-maker’s preferred learning strategy. As the decision-maker’s interim expected payoff VD(τ) is strictly increasing and concave in τ1and τ2(Lemma 7), this implies that VD(τ) is weakly increasing in β for γ= 0. Suppose from this point onwards that γ > 0 (case γ < 0 is completely analogous). Then, for the researcher’s weights to be weakly positive, we must have β≥β:= γαD 2 αD 1 . The equilibrium test allocation is given by Theorems 1 and 2, where we can express ˆαkfor k= 1,2 as (34),22 and so λ1= (2β−1)(αD 1)2−2γαD 1αD 2, λ2= (2β−1)(αD 2)2+ 2γαD 1αD 2. Note that λkfor k= 1,2 is strictly increasing in β. Therefore, there exist β1and β2such that ˆαk= 0 for β≤βk, and ˆαkis strictly positive and strictly increasing for β > βk. These values are given by the respective roots of λk= 0:23 β1:= 1 2+γαD 2 αD 1 , β2:= 1 2−γαD 1 αD 2 . Observe that β2<1/2< β1and β < β1. Theorems 1 and 2 then imply that if β > β2, then ∃k:τ∗ k(β, γ)>0 and τ∗ 1(β, γ) + τ∗ 2(β, γ) = T(since at least one of the auxiliary weights ˆαkis then strictly positive). Therefore, we can conclude that if β > β2, then τ∗(β, γ) = (0, T) for all β∈[β, β1], and if β < β2, then τ∗(β, γ) = (0,0) for β∈[β, β2], (0, T) for β∈(β2, β1]. As VD(τ) is strictly increasing in τ∗ 2, it follows that VD(τ) is weakly increasing in βin either of the two cases. Consider now the case when β > β1, so we have ˆα1,ˆα2>0. Theorems 1 and 2imply 22While Lemma 6 is only stated for β > 1/2, expression (34) is well-defined for all β > β. 23Note that these βkare different from the ones defined in the proof of Lemma 5. 33
that the equilibrium test allocation is given by τ∗ 1= max (0,min (ˆα1 ˆα2Σ0 1−Σ0 2 Σ0 1Σ0 2(ˆα1 ˆα2+ 1) + ˆα1 ˆα2 ˆα1 ˆα2+ 1T, T)), τ∗ 2=T−τ∗ 1. (35) Note that the expression above (and, hence, VD(τ)) only depends on βthrough ˆα1 ˆα2. Observe that lim β→+∞ ˆα1 ˆα2=αD 1 αD 2 , so that lim β→+∞τ∗(γ, β) = τ∗(αD), so the equilibrium test allocation converges to the decision-maker’s optimal (payoff-maximizing) test allocation as β→ ∞. Further, ˆα1 ˆα2is strictly increasing in βin the case considered: ∂ ∂β ˆα1 ˆα2 =1 ˆα2 ∂ˆα1 ∂β −ˆα1 ˆα2 2 ∂ˆα2 ∂β =ˆα1 ˆα2 αD 1 ˆα12 −αD 2 ˆα22!>0, where the inequality follows from ˆαk=√λkand αD 1 αD 22 >ˆα1 ˆα22 =(2β−1)(αD 1)2−2γαD 1αD 2 (2β−1)(αD 2)2+ 2γαD 1αD 2 =(αD 1)2−2γ 2β−1αD 1αD 2 (αD 2)2+2γ 2β−1αD 1αD 2 . Together with the limit result above, this implies that ˆα1 ˆα2and τ∗(γ, β)monotonically converge to αD 1 αD 2 and τ∗(αD), respectively, as β→ ∞. As VD(τ) is concave in τ(see Lemma 7) and maximized by τ∗(αD), it must be weakly increasing in β∈(β1,+∞). A.4 Proofs for Section 6.2: Discrimination, welfare and inequality A.4.1 Proof of Lemma 3 For the first part of the statement, observe that the maximization problem of an impartial advisor is max τ∈T 1 2Ehu1˜ d∗(τ),˜ θ1+u2˜ d∗(τ),˜ θ2i. By the definition of welfare, an impartial advisor thus chooses the welfare-maximizing test allocation in equilibrium, since the objectives are scaled versions of each other. Let us turn to the second part of lemma. Fix δ∈(0,1). Using Theorem 2, we can solve 34
the discrimination model as a solution to a single-agent model with weights ˆα1(p, δ) := 1 2p(1 + δ)(1 −2p−δ) if p < 1−δ 2 0 otherwise (36) ˆα2(p, δ) := 1 2p(1 −δ)(1 + 2p+δ) if p > −1+δ 2 0 otherwise (37) Let p= 0. Then ˆα1(0, δ) = ˆα2(0, δ) = 1 2√1−δ2. The equilibrium test allocation depends on the weights only through their ratios ˆα1(0,δ) ˆα1(0,δ)+ˆα2(0,δ)=1 2and ˆα2(0,δ) ˆα1(0,δ)+ˆα2(0,δ)= 1 2. These ratios are independent of the level of discrimination δ, which completes the proof. A.4.2 Proof of Lemma 4 We first define several new notions. Given the politician’s discrimination δ∈(0,1) and the advisor’s partiality p≥0, let τ∗(p, δ)=(τ∗ 1(p, δ), τ∗ 2(p, δ)) denote the respective equilibrium test allocation. Note it holds τ∗ 1(p, δ) + τ∗ 2(p, δ) = 1, i.e., the budget is always exhausted in equilibrium.24 Given a particular test allocation τand equilibrium decision strategy, let Vk(τ, δ) denote the ex ante expected utility of group k, i.e., Vk(τ, δ) = Eh−(˜ d∗D(τ;δ)−˜ θk)2i. Under the additional constraint that the budget is exhausted, and so the test allocation satisfies τ1= 1 −τ2, let us define the difference between the expected utilities of group 1 from group 2 as ∆(τ2, δ) := V1((1 −τ2, τ2), δ)−V2((1 −τ2, τ2), δ).(38) Note that inequality is then given by I(p, δ) = |∆(τ∗ 2(p, δ), δ)|. 24This result follows from the assumption that the sums of the weights of each player are equal: PkαR k(p) = PkαD k(δ). Thus, it cannot simultaneously hold αR k(p)≤αD kfor both k(which is necessary and sufficient for ˆα1(p, δ) = ˆα2(p, δ) = 0), and there is thus always learning in equilibrium. 35
We now state two intermediary results (proved in Online Appendix C.3), used to prove Lemma 4. Lemma 9. For each δ∈(0,1), there exists a threshold partiality level paux(δ)∈0,1−δ 2 such that the equilibrium learning strategy satisfies τ∗ 2(p, δ) = 1 for advisors with p≥ paux(δ) and τ∗ 2(p, δ) = τaux 2(p, δ) := 2ˆα2(p, δ)−ˆα1(p, δ) ˆα1(p, δ) + ˆα2(p, δ)for advisors with p∈[0, paux(δ)], where ˆα1(p, δ) and ˆα2(p, δ) are given by equations (36) and (37). Lemma 10. For each δthere exists a unique value of ˆτ2(δ)∈(1/2,1) for which it holds ∆(ˆτ2(δ); δ) = 0. Given the politician’s level of discrimination δand her equilibrium decision strategy, the value ˆτ2(δ) is the amount of tests allocated to learn about group 2 that guarantees that the expected utilities of both social groups are the same. From the proof of Lemma 10, we obtain ∆(τ2, δ) = (1 + δ)(1 −τ2) 1+1−τ2−(1 −δ)τ2 1 + τ2 . Solving ∆(ˆτ2(δ), δ) = 0 for ˆτ2(δ)≥0, we obtain ˆτ2(δ) = −(1 −δ) + √1+3δ2 2δ.(39) From Lemma 9, the equilibrium test allocation satisfies: τ∗ 2(p, δ)=1/2 at p= 0, τ∗ 2(p, δ) = 1 at p≥paux(δ), τ∗ 2(p, δ) is continuous on p∈[0,1] and it is strictly increasing in p∈[0, paux(δ)). Since ˆτ2(δ)∈(1/2,1), we thus have that for any level of discrimination δ∈(0,1) there exists a unique partiality level ˆp(δ)∈(0, paux(δ)) such that ˆτ2(δ) = τ∗ 2(ˆp(δ), δ). Solving this equation, we obtain ˆp(δ) = δ1 √1+3δ2−1 2, 36
which is continuous on δ∈(0,1). Taking the first derivative, we get ˆp′(δ) = 1 (1 + 3δ2)3/2−1 2 >0 if δ < δ = 0 if δ=δ <0 if δ > δ where δ:= 1 31−21/2+ 22/3. A.4.3 Proof of Proposition 4 The first part of the statement follows directly from the following Lemma 11, which is proved in Online Appendix C.3. Lemma 11. Welfare W(p, δ) and inequality I(p, δ) are both (i) continuous in p; and (ii) strictly decreasing on p∈(0,ˆp(δ)), where ˆp(δ) restores equality. Furthermore, welfare strictly decreases and inequality strictly increases on p∈(ˆp(δ), paux(δ)), where paux(δ) is characterized in Lemma 9; and they are both constant on p≥paux(δ). Hence, the advisor with ˆp(δ) strictly improves both welfare and inequality outcomes compared to any advisor with p > ˆp(δ), and hence the advisors with p > ˆp(δ) do not constitute the Pareto frontier. For the second part of the statement, we build on the following Lemma 12, which is proved in Online Appendix C.3. For a given δ, let (1 −¯τ2(δ),¯τ2(δ)) denote the optimal learning strategy of a politician in the unchecked discrimination, and let (1−ˆτ2(δ),ˆτ2(δ)), where ˆτ2(δ) is given by equation (39), denote the equilibrium test allocation chosen by the advisor ˆp(δ). Lemma 12. Fix δ. Then it holds |τ2(δ)−1/2|>|ˆτ2(δ)−1/2|. 37
Let ω(τ2, δ) := V1((1 −τ2, τ2), δ) + V2((1 −τ2, τ2), δ) =−3−δ2+1/2(1 + δ2) + (1 + δ)(1 −τ2) 2−τ2 +1/2(1 + δ2) + (1 −δ)τ2 1 + τ2 denote the sum of the expected utilities of the two groups as a function of test allocation τ= (τ1, τ2) under the constraint that the budget is fully used, τ1= 1 −τ2, and where Vk(τ, δ) is given by equation (C.3.2). Since the budget is exhausted in equilibrium (as shown in the proof of Lemma 4), welfare is thus given by W(p, δ) = ω(τ∗ 2(p, δ), δ). Observe the following symmetry property: ω(τ2, δ) = ω(1 −τ2, δ). Moreover, for any δ∈(0,1), the function ω(τ2, δ) is strictly concave and maximized at τ2= 1/2, when the test budget is split equally between the two groups. Therefore, the function ω(τ2, δ) is strictly decreasing in |τ2−1/2|. Lemma 12 then implies that welfare with an advisor ˆp(δ) is strictly higher than welfare under unchecked discrimination. References C. Aina. Tailored Stories. Working paper., 2024. URL https://chiaraaina.github. io/files/TailoredStories_Aina.pdf. I. Ball and X. Gao. Benefitting from Bias: Delegating to encourage information acquisition. Journal of Economic Theory, 217:105816, 2024. doi: 10.1016/j.jet.2024.105816. A. Bardhi. Attributes: Selective Learning and Influence. Econometrica, 92(2):311–353, 2024. doi: 10.3982/ECTA18355. A. Bardhi and N. Bobkova. Local Evidence and Diversity in Minipublics. Journal of Political Economy, 131(9):2451–2508, September 2023. doi: 10.1086/724322. S. Callander. Searching and Learning by Trial and Error. The American Economic Review, 101(6):2277–2308, 2011. doi: 10.1257/aer.101.6.2277. C. Carnehl and J. Schneider. A quest for Knowledge. arXiv working paper, December 2024. doi: 10.48550/arXiv.2102.13434. J. D. Carrillo and T. Mariotti. Strategic ignorance as a self-disciplining device. The Review of Economic Studies, 67(3):529–544, 2000. 38
Y.-K. Che and N. Kartik. Opinions as Incentives. Journal of Political Economy, 117 (5):815–860, October 2009. doi: 10.1086/648432. L. D ' Amico and G. Tabellini. Disengaging from Reality: Online Behavior and Unpleasant Political News. SSRN Electronic Journal, 2022. doi: 10.2139/ssrn.4305223. I. Deimen and D. Szalay. Delegated Expertise, Authority, and Communication. American Economic Review, 109(4):1349–1374, April 2019. doi: 10.1257/aer.20161109. J. S. Demski and D. E. Sappington. Delegated expertise. Journal of Accounting Research, 25(1):68–89, 1987. doi: 10.2307/2491259. F. Echenique and A. Li. Rationally Inattentive Statistical Discrimination: Arrow Meets Phelps. arXiv working paper, November 2023. doi: 10.48550/arXiv.2212.08219. Facebook. Working to stop misinformation and false news. https://www.facebook. com/formedia/blog/working-to-stop-misinformation-and-false-news, 2017. Accessed: 2024-07-30. M. Fosgerau, R. Sethi, and J. Weibull. Equilibrium Screening and Categorical Inequality. American Economic Journal: Microeconomics, 15(3):201–242, August 2023. doi: 10. 1257/mic.20210391. R. G. Fryer, Jr, P. Harms, and M. O. Jackson. Updating Beliefs when Evidence is Open to Interpretation: Implications for Bias and Polarization. Journal of the European Economic Association, 17(5):1470–1501, October 2019. doi: 10.1093/jeea/jvy025. F. Germano, V. G´omez, and F. Sobbrio. Crowding Out the Truth? A simple Model of Misinformation, Polarization and Meaningful Social Interactions. arXiv working paper, 2022. doi: 10.48550/arXiv.2210.02248. P. Ilinov, A. Matveenko, M. Senkov, and E. Starkov. Optimally Biased Expertise. arXiv working paper, 2022. doi: 10.48550/arXiv.2209.13689. M. Kirneva. Informing to divert attention. Working paper, 2023. A. Liang, X. Mu, and V. Syrgkanis. Dynamically Aggregating Diverse Information. Econometrica, 90(1):47–80, 2022. doi: 10.3982/ECTA18324. A. Liang, J. Lu, and X. Mu. Algorithm design: A fairness-accuracy frontier. accepted at Journal of Political Economy, 2024. 39
We first characterize ˜ θk(s), the voter’s posterior belief about θkgiven some signal realizations sA k, sB k. It is immediate to verify directly from normal p.d.f.s that if ˜ X∼ N(µ, σ2 X), ˜ε∼ N(0, σ2 ε) and Y=X+ε, then ˜ X|Y∼ N ˆµ, ˆσ2, where ˆµ= 1 σ2 X 1 σ2 X +1 σ2 ε µ+ 1 σ2 ε 1 σ2 X +1 σ2 ε Y, ˆσ2=1 1 σ2 X +1 σ2 ε . This directly implies characterization (4)–(6) in Section 2.2. In the context of this proof, this gives first that ˜ θk|sA k∼ N(ˆµk,A,ˆ Σk,A) with ˆµk,A = 1 Σ0 k 1 Σ0 k +tAqA k µ0 k+tAqA k 1 Σ0 k +tAqA k sA k,ˆ Σk,A =1 1 Σ0 k +tAqA k , and then, subsequently, that (˜ θk|sA k)|sB k∼ N(ˆµk,ˆ Σk), where ˆµk= 1 ˆ Σk,A 1 ˆ Σk,A +tBqB k ˆµk,A +tBqB k 1 ˆ Σk,A +tBqB k sB k= 1 Σ0 k µ0 k+tAqA ksA k+tBqB ksB k 1 Σ0 k +tAqA k+tBqB k = 1 Σ0 k 1 Σ0 k +τk µ0 k+1 1 Σ0 k +τktAqA ksA k+tBqB ksB k,(B.1.2) ˆ Σk=1 1 ˆ Σk,A +tBqB k =1 1 Σ0 k +tAqA k+tBqB k =1 1 Σ0 k +τk .(B.1.3) Comparing (B.1.3) with its counterpart (6) in the baseline model, we note that they coincide. Next, compare (B.1.2) with its counterpart (5) in the baseline model. In order to conclude that the distribution of posteriors is the same as in the baseline model, we need to show that the random variable defined as ˜sf k:= 1 τktAqA k˜sA k+tBqB k˜sB k=˜ θk+1 τktAqA k˜εA k+tBqB k˜εB k has the same distribution as ˜skfrom the baseline model. Note that ˜sf k|θk∼ N θk,1 τk since E[˜εm k] = 0 and var tmqm k τk˜εm k=1 τk, and then ˜sf k∼ N µ0 k,Σ0 k+1 τk. For a given τ, these coincide with the distributions of ˜sk|θkand ˜skfrom the baseline model, respectively. Therefore, the distribution of ˆµkabove coincides with that of (4), so the distribution of ˜ θk(sf k) in the media model coincides with the distribution of ˜ θk(sk) in the baseline model when sf k=sk. Together, the two latter facts imply that the distribution of ˜ θk(˜sf k) in the media model coincides with the distribution of ˜ θk(˜sk) in the baseline 6
model. Our next lemma below show that all players’ respective ex ante expected utilities only depend on (t, qA, qB) through τ=τ(t, qA, qB). Specifically, define player i’s value function as Vi(τ) := E"− 2 X k=1 αi kd∗(˜sf, τ)−˜ θk(˜sf, τ)2#(B.1.4) for i=A, B, V , where d∗(˜sf, τ) is the voter’s equilibrium decision strategy given τ= τ(t, qA, qB), and sf= (sf 1, sf 2) is a vector of fictitious “aggregate” signals introduced in the proof of Lemma 13. Lemma 14. In the media model, given (t, qA, qB), the expectations of all players’ utilities (B.1.1) are given by values (B.1.4) and only depend on τt, qA, qB: Ehuid∗˜s, t, qA, qB,˜ θ˜s, t, qA, qBi=Viτt, qA, qB Proof. Rewriting the expected utility (the left-hand side of the equality in the lemma) using definition (B.1.1), we get E"− 2 X k=1 αi kd∗˜s, t, qA, qB−˜ θk˜s, t, qA, qB2#.(B.1.5) The voter’s equilibrium decision for any signal realization sis given by d∗s, t, qA, qB= 2 X k=1 αV kEh˜ θks, t, qA, qBi. By Lemma 13 we know that ˜ θk˜s, t, qA, qB∼˜ θk˜sf, τ(the two have the same distributions), where ˜sf= (˜sf 1,˜sf 2) with ˜sf k=1 τktAqA k˜sA k+tBqB k˜sB k. This implies that d∗˜s, t, qA, qB∼d∗(˜sf, τ). Then we can write (B.1.5) as E"− 2 X k=1 αi kd∗(˜sf, τ)−˜ θk˜sf, τ2#, which is exactly (B.1.4), the definition of Vi(τ). We now make a statement about the media model equilibria in cases not covered in 7
Proposition B.1.1, and prove both statements simultaneously. Proposition B.1.2. In each of the cases listed below, all equilibria are payoff-equivalent to the following: 1. if αA 1> αV 1> αB 1and Tis not large enough: both media outlets choose the same extreme coverage, either qA=qB= (1,0), or qA=qB= (0,1), and the voter achieves her best information; 2. if αV 1≥αA 1: the voter only follows outlet A,t= (T, 0), and outlet Aachieves its best information;32 3. if αB 1≥αV 1: the voter only follows outlet B,t= (0, T), and outlet Bachieves its best information. Proof of Propositions B.1.1 and B.1.2. Lemmas 13 and 14 imply that the media setting with four signals s= (sA 1, sA 2, sB 1, sB 2) can, without loss, be replaced by the baseline model with fictitious aggregate signals sf= (sf 1, sf 2), which have the same distribution conditional on aggregate attention allocation τt, qA, qBas the signals in our main model of Section 2. For brevity, we often drop the arguments of the aggregate attention allocation and write it as simply τ. Proceeding by backwards induction, we first note that the voter’s equilibrium decision d∗sf, τis analogous to the single-player and strategic versions of our model, see the proof of Lemma 14. In the remainder of this proof, we assume the voter follows this equilibrium decision strategy d∗sf, τ. Proceeding further, Lemma 14 implies that the outlets’ problem of choosing coverage qmand the voter’s problem of choosing attention allocation treduce to the problem of maximizing value function (B.1.4) over τt, qA, qB. Fix some tand consider the outlets’ problem. Let TV(t) := t1+t2denote the total amount of attention devoted to the two outlets by the voter. Let τi∗(t) := arg maxτ∈R2 +Vi(τ) s.t. τ1+τ2≤TV(t) for all i=A, B, V . Note that player i’s best information, by definition, is given by maxt∈R τi∗(t). For both m=A, B,τm∗(t) is then described by Lemma 2 and Theorem 2 with T=TV(t). By Lemma 2, if m’s expected payoff (12) is weakly increasing in 32We define outlet m’s best information by analogy with the voter’s best information in the text, i.e., (t, qA, qB) must jointly maximize m’s expected payoff given the voter’s equilibrium decision strategy. 8
τk, it is also strictly concave in τk.33 From Theorems 1 and 2, we obtain a closed-form solution for τi∗(t): τi 1∗(t) = max 0,min ˆαi 1Σ0 1−ˆαi 2Σ0 2 Σ0 1Σ0 2(ˆαi 1+ ˆαi 2)+ˆαi 1 ˆαi 1+ ˆαi 2 TV(t), TV(t), τi 2∗(t) = TV(t)−τi 1∗(t), (B.1.6) where ˆαi k:= qmax 0,(αi k)2−(αi k−αV k)2for all k= 1,2 and i=A, B, V (note that ˆαV k=αV k, so for i=V, (B.1.6) coincides with the solution in Theorem 1). Note that τi 1∗(t) is weakly increasing in αi 1(subject to the αi 1+αi 2= 1 constraint). This characterization implies that since αi 1+αi 2= 1 for all i=A, B, V , the auxiliary weights ˆαm kcannot simultaneously be zero for both k= 1,2 for any m=A, B, so τm 1∗(t) + τm 2∗(t) = TV(t) for all t(the media want no attention wasted). We note that subject to this constraint, value Vi(τ1, TV(t)−τ1) is either strictly monotone, or strictly concave in τ1for all i=A, B, V . The characterization in Theorem 2 also implies that given t, if ˆαm 1Σ0 1>ˆαm 2Σ0 2and TV(t)≤ˆαm 1 ˆαm 2Σ0 2−1 Σ0 1 (or for all TV(t) if ˆα2= 0), then τm∗(t) = (TV(t),0), and vice versa: if ˆαm 2Σ0 2>ˆαm 1Σ0 1and TV(t)≤ˆαm 2 ˆαm 1Σ0 1−1 Σ0 2 (or for all TV(t) if ˆα1= 0), then τm∗(t) = (0, TV(t)). If none of these conditions apply, then τm∗(t) is interior (and strictly monotone in αm 1subject to the αm 1+αm 2= 1 constraint). Therefore, if we let ˆ T:= max 0,min m∈{A,B}ˆαm 1 ˆαm 2Σ0 2−1 Σ0 1,min m∈{A,B}ˆαm 2 ˆαm 1Σ0 1−1 Σ0 2, then whenever TV(t)≤ˆ T, both media outlets optimally prefer the same extreme coverage: τA 1∗(t) = τB 1∗(t)∈ {0, TV(t)}. Note that if ˆαA 1Σ0 1>ˆαA 2Σ0 2and ˆαB 1Σ0 1<ˆαB 2Σ0 2, then ˆ T= 0, and such a preference profile (τA 1∗(t) = τB 1∗(t)∈ {0, TV(t)}) cannot arise. Further, if TV(t)>ˆ T, then τm∗(t) is interior for at least one outlet m, and since αA 1> αB 1 (and αA 2< αB 2), we then have τA 1∗(t)> τB 1∗(t). Next, we claim that if τ=τm∗(t) for some outlet min equilibrium, then mis either polarized (i.e., qm∈ {((1,0),(0,1)}), or ignored (tm= 0). To see this, proceed by contradiction and consider a strategy profile t, qA, qBsuch that τkt, qA, qB< τm k∗(t), qm k<1, and tm>0 for some k= 1,2 and m=A, B. Then increasing qm k(while 33This is most evident from (30) in the proof of Theorem 2: one can readily see by analogy that both monotonicity and concavity in this problem depend on whether λi k:= αV k2αm k−αV k≷0. 9
decreasing the other weight qm −k) brings the aggregate test allocation τcloser to m’s optimum τm∗(t), which strictly increases m’s expected payoff (12).34 This is a profitable deviation for outlet m, and the original strategy profile thus cannot be an equilibrium. The same is true of the case when τk> τm k∗(t) and qm k>0 and tm>0 for some k= 1,2, m=A, B. The logic above also implies that as αA 1> αB 1, there is no equilibrium with τsuch that τ1> τA 1∗(t) or τB 1∗(t)> τ1, since in the former case both outlets mwant to decrease qm 1and increase qm 2(and at least one has the scope for such a deviation), and in the latter they want to do the opposite. Therefore, in equilibrium, the aggregate attention allocation τ=τt∗, qA∗, qB∗must be such that τ1∈τB 1∗(t∗), τA 1∗(t∗)and τ2=TV(t∗)−τ1∈τA 2∗(t∗), τB 2∗(t∗). It remains to consider the voter’s attention allocation problem. Since τm∗(t) only depend on tthrough TV(t), the voter’s problem can be split into first choosing the total amount of attention to devote to media, TV, and then choosing how to split it by way of choosing τsuch that τ1∈τB 1∗(TV), τA 1∗(TV)and τ2=TV−τ1. Note that in the second step, all aggregate allocations in this interval are available to the voter for a given TV. By setting tm=TVfor some mand t−m= 0, she can induce τ=τm∗(TV). Any allocation τin the interior of the interval can be achieved by setting tA=τ1and tB=τ2, in which case the unique best response for outlet Ais qA= (1,0), and the unique best response for outlet Bis qB= (0,1), as argued above. We shall thus refer to τthat satisfy these requirements as feasible given some TV. The voter never wants any attention wasted, since her payoff is strictly increasing in τk for both k. From (B.1.6) we can see that d dTVτm k∗(TV)∈[0,1] for all m, k, hence for any TV, any τathat is feasible given TV, and any ε > 0, there exists τbthat is feasible given TV+εand is such that τb k≥τa kfor both k, with at least one inequality being strict.35 Therefore, increasing TValways offers a strict improvement the voter, so in equilibrium, TV(t∗) = T. We now move on to the second stage of the voter’s problem. If T≤ˆ T, then τA 1∗(T) = 34 “Closer” here is used in the sense of the new aggregate attention allocation ¯τbeing a convex combination of the original τand the optimal τm∗(t). The payoff function being concave on the relevant interval, and τm∗(t) being its maximizer then directly imply that ¯τyields a higher payoff than the original τ. 35A way to see this is to notice that an increase in TVimplies an increase in the strong set order of the intervals of feasible in τ1and τ2. 10
τB 1∗(T)∈ {0,1}, so the set of available τis a singleton, and the voter is then indifferent between all attention allocations. Suppose now that T > ˆ T. If αA 1> αV 1> αB 1, then τV 1∗(T)∈τB 1∗(T), τA 1∗(T), so the voter can attain her optimum τV∗(T) by following the strategy described above. It is immediate from (B.1.6) that in this case there exists ˆ TV≥ˆ Tsuch that τV∗(T) is interior if and only if T > ˆ TV. For T > ˆ TV,τV∗(T) is uniquely optimal for the voter, hence t= (τV 1∗(T), τV 2∗(T)), qA= (1,0), and qB= (0,1) is the unique equilibrium, proving Proposition B.1.1. For T≤ˆ TV, as argued above, the voter can still achieve τV∗(T), but the equilibrium may not be unique. This proves case 1 of Proposition B.1.2. In case αV 1/∈αA 1, αB 1, since the voter’s value function (11) is strictly concave in τ, she will optimally choose the available allocation closest (in the sense of Footnote 34) to τV∗(T). In particular, if αB 1≥αV 1, then τV 1∗(T)≤τB 1∗(T), so the voter will choose t= (0, T), resulting in τ=τ∗ B(T). And if αV 1≥αA 1, then τV 1∗(T)≥τA 1∗(T), and the voter will choose t= (T, 0), resulting in τ=τ∗ A(T). This proves parts 2 and 3 of Proposition B.1.2 and completes this proof. Finally, we state the following general corollary before proving it along with Corollary B.1.1. Corollary B.1.2. The voter weakly prefers the equilibrium when both media outlets m=A, B are available to the equilibrium when only one outlet m=A, B is present in the market. Proof of Corollaries B.1.1 and B.1.2. In the monopoly case, if only outlet mis available, then the voter always lends it her full attention, tm=T, since her value (B.1.4) reduces to (11), which is strictly increasing and strictly concave in τ. Therefore, mcan choose aggregate attention allocation τfreely subject to τ1+τ2=T. Since αi 1+αi 2= 1 for all i, Theorem 2 implies that mdoes not want to waste any of the voter’s attention. The unique equilibrium aggregate attention allocation is, therefore, τm∗(T), as defined in the proof of Propositions B.1.1 and B.1.2, which is m’s best information. Propositions B.1.1 and B.1.2 above shows that under media duopoly, the equilibrium aggregate attention allocation τ∗is given by τV∗(T) if αA 1> αV 1> αB 1, by τA∗(T) if αA 1≤αV 1, and by τB∗(T) if αV 1≤αB 1. Note that in the two latter cases, it is the voter’s preferred media outlet that has the monopoly power, which proves Corollary B.1.2 for those cases. In the former case, τV∗(T) being its maximizer of the voter’s value 11
function imply that the voter prefers duopoly, concluding the proof of Corollary B.1.2. If τV∗(T)∈τB∗(T), τA∗(T), which is the case if and only if T > ˆ TV, as defined in the proof of Propositions B.1.1 and B.1.2, then it is also the unique maximizer (since the value function is strictly concave on that interval in that case), which proves Corollary B.1.1. B.2 Changing Preferences In this second extension, we study how a (potential) change in an agent’s preferences between learning and decision-making affects his learning strategy. Following the literature on time-inconsistent preferences (e.g., O’Donoghue and Rabin, 1999), we consider sophisticated and naive agents and investigate the impact of such a change in preferences on the learning strategy and on the agent’s welfare. We model this similarly to our baseline model by considering a single agent facing a two-stage problem. In the first stage, the agent acquires information about unknown attributes. In the second stage, the agent takes a decision based on the acquired information. The agent’s preference may change between the two stages. For instance, this change may occur due to a change in the agent’s circumstances, such as losing the job or falling ill. Alternatively, the change in preferences may be due to the agent’s self-control problems, such as succumbing to temptation. In the first stage, when acquiring information, the agent has utility function uR(d, θ) = −d−αR 1θ1−αR 2θ22,(B.2.1) where we assume αR k>0 for k= 1,2. Afterwards, the agent’s preference may change, which we model as a change in the weight of attribute ˜ θ2occurring with (ex ante) probability p∈(0,1). In the second stage, when making the decision, the agent has utility function uD(d, θ) = −d−αR 1θ1−αR 2θ22with probability 1 −p, −d−αR 1θ1−cαR 2θ22with probability p, 12
where c > 0 and c= 1.36 The naive agent believes her preferences will stay the same across both stages, while the sophisticated agent understands they might change. Throughout this section, we simplify the analysis by assuming µ0= (0,0), so that, absent any new information, the researcher and both types of decision-makers agree on the optimal decision d∗= 0. Additionally, we fix Σ0 1= Σ0 2=T= 1. As the naive agent incorrectly assumes no possibility of change, the naif’s learning problem coincides with that of a single player in Section 3, and the equilibrium test allocation follows from Theorem 1 with weights αR. In contrast, for the sophisticate, a strategic game similar to that in Section 4 ensues. However, Theorem 2 does not apply directly, as the decision-maker’s preferences are stochastic. By appropriately adapting Lemma 2 and then applying the same logic as in the proof of Theorem 2, we find that the sophisticate’s equilibrium test allocation can nevertheless be expressed as the solution to a single-player problem with auxiliary weights37 (ˆα1,ˆα2) := αR 1, αR 2pmax{1−p+p(c(2 −c)),0}. Thus, the potential change in preferences reduces the weight put on attribute ˜ θ2, as the sophisticate anticipates the misalignment between his current and future self. Notably, the next result shows that this reduction may even induce the sophisticate to engage in “strategic ignorance” (see, e.g., Carrillo and Mariotti, 2000). Proposition B.2.1. The sophisticate avoids learning about attribute ˜ θ2for any T > 0 if and only if √p(c−1) ≥1. In the presence of a potential shift in preferences, represented by c, the sophisticate may choose strategic ignorance regarding attribute ˜ θ2. If the change reduces the weight on ˜ θ2 or does not increase it too much, the agent still learns about it, but with less intensity (ˆα2< αR 2). In contrast, the sophisticate avoids learning about attribute ˜ θ2if the change is sufficiently high (c > 2) and the probability of change is not too small. In short, when the sophisticate expects the future self to overreact to information about ˜ θ2, she may opt for strategic ignorance. 36This setting extends the baseline model by allowing for uncertainty about the decision-maker’s preferences. An alternative interpretation is that there are multiple potential decision-makers, and the researcher is uncertain about who will make the decision. We extend our equilibrium concept to this setting in Section B.2.1. 37The formal steps are contained in the proof of Proposition B.2.1. 13
Having established the differences between the naif and the sophisticate’s equilibrium learning strategies, we want to understand the implications for the agent’s welfare. In this context, the choice of the appropriate welfare criterion is unclear. For instance, if the change in preferences is triggered by succumbing to temptation, the initial utility function uRseems most appropriate. However, if the change in preferences is due to altered circumstances, such as illness or parenthood, then the changed utility function uDseems more appropriate. We remain agnostic about the welfare criterion and consider both alternatives. Proposition B.2.2. If αR 1/∈(ˆα2/2,2αR 2), the expected payoff of the naif and the sophisticate coincide for both welfare criteria. Otherwise, if the welfare criterion is the initial utility, uR, the welfare of the sophisticate exceeds that of the naif. If the welfare criterion is the changed utility, uD, the naif’s welfare exceeds that of the sophisticate for c > 1 and vice versa for c < 1. The interesting case arises when a potential change in preferences influences equilibrium learning strategies. As shown in the appendix, the sophisticate allocates (at least weakly) more tests to attribute ˜ θ1than the naif. Anticipating a shift in preferences, the sophisticate reduces emphasis on attribute ˜ θ2to avoid misalignment with the future self’s choices. When initial utility is used as the welfare criterion, the sophisticate’s expected payoff is higher because they optimize learning, aware of possible preference changes, while the naif makes suboptimal choices. However, when the welfare criterion is the changed utility, the naif can outperform the sophisticate. If c > 1, the naif learns more about ˜ θ2, benefiting whether preferences change or not: If a change occurs, the naif has learned more about the now more important attribute ˜ θ2than the sophisticate (in extreme case, the sophisticate engages in strategic ignorance and learns nothing about ˜ θ2); if no change occurs, the naif has implemented the optimal test allocation due to her naivete. In contrast, the sophisticate’s “hedging” learning strategy against the potential change, which underweighs ˜ θ2, is less effective. However, this is reversed when c < 1. Then, the naif’s learning strategy is still optimal when no change occurs but (substantially) worse when a change occurs. The result in Proposition B.2.2 has a parallel to O’Donoghue and Rabin’s (1999) findings in their “doing it now or later” framework. They find that in case of immediate costs the sophisticate is always better off than the naif (as in our case with the initial 14
utility). Conversely, in case of immediate gratification, either type of the agent can be better off (as in our case with changed utility). In both, O’Donoghue and Rabin (1999) and our model, the sophisticate may engage in a form of overcompensation. In our case it happens by putting less weight on attribute ˜ θ2to hedge against preference changes, which leads to a worse outcome irrespective of whether preferences actually change. Summarizing, both models highlight that naive strategies can sometimes lead to better outcomes if future preferences or circumstances validate those strategies. This underscores the complexity of modeling changing preferences and how different assumptions about future behavior or the reason for change can impact welfare. B.2.1 Proofs for the extension on changing preferences Equilibrium concept. A weak Perfect Bayesian Equilibrium of the game with potentially changing preferences (hereinafter referred to simply as “equilibrium”) is a tuple (τ∗, d∗ c(s, τ), d∗ u(s, τ),˜ θ(s, τ)) such that: 1. the researcher’s test allocation strategy τ∗∈ T maximizes his expected payoff given the decision-maker’s strategies d∗ c(s, τ) and d∗ u(s, τ); 2. the decision-maker’s decision strategies d∗ c(s, τ) : R2×T → Rand d∗ u(s, τ) : R2× T → Rmaximize her expected payoff given changed and unchanged preferences, respectively, and her posterior beliefs ˜ θ(s, τ); 3. the decision-maker’s posterior beliefs ˜ θ(s, τ) : R2× T → ∆(R2) are obtained via Bayes’ rule given the signal realizations sand the researcher’s choice τ. Proof of Proposition B.2.1. Observe that the decision-maker’s equilibrium decision is given by d∗ u(s, τ) = αR 1ˆµ1(s, τ) + αR 2ˆµ2(s, τ) with probability 1 −p, d∗ c(s, τ) = αR 1ˆµ1(s, τ) + cαR 2ˆµ2(s, τ) with probability p. Thus, for any test allocation τ, the uncertainty about the decision-maker’s preferences translates accordingly into uncertainty about the decision-maker’s decision. Hence, we 15
proof of Lemma 5, it holds that αR 1(β, γ)≤1/2αD 1if and only if γ≥γ1(β); and αD 2(β, γ)≤1/2αD 2if and only if γ≤γ2(β). By Theorem 2, the analyst’s equilibrium test allocation satisfies τ∗(β, γ) = (τ∗ 1(β, γ), τ∗ 2(β, γ)) = (T, 0) if γ≤γ2(β), (0, T) if γ≥γ1(β). (C.2.5) Now suppose γ∈[γ2(β), γ1(β)]. Let ˜τk(β, γ) and ˆαk(β, γ) for k= 1,2 be given by equations (33) and (34). Note that expressions in (34) are well-defined for γ∈[γ2(β), γ1(β)]. Furthermore, we have ˜τ1(β, γ) = 1 Σ0 2 +T > T at γ=γ2(β), −1 Σ0 1 <0 at γ=γ1(β), since ˆα2(β, γ2(β)) = 0, ˆα1(β, γ2(β)) >0, and ˆα1(β, γ1(β)) = 0, ˆα2(β, γ1(β)) >0. Further, ˜τ1(β, γ) is continuous and strictly decreasing in γ. To see the strict monotonicity, let us take the partial derivative of ˜τ1(γ) with respect to γ. For ease of notation, we suppress the dependence on βand γfrom the right-hand side of the expression. We get ∂˜τ1(β, γ) ∂γ = (Σ0 1+ Σ0 2)ˆα2∂ˆα1 ∂γ −ˆα1∂ˆα2 ∂γ Σ0 1Σ0 2(ˆα1+ ˆα2)2+ˆα2∂ˆα1 ∂γ −ˆα1∂ˆα2 ∂γ (ˆα1+ ˆα2)2<0, where we used ∂ˆα1(β,γ) ∂γ =−αD 2 √ˆα1(β,γ)<0 and ∂ˆα2(γ) ∂γ =αD 1 √ˆα2(β,γ)>0 to determine the sign. By the intermediate value theorem, there exist γI(β), γII (β)∈(γ2(β), γ1(β)) such that ˜τ1(β, γ) = Tat γ=γI(β), 0 at γ=γII (β). (C.2.6) From the strict monotonicity of ˜τ1(β, γ), it further holds that ˜τ1(β, γ) > T if and only if γ < γI(β), <0 if and only if γ > γII (β). (C.2.7) Finally, note that ˆαk(β, γ) as given by (34) are equivalent to (13) and (14) given de22
composition (18). By Theorems 1 and 2, the analyst’s equilibrium test allocation for γ∈[γ2(β), γ1(β)] is τ∗ 1(β, γ) = max {0,min {˜τ1(β, γ), T}} (C.2.8) τ∗ 2(β, γ) = max {0,min {˜τ2(β, γ), T}} (C.2.9) Combining the results (C.2.8), (C.2.9) with (C.2.7) and (C.2.5), together with the strict monotonicity result of ˜τ1(β, γ), we obtain the claim of Lemma 6. C.2.3 Proof of Lemma 7 Given a test allocation τ= (τ1, τ2), the decision-maker’s interim expected payoff is VD(τ) = −αD 12Σ0 1 1 + τ1Σ0 1−αD 22Σ0 2 1 + τ2Σ0 2 .(C.2.10) The first, the second, and the cross derivates of the decision-maker’s interim expected payoff with respect to τk≥0 and τl=τkare ∂V D(τ) ∂τk=αD kΣ0 k 1+τkΣ0 k2>0, ∂2VD(τ) ∂τ2 k = −2αD k2Σ0 k 1+τkΣ0 k3<0, and ∂2VD(τ) ∂τk∂τl= 0. Hence, VD(τ) is strictly increasing and strictly concave function. C.2.4 Proof of Lemma 8 Fix the test budget T > 0 and the decision-maker’s weight vector αD∈R2 ++. Suppose γ= 0 (the agent is not distorted). (i) If β≤1/2 (the agent is too insensitive), then the agent optimally chooses no testing: τ∗(β, 0) = (0,0). (ii) If β > 1/2 (the agent is sufficiently sensitive), then the agent optimally chooses the decision-maker’s most preferred test allocation: τ∗(β, 0) = τ∗(1,0). Suppose the premise of Lemma 8 holds. Part (i) immediately follows from Lemma 5. For part (ii) take β > 1/2 and let ˜τ1(β, γ) and ˜τ2(β, γ) be given by (33). At γ= 0, we 23
get ˜τ1(β, 0) = √2β−1αD 1Σ0 1−√2β−1αD 2Σ0 2 Σ0 1Σ0 2√2β−1αD 1+αD 2+√2β−1αD 1 √2β−1αD 1+αD 2T =αD 1Σ0 1−αD 2Σ0 2 Σ0 1Σ0 2(αD 1+αD 2)+αD 1 αD 1+αD 2 T Lemma 6 and Theorem 1 then imply that the agent’s equilibrium test allocation coincides with the optimal test allocation of a single player with weight vector α=αD. C.3 Details of the Proofs in Appendix A.4 C.3.1 Proof of Lemma 9 Let us set T= 1 and Σ0 1= Σ0 2= 1. From Theorems 1 and 2, the equilibrium test allocation satisfies τ∗ 1(p, δ)=1−τ∗ 2(p, δ) and τ∗ 2(p, δ) = min ˆα2(p, δ)−ˆα1(p, δ) ˆα1(p, δ) + ˆα2(p, δ)+ˆα2(p, δ) ˆα1(p, δ) + ˆα2(p, δ),1= min2ˆα2(p, δ)−ˆα1(p, δ) ˆα1(p, δ) + ˆα2(p, δ) | {z } τaux 2(p,δ):= ,1 where ˆα1(p, δ) and ˆα2(p, δ) are given by equations (36) and (37), respectively. The function τaux 2(p, δ) is well-defined, since ˆα1(p, δ) + ˆα2(p, δ)= 0 for all values of pand δ. Note that τaux 2(p, δ)=1/2 when p= 0 and τaux 2(p, δ) = 2 when p≥1−δ 2. Furthermore, when p∈0,1−δ 2we have ∂τaux 2(p, δ) ∂p =2∂ˆα2 ∂p −∂ˆα1 ∂p (ˆα1+ ˆα2)−(2ˆα2−ˆα1)∂ˆα1 ∂p +∂ˆα2 ∂p (ˆα1+ ˆα2)2 =3 4 (1 −δ2) ˆα1ˆα2(ˆα1+ ˆα2)2>0.(C.3.1) Note that τaux 2(p, δ) is continuous in pfor p≥0.39 Furthermore, τaux 2(p, δ) is strictly increasing in pon p∈[0,(1 −δ)/2). At p= 0 we have τaux 2(p, δ)=1/2<1 and at p≥(1 −δ)/2 we have τaux 2(p, δ) = 2 >1. Hence, for each value of discrimination δ∈ (0,1) there exists a threshold partiality level of the advisor paux(δ)∈0,1−δ 2, defined implicitly by the equation τaux 2(paux(δ), δ) = 1, such that the equilibrium testing strategy 39Since ˆα1(p, δ) and ˆα2(p, δ) are both continuous in p, the only possible point of discontinuity of τaux 2(p, δ) is if ˆα1(p, δ) + ˆα2(p, δ) = 0. Equations (36) and (37) imply this never happens for all feasible values of pand δ. Hence, τaux 2(p, δ) is continuous on p≥0. 24
satisfies τ∗ 2(p, δ) = τaux 2(p, δ) for p∈[0, paux(δ)] and τ∗ 2(p, δ) = 1 for p≥paux(δ). C.3.2 Proof of Lemma 10 The ex ante expected utility of group k,Vk(τ, δ), can be derived from the expected payoff of a researcher VR(τ) in the baseline model in the proof of Lemma 2, where we set the decision-maker’s weights to αD 1(δ), αD 2(δ) and the researcher’s weights (now group kweights) to αR k= 1 and αR −k= 0. In other words, given allocation τand the politician’s equilibrium strategy d∗(s, τ;δ), the ex ante expected payoff of group kis the same as that of the researcher in the baseline model who solely cares about attribute ˜ θk (with weight αR k= 1) and not about the other attribute (with weight αR −k= 0). We get Vk(τ, δ) = Eh−(˜ d∗D(τ;δ)−˜ θk)2i =−vD 0(δ)−µ0 k2−σ2,D 0(δ)+Σ0 k+ ˆσ2,D(τ;δ) + 2 cov ˜ d∗D(τ;δ),˜ θk =−αD 1(δ)µ0 1+αD 2(δ)µ0 2−µ0 k2−αD 1(δ)2Σ0 1+αD 2(δ)2Σ0 2+ Σ0 k +Σ0 k 1 + τkΣ0 kαD k(δ)2+ 2αD k(δ)Σ0 kτk+Σ0 −k 1 + τ−kΣ0 −kαD −k(δ)2 Using µ0 1=µ0 2= 0, Σ0 1= Σ0 2= 1, we get Vk(τ, δ) = −αD 1(δ)2+αD 2(δ)2+ 1 +1 1 + τkαD k(δ)2+ 2αD k(δ)τk+1 1 + τ−kαD −k(δ)2.(C.3.2) The difference between the expected utilities of the two groups is V1(τ, δ)−V2(τ, δ) = 2αP 1(δ)τ1 1 + τ1−2αP 2(δ)τ2 1 + τ2 =(1 + δ)τ1 1 + τ1−(1 −δ)τ2 1 + τ2 where we used αP 1(δ) = 1 2(1 + δ) and αP 2(δ) = 1 2(1 −δ). The difference between the expected utilities of the two groups under the additional constraint that τ1+τ2= 1 is ∆(τ2, δ) = V1((1 −τ2, τ2), δ)−V2((1 −τ2, τ2), δ) = (1 + δ)(1 −τ2) 1+1−τ2−(1 −δ)τ2 1 + τ2 . The function ∆(τ2, δ) is continuous and strictly decreasing in τ2since δ∈(0,1) and thus ∂∆(τ2,δ) ∂τ2=−(1+δ) (1+1−τ2)2−(1−δ) (1+τ2)2<0. Furthermore, we have ∆(1/2, δ)>0 and 25
∆(1, δ)<0. Hence, for each δthere exists a unique value of ˆτ2(δ)∈(1/2,1) for which it holds ∆(ˆτ2(δ); δ) = 0. C.3.3 Proof of Lemma 11 Let ω(τ2, δ) := V1((1 −τ2, τ2), δ) + V2((1 −τ2, τ2), δ) =−3−δ2+1/2(1 + δ2) + (1 + δ)(1 −τ2) 2−τ2 +1/2(1 + δ2) + (1 −δ)τ2 1 + τ2 denote the sum of the expected utilities of the two groups as a function of test allocation τ= (τ1, τ2) under the constraint that the budget is fully used, τ1=T−τ2 with T= 1, and where Vk(τ, δ) is given by equation (C.3.2). Since the budget is exhausted in equilibrium (as shown in the proof of Lemma 4), welfare is thus given by W(p, δ) = ω(τ∗ 2(p, δ), δ). The partial derivative of welfare with respect to p > 0, whenever differentiable, is then ∂W(p, δ) ∂p =∂τ∗ 2 ∂p ∂ω(τ∗ 2, δ) ∂τ2 =∂τ∗ 2 ∂p 1/2(1 + δ2)−(1 + δ) (2 −τ∗ 2)2 | {z } <0,since δ<1 +−1/2(1 + δ2) + (1 −δ) (1 + τ∗ 2)2 where we suppressed the dependence of τ∗ 2(p, δ) on (p, δ) from the RHS for brevity. Note that 1/2(1 + δ2)−(1 + δ)>−1/2(1 + δ2) + (1 −δ)and that (2 −τ∗ 2)2<(1 + τ∗ 2)2, since τ∗ 2(p, δ)>1/2 for p > 0. Hence, ∂ω(τ∗ 2,δ) ∂τ2<0. We thus have sign ∂W(p,δ) ∂p = −sign ∂τ∗ 2 ∂p . Let τaux 2(p, δ), ˆp(δ) and paux(δ) with ˆp(δ)< paux(δ), be the variables defined in the proof of Lemma 4. From equation (C.3.1), it follows sign ∂W (p, δ) ∂p =−sign ∂τ∗ 2(p, δ) ∂p = −sign τaux 2(p,δ) ∂p <0p∈(0, paux(δ)) 0p>paux(δ) , and the derivative does not exist at p=paux(δ). We thus have that welfare W(p, δ) is (i) continuous in p(since Vk(τ∗(p, δ), δ) is continuous in p); (ii) strictly decreasing on p∈(0, paux(δ)); and (iii) constant on p≥paux(δ), where all the advisors use the entire budget to learn exclusively about group 2. Furthermore, the equalizing partiality level 26
ˆp(δ)< paux(δ). Next, inequality is given by I(p, δ) = |∆(τ∗ 2(p, δ), δ)| where ∆(τ∗ 2(p, δ), δ) is given by equation (38). As we have shown in the proof of Lemma 4, the function ∆(τ∗ 2(p, δ), δ) is continuous in τ∗ 2(p, δ), strictly decreasing in τ∗ 2(p, δ), and is zero at τ∗ 2(ˆp(δ), δ), where τ∗ 2(ˆp(δ), δ)< T (the advisor who restores equality, ˆp(δ), does not use the entire test budget to learn exclusively about group 2). As shown above, we have: the equilibrium allocation of tests to group 2, τ∗ 2(p, δ), is continuous in p, strictly increasing on p<paux(δ), and constant on p≥paux(δ), where the advisors use the entire test budget to learn exclusively about group 2: τ∗ 2(p, δ) = T. Hence, inequality I(p, δ) is (i) continuous in p; (ii) strictly decreasing on p∈(0,ˆp(δ)) (iii) zero at p= ˆp(δ); (iv) strictly increasing on p∈(ˆp(δ), paux(δ)); and (v) constant on p≥paux(δ). Therefore, welfare W(p, δ) and inequality I(p, δ) are both (i) continuous in p; and (ii) strictly decreasing on p∈(0,ˆp(δ)). Furthermore, welfare strictly decreases and inequality strictly increases on p∈(ˆp(δ), paux(δ)); and they are both constant on p≥paux(δ). C.3.4 Proof of Lemma 12 By Theorem 1, the politician’s optimal learning strategy under unchecked discrimination yields ¯τ2(δ) = 1−3δ 2if δ < 1 3, 0 if δ≥1 3. From the proof of Lemma 4, equation (39), the test allocation restoring equality ˆτ2(δ) = −(1−δ)+√1+3δ2 2δ. Fix δ∈[1/3,1). Then |τ2(δ)−1/2|>|ˆτ2(δ)−1/2|, since ˆτ2(δ)∈(1/2,1), as shown in 27
the proof of Lemma 4. Next, fix δ∈(0,1/3). Then we have, equivalently, |τ2(δ)−1/2|=1 2−1−3δ 2>−1 + δ+√1+3δ2 2δ−1 2=|ˆτ2(δ)−1/2| 1+3δ2>p1+3δ2 which holds for any δ∈(0,1/3). 28