Contracting over persistent information
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Zhao, Wei; Mezzetti, Claudio; Renou, Ludovic; Tomala, Tristan Article Contracting over persistent information Theoretical Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Zhao, Wei; Mezzetti, Claudio; Renou, Ludovic; Tomala, Tristan (2024) : Contracting over persistent information, Theoretical Economics, ISSN 1555-7561, The Econometric Society, New Haven, CT, Vol. 19, Iss. 2, pp. 917-974, https://doi.org/10.3982/TE5056 This Version is available at: https://hdl.handle.net/10419/320256 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Theoretical Economics 19 (2024), 917–974 1555-7561/20240917 Contracting over persistent information Wei Zhao School of Economics, Renmin University of China Claudio Mezzetti School of Economics, The University of Queensland Ludovic Renou School of Economics and Finance, Queen Mary University of London, CEPR, and Department of Economics, University of Adelaide Tristan Tomala HEC Paris and GREGHEC-CNRS We consider a dynamic principal-agent problem, where the sole instrument the principal has to incentivize the agent is the disclosure of information. The principal aims at maximizing the (discounted) number of times the agent chooses the principal’s preferred action. We show that there exists an optimal policy, where the principal recommends its most preferred action and discloses information as a reward in the next period, until either this action becomes statically optimal for the agent or the agent perfectly learns the state. Keywords. Dynamic, contract, information, revelation, disclosure, sender, receiver, persuasion. JEL classification. C73, D82. 1. Introduction We consider a dynamic “principal-agent” model, where the sole instrument the principal has is information.1Principal and agent are engaged in a long-term relationship. The principal aims at inducing the agent to choose an action—the principal’s most preferred action—as often as possible, and can only do so by disclosing information about Wei Zhao: [email protected] Claudio Mezzetti: [email protected] Ludovic Renou: [email protected] Tristan Tomala: [email protected] Wei Zhao gratefully acknowledges the support of the HEC Foundation. Claudio Mezzetti thankfully acknowledges financial support from the Australian Research Council Discovery grant DP190102904. Ludovic Renou gratefully acknowledges the support of the Agence Nationale pour la Recherche under grant ANR CIGNE (ANR-15-CE38-0007-01) and through the ORA Project “Ambiguity in Dynamic Environments” (ANR-18-ORAR-0005). Tristan Tomala gratefully acknowledges the support of the HEC Foundation and ANR/Investissements d’Avenir under grant ANR-11-IDEX-0003/Labex Ecodec/ANR-11-LABX-0047. 1That is, the principal cannot make transfers, terminate the relationship, choose allocations, or constrain the agent’s choices. ©2024 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at https://econtheory.org.https://doi.org/10.3982/TE5056
918 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) an unknown state. To give examples, the principal is: (i) an external consultant with a clear agenda about what a company (the agent) should do, (ii) a department in a corporation aiming to maintain a central role while advising the CEO, (iii) a technology leading, multinational firm in a joint venture with a local firm in a less developed country; (iv) a lobbyist attempting to influence a politician. We assume that the principal commits to a disclosure policy, which we refer to as the offer of a “contract.” The dynamic contracting problem we study is, therefore, a dynamic persuasion problem. The standard approach in the study of dynamic contracting models (e.g., Spear and Srivastava (1987)) is to use the agent’s continuation value, or promised utility, as a state variable. The principal’s Bellman equation is then the fixed point of an operator, which satisfies a promise-keeping constraint in addition to incentive constraints. However, in dynamic persuasion models, there are additional complications. First, since the belief of the agent changes over time due to information disclosure, we must treat it as an additional state variable. This increases the dimensionality of the principal’s problem. Second, any information disclosure policy, to which the principal commits, generates a martingale of beliefs. We must therefore impose the constraint that the belief process is a martingale. To the best of our knowledge, we are the first to be able to provide a complete characterization of an optimal contract by solving for the fixed point of a Bellman equation with two state variables tracking the evolution of the agent’s beliefs and of his promised utility. We now illustrate the general properties of our optimal policy. First, the principal uses information disclosure as a “carrot” to motivate the agent to take the principal’s most preferred action until either the agent perfectly learns the state, or choosing the principal’s most preferred action becomes statically optimal. Moreover, if the agent learns the state, he learns it in finite time. After the agent has learned the state, he will take his optimal action in that state. Alternatively, as long as the agent keeps getting pieces of information from the principal (and thus, has not learned the state yet), he will take the principal’s preferred action. By trickling down bits of information, the principal is able to induce the agent to delay moving away from his favorite course of action. In some instances, the principal will promise eventual full disclosure of the state with probability one. In other instances, the principal will be able to stir the agent’s beliefs so that, with positive probability, the agent will take the principal’s favorite action forever. We provide a characterization of when this occurs. Define the agent’s opportunity cost at a state as the difference between the agent’s stage payoff at his optimal action and the stage payoff when taking the principal’s preferred action. Generically, the agent’s opportunity cost, relative to the principal’s benefit from his preferred action, is different in different states, and our optimal policy exploits these differences. The second property of our optimal policy is that, along the paths at which the agent plays the principal’s most preferred action, his belief about the likelihood of the “high opportunity cost” state is decreasing. Intuitively, the optimal contract 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 919 Figure 1. Evolution of actions and beliefs over time. exploits the asymmetry in opportunity costs and lowers the agent’s expected opportunity cost—hence making it easier to incentivize the agent—by biasing information disclosure in the direction of informing him when the opportunity cost is high.2 Figure 1plots four representative evolutions of the agent’s belief about the high opportunity cost state. In each panel, the grey region “OPT” indicates the region at which choosing the principal’s most preferred action is optimal for the agent. An arrow pointing from one belief to another indicates how the agent revises his belief within the period following a signal’s realization. Multiple arrows originating from the same point thus represent the information disclosed by the policy. Within a period, the agent takes a decision after having revised his beliefs. Arrows have different colors/patterns. At all beliefs at the end of continuous black arrows, the agent chooses the principal’s most preferred action. At all beliefs at the end of dotted magenta arrows, he chooses what is best given his current belief. Third, in panels (a), (b), and (c), the policy does not disclose information to the agent at the first period. Starting from the second period, the policy discloses just enough information to compensate the agent for the opportunity cost of choosing the principal’s preferred action; no rent is left to the agent. However, as panel (d) illustrates, in some cases the policy discloses information in the first period, which may leave a strictly positive rent to the agent. For instance, it does so if the promise of full information disclosure at the next period would not incentivize the agent to choose the principal’s preferred 2To be precise, under our policy, upon receiving the signal “the opportunity cost is high,” the agent learns that this is indeed true. However, the signal is not sent with probability one. This corresponds to the (magenta/dotted) arrows pointing at 1 in Figure 1. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
920 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) action. Disclosing information at the first period may also be necessary to reduce the agent’s expected opportunity cost of following the principal’s recommendation. Finally, with the exception of panel (b), the policy does not induce the agent to believe that playing the principal’s most preferred action is optimal. This is markedly different from what we would expect from the static analysis of Kamenica and Gentzkow (2011). Intuitively, the “static” persuasion policy is suboptimal because it does not extract all the information surplus it creates. Even in panel (b), the beliefs do not jump immediately to the “OPT” region. In fact, the belief process may approach the “OPT” region only asymptotically. These properties highlight that in our dynamic environment, information is used as a compensation tool for creating intertemporal incentives, more than as a persuasion tool to affect the agent’s myopic incentives. Related literature The paper is part of the literature on Bayesian persuasion, pioneered by Kamenica and Gentzkow (2011), and recently surveyed by Kamenica (2019). The three most closely related papers are Ball (2023), Ely and Szydlowski (2020), and Orlov, Skrzypacz, and Zryumov (2020). In common with our paper, these papers study the optimal disclosure of information in dynamic games and show how the disclosure of information can be used as an incentive tool. The observation that information can be used to incentivize agents is not new and dates back to the literature on repeated games with incomplete information, for example, Aumann, Maschler, and Stearns (1995). See Garicano and Rayo (2017)andFudenberg and Rayo (2019) for some more recent papers exploring the role of information provision as an incentive tool. The classes of dynamic games studied differ considerably from one paper to another, and this makes comparisons difficult. In Ely and Szydlowski (2020), the agent has to repeatedly decide whether to continue working on a project or to quit (i.e., unlike our paper, there are only two actions); quitting ends the game. The principal aims at maximizing the number of periods the agent works on the project and can only do so by disclosing information about its complexity, modeled as the number of periods required to complete the project. Thus, their dynamic game is a quitting game, while ours is a repeated game. When the project is either easy or difficult (i.e., when there are two states), the optimal disclosure policy initially persuades the agent that the task is easy, so that he starts working. (Naturally, if the agent is sufficiently convinced that the project is easy, there is no need to persuade him initially.) If the project is in fact difficult, the policy then discloses it at a later date, when completing the project is now within reach. A main difference with our optimal disclosure policy is that information comes in lumps in Ely and Szydlowski (2020), that is, information is disclosed only at the initial period and at a later period, while information is repeatedly disclosed in our model.3Another main difference is as follows. In Ely and Szydlowski, only when the promise of full information disclosure at a later date is not enough to incentivize the agent to start working 3When there are more than two states, the optimal policy discloses information more frequently in Ely and Szydlowski (2020). The frequency of disclosure is thus a consequence of the dimensionality of the state space in their model, while it is not so in our model. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 921 does the principal persuade the agent initially. This is not so with our policy: the principal persuades the agent in a larger set of circumstances. This initial persuasion reduces the cost of incentivizing the agent in future periods. Orlov, Skrzypacz, and Zryumov (2020) also consider a quitting game, where the principal aims at delaying the quitting time as far as possible.4The quitting time is when the agent decides to exercise an option, which has different values to the principal and the agent. The principal chooses a disclosure policy informing the agent about the option’s value. When the principal is able to commit to a long-run policy, it is optimal to fully reveal the state with some delay. This policy is not optimal in Ely and Szydlowski (2020), or in our paper. See Au (2015), Bizzotto, Rüdiger, and Vigier (2021), Che, Kim, and Mierendorff (2023), Henry and Ottaviani (2019), and Smolin (2021) for other papers on information disclosure in quitting games, where the agent either waits and obtains additional information, or takes an irreversible action and stops the game. Ball (2023) studies a continuous time model of information provision, where the state changes over time and payoffs are the ones of the quadratic example of Crawford and Sobel (1982). Ball shows that the optimal disclosure policy requires the sender to disclose the current state at a later date, with the delay shrinking over time. The main difference between his work and ours is the persistence of the state (also, we consider two different classes of games). When the state is fully persistent, as in Ely and Szydlowski (2020) and our model, full information disclosure with delay is not optimal in general. (See the discussion of Example 1in Section 3.) Finally, there are a few papers on dynamic persuasion, where the agent takes an action repeatedly. However, either the agent is myopic, for example, Ely (2017)andRenault, Solan, and Vieille (2017), or the principal cannot commit, for example, Escude and Sinander (2023). 2. The problem 2.1 The model A principal and an agent interact over an infinite number of periods, indexed by t∈ {1, 2, }. At the first period, the principal learns a payoff-relevant state ω∈= {ω0,ω1}, while the agent remains uninformed. The prior probability of ωis p0(ω)>0. At each period t, the principal sends a signal s∈Sand, upon observing s, the agent takes decision a∈A.ThesetsAand Sare finite. The cardinality of Sis as large as necessary for the principal to be unconstrained in his information disclosure policy.5Throughout, we interchangeably use the words “period” and “stage.” We assume that there exists a∗∈Asuch that the principal’s stage payoff is strictly positive whenever a∗is chosen, and zero otherwise. The principal’s stage payoff function is thus v:A×→R,withv(a∗,ω0)>0, v(a∗,ω1)>0, and v(a,ω0)=v(a,ω1)=0forall a∈A\{a∗}. The agent’s stage payoff function is u:A×→R. The (common) discount factor is δ∈(0, 1). 4We refer to what Orlov, Skrzypacz, and Zryumov (2020) call the agent as the principal, and vice versa. 5From Makris and Renou (2023), it is enough to have the cardinality of Sas large as the cardinality of A. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
922 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) We write At−1for A×···×A t−1times and St−1for S×···×S t−1times , with generic elements atand st, respectively. A behavioral strategy for the agent is a collection of maps σ=(σt)∞ t=1 with σt:At−1×St→(A). Before learning the state, the principal commits to a strategy, or contract, specifying, as a function of the state, the information to be disclosed (i.e., the statistical experiment to be conducted) at each history of realized signals and actions. Formally, the principal commits to a collection of maps (a contract) τ=(τt)∞ t=1,withτt:At−1×St−1×→(S). The contract enables the principal to use information disclosures to reward or punish the agent for choosing the “right” or the “wrong” action. We denote by V(τ,σ)and U(τ,σ)the principal’s and the agent’s overall expected payoff under the profile (τ,σ).LetPτ,σ(·|ω)be the distribution over sequences of signals and actions induced by (τ,σ)conditional on ω. The principal’s expected payoff V(τ,σ)is ω p0(ω) t st,at−1 (1−δ)δt−1Pσ,τst−1,at−1|ωτtst|st−1,at−1,ωσta∗|st,at−1 ×va∗,ω.(1) The agent’s expected payoff is defined similarly. The objective is to characterize the maximal expected payoff Vmax the principal can achieve by committing to a contract τbefore learning the state, that is, Vmax =⎧ ⎨ ⎩ sup (τ,σ) V(τ,σ) subject to U(τ,σ)≥Uτ,σfor all σ. Several comments are worth making. First, an alternative interpretation of our model is that neither the principal nor the agent know the state, but the principal has the ability to conduct statistical experiments contingent on the state and past signals and actions. Second, the only additional information the agent obtains each period is the outcome of the statistical experiment. Third, the state is fully persistent and the principal perfectly monitors the action of the agent. Finally, the only instrument available to the principal is information. The principal can neither remunerate the agent nor terminate the relationship nor allocate different tasks to the agent. We purposefully make all these assumptions to address our main question of interest: What is the optimal way to incentivize the agent with information only? 2.2 An example Throughout the paper, we illustrate our results with the help of the following example. Example 1. The agent has three possible actions a0,a,1and a∗,witha0(resp., a1)the agent’s optimal action when the state is ω0(resp., ω1). The prior probability of ω1is 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 923 Table 1. Payoff table. a0a1a∗ ω00, 1 0, 0 1, 1/2 ω10, 0 0, 2 1, 1/2 1/3 and the discount factor is 1/2. The payoffs are in Table 1, with the first coordinate corresponding to the principal payoff. We start with few preliminary observations. First, regardless of the agent’s belief, action a∗is never optimal. Second, the opportunity cost of playing a∗is higher when the state is ω1than ω0,thatis,u(a1,ω1)−u(a∗,ω1)>u (a0,ω0)−u(a∗,ω0). It is, therefore, harder to incentivize the agent to play a∗when he is more confident that the state is ω1. As we shall see, the optimal policy exploits this asymmetry. We now consider some simple strategies the principal may commit to. To start with, assume that the principal commits to disclose information at the initial stage only. We call it the KG policy, in reference to Kamenica and Gentzkow (2011). Clearly, since a∗is never optimal, the principal’s payoff is 0. To obtain a positive payoff, the principal must condition his information disclosure on the agent’s actions. The simplest such policy is to “reward” the agent with full disclosure of the state for playing a∗at the beginning of the relationship, say up to period T∗. If the agent deviates, the harshest punishment the principal can impose is to reveal no information in subsequent periods, inducing a normalized expected payoff of 2/3. We are thus looking for the largest T∗such that (1−δ)1 2δ0+δ1+···+δT∗−1+1 3·2+2 3·1δT∗+···≥2 3, which is T∗=ln(5)/ln(2)=2, yielding the principal a payoff of (1−1 2)·(1+1 2)=3 4. Another simple strategy the principal can commit to is a “random full-disclosure policy,” where he fully discloses the state with probability αat period t(and withholds all information with the complementary probability) if the agent plays a∗at period t−1.6 (Again if the agent deviates, the harshest punishment is to withhold all information in all subsequent periods.) Thus, if we write V(resp., U) for the principal (resp., agent) payoff, the best recursive policy is to choose αso as to maximize V=1 21+1 2(1−α)V,subjectto: U=1 21 2+1 2(1−α)U+α4 3≥2 3. The principal’s best payoff is V=4/5withα=1/4. The random full-disclosure policy does better than the policy of fully disclosing the state with delay since it circumvents 6Full information with delay plays an important role in the work of Ball (2023) and Orlov, Skrzypacz, and Zryumov (2020). 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
924 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) the integer constraint on T. Intuitively, it makes it possible to incentivize the agent to play a∗a discounted number of periods slightly larger than 2, namely ln(5)/ln(2). As we will see in Section 3.5, the random full-disclosure policy is still suboptimal since it does not exploit the asymmetry in the agent’s opportunity cost of choosing a∗in the two states. The optimal policy exploits such asymmetry by disclosing no information in the first period and then either revealing that the state is ω1, the high opportunity cost state, or lowering the agent’s belief that the state is ω1. By doing so, the policy incentivizes the agent to take action a∗for a longer expected time. ♦ 3. Optimal contracts This section characterizes optimal contracts and discusses their most salient properties. 3.1 A recursive formulation The first step toward characterizing optimal contracts is to reformulate the principal’s problem as a recursive problem. To do so, we introduce two state variables. The first state variable is promised payoff. It is well known that classical dynamic contracting problems admit recursive formulations if we introduce promised payoff as a state variable and impose promise-keeping constraints, for example, Spear and Srivastava (1987). The second state variable we introduce is beliefs. We now turn to the formal reformulation of the problem. We first need some additional notation. We denote by p∈[0, 1]a generic belief, with pthe probability of ω1.Weletu(a,p):=p[u(a,ω1)−u(a,ω0)] +u(a,ω0)be the agent’s expected stage payoff of choosing awhen his belief is p. We define m(p):= maxa∈Au(a,p)as the agent’s optimal stage payoff when his belief is p,andM(p):= p[m(1)−m(0)] +m(0)as the agent’s expected stage payoff if he learns the state prior to choosing an action. Note that mis a piecewise linear convex function that Mis linear and that m(p)≤M(p)for all p. Similarly, we let v(a,p)be the principal’s expected stage payoff when the agent chooses aand the principal’s belief is p. Finally, let P:={p∈ [0, 1]:m(p)=u(a∗,p)}be the set of beliefs at which a∗is optimal. If nonempty, the set Pis the closed interval [p,¯ p]. Let W⊆[0, 1]×Rbe such that (p,w)∈Wif and only if w∈[m(p),M(p)]. Throughout, we consider the complete metric space of bounded, continuous functions V:W→ R, with the interpretation that V(p,w)is the principal’s payoff if he promises a payoff of wto the agent when the agent’s current belief is p. Consider the following maximization program: T(V)(p,w):= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ max ((λs,(ps,ws),as)∈[0,1]×W×A)s∈S s∈S λs(1−δ)v(as,ps)+δV (ps,ws), subject to: (1−δ)u(as,ps)+δws≥m(ps)for all ssuch that λs>0, s∈S λs(1−δ)u(as,ps)+δws≥w, s∈S λsps=p, s∈S λs=1. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 931 (ii) When v(a∗,0) v(a∗,1)=m(0)−u(a∗,0) m(1)−u(a∗,1), the random full-disclosure policy is optimal. We hasten to stress that both the KG and random full-disclosure policies are not optimal in general, as Example 1demonstrates. See Section 3.5. 3.4 Optimal policy: A formal description We define a family of policies (τq)q∈[q1,q1]indexed by a belief q, and prove later the existence of q∗∈[q1,q1]such that the policy τq∗is optimal. At each (p,w)∈W,apolicy prescribes a feasible tuple (λs,(ps,ws),as)s∈S, that is, a splitting (λs,ps)s∈S, a profile of recommendations (as)s∈Sand a profile of continuation payoffs (ws)s∈S.Therearefour different types of prescription, depending on which of four regions the state variables (p,w)belong to; the belief qparameterizes these regions. The four regions are W1 q:=(p,w):p∈0, q1),w≤q1−p q1m(0)+p q1mq1, W2 q:=(p,w):p∈(q,1 ,1−p 1−qm(q)+p−q 1−qm(1)<w≤1−p 1−q1mq1+p−q1 1−q1m(1) ∪(p,w):p∈q1,q,w≤1−p 1−q1mq1+p−q1 1−q1m(1), W3 q:=(p,w):p∈(q,1 ],w≤1−p 1−qm(q)+p−q 1−qm(1), W4 q:=W\W1 q∪W2 q∪W3 q. Figure 3illustrates the four regions, with W1 qthe black region, W2 qthe region with vertical lines, W3 qthe gray region, and W4 qthe region with slanted lines. Observe that regions W1 q and W4 qdo not depend on the parameter q, while the other two do. We begin with an informal overview of our optimal policy. In region W2 q, the principal recommends a∗and discloses information in the next period as a reward. The belief pdecreases over time, until either it reaches a point at which the agent will choose a∗ forever, or it enters region W4 q. Figure 3. Regions W1 q(black), W2 q(vertical lines), W3 q(gray), and W4 q(slanted lines). 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
932 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) In region W4 q, the principal discloses the state with sufficiently high probability so that, when disclosure does not occur, the agent’s belief is q1. At the belief q1, the agent plays a∗one final time. (Recall that at q1, the agent plays a∗if rewarded with full information disclosure at the next period.) In region W1 q, the belief pis so low that even the promise of full information disclosure at the next period does not incentivize the agent to play a∗, not even once. In this region, the principal sends either the signal s∗or the signal s0.Thesignals0perfectly informs the agent that the state is ω0, while the signal s∗induces the belief q1,atwhich the agent plays a∗one final time. In region W3 q, the belief pis higher than qand, even possibly, higher than q1.Inthis region, the principal sends either the signal s∗or the signal s1.Thesignals1perfectly informs the agent that the state is ω1, while the signal s∗induces the belief q≤q1,at which the agent plays a∗. It is almost the mirror image of what the policy does in region of W1 q; the only conceptual difference is that the policy induces the belief qrather than q1, the natural counterpart of q1. This asymmetry is a consequence of trying to induce a∗at the lowest possible average belief. We now formally define the policy τq, starting with region W2 q. Define the functions λ:W→[0, 1]and ϕ:W→[0, 1]so that (λ(p,w),ϕ(p,w)) is the unique solution of p w=λ(p,w)ϕ(p,w) mϕ(p,w)+1−λ(p,w)1 m(1)(5) for all w>m (p)and (λ(p,m(p)),ϕ(p,m(p))) =(1, p).When (p,w)is in region W2 q, the policy splits pinto two beliefs ϕ(p,w)and 1, with probability λ(p,w)and 1−λ(p,w), respectively. When the posterior belief is ϕ(p,w), the policy recommends a∗and promises the continuation payoff w(ϕ(p,w)) if the recommendation is followed. Therefore, if the agent follows the recommendation, his discounted expected payoff is m(ϕ(p,w)) =(1−δ)u(a∗,ϕ(p,w)) +δw(ϕ(p,w)). When the posterior belief is 1, the policy recommends a1and promises the continuation payoff m(1),witha1an optimal action at state ω1. Therefore, if the agent follows the recommendation, he achieves the discounted expected payoff m(1). Note that when w=m(p), the principal recommends a∗with probability one, and promises the continuation payoff w(p)in the future. Upon following the recommendation, the agent achieves the discounted expected payoff m(p). The key feature of the policy in region W2 qis to disclose, with some probability, that the state is ω1. As we already suggested, the rationale for disclosing when the state is ω1is two-fold. First, the lower the agent’s belief, the lower the cost of incentivizing the agent to play a∗relative to the principal’s benefit. Second, to satisfy the promise-keeping constraint, the policy needs to compensate the agent for playing a∗. Since the principal’s payoff is zero when the agent takes any action different from a∗, the best is to choose a compensation, which guarantees the highest probability of playing a∗. Putting these two observations together, at (p,w),policyτq(p,w)finds two beliefs (p,p )such that (i) the agent is asked to play a∗at p, (ii) p<psince the agent should play a∗at the lowest belief, and (iii) the probability of pis as high as possible. The best splitting is to 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 933 Figure 4. Construction of λand ϕ:p=λϕ +(1−λ)1; w=λm(ϕ)+(1−λ)m(1). have pas close as possible to pand p as far as possible, that is, equal to 1. Observe that since (1−λ(p,w))m(1)+λ(p,w)m(ϕ(p,w)) =w, the promise-keeping constraint binds in region W2 q. See Figure 4for an illustration. Note that starting with (p,w)∈W2 q, the decreasing sequence of beliefs (ϕ(p,w), ϕ2(p,w),)(and corresponding payoffs) reaches either region W4 q—as in panels (A) and (C)of Figure 1—or a belief in Pat which it is statically optimal for the agent to play a∗—as in panel (B)of Figure 1.14 In the latter case, the policy recommends a∗and stops disclosing information (i.e., the belief stays constant). When (p,w)is in region W4 q, the agent cannot be incentivized to play a∗at (p,w).15 In that case, the policy splits pinto posteriors 0, q1, and 1 with respective probabilities λ0,λq1,andλ1. Conditional on 0 (resp., 1), the policy recommends an action optimal at 0, (resp., an action optimal at 1), and promises a continuation payoff of m(0)(resp., m(1)). Conditional on q1, the policy recommends action a∗and promises a continuation payoff of w(q1). Doing so, the principal ensures that the agent plays a∗one more time. The probabilities (λ0,λq1,λ1)∈R+×R+×R+are the unique solution to λ0⎛ ⎜ ⎝ 0 m(0) 1⎞ ⎟ ⎠+λq1⎛ ⎜ ⎝ q1 mq1 1 ⎞ ⎟ ⎠+λ1⎛ ⎜ ⎝ 1 m(1) 1⎞ ⎟ ⎠=⎛ ⎜ ⎝ p w 1⎞ ⎟ ⎠. A solution exists since W4 qis the convex hull of (0, m(0)),(q1,m(q1)),and(1, m(1)).In this region, the promise-keeping constraint is also binding. When (p,w)is in region W1 q, the policy splits pinto 0 (i.e., discloses that the state is ω0)andq1with respective probabilities q1−p q1and p q1. If the realized belief is 0, the policy recommends an action optimal at 0 and promises a continuation payoff of m(0).If the realized belief is q1, the policy recommends a∗and promises a continuation payoff of w(q1). The agent is thus made indifferent between playing a∗and receiving w(q1)in the future, and playing a best reply to the belief q1forever. Intuitively, in region W1 q,the principal cannot incentivize the agent to take action a∗by promising future information 14We write ϕ2(p,w)for ϕ(ϕ(p,w),m(ϕ(p,w))). 15Recall that q1is the lowest belief at which the agent can be incentivized to play a∗. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
934 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) disclosure (since p<q 1). Hence, the principal must first persuade the agent by disclosing some information. Note that the promise-keeping constraint is slack in this region whenever (p,w)satisfies q1−p q1m(0)+p q1m(q1)>w. When (p,w)is in region W3 q, the policy splits pinto qand 1 with respective probabilities 1−p 1−qand p−q 1−q. Conditional on 1, the policy recommends an action optimal at 1 and promises a continuation payoff of m(1). Conditional on q, the policy recommends a∗and promises a continuation payoff of w(q). The agent is thus made indifferent between playing a∗and receiving w(q)in the future, and playing a best-reply to the belief q forever. The policy in this region is analogous to the one in region W1 q—the policy starts by disclosing some information. When q=q1, the reason for the analogy is immediate, as q1is the highest belief at which the agent is willing to take action a∗at the current period in exchange for full information at the next period. As we shall see later, the optimal policy τq∗may require q∗< q1, in order to guarantee that the principal’s value function is concave, a necessary requirement to minimize the cost of incentivizing the agent relative to the benefit to the principal. As in region W1 q, the promise-keeping constraint is also slack in this region whenever (p,w)satisfies w<1−p 1−qm(q)+p−q 1−qm(1).This completes the description of the policy τq. Before moving on, we first verify that our policy τq∗is optimal under the two benchmark scenarios discussed in Section 3.3. Given the value function, we just need to check whether τq∗solves the Bellman equation. Corollary 2. The policy τq∗is optimal both when |A|=2and when v(a∗,0) v(a∗,1)= m(0)−u(a∗,0) m(1)−u(a∗,1). We now illustrate our construction by revisiting Example 1. 3.5 Example 1revisited We have that M(p)=1+p,m(p)=max(1−p,2p)and w(p)=2max(2p,1−p)−(1/2). Therefore, Q1=[1/6, 1/2]. Assume that q=1/3 (we will show that this choice is the optimal’s one). Remember that the prior probability of ω1is 1/3 and the discount factor is 1/2. Let us start with the pair (p,m(p)) =(1/3, 2/3), which is in region W2 1/3.The policy recommends a∗to the agent and promises a continuation payoff of w(1/3)=5/6. The next value of the state variables is therefore (1/3, 5/6), which is again in W2 1/3.Ifthe agent had been obedient, the policy then splits the prior probability 1/3 into 3/11 and 1 with probability 22/24 and 2/24, respectively. Indeed, we have ⎛ ⎜ ⎝ 1 3 5 6 ⎞ ⎟ ⎠=22 24 ⎛ ⎜ ⎝ 3 11 m3 11⎞ ⎟ ⎠+2 24 1 m(1). Conditional on the posterior 3/11, the policy recommends a∗to the agent and promises a continuation payoff of w(3/11)=21/22. Conditional on the posterior 1, the 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 935 Figure 5. Evolution of the beliefs. policy recommends a1and promises a continuation payoff of m(1)=2. Therefore, the next value of the state variables is either (3/11, 21/22)or (1, 2), with the former again in W2 1/3. If the value of the state variables is (1, 2), the policy yet again recommends a1and a continuation payoff of 2. If the value of the state variables is (3/11, 21/22), the policy splits 3/11 into 7/39 and 1, with probability 39/44 and 5/44, respectively. Conditional on the posterior 7/39, the policy recommends a∗to the agent and promises a continuation payoff of w(7/39)=89/78. Conditional on the posterior 1, the policy recommends a1 and promises a continuation payoff of m(1)=2. Finally, at the state variables value of (7/39, 89/78), which is in region W4 1/3, the policy does a penultimate split of 7/39 into 0, 1/6 and 1 with probability 113/156, 18/156 and 25/156, respectively. Conditional on the posterior 1/6, the policy recommends a∗ and promises a continuation payoff of 7/6, that is, full information disclosure at the next period. The policy fully discloses the state in finite time to the agent. See Figure 5for the evolution of the beliefs at the beginning of each period. At all beliefs other than 0 and 1, the agent is recommended to play a∗. The principal’s expected payoff is 1285/1536, that is, about 0.83. Our optimal policy performs strictly better than the random full-disclosure policy because it exploits the asymmetry in the agent’s opportunity cost of choosing a∗in the two states. At each period in which information is disclosed and a∗is played, our policy decreases the belief at which a∗is played; the average discounted beliefs is p∗≈0.197 < 1/3. On the contrary, the random full-disclosure policy does not alter the belief that the state is ω1when a∗is played; the belief stays fixed at the prior p0=1/3. It remains to explain how to choose the parameter q∗to guarantee the optimality of τq∗. 3.6 Construction of q∗and optimality For all q∈[q1,q1],letVq:W→Rbe the value function induced by the policy τq.For all q,notethatVq(1, m(1)) =0 since a∗is not optimal at p=1, and Vq(0, m(0)) =0if a∗is not optimal at p=0(resp.,Vq(0, m(0)) =v(a∗,0 )if a∗is optimal at p=0). Also, Vq(q1,m(q1)) =(1−δ)v(a∗,q1)if q1>0(resp.,Vq(0, m(0)) =v(a∗,0 )if q1=0, since a∗ is then optimal at p=0). Therefore, any two policies τqand τqinduce the same values at all (p,w)∈W1 q∪W4 q=W1 q∪W4 q. (Remember that the regions W1 qand W4 qdo not vary with q; see Figure 3.) 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
936 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) Similarly, any two policies τqand τqinduce the same values at all (p,w)∈W2 min(q,q). Thus, in particular, τqand τq1induce the same values at all (p,w)∈W\W3 q. Finally, at all (p,w)∈W3 q,Vq(p,w)=1−p 1−qVq(q,m(q)) =1−p 1−qVq1(q,m(q)). Hence, characterizing Vq1is enough to characterize Vq. (See Appendix Bfor more details.) Recall that V∗is the unique solution to the fixed-point problem—to be optimal, a policy must therefore induce the value function V∗.Let q∗=supp∈q1,q1:Vq1p,m(p)≥Vq1(p,w)for all w. We are now ready to state our main result. Theorem 1. The policy τq∗is optimal: Vq∗=V∗. To understand the role of q∗, recall that for all p∈[q∗,1 ], the policy leaves rents to the agent.16 To minimize these rents, the principal therefore would like to have q∗as high as possible, that is, equal to q1, the highest belief at which the agent is willing to play a∗in exchange for full information disclosure at the next period. However, Vq1(·,m(·)) is not guaranteed to be concave in p, a necessary condition for optimality. To see that V∗(·,m(·)) must be concave in p, consider any pair (p,p)∈[0, 1]×[0, 1]and α∈[0, 1]. We have αV ∗p,m(p)+(1−α)V∗p,mp≤V∗αp +(1−α)p,αm(p)+(1−α)mp ≤V∗αp +(1−α)p,mαp +(1−α)p, where the first inequality follows from the concavity of V∗in both arguments and the second from V∗decreasing in wand the convexity of m. The optimal choice of q∗is thus the largest q, which guarantees Vq(·,m(·)) to be concave. More precisely, as we show in Appendix A.5, the definition of q∗guarantees that Vq∗ is concave in both arguments and decreasing in w,sothatVq∗(·,m(·)) is a concave function of p.WealsoprovethatVq∗(p,m(p)) ≥Vq1(p,m(p)) for all p. Since it is clearly the smallest such function, Vq∗is the concavification of Vq1.Inparticular,q∗=q1if Vq1(·,m(·)) is already concave. Figure 6illustrates the concavification for Example 1.In dashed red is the value function of policy τq1; in solid blue its concavification—the value function of policy τq∗,withq∗=1 3. The policy τq∗leaves rents to the agent, that is, the (ex ante) participation constraint does not bind, for all priors in [0, q1)∪(q∗,1 ]. Thisisquitenaturalforallpriorsin [0, 1]\Q1since the agent cannot be incentivized to play a∗even once. In the language of Ely and Szydlowski (2020), “the goalposts need to move,” that is, one needs to disclose information at the ex ante stage to persuade the agent to play a∗. However, our policy also leaves rents for all priors in (q∗,q1]. The intuitive reason is that the initial information disclosure reduces the cost of incentivizing the agent in subsequent periods sufficiently enough to compensate for the initial loss. (When the realized posterior is 1, the agent never plays a∗, thus creating the loss.) 16That is, the agent is promised a payoff of 1−p 1−q∗m(q∗)+p−q∗ 1−q∗m(1)>m (p). 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 937 Figure 6. The concavification of Vq1(·,m(·)) in Example 1. 4. Evolution of beliefs in the optimal policy The optimal policy discloses information gradually over time, with beliefs evolving until either the agent learns the state or believes that a∗is statically optimal. We can be more specific. First, we consider the instances when the policy converges with positive probability to a belief p∈P=[p,p], the set of beliefs at which a∗is optimal. Let Q∞=[p,q∞], with q∞the solution to mq∞=(1−δ)ua∗,q∞+δ1−q∞ 1−pm(p)+q∞−p 1−pm(1), if Pis nonempty, and Q∞=∅,otherwise.NotethatP⊆Q∞. See Figure 7for a graphical illustration. Intuitively, the set Q∞has the “fixed-point property,” that is, if one starts with a belief p∈Q∞and promised utility w(p), then the belief ϕ(p,w(p)) ∈Q∞.Toseethis,note that the pair (p,w(p)) is in region W2 q. Since ϕ(p,w(p)) ≤p(with a strict inequality if p/∈P), we then have a decreasing sequence of beliefs converging to an element in P. This is because, at all beliefs p∈Q∞, the policy splits pinto p=ϕ(p,w(p)) and 1, Figure 7. Construction of q∞. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
938 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) then splits pinto p =ϕ(p,w(p)) and 1, etc. The decreasing sequence (p,p,p,) converges, either in finite time or asymptotically, to a belief in P, at which no further splitting occurs and the agent plays a∗forever. See panel (b) of Figure 1for an an illustration. Recall that if the prior p0is larger than q∗, the policy first splits p0into q∗and 1. Hence, if q∗≤q∞, the agent’s belief enters the set Q∞with strictly positive probability.17 Therefore, if the agent’s prior belief is in the set Q∞ q∗, then there is a strictly positive probability that the agent chooses action a∗forever, where Q∞ q∗:=!Q∞if q∗> q∞, [p,1 )otherwise. Second, at all priors in [0, 1]\Q∞ q∗,thereexistsTδ<∞such that the belief process is absorbed in the degenerate beliefs 0 or 1 after at most Tδperiods. In other words, the agent learns the state for sure in finite time. The number of periods Tδcorresponds to the maximal number of periods the agent can be incentivized to play a∗.Weprovidean explicit computation in Appendix B.InExample1,Tδ=3. Moreover, the number Tδis increasing in δand converges to +∞ as δconverges to 1. (Note that the convergence is uniform in that it does not depend on p0∈[0, 1]\Q∞ q∗.) Thus, we have the following corollary. Corollary 3. Under the optimal disclosure policy τq∗, there is a strictly positive probability that the agent chooses action a∗forever if, and only if, p0∈Q∞ q∗. Alternatively, if p0/∈Q∞ q∗, then there exists Tδsuch that the agent perfectly learns the state (i.e., preaches either 0 or 1) with probability 1after at most Tδperiods. The interval Q∞ q∗includes the subinterval [p,¯ p], where the agent takes action a∗with probability one. In the complementary set Q∞ q∗\[p,¯ p], the probability that the agent takes action a∗forever is strictly less than 1. That is, the principal discloses the state with positive probability, and with the complementary probability he lowers the agent’s belief so that it converges to the region where taking action a∗is statically optimal. Convergence may be asymptotic or may happen in finite time. As already mentioned, the promise-keeping constraint binds in regions W2 q∗and W4 q∗, but may not bind in the other two regions. We now argue that under our policy τq∗, the promise-keeping constraint can only be slack in the first period. In other words, the promised-keeping constraint binds from period two onwards. To see this, suppose that (p0,m(p0)) is in region W3 q∗, hence the prior belief p0∈(q∗,1 ). What the policy τq∗ does is to split p0into q∗and 1, so that the state variable transit to either (q∗,m(q∗)) or (1, m(1)). In the latter case, the promise-keeping constraint clearly binds and will continue to bind in all subsequent periods, since the agent has learned that the state is ω1. In the former case, since (q∗,m(q∗)) ∈W2 q∗, the promise-keeping constraint binds and will continue to bind in all subsequent periods since the subsequent state variables will 17From the definition of q∗,wehavethatq∗≥psince Vq1(p,m(p)) =u(a∗,p)for all p∈P. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 939 either be in regions W2 q∗or W4 q∗or equal to (1, m(1)). A symmetric argument holds when (p0,m(p0)) is in region W1 q∗. Corollary 4. Under the optimal policy τ∗ q, the promise-keeping constraint can only be slack in the first period. All in all, information disclosure plays two roles in our optimal policy. First, the promise of future information disclosure motivates the agent to take action a∗in early periods. The intertemporal incentives make it possible to motivate to play a∗at beliefs outside P. Second, information disclosure decreases the discounted average belief that the state is the high opportunity cost state ω1and, therefore, makes it easier to incentivize the agent to take action a∗for a longer expected time. Appendix A: Proofs A.1 Mathematical preliminaries We collect without proofs some useful results about concave functions. Let f:[a,b]→R be a concave function and a≤x<y<z≤b. The following properties hold: (a) f(y)−f(x) y−x≥f(z)−f(y) z−y. (b) f(y)−f(a) y−a≥f(z)−f(a) z−a. (c) f(b)−f(x) b−x≥f(b)−f(y) b−y. (d) f(y)−f(x) y−x≥f(y+)−f(x+) y−xfor all ≥0suchthaty+≤b. Note that property (a) implies (d) and is true irrespective of whether x+y.Wewill repeatedly use these properties in most of the following proofs. To prove Lemma 3, we will use the following property: if f:[a,b]→Rsatisfies f(x)−f(a) x−a≥f(y)−f(a) y−afor all a<x≤y≤b,thenfis concave. A.2 Proposition 2 Proof of Proposition 2(i). By contradiction, assume that there exists s∈Ssuch that λs>0and (1−δ)v(as,ps)+δV ∗(ps,ws)<V∗ps,(1−δ)u(as,ps)+δws. Let (λ∗ s,p∗ s,w∗ s,a∗ s)s∈Sbe the policy, which achieves V∗(ps,(1−δ)u(as,ps)+δws),and consider the new policy (λs,ps,ws,as)s∈S\{s},λsλ∗ s,p∗ s,w∗ s,a∗ ss∈S. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
940 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) By construction, the new policy is feasible. Moreover, we have that s∈S\{s} λs(1−δ)v(as,ps)+δV ∗(ps,ws)+λs s∈S λ∗ s(1−δ)va∗ s,p∗ s+δV ∗p∗ s,w∗ s = s∈S\{s} λs(1−δ)v(as,ps)+δV ∗(ps,ws)+λsV∗ps,(1−δ)u(as,ps)+δws > s∈S λs(1−δ)v(as,ps)+δV ∗(ps,ws), a contradiction with the optimality of (λs,ps,ws,as)s∈S. Thus, we must have (1−δ)v(as, ps)+δV ∗(ps,ws)≥V∗(ps,(1−δ)u(as,ps)+δws)for all ssuch that λs>0. Since the fixed point satisfies V∗(ps,(1−δ)u(as,ps)+δws)≥(1−δ)v(as,ps)+ δV ∗(ps,ws),wehavethedesiredresult. Proof of Proposition 2(ii). Let s∈Ssuch that λs>0andas= a∗.Wehave (1−δ)v(as,ps)+δV ∗(ps,ws)=δV ∗(ps,ws)≥V∗ps,(1−δ)u(as,ps)+δws ≥V∗(ps,ws), where the first inequality follows from Proposition 2(i) and the second follows from V∗ decreasing in wand ws≥u(as,ps)for (1−δ)u(as,ps)+δws≥m(ps), to hold. It follows that V∗(ps,ws)=0. Proof of Proposition 2(iii). The proof is by contradiction. Suppose to the contrary that V∗(ps,ws)=V∗(ps,w s)for some w s∈(ws,M(ps)] and as=a∗. By Proposition 2(i), we have V∗ps,(1−δ)ua∗,ps+δws=(1−δ)va∗,ps+δV ∗(ps,ws) =(1−δ)va∗,ps+δV ∗ps,w s ≤V∗ps,(1−δ)ua∗,ps+δw s. Since V∗is decreasing in w, the inequality cannot be strict, hence V∗ps,(1−δ)ua∗,ps+δws=V∗ps,(1−δ)ua∗,ps+δw s.(6) Wenowshowthat V∗ps,(1−δ)ua∗,ps+δws=V∗(ps,ws),(7) hence V∗ps,(1−δ)ua∗,ps+δw s=V∗ps,w s=va∗,ps, 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 947 This implies that ¯ p 1−¯ p=ua∗,0 −ua†,0 ua∗,1 −ua†,1 . Assuming p0>¯ p, under our policy, the principal recommends the agent to take a∗in the first period and promises to split p0between 1 and ˜ pwith probability λin the second period, where (λ,˜ p)solves ⎧ ⎪ ⎨ ⎪ ⎩ λ˜ p+(1−λ)1=p0 λua∗,˜ p+(1−λ)ua†,1 =w(p0)=ua†,p0−(1−δ)ua∗,p0 δ. Replacing λ˜ p=λ−(1−p0)and λ(1−˜ p)=(1−p0)into the second equation yields λua∗,˜ p+(1−λ)ua†,1 =λua∗,1 −(1−p0)ua∗,1 +(1−p0)ua∗,0 +(1−λ)ua†,1 =ua∗,p0+(1−λ)ua†,1 −ua∗,1 =ua†,p0−(1−δ)ua∗,p0 δ =⇒ λ=1−ua†,p0−ua∗,p0 δua†,1 −ua∗,1 =1−p0 δ−1−p0 δ ¯ p 1−¯ p. Then it follows that the principal’s payoff is V=(1−δ)va∗,p0+δλva∗,˜ p =(1−δ)p0+δλ ˜ pva∗,1 +(1−δ)(1−p0)+δλ(1−˜ p)va∗,0 =p0−δ(1−λ)va∗,1 +(1−p0)va∗,0 =va∗,p0−δ(1−λ)va∗,1 =va∗,p0−p0+(1−p0)¯ p 1−¯ pva∗,1 =1−p0 1−¯ pva∗,¯ p, which is exactly the payoff under the KG policy, which splits the initial belief p0into ¯ p with probability 1−p0 1−¯ pand 1 with the complementary probability. Let VRbe the value function under the random full-disclosure policy. To show that our policy τis also optimal when v(a∗,0) m(0)−u(a∗,0)=v(a∗,1) m(1)−u(a∗,1), we need to verify that VR(p,w)= ps∈supp(τ) τ(ps)(1−δ)v(as,ps)+δV R(ps,ws),∀(p,w)∈W. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
948 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) Note that VR(p,w)=M(p)−w M(p)−u(a∗,p)v(a∗,p), since the probability of full disclosure αsatisfies αM(p)+(1−α)u(a∗,p)=w.Hence, ps∈supp(τ) τ(ps)(1−δ)v(as,ps)+δV R(ps,ws) =λ·(1−δ)va∗,ˆ p+δM(ˆ p)−w(ˆ p) M(ˆ p)−ua∗,ˆ pva∗,ˆ p =λM(ˆ p)−m(ˆ p) M(ˆ p)−ua∗,ˆ pva∗,ˆ p, where (λ,ˆ p)solves !λˆ p+(1−λ)1=p λm(ˆ p)+(1−λ)m(1)=w. Since v(a∗,0) m(0)−u(a∗,0)=v(a∗,1) m(1)−u(a∗,1),wehave v(a∗,ˆ p) M(ˆ p)−u(a∗,ˆ p)=v(a∗,p) M(p)−u(a∗,p)=v(a∗,1) M(1)−u(a∗,1). Therefore, recalling that v(a∗,1 )=0, we have λM(ˆ p)−m(ˆ p) M(ˆ p)−ua∗,ˆ pva∗,ˆ p =λM(ˆ p)−m(ˆ p) M(ˆ p)−ua∗,ˆ pva∗,ˆ p+(1−λ)M(1)−m(1) M(1)−ua∗,1 va∗,1 =va∗,p M(p)−ua∗,pλM(ˆ p)−m(ˆ p)+(1−λ)M(1)−m(1) =va∗,p M(p)−ua∗,pM(p)−w=VR(p,w). A.5 Theorem 1 To prove Theorem 1, we first introduce the following lemma. Lemma 2. Consider any feasible policy inducing the value function ˜ V.If ˜ Vis concave in both arguments, decreasing in wand satisfies ˜ Vp,m(p)≥(1−δ)va∗,p+δ˜ Vp,w(p), for all p∈Q1, then the policy is optimal. Proof. We argue that ˜ Vis the fixed point of the operator T,hence ˜ V=V∗.Let (λs,ps,ws,as)s∈Sbe a solution to the maximization problem T(˜ V)(p,w).Westartby the following observation. Consider any ssuch that as= a∗.Wehave (1−δ)v(as,ps)+δ˜ V(ps,ws)=δ˜ V(ps,ws)≤˜ V(ps,ws)≤˜ Vps,(1−δ)u(as,ps)+δws, 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 949 where the last inequality follows from the fact that ˜ Vis decreasing in wand m(ps)≤ (1−δ)u(as,ps)+δws≤(1−δ)m(ps)+δws≤ws. Consider now any ssuch that as=a∗. Since (λs,ps,ws,as)s∈Sis feasible, we have (1−δ)ua∗,ps+δws≥m(ps), hence ps∈Q1and, therefore, ˜ Vps,m(ps)≥(1−δ)va∗,ps+δ˜ Vps,−(1−δ)ua∗,ps+m(ps) δ w(ps) . The concavity of ˜ Vimplies that ˜ Vps,(1−δ)ua∗,ps+δws−˜ Vps,m(ps)≥δ˜ V(ps,ws)−˜ Vps,w(ps), whereweusetheidentity(1−δ)u(a∗,ps)+δws−m(ps)=δ(ws−w(ps)) and observation (a) about concave functions in Section A.1. Combining the above two inequalities implies ˜ Vps,(1−δ)ua∗,ps+δws≥(1−δ)va∗,ps+δ˜ V(ps,ws). It follows that T(˜ V)(p,w)= s∈S λs(1−δ)v(as,ps)+δ˜ V(ps,ws) ≤ s∈S λs˜ Vps,(1−δ)u(as,ps)+δws ≤˜ V s∈S λsps, s∈S λs(1−δ)u(as,ps)+δws) ≤˜ V(p,w), where the second inequality follows from the concavity of ˜ Vand the third inequality from ˜ Vbeing decreasing in w. Conversely, since the policy inducing ˜ Vis feasible, we must have that T(˜ V)(p,w)≥ ˜ V(p,w)for all (p,w). This completes the proof. Invoking Lemma 2, we only need to prove the following proposition to prove Theorem 1. Proposition 4. Let Vq∗be the value function induced by the policy τ∗,with q∗=supp∈Q1:Vq1p,m(p)≥Vq1(p,w)for all w. Then Vq∗is concave in (p,w), decreasing in w, and satisfies Vq∗p,m(p)≥(1−δ)va∗,p+δVq∗p∗,w(p), for all p∈Q1. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
950 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) Figure 8. The function mq. Proving Proposition 4requires to construct the value function Vqinduced by the policy τq. The construction is tedious, and we postpone it to Appendix B.Intherestofthis section, we only report the properties we need to prove Proposition 4. We start with an important identity, which we will use throughout. For any q∈ [q1,q1], define the function mq:[0, 1]→Ras ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ 1−p q1m(0)+p q1mq1if p∈0, q1, m(p)if p∈(q1,q], 1−p 1−qm(q)+p−q 1−qm(1)if p∈(q,1 ]. Note that mqis convex, mq(p)≥m(p)for all p∈[0, 1],mq(0)=m(0),andmq(1)= m(1). For a graphical illustration, see Figure 8. It is straightforward to check that we have the following identity: Vq(p,w)=λ(p,w)Vq(ϕ(p,w),mqϕ(p,w),(9) where the functions λand ϕare defined as in the main text, but with mqinstead of m; see equation (5). This identity states that knowing Vqon the set {(p,w)∈W:(p,w)= (p,mq(p))}suffices to reconstruct Vqat all points on its domain. We now make two additional observations. Observation A. For all q∈[q1,q1], we have the following identity: Vq(p,w)=1−p 1−pVqp,1−p 1−pw+p−p 1−pmq(1). Proof of Observation A. Let w=1−p 1−pw+p−p 1−pmq(1). Assume that w> mq(p). Since λp,wϕp,w mqϕp,w+1−λp,w1 mq(1)=p w, we have 1−p 1−pλp,wϕp,w mqϕp,w+1−1−p 1−pλp,w1 mq(1)=p w. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 951 Therefore, λ(p,w)=1−p 1−pλ(p,w)and ϕ(p,w)=ϕ(p,w)since the solution (λ(p,w), ϕ(p,w)) is unique when w>m q(p). The statement then follows from equation (9). Assume that w=mq(p). From the convexity of mq, this requires that w=mq(p),so that mq(p)=1−p 1−pmq(p)+p−p 1−pmq(1). The result follows from continuity as Vqp,mq(p)=lim w→mq(p)Vq(p,w), =lim w→mq(p) 1−p 1−pVqp,1−p 1−pw+p−p 1−pmq(1), =1−p 1−pVqp,1−p 1−pmq(p)+p−p 1−pmq(1), =1−p 1−pVqp,mqp. Note that this implies that Vqp,w(p)+c=λp,w(p)Vqϕp,w(p),mqϕp,w(p)+c λp,w(p), where cis a positive constant. Observation B. The value function Vq1(p,·):[mq1(p),M(p)] →Ris concave in w, for each p. See Lemma 3in Section B.2. A.5.1 Proposition 4(a) We prove that Vq∗is decreasing in w. To start with, fix p∈[0, 1] and (w,w)∈[mq∗(p),M(p)] ×[mq∗(p),M(p)],withw>w. First, assume that p≤q∗.Ifw=mq∗(p),thenVq∗(p,w)≤Vq∗(p,w)by construction of q∗.Ifw>mq∗(p),wehavethat Vq∗p,w−Vq∗(p,w) w−w=Vq1p,w−Vq1(p,w) w−w ≤Vq1(p,w)−Vq1p,mq∗(p) w−mq∗(p) =Vq∗(p,w)−Vq∗p,mq∗(p) w−mq∗(p)≤0, where the inequality follows from the concavity of Vq1with respect to w,forallw≥ mq1(p).(Recallthatmq∗(p)=mq1(p)for all p≤q∗.) Second, assume that p>q ∗. We show in detail how to make use of Observation A to deduce the result. We repeatedly use similar computations later on. We have Vq∗p,w=λp,wVq∗ϕp,w,mq∗ϕp,w =λp,w1−ϕp,w 1−ϕ(p,w)Vq∗ϕ(p,w),1−ϕ(p,w) 1−ϕp,wmq∗ϕp,w 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
952 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) +1−1−ϕ(p,w) 1−ϕp,wmq∗(1) =λ(p,w)Vq∗ϕ(p,w),λp,w λ(p,w)mq∗ϕp,w+1−λp,w λ(p,w)mq∗(1) =λ(p,w)Vq∗ϕ(p,w),mq∗ϕ(p,w)+w−w λ(p,w), where the first line follows from the construction of Vq∗, the second line from Observation A, the third line from the definition of the functions λand ϕ, and the last line from the following computations: λp,w λ(p,w)mq∗ϕp,w+1−λp,w λ(p,w)mq∗(1) =1 λ(p,w)w+1−1 λ(p,w)mq∗(1) =1 λ(p,w)w+1−1 λ(p,w)w−λ(p,w)mq∗ϕ(p,w) 1−λ(p,w) =mq∗ϕ(p,w)+w−w λ(p,w). Thus, we are able to express Vq∗(p,w)as λ(p,w)Vq∗(ϕ(p,w),˜ w),with ˜ wthe above expression. Moreover, ϕ(p,w)≤q∗as w≥mq∗(p). We can use the (already established) concavity of Vq∗in wfor each p≤q∗to deduce the desired result. More precisely, we have that Vq∗p,w−Vq∗(p,w) w−w = λ(p,w)Vq∗ϕ(p,w),mq∗ϕ(p,w)+w−w λ(p,w)−Vq∗ϕ(p,w),mq∗ϕ(p,w) w−w ≤0, where the inequality follows from the concavity of Vq∗in wat all p≤q∗. Lastly, since Vq∗(p,w)=Vq∗(p,mq∗(p)) for all w∈[m(p),mq∗(p)], the result immediately follows for all (w,w),withw∈[m(p),mq∗(p)]. A.5.2 Proposition 4(b) We prove the concavity of Vq∗with respect to both arguments (p,w). Let W={(p,w):w≥mq∗(p)}.Let (p,w)∈W,(p,w)∈Wand α∈[0, 1].Write (pα,wα)for αp w+(1−α)p w. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 953 Without loss of generality, assume that p≤p.Wehavethat αVq∗(p,w)+(1−α)Vq∗p,w =α1−p 1−pVq∗p,1−p 1−pw+p−p 1−pmq∗(1) ≥mq∗(p) +(1−α)Vq∗p,w ≤α1−p 1−p+(1−α)Vq∗⎛ ⎜ ⎜ ⎝p, α1−p 1−p1−p 1−pw+p−p 1−pmq∗(1)+(1−α)w α1−p 1−p+(1−α) ⎞ ⎟ ⎟ ⎠ =1−pα 1−pVq∗p,1−p 1−pα wα+p−pα 1−pα mq∗(1) =Vq∗(pα,wα), where the inequality follows from the concavity of Vq1with respect to wfor each pand the property that Vq∗(p,w)=Vq1(p,w)for all (p,w)such that w≥mq∗(p).Noticethat we use twice Observation A. Finally, for all (p,w)∈W,forall(p,w)∈Wand for all α,wehavethat αVq∗(p,w)+(1−α)Vq∗p,w =αVq∗p,maxw,mq∗(p)+(1−α)Vq∗p,maxw,mq∗p ≤Vq∗pα,αmaxw,mq∗(p)+(1−α)maxw,mq∗p ≤Vq∗(pα,wα), since αmax(w,mq∗(p))+(1−α)max(w,mq∗(p)) ≥wαand the fact that Vq∗is decreasing in wfor all p. This completes the proof of concavity. A.5.3 Proposition 4(c) We prove that Vq∗(p,m(p)) ≥(1−δ)v(a∗,p)+δVq∗(p,w(p)) for all p∈Q1. The statement is true for all p≤q∗by definition since Vq∗(p,w)=Vq1(p,w)for all w. Assume that p>q ∗. From Lemma 4,thereexistsqsuch that ϕ(p,w(p)) ≥ ϕ(p,w(p)) for all p≥p≥q. Moreover, it follows from A.6.3 and A.6.4 that V(p, m(p)) ≥V(p,w)for all w,forallp≤q. Therefore, we must have that q∗≥q.Itfollows that ϕ(p,w(p)) <ϕ (q∗,w(q∗)) ≤q∗,hencew(p)≥mq∗(p). We therefore have that Vq∗(p,w(p)) =Vq1(p,w(p)). Since Vq1(p,m(p)) =(1−δ)v(a∗,p)+δVq1(p,w(p)) for all p∈Q1and Vq∗(p, m(p)) =Vq∗(p,mq∗(p)) =Vq1(p,mq∗(p)), it is enough to prove that Vq1(p,mq∗(p)) ≥ Vq1(p,m(p)). Clearly, there is nothing prove if mq∗(p)=m(p)for all p∈Q1,thatis,ifq∗=q1(remember that mq1(p)=m(p)for all p∈Q1). 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
954 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) So, assume that mq∗(p)>m (p)for some p∈(q∗,q1),hencemq∗(p)>m (p)for all p∈(q∗,q1). We now argue that if Vq1(p,w)>V q1(p,m(p)) for some w≥mq∗(p),then Vq1p,mp<1−p 1−pVq1(p,w), for all p>p. To see this, observe that w>m (p)and, accordingly, 1−p 1−pw+p−p 1−pm(1)−mp>0, since mis convex. Hence, 0<Vq1(p,w)−Vq1p,m(p) w−m(p) = 1−p 1−pVq1p,1−p 1−pw+p−p 1−pm(1)−Vq1p,1−p 1−pm(p)+p−p 1−pm(1) w−m(p) ≤ Vq1p,1−p 1−pw+p−p 1−pm(1)−Vq1p,mp 1−p 1−pw+p−p 1−pm(1)−mp, where the equality follows Observation A and the inequality from the concavity of Vq1in wfor each p. Since Vq1(p,w)=1−p 1−pVq1p,1−p 1−pw+p−p 1−pm(1), we have the desired result. Finally, from the definition of q∗,foralln>0, there exist pn∈(q∗,min(q∗+1 n,q1)] and wn≥m(pn)such that Vq1(pn,m(pn)) <V q1(pn,wn). From the concavity of Vq1in w for all p,Vq1(pn,m(pn)) <V q1(pn,mq∗(pn)) for all n. From the above argument, for all p,forallnsufficiently large, that is, such that pn< p,wehavethat Vq1p,m(p)<1−p 1−pn Vq1pn,mq∗(pn). Taking the limit as n→∞,weobtainthat Vq1p,m(p)<1−p 1−q∗Vq1q∗,mq∗q∗=Vq1p,mq∗(p), which completes the proof. Appendix B: Constructing the value function This section characterizes the value function Vqinduced by the policy τq. As explained in the text, it suffices to characterize Vq1since Vq(p,w)=Vq1(p,w)for all (p,w)∈W\W3 q 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 955 Figure 9. Construction of the thresholds. and Vq(p,w)=1−p 1−qVq1(q,m(q)) for all (p,w)∈W3 q. We first start with the definition of important subsets of [0, 1]. B.1 Construction of the sets Qk Let Q0:=[0, 1]. We define inductively the set Qk⊆[0, 1],k≥0. We write qk(resp., qk) for infQk(resp., supQk). For any k≥0, define the function Uk:[qk,1 ]→R: Uk(q):=1−q 1−qkmqk+q−qk 1−qkm(1), with the convention that Uk≡m(1)if qk=1. Note that U0(q)=M(q)and Uk(q)≥m(q) for all k. We define Qk+1as follows: Qk+1=q∈Qk:(1−δ)ua∗,q+δUk(q)≥m(q). For a graphical illustration, see Figure 9. Few observations are worth making. First, we have that P⊆Qkfor all k.Second, we have a decreasing sequence, that is, Qk+1⊆Qkfor all k.Third,ifQkand Pare nonempty, then they are closed intervals. Fourth, the limit Q∞=limk→∞ Qk=#kQk exists and includes P.Moreover,ifP= ∅,thenq∞=p,wherep:=infP.IfP=∅,then Q∞=∅. Consequently, there exists k∗<∞such that ∅=Qk∗+1⊂Qk∗= ∅. The first to the third observations are readily proved, so we concentrate on the proof of the fourth observation. The limit exists as we have a decreasing sequence of sets. We prove that if P=∅,thenQ∞=∅. So, assume that P=∅. We first argue that it cannot be that Qk=Qk−1= ∅ for some k≥0. To the contrary, assume that Qk=Qk−1= ∅for some k≥0, hence Qk=Qk−1for all k≥k. From the convexity and continuity of mand the linearity of u,Qk−1is the closed interval [qk−1,qk−1], with the two boundary points solution to (1−δ)ua∗,q+δUk−2(q)=m(q). Therefore, if (qk,qk)=(qk−1,qk−1),wehavethat mqk−1=(1−δ)ua∗,qk−1+δmqk−1, 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
956 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) mqk−1=(1−δ)ua∗,qk−1+δ1−qk−1 1−qk−1mqk−1+qk−1−qk−1 1−qk−1m(1), ≤(1−δ)ua∗,qk−1+δmqk−1. This implies that u(a∗,qk−1)=m(qk−1)and u(a∗,qk−1)=m(qk−1)and, therefore, ∅= Qk−1⊆P, a contradiction. We thus have an infinite sequence of strictly decreasing nonempty closed intervals. Let ε:=minp∈[0,1]m(p)−u(a∗,p). Since P=∅,wehavethatε>0. For all p∈Q∞,for all k, m(p)≤(1−δ)ua∗,p+δUk(p) ≤(1−δ)m(p)−ε+δUk(p). Assume that Q∞is nonempty and let q∞its greatest lower bound. Since q∞∈Qkfor all k,wehavethatUk(q∞)≥m(q∞)+ε(1−δ)/δ for all k. Since limk→∞ Uk(q∞)=m(q∞), we have that m(q∞)≥m(q∞)+ε(1−δ)/δ, a contradiction. We now prove that if P= ∅,thenq∞=p.Fromabove,wehavethatifQk=Qk−1= ∅ for some k≥0, hence Qk=Qk−1for all k≥k,thenP=Qksince P⊆Qk.Ifwehavean infinite sequence of strictly decreasing sets, for all q∈Q∞, (1−δ)ua∗,q+δ1−q 1−q∞mq∞+q−q∞ 1−q∞m(1)≥m(q). Taking the limit q↓q∞,weobtainthatu(a∗,q∞)=m(q∞),thatis,q∞∈P.Hence,q∞= p. B.1.1 Derivation of Vq1We first derive Vq1for all (p,w)∈W\W2 q1. To start with, Vq1(1, m(1)) =0 since a∗is not optimal at p=1. Similarly, Vq1(0, m(0)) =0ifa∗is not optimal at p=0, while Vq1(0, m(0)=v(a∗,0 )if a∗is optimal at p=0. Also, Vq1(q1,m(q1)) =(1−δ)v(a∗,q1)if q1>0; while Vq1(0, m(0)) =v(a∗,0 )if q1=0, since a∗is then optimal at p=0. With the function Vq1defined at these three points, it is then defined at all points (p,w)in W1 q1∪W4 q1. In particular, it is easy to show that Vq1q1,w=Mq1−w Mq1−mq1(1−δ)va∗,q1=Mq1−w Mq1−ua∗,q1va∗,q1, for all w∈[m(q1),M(q1)]. At all points (p,w)∈W3 q1, Vq1(p,w)=1−p 1−q1Vq1q1,mq1. Therefore, Vq1is well-defined at all (p,w)∈W\W2 q1. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 963 Since ϕηis decreasing in η,wehaveϕη≤ϕηwhen η>η, and hence ϕ(ϕη,w(ϕη)) ≤ ϕ(ϕη,w(ϕη)),asϕη≤ϕη≤p≤q. Similarly, since ϕη<p≤q,wehavethat ϕ(ϕη,w(ϕη)) ≤ϕ(p,w(p)) and, therefore, η−(1−δ)(1−λη) δ>0. We now return to the computation of the gradient. We have =(1−δ)va∗,p+δVq1p,w(p) −(1−δ)ληva∗,ϕη+δVq1p,w(p)+η−(1−δ)(1−λη) δm(1)−ua∗,1 ×η−1 =(1−δ) ηva∗,p−ληva∗,ϕη +δ ηVq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δm(1)−ua∗,1 =(1−δ) η(1−λη)va∗,1 +δ ηVq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δm(1)−ua∗,1 . (10) We further develop the above expression. To ease notation, we write (ϕ(p),λ(p)) for (ϕ(p,w(p)),λ(p,w(p))).Notethatϕ(p)∈(qk−1,qk], since p∈(qk,qk+1].As η−(1−δ)(1−λη) δ>0, we have that =(1−δ) η(1−λη)va∗,1 +δ ηVq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δm(1)−ua∗,1 =(1−δ) η(1−λη)va∗,1 +δ η η−(1−δ)(1−λη) δ × Vq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δm(1)−ua∗,1 η−(1−δ)(1−λη) δ =(1−δ) η(1−λη)va∗,1 +1−(1−δ)(1−λη) η ×λ(p)Vq1ϕ(p),mq1ϕ(p) −Vq1ϕ(p),mq1ϕ(p)+η−(1−δ)(1−λη) δλ(p)m(1)−ua∗,1 ×η−(1−δ)(1−λη) δ−1 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
964 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) =(1−δ) η(1−λη)va∗,1 +1−(1−δ)(1−λη) η ×Vq1ϕ(p),mq1ϕ(p) −Vq1ϕ(p),mq1ϕ(p)+η−(1−δ)(1−λη) δλ(p)m(1)−ua∗,1 ×η−(1−δ)(1−λη) δλ(p)−1 ≥(1−δ) η(1−λη)va∗,1 +1−(1−δ)(1−λη) ηva∗,1 =va∗,1 , where we use Observation A and the induction step. We now show that the gradient is increasing in η. To start with, note that η−(1−δ)(1−λη) δ is increasing in ηsince 1−λη ηis decreasing in η(see Lemma 5). For any η>η ,wehave the following: Vq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δmq1(1)−ua∗,1 η−(1−δ)(1−λη) δ =λ(p)Vq1ϕ(p),mq1ϕ(p) −λ(p)Vq1ϕ(p),mq1ϕ(p)+η−(1−δ)(1−λη) δλ(p)mq1(1)−ua∗,1 ×η−(1−δ)(1−λ) δ−1 =Vq1ϕ(p),mq1ϕ(p) −Vq1ϕ(p),mq1ϕ(p)+η−(1−δ)(1−λη) δλ(p)mq1(1)−ua∗,1 ×η−(1−δ)(1−λ) δλ(p)−1 ≥Vq1ϕ(p),mq1ϕ(p) −Vq1ϕ(p),mq1ϕ(p)+η−(1−δ)(1−λη) δλ(p)mq1(1)−ua∗,1 ×η−(1−δ)(1−λη) δλ(p)−1 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 965 = Vq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δmq1(1)−ua∗,1 η−(1−δ)(1−λη) δ , where the inequality follows from the fact that ϕ(p)∈(qk−1,qk]and, therefore, the gradient G(ϕ(p);η)being increasing in ηby the induction hypothesis. Finally, we have that 1 ηVq1p,mq1(p)−Vq1p,w(p;η) =(1−δ)(1−λη) ηva∗,1 +1−(1−δ)(1−λη) η × Vq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δm(1)−ua∗,1 η−(1−δ)(1−λη) δ ≥(1−δ)(1−λη) ηva∗,1 +1−(1−δ)(1−λη) η × Vq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δmq1(1)−ua∗,1 η−(1−δ)(1−λη) δ =(1−δ)(1−λη) ηva∗,1 +1−(1−δ)(1−λη) η × Vq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δmq1(1)−ua∗,1 η−(1−δ)(1−λη) δ +(1−δ)(1−λη) η−(1−δ)(1−λη) η ×Vq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δmq1(1)−ua∗,1 η−(1−δ)(1−λη) δ −va∗,1 ≥1 ηVq1p,mq1(p)−Vq1p,wp;η +(1−δ)(1−λη) η−(1−δ)(1−λη) η 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
966 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) ×Vq1ϕ(p),mq1ϕ(p) −Vq1ϕ(p),mq1ϕ(p)+η−(1−δ)(1−λη) δλ(p)mq1(1)−ua∗,1 ×η−(1−δ)(1−λη) δλ(p)−1 −va∗,1 ≥1 ηVq1p,mq1(p)−Vq1p,wp;η. The last inequality follows from the fact that the gradient in the second bracket is weakly larger than v(a∗,1 )by the induction hypothesis and the fact that 1−λη η<1−λη η (Lemma 5). Since limk→∞ qk=pwhen P= ∅, this completes the proof that the gradient is greater than v(a∗,1 )for all p∈[0, p]. Fact 2: For all p∈I2,G(p;η)is increasing in η.We first treat the case P= ∅. Recall that for all p∈(p,q∞], we have an explicit definition of the value function Vq1(p,mq1(p)) as va∗,p−mq1(p)−ua∗,p mq1(1)−ua∗,1 va∗,1 . Define ¯η(p)as the solution to ϕ¯η(p)=ϕ(p,w(p;¯η(p))) =p. Note that for any p∈ (p,q∞],foranyη≤¯η,ϕη∈[p,q∞]. Therefore, Vq1p,w(p;η)=ληVq1ϕη,mq1(ϕη)=ληva∗,ϕη−mq1(ϕη)−ua∗,ϕη mq1(1)−ua∗,1 va∗,1 =va∗,p−w(p;η)−ua∗,p mq1(1)−ua∗,1 va∗,1 . It follows that the gradient is equal to v(a∗,1 )for all p∈(p,p∗],forallη≤¯η. Consider now η> ¯η. We rewrite the gradient G(p;η)as follows: Vq1p,mq1(p)−Vq1p,w(p;η) η =Vq1p,mq1(p)−Vq1p,wp;η1(p) η +Vq1p,wp;η1(p)−Vq1p,w(p;η) η =η1(p) η Vq1p,mq1(p)−Vq1p,wp,η1(p) η1(p) +η−η1(p) η Vq1p,wp;η1(p)−Vq1p,w(p;η) η−η1(p) 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 967 =η1(p) ηva∗,1 +η−η1(p) η 1−p 1−pVq1p,mq1(p)−Vq1p,wp;η−η1(p) 1−p 1−p η−η1(p) =η1(p) ηva∗,1 +η−η1(p) ηGp;η−η1(p) 1−p 1−p. Since we have already shown that G(p;η)is increasing in ηand weakly larger than v(a∗,1 ), we have that the gradient G(p;η)is also weakly increasing in η(and greater than v(a∗,1 )). We now treat the case P=∅. Define ¯η(p)as the solution to ϕ¯η(p)=ϕ(p,w(p; ¯η(p))) =q. Note that for any p∈[q,q],foranyη≤¯η,ϕη∈[q,q]. Therefore, for all η≤¯η, η=(1−δ)(1−λη)since the ratio mq1(1)−w(ϕη) 1−ϕηis constant in ηand so is ϕ(ϕη,w(ϕη)). (Recall that we vary ηat a fixed p.) It follows then from equation (10)that G(p;η)=(1−δ) η(1−λη)va∗,1 +δ ηVq1p,w(p)−Vq1p,w(p)+η−(1−δ)(1−λη) δm(1)−ua∗,1 =(1−δ) η(1−λη)va∗,1 =va∗,1 . We have that the gradient G(p;η)is equal to v(a∗,1 )for all p∈(q,q],forallη≤¯η.Finally, when η> ¯η, the same decomposition as in the case P= ∅ completes the proof. Fact 3: For all p∈I3,thegradientG(p;η)is increasing in η.We only treat the case P= ∅.(ThecaseP=∅is treated analogously.) Define ¯η(p)as the solution to ϕ¯η(p)= ϕ(p,w(p;¯η(p))) =q∞. By construction, for all p∈(q∞,1 ],forallη≤¯η(p),wehavethat ϕη∈(q∞,1 ]. Therefore, ϕη> q. Choose ¯η(p)≤η≤η.Wehavethatϕη≥ϕη≥qsince q∞≥qand, therefore, ϕp,w(p)+η−(1−δ)(1−λη) δmq1(1)−ua∗,1 =ϕϕη,w(ϕη) ≥ϕ(ϕη,w(ϕη)=ϕp,w(p)+η−(1−δ)(1−λη) δmq1(1)−ua∗,1 . Also, since q≤ϕη≤p,wehavethatϕ(ϕη,w(ϕη)) ≥ϕ(p,w(p)) and, therefore, η−(1−δ)(1−λη) δ≤0.Thesameappliestoη. Finally, as already shown, η−(1−δ)(1−λη) δ<η−(1−δ)(1−λη) δ. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
968 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) To ease notation, define (˜ λη,˜ϕη)as follows: ⎧ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎩ ˜ λη=λp,w(p)−(1−δ)(1−λη)−η δm(1)−ua∗,1 ˜ϕη=ϕp,w(p)−(1−δ)(1−λη)−η δm(1)−ua∗,1 (11) Notice that ˜ϕη=ϕ(ϕη,w(ϕη)) ∈I1since ϕη> q∞. The rest of the proof is purely algebraic and mirrors the case p∈I1.First,wehave the following: Vq1p,w(p)−Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 (1−δ)(1−λη)−η δ =˜ ληVq1˜ϕη,mq1(˜ϕη)+(1−δ)(1−λη)−η δ˜ ληmq1(1)−ua∗,1 −˜ ληVq1˜ϕη,mq1(˜ϕη)(1−δ)(1−λη)−η δ−1 = Vq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη−Vq1˜ϕη,mq1(˜ϕη) (1−δ)(1−λη)−η δ˜ λη , where we again use Observation A. Similarly, we have Vq1p,w(p)−Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 (1−δ)(1−λη)−η δ =˜ ληVq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη −˜ ληVq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη −(1−δ)(1−λη)−η δ˜ λη ×(1−δ)(1−λη)−η δ−1 =Vq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη −Vq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη −(1−δ)(1−λη)−η δ˜ λη 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 969 ×(1−δ)(1−λη)−η δ˜ λη−1 , where again we use Observation A and the fact (1−δ)(1−λη)−η δ˜ λη >(1−δ)(1−λη)−η δ˜ λη . Since ˜ϕη∈I1,wehavethat Vq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη −Vq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη −(1−δ)(1−λη)−η δ˜ λη ×(1−δ)(1−λη)−η δ˜ λη−1 ≤ Vq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη−Vq1˜ϕη,mq1(˜ϕη) (1−δ)(1−λη)−η δ˜ λη , where the inequality follows from our previous argument on the interval I1. It follows that Vq1p,w(p)−Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 (1−δ)(1−λη)−η δ ≤ Vq1p,w(p)−Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 (1−δ)(1−λη)−η δ . From equation (10), we then have that 1 ηVq1p,mq1(p)−Vq1p,w(p;η) =(1−δ)(1−λη) ηva∗,1 +(1−δ)(1−λη) η−1 × Vq1p,w(p)−Vq1p,w(p)−(1−δ)(1−λη)−η δm(1)−ua∗,1 (1−δ)(1−λη)−η δ 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
970 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) ≥(1−δ)(1−λη) ηva∗,1 +(1−δ)(1−λη) η−1 × Vq1p,w(p)−Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 (1−δ)(1−λη)−η δ =(1−δ)(1−λη) ηva∗,1 +1−(1−δ)(1−λη) η × Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 −Vq1p,w(p) (1−δ)(1−λη)−η δ =(1−δ)(1−λη) ηva∗,1 +1−(1−δ)(1−λη) η × Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 −Vq1p,w(p) (1−δ)(1−λη)−η δ +(1−δ)(1−λη) η−(1−δ)(1−λη) η ×%Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 −Vq1p,w(p) (1−δ)(1−λη)−η δ −va∗,1 ≥1 ηVq1p,mq1(p)−Vq1p,wp;η, where the last inequality follows from Vq1p,w(p)−(1−δ)(1−λη)−η δmq1(1)−ua∗,1 −Vq1p,w(p) (1−δ)(1−λη)−η δ = ˜ ληVq1˜ϕη,mq1(˜ϕη)−˜ ληVq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη (1−δ)(1−λη)−η δ ≥va∗,1 . 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 19 (2024) Contracting over persistent information 971 We now show that the the gradient G(p;η)is smaller than v(a∗,1 )for any η≤¯η(p). From equation (10), we have that 1 ηVq1p,mq1(p)−Vq1p,w(p;η) =(1−δ)(1−λη) ηva∗,1 −(1−δ)(1−λη) η−1 × Vq1p,w(p)−(1−δ)(1−λη)−η δm(1)−ua∗,1 −Vq1p,w(p) (1−δ)(1−λη)−η δ =va∗,1 −(1−δ)(1−λη) η−1 ×%Vq1p,w(p)−(1−δ)(1−λη)−η δm(1)−ua∗,1 −Vq1p,w(p) (1−δ)(1−λη)−η δ −va∗,1 =va∗,1 −(1−δ)(1−λη) η−1 ×%˜ ληVq1˜ϕη,mq1(˜ϕη)−˜ ληVq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη (1−δ)(1−λη)−η δ −va∗,1 & =va∗,1 −(1−δ)(1−λη) η−1 ≥0 ×%Vq1˜ϕη,mq1(˜ϕη)−Vq1˜ϕη,w˜ϕη;(1−δ)(1−λη)−η δ˜ λη (1−δ)(1−λη)−η δ˜ λη −va∗,1 & ≥0 ≤va∗,1 , where the inequality follows from the fact that ˜ϕη≤p(therefore, from our arguments on the interval I1, where we show that the gradient is larger than v(a∗,1 )). Finally, we can use a similar decomposition as in the case p∈I2to prove that the gradient is increasing for all η. 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
972 Zhao, Mezzetti, Renou, and Tomala Theoretical Economics 19 (2024) Appendix C C.1 Recursive formulation: A proof Ely (2015) proves that the principal’s maximal payoff is maxw∈[m(p0),M(p0)] V∗(p0,w), with V∗the unique fixed point of the contraction T, with the operator Tdiffering from the operator Tin that the promise-keeping constraint is written as an equality in all maximization problems T(V)(p,w); all other constraints are the same. Note that, like T, the operator Tis monotone. For any (p,w)∈W,let" V∗(p,w):=max w∈[w,M(p)] V∗(p, w)and " w∗(p,w)amaximizer. (If there are multiple maximizers, choose an arbitrary one.) We prove that V∗=" V∗.Todoso,weprovethatT(" V∗)=" V∗. Since Tis a contraction, hence has a unique fixed point, it follows that V∗=" V∗. (Note that we are not arguing that T= T.) We start with two simple observations: (i) T(V)(p,w)≥ T(V)(p,w)for all (p,w)∈ W,forallV, and (ii) " V∗(p,w)≥ V∗(p,w)for all (p,w)∈W. The first observation follows from the fact the promised-keeping constraint is an equality in T(V)(p,w), while it is an inequality in T(V)(p,w). The second observation follows immediately from the definition of " V∗. We now prove that T(˜ V∗)≥˜ V∗.Forall(p,w)∈W,wehave " V∗(p,w)= V∗p," w∗(p,w)= T V∗p," w∗(p,w), ≤T V∗p," w∗(p,w), ≤T V∗(p,w), ≤T" V∗(p,w), where the first line follows from the definitions of " V∗, V∗,and T,andthefactthat V∗= T( V∗); the second line from observation (i); the third line from the fact that " w∗(p,w)≥ w, so that all feasible solutions to T( V∗)(p," w∗(p,w)) are also feasible for T( V∗)(p,w); and the fourth line from observation (ii) and the definition of T(V)(p,w),V= V∗," V∗. We next prove that T(" V∗)≤" V∗. By contradiction, suppose that there exists (p,w)∈ Wand a feasible policy (λs,ps,as,ws)s∈Ssuch that " V∗(p,w)< s∈S λs(1−δ)v(as,ps)+δ" V∗(ps,ws). Moreover, we have that s∈S λs(1−δ)v(as,ps)+δ" V∗(ps,ws)= s∈S λs(1−δ)v(as,ps)+δ V∗ps," w∗(ps,ws) ≤ V∗p, s∈S(1−δ)u(as,ps)+δ" w∗(ps,ws) ≤" V∗(p,w), 15557561, 2024, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE5056 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License