Dynamic delegation with a persistent state
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Chen, Yi Article Dynamic delegation with a persistent state Theoretical Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Chen, Yi (2022) : Dynamic delegation with a persistent state, Theoretical Economics, ISSN 1555-7561, The Econometric Society, New Haven, CT, Vol. 17, Iss. 4, pp. 1589-1618, https://doi.org/10.3982/TE4710 This Version is available at: https://hdl.handle.net/10419/296394 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Theoretical Economics 17 (2022), 1589–1618 1555-7561/20221589 Dynamic delegation with a persistent state YiChen Johnson College of Busines, Cornell University In this paper, I study the dynamic delegation problem in a principal–agent model wherein an agent privately observes a persistently evolving state, and the principal commits to actions based on the agent’s reported state. There are no transfers. While the agent has state-independent preferences, the principal wants to match a state-dependent target. I solve the optimal delegation in closed form, which sometimes prescribes actions that move in the opposite direction of the target. I provide a simple necessary and sufficient condition for that to occur. Generically, the principal fares strictly better in the optimal delegation than in the babbling outcome. Over time, the principal is worse off in expectation, but the agent is better or worse off depending on the shape of the principal’s state-dependent target. Keywords. Communication, dynamic delegation, contrarian, quota mechanism, Brownian motion. JEL classification. D82, D83, D86. 1. Introduction In many organizations, decisions are not made by the person who actually holds and understands the most relevant information. Moreover, the transmission of this information is often hindered by conflicts of interest. The informed party may have a selfinterested motive to mislead the uninformed party, who then takes action based on information that may have been miscommunicated. A firm’s headquarters, for example, will allocate resources to a division manager over time. The headquarters metes out resources in order to hit a target amount of allocation that depends on some state of the project, say, profitability, consumer taste, or technical Yi Chen: [email protected] I am indebted to Dirk Bergemann, Johannes Hörner, and Larry Samuelson for their constant academic support. I am grateful to Marco Battaglini, Yeon-Koo Che, Chen Cheng, Eduardo Faingold, Simone Galperti, Yuan Gao, Marina Halac, Ryota Iijima, Justin Johnson, Navin Kartik, Nicolas Lambert, Fei Li, Elliot Lipnowski, Chiara Margaria, Dmitry Orlov, Gregory Pavlov, Ennio Stacchetti, Philipp Strack, Juuso Välimäki, Zhe Wang, Yiqing Xing, Xiye Yang, and John Zhu for extended discussions. I also thank the conference and seminar participants at Penn State University, Cornell University, Johns Hopkins University, University of Southern California, University of Arizona, Western University, Peking University HSBC Business School, Southern University of Science and Technology, the 2018 International Conference on Game Theory at Stony Brook, the 2018 Southern Economic Association Annual Conference, the 2019 North American Summer Meeting of the Econometric Society, the 2019 Osaka Workshop on Economics of Institutions and Organizations, and the 2020 North American Winter Meeting of the Econometric Society for helpful comments and suggestions. All errors are mine. ©2022 The Author. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at https://econtheory.org.https://doi.org/10.3982/TE4710
1590 Yi Chen Theoretical Economics 17 (2022) parameters, which develops slowly over time. Only the division manager directly observes this state, but he wants to receive more resources regardless of the state. Worse still, the headquarters receives profits only after a long lag, making it difficult to detect misrepresentations in the manager’s reports. These severe conflicts of interest beg the question, “Does the headquarters benefit at all from the manager’s information with a dynamic contract?” If so, how should it optimally act on the manager’s reports to utilize information? To answer these questions, I use a dynamic principal–agent model to investigate how the agent’s persistent but changing private information can be best elicited and put to use. The agent privately observes a state, which evolves as a Brownian motion with a drift. The agent continuously reports the state to the principal, but has the ability to inflate or shade the report at any time. The principal observes nothing but the agent’s reports and so commits to a dynamic contract specifying actions over time based on the reported history. The principal cares about the state because she incurs a flow cost that is quadratic in the gap between the action and her target, i.e., her state-dependent favorite action. The agent, by contrast, has a transparent motive independent of the state: the higher the action, the better. Communication is ineffective in a static contracting environment where the agent has severely misaligned, state-independent preferences. Indeed, to neutralize the misreporting motive of the agent, the contract would have to assign the same expected action for all reported states. When it comes to long-term relationships with an evolving state, then the prospect of communication is better. In order to elicit truthful reports, the principal only needs to ensure that the agent’s continuation payoff is independent of his current report. In other words, the principal commits to a fixed quota (Jackson and Sonnenschein (2007)), which is the discounted sum of actions. This leaves the principal with the optimization problem of intertemporally reallocating the quota to make the best use of the agent’s information. Using a recursive method, I solve the optimal contract in closed form. While tractability is usually difficult to obtain in a dynamic setting with persistent information, I address this obstacle by reducing the dimensionality of the recursive problem and transforming the nonlinear partial differential equation into two ordinary differential equations. The closed-form solution allows for comprehensive analysis of the optimal contract. First, the optimal contract prescribes how the action responds to information at any time. If a contract stipulates an action moving in the same direction as the target, it is said to exhibit a conformist pattern; if instead the action moves in the opposite direction from the target, it exhibits a contrarian pattern. I find a simple necessary and sufficient condition for either pattern to occur, which involves the third derivative of the target function. To understand the intuition, suppose the state process has no drift and the target function is increasing. If a positive shock to the state boosts the current target, the principal will be inclined to raise the current level of action. Meanwhile, due to the persistence of the state, future targets are also expected to increase, tempting the principal to take higher actions in the future. Unfortunately, the agent’s incentive constraints cannot allow both. Any increase in the current action necessitates lower future actions
Theoretical Economics 17 (2022) Dynamic delegation 1591 or vice versa. Whether a conformist or contrarian pattern emerges hinges on the tradeoff between hitting the current target on the one hand and hitting future targets on the other. If the former goal dominates, a conformist pattern emerges; otherwise, we should observe a contrarian pattern. When the target function has a positive third derivative, the expected future target is more sensitive to state shocks than the current target is. Then, in the intertemporal trade-off, the principal optimally sacrifices the goal of hitting the current target in exchange for a better chance of hitting future targets, leading to a contrarian pattern. Second, I show that communication generically improves upon babbling in the optimal dynamic contract. Communication is effective as long as the action is responsive to the reported state, whether in a conformist or a contrarian pattern. When the drift of the state is zero, the knife-edge exceptions occur when the target function is linear or quadratic. In these cases, a shock to the current state justifies an increase of current action as much as it demands a rise in the expected future actions; therefore, the principal’s optimal choice is to stay put despite her degree of freedom in responding to information. The result implies that the curvature of the target function is not sufficient to guarantee effective communication; instead, a nonzero third derivative is required. Third, the model delivers predictions on the welfare of the two parties as the state unfolds. The principal expects an ever-increasing cost over time, indicating an inevitable worsening of the match between the target and the action. This cost–backloading result holds even for patient players: the incentive constraints cause distortions to accumulate without bound, which weighs heavily on the principal precisely because of her patience. The agent, on the other hand, is not necessarily immiserated. The trend of his continuation payoff depends on the shape of the target function as well as on the state process. I also explore three extensions of the main model. The first enriches the state process by allowing for mean reversion and, accordingly, weaker persistence of the state. Mean reversion undermines the responsiveness of the expected future target to the current state, because any shock to the current state decays over time. Consequently, a contrarian pattern is less likely to emerge as the principal finds it less appealing to sacrifice hitting the current target for the sake of hitting future ones. The second extension examines the optimal contract for a finite horizon. Even adding calendar time as an additional state variable, the problem is still solvable. While the basic insights remain, the finite horizon introduces a deadline effect. In the beginning, when the deadline is far in the future, the contract behaves similarly to the main model. As time goes on, the action becomes less and less responsive to information, and the agent’s influence gradually vanishes. The third and final extension explores the implication of having a less patient agent relative to the principal. The difference in discount rates generally results in payoff front-loading for the agent, which is more likely to cause immiseration. Moreover, hitting future targets becomes less costly for the principal; thus, the intertemporal trade-off is more inclined to favor future targets and the contract appears more contrarian.
1592 Yi Chen Theoretical Economics 17 (2022) 1.1 A two-period example To illustrate the key trade-offs in a dynamic contracting problem, I present a two-period example. A state θt∈Rfollows a random walk. In period t=1, θ1is drawn from N(0, 1). In period t=2, θ2evolves from θ1with noise, θ2=θ1+ε, where ε∼N(0, 1)is independent of θ1. In each period t, the agent privately learns θt and reports ˆ θtto the principal, who takes action xt∈R. Monetary transfers are not available. The principal’s total cost is (x1−f(θ1))2+(x2−f(θ2))2,wheref(·)is a time-invariant target function. The agent’s total payoff is x1+x2. A contract is a pair (x1(ˆ θ1),x2(ˆ θ1,ˆ θ2)), mapping report histories into actions. Based on the revelation principle, I focus on truthful contracts. The principal solves min x1(·),x2(·,·) Ex1−f(θ1)2+x2−f(θ2)2 subject to x1(θ1)+Ex2(θ1,θ2)|θ1≥x1(ˆ θ1)+x2(ˆ θ1,ˆ θ2)∀θ1,ˆ θ1,ˆ θ2(1) x2(θ1,θ2)≥x2(θ1,ˆ θ2)∀θ1,θ2,ˆ θ2.(2) Condition (2) requires that truth-telling in period 2 is optimal for the agent after a truthful report in period 1. Condition (1) governs the truth-telling incentive in period 1 for the agent, who weighs all possible reporting strategies. Since the agent’s payoff is state-independent, condition (2) implies that x2(θ1,θ2) does not depend on θ2. Writing x2(θ1,θ2)=x2(θ1)for short, condition (1) is simplified to x1(θ1)+x2(θ1)≥x1(ˆ θ1)+x2(ˆ θ1)for all θ1and ˆ θ1. To satisfy this condition, we must have x1(θ1)+x2(θ1)≡W,(3) where Wis a constant, interpreted as the quota (total payoff) promised to the agent. This quota, as well as how x1and x2jointly respond to θ1, are optimally chosen by the principal to minimize cost. With the simplified incentive constraints, the optimal two-period contract is obtainable for general f(·)(see Appendix A.1). Specifically, for a linear target f(θ)=θ,wehave x1(θ1)=x2(θ1)=0, i.e., the outcome is “babbling” as the actions do not reflect information about the state. A quadratic target f(θ)=θ2does not lead to better utilization of information, as x1(θ1)=1andx2(θ1)=2forallθ1. An interesting case arises when f(θ)= eθ,forwhichx1(θ1)=1 2(e+√e)−1 2(√e−1)eθ1and x2(θ1)=1 2(e+√e)+1 2(√e−1)eθ1. As the first-period target eθ1increases, the corresponding action x1decreases, in order for x2to increase in the next period. Why does this pattern emerge? The answer lies in the shape of the target function. At an arbitrary state θ1, suppose the optimal contract specifies actions x1and x2over the two periods. Given the principal’s quadratic cost function, a marginal increase in
Theoretical Economics 17 (2022) Dynamic delegation 1593 x1brings a marginal benefit of 2(f(θ1)−x1)to the principal. Meanwhile, the quota mechanism forces x2to decrease, which imposes a marginal cost of 2(E[f(θ2)|θ1]−x2). Optimality requires them to cancel out. Now, a higher θ1raises the marginal benefit by 2f(θ1), but also raises the marginal cost by 2E[f(θ2)|θ1].1When the slope of the target function is convex (e.g., f(θ)=eθ), the latter effect dominates because E[f(θ2)|θ1]> f(E[θ2|θ1]) =f(θ1). To restore optimality, the principal should shift some quota from x1to x2despite the increased current target f(θ1). In Section 2, I lay out the formal model in continuous time with an infinite horizon. Continuous time allows for the gradual arrival of information and thus closedform analysis. An infinite horizon avoids the deadline effect, keeps the stationarity of the problem, and enables the study of the asymptotics. 1.2 Related literature This paper contributes to a closely related literature on allocation problems without monetary transfer. In a static setting, Jackson and Sonnenschein (2007)study the decision rule facing many replicas of the same allocation problem and propose a “quota mechanism” that links all allocations together, where efficiency is asymptotically achieved as the number of replicas increases. In a dynamic setting, repeated allocation games (e.g., Renault, Solan, and Vieille (2013), Margaria and Smolin (2018), Lipnowski and Ramos (2020)) feature an informed sender and an uninformed receiver, where the sender observes independent and identically distributed (i.i.d.) or persistent information and reports to the receiver for decision-making. When the receiver is assumed to have intertemporal commitment power, Frankel (2016)andGuo and Hörner (2020), among others, analyze dynamic allocation problems. In these models, dynamic versions of the quota mechanism arise in equilibrium strategy or optimal contract, where the quota is cashed out in a conformist pattern. My model also exhibits a dynamic quota for the agent, but the combination of a persistent state process and a general target function allows for intertemporal trade-offs that lead to potentially contrarian patterns. Such a counterintuitive pattern cannot arise in a static, multidimensional setting or in a dynamic model with finite state Markov chain and linear preferences. This paper also builds on the literature on communication. Since Crawford and Sobel (1982)andGreen and Stokey (2007) pioneered the field of sender–receiver games, a large body of scholarship has been produced (see Sobel (2013) for a comprehensive summary). Meanwhile, communication with a committed receiver has inspired the literature on delegation (see Holmstrom (1977), Melumad and Shibano (1991), Alonso and Matouschek (2008), Amador and Bagwell (2013)). This paper features a committed receiver and a privately informed sender, with evolving information and without transfers. It therefore expands on the models of dynamic delegation. More broadly, other related works tackle various allocation problems with similar, but different, settings.2Bird and Frug (2019) study a dynamic contracting problem without transfer and implement the unique optimal contract by a deadline to earn rewards. 1It holds that dE[f(θ2)|θ1]/dθ1=E[f(θ2)|θ1]due to the random walk assumption. 2Models of multiple competing agents (e.g., Ben-Porath, Dekel, and Lipman (2014) and de Clippel, Eliaz, Fershtman, and Rozen (2021)) find “strategic favoritism” as the optimal mechanism. There is also a larger
1594 Yi Chen Theoretical Economics 17 (2022) Boleslavsky and Lewis (2016)andMalenko (2019) study dynamic mechanisms of influence with either costly verification or noisy observation of the state. The remainder of the paper is organized as follows. Section 2lays out the setting for the continuous-time model. Section 3reduces the agent’s incentive constraint to a necessary condition and a stronger, sufficient condition. Section 4solves and analyzes the optimal contract. Section 5discusses three extensions of the main model, and Section 6 concludes. 2. The model There is a principal (she) and an agent (he). Time t≥0 is continuous. A state θevolves over time, but is observable only to the agent. The agent continuously makes potentially manipulated report ˆ θof the true state to the principal, who commits to action x∈Rat all times based on the history of reports. The state process θstarts at zero and evolves according to θt=μt +Zt, where Zis the standard Brownian motion on the probability space (,F,P).Theconstant μis the drift of the process. The volatility is constant and normalized to 1. The law of motion of θis common knowledge. The state process is highly persistent as it features independent increments. In Section 5.1, I introduce mean reversion into the process as a less persistent counterpart. Interests are misaligned. While the principal’s favorite action is state-dependent, the agent only wishes to induce actions as high as possible. Specifically, given a state–action pair (θ,x), the principal suffers a quadratic flow cost (x−f(θ))2from the gap between the action xand a state-dependent target f(θ). For ease of analysis, the function fis assumed to be piecewise C2. The agent’s flow payoff is simply x, independent of the state. Intertemporally, the players share the same discount rate r>0. A strategy mof the agent is a θ-measurable process, such that his reported process ˆ θ follows dˆ θt=mtdt+dθt, where mt∈Rrepresents the “intensity of misreporting” at instant t. The space of feasible strategies is set to M≡m:Ee2αt 0msds<∞∀t,and lim t→∞e−rtEe2αt 0msds=0 to exclude Ponzi-type strategies, where α≡1 22r+μ2−|μ|(4) is a positive constant. I show in Appendix A.11 that this restriction is not essential. literature on dynamic contract with transfers (i.e., Fernandes and Phelan (2000), Battaglini (2005), Sannikov (2008), Williams (2011), Kapiˇ cka (2013), DeMarzo and Sannikov (2016)), but the lack of transfers in my setting aggravates the agency problem because the continuation payoff of the agent can only be promised by a sequence of future allocations.
Theoretical Economics 17 (2022) Dynamic delegation 1595 In the beginning, the principal commits to a contract x. It is a process adapted to the information generated by ˆ θ, specifying, at any time t,anactionxt∈Ras a function of the history of reports ˆ θtup to time t. There are no monetary transfers. Given a contract–strategy pair (x,m), the total expected cost and payoff are, respectively, UP(x,m)=Em∞ 0 re−rtxt−f(θt)2dt UA(x,m)=Em∞ 0 re−rtxtdt, where Emdenotes the expectation induced by strategy m. Hereafter, “payoff” and “cost” refer to the agent’s total expected payoff and the principal’s total expected cost, unless otherwise noted. The following regularity condition ensures finiteness of the above cost and payoff. Assumption 1 (Regularity). There exists α0>0and α1∈[0, α)such that |f(θ)|≤ α0eα1|θ|. Intuitively, the condition prevents the target function from growing too exponentially in both directions. The constant αis defined in (4). This condition is not too restrictive, as it allows for all piecewise polynomials and piecewise continuous bounded functions, among others. The agent chooses a strategy mto maximize his payoff from a given contract x.The principal designs a contract xto minimize her cost given the agent’s strategy in reaction to the contract. With the usual revelation principle argument (see Appendix A.2 for a formal proof), it is without loss of generality to focus on truthful contracts, i.e., those that make truth-telling (m≡0) optimal for the agent among all strategies. Moreover, I focus on deterministic mechanisms, but, as is verified later, such a restriction is without loss of generality. Therefore, the principal solves min (xt(·))t≥0 E∞ 0 re−rtxtθt−f(θt)2dt(5) subject to E∞ 0 re−rtxtθtdt≥E∞ 0 re−rtxtˆ θtdt, where ˆ θt≡θt+t 0 msds∀m∈M.(6) The incentive constraint (6) stipulates that truth-telling leads to the highest payoff. While the constraint is expressed as of time zero, it also implies incentive compatibility at all later times, since the agent faces a decision problem with time-consistent preferences. This constraint implicitly assumes that the payoff of the agent is well defined on and off equilibrium, but such an assumption is innocuous because if a contract generates non-integrable payoffs for the agent, it must bring infinite cost to the principal, which is clearly suboptimal.
1596 Yi Chen Theoretical Economics 17 (2022) I will end this section with a few comments on the model. First, the quadratic-cost assumption is not essential for the qualitative results, but it greatly simplifies analysis because the cost-minimizing action facing an uncertain target is the mean of the target. The separability result (Lemma 2) directly benefits from this assumption. Second, one can alternatively define ft≡f(θt)as the state. Here I choose to keep the process simple, while summarizing everything in f. Third, the insatiable preferences of the agent can be interpreted in Crawford and Sobel (1982) as taking the bias to infinity and, therefore, represent severely misaligned interests. 3. Incentives of the agent This section reduces the incentive compatibility of the agent into a tractable form, in preparation for deriving the optimal contract. Specifically, a necessary condition for incentive compatibility is obtained in Section 3.1 from the first-order approach. It is then augmented to a sufficient condition in Section 3.2, which is later invoked to verify the optimality of the candidate solution in Section 4. 3.1 Incentive compatibility: Necessary condition The dynamic first-order approach (Williams (2011), Kapiˇ cka (2013), Pavan, Segal, and Toikka (2014), DeMarzo and Sannikov (2016)) derives a local version of the incentive constraints. To apply this method, I define a process W=(Wt)t≥0for any contract x, Wt≡Et∞ t re−r(s−t)xsds, as the agent’s on-path expected continuation payoff. The expectation Etis conditional on the information generated by the state process up to time t. To use the recursive method, I need a few state variables to summarize the history. The current state θtnaturally serves as a state variable. The continuation payoff Wtis commonly used as another state variable in the literature (see Abreu, Pearce, and Stacchetti (1986)andThomas and Worrall (1990), among others). Furthermore, due to the persistent private information, a third state variable, called the continuation marginal payoff, is usually required (Fernandes and Phelan (2000), Williams (2011), Kapiˇ cka (2013), Guo and Hörner (2020)). However, in this paper I can drop the third state variable despite the persistence of information. This is because the agent’s payoff is independent of the state, and, hence, his flow and continuation payoffs, both on and off path, are common knowledge. Even if the agent used strategy m=0andhisprivatebe- lief about the state diverged from the principal’s, the continuation payoff would evolve as if ˆ θt=θt+t 0msdswas the realization of the true state and the agent had reported truthfully. Given any contract x, the implied evolution of Wcanbewrittenasadiffusionprocess according to Lemma 1below.
Theoretical Economics 17 (2022) Dynamic delegation 1603 Figure 2. Exponential target: r=1, μ=0, and f(θ)=−e−0.7θ. The key property of an exponential target fis that f and falways have the same sign, so that γf is an amplified version of f. Whenever the current target increases, its future counterpart increases by even more. Example 4 (Kinked). Consider a kinked target function f(θ)=b0θif θ<0andf(θ)= b1θif θ≥0, where b0,b1>0andb0=b1. Then the expected future target is γf(θ)= f(θ)+e−√2r|θ|(b1−b0)/(2√2r). These functions are shown in Figure 3(a) for b1=1 3< 1=b0. Since (γf)−f=1 2(b1−b0)e−√2r|θ|for θ>0and(γf)−f=1 2(b0−b1)e−√2r|θ| for θ<0, according to Theorem 2(ii), the action is contrarian at all states where fis the smaller between b0and b1, and is conformist otherwise. Figure 3(b) plots the simulated paths for the target f(θt)and the action xt, where the two paths co-move whenever the target f(θ)is below its kink (i.e., on its steeper segment) and move out of phase otherwise. ♦ Figure 3. Kinked target: r=1, μ=0, and f(θ)=θ,ifθ<0 1 3θ,ifθ≥0.
1604 Yi Chen Theoretical Economics 17 (2022) Figure 4. Binary target: r=1, μ=0, and f(θ)=1{θ≥0.4}. Intuitively, the target’s slope ftakes only two values, and, therefore, the expected future target γf must have a slope in between. On the flatter (steeper) segment of the target function, the current target is less (more) responsive to shocks than its future counterpart. This explains the coexisting patterns in the same contract. Notably, both a kinked target and an exponential target can be increasing and concave, but the optimal contracts are qualitatively different. Therefore, concavity or convexity of the target function alone is insufficient to determine the pattern of the contract. Example 5 (Binary). We revisit Example 2for the binary target. The expected future target can be rewritten as γf (θ)=f(θ)−1 2e−√2r|θ−θ|sgn(θ−θ).Forθ=0.4, these functions are plotted in Figure 4(a). Since fis either zero or undefined, Definition 1 does not apply. For θ= θ,(γf)−f=√r/2e−√2r|θ−θ|>0, so that the action always moves in the opposite direction of the state. At θ=θ,(13) implies that the action jumps along with the target, which is “conformist” in a broader sense. Figure 4(b) simulates the time paths for the target and the action. Every time the target jumps, the optimal action follows suit. At other times, the action still responds to the state despite the constant target. ♦ This result obtains from the nature of the expected future target. The expected future target γf is everywhere increasing because it factors in the upward jump at θ=θ. A marginal increase in the state, although not affecting the target, increases the expected future target and hence demands a shift of resources from present to future. Below, Theorem 3captures the knife-edge cases where the contract becomes babbling, in which the principal optimally stays unresponsive to the agent’s reports. Such communication failures arise only for a nongeneric set of target functions. Otherwise, the principal always finds a direction to intertemporally reallocate actions to reduce cost. Theorem 3 (Impossibility). (i) For μ=0, the contract is babbling if and only if the target is almost everywhere identical to c0+c1θ+c2θ2for some constants c0,c1,c2.
Theoretical Economics 17 (2022) Dynamic delegation 1605 (ii) For μ= 0, the contract is babbling if and only if the target is almost everywhere identical to c0+c1θ+c2e−2μθ for some constants c0,c1,c2. According to the theorem, the curvature of the target function is not sufficient to guarantee the gains from information transmission. For example, when μ=0, effective information transmission requires a nonzero curvature of the information sensitivity or, equivalently, a nonzero third derivative of the target function. This is why quadratic target functions lead to babbling in part (i) of the theorem. Although linear and quadratic target functions do not admit effective communication when the state is a Brownian motion without drift, this is specific to the state process. When the state follows some other process, say the Ornstein–Uhlenbeck process (see Section 5.1), information is utilized even with these target functions. 4.3 Evolution of the contract Next, I discuss the stochastic evolution of the optimal contract on path, characterizing the dynamics of cost and payoff as the contract is executed over time. Proposition 3 (Cost and PayoffDynamics). (i) The principal’s continuation cost is a submartingale, i.e., Et[dC(θt,Wt)]/dt≥0. (ii) The agent’s continuation payoff monotonically increases (resp. decreases) over time if γf(θ)−f(θ)>0(resp. <0) for all θ. Part (i) of Proposition 3claims that the principal faces statistically growing continuation costs Ct≡C(θt,Wt)as the contract is executed over time. Intuitively, this is because the current state is realized while future states can only be predicted. At the beginning, the minimized cost is C(0, γf(0)). If, at a later time t>0, the state becomes zero again, then the continuation payoff Wtgoverned by the incentive constraints will have almost surely wandered away from γf(0), and the continuation cost will have increased. The back-loading of the principal’s cost is typical in dynamic mechanism design without transferable utilities (e.g., Guo and Hörner (2020)). Part (ii) implies that the agent does not necessarily end up immiserated; instead, the trajectory of his payoffs depends on the shape of the target function. When the target function is convex, the expected future target γf is always larger than the current target f. Therefore, the agent’s continuation payoff increases over time because the principal wants the action path to accommodate such an overall trend. When the target function is concave, the opposite is true and the agent is immiserated. In sum, the front-loading or back-loading of the agent’s payoff is driven by the expected evolution of the target function. 5. Extensions This section extends the main model in three directions. First, I consider a less persistent state process by introducing mean reversion. Second, I consider a finite time horizon to explore the nonstationary behavior of the contract. Finally, I study the effect of having an agent who is less patient than the principal.
1606 Yi Chen Theoretical Economics 17 (2022) 5.1 Mean reverting state process The persistence of the state is demonstrably important to the intertemporal trade-off: future information sensitivity can sometimes outweigh its current counterpart because a shock to the state will echo in the distant future. In the main model, the persistence is high in the sense that any shock is permanent without decay. In this section, I weaken the persistence and allow for mean reversion. This change in the state process has two implications: that the state has a stationary distribution and that the increment of the state is negatively correlated with the state. The mean reversion setting brings the model closer to the common wisdom found in the existing literature on dynamic allocation and explains why contrarian patterns rarely arise there. The persistence is weaker when the state exhibits mean reversion. In this subsection, I consider an Ornstein–Uhlenbeck process of the form dθt=−φθtdt+dZt, where φ≥0 is a constant representing the strength of mean reversion. When φ=0, the process reduces to a special case of the main model. It can be shown that Proposition 1 still holds as a necessary condition. With a procedure similar to that used in the main model, the cost and policy functions are obtained as C(θ,W)=W−γφ◦f(θ)2+1 rγφ◦(γφ◦f)2(θ),x(θ,W)=W+f(θ)−γφ◦f(θ), where the operation γφ◦fproduces the unique solution g=γφ◦fto the second-order differential equation g(θ)−2φθg(θ)−2rg(θ)=−2rf (θ),lim θ→±∞e−√r 2|θ|g(θ)=0. (15) While an explicit solution to (15) is not obtainable in general, one can derive an alternative expression for γφ◦fby means of a forward stochastic differential equation γφ◦f=E∞ 0 re−rtf(θt)dtθ0=θ,dθt=−φθtdt+dZt. Even with mean reversion, the term γφ◦f(θ)is again interpreted as the expected discounted future target, similar to that in the main model. Whether the contract is conformist or contrarian depends now on the comparison between fand (γφ◦f). It is not easy to directly compare optimal contracts with different parameters φ.That being said, comparison is possible in the special case where fis a polynomial. When f(θ)=n k=0bkθk, the future target function γφ◦f(θ)is a polynomial of the same order, but the coefficient on the highest order nis dampened toward zero and becomes bn· r/(r+nφ). Therefore, lim θ→±∞ (γφ◦f)(θ) f(θ)=r r+nφ <1, meaning that the contract is conformist whenever the state is sufficiently far from zero. Intuitively, states far away from zero have a strong tendency to drift back and, hence,
Theoretical Economics 17 (2022) Dynamic delegation 1607 Figure 5. The effect of mean reversion (r=1, φ=0.5). (a) The target f(θ)=θ(solid) and the future target γφ◦f(θ)(dashed). (b) The target f(θ)=θ2(solid) and the future target γφ◦f(θ) (dashed). weigh less in the expected future target than they do in the case of zero mean reversion. As a result, the future information sensitivity (γφ◦f)(θ)gives disproportionally large probability weight to states near zero, attenuating itself below f(θ)when |θ|is large. Figure 5shows two simple examples where fis a polynomial. Figure 5(a) features a linear target where γφ◦fis flatter than fand, hence, the contract is conformist everywhere. According to Theorem 3, without mean reversion, we would have ended up with a babbling outcome. Figure 5(b) plots the case of a quadratic target. The future target γφ◦fhas a dampened slope compared to f, and the contract is conformist at all states except zero. Again, the contract does not lead to babbling, although it would have if φ=0. 5.2 Finite horizon In some economic applications, the time horizon for a contract is relatively short, because the principal is either unable or unwilling to commit for a long period. The contracting horizon can also be short because the state or cost becomes visible to the principal after some period of time. The finite horizon introduces nonstationarity to the contracting environment and creates a deadline effect on top of the patterns found in the main model. For simplicity, we assume in this subsection that μ=0. The contracting horizon is T>0, and both players discount at rate r>0. Let Wt≡Et[T te−r(s−t)xsds]/(1−e−r(T−t)) be the normalized expected continuation payoff of the agent at time t, and write C(θ,W,t)as the principal’s normalized cost function when the state is θ, the agent’s continuation payoff is W, and the calendar time is t. The HJB equation is modified to rC(θ,W,t)=min xrx−f(θ)2+r(W−x)CW(θ,W,t) +1−e−r(T−t)Ct(θ,W,t)+1 21−e−r(T−t)Cθθ(θ,W,t).
1608 Yi Chen Theoretical Economics 17 (2022) Figure 6. The change of γtover time, showing the deadline effect. Parameters: r=1, μ=0, T=4, t=0 (dashed), and t=3.6 (dotted). Even with three state variables, the above system is still solvable thanks to the quadratic flow cost. Following the same procedure as in the main model, I find the cost and policy functions taking similar forms C(θ,W,t)=W−γtf(θ)2+1−e−r(T−t) rγt(γtf)2(θ) x(θ,W,t)=W+f(θ)−γtf(θ), where γt(z)≡√r 2√21−e−r(T−t) ·e−√2r|z|Erfc√2|z|−2√r(T−t) 2√T−t−e√2r|z|Erfc√2|z|+2√r(T−t) 2√T−t is a time-dependent kernel. Figure 6(a) plots the kernel at different calendar times. As tincreases, the kernel is gradually concentrated around zero. In the limit as t→T,the kernel collapses to a Dirac delta function. Holding tfixed while extending Tto infinity, the kernel converges to a Laplace distribution as in the main model. Once again, γt f(θ)=E[T te−r(s−t)f(θs)ds|θt=θ]is the expected discounted future targets. Figure 6(b) shows this γtf evaluated at different times. As time passes, the “future” is shorter, and, therefore, less probability weight is given to states far from θin the above expectation. In the limit, we have limt→Tγtf(θ)=f(θ)for all θat which fis continuous. How does the response to information change over time as the contract approaches the end of the horizon? This is determined by the comparison between fand (γtf), according to the policy function. With a finite horizon, the expected future information sensitivity (γtf)depends not only on the state, but also on calendar time t.Dueto
Theoretical Economics 17 (2022) Dynamic delegation 1609 the required smoothness of function f,limt→T(γtf)(θ)=f(θ)almost everywhere. In other words, the gap f−(γtf)tends to zero and the responsiveness to information vanishes as the deadline approaches. As a result, the agent loses his influence on the action over time, and the information transmission gradually reduces to babbling. This is consistent with the two-period example, wherein the second-period information is disregarded. 5.3 Less patient agent In some agency problems, the principal has a longer horizon and is thus more patient than the agent. Krasikov, Lamba, and Mettral (2020) study a dynamic contracting model with unequal discounting, where the patience gap generates a front-loading motive for the agent’s payoff that is interacting with the initial back-loading force. In this extension, I study how the patience gap affects the optimal response to information. Suppose the principal discounts at rate r>0, while the agent has a higher rate ρ>r. To keep notation simple, I set μ=0. Since the incentive constraint is evaluated on behalf of the agent, the first-order condition (FOC) now requires dWt=ρ(Wt−xt)dt.IntheHJB equation for the cost function, however, it is the principal’s discount rate rthat takes place: rC(θ,W)=min xrx−f(θ)2+ρ(W−x)CW(θ,W)+1 2Cθθ(θ,W). The policy function can be similarly obtained, leading to the Euler equation x(θ,W)−f(θ)=2ρ−r ρW−γρf(θ), where γρhas the same expression as γin the main model except that μ=0andris replaced by ρ. Two immediate changes arise from the policy function. First, the current action no longer cashes out the promised continuation payoff Wproportionally. Instead, since (2ρ−r)/ρ > 1, the principal front-loads payoffs to the agent to exploit the difference in patience. Second, the response to information is now determined by ∂x/∂θ =f(θ)−γρ f(θ)·(2ρ−r)/ρ. Directly, the future information sensitivity is amplified by a factor (2ρ−r)/ρ > 1, making it easier to overtake the current information sensitivity. Indirectly, the distribution γρis more concentrated around zero due to the higher ρ, pulling the future information sensitivity toward its current counterpart. While the overall effect is ambiguous, the direct effect dominates in many cases. As an example, let f(θ)=θ. The contract in the main model is babbling, but now since ρ>r, the principal can benefit from the information to some extent. Since (f(θ)−2ρ−r ργρf (θ))/f (θ)= 1−(2ρ−r)/ρ < 0, the contract is always contrarian.
1610 Yi Chen Theoretical Economics 17 (2022) 6. Concluding remarks The fact that the optimal contract can, as demonstrated in this paper, behave in a contrarian manner is a noteworthy departure from the existing literature that explores similar problems. The contrarian pattern is unique to an intrinsically dynamic situation, wherein information unfolds over time. For contrast, consider a static setting wherein the agent observes the realization of the entire state path before reporting the path to the principal at once. In that case, the principal always responds to information in a conformist pattern, because with all states available, the current state no longer plays the dual role of justifying current action and predicting future actions. The contrarian pattern of the contract can be viewed as a new implication for agency problems; it never arises if there are no conflicts of interest. The more aligned the preferences are, the less likely it is that the optimal contract will be contrarian. The principal’s action moving in the opposite direction of the agent’s report should not be interpreted as distrust or punishment; instead, it can be understood as the principal’s efficient way of utilizing information in the presence of conflicting interests. Appendix A.1 Solving the two-period contract The IC’s are simplified to one equation: x1(θ1)+x2(θ1)=W. Plug this back into the objective to obtain the unconstrained problem min x1(·),W Ex1(θ1)−f(θ1)2+EW−x1(θ1)−f(θ2)2|θ1. For every θ1, the FOC with respect to x1(θ1)gives x1(θ1)=1 2W+1 2(f(θ1)−E[f(θ2)|θ1]). Plugging this into the objective and taking FOC with respect to W,wehaveW=Ef(θ1)+ Ef(θ2). Replacing Win the expression of x1(θ1)with the above, we have the solution for x1:x∗ 1(θ1)=1 2(f(θ1)−E[f(θ2)|θ1]) +1 2(Ef(θ1)+Ef(θ2)). We then find x2from the incentive compatibility (IC) condition: x∗ 2(θ1)=−1 2(f(θ1)−E[f(θ2)|θ1]) +1 2(Ef(θ1)+ Ef(θ2)). A.2 Revelation principle Lemma 3 (Revelation Principle). Given any contract xthat implements a mapping from state paths into action paths, there exists a truthful contract x†that implements the same mapping. Proof. Suppose the given contract xinduces a (not necessarily truthful) strategy m∈ M, which generates a mapping from state paths into action paths. Let Mt≡t 0msds be the accumulated misreporting. Consider a new contract x†such that x† t(ˆ θt)≡ xt(( ˆ θ+M)t). If the truth-telling strategy m†is not optimal for the agent under this contract, there exists a strategy m∈Malong with M t≡t 0m sdssuch that E[∞ 0re−rtx† t((θ+
Theoretical Economics 17 (2022) Dynamic delegation 1611 M)t)dt]>E[∞ 0re−rtx† t(θt)dt]. Contradiction arises as m+m∈Moutperforms min the original contract: E∞ 0 re−rtxtθ+M+Mtdt=E∞ 0 re−rtx† tθ+Mtdt >E∞ 0 re−rtx† tθtdt =E∞ 0 re−rtxt(θ+M)tdt. The new contract x†implements the original mapping from θtto xtby construction, ∀t. A.3 Proof of Lemma 1 Given a contract x, define the process of the agent’s total payoff evaluated at time 0 but with information at time t, ˆ W0 t≡t 0 re−rsxsds+e−rtWt, which is a martingale because for any 0 ≤t≤t, Etˆ W0 t=t 0 re−rsxsds+Ett t re−rsxsds+e−rtEt∞ t re−r(s−t)xsds=ˆ W0 t. By Theorem 1.3.13 in Karatzas and Shreve (1991), the martingale ˆ W0 thas a rightcontinuous-with-left-limit (RCLL) modification. Therefore, by Theorem 3.4.15 in the same book, the martingale has a representation: ˆ W0 t=ˆ W0 0+t 0 re−rsβsdZs∀t≥0. Subtracting the two expressions for ˆ W0 tand then differentiating with respect to t,we have dWt=r(Wt−xt)dt+rβtdZt=r(Wt−xt)dt+rβt(dˆ θt−μdt), which has an equivalent integral form Wt=W0+t 0r(Ws−xs)ds+t 0rβsdZs. A.4 Proof of Proposition 1 For any strategy m∈M, Novikov’s condition is satisfied. By the Girsanov theorem, there exists a martingale Ywith Yt≡et 0msdZs−1 2t 0m2 sds, serving as the Radon–Nikodym derivative between the measure induced by mand the measure under truth-telling. It evolves according to dYt=YtmtdZtwith Y0=1. Besides Yt, the cumulative misreporting Mt=t 0msdsis also a state variable, with evolution dMt=mtdt. Then the agent’s payoff from a strategy m∈Mis E[∞ 0re−rtYtxtdt].
1612 Yi Chen Theoretical Economics 17 (2022) Let pYbe the costate variable for the drift of Yand let qYbe the costate for the volatility of Y.LetpMand qMbe the counterparts for M. The agent’s current value Hamiltonian is rYx +qYYm+pMm. The first-order condition for m=0 to be optimal, evaluated at m=0, Y=1, is qY+pM=0. (16) The Euler equations for Yand M, evaluated at m=0, Y=1, are dpY=rpY−xdt+qYdZt dpM=rpMdt+qYdZt, (17) with transversality conditions limt→∞pY te−rt =0andlimt→∞ pM te−rt =0. The solution to the above backward stochastic differential equations are pY t=Et[∞ tre−r(s−t)xsds]= Wtand pM t=0, where Wtis the agent’s continuation payoff defined in Section 3.1. Hence, by comparing (17)and(7), we have qY=rβ. Plugging this back into (16)and using the fact that pM=0, we have the necessary condition β=0. A.5 Proof of Proposition 2 Suppose the agent uses an arbitrary strategy m∈M, so that the reported process is ˆ θ= θ+M,whereMt=t 0mtdt. The resulting action and continuation payoff processes are denoted as xmand Wm. Because ˆ θis in the support of of θ, these two processes evolve as if ˆ θwas the true state and the agent reported truthfully. Therefore, plugging in βt≡0, we have dWm t=r(Wm t−xm t)dtand, thus, xm tdt=−(ert/r)d(e−rtWm t). The agent’s payoff from strategy mis lim t→∞ Et 0 re−rsxm sds=W0−lim t→∞e−rtEWm t, whichmeansthataslongaslimt→∞e−rtEWm t=0 for all m∈M, the agent’s payoff is always W0regardless of his strategy. A.6 Proof of Lemma 2 Assume toward a contradiction that another contract ˆ x+a=x+aachieves E[∞ 0re−rt(ˆ xt+ a−f(θt))2]<E[∞ 0re−rt(xt+a−f(θt))2]and ∞ 0re−rt(ˆ xt+a)=W0+a. This means ∞ 0re−rt ˆ xt=W0and E∞ 0 re−rtˆ xt−f(θt)2 =E∞ 0 re−rtˆ xt+a−f(θt)2−a2−2aE∞ 0 re−rtˆ xt−f(θt) <E∞ 0 re−rtxt+a−f(θt)2−a2−2aE∞ 0 re−rtxt−f(θt) =E∞ 0 re−rtxt−f(θt)2,