scieee AI-readable full text Open interactive document viewer

Persistence in a dynamic moral hazard game

Bohren, J. Aislinn

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Bohren, J. Aislinn Article Persistence in a dynamic moral hazard game Theoretical Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Bohren, J. Aislinn (2024) : Persistence in a dynamic moral hazard game, Theoretical Economics, ISSN 1555-7561, The Econometric Society, New Haven, CT, Vol. 19, Iss. 1, pp. 449-498, https://doi.org/10.3982/TE2680 This Version is available at: https://hdl.handle.net/10419/296464 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Theoretical Economics 19 (2024), 449–498 1555-7561/20240449 Persistence in a dynamic moral hazard game J. Aislinn Bohren Department of Economics, University of Pennsylvania This paper explores how the persistence of past choices creates incentives in a continuous time stochastic game involving a large player (e.g., a firm) and a sequence of small players (e.g., customers). The large player faces moral hazard and her actions are distorted by a Brownian motion. Persistence refers to how actions impact a payoff-relevant state variable (e.g., product quality depends on past investment). I characterize actions and payoffs in Markov perfect equilibria (MPE) for a fixed discount rate, show that the perfect public equilibrium (PPE) payoff set is the convex hull of the MPE payoff set, and derive sufficient conditions for a MPE to be the unique PPE. Persistence can serve as an effective channel for intertemporal incentives in a setting where traditional channels fail. Applications to persistent product quality and policy targeting demonstrate the impact of persistence on equilibrium behavior. Keywords. Continuous time games, stochastic games, moral hazard. JEL classification. C73, L1. 1. Introduction This paper studies how the persistence of past choices can be used to create incentives in a continuous time stochastic game in which a large player interacts with a sequence of small players. Persistence refers to the impact that actions have on a payoff-relevant state variable, such as a worker’s rating, a firm’s product quality, or a government’s key economic variables. It can capture exogenous features of the environment, such as how past investment influences current quality or how past policy choices map into the current level of an economic variable. It can also capture endogenous design choices, such as how a rating system aggregates past reviews and rewards a worker based on her rating. The large player faces moral hazard and her past actions are not perfectly observed by consumers: they are distorted by a Brownian motion. Incentives can depend on the noisy signal of action choices as well as on how persistence influences future payoffs through the impact that actions have on the state. The goal of this paper is to determine whether and how persistence strengthens incentives to overcome moral hazard. J. Aislinn Bohren: [email protected] Earlier versions of this paper were circulated under the titles “Stochastic games in continuous time: Persistent actions in long-run relationships” and “Using persistence to generate incentives in a dynamic moral hazard problem.” I thank Simon Board, Matt Elliott, Jeff Ely, Ben Golub, Alex Imas, Bart Lipman, David Miller, George Mailath, Markus Mobius, Paul Niehaus, Andrew Postlewaite, Yuliy Sannikov, Andy Skrzypacz, Joel Sobel, Jeroen Swinkels, Joel Watson, Alex Wolitzky, and especially S. Nageeb Ali for useful comments. I also thank numerous seminar participants for helpful feedback. ©2024 The Author. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at https://econtheory.org.https://doi.org/10.3982/TE2680 450 J. Aislinn Bohren Theoretical Economics 19 (2024) The framework captures many economic settings in which past choices shape key features of current and future interactions. For example, a worker’s rating on a platform depends on the quality of service she has provided to previous customers. She may be rewarded for earning a good rating and punished for poor performance. This provides an incentive for her to earn and maintain a good rating. Similarly, a firm’s ability to make a high quality product is a function not only of its effort today, but also its past investments in developing technology and training its workforce. Quality today is linked to a firm’s future quality, in that customers experience similar quality across time due to the persistence of investment. When customers are willing to pay a higher price or buy a larger quantity of a high quality good, persistence provides an incentive for the firm to invest in developing a high quality product. Finally, a government’s success in reaching the target level for an economic variable depends on both past and current policy choices. When past policy choices impact the future value of an economic variable, the government may be willing to undertake more costly actions today, since the benefit of such actions continue to accrue in future periods. I study perfect public equilibria (PPE) in this framework; that is, equilibria in which strategies depend only on public information. I establish that the PPE payoff set is equal to the convex hull of the Markov perfect equilibrium (MPE) payoff set. In a MPE, equilibrium actions and payoffs depend only on the payoff-relevant components of the game— in this case, the observable state. Any paths of information that lead to the same current state prescribe the same continuation play. In contrast, a PPE can depend on past information in an arbitrary way. The intuition for this result stems from the type of incentives that are possible in games with small players and Brownian information. In a stochastic game, dynamic incentives can either be informational—signals are used to coordinate future equilibrium play—or structural—actions impact the structure of future interactions through their impact on the state, including both the state’s direct impact on future feasible payoffs and its indirect impact through its effect on future equilibrium play. There are two main forms of informational incentives: burning value, where incentives are created by the threat of switching to an inefficient action profile, and transferring continuation payoffs tangent to the set of equilibrium payoffs. It is not possible to provide incentives with transfers when facing small players, and Brownian information is too noisy to create effective incentives via value-burning (Sannikov and Skrzypacz (2010)). Therefore, any nontrivial incentives in games with small players and Brownian information must be structural. This is precisely the channel for incentives in a MPE, as informational channels are precluded by definition.1 In establishing this result, I characterize equilibrium payoffs and actions in MPE for a fixed discount rate. This characterization yields sharp insights. It shows that whether persistence allows the large player to overcome moral hazard depends on the marginal 1In earlier related work, Faingold and Sannikov (2011) establish a similar result when small players have incomplete information about the large player’s type and the state is the belief that the large player is committed to choosing a certain action. Stochastic games with multiple large players, Brownian information, and a failure of identifiability will likely have a similar equilibrium characterization to this paper, as such games face similar issues with informational incentives (Abreu, Milgrom, and Pearce (1991), Sannikov and Skrzypacz (2007)). Theoretical Economics 19 (2024) Dynamic moral hazard game 451 impact of its action on the state and how sensitive the continuation payoff is to changes in the state. In contrast to a folk theorem, it determines what type of equilibria one expects to emerge and what pattern of behavior generates a given payoff. It shows how the dynamics of behavior depend on observable outcomes (e.g., a restaurant rating or an economic variable) and how incentives and payoffs depend on key parameters of the model (e.g., the depreciation rate of investment). The characterization of the continuation payoff captures both the direct and equilibrium channels for structural incentives. For example, when consumers have observed a given level of quality in the recent past, their willingness to pay for a product depends on both the persistence of this quality— the direct channel—as well as their belief about how quality influences the firm’s current investment choice—the equilibrium channel. The interaction of these two channels can significantly strengthen or dampen incentives, depending on the structure of the game. The second main result determines when a MPE emerges as the unique PPE. This result relies on determining when there is a unique MPE in the class of Markov equilibria. When there is a unique MPE, the result described above establishes that this will also be the unique PPE. Uniqueness depends on incentives as the state approaches the boundary of the state space. If boundary incentives are unique, e.g., it is possible to sustain a unique equilibrium action profile and payoff at the boundary, then from the MPE characterization, incentives must also be unique on the interior of the state space. I present sufficient conditions for uniqueness in two cases: (i) an unbounded state space and (ii) a bounded state space. In case (i), these conditions rule out complementarities between the direct and the equilibrium channels for incentives near the boundary, such as multiple optimal action profiles due to coordination motives. In case (ii), these conditions ensure that incentives collapse as the state approaches the boundary, which rules out the possibility of sustaining multiple equilibrium action profiles at the boundary. Several applications illustrate how persistence can be used to create effective incentives. The first application modifies the canonical product choice setting to allow a firm’s effort to have a persistent effect on the quality of its product. I show that persistence provides effective incentives for the firm to invest in building a high quality product. These incentives are present in the long run, in that the firm continues to choose a positive level of investment as the time period grows large. I also consider a variation of the product choice game in which the marginal return to quality is non-monotonic and show that this can lead to firms specializing in low or high quality. In the second application, constituents elect a board to implement a policy that targets an economic variable. The level of the variable depends on current and past decisions by the board. For example, the Federal Reserve targets an interest rate or a board of directors sets a growth target for a company. I show that the board’s incentive to undertake costly intervention is strongest when the economic variable is an intermediate distance from its target; when it is far from its target, the benefit of intervention is significantly delayed, while when it is close to its target, the benefit of further intervention is small. In the final application, a government and a group of innovators invest in intellectual capital, and there is a strategic complementarity between their investments. This complementarity 452 J. Aislinn Bohren Theoretical Economics 19 (2024) gives rise to multiple Markov equilibria, including one in which neither party invests and several that sustain a positive level of investment. The equilibrium characterization in each application can be used to address important design questions. For example, a comparative static on how a firm’s payoff varies with the persistence of its effort provides insight into the optimal durability for a production technology, while a comparative static on how a worker’s effort varies with the persistence of its rating is useful for designing rating systems. 1.1 Literature Recent results on repeated games between a long-run/large and short-run/small players show that the intersection of noise in monitoring and instantaneous adjustment of actions creates a genuine challenge in providing intertemporal incentives (Fudenberg and Levine (2007,2009), Faingold and Sannikov (2011)).2In the analogue of this paper with no persistence, the large player cannot earn an equilibrium payoff above the best static Nash payoff.3In contrast, the equilibrium characterization in this paper demonstrates that persistence can lead to effective intertemporal incentives and enable the large player to overcome moral hazard. The literature on reputation with behavioral types is another important and well understood mechanism to overcome moral hazard in similar settings (Fudenberg and Levine (1989,1992), Faingold and Sannikov (2011), Faingold (2020)). If consumers believe that there is a chance that the firm is committed to choosing high effort, then the firm will be able to charge a higher price for its product. Incomplete information about the firm’s type creates a form of persistence, as consumers’ beliefs depend on past effort choices. However, fixing a strategic firm’s patience, such reputation effects vanish in the ex ante probability of behavioral types, and so the effectiveness of persistence via incomplete information requires a nontrivial fraction of behavioral types.4 The connection with the reputational literature motivates several key insights. First, when the firm is known to be strategic, this paper shows that other forms of persistence can overcome moral hazard.5Second, in contrast to the temporary incentives in reputation models (Cripps, Mailath, and Samuelson (2004), Faingold and Sannikov (2011)), the incentives in a stochastic game persist in the long run.6Finally, at a theoretical level, this 2Abreu, Milgrom, and Pearce (1991) first examined incentives in repeated games with imperfect monitoring and frequent actions. They established that shortening the period between actions has a crucial impact on the ability to structure effective incentives. 3Sannikov and Skrzypacz (2007) show that this is also the case in games between multiple long-run players in which deviations between individual players are indistinguishable. 4Kreps, Milgrom, Roberts, and Wilson (1982), Kreps and Wilson (1982), and Milgrom and Roberts (1982) first demonstrated that reputation, in the form of incomplete information about a player’s type, has a dramatic effect on equilibrium behavior. Mailath and Samuelson (2001) show that reputational incentives can also come from a firm’s desire to separate itself from an incompetent type. 5Along these lines, Dilmé (2019) shows that adjustment costs can help a firm overcome moral hazard by endogenously creating persistence. 6Long-run reputation effects are also possible in models with behavioral types when consumers cannot observe all past signals (Ekmekci (2011)) or the type of the firm is replaced over time. Theoretical Economics 19 (2024) Dynamic moral hazard game 453 paper explores the general properties of a stochastic game that has powerful intertemporal incentives. The reputational game can be viewed as a specific type of stochastic game. For instance, if instead of influencing the uncertainty about whether it is a behavioral type, a strategic firm makes a costly initial investment in a new production technology that benefits customers today and in the future, we observe similar intertemporal incentives in the resulting stochastic game. This final point merits a closer comparison with Faingold and Sannikov (2011), who characterize the unique MPE in the stochastic game that corresponds to a continuous time reputation model. In their paper, payoffs and the evolution of the state take a specific form due to Bayesian updating. My characterization builds on the techniques in their paper to understand more generally what properties of stochastic games are needed for uniqueness of MPE and nondegenerate intertemporal incentives. I analyze a general class of stochastic games that places few restrictions on the process governing the evolution of the state and the structure of payoffs. The key technical advancement, relative to their paper, is for the case of an unbounded state space and payoff for the large player, as it requires significantly different techniques to complete the analysis. Beyond reputation models with behavioral types, a rich literature analyzes dynamic games with a state variable in which effort is directly linked to future payoffs via the state. Ericson and Pakes (1995) were the first to analyze hidden investment and stochastic capital accumulation (the state) in a model that is similar in spirit to the quality example presented in Section 2. They study firm and industry dynamics, and establish equilibrium existence. Doraszelski and Satterthwaite (2010)modifyEricson and Pakes (1995)to guarantee the existence of a pure strategy MPE, which is computationally tractable. Neither paper establishes uniqueness, but instead focus on the dynamics associated with a particular MPE. More broadly, MPE is the workhorse solution concept across industrial organization and political economy. A comprehensive review of this literature is beyond the scope of this paper. This paper also relates to a literature on stochastic games with an unobservable state. In these games, incentives stem from the large player’s ability to manipulate the public belief about the state through her effort choice. Cisternas (2018) characterizes necessary conditions for the existence of Markov equilibria in a continuous time stochastic game with an unobservable state and sufficient conditions in two more restrictive classes of games. Hidden states significantly complicate the model, and it is not possible to establish uniqueness results or a full equilibrium characterization. Board and Meyer-ter-Vehn (2013) study a setting in which a firm’s hidden quality depends on past effort and consumers learn about this quality from noisy signals. My paper differs in focus in that there is no adverse selection, there is strategic interaction between the large and small players, and it allows for a richer class of stage game payoffs. Several folk theorems exist for discrete time stochastic games with observable states, beginning with a perfect monitoring setting in Dutta (1995) and extending to imperfect monitoring environments in Fudenberg and Yamamoto (2011) and Hörner, Takuo, Satoru, and Vieille (2011). My setting differs in that there is a single large player and information follows a diffusion process. It is already known that these two changes significantly alter incentives in standard repeated games (compare the folk theorem in Fudenberg, Levine, and Maskin (1994) to the equilibrium degeneracy in Fudenberg and 454 J. Aislinn Bohren Theoretical Economics 19 (2024) Levine (2007,2009) and Faingold and Sannikov (2011)). The intuition is similar for the discrete time stochastic game folk theorems compared to the MPE uniqueness result in this paper.7 The organization of the paper proceeds as follows. Section 2presents a product choice example to motivate the model. Section 3sets up the model and characterizes the structure of PPE. Section 4presents the three main results: existence of a Markov equilibrium, characterization of the PPE payoff set, and uniqueness of a Markov equilibrium in the class of all PPE. Section 5presents structural results on the shape of equilibrium payoffs. Section 6explores several applications. All proofs are provided in the Appendix. 2. Example 1: Product choice with persistent quality Consider a variation of the canonical product choice setting in which a monopolist firm provides a product to consumers and the firm’s effort has a persistent effect on the quality of the product. At each time t, the firm chooses an unobservable effort level at∈[0, a], where a>0. The quality of the firm’s product at time tdepends on both current and past effort, q(at,Xt)=(1−λ)at+λXt, where past effort influences quality through the observable stock quality Xt=t 0 e−θ(t−s)(asds +dZs), θ>0 determines the decay rate of past effort, (Zt)t≥0is a standard Brownian motion, and λ∈[0, 1]captures the relative importance of past effort in determining current quality.8Effort increases quality both today and in the future. There is a continuum of identical consumers of unit mass. Consumers value quality: when they believe that the firm will choose effort level ˜ atat time t, they are willing to pay q(˜ at,Xt)for one unit of the product. Each consumer purchases the product for a price equal to her willingness to pay when it is positive and otherwise does not purchase. Therefore, the firm earns a flow revenue of bt=q(˜ at,Xt)when q(˜ at,Xt)>0andbt=0 otherwise. This exact form of revenue is chosen for simplicity; the important feature is that the flow revenue is increasing in quality and independent of the true current effort choice. Effort has flow cost a2 t/2 and the firm discounts at rate r>0. Therefore, the firm’s average discounted payoff equals r∞ 0 e−rtbt−a2 t/2dt. In the unique PPE with no persistence, λ=0, the firm exerts zero effort, quality is equal to zero, and the firm earns zero profit (this is a direct application of Theorem 3 7The paper also relates to an older literature on stochastic games and the existence of Markov equilibria in discrete time, including Shapley (1953), Dutta and Sundaram (1992), Nowak and Raghavan (1992), Duffie, John Geanakoplos, and McLennan (1994). 8In a slight abuse of notation, the Lebesgue integral and the stochastic integral are placed under the same integral sign. Theoretical Economics 19 (2024) Dynamic moral hazard game 455 from Faingold and Sannikov (2011)). Intertemporal incentives break down, despite the fact that the firm would earn higher profits if it could commit to higher effort.9 In this paper, I show that persistent quality incentivizes the firm to choose a positive level of effort and earn positive profits. Theorems 1to 3establish that there is a unique PPE, which is Markov in the stock quality Xt. The effort level and profit in this unique equilibrium are characterized as a function of the impact of past effort on current quality λ, the depreciation rate of quality θ, and the discount rate r.Foranyλ>0, the firm chooses a positive level of effort and earns positive profits at positive and some (possibly all) negative levels of stock quality. Further, the firm has a long-run incentive to choose high effort. This contrasts with models in which the incentive to produce high quality is derived from consumers’ uncertainty over the firm’s payoffs and long-run effort converges to zero (Cripps, Mailath, and Samuelson (2004), Faingold and Sannikov (2011)). Persistence increases the firm’s payoffs through two complementary structural channels. First, the firm’s effort increases the stock quality, which increases future revenue through its impact on future prices. This is the direct effect of persistence, as discussed in the Introduction. Second, persistence creates a link with future payoffs, which allows the firm to credibly choose a positive level of effort today, thereby increasing the current price, and hence, revenue. This second channel arises from the strategic interaction between the firm and consumers: it is the equilibrium effect discussed in the Introduction. When quality is high, the continuation value is approximately linear and it is possible to quantify the share of profit arising from each of these channels. The present value of the direct effect minus the cost of effort is approximately λ2/2(r+θ)2, which is higher when past effort plays a larger role in determining current quality (higher λ), quality depreciates at a lower rate (lower θ), or the firm is more patient (lower r). The present value of the equilibrium effect is approximately (1−λ)λ/(r+θ),whichisalso higher when quality depreciates at a lower rate or the firm is more patient. In contrast to the direct effect, the equilibrium effect is largest for intermediate values of λ.Thisis because the incentive to exert effort is increasing in λwhile the impact of effort on the current price is increasing in 1 −λ. This example will be used throughout the paper to demonstrate the results. The product choice framework lends itself to other variations, several of which are discussed in Section 6.1. 3. Model 3.1 Model setup States and actions A large player and a continuum I≡[0, 1]of identical small players, indexed by i, play a continuous time stochastic game with imperfect monitoring. At each instant of time t∈[0, ∞), a publicly observable state variable Xtin nonempty closed interval X⊂Rdetermines the action set and feasible flow payoffs. If Xis 9In contrast to Abreu, Milgrom, and Pearce (1991) and Sannikov and Skrzypacz (2007), this breakdown of incentives takes place despite there being no failure of identifiability. 456 J. Aislinn Bohren Theoretical Economics 19 (2024) bounded, denote the upper and lower boundary states by X≡supXand X≡infX,respectively, and assume X0∈(X,X). Large and small players simultaneously choose actions atfrom Aand bi tfrom B(Xt), respectively, where Ais a nonempty compact subset of a Euclidean space and B(X)is a nonempty compact subset of a closed Euclidean space Bwith continuous correspondence X→ B(X). Denote the set of feasible pairs of small player actions and states as E≡{(b,X)∈B×X|b∈B(X)}. Assume that the boundary of the feasible set of actions for small players grows at most linearly with the state; that is, there exists a Kb,cb>0suchthatforall(b,X)∈E,|b|≤Kb|X|+cb.10 Individual actions are privately observed. Players observe the aggregate distribution of small players’ actions, bt∈B(Xt), and do not observe the large player’s action. Given initial state X0, the state evolves stochastically according to dXt=μ(at,bt,Xt)dt +σ(bt,Xt)dZt,(1) where (Zt)t≥0is a one-dimensional Brownian motion, and the drift and volatility are determined by Lipschitz continuous functions μ:A×E→Rand σ:E→R,whichare linearly extended to A×{(b,X)∈B ×X|supp b⊂B(X)}and {(b,X)∈B ×X|supp b⊂ B(X)}, respectively.11 The drift depends on the large player’s action, the aggregate action of the small players, and the state. Volatility is independent of the large player’s action to maintain the assumption that it is not perfectly observed. If the state space is bounded, then to prevent the state from escaping its boundary and maintain imperfect monitoring at the boundary, the volatility must be zero at the upper and lower bounds, σ(b,X)=0forallb∈B(X)and σ(b,X)=0forallb∈B(X), and the drift must be weakly negative at the upper bound, weakly positive at the lower bound, and independent of (a,b)at both bounds, μ(a,b,X)=m≤0forall(a,b)∈A×B(X)and μ(a,b,X)=m≥0forall(a,b)∈A×B(X). To ensure that the future path of the state is stochastic, except at boundary states, assume that its volatility is positive at all interior states. Assumption 1 (Positive Volatility). When X=R,infEσ(b,X)>0.WhenXis compact, there exists a C>0such that σ(b,X)≥C(X−X)(X−X)for all (b,X)∈E. This assumption rules out interior absorbing states, where state Xis absorbing if the drift and volatility are equal to zero, μ(a,b,X)=0andσ(b,X)=0forall(a,b)∈ A×B(X). The path of the state provides a public signal of the large player’s action. There are no additional public signals. This is without loss of generality, as additional public signals have no effect on the equilibrium characterization (see the discussion in Section 3.3). Let (Ft)t≥0represent the filtration generated by the public information (Xt)t≥0.Small players observe no information about the large player’s action beyond what is contained in (Ft)t≥0. 10Iuse|·|to denote the Euclidean norm for vectors.. 11Functions μand σare extended to distributions as B(X)μ(a,b,X)db(b)and B(X)σ(b,X)2db(b). Theoretical Economics 19 (2024) Dynamic moral hazard game 463 4.1 Existence of Markov equilibria In a Markov equilibrium, the continuation value and actions depend solely on the current value of the state; they are independent of the past path of the state. Since the path of the state provides a signal of the large player’s action, using it to punish or reward the large player could give rise to PPE in which different paths of the state specify different continuation payoffs and equilibrium actions, even when these paths map to the same current state. In a Markov equilibrium, this is not allowed. Theorem 1establishes existence of a Markov equilibrium and characterizes equilibrium behavior and payoffs in Markov equilibria. The continuation value is characterized as the solution(s) U:X→Rto an ordinary differential equation that maps each state to a payoff. If there are multiple solutions, then each solution characterizes a Markov equilibrium (see Section 6.3 for an illustration of a setting with multiple Markov equilibria). Given a solution U, the corresponding Markov equilibrium action profile is the sequentially rational action profile at state Xand incentive weight U(X)/r. The large player has nondegenerate incentives at any state with U(X)= 0. Theorem 1. Assume Assumptions 1to 3. Given initial state X0,ifUisasolutiontothe optimality equation rU(X)=rg∗X,U(X)+U(X)μ∗X,U(X)+1 2U(X)σ∗X,U(X)2(7) on X(on (X,X)if Xis compact) and Uhas linear growth (is bounded if gis bounded), then Ucharacterizes a Markov equilibrium with the following payoffs and actions: (i) Equilibrium payoff U(X0). (ii) Continuation values (Wt)t≥0=(U(Xt))t≥0. (iii) Equilibrium actions (at,bt)t≥0=(S∗(Xt,U(Xt)))t≥0,whereS∗(X,U(X)) is the unique solution to (6)at state Xand incentive weight U(X)/r. The optimality equation has at least one twice continuously differentiable solution that lies in the range of feasible payoffs for the large player and has linear growth (is bounded if gis bounded). Thus, there exists at least one Markov equilibrium. From the optimality equation, the continuation value U(X)is equal to the sum of the equilibrium flow payoff g∗(X,U(X)) and the expected change in the continuation value. This expected change has two components: (i) the interaction between the slope of the continuation value and the drift of the state, U(X)μ∗(X,U(X))/r, and (ii) the interaction between the concavity of the continuation value and the volatility of the state, U(X)σ∗(X,U(X))2/2r. In relation to the discussion in Section 3.3, a nontrivial incentive weight is possible at some states without the continuation value escaping the payoff set. Theorem 1 shows that the volatility of the continuation value in a Markov equilibrium is equal to its slope, rβt=U(Xt). At any interior state X∗that yields the maximum continuation value 464 J. Aislinn Bohren Theoretical Economics 19 (2024) across all states, U(X∗)=0. Therefore, when Xt=X∗, the volatility of the continuation value is zero, rβt=0, which ensures that the continuation value does not escape the payoff set. In these periods, the large player acts myopically and earns the static Nash payoff in state X∗. At other states, the continuation value can be sensitive to changes in the state, U(X)= 0, generating nontrivial incentives. Outlineofproof In a Markov equilibrium, continuation values take the form of Wt= U(Xt)for some function U. Assuming that Uis twice continuously differentiable, by Ito’s formula the continuation value must follow the law of motion, dU(Xt)=U(Xt)μa∗ t,b∗ t,Xtdt +1 2U(Xt)σb∗ t,Xt2dt +U(Xt)σb∗ t,XtdZt. By Lemma 1, the continuation value must also follow the law of motion in (5). Matching the drifts of these two laws of motion yields the optimality equation, while matching the volatilities yields the equilibrium volatility of the continuation value, rβt=U(Xt). Showing that the optimality equation has at least one twice continuously differentiable solution that lies in the range of feasible payoffs for the large player establishes existence. Faingold and Sannikov (2011) follow similar steps to derive a Markov equilibrium in a game of incomplete information. Relative to their derivation, the innovative part of my proof lies in establishing existence of a solution to the optimality equation when the state space is unbounded, particularly when gis also unbounded. I show by construction that there exist lower and upper solutions to the optimality equation, α:X→Rand α:X→R, that have linear growth. This is only possible when the maximum drift of the state has linear growth at rate less than r(Assumption 2). The lower and upper solutions characterize bounds on the solution to the optimality equation, α(X)≤U(X)≤α(X) for all X. Next I show that the bound on the optimality equation grows linearly with respect to U(X)and, therefore, the optimality equation does not grow too quickly (technically speaking, it satisfies a growth condition on any compact subset of the state space). These conditions establish that the optimality equation has a twice continuously differentiable solution with linear growth. When gis bounded, the lower and upper solutions are constant, which establishes existence of a bounded solution. The final step is to show that the continuation value and actions characterized above constitute a Markov equilibrium. Given a solution U(X)and an action profile uniquely specified at state Xtby (a∗ t,b∗ t)=S∗(Xt,U(Xt)) (where uniqueness follows from Assumption 3), the state variable evolves uniquely according to (1), the continuation value (U(Xt))t≥0satisfies the law of motion (5), and the action profile satisfies the conditions for sequential rationality (6). Therefore, (a∗ t,b∗ t,U(Xt)) constitute a PPE. For a given solution U(X), the state evolves uniquely and actions are uniquely specified as a function of the state. Therefore, each solution to the optimality equation characterizes a unique Markov equilibrium. If there are multiple solutions, then there will be multiple Markov equilibria. Example 1 (Product Choice, cont.). Given a(X,z)and b(X,z)characterized in Section 3.2, any solution to rU(X)=rbX,U(X)−r 2aX,U(X)2+U(X)aX,U(X)−θX+1 2U(X)(8) Theoretical Economics 19 (2024) Dynamic moral hazard game 465 with linear growth as X→∞and bounded as X→−∞characterizes a Markov equilibrium with equilibrium actions a(X,U(X)) and b(X,U(X)).♦ 4.2 The PPE payoff set Let ξ:X⇒Rdenote the correspondence that maps each state onto the corresponding set of PPE payoffs for the large player, and let ϒ:X⇒Rdenote the analogous correspondence for the Markov equilibrium payoffs characterized by the optimality equation in Theorem 1.Theorem2shows that in any PPE, the large player cannot achieve a payoff above the highest or below the lowest Markov equilibrium payoff in ϒ. Theorem 2. Assume Assumptions 1to 3. Then for any state X∈X(state X∈(X,X)if X is compact), the set of PPE payoffs of the large player at state Xis equal to the convex hull of the set of Markov equilibrium payoffs at state X,ξ(X)=co(ϒ(X)). The impossibility of the large player achieving a PPE payoff above the highest Markov payoff in ϒyields insight into the type of incentives generated by persistence. As discussed in the Introduction and Section 3.3, incentives can be either informational or structural. When a Markov equilibrium yields the highest equilibrium payoff, it precludes the existence of equilibria that achieve higher payoffs using informational incentives. Therefore, any nontrivial incentives arising from persistence are structural. Outlineofproof The key argument in the proof shows that any PPE with an initial payoff above the highest Markov equilibrium payoff in ϒwill eventually yield a continuation value that lies outside the set of feasible payoffs for the large player, which is a contradiction. Suppose that a PPE with continuation values (Wt)t≥0yields a payoff higher than the maximum Markov equilibrium payoff in ϒat state X0.LetDt≡Wt−U(Xt)be the difference between the continuation values in these two equilibria at time t.Ishowthat whenever D0>0, Dtwill grow arbitrarily large with positive probability, independent of Xt. By Lemma 1,|Wt(S)|is bounded with respect to Xt.Thus,Dtcan only grow arbitrarily large when Xtgrows arbitrarily large, so it cannot be that D0>0. This escape argument is similar to other papers in the literature, in particular Faingold and Sannikov (2011). Their proof relies on the compactness of the state space to show that the volatility of Dtis bounded away from zero and relies on the boundedness of the flow payoff to reach a contradiction when Dtgrows arbitrarily large. Therefore, their proofs do not trivially extend to an unbounded state space or an unbounded flow payoff. The innovative parts of this proof are to establish that the volatility of Dtis bounded away from zero on an unbounded state space and to show that when Dtgrows arbitrarily large, it can jump outside of the feasible payoff set (a contradiction) provided the state does not grow too quickly. Equilibrium degeneracy without persistent actions If the state evolves independently of the large player’s action, then there is no link between the current action and the continuation value. It is not possible to generate effective intertemporal incentives and the large player acts myopically. In the unique PPE, both players play the static Nash equilibrium action profile S∗(X,0 )at all states X. 466 J. Aislinn Bohren Theoretical Economics 19 (2024) Corollary 1. Assume Assumptions 1to 3and suppose μis independent of afor all X. Then in the unique PPE, (at,bt)=S∗(Xt,0 )for all t≥0and the continuation value is characterized by the unique solution to the optimality equation (7). This is the stochastic game analogue of the equilibrium degeneracy result in repeated games with a long-run player and short-run/small players (Fudenberg and Levine (2007,2009), Faingold and Sannikov (2011)). 4.3 Equilibrium uniqueness This section establishes sufficient conditions for there to be a unique PPE, which is Markov. The main step is to determine when the optimality equation has a unique feasible solution. When this is the case, Theorem 2establishes that PPE payoffs are uniquely specified as the payoffs in this unique Markov equilibrium. The behavior of the optimality equation as the state approaches its boundary plays a key role in establishing when it has a unique solution. Any two feasible solutions that satisfy the same boundary conditions cannot differ on the interior of the state space: they must be equivalent (see Lemma 7in Appendix A.4). Therefore, establishing that all feasible solutions satisfy the same boundary conditions is necessary and sufficient to establish a unique solution. I outline a set of sufficient conditions to guarantee this when X=R; the case of a compact state space requires no additional conditions. The application in Section 6.3 illustrates how multiple Markov equilibria can arise when this condition fails. 4.3.1 Unbounded state space (X=R)Assumption 4(below) outlines a set of sufficient conditions for a unique Markov equilibrium when X=R. The first condition requires the large player’s equilibrium flow payoff and the equilibrium drift to be additively separable in the state Xand incentive weight zas Xapproaches ∞and −∞.Thisrulesout complementarities between the direct and equilibrium channels for incentives near the boundary, which prevents multiple equilibrium incentive weights—and hence, equilibrium action profiles—at a given state. It is used to establish that the slope of the continuation value converges to the same limit in all Markov equilibria. The second condition relates to the volatility: it is a technical condition that helps establish that two distinct solutions to the optimality equation cannot have the same limit slope. The third condition applies to a growth model where the drift of the state approaches infinity as X→∞ (or approaches negative infinity as X→−∞); it ensures that the volatility does not also grow arbitrarily large. It is also used to pin down a unique boundary continuation value. Assumption 4. (i) Additive Separability Near Boundary. There exists a δ>0and continuously differentiable functions g1,μ1:X→Rand g2,μ2:R→Rwith μ1monotone such that for |X|>δ,g∗(X,z)=g1(X)+g2(z)and μ∗(X,z)=μ1(X)+μ2(z). (ii) Volatility. The function σ∗(X,z)2is Lipschitz continuous. Theoretical Economics 19 (2024) Dynamic moral hazard game 467 (iii) Growth Case. When limX→∞ μ1(X)=∞, then there exists an ε,δ>0such that for X>δand z∈R,|μ1(X)|/σ∗(X,z)2>ε, and similarly when limX→−∞ μ1(X)= −∞.23 Given Assumption 4(i), select g1(X)and g2(z)such that g2(z)contains any constant term in g∗(X,z)to uniquely pin down each function, and similarly for μ1(X)and μ2(z). When gis bounded, it is possible to establish uniqueness without additive separability; Assumption 5 in Supplemental Appendix D.3 (available at http://econtheory.org/supp/ 2680/supplement.pdf) presents an alternative condition. Theorem 3establishes uniqueness and characterizes the limit of the continuation value and its slope as the state grows large. Theorem 3. Suppose X=Rand assume Assumptions 1to 4. For each initial state X0∈ X, there exists a unique PPE that is Markov and characterized by the unique solution Uof (7)on Xwith linear growth (bounded when gis bounded). The slope of the continuation value converges to a constant, lim X→xU(X)=zxwhere zx≡lim X→xrg1(X)/rX −μ1(X),(9) and the continuation value converges to lim X→xU(X)−y(X)=g2(zx)+zxμ2(zx)/r (10) for x∈{−∞,∞},wherey(X)≡−φ(X)(rg1(X)/φ(X)μ1(X))dX and φ(X)≡ exp((r/μ1(X))dX)when limX→xμ1(x)= 0,andy(X)≡g1(X)when limX→xμ1(X)= 0.Whengis bounded, this implies the continuation value converges to the limit static Nash equilibrium payoff and the slope of the continuation value converges to zero: for x∈{−∞,∞}, lim X→xU(X)−g∗(X,0 )=0and lim X→xU(X)=0. (11) Theorem 3establishes that the slope of the continuation value converges to a unique limit slope, which is equal to the ratio of the growth rate of the flow payoff to the growth rate of the drift with respect to the state. Given this slope, the boundary condition (10) highlights the impact of structural incentives on the continuation payoff. Repeated play of the static Nash equilibrium profile yields a payoff UNE that satisfies limX→xUNE(X)−y(X)=g2(0)+zxμ2(0)/r. Therefore, from (10), the continuation value approaches the sum of this repeated static Nash payoff and a constant g2(zx)−g2(0)+zx(μ2(zx)−μ2(0))/r. This constant determines the extent to which structural incentives persist at the boundary of the state space. The first term, g2(zx)−g2(0), captures the equilibrium effect of persistence. It is the portion of the equilibrium flow payoff that arises from future strategic interaction; it captures the effect of the large player’s action on the small players’ actions, net of the cost of a.The 23Part (ii) is unnecessary when gis bounded. Note that part (iii) holds trivially when σ(b,X)is bounded. 468 J. Aislinn Bohren Theoretical Economics 19 (2024) second term, zx(μ2(zx)−μ2(0))/r, captures the direct effect of persistence on future feasible payoffs, measured by how the continuation value changes with respect to the state and how the state changes with respect to the large player’s equilibrium action relative to the static Nash action. If this constant is positive, then as the state becomes large, structural incentives provide the large player with a payoff that is strictly higher than the payoff from playing the static Nash profile at each state. When the asymptotic slope zxis nonzero, it is possible to sustain nontrivial intertemporal incentives as the state grows large. This is an important and novel insight of this paper. If it is possible to sustain nontrivial incentives at the boundary of the state space, then incentives are permanent in the sense that they do not dissipate with time, regardless of the asymptotic behavior of the state with respect to time. In the case of a bounded flow payoff, zx=0. Therefore, incentives collapse at the boundary and the continuation value converges to the limit of the static Nash payoff. However, this does not preclude the existence of long-run incentives: even when incentives collapse at the boundary, the state does not necessarily converge to a boundary state as t→∞. Therefore, it can be possible to sustain nontrivial incentives in the long run. The continuation of Example 1below illustrates how to verify Assumption 4and derive the boundary conditions in Theorem 3when the flow payoff is unbounded, while Section 6.1 illustrates how to do so for a bounded flow payoff. Section 6.3 shows that there can be multiple MPE in an application in which Assumption 4(specifically, additive separability) fails. Outline of proof I first show that all solutions to the optimality equation have the same boundary conditions. Faingold and Sannikov (2011) also characterize boundary conditions as a step towards establishing that the optimality equation in their paper has a unique solution. Relative to their result, the innovative part of my proof is in establishing boundary conditions for an unbounded flow payoff and state space, as I next describe. Let ψ(X,z)≡g∗(X,z)+zμ∗(X,z)/r be the sum of the large player’s flow payoff and return on effort at the sequentially rational action profile (a(X,z),b(X,z)),andletU(X) be a solution to the optimality equation. Suppose that U(X)does not converge as X→ ∞. Then for any slope zsuch that the continuation value has slope zinfinitely often at large X,U(X)will alternate between being convex and concave at slope z.Fromtheoptimality equation, ψ(X,z)will lie above U(X)when it is concave at slope zand will lie below U(X)when it is convex at slope z. Therefore, the oscillation of ψ(X,U(X)) is at least as large as the oscillation of U(X). This violates the monotonicity of ψ,soitmust be that U(X)has a limit z∞∈R. Since U(X)has linear growth (by Theorem 1), this limit must be finite. Moreover, it is equal to limX→∞ U(X)/X. Given additive separability, as well as the Lipschitz continuity and monotonicity of μ1and g1, the limits of ψ(X,z)/X and ψ(X,z)exist and are equal as X→∞. Denote these limits by ψ∞(z).Weusethese properties and the optimality equation to show that limX→∞ σ∗(X,U(X))2U(X)/X = 0 and, therefore, limX→∞ U(X)/X −ψ(X,U(X))/X =0. This establishes that the limit slope z∞is a fixed point of ψ∞(z). The additively separable assumption on g∗ and μ∗issufficienttoensurethatψ∞(z)has a unique fixed point, which is equal to z∞=limX→∞ rg1(X)/(rX −μ1(X)). This guarantees that all solutions to the optimality equation have the same limit slope. Theoretical Economics 19 (2024) Dynamic moral hazard game 469 Using the characterization of the limit slope, it can be shown that any solution U(X)to the optimality equation satisfies limX→∞ U(X)−U(X)μ1(X)/r −g1(X)= g2(z∞)+z∞μ2(z∞)/r. Consider the linear first-order differential equation (FODE) y(X)−y(X)μ1(X)/r −g1(X)=0. Establishing that any solution U(X)satisfies limX→∞ U(X)−y(X)=g2(z∞)+z∞μ2(z∞)/r for any linear growth solution yto this FODE yields the boundary condition for U(X)(i.e., (10)). Therefore, all solutions to the optimality equation approach the same value and slope as the state grows large or small. Finally, I show that any two such solutions Uand Vcannot differ on the interior of the state space. Similar to Faingold and Sannikov (2011), if there exists an Xsuch that U(X)−V(X)>0, the structure of the optimality equation prevents these solutions from satisfying the same boundary conditions for at least one boundary. Example 1 (Product Choice, cont.). This example satisfies Assumption 4.Fromthe characterization in Section 3.2, the sequentially rational effort a(X,z)is independent of Xand the consumers’ willingness to pay b(X,z)is additively separable in (X,z). Therefore, the flow payoff g∗(X,z)=b(X,z)−a(X,z)2/2andthedriftμ∗(X,z)= a(X,z)−θX are additively separable in (X,z). From these expressions, g1(X)=λX for X>0andg1(X)=0forX<0, while μ1(X)=−θX for all X, which is monotone. Finally, σ∗(X,z)2=1 trivially satisfies Lipschitz continuity, and the growth condition is not relevant since limX→∞ μ1(X)=−∞and similarly for X→−∞. From Theorem 3, the limit slopes are z∞=rλ/(r+θ)and z−∞ =0. Therefore, equilibrium effort approaches a(X,z∞)=λ/(r+θ)as Xgrows large, which is strictly positive. As discussed in Section 2, this contrasts with settings in which effort does not have a persistent effect on quality and long-run effort converges to zero (Cripps, Mailath, and Samuelson (2004), Faingold and Sannikov (2011)). From (10), for large Xthe continuation value approximates U(X)≈rλ r+θX+(1−λ)λ r+θ+λ2 2(r+θ)2, where the first term is the payoff from repeated play of the static Nash equilibrium profile, and the second and third terms capture the impact of structural incentives on the equilibrium payoff: the equilibrium effect of persistence stemming from future strategic interaction between the firm and consumers, and the direct effect of persistence on future payoffs via the stock quality, respectively.24 In contrast, as Xapproaches −∞, equilibrium effort approaches zero and the continuation value converges to zero, limX→−∞ U(X)=0. Therefore, at large negative values of the state, incentives collapse. ♦ 24Given the expression for a(X,z)above, g2(z)=(1−λ)a(X,z)−a(X,z)2/2 and μ2(z)=a(X,z),the constant on the right hand side of (10)is(1−λ)λ/(r+θ)+λ2/2(r+θ)2. The payoff from repeated play of the static Nash equilibrium profile, y(X)=rλX/(r+θ), is calculated from the expression for y(X)in Theorem 3, using the expressions for g1(X)and μ1(X)above and φ(X)=exp(−(r/θX)dX)=X−r/θ. 470 J. Aislinn Bohren Theoretical Economics 19 (2024) 4.3.2 Bounded state space (Xcompact) When Xis compact, uniqueness follows from Assumptions 1to 3. No additional conditions are needed as in Theorem 3, as Lipschitz continuity together with the conditions on the drift and volatility that prevent the state from escaping its boundary (i.e., positive drift and zero volatility at X, and analogously for X) establish that the large player plays a unique action at the boundary and pin down a unique boundary continuation value. Theorem 4establishes uniqueness when the state space is compact, and characterizes the limit of the continuation value and the large player’s incentive constraint.25 Theorem 4. Suppose Xis compact and assume Assumptions 1to 3. For each initial state X0∈X, there exists a unique PPE that is Markov and characterized by the unique bounded solution Uof (7)on (X,X). When the boundary states are absorbing, the continuation value converges to the static Nash equilibrium payoff and intertemporal incentives collapse at the boundary, lim X→xU(X)−g∗(X,0 )=0and lim X→xμ∗X,U(X)U(X)=0 (12) for x∈{X,X}. When the boundary states are not absorbing, lim X→XU(X)=g∗(X,0 )+mu/r and lim X→X U(X)=g∗(X,0 )+mu/r (13) given unique finite limit slopes u≡limX→XU(X)and u≡limX→XU(X). The continuation value at a boundary state depends on whether the boundary state is absorbing or reflecting. When the boundary is absorbing, the state remains at the boundary once it is reached and, therefore, the continuation value converges to the static Nash payoff. When the boundary is reflecting, the limit of the continuation value also depends on its (unique) limit slope and the boundary drift, which captures how quickly the state moves away from the boundary and how the continuation value changes as the state changes. In either case, the impact of the long-run player’s action on the drift of the state converges to zero at the boundary. Therefore, incentives collapse and the equilibrium action profile converges to the static Nash action profile. This rules out the possibility of sustaining multiple equilibrium action profiles at the boundary, a key step in establishing uniqueness. An important difference from Theorem 3is that incentives collapse even if the slope of the continuation value does not converge to zero. This stems from the requirement that the boundary drift is independent of the 25Theorems 1,2, and 4also hold for an alternative version of Assumption 3when the static Nash payoff g∗(X,0 )is increasing in X. Specifically, assume that the restriction of S∗to X×[0, ∞)is nonempty, is single-valued, and returns ¯ b=δbfor some b∈B(X),whereδbis the Dirac measure on action b,S∗is Lipschitz continuous on every bounded subset of X×[0, ∞)when Xis compact and on X×[0, ∞)when X=R, and when X=R,thereexistsaδ>0suchthatforall|X|>δand z∈[0, ∞), the rate of change of g(S∗(X,z),X)+zμ(S∗(X,z),X)/r with respect to Xis monotone in Xand σ(S∗(X,z),X)is monotone in Xand constant in z. Change the definition of S∗to set S∗(X,z)=S∗(X,0 )for z<0. By Proposition 2,the solution U(X)is increasing, and the values for z<0 are irrelevant. An analogous restriction to (−∞,0 ]is possible when g∗(X,0 )is decreasing in X. Theoretical Economics 19 (2024) Dynamic moral hazard game 471 long-run player’s action—in order to maintain imperfect monitoring when the volatility is zero—combined with continuity as the drift approaches its boundary. As in Theorem 3, when incentives collapse at the boundary, this does not preclude the existence of long-run incentives, as the state does not necessarily converge to a boundary state as t→∞.Section6.2 provides an illustration of Theorem 4. 5. Properties of equilibrium payoffs The optimality equation yields rich insights into how the correspondence of PPE payoffs is tied to the underlying structure of the game. Propositions 1and 2show that the shape of the static Nash equilibrium payoff g∗(X,0 )is a key determinant of the shape of the Markov equilibrium continuation value. Note that g∗(X,0 )is straightforward to derive from the primitives of the game. Proposition 1relates the number and type of extrema for U(X)to the shape of g∗(X,0 ). Given a solution U(X)to the optimality equation, define an interval minimum of Uon a closed proper interval I⊂Xas [Xa,Xb]⊂int Isuch that U(X)=0 for all X∈[Xa,Xb],andthereexistsanε>0suchthatU(Xa)<U (X)for all X∈ (Xa−ε,Xa)∪(Xb,Xb+ε), with an analogous definition for interval maximum.26 Proposition 1. Assume Assumptions 1to 3.LetI⊂Xdenote a closed proper interval of states and let U(X)denote a linear growth or bounded (when gbounded) solution to (7). (i) If g∗(X,0 )is constant on I,thenU(X)has at most one interval extremum on I. (ii) If g∗(X,0 )is strictly monotone on I,thenU(X)has at most two interval extrema on Iand is not constant on I.Ifg∗(X,0 )is strictly increasing (decreasing) on I, and U(X)has an interval minimum [X1a,X1b]and maximum [X2a,X2b],then X1b<X 2a(X2b<X 1a). (iii) If g∗(X,0 )has ninterval extrema on I,thenU(X)has at most n+2interval extrema on I. The intuition for Proposition 1stems from the behavior of the continuation value at interior extrema. Given solution U(X), if there is an extremum at state X,thenU(X)= 0 and the optimality equation simplifies to U(X)=g∗(X,0 )+U(X)σ∗(X,0 )2/2r.If the extremum is a minimum, U(X)≥0, and, therefore, U(X)≥g∗(X,0 ). Similarly, at a maximum, U(X)≤0, and, therefore, U(X)≤g∗(X,0 ). Hence, the oscillation of U(X) is bounded by the oscillation of g∗(X,0 ). When the continuation value converges to the static Nash payoff at boundary states, then it is possible to characterize additional results on the shape of payoffs across the entire state space. Proposition 2relates the monotonicity or single-peakedness of U(X) to the monotonicity or single-peakedness of g∗(X,0 ). 26Note that since Uis twice continuously differentiable, if U(X)=0forallXin some open interval (Xa,Xb),thenU(Xa)=U(Xb)=0 and, therefore, U(X)=0onclosedinterval[Xa,Xb].Inthecaseof Xa=Xb, this definition corresponds to a strict extremum point. 472 J. Aislinn Bohren Theoretical Economics 19 (2024) Proposition 2. Assume Assumptions 1to 3and gbounded. When X=R, assume Assumption 4,andwhenXis compact, assume the boundary states {X,X}are absorbing. Let U(X)denote the unique bounded solution to (7). (i) The term g∗(X,0 )is constant on Xif and only if U(X)is constant on X. (ii) If g∗(X,0 )is monotonically increasing (decreasing) on X,thenU(X)is monotonically increasing (decreasing) on X. (iii) If g∗(X,0 )is single-peaked with a unique interval maximum (minimum) and g∗(X,0 )=g∗(X,0 )(or, in the case of X, unbounded, limX→∞ g∗(X,0 )= limX→−∞ g∗(X,0 )), then U(X)is single-peaked with a unique interval maximum (minimum). (iv) If g∗(X,0 )has Ninterval extrema on X,thenU(X)has at most Ninterval extrema on X. Applying Propositions 1and 2to specific applications will yield structural empirical predictions about how equilibrium behavior and payoffs change with the state. This is illustrated in Section 6.1 when the static Nash payoff is monotonic and in Section 6.2 when the static Nash payoff is single-peaked. Proposition 3establishes a bound on the PPE payoff across all states when the continuation value converges to the static Nash payoff at the boundary states. Let W≡supX∈XU(X)and W≡infX∈XU(X)be the least upper bound and greatest lower bound of the large player’s PPE payoff across all states, and let XHand XLdenote the sets of states that yield these payoffs (where, in a slight abuse of notation, I say ∞∈XH if limX→∞ U(X)=Wand similarly for −∞ and the case of XL). The following result shows that the smallest static Nash payoff in XHbounds the PPE payoff from above and, similarly, the largest static Nash payoff in XLbounds the PPE payoff from below.27 Proposition 3. Assume Assumptions 1to 3and gbounded. When X=R, assume Assumption 4,andwhenXis compact, assume the boundary states {X,X}are absorbing. Then the PPE payoff is bounded above (below) by the least (greatest) static Nash payoff at the states that yield the highest (lowest) PPE payoff, sup X∈XL g∗(X,0 )≤W≤W≤inf X∈XH g∗(X,0 ), where, in a slight abuse of notation, if X∈{−∞,∞},theng∗(X,0 )corresponds to limx→Xg∗(x,0 ). These bounds follow directly from the optimality equation. To see this, consider the case in which there is an interior state XHsuch that W=U(XH).ThenU(XH)= 27In general, it may be difficult to characterize XHfrom the primitives of the game, as XHdoes not necessarily correspond to the set of states that maximizes the static Nash payoff. A weaker bound that can be easily characterized is the highest static Nash payoff across all states, W≤supXg∗(X,0 ), and similarly, W≥infX∈Xg∗(X,0 ). Theoretical Economics 19 (2024) Dynamic moral hazard game 479 7. Conclusion This paper shows that persistence provides an important channel for intertemporal incentives and develops a tractable method to characterize Markov equilibrium behavior and payoffs. The tools developed in this paper will yield insights into equilibrium behavior in a broad range of settings, from industrial organization to political economy to macroeconomics. Once functional forms are specified for payoffs and the evolution of the state, it is straightforward to use Theorem 1to construct Markov equilibria. This in turn can be used to derive empirically testable comparative statics and predictions about the dynamics of equilibrium behavior based on observable features of the environment. Future research can use this framework to address design questions in specific applications, such as determining the optimal structure of persistence in a rating mechanism. Furthermore, the equilibrium characterization can be used for structural estimation. Markov equilibria have an intuitive appeal in empirical work due to their simplicity and dependence on payoff-relevant variables to structure incentives. Players do not need to condition on past behavior in a complex way, as actions and payoffs are fully determined by the current value of the state. Establishing that a Markov equilibrium exists and is unique provides a strong justification for focusing on this equilibrium concept, while the equilibrium characterization yields expressions for payoffs and actions that can calibrated and estimated. Appendix A: Proofs A.1 Proof of Lemma 1 I first show that (Vt(S))t≥0is a martingale and (Wt(S))t≥0is bounded with respect to (Xt)t≥0. Claim 1. Under Assumption 2, for any public strategy profile S=(at,bt)t≥0, initial state X0, and path of the state variable (Xt)t≥0that evolves according to (1)given S,Vt(S)is a martingale and there exists a KW>0such that |Wt(S)|≤KW(1+|Xt|)for all t≥0. Suppose gis unbounded. By Assumption 2,thereexistsak∈[0, r)and c>0 such that for all (a,b,X)∈A×E,ifX≥0, then μ(a,b,X)≤kX +c, and if X≤0, then μ(a,b,X)≥kX −c. Lipschitz continuous functions have linear growth. Therefore, by Lipschitz continuity of gand σ, the compactness of A, and the assumption that |b|≤Kb|X|+cbfor all (b,X)∈E,thereexistsaKg,Kσ,c>0suchthatforall (a,b,X)∈A×E,|g(a,b,X)|≤Kg(c k+|X|)and |σ(b,X)|≤Kσ(1+|X|). I first derive a bound on Eτ|g(at,bt,Xt)|, the expected flow payoff at time tconditional on available information at time τ≤t. This bound will be independent of the 480 J. Aislinn Bohren Theoretical Economics 19 (2024) strategy profile. Define f:X→Ras f(X)≡⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ Kgc k−Xif X≤−1 −1 8KgX4+3 4KgX2+3 8Kg+Kg c kif X∈(−1, 1) Kgc k+Xif X≥1. Note that f∈C2,f≥0, |f|≤Kg,and f(X)=⎧ ⎨ ⎩ 0if|X|≥1 3 2Kg1−X2if |X|<1. Ito’s lemma holds for any C2function. Given a strategy profile S=(at,bt)t≥0, initial state Xτ<∞, and path of the state variable (Xt)t≥τthat evolves according to (1), f(Xt)=f(Xτ)+t τf(Xs)μ(as,bs,Xs)+1 2f(Xs)σ(bs,Xs)2ds +t τ f(Xs)σ(bs,Xs)dZs ≤f(Xτ)+t τKgk|Xs|+c+3KgK2 σds +KgKσt τ1+|Xs|dZs ≤f(Xτ)+kt τ f(Xs)ds +3KgK2 σ(t−τ)+KgKσt τ1+|Xs|dZs for all t≥τ, where the first inequality follows from f(X)μ(a,b,X)≤Kg(k|X|+c), 1 2f(X)σ(b,X)2≤3KgK2 σ,andf(X)σ(b,X)z≤KgKσ(1+|X|)zfor all z∈Rand for all (a,b,X)∈A×E, and the second inequality follows from the definition of f.Theaddition of the absolute value sign in f(X)μ(a,b,X)≤Kg(k|X|+c)follows from the sign of f, and the bound on 1 2f(X)σ(b,X)2follows from f(X)σ(b,X)2=0if|X|≥1and f(X)σ(b,X)2=3 2Kg1−X2σ(b,X)2≤3 2Kg1−X2K2 σ1+|X|2≤6KgK2 σ if |X|<1. Taking expectations and noting that (1+|Xs|)is square-integrable on [τ,t], so the expectation of the stochastic integral is zero, Eτf(Xt)≤f(Xτ)+3KgK2 σ(t−τ)+kt τ Eτf(Xs)ds ≤f(Xτ)+3KgK2 σ(t−τ)ek(t−τ), where the last line follows from Gronwall’s inequality. Note that |g(a,b,X)|≤f(X)for all (a,b,X)∈A×E. Therefore, e−r(t−τ)Eτg(at,bt,Xt)≤e−r(t−τ)Eτf(Xt)≤f(Xτ)+3KgK2 σ(t−τ)e−(r−k)(t−τ). Theoretical Economics 19 (2024) Dynamic moral hazard game 481 I next show that if Xt<∞,thenWt(S)<∞, Wt(S)=Etr∞ t e−r(s−t)g(as,bs,Xs)ds ≤r∞ t e−r(s−t)Etg(as,bs,Xs)ds ≤r∞ tf(Xt)+3KgK2 σ(s−t)e−(r−k)(s−t)ds =r r−kf(Xt)+3rKgK2 σ (r−k)2, which is finite for any Xt<∞and k<r. Also, given that fhas linear growth, there exists aKW>0suchthat|Wt(S)|≤KW(1+|Xt|). By similar reasoning, E|Vt(S)|<∞for any X0<∞since EVt(S)=EEtr∞ 0 e−rsg(as,bs,Xs)ds≤Er∞ 0 e−rsg(as,bs,Xs)ds is finite for any X0<∞and k<r. Finally, EtVt+k(S)=Etrt+k 0 e−rsg(as,bs,Xs)ds +e−r(t+k)Wt+k(S) =rt 0 e−rsg(as,bs,Xs)ds +Etrt+k t e−rsg(as,bs,Xs)ds +e−r(t+k)Et+kr∞ t+k e−r(s−(t+k))g(as,bs,Xs)ds =rt 0 e−rsg(as,bs,Xs)ds +e−rtWt(S)=Vt(S). Taken together, this implies Vt(S)is a martingale and establishes Claim 1 for the case of gunbounded. If gis bounded, then trivially, Wt(S)<∞and E|Vt(S)|<∞for all t≥0 and X0∈X, and only the final step is needed to establish the claim. I next derive the evolution of the continuation value. This part of the proof follows from almost identical reasoning to the proof of Proposition 2 in Faingold and Sannikov (2011). The derivative of Vt(S)with respect to tis: dVt(S)=re−rtg(at,bt,Xt)dt −re−rtWt(S)dt +e−rt dWt(S). By the martingale representation theorem (Karatzas and Shreve (1991)), there exists a progressively measurable process (βt)t≥0such that Vtcan be represented as dVt(S)= 482 J. Aislinn Bohren Theoretical Economics 19 (2024) re−rtβtσ(bt,Xt)dZt. Combining these two expressions for dVt(S)yields the law of motion for the continuation value, dWt(S)=rWt(S)−g(at,bt,Xt)dt +rβtσ(bt,Xt)dZt =rWt(S)−g(at,bt,Xt)dt +rβtdXt−μ(at,bt,Xt)dt, where βtcaptures the sensitivity of the continuation value to the state variable. As shown above, any continuation value has linear growth with respect to Xtand is bounded when gis bounded. Finally, I establish sequential rationality. This part of the proof follows from almost identical reasoning to the proof of Proposition 3 in Faingold and Sannikov (2011). Consider strategy profile (at,bt)t≥0played from period τonward and alternative strategy ( at,bt)t≥0played up to time τ. Recall that all values of Xtare possible under both strategies, but that each strategy induces a different measure over sample paths (Xt)t≥0.At time τ, the state variable is equal to Xτ.Actionaτwill induce dXτ=μ(aτ,bτ,Xτ)dt + σ(bτ,Xτ)dZτ, whereas action aτwill induce dXτ=μ( aτ,bτ,Xτ)dt +σ(bτ,Xτ)dZτ.Let  Vτbe the expected average payoff conditional on information at time τwhen the large player follows  aup to τand aafterward, and let Wτbe the continuation value when the large player follows strategy (at)t≥0starting at time τ:  Vτ=rτ 0 e−rsg( as,bs,Xs)ds +e−rτWτ. Consider changing τso that the large player plays strategy ( at,bt)for another instant: d Vτis the change in average expected payoffs when the large player switches to (at)t≥0 at τ+dτ instead of τ. When the large player switches strategies at time τ, d Vτ=re−rτg( aτ,bτ,Xτ)−Wτdτ +e−rτ dWτ =re−rτg( aτ,bτ,Xτ)−g(aτ,bτ,Xτ)dτ +re−rτβτdXτ−μ(aτ,bτ,Xτ)dτ =re−rτg( aτ,bτ,Xτ)−g(aτ,bτ,Xτ)+βτμ( aτ,bτ,Xτ)−βτμ(aτ,bτ,Xτ)dτ +re−rτβτσ(bτ,Xτ)dZτ. There are two components to this strategy change: how it affects the immediate flow payoff and how it affects the future state Xt, which impacts the continuation value. The profile ( at,bt)t≥0yields the large player a payoff of  W0=E0[ V∞]=E0 V0+∞ 0 d Vt =W0+E0r∞ 0 e−rtg( at,bt,Xt)+βtμ( at,bt,Xt)−g(at,bt,Xt) −βtμ(at,bt,Xt)dt. If g(at,bt,Xt)+βtμ(at,bt,Xt)≥g( at,bt,Xt)+βtμ( at,bt,Xt) Theoretical Economics 19 (2024) Dynamic moral hazard game 483 holds for all t≥0, then W0≥ W0and deviating to S=( at,bt)is not a profitable deviation. A strategy (at)t≥0is sequentially rational for the large player if, given (βt)t≥0,forallt, at∈argmax a∈Ag(a,bt,Xt)+βtμ(a,bt,Xt). A.2 Proof of Theorem 1 In a Markov equilibrium, the continuation value and equilibrium actions are characterized as a function of the state variable as Wt=U(Xt),a∗ t=a(Xt),andb∗ t=b(Xt).ByIto’s formula, if a Markov equilibrium with a twice continuously differentiable continuation value exists, the continuation value will evolve according to dU(Xt)=U(Xt)dXt+1 2U(Xt)σb∗ t,Xt2dt =U(Xt)μa∗ t,b∗ t,Xtdt +1 2U(Xt)σb∗ t,Xt2dt +U(Xt)σb∗ t,XtdZt. (16) Similar to the derivation of a Markov equilibrium in Faingold and Sannikov (2011), matching the drift of (16) with the drift of the continuation value characterized in (5) yields the optimality equation U(X)=2rU(X)−ga(X),b(X),X σb(X),X2−2μa(X),b(X),XU(X) σb(X),X2, (17) which is a second-order nonhomogenous differential equation, and matching the volatilities characterizes the process governing incentives, rβt=U(Xt). Substituting this expression into the condition for sequential rationality characterized in (6)yields the Markovian action profile (a(X),b(X)) =S∗(X,U(X)) (by Assumption 3,S∗is single-valued.) Plugging this into (17)yields(7). I first establish that (7) has at least one solution U∈C2that takes on values in the interval of feasible payoffs for the large player. In the case of an unbounded state space, Theorem 5.6 from De Coster and Habets (2006) gives sufficient conditions for the existence of a solution to a second-order differential equation defined on R3.Iconstruct upper and lower solutions to (17) at action profile S∗(X,U(X)) to show that these conditions are satisfied. This leads to the following lemma, which is the innovative part of this proof and is proven in Supplemental Appendix B. Lemma 2. If X=R,then(7)has at least one solution U∈C2on Xthat lies in the range of feasible payoffs for the large player. In the case of a bounded state space, I use an extension of a standard existence result from De Coster and Habets (2006), which was developed in Faingold and Sannikov (2011). The extension is necessary because (7) is undefined at the boundary of the state space, {X,X}. This leads to the following lemma, which is proven in Supplemental Appendix B. 484 J. Aislinn Bohren Theoretical Economics 19 (2024) Lemma 3. If Xis compact, then (7)has at least one solution U∈C2on (X,X)that lies in the range of feasible payoffs for the large player. Finally, I construct a Markov equilibrium that yields payoff U(X0),whereUis a solution to (7). The function X→ S∗(X,U(X)) is Lipschitz continuous, as are X→ μ∗(X,U(X)) and X→ σ∗(X,U(X)). Therefore, the state variable starts at X0and evolves according to the unique strong solution (Xt)t≥0to the stochastic differential equation dXt=μ∗Xt,U(Xt)dt +σ∗Xt,U(Xt)dZt. Moreover, dU(Xt)=U(Xt)μ∗Xt,U(Xt)dt +1 2U(Xt)σ∗Xt,U(Xt)2dt +U(Xt)σ∗Xt,U(Xt)dZt =rU(Xt)−g∗Xt,U(Xt)dt +U(Xt)σ∗Xt,U(Xt)dZt and, therefore, the process of continuation values Wt=U(Xt)satisfies (5)withprocess of incentive weights βt=U(Xt)/r. Finally, the strategy profile (a∗ t,b∗ t)t≥0satisfies (6) given (βt)t≥0with βt=U(Xt)/r. Therefore, (a∗ t,b∗ t)t≥0is a PPE yielding equilibrium payoff U(X0). A.3 Proof of Theorem 2 Let Ube the linear growth (when gis unbounded) or bounded (when gis bounded) solution to (7) that yields the highest MPE payoff at X0. Suppose there exists an initial state X0∈Xand a PPE strategy profile S=(at,bt)t≥0that yields an equilibrium payoff W0>U(X0). In such a PPE, the state (Xt)t≥0evolves according to (1)givenS=(at,bt)t≥0 and, by Lemma 1, the continuation value evolves according to dWt(S)=rWt(S)−g(at,bt,Xt)dt +rβtdXt−μ(at,bt,Xt)dt(18) for some process (βt)t≥0. By Assumption 3, a unique action profile satisfies (6) at each (X,rβ). Therefore, by Lemma 1, equilibrium actions satisfy (at,bt)=S∗(Xt,rβt).By Ito’s formula, the process (U(Xt))t≥0evolves according to dU(Xt)=U(Xt)μ∗(Xt,rβt)dt +1 2U(Xt)σ∗(Xt,rβt)2dt +U(Xt)σ∗(Xt,rβt)dZt. (19) DefineaprocessDt≡Wt(S)−U(Xt)with initial condition D0=W0(S)−U(X0)>0. Then Dtevolves according to dDt=dWt(S)−dU(Xt). Plugging in Eqs. (18)and(19), the process has volatility f(Xt,βt),wheref(X,β)≡(rβ −U(X))σ∗(X,rβ),andhas Theoretical Economics 19 (2024) Dynamic moral hazard game 485 drift rDt+d(Xt,βt),where d(X,β)≡rU(X)−g∗(X,rβ)−U(X)μ∗(X,rβ)−U(X)σ∗(X,rβ)2/2 =rg∗X,U(X)−g∗(X,rβ)+U(X)μ∗X,U(X)−μ∗(X,rβ) +U(X)σ∗X,U(X)2−σ∗(X,rβ)2/2, and the second line follows from substituting the right hand side of (7)forU(X). Lemma 4. If f(X,β)=0and σ(X,rβ)>0,thend(X,β)=0. Proof. Suppose f(X,β)=0forsome(X,β)and σ(X,rβ)>0. Then rβ =U(X).The action profile associated with S∗(X,U(X)) corresponds to the actions played in the Markov equilibrium with continuation value U(X)at state X. Therefore, d(X,β)=0. Lemma 5. For every ε>0,thereexistsaη>0such that either d(X,β)>−εor |f(X,β)|>η. Proof. Suppose the state space is unbounded, X=R. Note that in this case, σ∗(X,rβ) is bounded away from 0 by Assumption 1, so, by Lemma 4,iff(X,β)=0, then d(X,β)= 0. First show that there exists an M>0 such that this is true for (X,β)∈a≡{X×R: |β|>M}. Since Uis bounded by Lemma 9inthecaseofgunbounded or Lemma 26 in Supplemental Appendix D.2 in the case of gbounded (note that neither lemma requires Assumption 4), and since σ∗(X,rβ)is bounded away from 0, there exists an M>0and η1>0suchthat|f(X,β)|>η 1for all |β|>Mand X∈X, regardless of d. Next show that there exists an δ>0 such that this is true for (X,β)∈b≡{X×R:|β|≤M,|X|>δ }. Consider the set b⊂bwith d(X,β)≤−ε. Itmustbethatβis bounded away from U(X)/r on b. Suppose not. Then either (i) there exists some (X,β)∈bwith β=U(X)/r, which implies f(X,β)=0 and, therefore, d(X,β)=0—a contradiction— or (ii) as Xbecomes large, the boundary of the set bapproaches β=U(X)/r. The latter implies that for any δ1>0, there exists an (X,β)∈bwith rβ −U(X)<δ 1. Choose δ1so that |g∗(X,U(X)) −g∗(X,rβ)|<ε/4r,|U(X)||μ∗(X,U(X)) −μ∗(X,rβ)|<ε/4, and |U(X)||σ∗(X,U(X))2−σ∗(X,rβ)2|=0, which is possible given that g∗and μ∗ are Lipschitz, Uis bounded, and σ∗is independent of zfor large X.Then|d(X,β)|< ε/4+ε/4+ε/4=3ε/4, which is a contradiction. Therefore, there exists a η2such that |f(X,β)|>η 2on b. Thenonthesetb,ifd(X,β)≤−ε,then|f(X,β)|>η 2.Finally show this is true for (X,β)∈c≡{X×R:|β|≤Mand |X|≤δ}. Consider the set c⊂cwhere d(X,β)≤−ε.Thefunctiondis continuous and cis compact, so cis compact. The function |f|is continuous and, therefore, achieves a minimum η3on c. If η3=0, then d=0 by Lemma 4—a contradiction. Therefore, η3>0and|f(X,β)|>η 3 for all (X,β)∈c.Takeη≡min{η1,η2,η3}. Then when d(X,β)≤−ε,|f(X,β)|>η. The proof for a bounded state space is analogous (see Supplemental Appendix C). Lemma 6. Given X0, any PPE payoff W0is such that W0≤U(X0). 486 J. Aislinn Bohren Theoretical Economics 19 (2024) Proof. Choose ε=rD0/4 and suppose Dt≥D0/2. Then, by Lemma 5,thereexistsa η>0 such that whenever the drift of Dtis less than rDt−ε>rD 0/2−rD0/4=rD0/4>0, |f(Xt,βt)|>η. Thus, as long as Dt≥D0/2>0, it has either positive drift or positive volatility. This implies it grows arbitrarily large with positive probability, irrespective of Xt. This is a contradiction, since in the case that gis unbounded, by Lemma 1,Dtis the difference of two processes that are bounded with respect to Xt, and in the case that gis bounded, Dtis the difference of two bounded processes. Thus, it cannot be that D0>0 anditmustbethecasethatW0≤U(X0). Letting Ube the linear growth (when gis unbounded) or bounded (when gis bounded) solution to (7) that yields the lowest MPE payoff at X0, by analogous reasoning it is not possible to have D0=W0(S)−U(X0)<0, implying W0≥U(X0). The proof of Theorem 2immediately follows from Lemma 6, the analogue for W0≥U(X0),andthe fact that at any state X∈X, it is possible for the large player to achieve any payoff in the convex hull of the set of Markov equilibrium payoffs at state Xby randomization at time zero. Proof of Corollary 1The existence of a Markov equilibrium follows from Theorem 1. When μis independent of a, the sequential rationality condition (6)inaMarkovequilibrium collapses to maximizing the static flow payoff, and the large player plays the unique static Nash action profile S∗(X,0 )in each state. Therefore, any solution to (7) must satisfy U(Xt)=Etr∞ t e−rsg∗(Xs,0 )dt, (20) where the measure over the state is independent of the solution Usince equilibrium actions are independent of U. Given that the right hand side of (20) is independent of U,(7) must have a unique solution and there is a unique Markov equilibrium. By Theorem 2, this is also the unique PPE. The solution to (7) evaluated at state Xtanalytically characterizes (20). A.4 Proof of Theorems 3and 4 IproveTheorems3and 4simultaneously. The proof proceeds in three steps: Step 1. Any solution to the optimality equation has the same boundary conditions. Step 2. If all solutions have the same boundary conditions, then there is a unique linear growth (bounded) solution. Step 3. When there is a unique solution, then there is a unique PPE. Let ψ(X,z)≡g∗(X,z)+zμ∗(X,z)/r be the value of the large player’s incentive constraint at the sequentially rational action profile for incentive weight z/r. All intermediate theorems and lemmas maintain Assumptions 1to 3and, as stated, Assumption 4.As a reminder, |·|denotes the Euclidean norm for vectors. I first present an intermediate result that will be used in Steps 1 and 2. Theoretical Economics 19 (2024) Dynamic moral hazard game 487 Lemma 7. Suppose Uand Vare both linear growth (bounded) solutions to (7),with U(X)<V(X)for some interior state X∈X.ThenV−Udoes not have an interior maximum and is monotone for large |X|. Proof. First suppose Xis compact. It follows from identical reasoning to Lemma C.7 in Faingold and Sannikov (2011)thatifUand Vare two linear growth (bounded) solutions of (7)suchthatU(X0)≤V(X0)and U(X0)≤V(X0), with at least one strict inequality, then U(X)<V(X)and U(X)<V(X)for all X∈(X0,X).31 Similarly if U(X0)≤V(X0) and U(X0)≥V(X0), with at least one strict inequality, then U(X)<V(X)and U(X)> V(X)for all X∈(X,X0). Suppose Uand Vare both bounded solutions to (7), with U(X)<V(X)for some X∈(X,X). Suppose V−Uhas an interior maximum at some X∗∈(X,X).Thenby continuity, this implies that U(X∗)=V(X∗).IfU(X∗)<V(X∗), then by the above statement, U(X)<V(X)for all X>X ∗and, therefore, V(X)−U(X)is strictly increasing for X>X ∗. This contradicts that X∗is an interior maximum. If U(X∗)>V(X∗), then by the above statement, U(X)>V(X)and U(X)>V(X)for all X>X ∗,and U(X)>V(X)and U(X)<V(X)for all X<X ∗. Therefore, X∗is a global maximum. This contradicts U(X)<V(X)for some X∈(X,X). Therefore, V−Udoes not have an interior maximum. Given this, V−Uhas at most one interior minimum. Therefore, there exists a δ>0suchthatV−Uis monotone for |X−X|<δand |X−X|<δ.The proof for the case of X=Ris analogous, replacing Xand Xwith ∞and −∞, respectively. Step 1: Boundary conditions Lemmas 8to 19 as well as Lemmas 26 and 27 in Supplemental Appendix D.2 establish the following boundary conditions for the case of X=R. When gis unbounded, any solution Uof (7) with linear growth satisfies limX→pU(X)− yL(X)=g2(zp)+zpμ2(zp)/r,limX→pU(X)=zp,andlimX→pσ(X,U(X))2U(X)=0 for p∈{−∞,∞},wherezp≡rgp/(r−μp)given μp≡limX→pμ∗(X,z)/X and gp≡ limX→pg∗(X,z)/X, which exist and are finite, and yL(x)≡−f(x)rg1(x)/f (x)μ1(x)dx with integrating factor f(x)≡exp(r/μ1(x)dx)when limx→pμ1(x)= 0andyL(x)≡ g1(x)when limx→pμ1(x)=0. When gis bounded, this simplifies to limX→pU(X)=gp, where gp≡limX→pg∗(X,0 ),andlimX→pU(X)=0. Supplemental Appendix D.1 establishes analogous boundary conditions for the case of Xcompact, and Supplemental Appendix D.3 establishes the same boundary conditions for the case of X=Rand g bounded under an alternative to Assumption 4. Define ψ(X,z)≡ψ(X,z)/X and U(X)≡U(X)/X.Letψand ψdenote the partial derivatives of ψand ψwith respect to X.Letδ0>0 denote the lower bound above which the large |X|properties of Assumptions 3and 4hold. Several lemmas use the property that g∗(X,z),μ∗(X,z),andσ∗(X,z)are bounded in z, which follows from the compactness of Aand B(X). The Lipschitz continuity of g1,μ1,g2,andμ2is also used, which follows from the Lipschitz continuity of g∗(X,z)and μ∗(X,z). The following series of lemmas are stated for an unbounded flow payoff g;toapplythemtoabounded 31Analogous to the definition of φ1in their result, set X1≡inf{X∈[X0,X):U(X)≥V(X)}and apply the same reasoning. 488 J. Aislinn Bohren Theoretical Economics 19 (2024) flow payoff, simply substitute “bounded solution to (7)” for “linear growth solution to (7)” throughout. Lemma 8. Suppose X=R.Givenp∈{−∞,∞},μp≡limX→pμ∗(X,z)/X and gp≡ limX→pg∗(X,z)/X exist and are finite. Moreover, limX→pψ(X,z)=limX→pψ(X,z)= ψp(z)for all z∈R,whereψp(z)≡gp+zμp/r. Proof.Letp=∞and fix z∈R. Given Assumption 4(i), ψ(X,z)=g 1(X)+zμ 1(X)/r for X>δ 0. By the Lipschitz continuity of g1and μ1,g 1and μ 1are bounded, and, therefore, ψ(·,z)is bounded for any z∈R. By Assumption 3,ψ(·,z)and g 1are monotone for large X(the latter follows from the assumption holding at z=0). Therefore, by the monotone convergence theorem, ψ∞(z)≡limX→∞ ψ(X,z)and g∞≡limX→∞ g 1(X) exist and are finite. Given that ψand g 1have well defined limits and ψ(X,z)=g 1(X)+ zμ 1(X)/r for large X,μ∞≡limX→∞ μ 1(X)exists and is finite. Moreover, ψ∞(z)= g∞+zμ∞/r.Wheng1and μ1are unbounded, then by l’Hopital’s rule, limX→∞ ψ(X,z)= ψ∞(z),limX→∞ g1(X)/X =g∞,andlimX→∞ μ1(X)/X =μ∞.Inthecasewhereg1or μ1is bounded, this immediately follows from g∞=0orμ∞=0. Given that g2(z)and μ2(z)are independent of X,limX→∞ g2(z)/X =0andlimX→∞ μ2(z)/X =0. This implies limX→∞ g∗(X,z)/X =g∞and limX→∞ μ∗(X,z)/X =μ∞.Notethatμ∞<rby Assumption 2. The proof for p=−∞is analogous. Lemma 9. Suppose X=Rand Uis a solution of (7)with linear growth. Then for p∈ {−∞,∞}, there exists a finite U p∈Rsuch that limX→pU(X)=limX→pU(X)=U p. Proof.Letp=∞ and let Ube a solution of (7) with linear growth. Suppose liminfX→∞ U(X)= lim supX→∞ U(X).Thenforallδ>0, by the continuity of U,there exists a zand an increasing sequence (Xn)n∈Nof alternating consecutive Xsuch that X1>δ,U(Xn)=z,andU(Xn)≤0fornodd, and U(Xn)=zand U(Xn)≥0forn even, with one inequality for U strict. From (7), this implies U(Xn)≤ψ(Xn,z)for n odd and ψ(Xn,z)≤U(Xn)for neven. Thus, the oscillation of ψ(X,z)is at least as large as the oscillation of U. But by Assumption 3,ψ(X,z)is monotone for X>δ 0.Therefore,itmustbethatlim infX→∞ U(X)=lim supX→∞ U(X).LetU ∞denote this limit. Given Uhas linear growth, |U ∞|<∞and, by l’Hopital’s rule, limX→∞ U(X)=U ∞.In thecaseofgbounded, Ubounded implies U ∞=0andlimX→∞ U(X)=0. The proof for p=−∞is analogous. Lemma 10. Suppose X=Rand Uisasolutionof (7)with linear growth. Then limX→pψ(X,U(X)) =ψp(U p)for p∈{−∞,∞},whereU p≡limX→pU(X). Proof.Letp∈{−∞,∞}and let Ubeasolutionof(7) with linear growth. Given μ∗and g∗are Lipschitz continuous and additively separable in (X,z)for |X|>δ 0,thereexistsa M1,M2,M3,c>0andδ>δ 0such that for |X|>δ, ψ(X,z1)−ψ(X,z2)≤M1|z1−z2|+M2|z1||z1−z2|+M3|z1−z2||X|+|z2|. Theoretical Economics 19 (2024) Dynamic moral hazard game 495 Proof of Proposition 2.LetUbe the unique bounded solution to (7). At a state X corresponding to an interior extremum on X,U(X)=0. From (7), if Xis an interior minimum, g∗(X,0 )≤U(X), and if Xis an interior maximum, U(X)≤g∗(X,0 ).Letn denote the number of (strict interior) interval extrema of Uon Xand let ngdenote the number of (strict interior) interval extrema of g∗(X,0 )on X. First consider Xcompact. Item (i). Suppose g∗(·,0 )is constant on X.Thenng=0andthereexistsac∈Rsuch that g∗(X,0 )=cfor all X∈X. By part (iv) (see below for proof), ng=0impliesn=0. By the boundary conditions from Theorem 4,U(X)=cand U(X)=c, which implies U(X)=U(X). Combined with n=0, this implies that Uis constant on X. To establish the other direction, suppose g∗(·,0 )is not constant on X. Then there exists a proper interval I1⊂Xsuch that g∗(·,0 )is strictly monotone on I1. Take a closed proper subset I2⊂I1. By Proposition 1(ii), Uis not constant on I2. Therefore, Uis not constant on X. Item (ii). Suppose g∗(·,0 )is monotonically increasing on Xand Uis not monotonically increasing. Then ng=0andU(X)<0forsomeX∈X. By Proposition 2(iv) (see below for proof), ng=0impliesn=0. Therefore, it must be that Uis monotonically decreasing on X, i.e., U(X)≤0forallX∈X.GivenU(X)<0forsomeX∈X, this implies that U(X)<U(X). By the boundary conditions from Theorem 4,U(X)=g∗(X,0 )and U(X)=g∗(X,0 ), and by the monotonicity of g∗(·,0 ),g∗(X,0 )≤g∗(X,0 ). This implies U(X)≤U(X), a contradiction. Therefore, Uis monotonically increasing. The proof for Umonotonically decreasing is analogous. Item (iii). Suppose g∗(X,0 )=g∗(X,0 )and suppose g∗(·,0 )is single-peaked with a unique interval extremum, a maximum. Let Xg 1denote the interval of states corresponding to this extremum. By Proposition 2(iv), ng=1impliesn≤1. Given that there is a unique interval extremum and it is a maximum, g∗(·,0 )is monotonically increasing on I1=[X,infXg 1]and strictly so on some proper interval I 1⊂I1. Therefore, by Proposition 1(ii), Uis not constant on I1. Similarly, g∗(·,0 )is monotonically decreasing on I2=[supXg 1,X]and strictly so on some proper interval I 2⊂I2,soUis not constant on I2. From the boundary conditions, U(X)=g∗(X,0 )and U(X)=g∗(X,0 ). Therefore, U(X)=U(X). Since Uis not constant and U(X)=U(X), by continuity Umust have at least one interval extremum, n≥1. Given that it was already established that n≤1, it must be that n=1. Suppose the unique interval extremum for Uis a minimum. Let X1denote the interval of states corresponding to this extremum. Then g∗(x1,0 )≤U(x1)for all x1∈X1. Given that X1is a minimum and it is the unique interval extremum, Uis monotonically decreasing on [X,inf X1]and strictly so on some some proper interval I⊂[X,inf X1]. This implies that for all x1∈X1,U(x1)<U (X)=g∗(X,0 ). Therefore, for all x1∈X1, g∗(x1,0 )≤U(x1)<g ∗(X,0 ). Further, since Xg 1is the unique interval extremum of g∗(·,0 )and a maximum, for all xg 1∈Xg 1,g∗(X,0 )=g∗(X,0 )<g ∗(xg 1,0 ).Buttheng∗(·,0 ) must have two interval extrema, since g∗(x1,0 )<g ∗(X,0 )=g∗(X,0 )for x1∈X1and g∗is continuous. This is a contradiction. Therefore, Uis single-peaked with a unique interval maximum. The proof for Usingle-peaked with a minimum is analogous. Item (iv). This follows directly from U(X)≥g∗(X,0 )at an interior minimum of U, U(X)≤g∗(X,0 )at an interior maximum of U,U(X)=g∗(X,0 ),U(X)=g∗(X,0 ),and the Lipschitz continuity of g∗. 496 J. Aislinn Bohren Theoretical Economics 19 (2024) For the case of X=R,replaceg∗(X,0 )and U(X)with limX→∞ g∗(X,0 )and limX→∞ U(X), and analogously for X. These limits exist by the proof of Theorem 4. Proof of Proposition 3.LetUbe the unique bounded solution to (7). Then Uis continuous and bounded on a closed set. Therefore, either Uattains a global maximum on X, in which case W=U(XH)for some XH∈X,orinthecasewhereXis unbounded, it is also possible that W=lim supX→XHU(X)for either XH=−∞or XH=∞.Suppose Uattains a global maximum at an interior state XH∈int X.ThenU(XH)=0and U(XH)≤0. From (7), this implies U(XH)=2rW−g∗(XH,0 ) σ∗(XH,0 )2≤0, and, therefore, W≤g∗(XH,0 ).IfXis bounded and Uattains a global maximum at boundary state XH=Xor XH=X, then by Theorem 4,W=g∗(XH,0 ). Similarly, if X is unbounded and limX→XHU(X)=Wfor either XH=−∞or XH=∞, then by Theorem 4,W=g∗(XH,0 ). Therefore, W≤infXH∈XHg∗(XH,0 ). The proof for Wis analogous. References Abreu, Dilip, Paul Milgrom, and David Pearce (1991), “Information and timing in repeated partnerships.” Econometrica, 59, 1713–1733. [450,452,455] Abreu, Dilip, David Pearce, and Ennio Stacchetti (1990), “Toward a theory of discounted repeated games with imperfect monitoring.” Econometrica, 58, 1041–1063. [458] Board, Simon and Moritz Meyer-ter-Vehn (2013), “Reputation for quality.” Econometrica, 81, 2381–2462. [453] Bohren, J. Aislinn (2016), “Using persistence to generate incentives in a dynamic moral hazard problem.” PIER Working Paper 16-024. [462] Cisternas, Gonzalo (2018), “Two-sided learning and the ratchet principle.” The Review of Economic Studies, 85, 307–351. [453] Cripps, Martin, George Mailath, and Larry Samuelson (2004), “Imperfect monitoring and impermanent reputations.” Econometrica, 72, 407–432. [452,455,469] De Coster, Colette and Patrick Habets (2006), Two-Point Boundary Value Problems: Lower and Upper Solutions. Elsivier. [483] Dilmé, Francesc (2019), “Reputation building through costly adjustment.” Journal of Economic Theory, 181, 586–626. [452] Doraszelski, Ulrich and Mark Satterthwaite (2010), “Computable Markov-perfect industry dynamics.” RAND Journal of Economics, 41, 215–243. [453] Duffie, Darrell, Andreu Mas-Colell John Geanakoplos, and Andrew McLennan (1994), “Stationary Markov equilibria.” Econometrica, 62, 745–781. [454] Theoretical Economics 19 (2024) Dynamic moral hazard game 497 Dutta, Prajit K. (1995), “A folk theorem for stochastic games.” Journal of Economic Theory, 66, 1–32. [453] Dutta, Prajit K. and Rangarajan Sundaram (1992), “Markovian equilibrium in a class of stochastic games: Existence theorems for discounted and undiscounted models.” Economic Theory, 2, 197–214. [454] Ekmekci, Mehmet (2011), “Sustainable reputations with rating systems.” Journal of Economic Theory, 146, 479–503. [452] Ericson, Richard and Ariel Pakes (1995), “Markov-perfect industry dynamics: A framework for empirical work.” The Review of Economic Studies, 62, 53–82. [453] Faingold, Eduardo and Yuliy Sannikov (2011), “Reputation in continuous time games.” Econometrica, 79, 773–876. [450,452,453,454,455,458,460,461,462,464,465,466,468, 469,481,482,483,487] Faingold, Eduardo (2020), “Reputation and the Flow of Information in Repeated Games.” Econometrica, 88, 1697–1723. [452] Fudenberg, Drew and David Levine (1989), “Reputation and equilibrium selection in games with a patient player.” Econometrica, 57, 759–778. [452] Fudenberg, Drew and David Levine (1992), “Maintaining a reputation when strategies are imperfectly observed.” Review of Economic Studies, 59, 561–579. [452] Fudenberg, Drew and David Levine (2007), “Continuous time limits of repeated games with imperfect public monitoring.” Review of Economic Dynamics, 10, 173–192. [452, 453,454,462,466] Fudenberg, Drew and David Levine (2009), “Repeated games with frequent signals.” Quarterly Journal of Economics, 233–265. [452,454,462,466] Fudenberg, Drew, David Levine, and Eric Maskin (1994), “The folk theorem with imperfect public information.” Econometrica, 62, 997–1039. [453] Fudenberg, Drew and Yuichi Yamamoto (2011), “The folk theorem for irreducible stochastic games with imperfect public monitoring.” Journal of Economic Theory, 146, 1664–1683. [453] Hörner, Johannes, Takuo Sugaya, Satoru Takahashi, and Nicolas Vieille (2011), “Recursive methods in discounted stochastic games: An algorithm and a folk theorem.” Econometrica, 79, 1277–1318. [453] Karatzas, Ioannis and Steven Shreve (1991), Brownian Motion and Stochastic Calculus. Springer-Verlag, New York. [481] Kreps, David, Paul Milgrom, John Roberts, and Robert Wilson (1982), “Rational cooperation in the finitely dilemma repeated Prisoners’ dilemma.” Journal of Economic Theory, 27, 245–252. [452] Kreps, David and Robert Wilson (1982), “Reputation and imperfect information.” Journal of Economic Theory, 27, 253–279. [452] 498 J. Aislinn Bohren Theoretical Economics 19 (2024) Mailath, George J. and Larry Samuelson (2001), “Who wants a good reputation?” Review of Economic Studies, 68, 415–441. [452,473] Milgrom, Paul and John Roberts (1982), “Predation, reputation and entry deterrence.” Journal of Economic Theory, 27, 280–312. [452] Nowak, Andrzej S. and T. E. S. Raghavan (1992), “Existence of stationary correlated equilibria with symmetric information for discounted stochastic games.” Math. Oper. Res., 17, 519–526. [454] Sannikov, Yuliy (2007), “Games with imperfectly observable actions in continuous time.” Econometrica, 75, 1285–1329. [458] Sannikov, Yuliy and Andrzej Skrzypacz (2007), “Impossibility of collusion under imperfect monitoring with flexible production.” American Economic Review, 97, 1794–1823. [450,452,455,462] Sannikov, Yuliy and Andrzej Skrzypacz (2010), “The role of information in repeated games with frequent actions.” Econometrica, 78, 847–882. [450,462] Shapley, Lloyd S. (1953), “Stochastic games.” Proceedings of the National Academy of Sciences, 39, 1095. [454] Strulovici, Bruno and Martin Szydlowski (2015), “On the smoothness of value functions and the existence of optimal strategies.” Journal of Economic Theory, 159, 1016–1055. [458] Co-editor Thomas Mariotti handled this manuscript. Manuscript received 2 November, 2016; final version accepted 29 October, 2019; available online 3 March, 2023.