Subgame-perfect equilibrium in games with almost perfect information: Dispensing with public randomization
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Barelli, Paulo; Duggan, John Article Subgame-perfect equilibrium in games with almost perfect information: Dispensing with public randomization Theoretical Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Barelli, Paulo; Duggan, John (2021) : Subgame-perfect equilibrium in games with almost perfect information: Dispensing with public randomization, Theoretical Economics, ISSN 1555-7561, The Econometric Society, New Haven, CT, Vol. 16, Iss. 4, pp. 1221-1248, https://doi.org/10.3982/TE3243 This Version is available at: https://hdl.handle.net/10419/253456 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Theoretical Economics 16 (2021), 1221–1248 1555-7561/20211221 Subgame-perfect equilibrium in games with almost perfect information: Dispensing with public randomization Paulo Barelli Department of Economics, University of Rochester John Duggan Department of Political Science and Department of Economics, University of Rochester Harris, Reny, and Robson (1995) added a public randomization device to dynamic games with almost perfect information to ensure existence of subgame perfect equilibria (SPE). We show that when Nature’s moves are atomless in the original game, public randomization does not enlarge the set of SPE payoffs: any SPE obtained using public randomization can be “decorrelated” to produce a payoffequivalent SPE of the original game. As a corollary, we provide an alternative route to a result of He and Sun (2020) on existence of SPE without public randomization, which in turn yields equilibrium existence for stochastic games with weakly continuous state transitions. Keywords. Existence, subgame-perfect equilibrium, infinite-action games, stochastic games, public randomization. JEL classification. C72, C73. 1. Introduction A seminal result of Harris, Reny, and Robson (1995) (henceforth, HRR) ensures existence of subgame perfect equilibrium (SPE) in dynamic games with almost perfect information by augmenting such games with a public randomization device. That is, they assume that in addition to Nature’s moves in the original game, players observe a uniformly distributed, payoff-irrelevant public signal in every stage. This convexifies equilibrium probabilities over continuation paths in the extended game and allows them to use limiting arguments to deduce existence of a SPE; in the original game, their construction corresponds to a generalized strategy profile in which players’ actions are marked by a form of correlation. We focus on the subclass of games with atomless moves by Nature, and our main contribution is that in such games, one can dispense with public signals in HRR’s result: each SPE obtained by augmenting the original game to allow public signals can be “decorrelated” to produce a payoff-equivalent SPE of the original game involving no correlation or public signals. Therefore, for a large class of dynamic games, public randomization is without loss of generality; or put differently, in a world Paulo Barelli: [email protected] John Duggan: [email protected] We thank the three anonymous referees for their valuable feedback on the paper. ©2021 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at https://econtheory.org.https://doi.org/10.3982/TE3243
1222 Barelli and Duggan Theoretical Economics 16 (2021) with atomless moves by Nature, introducing “sunspots” does not enlarge the set of subgame perfect equilibrium payoffs. In particular, it cannot enlarge the empty set to a nonempty set, so a direct corollary is existence of SPE in games with atomless moves by Nature, providing an alternative route to existence of SPE to that recently taken by He and Sun (2020). The class of games with atomless moves by Nature is general enough to capture many applications of interest; in particular, it subsumes stochastic games with weakly continuous, atomless state transitions—in some respects, going well beyond the analysis in the classical literature. More formally, a stochastic game is played among nplayers; in each period, a state variable zis publicly observed; players then simultaneously choose actions yi;andgiventhestatezand action profile y=(y1,,yn), next period’s state is drawn from a transition probability μ(·|y,z). This process is repeated in discrete time over an infinite horizon. In their classic analysis, Mertens and Parthasarathy (2003) did not assume that states are atomlessly distributed, but their conditions for existence of SPE impose the strong condition of norm-continuity on the state transition. Letting Zdenote the set of states and (Z)the set of probability measures over states, the assumption of Mertens and Parthasarathy (which is standard in the literature) is that the mapping (y,z)→μ(·|y,z)is jointly measurable and continuous in ywith the total variation norm on (Z). This norm-continuity assumption precludes the possibility that states have a component that depends in a deterministic way on a continuous action. In contrast, our decorrelation approach does not impose norm-continuity; rather, adding atomless transitions to HRR’s framework, we require only that state transitions are weakly continuous, i.e., the mapping (y,z)→μ(·|y,z)is continuous with the weak* topology on (Z). Thus, a component of the state is permitted to vary in a deterministic, continuous way with respect to states and actions. Our analysis is also more general than that of Mertens and Parthasarathy in that it permits the players’ payoffs and Nature’s moves to depend on the entire history of play, but our payoff structure is less general in one respect: this dependence has to be jointly continuous, whereas Mertens and Parthasarathy allow for payoffs and moves by Nature that are continuous in the current action profile but merely measurable in the current state.1 Our analytical approach proceeds as follows. Given a game with atomless moves by Nature, HRR show that there is a SPE in the extended game, which adds payoff-irrelevant public signals in each period. For any such SPE strategy profile, we exploit nonatomicity to “decorrelate” the SPE in each period via repeated application of Mertens’ (2003) “measurable ‘measurable choice’ theorem.” In the first period, there is no previous public signal, so the players’ actions in the SPE are trivially uncorrelated. In the second and later periods of the extended game, however, players can condition on the public signal in the first period. We use Mertens’ theorem to replace conditioning of moves on the first-period public signal (which is payoff-irrelevant) in the extended game with conditioning on Nature’s move (which is nonatomically distributed and payoff-relevant) in the first period of the original game, in which public signals are unavailable. The key 1We also derive a result on existence of a correlated SPE that allows the transition probability to have atoms and assumes only weakly continuous transitions.
Theoretical Economics 16 (2021) Dispensing with public randomization 1223 is to do so in a way that is measurable and preserves the expected discounted payoff of each profile of the players’ actions in the first period, thereby maintaining equilibrium conditions on the players’ first-period choices. This step renders the first-period public signal moot, but in the third and later periods of the extended game, SPE strategies can still condition on the public signal in the second period. We then use Mertens’ theorem to replace conditioning on the secondperiod public signal with conditioning on Nature’s moves in the original game, again preserving expected payoffs from profiles of the players’ actions in the first two periods and thereby maintaining equilibrium conditions. By repeating this procedure, we inductively construct a new profile of strategies such that players’ actions after any history do not depend on the previous public signals, and such that expected payoffs in the first period are preserved from the original SPE. As a consequence, the new profile of strategies is a SPE of the original game that is payoff-equivalent to the SPE of the extended game. Thus, we dispense with HRR’s public randomization for a large class of games of interest, and we conclude that for the purpose of characterizing SPE payoffs of such games, the inclusion of an external public randomization device is without loss of generality. We emphasize that public randomization is payoff-irrelevant, whereas moves by Nature are payoff relevant, so that the substitution of conditioning on Nature’s moves in place of correlation requires delicate arguments. When removing correlation via the period tpublic signal, we must specify new equilibrium strategies in all future periods; moreover, this “cleansing” of correlation must be iterated an infinite number of times, once for each period t.2 Our existence corollary is consistent with an example of Luttmer and Mariotti (2003), in which there is no SPE in a game of perfect information with moves by Nature that are not always atomless. It is also obtained in Proposition 1 of He and Sun (2020), who establish existence of SPE without resorting to public signals under the assumption that Nature’s moves are atomless; their focus is on existence, and they do not provide results on “decorrelation” of equilibria. Our route to existence is comparatively short, but theirs is more direct, in that they do not rely on the results of HRR, but instead follow a backward-then-forward induction argument similar to that of HRR. Their use of Mertens’ (2003) theorem is in ensuring the existence of a suitable jointly measurable selection from next period’s equilibrium payoff correspondence in the forward induction argument, whereas we use Mertens’ (2003) theorem to replace conditioning on public signals with conditioning on moves by Nature.3Methodologically, this decorrelation approach based on nonatomicity is related to our previous work on dynamic games, where 2Observe that although one can work with payoffs or distributions over infinite histories in the extended game Mariotti (2000), the process of decorrelation involves a substitution of equilibrium play after the first period: we maintain the players’ moves in the first period, and we preserve payoffs of the original SPE evaluated at the beginning of the game, but players’ moves in subsequent periods are otherwise unrelated to the original profile. Thus, decorrelation generally produces a distinct distribution over paths of play. Of course, this is unavoidable, because correlation is removed by adding extra conditioning on the realizations of Nature’s moves. 3He and Sun (2020) covered games of perfect information, and they give results for a class of dynamic games that drops continuity of Nature’s moves in exchange for absolute continuity with respect to a fixed, atomless measure.
1224 Barelli and Duggan Theoretical Economics 16 (2021) payoff correspondences with convex values are often needed on technical grounds, but where correlation can be difficult to justify on economic grounds. Similar methods are used in Duggan (2012) to prove existence of stationary Markov perfect equilibria in noisy stochastic games: a correlated equilibrium is deduced from Nowak and Raghavan (1992), and a version of Mertens’ theorem is used to construct a payoff-equivalent equilibrium in which correlation is replaced by conditioning on the noise component of the state.4 Barelli and Duggan (2014) used this approach to prove that for every stationary correlated equilibrium in the Nowak–Raghavan sense, there is a payoff-equivalent stationary semi-Markov equilibrium obtained by replacing correlation with conditioning on the state and actions in the previous period.5 The remainder of the paper is organized as follows. In Section 2,wesetforththe framework of games of almost perfect information used by HRR, and we specialize this to games with atomless moves by Nature. In Section 3, we present our main result, which shows that we can dispense with public randomization in any game with atomless moves by Nature. In Section 4, we apply our result to stochastic games, and in Section 5, we prove the decorrelation theorem. Section 6concludes, and the Appendix shows the equivalence of geometric discounting (which we use) and the payoff formulation of HRR. 2. Games with atomless moves by nature We adopt the framework of HRR but for two modifications, both inconsequential. First, we omit the set Y0of starting points of the game, which was used by HRR to establish upper hemicontinuity of equilibria, whereas our focus is on de-correlation of equilibria. Second, we use a representation of payoff functions in terms of stage payoffs and geometric discounting, whereas HRR use continuous payoff functions defined on infinite histories. This formulation is convenient for our analysis, and we show in Proposition A.1,intheAppendix, that the discounting approach is without loss of generality; given the widespread use of geometric discounting in applied work, this result may be of independent interest. Data. The data of the HRR framework (with our two modifications) are as follows. •There is a finite, nonempty set N={1, ,n}of active players, indexed by ior j,and a passive player, “Nature,” denoted 0. Let N0=N∪{0}. •There is a countably infinite set T={1, 2, }of time periods, indexed by sor t.Let T0=T∪{0}. 4See also He and Sun (2017) for an alternative approach: instead of decomposing states into two components as in Duggan (2012), He and Sun (2017) focused on two sigma-algebras on states, and, assuming that transitions are measurable with respect to the (much) coarser of the two, apply ideas from Dynkin and Evstigneev (1976) to obtain the required convexity. 5Similar techniques are also applied in Barelli and Duggan (2015a) to show that stationary Markov perfect equilibria in noisy stochastic games are payoff-equivalent to equilibria that select only extreme points of equilibrium payoffs in induced games, and they are applied to purification of Bayes Nash equilibria in Barelli and Duggan (2015b).
Theoretical Economics 16 (2021) Dispensing with public randomization 1225 •For each t∈T0and each i∈N, there is a nonempty, complete, separable metric space Yt,iof actions denoted yt,i.SetYt=× i∈NYt,i, with elements yt=(yt,i)i∈N, and set Y=(Yt)t∈T0. •For each t∈T0, there is a nonempty, complete, separable metric space Ztof Nature’s actions denoted zt.SetZ=(Zt)t∈T0. •For each t∈T, there is a nonempty, closed subset Xt⊆×t s=1(Ys×Zs)of possible t-period histories with typical element of Xtdenoted xt(the structure of Xtis elaborated below). Designate a fixed x0∈Y0×Z0as the initial history at which the game begins, and set X0={x0}. Finally, set X=(Xt)t∈T0. •For each t∈Tand each i∈N, there is continuous correspondence At,i:Xt−1⇒Yt,i of feasible actions with nonempty, compact values. In addition, there is a continuous correspondence At,0:Xt−1⇒Ztwith nonempty, closed values. Set At= ×i∈N0At,iand A=(At)t∈T. Consistent with the interpretation of Xtas the set of possible t-period histories, we assume that for each t∈T,Xt=graph(At).Inparticular, the projection of Xton Yt×Ztis a product of sets for each player and Nature. •For each t∈T, there is a continuous mapping ϕt:Xt−1→(Zt),where(·)represents the set of Borel probability measures endowed with the weak* topology. Assume that for each xt−1∈Xt−1, the support of ϕt(xt−1)is contained in At,0(xt−1). Set ϕ=(ϕt)t∈T. •For each t∈Tand each i∈N, there is a bounded, continuous stage payoff function ut,i:Xt→R.Setut=(ut,i)i∈Nand u=(ut)t∈T. •For each i∈N, there is a discount factor δi∈[0, 1).Letδ=(δi)i∈N. These elements describe a game of almost perfect information (or simply, a game), denoted G=(N0,Y,Z,X,x0,A,ϕ,u,δ), in which players and Nature move simultaneously in each period. Given any history xt, we define the subgame at xt, denoted G(xt), in the obvious way. Infinite histories.Forasequencex∈×t∈T0(Yt×Zt)of action profiles in each period, given any t∈T0,wedenotebyxtthe truncation of x,whichprojectsxonto its first t+1 coordinates. An infinite history is a sequence x∈×t∈T0(Yt×Zt)such that for each t∈T0,wehavext∈Xt.LetX∞denote the set of infinite histories, which we endow with relative topology inherited from the product topology on ×t∈T0(Yt×Zt),alongwith the measurable structure generated by finite cylinder sets. Then (X∞)is the set of probability measures ξon infinite histories. Given any t∈Tand history xt∈Xt,let Ht(xt)={x∈X∞|x t=xt}denote the set of continuation histories at xt. Since Atis continuous for each t∈T, it follows that Ht:Xt⇒X∞is a continuous correspondence. Since Ytand Ztare metric spaces for each t∈T, it follows that the product topology on X∞is metrizable (Theorem 3.36, Aliprantis and Border (2006)). Stage payoffs. Assume that for all players i∈N, the discounted sum of stage payoffs, t∈Tδt−1 iut,i(xt), is bounded and continuous on X∞.6An alternative formulation of 6Continuity of discounted streams of payoffs would follow from continuity of each ut,iif stage payoffs were uniformly bounded across t, but we do not impose the latter assumption.
1226 Barelli and Duggan Theoretical Economics 16 (2021) payoffs, used by HRR, is to assume a bounded, continuous payoff function ui:X∞→R over infinite histories for each player. In the Appendix, we exploit history dependence of stage payoffs to show that the two formulations are equivalent; in fact, we show that it is without loss of generality to assume a common, positive discount factor in games of almost perfect information.7 Strategies.Astrategy for player i∈Nis a sequence fi=(ft,i)t∈Tof Borel measurable mappings ft,i:Xt−1→(Yt,i)such that ft,i(At,i(xt−1)|xt−1)=1 for each t∈Tand each xt−1∈Xt−1. Define ft=(ft,i)i∈N:Xt−1→× i∈N(Yt,i)as the profile of mappings for each player in period t.Astrategy profile is an ordered n-tuple f=(fi)i∈N.LetFidenote the set of strategies for player i,andletF=×i∈NFidenote the set of strategy profiles. Continuation payoffs.Givenf∈F,t∈T,i∈N,andxt−1∈Xt−1,playeri’s continuation payoff is the expected discounted payoff in the remainder of the game, following history xt−1, defined recursively by Ut,i(xt−1,f)=yzut,i(xt−1,y,z) +δiUt+1,i(xt−1,y,z),fϕt(xt−1)(dz) i∈N ft,i(xt−1)(dy), where we integrate over players’ and Nature’s actions in period t.Thisisastandardconstruction using dynamic programming techniques to establish the existence of mappings Ut:Xt−1×F→Rnfor each tsuch that, for each f, the continuation payoff Ut(xt−1,f)is Borel measurable as a function of xt−1.8 Equilibrium.Asubgame perfect equilibrium (SPE) is a strategy profile fsuch that for each t∈T, each i∈N, each xt−1∈Xt−1, and each ˜ fi∈Fi, Ut,i(xt−1,f)≥Ut,ixt−1,(˜ fi,f−i). Clearly, a strategy profile fis a SPE of Gif and only if for every history xt,thestrategies restricted to this subgame form a SPE of the subgame G(xt). Auxiliary games. Given a game G,anyt∈T,anyxt−1∈Xt−1, and any bounded, Borel measurable function V:Yt×Zt→Rn,theauxiliary game induced by Vat tgiven xt−1is the strategic form game Gt(xt−1,V)=(N,(At,i(xt−1))i∈N,(πi)i∈N)with player set N, strategy sets At,i(xt−1)with mixed strategies σi∈(At,i(xt−1)), and payoff functions πi(y)=zut,i(xt−1,y,z)+δiVi(y,z)ϕt(xt−1)(dz). 7This is not true in a stationary stochastic game, because stage payoffs are history-independent in that setting. 8Specifically, given xt−1,f, and ϕ, a standard argument (see, e.g., Bertsekas and Shreve (1996)) yields a unique probability measure Pf,ϕ(·|xt−1)over the Borel sets of ×s≥tXssuch that the mapping xt−1→ Pf,ϕ(·|xt−1)is Borel measurable and with the appropriate marginals over the factors; in particular, the marginal over Zt×Ytis i∈N0ft,i(xt−1).Theni’s continuation payoff at tgiven fis Ut,i(xt−1,f)= E[s≥tδs−1 ius,i],whereEdenotes expectation with respect to Pf,ϕ(·|xt−1).ObservethatUt,iis Borel measurable on xt−1, and the recursive definition above follows immediately.
Theoretical Economics 16 (2021) Dispensing with public randomization 1227 Here, the values V(y,z)stand in for the players’ expected future payoffs given action profile (y,z);notethatUt+1((xt−1,·),f)in the definition of SPE plays the same role as V(·)in the definition of auxiliary game. One-shot deviation principle.LetNt(xt−1,V)denote the set of mixed strategy Nash equilibria of the auxiliary game Gt(xt−1,V), i.e., σ=(σi)i∈N∈Nt(xt−1,V)if and only if for each i∈Nand each y i∈At,i(xt−1),wehave y πi(y) j∈N σj(dy)≥y−i πiy i,y−i j=i σj(dy−i). By the one-shot deviation principle, a strategy profile fis a SPE of Gif and only if for all t∈Tand all xt−1∈Xt−1, the profile ft(xt−1)=(ft,i(xt−1))i∈Nis a mixed strategy equilibrium of the auxiliary game Gt(xt−1,Ut+1((xt−1,·),f)).9 Extended games. Given a game G=(N0,Y,Z,X,x0,A,f0,u,δ),theextension of G is the game ˆ Gsuch that ˆ N=N, and for each t∈T, (i) for each i∈N,ˆ Yt,i=Yt,i, (ii) ˆ Zt= Zt×[0, 1], (iii) possible t-period histories are as in the original game with the addition of a signal ωs∈[0, 1]in each period t∈T, i.e., ˆ Xt=(y0,z0),(y1,z1,ω1),,(yt,zt,ωt)|xt∈Xt,(ωs)t s=1∈[0, 1]t, writing elements as ˆ xt=(xt,ω1,,ωt), (iv) for each ˆ xt∈ˆ Xt−1, the marginal of ˆ ft,0(ˆ xt−1)on Ztis ϕt(xt−1),andωtis drawn independently from the uniform distribution on [0, 1],thatis, ˆϕt(ˆ xt−1)=ϕt(xt−1)⊗λ,whereλis the uniform measure on [0, 1], (v) for each i∈N0and each ˆ xt−1∈ˆ Xt−1,wehave ˆ At,i(ˆ xt−1)=At,i(xt−1), and (vi) for each i∈Nand each t∈T,wehave ˆ ut,i(ˆ xt)=ut,i(xt).Thus,ωtis a payoff-irrelevant public signal. Without risk of confusion, given ˆ V:ˆ Yt׈ Zt→Rn, we shall use ˆ Gt(ˆ xt−1,ˆ V) to denote the corresponding auxiliary game at period t∈T,and ˆ Ut(ˆ xt−1,ˆ f)to denote continuation payoffs at t∈T. We now present the definition of games with atomless moves by Nature. The natural definition would be to have ϕt(xt−1)atomless for all t∈Tand all xt−1∈Xt−1,butwe will use a weaker definition that allows for the possibility that the distribution of Nature’s actions has an atom, as long as the active players have only trivial moves and Nature’s moves are atomless in the next period.10 This definition admits games in which the 9The one-shot deviation principle applies because the game is “continuous at infinity” (Blackwell (1965)). 10He and Sun (2020) assumed the stronger version that ϕt(xt−1)is atomless for all t∈Tto obtain their Theorem 1, and then give their Assumption 2, which allows atoms in periods with a single active player, to obtain their Proposition 1. Their class of games satisfying Assumption 2 is closely related to our games with atomless moves by Nature. Any game with atomless moves by Nature can be reformulated as one that satisfies He and Sun’s Assumption 2: if Nature’s move has an atom in period tin our setting, then period t+1 could be collapsed into period tto satisfy Assumption 2 in their framework; this transformation relies on the possibility that Nature’s move in a given period depends on the actions of players in that period, which He and Sun allow. Conversely, any game in which Nature’s moves are atomless and depend on players’ actions within a period can be transformed into a game with atomless moves by Nature by staggering Nature’s moves into a subsequent period. However, we do not allow for sequences of perfect information moves, as described in part (i) of their Assumption 2.
1228 Barelli and Duggan Theoretical Economics 16 (2021) active players and Nature move in alternate periods, a fact that allows us to apply our results to the class of stochastic games, in Section 4. Games with atomless moves by Nature.AgameGis a game with atomless moves by Nature if for all t≥2, and all xt∈Xt, the following holds: ϕt(xt−1)has an atom ⇒for all i∈N,At+1,i(xt)=1 and ϕt+1,0(xt)is atomless. Clearly, Gis a game with atomless moves by Nature if ϕt(xt−1)is atomless for all tand all histories xt−1, but our condition is strictly weaker than this requirement. 3. Dispensing with public randomization Theorem 4 of HRR establishes the following equilibrium existence result. Theorem 1 (Harris, Reny, and Robson). For each game G,theextension ˆ Gadmits a subgame perfect equilibrium. HRR consider several approaches to establishing existence of SPE in the original game G, without public randomization. They note that their theorem implies existence of SPE in games with finite action sets and in zero-sum games, where Nature’s role in the extended game is not crucial.11 We use the HRR theorem to analyze games without public randomization by imposing nonatomicity structure on Nature’s moves. The main result of this paper is the following theorem: for a game with atomless moves by Nature, every SPE ˆ fof the extended game can be “decorrelated” to produce a SPE fof the original game that preserves the players’ expected discounted payoffs from all action profiles at the initial history. Formally, we say fis payoff-equivalent to ˆ fif for all i∈Nand all y∈×i∈NA1,i(x0),wehave zu1,i(x0,y,z)+δiU2,i(x0,y,z),fϕ1(x0)(dz) =zu1,i(x0,y,z)+δiωˆ U2,i(x0,y,z),ω,ˆ fλ(dω)ϕ1(x0)(dz), where U2,iis the continuation payoff after period 1 in the original game and ˆ U2,iis the continuation payoff after period 1 in the extended game.12 Theorem 2. In a game Gwith atomless moves by Nature, every subgame perfect equilibrium of the extended game ˆ Gis payoff equivalent to a subgame perfect equilibrium of G. 11HRR also claimed that existence of SPE in games of perfect information is a consequence of their main result, but Luttmer and Mariotti (2003) provide a counterexample to this claim. 12In the proof of Theorem 2, we define a notion of payoff equivalence at an arbitrary history. Note, however, that the equivalence established in the theorem holds for payoffs calculated at the beginning of the game; in general, the process of de-correlation can change equilibrium play and payoffs in later subgames.
Theoretical Economics 16 (2021) Dispensing with public randomization 1235 Corollary 3. For every stochastic game Gsatisfying (i)–(iv), the associated game ˜ Gwith public randomization admits a subgame perfect equilibrium. By Corollary 3, any stochastic game satisfying (i)–(iv) admits a form of correlated SPE, complementing the well-known result of Nowak and Raghavan (1992), which establishes existence of a stationary correlated equilibrium under norm continuity of transitions.21 Importantly, by allowing for subgame perfect equilibria in which strategies are history-dependent, we extend this existence result to stochastic games Gin which the transition probability may have atoms and is merely weak* continuous. For example, it may be that the transition probability in Gis a deterministic, continuous function of the previous state and actions. Returning to Example 1, we can remove noise from the depreciation rates, simply fixing them at ¯ r1and ¯ r2, so that in any period, the capital stock levels, k1and k2, and consumption levels, c1and c2, determine new capital stock levels k i=Fi(k1,k2)+(1−ri)ki−ciin a deterministic way. Such deterministic transitions violate norm continuity, but Corollary 3delivers a SPE in the associated game, where the agents observe a single, one-dimensional public signal between periods. 5. Proof of Theorem 2 Let Gbe any game with atomless moves by Nature. Before proceeding to the proof, we extend the concept of payoff equivalence to any period t∈Tand history ˆ xt−1∈ˆ Xt−1 as follows. Given strategy profiles ˆ fand ˆ fin the extended game ˆ G, we say that ˆ fis payoff-equivalent to ˆ fat ˆ xt−1if for all i∈Nand all y∈×i∈NAt,i(xt−1),wehave zut,i(xt−1,y,z)+δiωˆ Ut+1,i(xt−1,y,z,ω),ˆ fλ(dω)ϕt(xt−1)(dz) =zut,i(xt−1,y,z)+δiωˆ Ut+1,i(xt−1,y,z,ω),ˆ fλ(dω)ϕt(xt−1)(dz), so that the expected discounted payoffs from ˆ fand ˆ f, calculated at ˆ xt−1,fromevery action profile are the same for every active player.22 In addition, as in HRR, let Et+1(xt) be the set of SPE payoffs in any subgame ˆ G(ˆ xt)of ˆ Gsuch that the history of actions in ˆ xt is xt. Because the signals ω1,,ωtare payoff irrelevant, this set is well-defined, and by Theorem 5 of HRR, the correspondence Et+1:Xt⇒Rnso-defined has a closed graph, and thus is lower measurable (Theorem 18.20, Aliprantis and Border (2006)). To prove Theorem 2,fixaSPE ˆ fof ˆ Gand set f0=ˆ f. In general, for a recursive construction, for each t∈T,wetakeft−1as a given SPE of ˆ Gsatisfying the following conditions:23 21They also assume that the transitions are absolutely continuous with respect to a fixed measure. See also Ja´ skiewicz and Nowak (2017) for an overview of the literature on existence of stationary equilibria in discounted stochastic games. 22Note that the initial definition of payoff equivalence compares a strategy profile fin Gwith a profile ˆ f in ˆ G, whereas payoff equivalence at a history compares to profiles in ˆ G. 23These conditions are vacuously satisfied by f0when t=1, and (C1t) is vacuously satisfied when t=2.
1236 Barelli and Duggan Theoretical Economics 16 (2021) (C1t) for all s∈Tand all ˆ xs−1∈ˆ Xs−1,ft−1 s(ˆ xs−1)is independent of ω1,,ωt−2, (C2t) for all s≥tand all ˆ xs−1∈ˆ Xs−1such that ϕt−1(xt−2)is atomless, ft−1 s(ˆ xs−1)is independent of ω1,,ωt−1, (C3t) for all ˆ xt−2∈ˆ Xt−2,ft−1is payoff-equivalent to ft−2at ˆ xt−2. That is, the strategy profile ft−1does not depend on public signals in periods up to and including t−1, unless Nature’s move ϕt−1(xt−2)hasanatomathistoryxt−2,inwhich case strategies are still independent of public signals in the first t−2 periods. Moreover, expected discounted payoffs from action profiles calculated in period t−1arethesame for ft−1and ft−2. We will construct a strategy profile ftin ˆ Gsatisfying (C1t+1)–(C3t+1)suchthatft shares the first tmoves with ft−1, i.e., ft 1,,ft t=ft−1 1,,ft−1 t. Later moves (ft t+1,ft t+2,)will be independent of ω1,,ωt−1; and if ϕt(xt−1)is atomless, they will be independent of ωtas well. The later moves will also preserve expected payoffs at each history ˆ xt−1from each action profile yt∈× i∈NAt,i(xt−1).In fact, if Nature’s move ϕt(xt−1)has an atom at any history ˆ xt−1, then we simply specify that (ft t+1,ft t+2,)restricted to ˆ G(ˆ xt−1)is identical to ft−1, i.e., for all s≥t+1and all ˆ x s−1∈ˆ Xs−1with ˆ x t−1=ˆ xt−1,ft s(ˆ x s−1)=ft−1 s(ˆ x s−1). Importantly, if Nature’s move ϕt(xt−1)is atomless at history ˆ xt−1, then (ft t+1,ft t+2,)is chosen so that it forms a SPE in the subgame ˆ G(ˆ xt−1), and these later moves will generally differ from those in ft−1. Then strategies restricted to earlier subgames remain subgame perfect by payoff equivalence. Because ft−1does not depend on public signals prior to t−1, by (C1t), we suppress the realizations of public signals prior to period t−1 and write ft−1 t(ˆ xt−1) as ft−1 t(xt−1,ωt−1).Whenϕt−1(xt−2)is atomless, (C2t) allows us to write this simply as ft−1 t(xt−1). The proof consists of a recursive construction in Steps 1–3, and a final Step 4 that produces the desired SPE of the original game G: Steps 1 and 2 perform the inductive construction, and Step 3 verifies that the inductively constructed profile of strategies satisfies (C1t+1)–(C3t+1); Step 4 is an independent verification step, taking place after Steps 1–3 have been repeated countably many times, and establishing that the construction in fact yields an SPE. For each t∈T,letX◦ tconsist of histories xt∈Xtof the original game Gsuch that ft+1,0(xt)is atomless. Note that the set of atomless probability measures is Borel measurable, and since ft+1,0 is Borel measurable, the set X◦ tis itself Borel measurable. In case xt−1/∈X◦ t−1, as mentioned in the preceding paragraph, we specify ftso that at all subsequent histories, play proceeds according to ft−1. To confirm that our conditions are satisfied in this case, note that since ft(xt−1)has an atom, the definition of game with atomless moves by Nature implies that Nature’s move ft−1(xt−2) priortothatisatomless,andthen(C2 t) implies for all s≥t,ft−1 s(ˆ xt−1)is independent of ω1,,ωt−1, so that our specification of ft=ft−1in later subgames satisfies (C1t+1).24 24The specification vacuously satisfies (C2t+1) and (C3t+1), since ϕt−1(xt−2)has an atom.
Theoretical Economics 16 (2021) Dispensing with public randomization 1237 Thus, our arguments focus on histories ˆ xt−1∈X◦ t−1×[0, 1]t−1at which Nature’s moves are atomless. Step 1 in the construction, next, is broken into two parts. The task is to disconnect continuation payoffs in period t+1 from public signals in preceding periods. In the first part, we consider a history ˆ xt−1∈X◦ t−1×[0, 1]t−1such that Nature’s move in the previous period was also atomless, in which case, by (C2t), we need address only dependence on public signals ωt,inperiodt. In the second part, we consider a history such that Nature’s move ϕt−1(xt−2)in period t−2 has an atom. In this case, (C1t) delivers independence from Nature’s moves in the first t−2 periods, but dependence on both ωt−1and ωtmust be addressed. Step 1.1: Disconnecting continuation payoffs after atomless moves in period t−1. Let X1.1 t−1={xt−1∈X◦ t−1|xt−2∈X◦ t−2}consist of (t−1)-period histories in the original game such that Nature’s moves, ϕt(xt−1)and ϕt−1(xt−2),inperiodstand t−1 are atomless. Consider a history ˆ xt−1∈X1.1 t−1×[0, 1]t−1. Since ft−1is a SPE of the extended game ˆ G, the profile (ft−1 t,i(ˆ xt−1))i∈Nis a Nash equilibrium of the auxiliary game ˆ Gt(ˆ xt−1,ˆ Ut+1((ˆ xt−1,·),ft−1)).By(C2 t), the strategies ft−1 sin periods s≥t+1 are independent of public signals ω1,,ωt−1,sowecanwriteplayeri’s continuation payoff at ˆ xt, namely ˆ Ut+1,i(ˆ xt,ft−1),as ˆ Ut+1,i((xt,ωt),ft−1). Then player i’s expected payoff at ˆ xt−1from an action profile y∈×i∈NAt,i(xt−1)simplifies to zut,i(xt−1,y,z)+δiωˆ Ut+1,i(xt−1,y,z,ω),ft−1λ(dω)ϕt(xt−1)(dz).(5) For all histories xt−1∈X1.1 t−1and all (y,z)∈At(xt−1), the continuation payoff ˆ Ut+1((xt−1, y,z,ω),ft−1)belongs to Et+1(xt−1,y,z)and, therefore, ωˆ Ut+1(xt−1,y,z,ω),ft−1λ(dω)∈ω Et+1(xt−1,y,z)λ(dω) =coEt+1(xt−1,y,z), where the equality follows from Lyapunov’s theorem (cf. part 2 of the theorem of Mertens (2003)). Integrating over z,wehave zωˆ Ut+1(xt−1,y,z,ω),ft−1λ(dω)∈z coEt+1(xt−1,y,z)ϕt(xt−1)(dz) =z Et+1(xt−1,y,z)ϕt(xt−1)(dz), where the equality follows from Lyapunov’s theorem, since ϕt(xt−1)is atomless and Et+1(xt−1,y,z)is closed. By part 3 of the theorem of Mertens (2003), there is a Borel measurable mapping t+1:{xt∈Xt|xt−1∈X1.1 t−1}→Rnsuch that for all (xt−1,y)∈X◦ t−1×Ytwith xt−2∈X◦ t−2 and y∈×i∈NAt,i(xt−1), the mapping t+1(xt−1,y,·)is a selection from Et+1(xt−1,y,·), and z t+1(xt−1,y,z)ϕt(xt−1)(dz)=zωˆ Ut+1(xt−1,y,z,ω),ft−1λ(dω).
1238 Barelli and Duggan Theoretical Economics 16 (2021) This gives us a selection of SPE payoffs that is independent of ω1,,ωtand such that after each history ˆ xt−1∈X1.1 t−1×[0, 1]t−1,playeri’s expected payoff from each action profile y∈×i∈NAt,i(xt−1)is z [ut,i(xt−1,y,z)+δit+1,i(xt−1,y,z)ϕt(xt−1)(dz) =zut,i(xt−1,y,z)+δiωˆ Ut+1,i(xt−1,y,z,ω),ft−1λ(dω)ϕt(xt−1)(dz),(6) which is just (5). Thus, the selection t+1preserves the expected payoff from ft−1for each action profile yin period tfollowing histories ˆ xt−1with xt−1∈X1.1 t−1. Step 1.2: Disconnecting continuation payoffs after moves with atoms in period t−1. Let X1.2 t−1={xt−1∈X◦ t−1|xt−2/∈X◦ t−2}consist of (t−1)-period histories in the original game such that Nature’s move ϕt(xt−1)in period tis atomless, but its move ϕt−1(xt−2)in period t−1 is not atomless. Consider a history ˆ xt−1∈X1.2 t−1×[0, 1]t−1,sothatϕt−1(xt−2) has an atom. By definition of a game with atomless moves by Nature, it follows that the active players’ moves are trivial in period tat history ˆ xt−1, in the sense that their actions are predetermined at this history. By (C1t), for all s≥t−1 and all histories ˆ x s with ˆ x t−2=ˆ xt−2,ft−1 s(ˆ x s)is independent of ω1,,ωt−2, but future actions may depend on the public signal ωt−1in period t−1. Since ϕt−1(xt−2)has an atom, the arguments from Step 1.1 cannot be used to eliminate dependence of ft−1 son ωt−1,butwecanuse a similar argument to replace dependence on (ωt−1,ωt)with an appropriate selection of equilibrium payoffs as a function of zt, the distribution of which, namely ϕt(xt−1),is atomless. Because action sets are singleton at ˆ xt−1, this selection may be constructed without concern for equilibrium incentives in period t, but we must preserve expected payoffs from action profiles in period t−1. Player i’s continuation payoff at ˆ xt−2from action profile y∈×i∈NAt−1,i(xt−2)is zut−1,i(xt−2,y,z)+δizut,ixt−2,y,z,y,z +δi(ω,ω)ˆ Ut+1,ixt−2,y,z,ω,y,z,ω,ft−1λ2dω,ωϕt(xt−2,y,z)dz ×ϕt−1(xt−2)(dz),(7) where yis the unique feasible action profile at (xt−2,y,z),andλ2is Lebesgue measure on the unit square [0, 1]2.Wecanwriteyexplicitly as a function αt(xt−1)of history; and because the feasible action correspondences At,iare continuous, the mapping αt:X1.2 t−1→Ytis continuous. Since the public signal is payoff irrelevant, it follows that for all (ω,ω),wehave ˆ Ut+1,ixt−2,y,z,ω,y,z,ω,ft−1∈Et+1xt−1,y,z,y,z, where again y=αt(xt−2,y,z),andthus (ω,ω)ˆ Ut+1,ixt−2,y,z,ω,y,z,ω,ft−1λ2dω,ω∈coEt+1xt−2,y,z,y,z,
Theoretical Economics 16 (2021) Dispensing with public randomization 1239 by Lyapunov’s theorem. Then the integral z(ω,ω)ˆ Ut+1,ixt−2,y,z,ω,y,z,ω,ft−1λ2dω,ωϕt(xt−2,y,z)dz belongs to z coEt+1xt−2,y,z,y,zϕt(xt−2,y,z)dz =z Et+1xt−2,y,z,y,zϕt(xt−2,y,z)dz, where the equality follows by Lyapunov’s theorem from the assumption that ϕt(xt−2, y,z)is atomless. By Mertens (2003), there is a Borel measurable mapping t+1:{xt∈Xt|xt−1∈ X1.2 t−1}→Rnsuch that for all (xt−1,y)∈X◦ t−1×Ytwith xt−2/∈X◦ t−2and y=αt(xt−1), the mapping t+1(xt−1,y,·)is a selection from Et+1(xt−1,y,·)and z t+1(xt−1,y,z)ϕt(xt−1)(dz) =z(ω,ω)ˆ Ut+1,ixt−2,y,z,ω,y,z,ω,ft−1λ2dω,ωϕt(xt−2,y,z)dz. This gives us a selection of SPE payoffs that is independent of ω1,,ωtand such that after each history ˆ xt−2/∈X◦ t−2×[0, 1]t−2,playeri’s expected payoff from each action profile y∈×i∈NAt−1(xt−2)is zut−1,i(xt−2,y,z)+δiz ut,ixt−2,y,z,y,zϕt(xt−2,y,z)dzϕt−1(xt−2)(dz) +δ2 iz t+1(xt−1,y,z)ϕt(xt−1)(dz),(8) where y=αt(xt−1), which is equivalent to (7). Thus, the selection t+1preserves payoffs from ft−1for each action profile yin period t−1 following histories ˆ xt−2with xt−2/∈X◦ t−2,asrequired. Having removed dependence of continuation payoffs in period t+1onNature’sprevious moves, we next describe equilibrium behavior that supports those payoffs. To this end, we assign a SPE of every subgame ˆ G(ˆ xt)with xt∈X◦ tto generate payoffs t+1(xt), and we must do so in a way that is measurable and independent of the history of public signals. We apply Proposition 10 of HRR, reproduced below for the reader’s convenience, to construct the desired equilibrium selection, say ˜ ft.25 Since actions are pinned down at any history ˆ xtwith xt/∈X◦ t, we then splice these together with ˜ ftto arrive at the desired strategy profile ft. 25Most of the notation in the proposition will be clear from its statement and from our explanations below, but Ct+1is an arbitrary upper hemicontinuous correspondence from Xtto Rnwith nonempty, closed values contained in Un,whereUis a compact set that contains the range of each player’s discounted payoffs. Then Ct+1is the set of payoff vectors for the period tstage game when continuation payoffs are chosen from Ct+1.
1240 Barelli and Duggan Theoretical Economics 16 (2021) Theorem 3 (Proposition 10 of HRR). Suppose that ct:ˆ Xt→Rnis a Borel measurable selection from Ct+1. Then there exist Borel measurable mappings ˜ ft,i:ˆ Xt−1→(Yt,i) for all i∈Nand a Borel measurable random selection ct+1:ˆ Xt→Rnfrom Ct+1,such that, for all ˆ x∈ˆ Xt−1: (i) (˜ ft,i(ˆ x)(·))i∈Nis a Nash equilibrium of the stage game when continuation payoff vectors are given by ct+1(ˆ x,·):ˆ At(ˆ x)→Rn;and (ii) ct(ˆ x)isthepayoffvectorofthisNashequilibrium. Step 2: From continuation payoffs to actions. To apply Proposition 10 of HRR, we set t=1 in their result, and we consider a game of almost perfect information with set of starting points equal to our X◦ t, and with subgames determined by each starting point xt∈X◦identical to the subgame G(xt)in our original game, G. We identify their correspondence C2with our Et+1, their mapping c1with our t+1, and their 1-period histories with our set {xt+1∈Xt+1|xt∈X◦ t}×[0, 1]of (t+1)-period histories such that Nature’s move is atomless in period t, together with a public signal in period t.Hence, c1=t+1maps each xt∈X◦ tto SPE payoff vectors in (C2)(xt)=Et+1(xt).ByHRR’s Proposition 10, there exist Borel measurable mappings ˜ ft t+1:X◦ t→(Yt+1,i)for each i∈Nand a Borel measurable selection c2:{xt+1∈Xt+1|xt∈X◦ t}×[0, 1]→Rnfrom Et+2such that for all xt∈X◦ t,(i) ˜ ft t+1(xt)=(˜ ft t+1,i(xt))i∈Nis a mixed strategy equilibrium of the auxiliary game ˆ G1(xt,c2), and (ii) equilibrium payoffs from ˜ ft t+1in the auxiliary game are c1(xt)=t+1(xt). Note that the domain of the mapping ˜ ft t+1is the set of starting points xt∈X◦ t, and so it is manifestly independent of public signals ω1,,ωt. However, we want to use ˜ ft t+1to describe moves in the extended game ˆ G,andsoweextend it to the domain X◦ t×[0, 1]tin the obvious way: for all ˆ xt∈ˆ Xtwith xt∈X◦ t,weset the value of ˜ ft t+1at ˆ xtto be equal to ˜ ft t+1(xt). This introduces nominal dependence on public signals that will be removed in Step 4. Applying HRR’s Proposition 10 recursively (as in the proof of Lemma 18 of HRR), we obtain a sequence ˜ ft t+1,˜ ft t+2, such that for all s∈Twith s>t, the mapping ˜ ft s is defined on histories ˆ xs−1∈ˆ Xs−1such that Nature’s move is atomless in period t, i.e., xt−1∈X◦ t−1.Moreover, ˜ ft s(ˆ xs−1)is independent of public signals ω1,,ωt,andtheprofile ˜ ft s(ˆ xs−1)is a mixed strategy equilibrium of the auxiliary game with continuation payoffs generated by ˜ ft s+1,˜ ft s+2,. We then define the strategy profile ft=(ft s)s∈Tso that in any period s>t, given any history ˆ xs−1∈ˆ Xs−1such that Nature’s move is atomless in period t, players use the strategies obtained via HRR’s Proposition 10; and otherwise, the players follow their strategies in ft−1. Formally, for all s∈Tand all ˆ xs−1∈ˆ Xs−1, (i) if s>tand xt−1∈X◦ t−1,thenft s(ˆ xs−1)=˜ ft s(ˆ xs−1); (ii) if s>tand xt−1/∈X◦ t−1,then ft s(ˆ xs−1)=ft−1 s(ˆ xs−1); and (iii) if s≤t,thenft s(ˆ xs−1)=ft−1 s(ˆ xs−1). Because X◦ t−1is Borel measurable, the mappings ft sso-defined are Borel measurable for all s∈T. Later, we will use the fact that for all ˆ xt∈ˆ Xtsuch that xt−1∈X◦ t−1, ˆ Ut+1ˆ xt,ft=t+1(xt),(9) which follows from (ii), above.
Theoretical Economics 16 (2021) Dispensing with public randomization 1241 We claim that for all s∈Tand all ˆ xs−1∈ˆ Xs−1, the players’ moves ft s(ˆ xs−1)are independent of public signals ω1,,ωt−1in the first t−1 periods. Indeed, in case (i), above, this is implied by the HRR construction. In cases (ii) and (iii), ft sspecifies the same actions as ft−1 s,and(C1 t) implies that actions are independent of ω1,,ωt−2. Clearly, the claim also holds for ωt−1if s<t,soconsiders≥t.IfNature’smoveϕt−1(xt−2)is atomless in period t−1, then the claim follows from (C2t); and otherwise, the definition of game with atomless moves by Nature implies that Nature’s move is atomless in period t, and thus case (i) applies. This establishes the claim, and for the remainder of Step 2, we omit dependence of strategies and subgames on public signals in the first t−1periods for notational simplicity. To verify that ftis a SPE, note that the construction of HRR implies that (ft t+1, ft t+2,)forms a SPE in each subgame ˆ G(xt−1)such that Nature’s move is atomless in period t, i.e., xt−1∈X◦ t−1. As well, for all xt−1/∈X◦ t−1, play in the subgame ˆ G(xt−1)proceeds according to ft−1, which is a SPE. For subgames starting at histories xt−2∈Xt−2, we consider two cases. First, if Nature’s move ϕt−1(xt−2)in period t−1 is atomless, then in Step 1.1, equation (6) implies that for all i∈Nand all y∈×i∈NAt,i(xt−1), z [ut,i(xt−1,y,z)+δiˆ Ut+1,i(xt−1,y,z),ftϕt(xt−1)(dz) =zut,i(xt−1,y,z)+δiωˆ Ut+1,i(xt−1,y,z,ω),ft−1λ(dω)ϕt(xt−1)(dz). Using ft t(xt−1)=ft−1 t(xt−1), we can integrate both sides over action profiles yto obtain ˆ Ut,i(xt−1,ft)=ˆ Ut,i(xt−1,ft−1). This, in turn, implies that the players’ payoff functions in the auxiliary games ˆ Gt−1(xt−2,ˆ Ut((xt−2,·),ft)) and ˆ Gt−1(xt−2,ˆ Ut((xt−2,·),ft−1)) are identical. Because ftspecifies the same mixtures over action as ft−1at xt−2,andft−1is a SPE, it follows that ft t−1(xt−2)=ft−1 t−1(xt−2)is a mixed strategy equilibrium of the auxiliary game. Second, if Nature’s move ϕt−1(xt−2)in period t−1 has an atom, then in Step 1.2, equality of (6)and(7), together with the fact that the players’ moves are pinned down in period t, again implies that ˆ Ut,i(xt−1,ft)=ˆ Ut,i(xt−1,ft−1), which implies that the auxiliary games are identical. Since ft−1is a SPE, it follows that ft t−1(xt−2)=ft−1 t−1(xt−2) is a mixed strategy equilibrium of the auxiliary game. In general, for any period s<t, suppose that for all xs∈Xs,wehave ˆ Us+1,i(xs,ft)= ˆ Us+1,i(xs,ft−1). This implies that for all xs−1∈Xs−1and all y∈×i∈NAs,i(xs−1),wehave z [us,i(xs−1,y,z)+δiˆ Us+1,i(xs−1,y,z),ftϕs(xs−1)(dz) =zus,i(xs−1,y,z)+δiωˆ Us+1,i(xs−1,y,z,ω),ft−1λ(dω)ϕs(xs−1)(dz). Using ft s(xs−1)=ft−1 s(xs−1), we then integrate both sides over action profiles yto obtain ˆ Us,i(xs−1,ft)=ˆ Us,i(xs−1,ft−1). By induction, it follows that for all s=0, 1, ,t−1and all xs−1∈Xs−1,wehave ˆ Us,i(xs−1,ft)=ˆ Us,i(xs−1,ft−1), implying coincidence of the auxiliary games ˆ Gs(xs−1,ˆ Us((xs−1,·),ft)) and ˆ Gs(xs−1,ˆ Us((xs−1,·),ft−1)), and since
1242 Barelli and Duggan Theoretical Economics 16 (2021) ft−1is a SPE, ft s(xs−1)=ft−1 s(xs−1)is a mixed strategy equilibrium of the auxiliary game. We conclude that ftis a SPE, as required. Next, we verify that the strategy profile ftspecified in Step 2 satisfies conditions (C1t+1)–(C3t+1). Step 3: Completing the induction. In Step 2, we have already shown that for all s∈Tand all ˆ xs−1∈ˆ Xs−1, the players’ moves ft s(ˆ xs−1)are independent of public signals ω1,,ωt−1, verifying (C1t+1). Now consider any period s≥t+1andhistory ˆ xs−1∈ˆ Xs−1such that xt−1∈X◦ t−1, so that Nature’s move in period t, namely ϕt(xt−1), is atomless. Then ft s(ˆ xs−1)=˜ ft s(ˆ xs−1), and the latter is independent of ω1,,ωt,by the HRR construction, fulfilling (C2t+1). To verify payoff equivalence, consider any ˆ xt−1∈ˆ Xt−2and any y∈× i∈NAt,i(xt−1). If Nature’s move has an atom in period t, i.e., xt−1/∈X◦ t−1, then ftcoincides with ft−1in the subgame ˆ G(ˆ xt),andthusftis payoff equivalent to ft−1at ˆ xt−1. If Nature’s move is atomless, i.e., xt−1∈X◦ t−1, then zut,i(xt−1,y,z)+δiωˆ Ut+1,i(xt−1,y,z,ω),ftλ(dω)ϕt(xt−1)(dz) =zut,i(xt−1,y,z)+δit+1,i(xt−1,y,z)ϕt(xt−1)(dz) =zut,i(xt−1,y,z)+δiωˆ Ut+1,i(xt−1,y,z,ω),ft−1λ(dω)ϕt(xt−1)(dz), where the first equality follows from (9) and the second equality from (6). This yields (C3t+1), as required. Finally, after Steps 1–3 have been performed countably many times, we construct a SPE of the original game with atomless moves by Nature, G, that is payoff-equivalent to ˆ f. Step 4: Construction of f.Let f∞=(ft−1 t)t∈Tbe the strategy profile in ˆ Gsuch that play in each period tis determined by ft−1. For each t∈T,considerany ˆ xt−1∈ˆ Xt−1. If Nature’s move is atomless in period t−1, i.e., xt−2∈X◦ t−2,then(C2 t) implies that f∞ t(ˆ xt−1)=ft−1 t(ˆ xt−1)is independent of ω1,,ωt−1. Otherwise, if Nature’s move has an atom in period t−1, then (C1t) implies that f∞ t(ˆ xt−1)=ft−1 t(ˆ xt−1)is independent of ω1,,ωt−2, and the definition of game with atomless moves by Nature implies that the players’ actions are pinned down by the history xt−1,sothattheyareindependent of ωt−1, as well. Thus, we remove the nominal dependence of strategies on public signals by simply projecting each ft−1 tonto Xt−1. Formally, for all i∈N,wedefine ft,i:Xt−1→(Yt,i)as follows: for all xt−1∈Xt−1, choose an arbitrary ˆ xt−1∈ˆ Xt−1 such that the history of actions is xt−1,andsetft,i(xt−1)=f∞ t,i(ˆ xt−1). Finally, we define ft=(ft,i)i∈Nand f=(ft)t∈T. To see that fis a SPE of G, recall that for each t∈T,ft−1 tis a SPE of ˆ G. Thus, for each ˆ xt−1∈ˆ Xt−1, the profile ft−1 t(ˆ xt−1)is a mixed strategy equilibrium of the auxiliary game ˆ Gt(ˆ xt−1,ˆ Ut+1((ˆ xt−1,·),ft−1)). For each s>t,(C s+1) implies that fsis payoff equivalent to fs−1at all ˆ xs−1∈ˆ Xs−1. This means that for all i∈Nand all y∈× i∈NAs,i(xs−1),we have zus,i(xs−1,y,z)+δiωˆ Us+1,i(xs−1,y,z,ω),fsλ(dω)ϕs(xs−1)(dz)
Theoretical Economics 16 (2021) Dispensing with public randomization 1243 =zus,i(xs−1,y,z)+δiωˆ Us+1,i(xs−1,y,z,ω),fs−1λ(dω)ϕs(xs−1)(dz). Arguing as in Step 2, we can integrate both sides by fs s(xs−1)=fs−1 s(xs−1)to obtain ˆ Us(xs−1,fs)=ˆ Us(xs−1,fs−1). This, in turn, implies that fsis payoff equivalent to fs−1at each ˆ xs−2∈ˆ Xs−2.By(C3 s), fs−1is payoff equivalent to fs−2at ˆ xs−2, and thus we similarly obtain ˆ Us−1xs−2,fs=ˆ Us−1xs−2,fs−1=ˆ Us−1xs−2,fs−2. Continuing in this way, we conclude that ˆ Ut(xt−1,fs)=ˆ Ut(xt−1,ft). Taking the limit as s→∞, continuity of payoffs implies that for all xt−1∈Xt−1,wehave ˆ Ut(xt−1,f∞)= ˆ Ut(xt−1,ft), which implies that f∞is payoff equivalent to ft−1at ˆ xt−1. Therefore, ft−1 t(ˆ xt−1)is, in fact, a mixed strategy equilibrium of the auxiliary game ˆ Gt(ˆ xt−1, ˆ Ut+1((ˆ xt−1,·),f∞)). By the one-shot deviation principle, we conclude that f∞is a SPE of ˆ G. Since the strategy profile f∞is independent of the payoff-irrelevant public signals, we conclude that fis a SPE of G. Finally, setting t=2intheabovediscussion,weobtain ˆ U2(x1,f∞)=ˆ U2(x1,f1),andthen(C2 2) implies that f∞is payoff-equivalent to ˆ fat ˆ x0 and, therefore, fis payoff equivalent to ˆ f. 6. Conclusions and variations For the class of dynamic games with atomless moves by Nature, we have shown that any SPE obtained in the extended game with public randomization is payoff-equivalent to a SPE of the original game. This has several implications. First, HRR’s public randomization device can be invoked without any loss of generality, as SPE payoffs are unaffected when we allow players to correlate their choices. Second, HRR’s public randomization can be viewed as a step in a proof scheme to ensure existence of an SPE of the original game. Third, although we fix the initial history x0in the analysis, our decorrelation result allows us to invoke HRR’s closed graph result for the extended game (Proposition 34, p. 537) to deduce a closed graph of SPE payoffs of the original game: if we vary the initial action so that xm 0→x0, and if we select a corresponding convergent sequence of SPE payoffs, then HRR show that the limit of those payoffs is a SPE payoff of the extended game at x0, and our de-correlation result implies that it is, in fact, a SPE payoff of the game at x0. In turn, this has implications for continuity of SPE payoffs when action sets vary in a continuous way, or when we approximate an infinite-horizon game by a sequence of finite-horizon games. It is known that SPE payoffs are not generally upper hemicontinuous in action sets, as correlation may be required in the limiting game to support the limit of SPE payoffs (Börgers (1991)), but this concession is unneeded in games with atomless moves by Nature. For example, in Remark 1below we illustrate that if we impose a finite “grid” on the action sets of the players (possibly including Nature) and compute SPE payoffs of the game as the grid becomes fine, then the sequence of payoffs will approach a SPE payoff of the original game with continuous actions.
1244 Barelli and Duggan Theoretical Economics 16 (2021) Remark 1. Let G=(N0,Y,Z,X,x0,A,ϕ,u,δ)be a game with atomless moves by Nature, and let {Gm}be a sequence with Gm=(N0,Y,Z,X,x0,Am,ϕ,u,δ)such that for all t∈T,alli∈N0,andallxt−1∈Xt−1, the sequence {Am t,i(xt−1)}converges to At,i(xt−1) in the Hausdorff metric. For each m,letpmbe a SPE payoff of Gm(x0), and assume pm→p.Thenpis a SPE payoff of Gat x0. The result follows from HRR’s Proposition 34 by formulating a game of almost perfect information, G, in their framework with the same primitives as G, but specifying their set of initial histories as a subset of the real line X0={1 m|m=1, 2, }∪{0}with the relative topology, and specifying their feasible action correspondences as follows: for all t∈T,alli∈N0,andall(t−1)-period histories xt−1, At,i(xt−1)=⎧ ⎨ ⎩ Am t,i(xt−1)if x0=1 m, At,i(xt−1)if x0=0, where xt−1is the history of our game that coincides with xt−1in periods 1, 2, ,t−1. By the assumption that feasible action sets converge in the Hausdorff metric, the feasible action correspondences in Gsatisfy the continuity assumption of HRR, and their proposition applies. Beyond merely technical interest, our remark can facilitate computation of SPE payoffs via finite approximation of games with atomless moves by Nature, without the need to invoke public randomization to support the limiting payoff. Appendix:Comparison of payoff formulations This Appendix compares the payoff formulation of HRR, which assumes payoffs defined directly on infinite histories, to the payoff formulation with geometric discounting, which we use in this paper. Given the prevalence of geometric discounting in applications, the next proposition, which establishes that the two payoff formulations are equivalent, may be of independent interest. The first part of the proposition, which states that discounting is a special case of the HRR formulation, is immediate; the contribution of the result is the converse. Of note, the proof of the converse direction relies on the nontrivial step that for a given player iand any period t, we can choose a path of play starting from each history xtsuch that player i’s expected payoff is a continuous function of the starting point. This step could be replaced by a theorem of the maximum argument if moves by Nature were restricted to compact sets, but we do not assume the sets At,0(xt−1)are compact.26 Since Nature’s moves are exogenously given, however, this noncompactness is not critical. The problem essentially boils down to continuity of the optimal value in a general, nonstationary Markov decision process (viewing the history xtas a parameter of the problem). For this, we in fact formulate the decision problem as a two-player, zero-sum 26It would actually suffice if there were a compact Kt⊂Ztfor each tsuch that Kt∩int At,0(xt−1)=∅for all xt−1. This assumption is quite weak, but we do not impose it.