Reputation building under uncertain monitoring
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Deb, Joyee; Ishii, Yuhta Article Reputation building under uncertain monitoring Theoretical Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Deb, Joyee; Ishii, Yuhta (2025) : Reputation building under uncertain monitoring, Theoretical Economics, ISSN 1555-7561, The Econometric Society, New Haven, CT, Vol. 20, Iss. 1, pp. 169-208, https://doi.org/10.3982/TE4758 This Version is available at: https://hdl.handle.net/10419/320284 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Theoretical Economics 20 (2025), 169–208 1555-7561/20250169 Reputation building under uncertain monitoring Joyee Deb Department of Economics, New York University Yuhta Ishii Department of Economics, Pennsylvania State University We study the standard reputation model with a long-run (LR) player facing a sequence of short-run (SR) opponents, with one difference: the SR players are uncertain about the monitoring structure, while the LR player knows it. We construct examples where the standard reputation result breaks down: Even if there is a possibility that the LR player is a commitment type who always plays the action to which he wants to commit, there exist “bad” equilibria in which the LR player gets payoffs substantially lower than his commitment payoffs. In contrast, if there is the possibility of dynamic commitment types who switch between “signaling” actions that help the SR players learn the monitoring structure and “collection” actions that are desirable for payoffs, our main theorem shows that a sufficiently patient LR player obtains payoffs of at least the commitment payoffs in each state in every equilibrium. Keywords. Reputation, monitoring, repeated games, learning. JEL classification. C73, L14. 1. Introduction Consider a long-run firm building a reputation for producing environmentally-friendly products. Such a reputation is valuable for the firm when consumers care about the environmental impact of their purchases and are often willing to pay more for green products. Consumers make purchase decisions based on whether products have “ecofriendly” labels, but are typically unsure of how much to trust the labels. Many of these labels are genuine certifications with stringent standards, but numerous others have been discredited as being fake. As a result, on seeing an eco-label, consumers are uncertain about its informational content, and may not be convinced about the product Joyee Deb: [email protected] Yuhta Ishii: [email protected] For helpful comments that significantly improved the paper, we thank the three anonymous referees, as well as Dilip Abreu, Heski Bar-Isaac, Martin Cripps, Mehmet Ekmekci, Drew Fudenberg, Johannes Hörner, Michihiro Kandori, Barry Nalebuff, Aniko Öry, Harry Pei, Andy Skrzypacz, Alex Wolitzky, and Jidong Zhou. We also thank Haoning Chen and Kirtivardhan Singh for excellent research assistance. Finally, we are grateful to seminar participants at Brown, Duke, ITAM, Oxford, Queen Mary University of London, University College London, University of Warwick, Yale, and the SITE Summer Workshop 2016 in Dynamic Games, Contracts, and Markets. ©2025 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at https://econtheory.org.https://doi.org/10.3982/TE4758
170 Deb and Ishii Theoretical Economics 20 (2025) being environmentally friendly.1But if consumers do not trust product labeling, a firm, even after honest investment in green products and after undergoing reliable labeling, may find it difficult to establish a positive reputation and convince consumers that its products are indeed environmentally friendly. This motivates the central question of the paper: Can reputations be built in environments with such uncertainty in monitoring? To start, consider reputation building in environments in the absence of such uncertainty. Canonical models of reputation (e.g., Fudenberg and Levine (1992)) consider a long-run (LR) agent (a firm) who repeatedly interacts with short-run (SR) opponents (consumers). There is incomplete information about the firm’s type: consumers entertain the possibility that the firm is of a “commitment” type that is committed to playing a particular action in every period. Even when the actions of the firm are noisily observed, the classical reputation result states that if a sufficiently rich set of commitment types occurs with positive probability, a patient firm can achieve payoffs arbitrarily close to their Stackelberg payoff of the stage game in every equilibrium.2Intuitively by mimicking a commitment type that always plays the Stackelberg action, a LR firm can eventually signal to the consumer its intention to play the Stackelberg action in the future and thus obtain high payoffs in any equilibrium. Importantly, this result remains valid even on introduction of other arbitrary commitment types. This intuition critically relies on the consumer’s ability to accurately interpret the noisy signals, but if monitoring is uncertain, the reputation builder may find it difficult to signal his intentions. To study the effect of uncertain monitoring, we also consider the canonical model of a LR firm facing a sequence of SR consumers, but with one key difference. At the beginning of the game, a persistent state (θ,ω)∈×is realized, which determines both the type of the firm, ω, and the monitoring structure, πθ:A1→(Y): a mapping from actions taken by the firm to distribution of signals, (Y), observed by consumers. We assume that the firm knows the state of the world, but the consumer does not. We first show in a simple example that uncertain monitoring can cause the traditional reputation result to break down: Even if consumers believe that the firm may be a commitment type that plays the Stackelberg action every period, there exist equilibria in which even a patient firm obtains payoffs far below its Stackelberg payoff. Such “bad equilibria” arise due to an identification problem that stems from the uncertainty about monitoring: Good actions in one state cannot be statistically distinguished from a bad action in a different state. Our simple example with such a bad equilibrium leads us to ask what might restore reputation building under uncertain monitoring in the face of such identification problems. Under an assumption that the action space is sufficiently rich, we construct a set of commitment types such that, if these types occur with positive probability, a sufficiently patient firm obtains payoffs arbitrarily close to the Stackelberg payoff in all equilibria, 1The Federal Trade Commission maintains, “Very few products, if any, have all the attributes consumers seem to perceive from such claims, making these claims nearly impossible to substantiate” (Source: E. Wyatt, “FTC Issues Guidelines for Eco-Friendly Labels,” New York Times, Oct 1, 2012). 2The Stackelberg payoff is the payoff that the LR player would get if he could commit to an action in the stage game, and the Stackelberg action is the corresponding commitment action. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 171 even when the consumers are uncertain about the monitoring environment.3Importantly, the result holds independent of the fine details of the type space in that it remains valid even if we include other arbitrary commitment types. The commitment types that we construct are committed to dynamic (time-dependent) strategies that switch infinitely often between signaling actions that help the consumer learn the unknown monitoring state and collection actions that are desirable for payoffs (the Stackelberg action). A key contribution is the construction of these dynamic commitment types that play periodic strategies. As we will discuss later, such dynamic commitment types are generally necessary for reputation building under uncertain monitoring, because signaling the unknown state and Stackelberg payoff collection may require the use of different actions in the stage game. The proof of the main result involves establishing two properties, which together imply that the LR player can guarantee payoffs close to Stackelberg payoffs in any equilibrium. First, we show that by mimicking any commitment type, the LR player can ensure in any equilibrium with high probability that the SR players’ predictions of the public signal distribution are close to the true distribution generated by this commitment type in all but a finite number of periods. This step demonstrates the classic result in the spirit of “merging of opinions,” à la Blackwell and Dubins (1962), and is proved using standard arguments from Gossner (2011).4In our setting, ensuring accurate predictions of the public signal distribution by the SR players is not sufficient for a reputation result due to potential identification problems across states. Second, we show that by mimicking the appropriate commitment type, the LR player can additionally ensure that the SR players learn the state at a rate that is uniform across all equilibria. We prove this by establishing a result on robust learning, which provides an easy-to-check sufficient condition that guarantees that an observer will learn the validity of an event at a uniform rate across a rich class of learning environments. The condition relates the uniform rate at which Hellinger transforms vanish across all learning environments in the class to uniform learnability of an event.5To the best of our knowledge, the robust learning theorem is a novel methodological contribution, which applies to general learning environments beyond the specific reputation context of this paper. A key feature of the constructed dynamic commitment types is that they return to the signaling phase infinitely often. One might reasonably conjecture that the inclusion of a commitment type that begins with a sufficiently long phase of signaling followed by a permanent switch to playing the Stackelberg action for the true state would suffice for reputation building. We show in examples that this is generally not sufficient. Also, while this paper is motivated by environments with uncertain monitoring, our model allows for uncertainty both about monitoring and about the payoffs of the reputation 3We can also interpret our model as one that represents subjective uncertainty that consumers have about the actual monitoring structure and the behavior of the reputation-building firm. We show that the firm can indeed effectively establish a reputation, as long as the consumers assign positive probability to the constructed commitment types and the correct monitoring structure. 4See also the discussion after Lemma 3. 5See Section 5.3 for precise statements of our sufficient condition, as well as Torgersen (1991) and Moscarini and Smith (2002) for illustrations of other applications of the Hellinger transform. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
172 Deb and Ishii Theoretical Economics 20 (2025) builder. Finally, our main result continues to hold even if the signals observed by the SR players are unobserved by the LR player. While the main result establishes a lower bound on the LR player’s equilibrium payoff, a natural question is whether the LR player can obtain payoffs much higher than the Stackelberg payoff. With uncertain monitoring, a patient LR player may be able to obtain payoffs that are strictly higher than the Stackelberg payoff of the true state. The reason is that the LR player may not find it optimal to signal the true state, but would rather block learning to attain payoffs that are higher than the Stackelberg payoff in the true state. Providing a general, sharp characterization of an upper bound on a patient LR player’s equilibrium payoffs is difficult, as it depends on the specific set of commitment types and the prior distribution over types.6Nevertheless, we provide a joint sufficient condition on the monitoring structure and stage game payoffs that ensures that the lower bound and the upper bound coincide: Loosely speaking, these are games in which state revelation is desirable for the LR player. 1.1 Related literature We contribute to the literature on reputation that started with Kreps and Wilson (1982) and Milgrom and Roberts (1982), and includes the canonical models of Fudenberg and Levine (1989,1992), and more recent contributions by Gossner (2011). As far as we know, this paper is the first to study reputation under uncertain monitoring. Aumann, Maschler, and Stearns (1995)andMertens, Sorin, and Zamir (2014)study repeated games with uncertainty in both payoffs and monitoring, but focus on zero-sum games. Wiseman (2005), Hörner and Lovo (2009), and Hörner, Lovo, and Tomala (2011) study payoff uncertainty in non-zero-sum repeated games, but do not allow uncertainty about the monitoring structure. Our framework is closest to Fudenberg and Yamamoto (2010), who study a repeated game in which there is uncertainty about both monitoring and payoffs. However, Fudenberg and Yamamoto (2010) focus on perfect public ex post equilibrium in which players play strategies whose best responses are independent of any belief about the state. As a result, in equilibrium, no player has an incentive to affect the beliefs of the opponents about the monitoring structure. We study more general equilibria where the LR player may have incentive to affect the beliefs of the SR players about the monitoring structure. The necessity of dynamic commitment types for reputation building due to identification problems is novel. Dynamic commitment types also arise in reputation building against LR opponents, as in Aoyagi (1996), Celentani, Fudenberg, Levine, and Pesendorfer (1996), and Evans and Thomas (1997), because establishing a reputation for carrying out punishments after certain histories can be beneficial for the reputation builder.7,8 6This is in contrast to the previous papers in the literature, where the payoff upper bound is generally independent of the fine details of the type space such as the relative probabilities of commitment types. 7Atakan and Ekmekci (2011,2015), and Ghosh (2014) also use similar ideas. 8In this literature, some papers do not require the use of dynamic commitment types by restricting attention to conflicting interest games. See, for example, Schmidt (1993) and Cripps, Dekel, and Pesendorfer (2005). 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 173 But in our setting with SR players, the threat of punishments has no bite. Dynamic commitment types turn out to still be necessary to resolve a trade-off between signaling the correct state and collecting the Stackelberg payoff, which are both desirable to the reputation builder. In a recent paper, Pei (2020) studies reputation with interdependent values. Pei (2020) restricts attention to perfect monitoring and a finite number of stationary commitment types, and studies the conditions under which the repeated game yields a reputation result. In contrast, we study a model where actions are imperfectly observed, but the observed public signals can potentially convey information about the state. We similarly show that reputation building can break down when the type space only consists of stationary commitment types, and further construct dynamic commitment types that would restore a reputation result given general type spaces that contain these dynamic commitment types in its support. Our negative examples demonstrate that reputation building may be fragile in the presence of uncertainty about monitoring, because multiple combinations of state and action lead to the same distribution over observed public signals. Identification problems can also give rise to long-run disagreements between different agents in Acemoglu, Chernozhukov, and Yildiz (2016), and can result in convergence to incorrect beliefs in dynamic games with learning, as in Fudenberg and Levine (1993a,1993b). The novel question that we address here is whether or not such identification problems can be circumvented by a patient long-lived player in a reputation setting. Finally, our robust learning theorem also relates to a recent literature that studies rates of learning in decision theoretic settings. Moscarini and Smith (2002)andMu, Pomatto, Strack, and Tamuz (2021) both provide exact characterizations of the speed of learning in decision theoretic settings, focusing on learning environments where the signals arrive in an independent and identically distributed (i.i.d.) manner conditional on the realized state. On the other hand, our robust learning theorem focuses only on a lower bound on the rate of learning, while allowing for signals that may exhibit arbitrary forms of serial correlation. Our robust learning result also relates loosely to ideas of uniform learning from Vapnik–Chervonenkis theory used, for example, in Al-Najjar (2009)andAl-Najjar and Pai (2014). These papers study the uniform learning of a rich class of events given any i.i.d. process. The main conceptual distinction of our robust learning result is that we study uniform learning of finitely many events, but allow for any arbitrary stochastic process that may involve arbitrary serial correlations. 2. Model 2.1 Notation We first introduce some notation that we use throughout the paper. Given a countable set X,let(X)denote the set of all probability measures on X.Let+(X)be the set of full support probability measures on X.ForanyB⊆X,weletBcdenote the complement of B. Given x,x∈Xand some real number λ∈[0, 1],weletλx ⊕(1−λ)x∈(X)denote the probability measure that assigns probability λto xand 1 −λto x.Ifν∈ 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
174 Deb and Ishii Theoretical Economics 20 (2025) (X1×···×Xn),thenmargXjνis the marginal distribution of νon Xj:margXjν(xj)= i=jν(xj,x−j). Given a probability measure ν∈(X)and some function g:X→R, define Eν[g(x)] to be the expectation of g(x)when xis distributed according to ν. Given a finite set Yand a countable set X, define S(Y,X)as the set of all possible stochastic processes over Y∞with state space Xas follows. Formally, an element s∈S(Y,X)is a sequence s={st}∞ t=0, where for each t,st∈(Yt×X)satisfies the consistency condition margYt−1×Xst=st−1. By Kolmogorov’s extension theorem, for any s∈S(Y,X), there exists some s∞∈(Y∞×X)such that margYt×Xs∞=stfor all t.For any s∈S(Y,X)and any subset C⊆X, we can also define sC∈S(Y,X)as the corresponding stochastic process conditional on C:sC=(st(·|C))∞ t=0. We use Nto represent the set of all natural numbers including zero and let N+:= N\{0}. Finally, we establish the convention that both inf∅=min∅=∞and sup∅= max∅=−∞. 2.2 Setting A long-run (LR) player, player 1, faces a sequence of short-run (SR) player 2s. Before the interaction begins, a pair (θ,ω)∈×of a state of the world and type of player 1 is drawn independently according to the product measure γ0:=ν0×μ0with ν0∈+() and μ0∈+(). We assume that is finite and enumerate :={θ0,,θm−1},but may possibly be countably infinite.9The realized pair of state and type (θ,ω)is then fixed for the entirety of the game. In each period t=0, 1, 2, , players simultaneously choose actions from their respective action spaces at 1∈A1and at 2∈A2. We assume A1and A2are finite. Let A=A1×A2.LetAi:=(Ai)be the set of mixed actions of player iwith typical element αi. In each period t≥0, after players have played action profile at∈A, a public signal yt is drawn from a finite signal space Yaccording to the probability measure, ψ(·|at,θ)∈ (Y). Note importantly that both the action profile chosen at time tand the state of the world θpotentially affect the signal distribution. The state of the world θrepresents the unknown monitoring structure. Denote by Ht:=Ytthe set of all t-period public histories with typical element ht=(y0,,yt−1)and assume by convention that H0:=∅. Let H:=∞ t=0Htdenote the set of all public histories of the repeated game. We assume that the LR player observes the realized state of the world θ∈perfectly so that his private history at time tis formally a vector, ht 1∈Ht 1:=×At 1×Yt. Meanwhile the SR player at time tobserves only the public signals up to time tand so his information coincides exactly with the public history Ht 2:=Ht. A strategy for player iis a map σi:∞ t=0Ht i→Ai. Denote the set of strategies of player iby i. Finally, let B1be the set of static state-contingent mixed actions of player 1, B1:={β1:→A1}with typical element β1. 9The assumption of allowing to be countably infinite is standard in the existing literature (e.g., Fudenberg and Levine (1992)) when the Stackelberg action of the stage game can be mixed. We do not know whether our arguments can be extended to the setting where ||is countably infinite. We leave this open for future research. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 175 2.3 Type space We assume that =com ∪{ωs},wherecom is the set of commitment types and ωsis a strategic type. Each commitment type ω∈com is associated with a strategy σω 1∈1 such that type ωalways plays σω 1. In contrast, type ωs∈is a strategic type who chooses a strategy σ1∈1to maximize payoffs, which we describe in the next subsection. Thus, a strategy profile, denoted σ=((σ1(ω))ω∈,σ2),isatupleforwhichσ1(ω)=σω 1for all ω∈com. 2.4 Payoffs and equilibrium Any strategy profile σtogether with the prior γinduces a unique stochastic process, (πσ t)∞ t=0∈S(Y×A,×)for all t. By the Kolmogorov extension theorem, there exists some πσ ∞∈(H∞×A∞××)such that for all t,margHt×At××πσ ∞=πσ t. To study SR players’ best responses, it will also be useful to define the following beliefs of the SR players after observing a public signal history: λσ t·|ht:=margA1×πσ t·|ht∈(A1×), γσ t·|ht:=marg×πσ t·|ht∈(×), νσ t·|ht:=margπσ t·|ht∈(), μσ t·|ht:=margπσ t·|ht∈(). Then SR players’ expected payoffs in any period depend on the belief, λ∈(A1×): u2(a2,λ):=Eλu2(a1,a2,θ)= a1∈A1,θ∈ u2(a1,a2,θ)λ(a1,θ). Thus, a strategy profile, σ, yields the expected payoff of u2(σ2(ht),λσ t(ht)) in period t after the public history ht.LetB2(λ)denote the mixed best responses of player 2, i.e., B2(λ):=argmaxα2∈A2u2(α2,λ). With a slight abuse of notation, we write B2(α1,θ)= B2(α1×1θ),where1θis the Dirac probability measure that assigns probability 1 to θ, and B2(β1,p)=B2(λβ1,p),whereforβ1∈B1and p∈(),λβ,p(a1,θ)=p(θ)β1(a1|θ). The payoff of the LR strategic type, ωs, in state θis given by U1(σ1,σ2,θ;δ):=Eπσ ∞(1−δ) ∞ t=0 δtu1at 1,at 2,θ|θ,ωs. Then the ex ante expected payoff of type ωsis U1(σ1,σ2;δ):=Eν0U1(σ1,σ2,θ;δ). Finally, we can define the statewise-Stackelberg payoff of the stage game. The Stackelberg payoff of player 1 in state θis given by u∗ 1(θ):=sup α1∈A1 inf α2∈B2(α1,θ)u1(α1,α2,θ). 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
176 Deb and Ishii Theoretical Economics 20 (2025) For each ε>0, let Sε θbe the set of ε-Stackelberg actions in state θ, which are the mixed actions that approximate u∗ 1(θ)up to εin θ∈: Sε θ:=α1∈A1:inf α2∈B2(α1,θ)u1(α1,α2,θ)>u ∗ 1(θ)−ε. We analogously define Sε⊆B1as Sε:=β1∈B1:β1(θ)∈Sε θfor all θ∈. Our analysis will focus on Bayes Nash equilibria; to shorten the exposition, subsequently we will refer to Bayes Nash equilibrium simply as equilibrium. We let BNEδdenote the set of all equilibria of the game.10 2.5 Information structure and key assumptions We now impose two key assumptions on the information structure, Assumptions 1and 2, which we maintain for the entirety of the paper. We start with a definition. Definition 1. A signal structure ψsatisfies action identification for (α1,θ)∈A1×if, for all α2∈A2, ψ(·|α1,α2,θ)=ψ·|α 1,α2,θ=⇒ α1=α 1. Let Bid ⊆B1be the set of all β1∈B1such that (β1(θ),θ)satisfies action identification for all θ∈. Assumption 1. For every ε>0,Sε∩Bid = ∅. In words, the above assumption holds if and only if in every state θ, there exists some ε-Stackelberg action in state θsuch that this action would be statistically identified from all other actions regardless of the actions played by the SR player. Note that this is generally a minimal condition that is required for a LR player to be able to guarantee Stackelberg payoffs in state θ, since without it, reputation building may be impossible even when θis common knowledge. While the above assumption concerns statistical identification of actions for a fixed state θ, this is generally not sufficient for a reputation theorem. We furthermore impose the following assumption, which concerns the statistical identification of actions across states. Assumption 2. For every θ= θ,thereexistsomeα1∈A1such that ψ(·|α1,α2,θ)= ψ·|α 1,α2,θ for all α 1∈A1and all α2∈A2. 10Our main theorems provide bounds on payoffs across all equilibria. So these bounds also apply even when restricting attention to more stringent solution concepts such as perfect Bayes Nash equilibria. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 183 Figure 6. The information structure. signaling actions that help the consumer learn the unknown monitoring state and collection actions that are desirable for payoffs of the LR player. Because of the necessity to play both types of actions, our commitment types are nonstationary, playing a periodic strategy that alternates between signaling phases and collection phases.23 Finally, as we have already emphasized, our reputation result does not depend on specific distributional assumptions on the type space. In particular, it remains valid even if we include other possibly bad commitment types, as richness of the type space (,μ)only requires the existence of types ωβ1, while placing no restrictions on the existence or absence of other commitment types. 4.3 Necessary characteristics of commitment types The commitment types, ωβ1, have two key features: (i) They switch play between signaling and collection phases, and (ii) they do so infinitely often. These two features are important and in some sense also necessary for reputation building, given the possibility of identification problems in the monitoring structure. Consider again the stage game from Figure 2and suppose that the information structure is now given by Figure 6. To highlight the importance of (i), we provide an example below in which the strategic LR player regardless of his discount factor obtains a low equilibrium payoff in state θ=bif all commitment types play stationary strategies. To highlight the importance of (ii), we consider type spaces in which all commitment types play strategies that front-load the signaling phases and again construct equilibria in which LR gets a payoff much below the Stackelberg payoff in state b. 4.3.1 Stationary commitment types Consider any arbitrary countable set ∗of commitment types, each of which is associated with the play of a state-contingent action β∈B1at all periods. For each ω∈∗,letβωbe the associated state-contingent mixed action plan of type ω. Notice that this type space contains only stationary commitment types. We now show that the existence of such types is generally not sufficient for reputation building. Formally, given any countable set of stationary commitment types, ∗,wecanconstruct a set of commitment types com ⊇∗and a probability measure μ∈+(com ∪ {ωs})such that there exists an equilibrium in which the strategic LR player obtains a payoff significantly below the Stackelberg payoff in state b. 23A similar reputation theorem can be proved also with stationary commitment types that have access to a public randomization device. In particular, we would need a rich space of stationary commitment types that each signal the state with different probabilities. We thank Johannes Hörner for pointing this out. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
184 Deb and Ishii Theoretical Economics 20 (2025) To simplify notation, let Ab:={α1(C)≥2/3}.NoticethatBis a best response to α1 in state bif and only if α1∈Ab.Letb:={ω∈∗:βω(b)∈Ab}. Given the information structure, ψ, for every α1∈Ab, there exists a corresponding bad action, α1, in state gsuch that ψ(·|α1,b)=ψ(·|α1,g). For every ω∈b,letωdenote atypewhoplaysβω(b)in state gand Din state b. Let the type space consist of =∗∪{ω:ω∈b}∪ωs. Claim 1. Suppose that γ0(ω,b)<2 3γ0(ω,g)for all ω∈b. Then for every δ∈(0, 1),it is a PBE for the LR to always play Dand the SR to always play N. In particular, this PBE yields a payoff of 0<u ∗ 1(b)=4/3to the LR in state θ=b. Proof.Letσdenote the above strategy profile and consider the belief, λσ t((C,b)|ht), that the SR assigns to the event (C,b)at a history ht. By construction, for any ω∈b, γσ t((ω,b)|ht)=γ0(ω,b) γ0(ω,g)γσ t((ω,g)|ht)for any ht. Therefore, λσ t(C,b)|ht= ω∈b γσ t(ω,b)|htβω(C|b)+ ω/∈b γσ t(ω,b)|htβω(C|b) < ω∈b 2 3γσ t(ω,g)|ht+ ω/∈b 2 3γσ t(ω,b)|ht≤2 3. Recall that it is a best response to play Nat a history if λσ t((C,b)|ht)<2/3. Hence, it is abestresponsefortheSRtoplayNat all histories. Then it is immediate that it is a best response for te LR to play Dat all histories. If ν0(b)=1, as long as the closure of Ab∩{βω(b):ω∈b}contains 2/3(themixed Stackelberg action), then a sufficiently patient player obtains payoffs close to 4/3inany equilibrium, since a deviation to mimicking one of the good commitment types in b guarantees such a high payoff. Now consider the case when ν0(b)=1/2. Consider again a deviation to mimicking a good type in b. Such a deviation no longer guarantees a high payoff, since there are now also bad commitment types in {¯ω:ω∈b}in state g that replicate exactly the same distribution over public signals as the good commitment types. As a result, SR players are never able to differentiate between these types, and if the prior places relatively higher weight on such types in state g, then the SR players will never become optimistic about the event b×{b}. 4.3.2 Type spaces with front-loaded signaling Next we present an example where each commitment type switches between signaling and collection, but not infinitely often; i.e., they can play signaling actions for at most Nperiods and then switch to collection forever. In such type spaces, we show that a reputation theorem again does not hold generally. Again consider the same stage game (Figure 2) and information structure (Figure 6) from the previous subsection. Let ωbbe a bad commitment type who always plays α∗ 1= 2 3C⊕1 3Din state gand always plays Din state b.Notethatψ(·|α∗ 1,g)=ψ(·|C,b).Let 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 185 ωtdenote a commitment type who plays Duntil period t(signaling phase) and thereafter switches to the action Cforever after (collection phase).24 For any N∈N+∪{∞}, consider the set of types N:={ωt:t∈N+,t≤N}∪{ωs,ωb}.25 We now show in the following claim that without further distributional assumptions on the type space, reputation building cannot be guaranteed. Claim 2. Let N∈N+∪{∞}and ν0(g)=ν0(b)=1/2. Then there exists some μ0∈+(N) such that for any δ∈(0, 1),itisaPBEfortheLRtoalwaysplayDand the SR to always play N. Moreover, this PBE yields a payoff of 0<4/3=u∗ 1(b)to ωsfor all discount factors in state b. Proof. Consider the probability distribution over types given by μ0ωt=κtε,μ0ωb=ε,μ0ωs=1− N τ=0 κτε. We assume that κ∈(0, 1/4]and ε∈(0, 1/2), in which case, μ0is a valid probability measure since N τ=0κτε<1. Let σdenote the above strategy profile. Consider the probability, λσ t((C,b)|ht). Since only types {ω1,,ωt}play Cin state bat such a history, λσ t(C,b)|ht=γσ tω1,,ωt−1×{b}|ht, but for each τ, the likelihood ratio between (ωτ,b)and (ωb,g)is given by γσ tωτ,b|ht γσ tωb,g|ht=μ0ωτ μ0ωb τ τ=1 ψ(yτ|D,b) ψ(yτ|α1,g)≤μ0ωτ μ0ωb3 2τ =3 2κτ Therefore, λσ t(C,g)|ht≤ t−1 τ=13 2κτ γσ tωb,g|ht≤ ∞ τ=13 8τ <2 3. Again recall that whenever λσ t((C,g)|ht)<2 3, the SR has a strict incentive to play N. Therefore, this shows that it is indeed optimal for the SR to play Nat all histories. Given this, it is immediate that it is a best response for ωsto always play D. Reputation building fails in this example because all signaling is front-loaded by all commitment types. To see the basic idea, consider again the deviation to a strategy of mimicking ωt. The hope under such a deviation for the LR is that the initial tperiods of signaling would be sufficient to convince the SR players that the state is bto a sufficient 24For expositional simplicity, we focus only on those commitment types who play the pure Stackelberg action in the collection phase. The example can be easily extended to settings where such types play mixed actions in the collection phase. 25When N=∞,N={ωt:t∈N+}∪{ωs,ωb}. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
186 Deb and Ishii Theoretical Economics 20 (2025) degree of confidence that it eliminates identification problems across states. However, the claim above shows that this is infeasible for the LR. The problem is that under the constructed belief, μ0∈+(), the likelihood of (ωt,b)is much smaller than the likelihood of (ωb,g), so that even after tperiods of signaling, the SR still maintains high probability on (ωb,g). Remark. If instead γ0(ωt,b)were sufficiently large relative to tfor every t,thenareputation result would hold in the example. However, as previously emphasized, this illustrates the dependence of reputation building on the fine details of the prior distribution over types (beyond just the support of the prior distribution) when signaling is frontloaded. Both of these issues highlighted in the above examples are no longer problematic given the commitment types constructed in the main theorem. First, because the commitment types enter signaling phases many times, there are no bad types in other states that can replicate similar distributions over public signals during these signaling phases for long periods of time. Second, the commitment types signal the state indefinitely so that a LR player who mimics such a commitment type can ensure eventual correct learning of the state even if the probability of such a commitment type is initially very small. Finally, while the above analysis demonstrates the necessity of dynamic commitment types in particular examples, there are many specific settings where either stationary commitment types or commitment types with front-loaded signaling suffice.26 An exact characterization in general games of when such simpler types suffice is beyond the scope of this paper.27 Despite this, we emphasize again that Theorem 1holds as long as richness is satisfied without any restrictions on what other types are or are not present. 5. Proving Theorem 1 We now return to prove Theorem 1. The overall structure of the proof follows the standard approach in the reputation literature. We show that for β1∈Sε, a sufficiently patient LR player, by playing the strategy, σβ1, associated with type ωβ1, can obtain payoffs at least u∗ 1(θ)−ρin any equilibrium. To show this, we prove two key properties that hold uniformly across all equilibria: For every ε>0, there exists some J(that can be chosen independent of the choice of equilibrium) such that by deviating to play σβ1in any equilibrium, the following statements hold: P1. The SR players’ predictions of the current period’s public signal distribution are approximately correct in all but Jperiods with probability at least 1 −ε(see Lemma 3for details).28 26See the discussion after Theorem 1and the remark above. 27The main obstacle is the difficulty of explicit construction of equilibria in general reputation games. 28By “approximately correct,” we mean that the SR players’ predictions will be close to the actual public signal distribution under σβ1and state θ. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 187 P2. The SR players’ beliefs assign probability at least 1 −εon the correct state θin all but Jperiods with probability at least 1 −ε(see Lemma 4for details). Property P1 holds in previous papers studying reputation building under imperfect monitoring such as Fudenberg and Levine (1986)andGossner (2011), and its proof follows these standard arguments. In those environments, with appropriate identification assumptions, P1 implies that with high probability, the SR players’ best responses will be approximately correct in all but J1periods. However, this property alone is inadequate for a reputation theorem in our environment. Even if the SR players’ predictions of today’s public signal distribution is exactly the same in all periods as that of σβ1in state θ, because of identification problems across states, this does not necessarily imply that the SR players’ beliefs are concentrated on the correct state θ. Property P2 addresses this issue. To prove it, we prove a theorem on robust learning (Theorem 2) that establishes a simple-to-check sufficient condition to ensure that an observer learns the relevant state at a rate that is uniform across a rich class of general learning environments. We then apply this theorem to the reputation setting to show that SR players learn the state θat a rate that is uniform across all equilibria. 5.1 Formal details of the proof of Theorem 1 We now provide details of the proof of Theorem 1. Proofs not provided in the text can be found in the Appendices. We first extend the notion of ε-entropy confirming best response of Gossner (2011) to our framework.29 To state this, first recall the definition of the Kullback–Leibler divergence of two probability measures: Given two probability measures P,Q∈(Y), D(PQ):= y∈Y P(y)logP(y) Q(y). Recall the basic properties of relative entropy that D(PQ)≥0forallP,Q∈(Y),and D(PQ)=0 if and only if P=Q. Definition 3. Let (κ,ε)∈[0, 1]2.Thenα2∈A2is a (κ,ε)-confirming best response at (α1,θ)if there exists some λ∈(A1×)such that (i) α2∈B2(λ) (ii) D(ψ(·|α1,α2,θ)ψ(·|λ,α2))≤ε(see footnote 30)30 (iii) margλ(θ)≥1−κ. We let CBRκ,ε(α1,θ)be the set of all (κ,ε)-confirming best responses at (α1,θ). 29Fudenberg and Levine (1992) provide a similar definition that uses the notion of total variational distance between probability measures instead of Kullback–Leibler divergence. 30We define ψ(·|λ,α2):=a1,a2,θψ(·|a1,a2,θ)λ(a1,θ)α2(a2). 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
188 Deb and Ishii Theoretical Economics 20 (2025) The following lemma motivates the definition of (κ,ε)-confirming best responses and shows that for εsmall, if short-run players play an (ε,ε)-confirming best response, then the LR player obtains payoffs close to those as if the SR player were best responding with perfect knowledge of both the LR player’s action and state. Lemma 1. For every α1∈A1, liminf ε→0inf α2∈CBRε,ε(α1,θ)u1(α1,α2,θ)≥inf α2∈B2(α1,θ)u1(α1,α2,θ). Proof. By Assumption 1and the property that D(P|Q)=0 if and only if P=Q,wehave that CBR0,0(α1,θ)=B2(α1,θ).Moreover,CBR ε,ε(α1,θ)is upper hemi-continuous with respect to ε, and so the inequality follows. Notice that a (1, ε)-confirming best response is essentially the extension of the idea of ε-entropy confirming best response in Gossner (2011) to the current setting. Under a (1, ε)-confirming best response, condition (iii) in Definition 3is trivially satisfied and so the definition only requires that the public signal distribution associated with the belief λrequired to sustain α2as a best response be ε-close in Kullback– Leibler divergence to the true distribution of public signals under the action profile (α1,α2)and state θ.Whenκis small, condition (iii) additionally requires that λindeed places large probability on the state θ. This additional requirement is important in Lemma 1, since generally, liminfε→0infα2∈CBR1,ε(α1,θ)u1(α1,α2,θ)may be strictly less than infα2∈B2(α1,θ)u1(α1,α2,θ). The following lemma constitutes the key step in the proof of the main theorem, which shows that if the LR deviates to play σβ1in any equilibrium, then the SR plays strategies consistent with (ε,ε)-confirming best responses in all but a finite number of periods with very large probability. Formally, define the following set of histories given an equilibrium σand a type ω∈who plays strategies that only depend on Ht×31 Mσ,(ω,θ)(J,κ,ε):=h∞∈H∞:t:σ2ht/∈CBRκ,εσ1ω,ht,θ,θ<J. These are the set of public histories, h∞,wheretypeωand the SR players together play action profiles that are (κ,ε)-confirming best responses at state θin all but Jperiods. The following lemma provides a lower bound on the probability of such histories that applies uniformly across all equilibria. Lemma 2. Suppose that μ(ωβ1)>0. Then for every ε>0, there exists some Jsuch that infσ∈BNEδπσ,(ωβ1,θ) ∞Mσ,(ωβ1,θ)(J,ε,ε)≥1−2ε.32 31In the analysis, we are concerned with these sets only for types ωβ1who play strategies that only depend on Ht×. Therefore, the restriction to such types is not restrictive. 32Notice that in this lemma, we do not necessarily require that β1∈Sεfor some ε>0 small. Later in the proof of Theorem 1, when we use this lemma, we will use Lemma 2for the particular case in which β1∈Sε for ε>0 small to ensure that by mimicking ωβ1, the LR can ensure high payoffs. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 189 There are two aspects of the lemma above that are worth emphasis. First is that the set of histories in Mσ,(ω,θ)(J,ε,ε)ensures that players play action profiles consistent with (ε,ε)-confirming best responses in all but Jperiods. One could weaken this to analyze the probability of the set of histories in Mσ,(ω,θ)(J,1,ε)that only require players to play action profiles consistent with (1, ε)-confirming best responses in all but Jperiods as in Gossner (2011). Indeed, the arguments of Fudenberg and Levine (1986)andGossner (2011) imply a uniform lower bound on the probability of such histories across all equilibria. However, this is insufficient for our reputation theorems since as previously discussed, Lemma 1does not apply to (1, ε)-confirming best responses. The conclusion of the above lemma does not hold for any arbitrary type ωand holds only for types ωβ1. This is again because the definition of Mσ,(ω,θ)(J,ε,ε)requires SR players to hold approximately correct beliefs on the state θin all but Jperiods. In particular, if ωwere a stationary commitment type, then πσ,(ω,θ) ∞(Mσ,(ω,θ)(J,ε,ε)|ω,θ)may actually be quite small for some equilibria, σ. We prove Lemma 2in Section 5.2. Before this, we present the proof of Theorem 1, which is now immediate. Proof of Theorem 1. Define u:=mina∈Aminθ∈u1(a,θ). and choose any θ.Wewill show that there exists some δ∗<1 such that whenever δ>δ ∗,U1(σ,θ;δ)>u ∗ 1(θ)−ρfor all σ∈BNEδ. This then proves the theorem, since there are finitely many states θ∈. First choose some ε∗>0suchthatforallε<ε ∗, (1−2ε)u∗ 1(θ)−ρ 4+2εu >u ∗ 1(θ)−ρ. By assumption, we can choose β1∈Sρ/8such that μ(ωβ1)>0. By Lemma 1,thereexists some ε∈(0, ε∗)such that u1β1(θ),α2,θ>min α2∈B2(β1(θ),θ)u1β1(θ),α2,θ−ρ 8≥u∗ 1(θ)−ρ 4 for all (β1(θ),α2)that is an (ε,ε)-confirming best-response at θ, where the last inequality follows from construction that β1∈Sρ/8. By Lemma 2, there exists some Jsuch that for every equilibrium, σ, πσ,(ωβ1,θ) ∞Mσ,(ωβ1,θ)(J,ε,ε)≥1−2ε. As a result, in any equilibrium, σ, by mimicking the strategy of the commitment type ωβ1, the LR player 1 obtains at least the payoff (1−2ε)1−δJu+δJu∗ 1(θ)−ρ 4+2εu. Then we can choose some δ∗<1suchthatforallδ>δ ∗, (1−2ε)1−δJu+δJu∗ 1(θ)−ρ 4+2εu >u ∗ 1(θ)−ρ. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
190 Deb and Ishii Theoretical Economics 20 (2025) 5.2 Proving Lemma 2 We now prove our key lemma, which follows in a straightforward manner from the following two lemmas. The complete proof of Lemma 2is provided in Appendix D.To simplify notation, given C⊆×, define φσ t·|ht=margYπσ t·|ht,φσ,C t·|ht=margYπσ t·|ht,C. In words, φσ t(·|ht)is the distribution over yt∈Yin period tin the equilibrium σ,conditional on the public history ht.33 Additionally, φσ,C tis this distribution when conditioned on the event (ω,θ)∈C. Lemma 3 (Merging). Suppose that γ0(ω,θ)>0. Then for every ε>0, there exists some J1 such that in every equilibrium σ, πσ,(ω,θ) ∞h∞∈H∞:t:Dφσ,(ω,θ) t·|htφσ t·|ht>ε <J 1≥1−ε. Lemma 4 (Uniform Learning). Suppose μ0(ωβ1)>0. Then for every ε>0,thereexists some J2such that for all σ∈BNEδ, πσ,(ωβ1,θ) ∞h∞∈H∞:t:νσ tθ|ht<1−ε<J 2≥1−ε. As in Fudenberg and Levine (1986)andGossner (2011), Lemma 3strengthens the classical merging results, e.g., Blackwell and Dubins (1962)andKalai and Lehrer (1993), by establishing a uniform upper bound across all equilibria on the probability of histories in which the SR player’s prediction of today’s public signal distribution, φσ t(·|ht), diverges substantially from the “true” public signal distribution, φσ,(ω,θ) t(·|ht),inmore than J1time periods, when LR plays σ(ω)in state θ. The proof follows using standard merging arguments of Gossner (2011), which we include for completeness in Appendix C. To prove Lemma 4, we show that in any state θ, by playing σβ1(θ), the LR player can ensure that the SR players learn the state θat a rate that is uniform across all equilibria. Indeed standard arguments immediately imply that in any equilibrium, SR players learn the true state θwhenever the LR player plays σβ1(θ). However, the additional uniformity requirement requires further analysis, which we now address in Section 5.3. 5.3 A robust learning theorem Consider the following general model of learning. There is a finite signal space Yand a countable state space .Alearning environment is some π∈S(Y,)for which π:= margπhas full support on . Recall that for any B⊆,πB∈S(Y,)denotes the stochastic process conditional on ξ∈B:πB=(πt(·|B))∞ t=0. Note that this allows the stochastic process, πξ,foranyξ∈, to be very general, which may potentially contain arbitrary forms of serial correlations. 33In fact, φσ t(·|ht)is the SR players’ subjective belief of the period tpublic signal after observing ht. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 191 To interpret, in a learning environment, at the beginning of each period t=1, 2, , an observer updates her beliefs about the true state ξ∈according to Bayes’ rule upon the realization of a history of signals ht=(y0,,yt−1).Letρπ t(·|ht)∈()denote the observer’s beliefs after observing ht. We now describe formally our definition of robust learning. Definition 4. Let ξ∗∈B⊆and S∗⊆S(Y,). Then we say that an observer S∗- robustly learns Bat ξ∗if for every κ∈(0, 1), there exists some Ksuch that inf π∈S∗π∞∞ t=Kh∞:ρπ tB|ht≥1−κ|ξ∗≥1−κ. Intuitively, S∗-robust learning requires an observer’s beliefs to concentrate on Bforever after period Kwith high probability for all learning environments in S∗. Our main theorem in this section establishes a simple sufficient condition on S∗ that guarantees S∗-robust learning of Bat ξ∗. To state it, we first need a few definitions that are well known from the theory of statistical experiments. First fix a learning environment π∈S(Y,),someξ∗∈,andB⊆. We now define the function Hπ t(·;B,ξ∗):[0, 1]→R, which is also known as the Hellinger transform. Formally this function is defined as Hπ tz;B,ξ∗:= ht∈HtπB thtzπξ∗ tht1−z=Eπξ∗ tπB tht πξ∗ thtz. This is the moment generating function of the (random) log-likelihood ratio at time t, log πB t(ht) πξ∗ t(ht),whenhtis distributed according to πξ∗ t. Toward our robust learning result, let us also define Hπ tB,ξ∗=inf z∈[0,1]Hπ tz;B,ξ∗∈[0, 1]. Roughly speaking, Hπ t(B,ξ∗)measures the informativeness of the learning environment at time twith respect to learning the relative likelihoods of Bvs. ξ∗.Noticethatby Jensen’s inequality, Hπ t(B,ξ∗)≤1. Intuitively, a completely uninformative learning environment attains this maximal value of Hπ t(B,ξ∗)=1. On the other hand, if the supports of πB tand πξ∗ tare disjoint so that the learning environment distinguishes Bfrom ξ∗perfectly, then Hπ t(B,ξ∗)=0. In Appendix A, we list some additional useful properties of the Hellinger transform.34 In the following theorem, we show that when the Hellinger transforms converge to zero (information converges to perfect information) at a fast enough rate uniformly across all learning environments, π∈S∗, then the observer S∗-robustly learns Bat ξ∗. 34See also Torgersen (1991) and Moscarini and Smith (2002) for more details on the Hellinger transform. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
192 Deb and Ishii Theoretical Economics 20 (2025) Theorem 2. Let S∗⊆S(Y,)and ξ∗∈B⊆. Suppose that infπ∈S∗π(ξ∗)>0and lim K→∞ sup π∈S∗ ∞ t=K Hπ tBc,ξ∗=0. Then an observer S∗-robustly learns Bat ξ∗.35 The following corollary will be useful: It shows that if we can guarantee S∗-robust learning of a finite collection of sets at ξ∗, then we can also guarantee S∗-robust learning of the intersection of these sets at ξ∗. Corollary 1. Let ξ∗∈B1,,Bn⊆and S∗⊆S(Y,). Suppose that infπ∈S∗π(ξ∗)> 0and that for all =1, 2, ,n, lim K→∞ sup π∈S∗ ∞ t=K Hπ tBc ,ξ∗=0. Then the observer S∗-robustly learns B1∩B2∩···∩Bnat ξ∗. 5.3.1 Uniform signaling of the state in reputation building Given any equilibrium, σ, the SR players face a learning environment about the state space =×along the samelinesasinSection5.3. Of course, when we view an equilibrium, σ, as a learning environment, we can also define the appropriate Hellinger transforms. Thus, for any equilibrium σand any event A⊆×, we define the Hellinger transform as Hσ tz;B,(ω,θ)= ht∈Htπσ,B thtzπσ,(ω,θ) tht1−z. We also accordingly define Hσ tB,(ω,θ)=inf z∈[0,1]Hσ tz;B,(ω,θ). Through a straightforward computation in Lemma 8in Appendix B, we show that for any θ= θ, lim K→∞ sup σ∈BNEδ ∞ t=K Hσ t×θ,ωβ1,θ=0. By Corollary 1, the SR players BNEδ-robustly learn θ=θ×(\{θ})=×{θ}at (ωβ1,θ), which proves Lemma 4. 35We leave open the question of whether this condition is also necessary for S∗-robust learning for future research. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 199 Lemma 10. Suppose that γ0(ω,θ)>0. Then for every σ∈BNEδ, πσ,(ω,θ) ∞h∞∈H∞:t:Dφσ,(ω,θ) t·|htφσ t·|ht>ε ≥J≤−logγ0(ω,θ) Jε . Proof. For every T, by the chain rule for Kullback–Leibler divergence, DmargHTπσ,(ω,θ) TmargHTπσ T=Eπσ,(ω,θ) TT t=0 Dφσ,(ω,θ) t·|htφσ t·|ht =Eπσ,(ω,θ) ∞T t=0 Dφσ,(ω,θ) t·|htφσ t·|ht. Moreover, D(margHTπσ,(ω,θ) TmargHTπσ t)≤−log γ0(ω,θ)by the previous lemma. Therefore, by the monotone convergence theorem, Eπσ,(ω,θ) ∞∞ t=0 Dφσ,(ω,θ) t·|htφσ t·|ht≤−logγ0(ω,θ). Then by Markov’s inequality, πσ,(ω,θ) ∞h∞∈H∞:t:Dφσ,(ω,θ) t·|htφσ t·|ht>ε >J ≤πσ,(ω,θ) ∞∞ t=0 Dφσ,(ω,θ) t·|htφσ t·|ht>Jε ≤−log γ0(ω,θ) Jε . The proof of Lemma 3is now immediate. Proof of Lemma 3. Choose J1sufficiently large such that −log γ0(ω,θ) J1ε<ε.Then Lemma 3is immediate from Lemma 10. Appendix D: Proof of Lemma 2 By Lemmas 3and 4,thereexistJsuch that for all σ∈BNEδ, 1−ε≤πσ,(ωβ1,θ) ∞h∞:t:Dφσ,(ω,θ) t·|htφσ t·|ht>ε <J, 1−ε≤πσ,(ωβ1,θ) ∞h∞:t:νσ tθ|ht<1−ε<J. Therefore, for all σ∈BNEδ, πσ,(ωβ1,θ) ∞Mσ,(ωβ1,θ)(2J,ε,ε) ≥πσ,(ωβ1,θ) ∞h∞:t:Dφσ,(ω,θ) t·|htφσ t·|ht>ε ,t:νσ tθ|ht<1−ε<J ≥1−2ε. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
200 Deb and Ishii Theoretical Economics 20 (2025) Appendix E: Proving Theorem 3 The proof of Theorem 3uses ideas from Mertens, Sorin, and Zamir (2014)withsome modifications. Let us begin with some notation. Given any probability vector x∈(), let xdenote the Euclidean norm: x2= θ∈ x(θ)2. Note that if player 1 plays a strategy that induces λ∈(A1×)as the joint distribution over A1×and player 2 plays a2, then player iobtains the expected utility ui(a2,λ):=Eλui(a1,a2,θ)= a1,θ ui(a1,a2,θ)λ(a1,θ). We now extend the definition of a best response to ε-best response: BRε 2(λ):=a2∈A2:max a 2∈A2 u2a 2,λ−u2(a2,λ)≤ε. Define for any ε≥0, Wε(λ)=max a2∈Bε 2(λ)u1(a2,λ). Finally, given λ∈(A1×),letq(·|y,λ)be the induced posterior belief about θafter observation of the signal y: q(θ|y,λ)= a1∈A1 λ(a1,θ)ψ(y|a1,θ) θ∈ a1∈A1 λa1,θψy|a1,θ. Proposition 1. For every ε>0, there exists some ρ>0such that E q(·|y,λ)−margλ 2<ρ⇒W0(λ)≤cavV(margλ)+ε. See Appendix Hfor the proof. The following lemma provides a uniform bound (across all equilibria) on the number of times where the expected movement (in terms of · 2distance) in the SR players’ beliefs is greater than ε. Lemma 11. For any σ∈BNEδand any ε>0, t:Eπσ ∞ νσ t+1ht+1−νσ tht 2≥ε≤1 ε. Proof. Consider any time t+1: Eπσ ∞ νσ t+1ht+1−ν0 2=Eπσ ∞ νσ t+1ht+1 2−ν02≤1. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 201 By the martingale property of beliefs, πσ ∞-almost surely, Eπσ ∞[νσ τ+1(hτ+1)|hτ]=νσ τ(hτ). Therefore, it is straightforward to show that 1≥Eπσ ∞ νσ t+1ht+1 2−ν02]= t τ=0 Eπσ ∞ νσ τ+1hτ+1−νσ τhτ 2. Sincetheaboveholdsforeveryt, it implies that ∞ τ=0 Eπσ ∞ νσ τ+1hτ+1−νσ τhτ 2≤1, which implies the claim. We can now prove Theorem 3. Proof of Theorem 3. We first provide an upper bound on Eπσ ∞(1−δ) ∞ t=0 δtu1at 1,at 2,θ that holds across all σ∈BNEδ. Notice that the above payoff is not equal to U1(σ,θ;δ), since the expectation does not condition on ωs. However, one can interpret the payoff above as follows. For any equilibrium σ∈BNEδ,let¯σ1denote the strategy, where the LR player fictitiously draws some ω∈according to μ0and plays the strategy in the equilibrium, σ, associated with that type for the entirety of the repeated game.41 Indeed U1(¯σ1,σ2;δ)corresponds to the payoff above. By Proposition 1, there exists some ρ>0suchthat Eλ q(·|y,λ)−p 2<ρ⇒W0(λ)≤cavV(margλ)+ε/8. Choose n∈Nsuch that 1 n(u−u)<ε/8andδ∗such that for all δ>δ ∗, 1−δnm ρu+δnm ρcavV(ν)+ε 4<cavV(ν)+ε 2.(2) For any σ∈BNEδ,let Tσ:=t:Eπσ ∞ νσ t+1·|ht+1−νσ t·|ht 2≥ρ n. For all t/∈Tσ, by Markov’s inequality, we have πσ ∞ νσ t+1·|ht+1−νσ t·|ht 2≥ρ≤1 n. 41For example, if the LR player draws a commitment type ω, then the LR player plays σω. If instead the LR player indeed draws ωs, then the LR player simply plays σ1. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
202 Deb and Ishii Theoretical Economics 20 (2025) Figure 7. Quality choice. Thus, at all t/∈Tσ, Eπσ ∞W0νσ t·|ht≤1 n(u−u)+Eπσ ∞cavVνσ t·|ht+ε/8≤cavV(ν0)+ε/4. Therefore, U1(¯σ1,σ2;δ)≤(1−δ) ∞ t=0 δtEπσ ∞W0(νσ t·|ht =(1−δ) t∈Tσ δtu+ t/∈Tσ δtEπσ ∞W0νσ t·|ht ≤(1−δ) t∈Tσ δtu+ t/∈Tσ δtcavV(ν0)+ε/4. By Lemma 11, for every σ∈BNEδ,|Tσ|≤nm/ρ. Therefore, for all δ>δ ∗and any σ∈ BNEδ, U1(¯σ1,σ2;δ)≤1−δnm/ρu+δnm/ρcavV(ν0)+ε/4<cavV(ν0)+ε/2. Finally, note that U1(¯σ1,σ2;δ)≥1−μ0cU1(σ1,σ2;δ)+μ0cu. Let χ∗>0besuchthatforallχ<χ ∗, 1 1−χcavV(ν0)+ε 2−χu<cavV(ν0)+ε. Thus, for all δ>δ ∗and μ0(c)<χ ∗, U1(σ1,σ2;δ)≤1 1−μccavV(ν0)+ε 2−μ0cu<cavV(ν0)+ε. Appendix F: Example The following example shows that the probability of commitment types matters for the upper bound even when δis close to 1. Consider the quality choice game with the stage game payoffs given by Figure 7. In the repeated game this stage game is repeatedly played and all payoffs are common knowledge. Note that the Stackelberg payoff of the above game is 3/2. Furthermore, note that Bis a best response for the SR player in the stage game if and only if α1(C)≥1/2. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 203 Figure 8. The information structure. There are two states ={1, −1}that only affect the signal distribution of the public signal. There are two types in the game: ={ωc,ωs}. The commitment type, ωc,inthis game is a type that always plays the mixed action, 2 3H⊕1 3L, regardless of the state.42 In particular, we assume that the probability of each state is identical and the probability of the commitment type is μ∈(0, 1). The signal space is binary, Y={¯ y,y}and the information structure is given by Fig. 8. Note that according to this information structure, (2 3H⊕1 3L,θ)is statistically indistinguishable from (L,−θ):ψ(·|2 3H⊕1 3L,θ)=ψ(·|L,−θ). In this example, we have the following observation. Claim 3. There exists μ∗such that for all μ>μ ∗and any δ∈(0, 1), there exists an equilibrium in which the strategic player obtains a payoff of 2in both states. Proof. Consider the candidate equilibrium strategy profile in which the strategic LR player always plays L. Choose μ∗=3 4. Then we will show that when μ>μ ∗, this strategy profile is indeed an equilibrium for any δ∈(0, 1). Consider the incentives of the SR player. To study this, we want to compute the probability that the SR player assigns to action Tgiven the candidate equilibrium strategy of the LR player: λσ tH|ht=2 3μσ tωc|ht=2 3γσ tωc,1|ht+γσ tωc,−1|ht. Consider the likelihood ratio γσ tωc,θ|ht γσ tωs,−θ|ht=γσ tωc,θ|h0 γσ tωs,−θ|h0=μ 1−μ. This then implies that for all ht,μ(ωc|ht)=μ,μ(ωs|ht)=1−μ. Thus, for all htand all μ>μ ∗, λσ tH|ht=2 3μ>1 2. This then implies that for all ht, the SR player’s best response is to play L. Furthermore, because the SR player is playing the same action at all histories, the strategic LR player’s best response is to play Bat all histories. Thus, the proposed strategy profile is indeed 42Note that this is in reality not the mixed Stackelberg action. However, by appropriately modifying the information structure, the same conclusions hold, even if the commitment type plays some other mixed action in every period. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
204 Deb and Ishii Theoretical Economics 20 (2025) an equilibrium. Furthermore, according to this strategy profile, the strategic LR player’s payoff is 2 in both states, concluding the proof. The above discussion shows that when the commitment type occurs with large probability, even an arbitrarily patient strategic LR player obtains a payoff strictly greater than the Stackelberg payoff in equilibrium. We now examine an upper bound when the commitment type probability is small. Claim 4. Let ε>0. Then there exists some μ∗>0and δ∗<1such that for all μ<μ ∗and δ>δ ∗,U1(σ,δ)<3/2+εfor all σ∈BNEδ. Proof.ConsiderV(p)for any p∈(). Because the stage game utilities are stateindependent, it is straightforward to show that V(p)≤sup α1∈A1 max a2∈B2(α1)u1(α1,a2)=3/2, where the equality follows from a straightforward calculation. The claim then follows from Theorem 3. Appendix G: Proof of Corollary 2 The lower bound is a consequence of Theorem 1. Letusnowshowtheupperbound. Choose some ν∈(0, minθ∈ν0(θ)). Suppose by way of contradiction that there exists some state θ∗∈and some sequence δn→1andσn∈BNEδnsuch that for all n,U1(σn,θ∗;δn)≥u∗ 1(θ∗)+ε.ByTheorem 3, ν0θ∗u∗ 1θ∗+ε+lim sup n→∞ θ=θ∗ ν0(θ)U1σn,θ;δn< θ∈ ν0(θ)u∗ 1(θ)+νε. Together with Theorem 1,wehave θ=θ∗ ν0(θ)u∗ 1(θ)≤lim sup n→∞ θ=θ∗ ν0(θ)U1σn,θ;δn≤ θ=θ∗ ν0(θ)u∗ 1(θ)−ν0θ∗−νε, but this is a contradiction. Appendix H: Proving Proposition 1 Let us first define the set ˆ NR(p):=λ∈(A1×):λ(·|θ)θ∈∈NR(p),margλ=p, ˆ NR := p∈() ˆ NR(p). 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 205 Notice that V(p)=supλ∈ˆ NR(p)W0(λ). Analogously, we can define for any ε>0, Vε(p)=sup λ∈ˆ NR(p) Wε(λ). Finally, define also for every ε≥0, ε(a2):=λ∈ˆ NR : a2∈Bε 2(λ). We begin with some lemmas. Lemma 12. Let ε>0. Then there exists some ρ>0such that for all λ∈(A1×), E q(·|y,λ)−margλ 2<ρ⇒inf ˆ λ∈ˆ NR λ−ˆ λ<ε. See Lemma V.3.6 in Mertens, Sorin, and Zamir (2014) for the proof. Lemma 13. Let ε>0. Then there exists some ρ>0such that for all λ,ˆ λ∈(A1×), λ−ˆ λ<ρ⇒W0(λ)≤Wε(ˆ λ)+ε. Proof.Letε>0. First choose ρ>0 sufficiently small such that λ−ˆ λ<ρ ⇒max a2∈A2u2(a2,λ)−u2(a2,ˆ λ)≤ε. Then there exists some ρ∈(0, ρ)such that λ−ˆ λ≤ρ=⇒ B0 2(λ)⊆Bε 2(ˆ λ). Therefore, whenever λ−ˆ λ≤ρ, W0(λ)=max a2∈B0 2(λ) u1(a2,λ)≤max a2∈Bε 2(ˆ λ) u1(a2,λ)≤max a2∈Bε 2(ˆ λ) u1(a2,ˆ λ)+ε=Wε(ˆ λ)+ε. Lemma 14. For every ε>0, there exists some ρ>0such that for all a2∈A2, λ∈ρ(a2)⇒inf λ∈0(a2) λ−λ <ε. Proof. Suppose otherwise. Then for some a2∈A2and ε>0, there exists some sequence ρn→0andλn∈ρn(a2)such that inf λ∈0(a2) λn−λ ≥ε.(3) By Bolzano–Weierstrass, without loss of generality, by replacing the original sequence with an appropriate subsequence, we can assume this sequence to be convergent to some limit λ. However, note that since λn→λand λn∈ρn(a2)for all n,λ∈0(a2). This contradicts (3). Lemma 15. For every ε>0, there exists some ρ∗>0such that for all λ∈ˆ NR and all ρ< ρ∗, Wρ(λ)≤cavV(margλ)+ε. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
206 Deb and Ishii Theoretical Economics 20 (2025) Proof. First, because cavVand u1(·,a2)are Lipschitz continuous for all a2∈A2,there exists some ε>0 such that whenever λ−λ<ε ,then cavV(margλ)−cavVmargλ,max a2∈A2u1(a2,λ)−u1a2,λ<ε/2. By the previous lemma, let ρ>0besuchthatforalla2∈A2, λ∈ρ(a2)⇒inf λ∈0(a2) λ−λ <ε . Recall that Wρ(λ)=max a2∈Bρ 2(λ) u1(a2,λ). Let aρ 2(λ)∈Bρ 2(λ)be the solution to the above maximization problem. Thus, for every λ∈ˆ NR, λ∈ρ(aρ 2(λ)). Therefore, for all λ∈ˆ NR, there exists some λ(λ)∈0(aρ 2(λ))with λ−λ(λ)≤ε. Then for any λ∈ˆ NR, Wρ(λ)=u1aρ 2(λ),λ≤max a2∈B0 2(λ(λ)) u1(a2,λ) ≤W0λ(λ)+ε/2 ≤cavVmargλ(λ))+ε/2≤cavV(margλ)+ε. We can now prove Proposition 1. Proof of Proposition 1. By Lemma 15, there exists some ρ∗∈(0, ε/3)such that for all ˆ λ∈ˆ NR, Wρ∗(ˆ λ)≤cavV(margˆ λ)+ε/3. By Lemma 13 and Lipschitz continuity of cavV, there exists some ρ>0suchthat λ−λ <ρ ⇒W0(λ)≤Wρ∗λ+ρ∗,cavV(margλ)−cavVmargλ<ε/3. By Lemma 12,thereexistsρ>0suchthatforallλfor which E[q(·|y,λ)−margλ2]< ρ,thereexistsˆ λ(λ)∈ˆ NR such that ˆ λ(λ)−λ<ρ . Thus, for any λin which E[q(·|y,λ)−margλ2]<ρ,wehave W0(λ)≤Wρ∗ˆ λ(λ)+ρ∗≤cavV(margˆ λ(λ)+2ε/3≤cavV(marg(λ)+ε. References Acemoglu, Daron, Victor Chernozhukov, and Muhamet Yildiz (2016), “Fragility of asymptotic agreement under Bayesian learning.” Theoretical Economics, 11, 187–225. [0173] Al-Najjar, Nabil I. (2009), “Decision makers as statisticians: Diversity, ambiguity, and learning.” Econometrica, 77, 1371–1401. [0173] 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Theoretical Economics 20 (2025) Reputation building 207 Al-Najjar, Nabil I. and Mallesh M. Pai (2014), “Coarse decision making and overfitting.” Journal of Economic Theory, 150, 467–486. [0173] Aoyagi, Masaki (1996), “Reputation and dynamic Stackelberg leadership in infinitely repeated games.” Journal of Economic Theory, 71, 378–393. [0172] Atakan, Alp E. and Mehmet Ekmekci (2011), “Reputation in long-run relationships.” The Review of Economic Studies, rdr037. [0172] Atakan, Alp E. and Mehmet Ekmekci (2015), “Reputation in the long-run with imperfect monitoring.” Journal of Economic Theory, 157, 553–605. [0172] Aumann, Robert J., Michael Maschler, and Richard E. Stearns (1995), “Repeated games with incomplete information.” MIT press. [0172] Blackwell, David and Lester Dubins (1962), “Merging of opinions with increasing information.” The Annals of Mathematical Statistics, 33, 882–886. [0171,0190] Celentani, Marco, Drew Fudenberg, David K. Levine, and Wolfgang Pesendorfer (1996), “Maintaining a reputation against a long-lived opponent.” Econometrica, 64, 691–704. [0172] Cripps, Martin W., Eddie Dekel, and Wolfgang Pesendorfer (2005), “Reputation with equal discounting in repeated games with strictly conflicting interests.” Journal of Economic Theory, 121, 259–272. [0172] Cripps, Martin W., George J. Mailath, and Larry Samuelson (2004), “Imperfect monitoring and impermanent reputations.” Econometrica, 72, 407–432. [0195] Ely, Jeffrey C., Drew Fudenberg, and David K. Levine (2008), “When is reputation bad?” Games and Economic Behaivor, 63, 498–526. [0179] Ely, Jeffrey C. and Juuso Välimäki (2003), “Bad reputation.” Quarterly Journal of Economics, 118, 785–814. [0179] Evans, Robert and Jonathan P. Thomas (1997), “Reputation and experimentation in repeated games with two long-run players.” Econometrica, 65, 1153–1173. [0172] Fudenberg, Drew and David K. Levine (1986), “Limit games and limit equilibria.” Journal of Economic Theory, 38, 261–279. [0187,0189,0190] Fudenberg, Drew and David K. Levine (1989), “Reputation and equilibrium selection in games with a patient player.” Econometrica, 57, 759–778. [0172] Fudenberg, Drew and David K. Levine (1992), “Maintaining a reputation when strategies are imperfectly observed.” Review of Economic Studies, 59, 561–579. [0170,0172,0174, 0187,0193] Fudenberg, Drew and David K. Levine (1993a), “Self-confirming equilibrium.” Econometrica, 61, 523–546. [0173] Fudenberg, Drew and David K. Levine (1993b), “Steady state learning and Nash equilibrium.” Econometrica, 61, 547–574. [0173] 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
208 Deb and Ishii Theoretical Economics 20 (2025) Fudenberg, Drew and Yuichi Yamamoto (2010), “Repeated games where the payoffs and monitoring structure are unknown.” Econometrica, 78, 1673–1710. [0172] Ghosh, Sambuddha (2014), “Multiple long-lived opponents and the limits of reputation.” Report. [0172] Gossner, Olivier (2011), “Simple bounds on the value of a reputation.” Econometrica, 79, 1627–1641. [0171,0172,0187,0188,0189,0190,0193,0198] Hörner, Johannes and Stefano Lovo (2009), “Belief-free equilibria in games with incomplete information.” Econometrica, 77, 453–487. [0172] Hörner, Johannes, Stefano Lovo, and Tristan Tomala (2011), “Belief-free equilibria in games with incomplete information: Characterization and existence.” Journal of Economic Theory, 146, 1770–1795. [0172] Kalai, Ehud and Ehud Lehrer (1993), “Rational learning leads to Nash equilibrium.” Econometrica, 61, 1019–1045. [0190] Kreps, David M. and Robert J. Wilson (1982), “Reputation and imperfect information.” Journal of Economic Theory, 27, 253–279. [0172] Mertens, Jean-François, Sylvain Sorin, and Shmuel Zamir (2014), Repeated Games,volume 55. Cambridge University Press. [0172,0193,0200,0205] Milgrom, Paul R. and John Roberts (1982), “Predation, reputation and entry deterrence.” Journal of Economic Theory, 27, 280–312. [0172] Moscarini, Giuseppe and Lones Smith (2002), “The law of large demand for information.” Econometrica, 70, 2351–2366. [0171,0173,0191,0195] Mu, Xiaosheng, Luciano Pomatto, Philipp Strack, and Omer Tamuz (2021), “From Blackwell dominance in large samples to Rényi divergences and back again.” Econometrica, 89, 475–506. [0173,0195] Pei, Harry (2020), “Reputation effects under interdependent values.” Econometrica,88, 2175–2202. [0173] Schmidt, Klaus M. (1993), “Reputation and equilibrium characterization in repeated games of conflicting interests.” Econometrica, 61, 325–351. [0172] Torgersen, Erik (1991), Comparison of Statistical Experiments, volume 36. Cambridge University Press. [0171,0191,0195] Wiseman, Thomas (2005), “A partial folk theorem for games with unknown payoff distributions.” Econometrica, 73, 629–645. [0172] Co-editor Simon Board handled this manuscript. Manuscript received 25 January, 2022; final version accepted 13 June, 2024; available online 25 June, 2024. 15557561, 2025, 1, Downloaded from https://onlinelibrary.wiley.com/doi/10.3982/TE4758 by ZBW Kiel - Hamburg (German National Library of Economics), Wiley Online Library on [04/07/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
