scieee AI-readable full text Open interactive document viewer

Pathwise concentration bounds for Bayesian beliefs

Fudenberg, Drew,Lanzani, Giacomo,Strack, Philipp

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Fudenberg, Drew; Lanzani, Giacomo; Strack, Philipp Article Pathwise concentration bounds for Bayesian beliefs Theoretical Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Fudenberg, Drew; Lanzani, Giacomo; Strack, Philipp (2023) : Pathwise concentration bounds for Bayesian beliefs, Theoretical Economics, ISSN 1555-7561, The Econometric Society, New Haven, CT, Vol. 18, Iss. 4, pp. 1585-1622, https://doi.org/10.3982/TE5206 This Version is available at: https://hdl.handle.net/10419/296448 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Theoretical Economics 18 (2023), 1585–1622 1555-7561/20231585 Pathwise concentration bounds for Bayesian beliefs Drew Fudenberg Department of Economics, MIT Giacomo Lanzani Department of Economics, MIT Philipp Strack Department of Economics, Yale University We show that Bayesian posteriors concentrate on the outcome distributions that approximately minimize the Kullback–Leibler divergence from the empirical distribution, uniformly over sample paths, even when the prior does not have full support. This generalizes Diaconis and Freedman’s (1990) uniform convergence result to, e.g., priors that have finite support, are constrained by independence assumptions, or have a parametric form that cannot match some probability distributions. The concentration result lets us provide a rate of convergence for Berk’s (1966) result on the limiting behavior of posterior beliefs when the prior is misspecified. We provide a bound on approximation errors in “anticipated-utility” models, and extend our analysis to outcomes that are perceived to follow a Markov process. Keywords. Misspecified learning, Bayesian consistency. JEL classification. C11, D81. 1. Introduction Learning from repeated observations is a key feature of many economic settings, and almost all economic studies of learning model it as Bayesian inference. To understand the medium and long-run implications of Bayesian learning, it is useful to know how quickly beliefs concentrate around the data generating processes that best explain the observations. Our main result, Theorem 1, shows that the probability the posterior assigns to distributions that do not approximately maximize the likelihood assigned to the data vanishes exponentially quickly. Importantly, we identify conditions for this to hold not only with high probability, but for every possible realization of the data. More specifically, Theorem 1establishes that for every ε>0 the posterior probability of the Drew Fudenberg: [email protected] Giacomo Lanzani: [email protected] Philipp Strack: [email protected] We thank Simone Cerreia-Vioglio, Xiaohong Chen, Roberto Corrao, Jana Gieselmann, Paul Heidhues, Yuhta Ishii, Mats Köster, Pooya Molavi, and Todd Sarver for helpful comments and conversations, and National Science Foundation grant SES 1951056, the Guido Cazzavillan Scholarship, and the Sloan Foundation for financial support. ©2023 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at https://econtheory.org.https://doi.org/10.3982/TE5206 1586 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) distributions that do not ε-minimize the Kullback–Leibler (KL) divergence vanishes at an exponential rate. In contrast to earlier pathwise concentration bounds, our result holds even if the agent’s prior does not have full support, or satisfies parametric restrictions, and regardless of the true data generating process. Our results generalize Diaconis and Freedman (1990), which showed that a φpositivity condition implies that Bayesian posteriors converge to the empirical distribution at a uniform exponential rate. This condition requires that the support of the agent’s prior includes every distribution over outcomes, and thus rules out many settings of economic interest in which the set of outcome distributions is naturally restricted. For example, it does not apply to agents whose prior has finite support (which is natural in settings such as urn problems with a finite number of balls), agents who each period observe a set of Bernoulli trials that they think are i.i.d., or agents who believe (mistakenly or not) that some variables are positively correlated. In addition, φ-positivity rules out all cases where the support of the agent’s prior does not contain the true data generating process, so that the agent is misspecified. Theorem 1guarantees that beliefs concentrate on the approximate KL minimizers for the empirical frequency. We show that this is equivalent to concentration on a ball around the exact KL minimizers when priors have full support, but not in general. Moreover, since the KL minimizer is not unique, the theorem does not imply that beliefs converge. We use Theorem 1to prove Theorem 2, which provides a rate of convergence for Berk’s (1966) result that posterior beliefs concentrate around the Kullback–Leibler minimizers with respect to the true data generating process. Berk’s result, like our paper, is stated for an exogenous data generating process. It was extended to learning from endogenous data by Esponda and Pouzo (2016), which led to a renewed interest in misspecified learning in the economics literature.1A key step in the analysis of such models is often to establish that Bayesian beliefs concentrate quickly around the KL minimizers, and as our Theorem 1holds pathwise it immediately implies such a result.2In Fudenberg, Lanzani, and Strack (2021b), we use the concentration result to characterize the long-run beliefs of a correctly specified agent who has an imperfect and selective memory. Theorem 3extends Theorem 1to beliefs that result from observing a Markov process whose transition probabilities are unknown. This complements recent work by Molavi (2019)andEsponda and Pouzo (2021), which studied analogs of Berk (1966)and Esponda and Pouzo (2016) for Markovian environments. 1Subsequent papers include Fudenberg, Romanyuk, and Strack (2017), Molavi (2019), Bohren and Hauser (2021), Fudenberg and Lanzani (2023), He and Libgober (2021), Esponda, Pouzo, and Yamamoto (2021), Heidhues, K˝ oszegi, and Strack (2021), Levy, Moreno de Barreda, and Razin (2021), He (2022), and Frick, Iijima, and Ishii (2023). Before this, Arrow and Green (1973) gave the first general framework for this problem, and Nyarko (1991) pointed out that the combination of misspecification and endogenous observations can lead to cycles. In a setting with finitely many states, Frick, Iijima, and Ishii (2022) provide a convergence-in-probability result on the relative speed at which beliefs converge to the truth for agents with different likelihood functions; Frick, Iijima, and Ishii (2021) extend this to endogenous data. 2For example, Theorem 1provides a shorter way to prove Proposition 1 of Fudenberg, Lanzani, and Strack (2021a). Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1587 Most of the paper assumes that the set of outcomes is finite, as it is in Diaconis and Freedman (1990). Section 6discusses the extent to which the results extend to settings with infinitely many outcomes. It also uses our results to show that the play of a Bayesian agent converges to that predicted by the anticipated utility model (e.g., Kreps (1998), Preston (2005), Eusepi and Preston (2018)), and quantifies its rate of convergence. This clarifies when the anticipated utility model is a good approximation of rational play, and complements numerical studies by Cogley, Colacito, and Sargent (2007), Cogley and Sargent (2008), and Cogley, Colacito, Hansen, and Sargent (2008). 1.1 Theimportanceofpathwiseconcentration The statistical literature has many concentration results for beliefs; see, e.g., Shen and Wasserman (2001). Our results differ in two important ways. First, ours hold for every sample realization, and thus even when the true data generating process is time-varying and endogenous, while the existing statistics results show that beliefs concentrate with probability converging to 1 with respect to a fixed-data generating process. Second, the statistics results are for concentration around the parameters or distributions that minimize the KL divergence from the true-data generating process, while our result are for concentration around the KL-minimizers with respect to an arbitrary empirical distribution.3 Pathwise concentration has played an important role in a number of economic applications, starting with the analysis of nonequilibrium learning in games in Fudenberg and Levine (1993).4The result has also been used to analyze selective attention (Schwartzstein (2014)), the merging of opinions (Acemoglu, Chernozhukov, and Yildiz (2016)), recursive utility functions (Al-Najjar and Shmaya (2019)), and persuasion (Schwartzstein and Sunderam (2021)). To help motivate our analysis, we describe why pathwise concentration (rather than concentration in probability) is needed in four papers on very different problems. Fudenberg and Levine (1993) study the steady states of a model of nonequilibrium learning. The uniform concentration result implies that agents play myopically at any information set hthat has been reached many times. Because which information sets are reached is endogenous, the proof of the main theorem uses the pathwise concentration to rule out the possibility that posteriors only concentrate conditional to histories that induce the player to make choices that prevent reaching hmany times. This is not ruled out by concentration in probability, because the sets of histories under which the information set is reached less than Ntimes could have probability approaching 1 as N grows. Al-Najjar and Shmaya (2019) provide a representation result for Epstein–Zin preferences over stochastic consumption streams for patient agents. To do so, they bound 3Most of these papers also assume that the prior is correctly specified, so that the unique KL minimizer is the true distribution, but see Kleijn and Van der Vaart (2012) and the references therein for generalizations to misspecified priors. 4Subsequent applications to learning in games are Fudenberg and Levine (2006), Fudenberg and He (2018), Gonçalves (2020), and Clark and Fudenberg (2021). Clark, Fudenberg, and He (2022) apply our generalization of the result. 1588 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) the distance between the certain equivalents of period-tconsumption with period t−1 and period 0 information. The (relative) impact of small-probability events on utility increases in the high-patience limit, so the representation result needs the posterior consumption variance to vanish uniformly over all sample paths, which follows from the uniform concentration of the posterior. Gonçalves (2020) introduces an equilibrium concept for games that allows the agents to sample from the opponents’ strategies at a cost before playing. To show existence of the equilibrium, the maximization problem of the agent is transformed into an optimal stopping problem. There the uniform concentration result guarantees that the stopping time is uniformly bounded by a deterministic horizon, thus transforming an infinite-horizon problem into a finite-horizon one that is then solved by backward induction. Finally, Theorem 1can be used to study the limit points of misspecified learning when the distribution of outcomes depends on the action played by an agent, and that action depends on the agent’s beliefs. For example, the agent might be a customer who wants to learn which of two products she prefers, decides every period which one to buy, and receives a signal about the product they bought. To understand if the action acan be played in the long run, we need to understand whether the resulting process of beliefs makes it optimal to play a.Fudenberg, Lanzani, and Strack (2021a) showed that a limit action must be a best reply to all of the associated KL minimizers when the prior has subexponential decay. An earlier version of this paper, Fudenberg, Lanzani, and Strack (2022), use the results here to give a simpler and more transparent proof of this result. 1.2 The importance of relaxing full support The following are examples of commonly studied situations where uniform concentration results that require a full support prior are not applicable, but our results apply. Finite support priors In some problems, the reasonable priors have finite support as, e.g., if outcomes correspond to the color of balls drawn with replacement from an urn with known size but with unknown composition. Correlation restrictions When the outcome space Yhas a product structure, and the agent’s prior imposes a qualitative restriction on how the components are correlated (e.g., that they are positively correlated), the Diaconis and Freedman (1990)resultdoes not apply, while ours does. This naturally arises in economic problems as, e.g., in Spiegler (2020), where the agent neglects the mediating role of expectations in the Phillips curve and is mistakenly convinced that money supply and output are positively correlated. Similarly, our model can be used to study situations where the agent mistakenly believes outcomes are independent as in, e.g., Enke and Zimmermann (2019). Moreover, whenever the agent observes the outcome of a game and believes (correctly or not) that the players have coordinated on a correlated equilibrium, the correlation structure is naturally restricted, as the possible joint distributions must satisfy the obedience constraint.5 5Formally, let Yibe the set of actions available to player i,lettheoutcomespacebeY=× n i=1Yi, and let (ui)n i=1be the payoff functions. If the agent is certain that the outcome corresponds to a correlated Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1589 Markov models Another context where support restrictions arise naturally is in the study of Markov models. For example, if an analyst assumes that beliefs follow Bayes rule or that a stock price process is a martingale (Bachelier (1900), Fama (1965)), the techniques of Diaconis and Freedman (1990) cannot be applied, as they would require full support over the set of transition matrices. However, it is easy to extend our analysis to analyze belief concentration in Markov models, as we do in Section 5.Andour Markov model can also be used to study mistaken beliefs about the correlation between signals and outcomes, as in Esponda (2008). Extending the applications of Section 1.1 Our results can be used to extend some past applications of the pathwise concentration results. They permit an extension of Al- Najjar and Shmaya’s (2019) representation theorem to beliefs about the consumption process that do not have full support, such as its illustrative example (which is not covered by the paper’s result), and also to Markovian consumption processes. For the experimentation in games considered by Gonçalves (2020), our extension allows, e.g., beliefs concentrated on the pure strategies for the opponents, or certainty that the opponent does not play a dominated strategy. 2. Setup We study Bayesian beliefs induced by a sequence of subjectively i.i.d. data. Let Ybe a finite set of possible outcomes, and let P=(Y)be the set of probability measures over Yendowed with ·, the total variation distance of probability measures.6 Let μ0∈(P)=((Y)) denote a prior belief over distributions of outcomes, and =suppμ0denote its support.7Adata set yt=(y1,y2,,yt)∈Ytis a vector of outcomes. For every data set yt,weletμtbe the posterior belief, which is required to satisfy Bayes rule whenever the denominator is different from 0: μt(C)=C t  τ=1 p(yτ)dμ0(p) P t  τ=1 p(yτ)dμ0(p) .(Bayesrule) The empirical distribution ft∈Pis ft(z)=1 t t  τ=1 I{z}(yτ). Our main result is that along any path of realized outcomes, the probability the posterior belief assigns to the outcome distributions that do not best approximate the empirical distribution converges to zero at a uniform and exponential rate. To state this equilibrium, then every p∈must satisfy y−i∈Y−iui(yi,y−i)p(y−i|yi)≥y−i∈Y−iui(y i,y−i)p(y−i|yi)for all i∈I,yi,y i∈Yiwith p(yi)>0. 6Formally, for every p,q∈P,p−q=supY⊆Y|p(Y)−q(Y)|. 7For every X⊆Rk,welet(X)denote the set of Borel probability distributions on X. 1590 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) conclusion formally, we adopt the convention that 0/0=0and0log 0=0, and define H:P×P→¯ Rto be the (possibly infinite) Kullback–Leibler divergence of qwith respect to p: H(q,p)= z∈Y q(z)logq(z) p(z). Let M:P⇒Pbe the correspondence that maps a distribution qto the set of minimizers of the Kullback–Leibler divergence over the support of the prior: M(q)=argmin p∈ H(q,p). The log-likelihood assigned to the data set ytunder outcome distribution pis logt  τ=1 p(yτ)= z∈Y tft(z)log p(z)=−tH(ft,p)+t z∈Y ft(z)logft(z).(1) Minimizing the Kullback–Leibler divergence relative to the empirical distribution is hence the same as maximizing the log-likelihood assigned to the data set, so the KL minimizers M(ft)at time tcorrespond to the outcome distributions that maximize the likelihood of yt. Throughout, Bε(D)denotes the ball of radius εaround a set D⊆P in total variation distance, and denote by Mε:P⇒Pthe correspondence that maps a distribution qto the distributions that come within εof the minimum KL divergence: Mε(q)=p∈:Hq,p≤min p∈H(q,p)+ε. 3. The rate of convergence of Bayesian beliefs To show that Bayesian beliefs concentrate around the empirical distribution at a uniform rate, Diaconis and Freedman (1990) used the following condition. Definition 1(φpositivity). The prior μ0is φpositive if for φ:R++ →R++, μ0(Bε(p)) ≥φ(ε)for every p∈Pand ε>0. Since φpositivity requires the prior to assign strictly positive probability to every ε ball, it requires the prior to have full support. Theorem A(Diaconis and Freedman (1990)). For every φ:R++ →R++ and every ε∈ (0, 1)there are ˜ A(ε)∈R++ and g(ε)∈R++ such that μtBε(ft) 1−μtBε(ft)≥˜ A(ε)expg(ε)t, for all φpositive μ0,t∈N,andft∈(Y). Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1591 Theorem Ashows that for φpositive priors, the probability that Bayesian beliefs assign to distributions that are more that εaway from the empirical distribution vanishes exponentially quickly, so it quantifies the speed at which a Bayesian with full support prior becomes more certain when observing i.i.d. data. The strength of this theorem is that it holds not only in probability, but for every realization of outcomes. Clearly, φpositivity plays a crucial role in Theorem A,asifthepriorisnotφpositive the empirical distribution need not be in its support, so beliefs cannot concentrate around it. However, requiring the prior to satisfy φpositivity rules out several practically relevant cases. For example, φpositivity cannot be satisfied if the prior has finite support, reduces the dimensionality of the problem, or is supported only on unimodal distributions. Moreover, models of misspecified learning suppose that the true data generating process is not in the support of the prior, which rules out φpositivity. We extend Theorem Ato cases where φpositivity fails. Loosely, we require that either the prior gives all neighborhoods of a distribution sufficient weight or the prior gives zero weight to a small neighborhood of the distribution. Definition 2(φpositivity on ). The prior μ0is φpositive on if for φ:R++ →R++, μ0(Bε(p)) ≥φ(ε)for every p∈and ε>0. Note that φpositivity on reduces to φpositivity when =P, i.e., the prior has full support. In Diaconis and Freedman (1990),Bayesruleiswell-definedeverywhere,but this is not true when the prior does not have full support. We define (Y)to be the (compact) set of empirical frequencies for which Bayesian updating is well-defined for a prior with support .8Theorem 1below establishes that if beliefs are φpositive on , for every ε∈(0, 1), the posterior concentrates on Mε(ft). Theorem 1. For every φ:R++ →R++,α∈(0, 1),andε∈(0, 1),thereisA(ε)>0such that μtMε(ft) 1−μtMε(ft)≥A(ε)exp(αεt) for all t∈N,ft∈(Y),andμ0that is φpositive on . Moreover, if q:=infq∈minz∈suppqq(z)>0, then we can set A(ε)=φminq/2, (1−α)εq/2. Theorem 1only requires φpositivity on ,incontrasttoTheoremA, which assumes φpositivity on the whole space of distributions. When the prior is not φpositive on all of (Y), beliefs need not concentrate around the empirical frequency, because this frequency might not be in the prior’s support. This is why Theorem 1bounds the probability assigned to the distributions in Mε(ft), which are the εminimizers of the KL divergence, while Theorem Abounds the probability assigned to Bε(ft),theεball around 8That is, (Y)={q∈(Y):∃p∈,suppq⊆supp p}. 1592 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) the empirical distribution. Moreover, as Example 1below shows, the theorem does not apply to the εball Bε(M(ft)) around the exact minimizers M(ft), because when is not convex points far from the minimizers can attain almost the same divergence. The theorem implies that the probability assigned to all distributions that do not ε best explain the empirical frequency ftvanishes at the exponential rate αε: μt\Mε(ft)≤1 A(ε)exp(−αεt). Notice that the multiplicative constant A(ε)depends on the prior μ0only through  and the function φ. The second part of the statement guarantees that in the widelystudied case of finite support priors, there is an explicit formula to compute the rate of convergence as a function of and φ, with the intuitive comparative statics that the rate of convergence improves when φis higher and when the support is smaller. The next example shows why Theorem 1doesnotapplytotheεball Bε(M(ft)) around the exact minimizers M(ft). Example 1. Let Y={0, 1}, identify each p∈(Y)with the probability of y=1, and let μ0({1/4})=μ0({3/4})=1/2. Consider the sequence of outcomes (yt)∞ t=1where yt=1if tis odd and yt=0iftis even. In the even periods 2t, the data is uninformative about the state, and both 1/4and3/4 are minimizers. At every odd period 2t+1, for every p∈, H(f2t+1,p)=Kt−t 2t+1log(1−p)−t+1 2t+1log(p), where the term Ktdoes not depend on p. Thus in the odd periods M(f2t+1)={3/4},so for ε<1/2, Bε(M(f2t+1)) ={3/4}. However, μ2t+1BεM(f2t+1) 1−μ2t+1BεM(f2t+1)=μ2t+1{3/4} μ2t+1{1/4}=μ0{3/4}(1/4)t(3/4)t+1 μ0{1/4}(1/4)t+1(3/4)t=3 so beliefs do not concentrate on the neighborhood of the KL minimizer.9The concentration result fails because the difference between the KL-divergences is (log(3/4)− log(1/4))/(2t+1), which converges to 0. Thus even a very large data set provides only weak evidence in favor of p=3/4. ♦ 3.1 Proof sketch of Theorem 1 The proofs of all our results are in the Appendix. The proof of Theorem 1has three steps. Step 1 proves a local Lipschitz property of the KL divergence, step 2 gives an explicit rate of concentration for each realized empirical frequency, while step 3 concludes by turning this explicit local rate of convergence into an exponential (but with possibly implicit constant) global rate of convergence. 9In this example, is not connected. Example 4in the Appendix shows that the same problem can arise when it is. The example there adds a third outcome to this one, and specifies a that connects 1/4 and 3/4 via distributions that are not relevant under the specified outcome sequence. Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1599 6. Infinite outcome spaces This section discusses pathwise concentration in the case of an infinite outcome space. Noncompact prior support The most common approach when dealing with infinitely many outcomes is to use a parametric description of the data generating process, as was done in Section 4for a finite outcome space. However, if the set of parameters  indexing the data generating process is not compact, it will be typically not be possible to satisfy φ-positivity. For example, if the prior is supported over the normal distributions with some fixed variance σ2and unknown mean θ∈R, and all values of the mean are considered possible, the prior cannot be φ-positive, as for any ε>0, μ(Bε(θ)) ≥φ(ε)> 0forallθ∈would imply that μ(R)=∞. Similarly, pathwise concentration fails: Pathwise concentration requires that a finite number of observations can outweigh the prior, but no fixed finite number of observations can outweigh the prior if the prior probability of the ε-minimizing set can be arbitrarily low. For this reason, pathwise concentration can be obtained only for priors with compact support. Divergence vs. likelihood The empirical distribution is always discrete, but the KL divergence from a discrete distribution to a nonatomic one is infinite. To handle this, we shift from concentration around the KL-minimizer to concentration around the maximizer of the empirical log-likelihood. Since the likelihood is the negative of the divergence plus a constant, these coincide in the case of a finite Y, but only the empirical log-likelihood maximizers are always well-defined for a continuum of outcomes. With this change, we now extend Lemma 1to the case of infinitely many outcomes. Let Ybe a metric space, and suppose that there exists a σ-finite measure ξon Ysuch that for every θ∈, the probability measure associated with θis absolutely continuous with respect to ξwith Radon–Nykodim derivative pθ∈RY.Let={pθ:θ∈}and Psbe the set of simple (finite support) distributions over Y. Balls in are taken with respect to the supnorm. For every θ∈and q∈Ps,let L(q||pθ)= y∈Y q(y)logpθ(y) be the empirical log-likelihood of the empirical distribution qunder pθ∈. Also, let Mε(q)=pθ∈:L(q||pθ)+ε≥max θ∈L(q||pθ) be the set of εmaximizers of the empirical log-likelihood. Recall that (Y):=f∈Ps:∃θ∈,∀y∈suppf,pθ(y)>0 is the set of empirical frequencies for which Bayesian updating is well-defined. Lemma 2. For every ε,ε,κ∈R+,t∈N,ft∈(Y),¯ q∈Mε(ft),withε+κ≤ε, μtMε(ft) 1−μtMε(ft)≥μ0Bκ/R(ft,κ,¯ q)(¯ q)expε−κ−εt, 1600 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) where R(ft,κ,¯ q)=maxmax q∈∩Bκ(¯ q) z∈suppft 1 q(z),1 . As in the finite case, m-convexity is useful for guaranteeing belief concentration, but it is harder to satisfy m-convexity when Yis infinite. For this reason, we generalize the m-convexity to only hold on a given set of empirical frequencies. Definition 4. Let m>0. Lis uniformly strongly m-concave on F⊆Psif for all f∈F, ∇θL(f||pθ)−∇θL(f||pθ)Tθ−θ≤−m θ−θ  2 2 for all θ,θ∈. Let θ∗(ft)denote the empirical likelihood maximizer. In the case of a singledimensional parameter, uniform strong m-concavity on Fis still enough to prove that posteriors concentrate on a neighborhood of the unique maximizer at a rate that is uniform over paths with empirical frequency in F.Example5in the Appendix shows that convergence need not be uniform over frequencies that do not make the likelihood function m-uniformly concave. The main difficulty is that we have little information about which εmovements from θ∗(ft)least decrease the empirical likelihood. The proof uses the fact that when the parameters are unidimensional there are at most two candidates for a best fitting parameter outside Bε(θ∗(ft)) (either θ∗(ft)−εor θ∗(ft)+ε)toovercomethisdifficulty. Proposition 2. If Lis uniformly strongly m-concave on F,then for every φ:R++ → R++, μtBεθ∗(ft) 1−μtBεθ∗(ft)≥φε 4exptε2m 2 for all μ0that are φpositive on ⊆R,ε∈(0, 1),t∈N,andft∈(Y)∩F. As an immediate corollary, there is pathwise belief concentration for an agent who believes the data are generated by a normal distribution with known variance and unknown mean. Corollary 1. Let σ2∈R++ and pθ(y)=1 2πσ2exp−(y−θ)2 2σ2. For every φ:R++ →R++, and every α∈(0, 1), μtBεθ∗(ft) 1−μtBεθ∗(ft)≥φε 4exptε2 2σ for all μ0that are φpositive on ⊆R,ε∈(0, 1),t∈N,andft∈(Y). Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1601 No concentration on the εmaximizers Without additional assumptions, Proposition 2cannot be strengthened to obtain pathwise concentration on the εmaximizers, as Example 6in Appendix A.11 shows. Intuitively, with infinitely many signals, the informativeness of a single signal may be unbounded, so that the set of εmaximizers after a single signal can be arbitrarily small. If the prior probability assigned to these sets vanishes at a sufficiently high exponential rate, their good match to the data does not guarantee that the posterior concentrates on them. More precisely, for some priors the conclusion of Theorem 1does not even hold for t=1. That is, it is not possible to have a concentration that holds uniformly over all the same-length realizations, let alone concentration rate that is uniform over same-length realizations and grows exponentially in the sample size.17 We leave for future work the challenge of determining just what sorts of restrictions on the prior would allow a uniform concentration result. Anticipated utility Much of the macroeconomics literature assumes that the data agents observe can take infinitely many different values. This is true in particular for the literature on “anticipated utility” (Kreps (1998)), which assumes that agents in the economy choose actions that maximize their payoff under a point estimate that maximizes the likelihood of their sample, ignoring uncertainty about the state. This is a simpler problem than the maximization of expected utility, and the reduction in complexity and dimension makes anticipated utility models more tractable and easier to analyze. However, it has not been clear how much error the approximation induces. For example, Cogley and Sargent (2008)wrote: Macroeconomists might justify anticipated-utility models as an approximation to a correctly formulated Bayesian decision problem... (the models) would be more compelling if one could also show that anticipated-utility decisions well approximate Bayesian decisions. As far as we know, no one has assessed the quality of the approximation... There is also a small literature that addresses this question using numerical simulations (Cogley, Colacito, and Sargent (2007), Cogley and Sargent (2008), Cogley et al. (2008)). Our result on Bayesian updating can be used to derive analytical results that complement these numerical studies. In particular, they imply that the long-run behavior under anticipated utility models converges to that of an expected utility maximizer. This provides a formal justification for the use of anticipated as an approximation of expected utility models in studies of long-run behavior. To develop this link, suppose that in each period t∈{1, 2, 3, }the agent chooses an action from A. We assume that Ais a convex set, endowed with a metric dthat makes it a compact set. The action does not affect the outcome distribution but influences the agent’s utility function u:A×Y→R, which is strictly concave in a.LetA∗(ν)denote the (unique) optimal action given belief ν, i.e., A∗(ν)=argmax a∈A Epθu(a,y)dν(θ), 17Fudenberg, He, and Imhof (2017) and Fudenberg, Lanzani, and Strack (2021a) point out other odd implications of priors that decay exponentially quickly. 1602 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) and suppose A∗is uniformly continuous when ()is endowed with the topology of weak convergence of measures.18 Let A∗(M(ft)) denote the action that is optimal for a point belief in the likelihood maximizer M(ft). Proposition 3. Suppose that ⊆Ris convex and that μ0is φpositive on .IfL is uniformly strongly m-concave on F,then for all ε>0there is a T∈Nsuch that d(A∗(μt),A∗(M(ft))) ≤εfor every t>Tand every ft∈F. 7. Conclusion We have shown that for every realization of the data, Bayesian beliefs concentrate exponentially quickly on the models that best explain the empirical frequency of outcomes. One implication of this concentration result is that optimal actions can be determined directly from the empirical frequency without computing beliefs. More precisely, once the sample is sufficiently large, neither the exact sample size nor calendar time is needed to compute the optimal action; the empirical frequency is sufficient. As the dynamics and distribution of the empirical frequency are well understood, this insight can greatly simplify the analysis of the long-run behavior of Bayesian agents. In addition to the applications developed in this paper, Theorem 1may allow generalizations of other results about misspecified Bayesian agents who learn from endogenous data. One recurrent theme in this literature is the possibility that when actions are endogenous, misspecified beliefs can lead to cycles in setting that would not occur with correctly specified beliefs, because repeated play of an action generates evidence in favor of another action.19 In such situations, our concentration result may be used to bound the number of periods spent in each phase of the cycles. This would complement Esponda, Pouzo, and Yamamoto (2021), which characterized the asymptotic frequencies of these cycles when the space of beliefs can be partitioned into a finite number of attracting sets and the support of the prior is one-dimensional. Our uniform speed of convergence result might be useful in extending this to more general settings. In addition, as we provide a concentration bound for every finite time, our result can be used to characterize behavior in the “medium-run” before the asymptotic results apply. Mazumdar, Pacchiano, Ma, Bartlett, and Jordan (2020) prove that with high probability, the posteriors of a correctly-specified Bayesian concentrate around the true parameter at rate √n, and use this result to study the long-run properties of Thompson sampling. The paper allows for infinitely many outcomes, but imposes additional strong conditions such as log-concavity of the true data generating process, and a prior density that is bounded away from 0. Our results enable extensions to Thompson sampling with less restricted priors in the finite outcome case. In settings where multiple agents choose their actions based on the same observables, our concentration results can be used to quantify the minimal extent of the differences in their prior beliefs needed to rationalize different choices. For example, Olea, 18A sufficient condition for this is that is compact and A∗is continuous. 19See, e.g., Nyarko (1991), Fudenberg, Romanyuk, and Strack (2017), Levy, Razin, and Young (2020), and Lanzani (2022). Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1603 Luis, Ortoleva, Pai, and Prat (2021) showed that when observing signals of an object’s value, misspecified agents with lower-dimensional models have a higher willingness to pay after the first few observations, while correctly specified agents have a higher willingness to pay in the long-run; our result on the speed of convergence may help to better identify the switching time. The learning in games literature has assumed correctly specified beliefs to appeal to Diaconis and Freedman (1990). Our generalization will facilitate the extension of the results from this literature to cases where the agents in the learning model have misspecified beliefs about the extensive form of the game. It will also enable extensions to incorrect beliefs about a complex network structure in Bowen, Dmitriev, and Galperti (forthcoming), and to overconfident agents as in Heidhues, K˝ oszegi, and Strack (2018). Appendix A.1 Properties of the KL divergence Lemma 3. For all p,˜ p,q∈P, H(q,p)−H(q,˜ p)≤2max z∈Ymaxq(z) p(z),q(z) ˜ p(z)p−˜ p. Proof.Let R:=max z∈Ymaxq(z) p(z),q(z) ˜ p(z), ˆ Y={z:p(z)>˜ p(z)}, and suppose without loss of generality that p(supp q)≥˜ p(suppq). Then H(q,p)−H(q,˜ p) = z∈suppqlogp(z) q(z)−log˜ p(z) q(z)q(z) = z∈suppqp(z)/q(z) ˜ p(z)/q(z) 1 rdrq(z) ≤ z∈suppq maxq(z) ˜ p(z),q(z) p(z) p(z) q(z)−˜ p(z) q(z) q(z) ≤R z∈suppq p(z) q(z)−˜ p(z) q(z) q(z) =R z∈suppq2Iˆ Y(z)−1p(z) q(z)−˜ p(z) q(z)q(z) =R z∈suppq2Iˆ Y(z)−1p(z)−R z∈suppq2Iˆ Y(z)−1˜ p(z) 1604 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) =2R z∈suppq Iˆ Y(z)p(z)− z∈suppq Iˆ Y(z)˜ p(z)+R˜ p(suppq)−p(suppq). As p(suppq)≥˜ p(suppq), the above term is bounded by ≤2R z∈suppq Iˆ Y(z)p(z)− z∈suppq Iˆ Y(z)˜ p(z)≤2R z∈ˆ Y p(z)− z∈ˆ Y˜ p(z) =2Rp−˜ p, where the last inequality follows from the definition of ˆ Yand the last equality by the definition of the total variation distance. Recall that a probability distribution p∈(Y)is absolutely continuous with respect to q∈(Y), denoted as pq,ifsuppp⊆supp q. Lemma 4. Let ε∈R+.ThenMε(·)={q∈:H(·,q)≤minq∈H(·,q)+ε}is nonemptyvalued and compact-valued. Moreover, for all p∈P,Mε(·)is upper hemicontinuous on Binfz∈supppp(z)/2(p)∩{q:qp}. Proof.IfH(p,q)=∞for all q∈,Mε(p)=is nonempty and compact. If there is ˆ qsuch that H(p,ˆ q)=K<∞,theset={q∈:H(p,q)≤K+ε}is compact by the continuity of H(p,·),andMε(p)⊆. So, the continuous and real-valued restriction of H(p,·)to has compact lower contour sets, and it attains a minimum. Thus Mε(p)⊇ M(p)=∅is nonempty and compact. For the second part of the statement, observe that if p/∈(Y),thenMε(p)=and, therefore, Mε(p)is trivially upper hemicontinuous at psince by definition Mε(p)⊆ for all p∈(Y). If instead p∈(Y),thereexistˆ q∈and K∈R+with H(p,ˆ q)=K. Moreover, the finiteness of H(p,ˆ q)=Kimplies that pˆ q.So,thereexistsK>0such that Hp,ˆ q≤K∀p∈Binfz∈supppp(z)/2(p)∩{q:qp}. We use this equation to show that there exists Csuch that r∈Mε(p),p∈ Binfz∈supppp(z)/2(p)∩{q:qp}implies r(y)≥Cfor all y∈supp p. Suppose by contradiction that this is not the case. Then there exist a convergent sequence (rn,pn)n∈N∈ (×(Binfz∈supppp(z)/2(p)∩{q:qp}))Nand an ˆ y∈supppwith rn∈Mε(pn)for all n∈N and limn→∞rn(ˆ y)=0. But we have H(pn,rn)≥ y∈Y pn(y)logpn(y)−pn(ˆ y)logrn(ˆ y)≥ y∈Y pn(y)logpn(y)−p 2logrn(ˆ y) and the right-hand side is diverging to ∞. So, eventually H(pn,rn)≥ε+K≥ε+ H(pn,ˆ q), a contradiction with rn∈Mε(pn). Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1605 This shows that for all p∈Binfz∈supppp(z)/2(p)∩{q:qp}, Mεp=r∈:Hp,r≤min ¯ r∈Hp,¯ r+ε =r∈:Hp,r≤min ¯ r∈Hp,¯ r+ε,r(y)≥C,∀y∈suppp.(4) Also, observe that the function His continuous on the set Binfz∈supppp(z)/2(p)∩{q:qp}×r∈:r(y)≥C,∀y∈supp p. Therefore, if we define G:(Binfz∈supp pp(z)/2(p)∩{q:qp})→Ras Gp=min {r∈:r(y)≥C,∀y∈suppp}Hp,r Gis continuous by the maximum theorem. Moreover, by equation (4)G(p)= minr∈H(p,r)for all Binfz∈supp pp(z)/2(p)∩{q:qp}, showing that minr∈H(·,r)is a continuous function when restricted on Binfz∈supp pp(z)/2(p)∩{q:qp}.Toconclude,we show that Mε(·)is upper hemicontinuous on Binfz∈supp pp(z)/2(p)∩{q:qp}by showing that it has a closed graph. Indeed, let (pn,rn)∈Binfz∈supp pp(z)/2(p)∩{q:qp}× be such that rn∈Mε(pn)for all n∈Nand limn→∞(rn,pn)=(ˆ r,ˆ p).Byequation(4), for all n∈N,wehavern∈{r∈:r(y)≥C,∀y∈suppp}and since this last set is close ˆ r∈{r∈:r(y)≥C,∀y∈suppp}. By the continuity of Hon (Binfz∈supppp(z)/2(p)∩{q:q p})×{r∈:r(y)≥C,∀y∈suppp},andofGon Binfz∈supp pp(z)/2(p)∩{q:qp},wehave min r∈H(ˆ p,r)−H(ˆ p,ˆ r)=G(ˆ p)−H(ˆ p,ˆ r)=lim n→∞G(pn)−H(pn,rn)≤ε proving that ˆ r∈Mε(ˆ p). Lemma 5 (Pinsker’s inequality). For every p,q∈(Y), p−q≤H(p,q) 2. A.2 Proof of Lemma 1 Lemma 1. If μ0is φpositive on , then for every ε,ε,κ∈R+,t∈N,ft∈(Y),and ¯ q∈Mε(ft)with ε+κ≤εand R(ft,κ,¯ q)<∞, we have μtMε(ft) 1−μtMε(ft)≥φκ/2R(ft,κ,¯ q)expε−κ−εt.(5) Proof. The proof uses the following bound on how much the Kullback–Leibler divergence can increase when moving from an εminimizer to a nearby distribution. Claim 1. For every p∈,f∈(Y),ε,κ∈R+,and¯ q∈Mε(f), p∈Bκ/2R(f,κ,¯ q)(¯ q)=⇒ Hf,p≤min p∈H(f,p)+ε+κ. 1606 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) Proof. For every two distributions f,q∈(Y), there is at least one outcome that is weakly more likely under fthan under q,so R(f,κ,¯ q)=max q∈∩Bκ(¯ q)max z∈Y f(z) q(z) is bounded below by 1. Thus p∈Bκ/2R(f,κ,¯ q)(¯ q)implies p∈Bκ(¯ q). Therefore, both p,¯ q are in ∩Bκ(¯ q), so from the definition of R, max z∈Ymaxf(z) p(z),f(z) ¯ q(z)≤R(f,κ,¯ q). Moreover, p∈Bκ/2R(f,κ,¯ q)(¯ q)implies that ¯ q∈Bκ/2R(f,κ,¯ q)(p)∩Mε(f), so by Lemma 3, Hf,p−H(f,¯ q)≤κ R(f,κ,¯ q)max z∈Ymaxf(z) p(z),f(z) ¯ q(z)≤κ R(f,κ,¯ q)R(f,κ,¯ q)=κ, and hence H(f,p)≤H(f,¯ q)+κ≤minp∈H(f,p)+ε+κ. We use Claim 1to provide a lower bound on the probability of the εminimizers given the empirical frequency ft. Observe that μtMε(ft) 1−μtMε(ft)=Mε(ft) exp−H(p,ft)tdμ0(dp) \Mε(ft) exp−H(p,ft)tdμ0(dp) ≥Mκ+ε(ft) exp−H(p,ft)tdμ0(dp) \Mε(ft) exp−H(p,ft)tdμ0(dp) ≥ exp−min p∈H(ft,p)+κ+εt exp−min p∈H(ft,p)+εt μ0Mκ+ε(ft) μ0\Mε(ft) =expε−κ−εtμ0Mκ+ε(ft) μ0\Mε(ft) ≥expε−κ−εtμ0Bκ/2R(ft,κ,¯ q)(¯ q) ≥expε−κ−εtφκ/2R(ft,κ,¯ q). The first equality follows from equation (1). The first inequality follows from ε+κ≤ ε, the second inequality from pointwise bounding the integrands and the definition of Mε, the third inequality from Claim 1, and the fourth because μ0is φpositive on . Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1607 A.3 Proof of Theorem 1 Theorem 1. For every φ:R++ →R++,α∈(0, 1),andε∈(0, 1)there is A(ε)such that μtMε(ft) 1−μtMε(ft)≥A(ε)exp(αεt) for all t∈N,ft∈(Y),andμ0that is φpositive on .Moreover,ifq:= infq∈minz∈suppqq(z)>0, then we can set A(ε)=φminq/2, (1−α)εq/2. We prove the theorem for φnondecreasing. This is without loss of generality, as if μ0 is φpositive on ,itisalso ˆ φpositive on where ˆ φ(ε)=sup ε≤ε φ(ε)∀ε∈R++. Clearly, ˆ φis nondecreasing and only depends on φ. Moreover, it also pointwise weakly dominates φ,sowhenq>0 , if we prove the statement for ˆ φwe have μt(Mε(ft)) 1−μt(Mε(ft)) ≥ˆ φmin{q/2, (1−α)ε}q/2exp(αεt) ≥φmin{q/2, (1−α)ε}q/2exp(αεt) proving the statement for φas well. We first show that if q:=infq∈minz∈suppqq(z)>0, Lemma 1yields the desired uniform rate of convergence. If (1−α)ε<q ,thenforallft∈(Y)and ¯ q∈M(ft),if p∈∩B(1−α)ε(¯ q),thensupp p=supp ¯ q⊆suppft,so 1 Rft,(1−α)ε,¯ q=max p∈∩B(1−α)ε(¯ q)max z∈Y ft(z) p(z)−1 ≥1 (1/q)=q. If instead (1−α)ε≥q,it is enough to observe that 1/R(ft,q/2, ¯ q)≥qfor all ft∈(Y), ¯ q∈M(ft). Now we move to the proof for the general case where some outcomes might have an arbitrarily low probability under data-generating processes in the support of the prior, i.e., qmight equal 0. Recall that (Y)={q∈(Y):∃p∈,suppq⊆suppp}is the set of distributions for which Bayes rule is well-defined, and that Theorem 1applies only to empirical distributions f∈(Y). To provide an upper bound on R, we show that the likelihood ratio f/q (which determines the value of R) can be uniformly bounded for all probability distributions qthat are sufficiently close to an (1−α)ε/2 minimizer of the Kullback–Leibler divergence. Intuitively, as f∈(Y)some distribution assigns nonvanishing probability to every outcome which has positive probability under f,and thus a distribution that assigns vanishing probability to some of these outcomes leads to an excessively low log-likelihood (and thus a high Kullback–Leibler divergence). 1608 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) Claim 2. For every ε∈(0, 1),thereexist¯κα(ε)∈(0, (1−α)ε/2]and c≥1such that for all κ≤¯καand f∈(Y),thereis¯ q∈M(1−α)ε/2(f)such that max y∈Y,q∈Bκ(¯ q) f(y) q(y)≤c. Proof. If not, then since (Y)and are compact, there is a sequence (f n,qn)∈ (Y)×with qn∈M(f n)that converges to (ˆ f,ˆ q),andsuchthat inf ¯ q∈M(1−α)ε 2 (f n)max y∈Y,q∈B1/n(¯ q) fn(y) q(y)≥n.(6) Since Yis finite, so is the set of possible supports, so there is a subsequence (fn)n∈Nsuch that each element of the sequence has common support, with (fn(y))nweakly decreasing for all y∈Y\supp ˆ f. Moreover, since  z∈Y fn(z)logfn(z)∈log1 |Y|,0 ∀n∈N the subsequence can also be taken such that z∈Yfn(z)log fn(z)converges. Since all the fnare in (Y)and have common support, qn∈M(fn),andlog(fn(z)) ≤ 0forallz∈Y,wehave H(fn,qn)≤H(fn,q1)= z∈suppf1 fn(z)logfn(z)− z∈suppf1 fn(z)logq1(z) ≤−  z∈suppf1 fn(z)logq1(z)≤− min z∈suppf1 logq1(z)<∞, so (H(fn,qn))n∈Nis bounded. Moreover, if there exist z∗∈Yand l∈R++ with limn→∞qn(z∗)=0andlimn→∞ fn(z∗)=l,then limsup n→∞ H(fn,qn)=limsup n→∞  y∈Y fn(y)logfn(y)−logqn(y) = y∈Y ˆ f(y)log ˆ f(y)+limsup n→∞ − y∈Y fn(y)logqn(y) ≥ y∈Y ˆ f(y)log ˆ f(y)+limsup n→∞ −fnz∗logqnz∗ = y∈Y ˆ f(y)log ˆ f(y)−lloglim n→∞qnz∗=∞, which contradicts (H(fn,qn))n∈Nbeing bounded. So, for all z∈Y, lim n→∞qn(z)=0=⇒ lim n→∞fn(z)=0=⇒ z/∈supp ˆ f. Thus G:=infn∈N,y∈supp ˆ flogqn(y)>−∞. Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1615 A.8 Proof of Lemma 2 Lemma 2. For every ε,ε,κ∈R+,t∈N,ft∈(Y),¯ q∈Mε(ft),withε+κ≤ε, μtMε(ft) 1−μtMε(ft)≥μ0Bκ/R(ft,κ,¯ q)(¯ q)expε−κ−εt, where R(ft,κ,¯ q)=maxmax q∈∩Bκ(¯ q) z∈suppft 1 q(z),1 . Proof. The proof follows from the following claims. Claim 5. For all p,ˆ p∈,q∈Ps, L(q||p)−L(q|| ˆ p)≤max z∈suppqmax1 p(z),1 ˆ p(z)p−ˆ p∞. Proof. L(q||p)−L(q|| ˆ p) = z∈suppqlogp(z)−logˆ p(z)q(z) = z∈suppqp(z) ˆ p(z) 1 rdrq(z)≤ z∈suppq max1 ˆ p(z),1 p(z)p(z)−ˆ p(z)q(z) ≤max z∈suppqmax1 p(z),1 ˆ p(z) z∈suppqp(z)−ˆ p(z)q(z) ≤max z∈suppqmax1 p(z),1 ˆ p(z)p−ˆ p∞, where the last equality follows from the definition of the supremum distance. Claim 6. For every p∈,f∈(Y),ε,κ∈R+,and¯ q∈Mε(f), p∈Bκ/R(f,κ,¯ q)(¯ q)=⇒ Lf||p+ε+κ≥max θ∈L(f||pθ). Proof. Since Ris bounded below by 1, p∈Bκ/R(f,κ,¯ q)(¯ q)implies p∈Bκ(¯ q).Therefore, both p,¯ qare in ∩Bκ(¯ q), so from the definition of R,maxz∈supp fmax{1 p(z),1 ¯ q(z)}≤ R(f,κ,¯ q).Moreover,p∈Bκ/R(f,κ,¯ q)(¯ q)implies that ¯ q∈Bκ/R(f,κ,¯ q)(p)∩Mε(f),soby Claim 5, L(f||¯ q)−Lf||p≤κ R(f,κ,¯ q)max z∈suppfmax1 p(z),1 ¯ q(z)≤κ R(f,κ,¯ q)R(f,κ,¯ q)=κ, 1616 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) and hence Lf||p≥L(f||¯ q)−κ≥max θ∈L(f||pθ)−ε−κ. To prove the lemma, observe that μtMε(ft) 1−μtMε(ft)≥Mκ+ε(ft) expL(ft||p)tdμ0(dp) \Mε(ft) expL(ft||p)tdμ0(dp) ≥ expmax θ∈L(ft||pθ)−κ−εt expmax θ∈L(ft||pθ)−εt μ0Mκ+ε(ft) μ0\Mε(ft) =expε−κ−εtμ0Mκ+ε(ft) μ0\Mε(ft) ≥expε−κ−εtμ0Bκ/R(ft,κ,¯ q)(¯ q). The first inequality follows from ε+κ≤ε, the second from pointwise bounding the integrands and the definition of Mε, and the third from Claim 6. A.9 Proof of Proposition 2 Proposition 2. If Lis uniformly strongly m-concave on F,then for every φ:R++ → R++, μtBεθ∗(ft) 1−μtBεθ∗(ft)≥φε 4exptε2m 2 for all μ0that are φpositive on ⊆R,ε∈(0, 1),t∈N,andft∈(Y)∩F. Proof. We claim first that for every θ∈and f∈(Y)∩F,∇θL(f||pθ∗(f))T(θ− θ∗(f)) ≤0. If not, 0<∇θL(f||pθ∗(f))Tθ−θ∗(f)=lim k→0 L(f||pθ∗(f)+k(θ−θ∗(f)))−L(f||pθ∗(f)) k. But this means that there is ˆ k∈(0, 1)such that L(f||pθ∗(f)+ˆ k(θ−θ∗(f)))−L(f||pθ∗(f))>0. As is convex, (1−ˆ k)θ∗(f)+ˆ kθ belongs to ,butthenθ∗(f)would not be a likelihood maximizer. Next, as Lis uniformly strongly m-concave, L(f||pθ)−L(f||pθ∗(f))≤∇θL(f||pθ∗(f))θ−θ∗(f)−m 2θ−θ∗(f)≤−m 2θ−θ∗(f). Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1617 The statement is trivially true if ⊆Bε(θ∗(f)). If not, since L(f||p(·))is concave and  is convex, at least one of θ+ε∈argmax θ:|θ−θ∗(f)|≥ε L(f||pθ), and θ−ε∈argmax θ:|θ−θ∗(f)|≥ε L(f||pθ) holds. We prove the result in the first case, the proof for the other case is symmetric. Let ¯ θ,θ∈be such that ¯ θ≥θ∗(f)+ε≥θ∗(f)+ε 2≥θ≥θ∗(f). We have L(f||p¯ θ)−L(f||pθ)≤Lθ(f||pθ)¯ θ−θ−m 2¯ θ−θ2 ≤Lθ(f||pθ∗(f))¯ θ−θ−m 2¯ θ−θ2≤−m 2¯ θ−θ2≤−ε2m 2, where the first inequality follows from the strong m-concavity of L,andthesecondby the concavity of Lin θ. Therefore, μtBεθ∗(f) 1−μtBεθ∗(f)≥ μtθ∗(f),θ∗(f)+ε 2 1−μtBεθ∗(f) ≥ μ0θ∗(f),θ∗(f)+ε 2 1−μ0Bεθ∗(f) exptε2m 2 ≥μ0θ∗(f),θ∗(f)+ε 2exptε2m 2 =φε 4exptε2m 2. A.10 Proof of Proposition 3 Proposition 3. Suppose that ⊆Ris convex and that μ0is φpositive on .IfL is uniformly strongly m-concave on F,then for all ε>0there is a T∈Nsuch that d(A∗(μt),A∗(M(ft))) ≤εfor every t>Tand every ft∈F. Proof. Since A∗is uniformly continuous, there exists ε∈R++ such that for all ν∈() and μt∈Bε(ν),d(A∗(μt),A∗(ν)) ≤ε. Since ⊆Rk, the topology of weak convergence 1618 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) on ()is metrized by the Lévy–Prokhorov metric. By the definition of this metric, ν− δp||LP ≤εwhenever νBε/2(p) 1−νBε/2(p)≥1−ε/2 ε/2. The statement follows from applying Proposition 2and choosing T≥ 2log1−ε/2 φε/8ε/2 ε2m/8 . A.11 Example 6 Example 6. (Unlimited Bernoulli trials) Suppose the outcome ycorresponds to the number of Bernoulli trials needed to get one success. If the agent believes the trials are i.i.d. with parameter pθ, their subjective distribution for outcome yis pθ(y)= θ(1−θ)y−1. Suppose that all success probabilities are considered possible, so that ={pθ:θ∈[0, 1]}.Then L(f||pθ)=∞  z=1 f(z)(z−1)log(1−θ)+log(θ)=log(1−θ)¯ z+log(θ)−log(1−θ). That Lis uniformly strongly 1-concave immediately follows from taking derivatives: ∂L(f||pθ) ∂θ =− ¯ z (1−θ)+1 θ+1 1−θ, ∂L(f||pθ)2 ∂2θ=− ¯ z (1−θ)2−1 θ2+1 (1−θ)2. The unique log-likelihood maximizing parameter is θ∗(f)=1/¯ z. Suppose that the prior belief μis such that lim θ→0 μ0, θ+2θ 1−θ expL(δ1/θp1/2)−L(δ1/θ||pθ)=0. (11) An example that satisfies the restriction is given by the CDF F(θ)=⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ exp−log(1−θ)1 θ−logθ 1−θ+1 θlog 1 22 for θ≤1/10, F1 10+10θ 9−1 91−F1 10 for θ>1/10. The restriction implies there cannot be Aand gsuch that μ1M1/2(δc) 1−μ1M1/2(δc)≥Aexp(g)∀c∈R++, (12) Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1619 so that the conclusion of Theorem 1does not hold for t=1. To see why equation (12) cannot be satisfied, observe that for every K>0, there is c>0suchthat μ1M1/2(δc) 1−μ1M1/2(δc)≤ μ10, 1 c+%2/c 1−1/c  μ11 4,1 2 ≤ μ0, 1 c+%2/c 1−1/c  μ1 4,1 2 expL(δc||p1/c ) expL(δc||p1 2) = μ0, 1 c+%2/c 1−1/c  μ1 4,1 2 expclog(1−1/c)−log1/c 1−1/c  expclog 1 2≤K, where the last inequality follows because equation (11) implies the left-hand side is arbitrarily close to 0 for sufficiently high c.♦ References Acemoglu, Daron, Victor Chernozhukov, and Muhamet Yildiz (2016), “Fragility of asymptotic agreement under Bayesian learning.” Theoretical Economics, 11, 187–225. [1587] Al-Najjar, Nabil and Eran Shmaya (2019), “Recursive utility and paramater uncertainty.” Journal of Economic Theory, 181, 274–288. [1587,1589] Arrow, Kenneth J. and Jerry R. Green (1973), “Notes on expectations equilibria in Bayesian settings.” Working Paper No. 33, Stanford University. [1586] Bachelier, Louis (1900), “Théorie de la spéculation.” Annales scientifiques de l’École normale supérieure., 17, 21–86. [1589] Berk, Robert H. (1966), “Limiting behavior of posterior distributions when the model is incorrect.” The Annals of Mathematical Statistics, 37, 51–58. [1585,1586,1594] Bohren, J. Aislinn and Daniel Hauser (2021), “Learning with model misspecification: Characterization and robustness.” Econometrica, 89, 3025–3077. [1586] Bowen, Renee, Danil Dmitriev, and Simone Galperti (2023), “Learning from shared news: When abundant information leads to belief polarization.” Quarterly Journal of Economics, 138, 955–1000. [1603] Clark, Daniel and Drew Fudenberg (2021), “Justified communication equilibrium.” American Economic Review, 111, 3004–3034. [1587] Clark, Daniel, Drew Fudenberg, and Kevin He (2022), “Observability, dominance, and induction in learning models.” Journal of Economic Theory, 206, 105569. [1587] Cogley, Timothy, Riccardo Colacito, Lars Peter Hansen, and Thomas J. Sargent (2008), “Robustness and US monetary policy experimentation.” Journal of Money, Credit and Banking, 40, 1599–1623. [1587,1601] 1620 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) Cogley, Timothy, Riccardo Colacito, and Thomas J. Sargent (2007), “Benefits from US monetary policy experimentation in the days of Samuelson and Solow and Lucas.” Journal of Money, Credit and Banking, 39, 67–99. [1587,1601] Cogley, Timothy and Thomas J. Sargent (2008), “Anticipated utility and rational expectations as approximations of Bayesian decision making.” International Economic Review, 49, 185–221. [1587,1601] Diaconis, Persi and David Freedman (1990), “On the uniform consistency of Bayes estimates for multinomial probabilities.” The Annals of Statistics, 18, 1317–1327. [1585,1586, 1587,1588,1589,1590,1591,1595,1596,1603] Dupuis, Paul and Richard S. Ellis (2011), A Weak Convergence Approach to the Theory of Large Deviations, volume 902. John Wiley & Sons. [1610] Enke, Benjamin and Florian Zimmermann (2019), “Correlation neglect in belief formation.” The Review of Economic Studies, 86, 313–332. [1588] Esponda, Ignacio (2008), “Information feedback in first price auctions.” The RAND Journal of Economics, 39, 491–508. [1589] Esponda, Ignacio and Demian Pouzo (2016), “Berk–Nash equilibrium: A framework for modeling agents with misspecified models.” Econometrica, 84, 1093–1130. [1586] Esponda, Ignacio and Demian Pouzo (2021), “Equilibrium in misspecified Markov decision processes.” Theoretical Economics, 16, 717–757. [1586] Esponda, Ignacio, Demian Pouzo, and Yuichi Yamamoto (2021), “Asymptotic behavior of Bayesian learners with misspecified models.” Journal of Economic Theory, 195, 105260. [1586,1602] Eusepi, Stefano and Bruce Preston (2018), “The science of monetary policy: An imperfect knowledge perspective.” Journal of Economic Literature, 56, 3–59. [1587] Fama, Eugene F. (1965), “The behavior of stock-market prices.” The journal of Business, 38, 34–105. [1589] Frankel, David M., Stephen Morris, and Ady Pauzner (2003), “Equilibrium selection in global games with strategic complementarities.” Journal of Economic Theory, 108, 1–44. Frick, Mira, Ryota Iijima, and Yuhta Ishii (2021), “A note on welfare comparisons for biased endogenous learning.” [1586] Frick, Mira, Ryota Iijima, and Yuhta Ishii (2022), “Welfare comparisons for biased learning.” [1586] Frick, Mira, Ryota Iijima, and Yuhta Ishii (2023), “Belief convergence under misspecified learning: A martingale approach.” Review of the Economic Studies, 90, 781–814. [1586] Fudenberg, Drew, Kevin He, and Lorens A. Imhof (2017), “Bayesian posteriors for arbitrarily rare events.” Proceedings of the National Academy of Sciences, 114, 4925–4929. [1601] Theoretical Economics 18 (2023) Pathwise concentration bounds for Bayesian beliefs 1621 Fudenberg, Drew and Kevin He (2018), “Learning and type compatibility in signaling games.” Econometrica, 86, 1215–1255. [1587] Fudenberg, Drew, Giacomo Lanzani, and Philipp Strack (2021a), “Limit points of endogenous misspecified learning.” Econometrica, 89, 1065–1098. [1586,1588,1601] Fudenberg, Drew, Giacomo Lanzani, and Philipp Strack (2021b), “Selective memory equilibrium.” [1586] Fudenberg, Drew, Giacomo Lanzani, and Philipp Strack (2022), “Pathwise concentration bounds for misspecified Bayesian beliefs.” Available at SSRN 3805083. [1588] Fudenberg, Drew and Giacomo Lanzani (2023), “Which misperceptions persist?” Theoretical Economics, 18, 1271–1315. [1586] Fudenberg, Drew and David K. Levine (1993), “Steady state learning and Nash equilibrium.” Econometrica, 547–573. [1587] Fudenberg, Drew and David K. Levine (2006), “Superstition and rational learning.” American Economic Review, 96, 630–651. [1587] Fudenberg, Drew, Gleb Romanyuk, and Philipp Strack (2017), “Active learning with a misspecified prior.” Theoretical Economics, 12, 1155–1189. [1586,1602] Gonçalves, Duarte (2020), “Sequential sampling and equilibrium.” https://bit.ly/ 3oPuSFL.[1587,1588,1589] He, Kevin (2022), “Mislearning from censored data: The gambler’s fallacy in optimalstopping problems.” Theoretical Economics, 17, 1269–1312. [1586] He, Kevin and Jonathan Libgober (2021), “Evolutionarily stable (mis)specifications: Theory and application.” arXiv:2012.15007.[1586] Heidhues, Paul, Botond K˝ oszegi, and Philipp Strack (2018), “Unrealistic expectations and misguided learning.” Econometrica, 86, 1159–1214. [1603] Heidhues, Paul, Botond K˝ oszegi, and Philipp Strack (2021), “Convergence in models of misspecified learning.” Theoretical Economics, 16, 73–99. [1586] Kleijn, Bas J. K. and Aad W. Van der Vaart (2012), “The Bernstein–von-Mises theorem under misspecification.” Electronic Journal of Statistics, 6, 354–381. [1587] Kreps, David M. (1998), “Anticipated utility and dynamic choice.” Econometric Society Monographs, 29, 242–274. [1587,1601] Lanzani, Giacomo (2022), “Dynamic concern for misspecification.” [1602] Levy, Gilat, Inés Moreno de Barreda, and Ronny Razin (2021), “Polarized extremes and the confused centre: Campaign targeting of voters with correlation neglect.” Quarterly Journal of Political Science, 16, 139–155. [1586] Levy, Gilat, Ronny Razin, and Alwyn Young (2020), “Misspecified politics and the recurrence of populism.” [1602] 1622 Fudenberg, Lanzani, and Strack Theoretical Economics 18 (2023) Mazumdar, Eric, Aldo Pacchiano, Yi-an Ma, Peter L. Bartlett, and Michael I. Jordan (2020), “On Thompson sampling with Langevin algorithms.” ArXiv preprint arXiv:2002.10002.[1602] Molavi, Pooya (2019), Macroeconomics With Learning and Misspecification: A General Theory and Applications.https://pooyamolavi.com/CREE.pdf.[1586] Montiel Olea, José Luis, Pietro Ortoleva, Mallesh Pai, and Andrea Prat (2021), “Competing models.” ArXiv preprint arXiv:1907.03809.[1603] Nyarko, Yaw (1991), “Learning in mis-specified models and the possibility of cycles.” Journal of Economic Theory, 55, 416–427. [1586,1602] Preston, Bruce (2005), “Learning about monetary policy rules when long-horizon expectations matter.” International Journal of Central Banking,1.[1587] Sanov, Ivan Nicolaevich (1961), “On the probability of large deviations of random variables.” Selected Translations in Mathematical Statistics and Probability, 1, 213–244. [1610] Schwartzstein, Joshua (2014), “Selective attention and learning.” Journal of the European Economic Association, 12, 1423–1452. [1587] Schwartzstein, Joshua and Adi Sunderam (2021), “Using models to persuade.” American Economic Review, 111, 276–323. [1587] Shen, Xiaotong and Larry Wasserman (2001), “Rates of convergence of posterior distributions.” The Annals of Statistics, 29, 687–714. [1587] Spiegler, Ran (2020), “Behavioral implications of causal misperceptions.” Annual Review of Economics, 12, 81–106. [1588] Co-editor Todd D. Sarver handled this manuscript. Manuscript received 18 February, 2022; final version accepted 15 November, 2022; available online 6 December, 2022.