Entropy regularization in mean-field games of optimal stopping
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Dianetti, Jodi; Dumitrescu, Roxana; Ferrari, Giorgio; Xu, Renyuan Working Paper Entropy regularization in mean-field games of optimal stopping Center for Mathematical Economics Working Papers, No. 755 Provided in Cooperation with: Center for Mathematical Economics (IMW), Bielefeld University Suggested Citation: Dianetti, Jodi; Dumitrescu, Roxana; Ferrari, Giorgio; Xu, Renyuan (2025) : Entropy regularization in mean-field games of optimal stopping, Center for Mathematical Economics Working Papers, No. 755, Bielefeld University, Center for Mathematical Economics (IMW), Bielefeld, https://nbn-resolving.de/urn:nbn:de:0070-pub-30076673 This Version is available at: https://hdl.handle.net/10419/333540 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
755 October 2025 Entropy Regularization in Mean-Field Games of Optimal Stopping Jodi Dianetti, Roxana Dumitrescu, Giorgio Ferrari and Renyuan Xu Center for Mathematical Economics (IMW) Bielefeld University Universit¨atsstraße 25 D-33615 Bielefeld ·Germany e-mail: [email protected] uni-bielefeld.de/zwe/imw/research/working-papers ISSN: 0931-6558 Unless otherwise noted, this work is licensed under a Creative Commons Attribution 4.0 International (CC BY) license. Further information: https://creativecommons.org/licenses/by/4.0/deed.en https://creativecommons.org/licenses/by/4.0/legalcode.en
ENTROPY REGULARIZATION IN MEAN-FIELD GAMES OF OPTIMAL STOPPING JODI DIANETTI, ROXANA DUMITRESCU, GIORGIO FERRARI, AND RENYUAN XU Abstract. We study mean-field games of optimal stopping (OS-MFGs) and introduce an entropyregularized framework to enable learning-based solution methods. By utilizing randomized stopping times, we reformulate the OS-MFG as a mean-field game of singular stochastic controls (SC-MFG) with entropy regularization. We establish the existence of equilibria and prove their stability as the entropy parameter vanishes. Fictitious play algorithms tailored for the regularized setting are introduced, and we show their convergence under both Lasry–Lions monotonicity and supermodular assumptions on the reward functional. Our work lays the theoretical foundation for model-free learning approaches to OS-MFGs. Keywords: Mean-field game of optimal stopping, singular stochastic control, entropy regularization, randomized stopping times, fictitious play algorithm. AMS subject classification: 91A16, 60G40, 93E20, 68T01 1. Introduction The study of mean-field games (MFGs) has become central to the analysis of large-population stochastic control systems, where individual agents interact through the empirical distribution of states and/or controls. For methodologies, techniques, and applications, we refer to the two-volume book [12]. In this paper, we focus on a specific class of such problems: mean-field games of optimal stopping (OS-MFGs). Optimal stopping has a wide range of applications in Economics, Finance, Operations Research, and other applied fields. Relevant examples include pricing American derivatives, entry-exit problems (real options models), and optimal timing for buying or selling an asset, among others. On the other hand, OS-MFGs have received attention only very recently. Among the fundamental contributions on this topic (without learning considered), we first mention [52], where the authors consider a specific game in which the interaction occurs through the number of players who have already stopped. In [13], the authors study an optimal stopping game with Brownian common noise, inspired by a model of bank runs. They prove the strong existence of mean-field equilibria in a setting with strategic complementarity, and weak existence in a setting with continuity. In [7], a purely analytical approach is adopted to solve the problem – namely, the study of the system of coupled Hamilton-Jacobi-Bellman and Fokker-Planck equations – and the existence of mixed solutions is proved. More recently, a compactification approach based on the linear programming formulation is developed in a series of papers [8], [33] and [34], in which the existence of relaxed solutions, interpreted via occupational measures, is established. These techniques have been then applied to solve entry-exit optimal stopping games in electricity markets in [1] and [5], and in the case with common noise in [35]. In [32], a rigorous connection between the occupation measures arising in the linear programming approach and randomized stopping is established. In Date: October 20, 2025. J. Dianetti: Department of Economics and Finance, University of Rome Tor Vergata. Email: jo[email protected]. R. Dumitrescu: ENSAE-CREST, Institut Polytechnique de Paris. Email: ro[email protected]. G. Ferrari: Center for Mathematical Economics, Bielefeld University. Email: [email protected]. R. Xu: Department of Finance and Risk Engineering, NYU. Email: [email protected]. 1
addition, [53] studies an optimal stopping game through the lens of the master equation. Another recent contribution is [46], where the authors develop a new approach to solving various types of MFGs (including optimal stopping) based on a mean-field version of the Bank-El Karoui representation theorem for stochastic processes. In the context of optimal stopping MFGs with partial information, [9,10,59] study the existence of the mean-field solution for a one-dimensional filtering problem under different model set-ups. Finally, related works on McKean-Vlasov optimal stopping are [61] and [21]. Despite their importance, OS-MFGs pose significant theoretical and computational challenges. First, they require more subtle techniques to be solved than MFGs with standard control because of the irregularity of the flow measure generated by the simultaneous exit of a significant number of players. This class of games becomes particularly challenging when the goal is to develop model-free, learning-based solution methods. The primary motivation of our work is to establish the foundations of a suitable OS-MFG framework that supports model-free solution methods. A central challenge in this direction is to rigorously formalize and analyze randomized and explorative strategies in games, which play a crucial role in the design of reinforcement learning (RL) algorithms. Establishing the existence and uniqueness of (mean-field) equilibria that accommodate such strategies is therefore essential—not only for theoretical completeness but also for enabling practical learning algorithms with provable guarantees. In the absence of such a foundation, rigorous analysis of multi-agent learning remains elusive. Moreover, the convergence of RL methods often depends on the properties of the underlying iterative schemes (e.g., fictitious play), whose well-posedness similarly hinges on the existence and uniqueness of equilibria in games that allow for randomized behavior (see, e.g., [40,41]). This need becomes even more pronounced in environments characterized by irregular or sparse decision-making, such as optimal stopping problems, where agents make a single, irreversible decision based on limited information. Unlike standard Markov Decision Processes (MDPs) [54,60], where frequent actions and rewards can reduce the reliance on explicit exploration, optimal stopping games require more deliberate and carefully designed exploration mechanisms to facilitate information acquisition [26]. On the technical side, however, a major challenge arises in applying RL to OS-MFGs. Standard equilibrium-finding algorithms in MFGs typically rely on the stability of optimal controls with respect to variations in the distribution of agents. This stability, however, is particularly difficult to ensure in the context of optimal stopping, where understanding how equilibrium stopping policies respond to small perturbations in the measure flow is highly nontrivial. To illustrate this, consider that in a Markovian setting, equilibrium stopping rules are often given by hitting times of the underlying state process at a free boundary. Determining the stability properties (e.g., Lipschitz continuity) of these free boundaries is already a technically intricate problem, even in single-agent optimal stopping settings (see, e.g., [23] and the discussion therein). Our contributions. To address these challenges, we develop a new theoretical framework that enables the use of model-free algorithms for OS-MFGs. The core idea is to reformulate the game as a MFG of singular stochastic controls (SC-MFG) by introducing randomized stopping times (see [30,62]). A randomized stopping time can be interpreted as the (conditional) probability of stopping before a certain time t. This interpretation leads naturally to an adapted, nonnegative, nondecreasing process with right-continuous sample paths, bounded by 1 – an object that, following terminology from control theory, we refer to as a singular control (see Chapter VIII in [37] for an introduction to singular stochastic controls). To promote exploration, we introduce an entropy regularization term into the stopping functional, as previously established in the single-agent setting of [30]. This regularization term is governed by a temperature parameter λ≥0and induces strict concavity in the performance criterion. This change enables optimizers to deviate from pure 0-1strategies (of not stopping or stopping) and instead favor reflecting-type controls. In the context of MFG, the equilibrium is characterized via a two-step procedure. First, given a fixed flow of measures 2
and a joint distribution, we determine the singular control that maximizes the regularized reward functional. Second, we impose the so-called consistency condition, which requires that the fixed input measures coincide with the distribution of the state process before stopping and the joint distribution of the (randomized) stopping time and the stopped state. Naturally, in this framework, the consistency conditions must be reformulated in terms of randomized stopping times – namely, in terms of the random Borel measure on the time axis induced by the singular control. The particular structure of these consistency conditions prevents us from applying existing existence results for mean-field equilibria in singular control games (see, e.g., [11,25,38,39]). We therefore establish a novel existence result for equilibria in our SC-MFG framework, which we believe is of independent theoretical interest (see Theorem 3.6). Having proved the existence of equilibria in randomized stopping times, we then move on to establishing λ-stability as the temperature parameter λ↓0. We show that any equilibrium of the SC-MFG approximates, as λ↓0(up to a subsequence), an equilibrium of the original MFG of optimal stopping without entropy regularization. Furthermore, under a Lasry-Lions monotonicity condition, we can prove a stronger result: the entire sequence of mean-field equilibria of the singular control game (parametrized by λ) converges to an equilibrium of the original optimal stopping game (see Theorem 3.10). This result ensures that the solutions of the regularized problems converge to equilibria of the original OS-MFG, providing a crucial theoretical justification for using entropybased approximations in RL methods. We then design novel fictitious-play algorithms tailored to the entropy regularized SC-MFG framework. We establish their convergence under both the Lasry–Lions monotonicity condition (see Theorem 4.5) and in the case of supermodular reward structures (see Theorem 4.9). It is important to note that the presence of the entropy regularization term introduces nonlinearity into the performance criterion, placing our model outside the scope of recent approaches based on linear programming for MFGs [8,33,34]. This nonlinearity necessitates the development of new analytical and numerical tools. The fictitious-play algorithm under the Lasry–Lions condition therefore extends previous algorithms developed under this monotonicity assumption to cases where the payoff functional is non-linear with respect to the control. Under the supermodularity condition, the fictitious-play algorithm is new in the literature. The theoretical foundation presented here will be complemented by a companion paper focused on the algorithmic design, related theoretical convergence analysis, and numerical validation of the proposed framework, which is currently in progress [28]. Related literature. Reinforcement learning applied to optimal stopping problems is closely tied to the broader challenge of learning in environments with sparse rewards [26,45]. In these settings, a reward is only granted at the stopping time, resulting in extreme sparsity and significant learning challenges compared to more conventional control tasks, whether formulated in continuous time or as classical MDPs. When the model is fully known, recent studies such as [56] and [58] have introduced deep learning techniques to approximate optimal stopping boundaries. In a related vein, [6] investigated randomized formulations of optimal stopping and provided convergence results for both forward and backward Monte Carlo-based optimization schemes. Similarly, [3] showcased the effectiveness of deep learning methods for solving high-dimensional singular control problems. In the realm of continuous-time RL, a notable advancement comes from [24], who proposed a general policy gradient framework applicable to a wide array of stochastic control problems, including optimal stopping, impulse control, and switching. Their approach leverages connections between stochastic control and its randomized counterparts. However, their framework currently lacks theoretical guarantees for convergence. The work [31] studies a regularized version of the onedimensional American Put option under a Shannon entropy framework. By employing an intensitybased control formulation, they introduce exploration and prove convergence of the policy iteration algorithm (PIA) for fixed temperature parameters. Nevertheless, the convergence of the optimal 3
policy as the temperature parameter λ↓0remains an open question. An alternative perspective was recently offered by [22], who reformulated the optimal stopping problem as a regular control problem with binary (0or 1) actions and applied entropy regularization. This transformed the problem into a classical entropy-regularized control setting, enabling the use of standard RL algorithms. The very recent paper [51] proposes a continuous-time reinforcement learning framework for a class of singular stochastic control problems without entropy regularization and generalizes the existing policy evaluation theories with regular controls to learn optimal singular control law and develop a policy improvement theorem. Another important research direction leverages non-parametric statistical methods for stochastic processes to design learning-based control algorithms [16,17,18,19]. In particular, [17] and [18] addressed one-dimensional singular and impulse control problems by learning critical thresholds using a non-parametric diffusion framework. This approach was extended to multivariate reflection problems in [19], while [16] contributed theoretical guarantees, including regret bounds and nonasymptotic PAC estimates. Entropy regularized MFGs and their reinforcement learning (RL) formulations have received increasing attention over the past year. We do not attempt to provide an exhaustive literature review here, but instead highlight a few key references: [15], which presents an RL approach to model-free MFGs for MDPs; [2], which discusses Q-learning in regularized MFGs for MDPs; [14], which provides a convergence analysis of machine learning algorithms for mean-field control and games over a finite time horizon in continuous time and space; and [36,42], which addresses entropy regularized MFGs in diffusion settings. As far as it concerns the more specific class of entropy-regularized MFGs of optimal stopping – as those considered in this paper – the only paper brought to our attention is [65]. This paper investigates a discrete-time major-minor MFG where the major player can choose to either control or stop. The goal is to find a relaxed (randomized) stopping equilibrium, formulated as a fixed point of a set-valued map, whose direct analysis is difficult due to the major player’s influence. To address this, the authors introduce entropy regularization for the major player’s problem and reformulate the minor players’ stopping problems using linear programming over occupation measures, in the spirit of [8]. They prove the existence of regularized equilibria via a fixed-point theorem and show that, as regularization vanishes, these equilibria converge to a solution of the original problem, thus establishing the existence of a relaxed equilibrium. Although there are clear connections between [65] and the present paper, the mathematical analysis is entirely different, as the problem in [65] involves a Stackelberg game (which is not present in our case) on top of the mean-field interaction, within a discrete-time and discrete-space setting – unlike our continuous-time and continuous-space framework. Organization of the paper. The rest of the paper is organized as follows. In Section 2we present the OS-MFG problem and introduce its entropy-regularized formulation through randomized stopping times. In Section 3we prove existence and stability of equilibria, while in Section 4we present the fictitious play algorithms and their convergence. Finally, Appendix Acollects results on a connection between the studied MFG of singular control and (another) auxiliary MFG of optimal stopping. 1.1. General notation. Let p≥1and (E, d)a non-empty Polish space. Define the set Psub p(E)of subprobability measures νon the Borel subsets of Esuch that REd(z, x0)pν(dz)<∞, for some (and thus all) x0∈E. The space Pp(E)represents the set of elements of Psub p(E)which are probability measures, and it is endowed with the Wasserstein distance dpdefined as dp(¯ν, ν) := inf π∈Π(¯ν,ν)ZE×E d(¯z, z)pπ(d¯z, dz),(1.1) 4
where Π(¯ν, ν)is the set of π∈ Pp(E×E)which has ¯νas first marginal and νas second marginal. Following [20, Section B], to define the Wasserstein distance d′ pon Psub p(E), we introduce a cemetery point ∂and obtain the enlarged space ¯ E:= E∪∂. By defining d(x, ∂) := d(x, x0)+1, we extend the definition of don ¯ E, in such a way that (¯ E, d)is still Polish. We define the classical Wasserstein distance dpon ¯ Eusing the same expression (1.1). The Wasserstein type distance d′ pon Psub p(¯ E)is given by d′ p(µ, ν) := dp(¯µ, ¯ν), where ¯µ(·) := µ(· ∩ E) + (1 −µ(E))δ∂,¯ν(·) := ν(· ∩ E) + (1 −ν(E))δ∂. Fix T > 0. We introduce the set Mp(E)of measurable functions m: [0, T ]→ Psub p(E)identified dt-a.e. on [0.T], endowed with the convergence in measure which is induced by the metric dM p(m, ν) := ZT 0 1∧d′ p(mt, νt)dt, m, ν ∈Mp(E).(1.2) A sequence (mn)nconverges to m(mn→m, in short) if for any ε > 0we have ZT 0 1{t:d′ p(mn t,mt)>ε}dt →0. Recall that, for any convergent sequence mn→m, there exists a subsequence (mnk)ksuch that (1.3) d′ p(mn t, mt)→0, dt-a.e. We refer to Appendix B in [34] for further details. Throughout the paper, C > 0will denote a generic constant that may change from line to line. 2. MFGs of optimal stopping and entropy regularization This section formally introduces the OS-MFG problem with randomization and entropy regularization, leading to a singular control formulation. More specifically, Section 2.1 introduces preliminaries and model assumptions; Section 2.2 describes the OS-MFG problem; and Section 2.3 develops the framework for randomized stopping times, entropy regularization, and the resulting singular control problem. 2.1. Preliminaries and assumptions. Let T∈(0,∞)be given and fixed throughout the rest of the paper. For d, d1∈N\ {0}, consider the continuous functions b: [0, T]×Rd→Rd, σ : [0, T]×Rd→Rd×d1, f : [0, T ]×Rd× Psub(Rd)→R,g:Rd× P(Rd)→R. Define Mp:= Mp(Rd)and M:= Mp× Pp([0, T]×Rd), endowed with the product topology. Elements of Mwill typically be denoted by couples (m, µ). Moreover, given a sequence (mn, µn)n⊂ M and a limit point (m, µ)⊂ M, we will write (mn, µn)n→(m, µ) to denote the convergence in the product topology. On a complete probability space (Ω,F,P), consider a d1-dimensional Brownian motion W:= (Wt)tand a Rd-valued square integrable F0-measurable random variable x0independent of W. Denote by Fx0,W := (Fx0,W t)tthe right-continuous extension of the filtration generated by Wand x0. For a time-horizon T∈(0,∞), let Tbe the set of Fx0,W -stopping times τsuch that τ≤Ta.s.. Assumption 2.1. The data of the problem (x0, b, σ, f, g)verify: (1) E[|x0|p]<∞. 5
(2) There exists a constant L > 0such that |b(t, x)|+|σ(t, x)| ≤ L(1 + |x|),∀t∈[0, T], x ∈Rd, |b(t, ¯x)−b(t, x)|+|σ(t, ¯x)−σ(t, x)| ≤ L|¯x−x|,∀t∈[0, T ], x, ¯x∈Rd. (3) There exists a constant K > 0such that |f(t, x, m)| ≤ K1 + |x|p+ZRd |z|pm(dz),∀(t, x, m)∈[0, T ]×Rd× Psub p(Rd), |g(x, µ)| ≤ K1 + |x|p+Z[0,T ]×Rd |z|pµ(ds, dz),∀(x, µ)∈Rd× Pp([0, T ]×Rd). (4) The functions fand gare continuous, and gis continuous in µ, locally uniformly with respect to x. That is, there exist a constant K > 0and a function wg:Pp([0, T]×Rd)× Pp([0, T]×Rd)→[0,∞)with limdp(µ,¯µ)→0wg(¯µ, µ)=0such that |g(x, ¯µ)−g(x, µ)| ≤ K(1 + |x|p)wg(¯µ, µ)for any x∈Rd, µ, ¯µ∈ Pp([0, T]×Rd). 2.2. The MFG of optimal stopping. Consider the mean-field game of optimal stopping (OSMFG, in short) in which, for a given behavior (m, µ)∈ M of the population of players, the representative player maximizes, over the choice of stopping times τ∈ T , the profit functional (2.1) J(τ, m, µ) := EZτ 0 f(t, Xt, mt)dt +g(Xτ, µ), subject to dXt=b(t, Xt)dt +σ(t, Xt)dWt, t ∈[0, T ], X0=x0. Thanks to Assumption 2.1, there exists a unique strong solution to the stochastic differential equation (SDE) for Xwith continuous paths a.s., and such that (2.2) Ehsup t∈[0,T ] |Xt|pi<∞. Moreover, the profit functional J(τ, m, µ)is well defined for any τ∈ T , the maximization problem has finite value (i.e., supτJ(τ, m, µ)<∞), and there exists an optimal stopping (i.e., τ∗∈ arg maxτJ(τ, m, µ)). We refer the interested reader to Appendix D in [49] for further details. We consider the following notion of equilibrium. Definition 2.2 (OS-MFG equilibrium with strict stopping).An optimal stopping MFG equilibrium (with strict stopping) is a triple (ˆτ, ˆm, ˆµ)∈ T × M such that (2.3) ˆτ∈arg maxτ∈T J(τ, ˆm, ˆµ), ˆmt(A) = P(Xt∈A, t ≤ˆτ),∀A∈ B Rd, t ∈[0, T ), ˆµ(B×[0, t]) = P(Xˆτ∈B, ˆτ≤t),∀B∈ B Rd, t ∈[0, T ]. The first condition is the optimality condition, while the second and the third condition correspond to the so-called consistency conditions. Under the equilibrium notion in Definition 2.2, Problem 2.1 introduces a MFG where agents interact through the running reward function fwith instantaneous mean-field measure mtand through the terminal cost gwith measure µquantifying the stopping time and stopping position. 2.3. Singular control formulation and entropy regularization. In this subsection, we reformulate the OS-MFG problem with the aim of providing a theoretical framework that enables agents to learn equilibria using RL algorithms. To this end, we modify the original OS-MFG framework by incorporating the exploration–exploitation trade-off, which is empirically known to enhance the convergence of learning methods. Fictitious play algorithms will be discussed later in Section 4. 6
2.3.1. Randomized stopping times and singular controls. We first introduce the concept of randomized stopping times. We borrow ideas from the literature in game theory (see, e.g., [62]) and from [30], where the entropy regularized version of an optimal stopping problem is considered. A randomized stopping time can be intrinsically connected to singular control when interpreted as the conditional probability of stopping before a given time t. To build this connection formally, define the set of stochastic processes A:= ξ: Ω ×[0, T]→[0,1],Fx0,W -adapted, nondecreasing, càdlàg, with ξ0−= 0 and ξT= 1. In the rest of this paper, with reference to the terminology of stochastic control theory, we shall refer to an element of Aas to a singular control (see, e.g., Chapter VIII in [37]). Consider a random variable U: Ω →[0,1] which is uniformly distributed and independent from Wand x0. Unless to enlarge the original probability space, we can assume Uto be defined on (Ω,F,P)as well, and to be F-measurable. Given ξ∈ A, a randomized stopping time is defined as τξ:= inf {t∈[0, T]|ξt> U}, with the convention inf ∅=T. Notice that τξis a stopping time with respect to the enlarged filtration Fx0,W,U , generated by x0, W and U. Hence τξis not necessarily an Fx0,W -stopping time. However, if τ∈ T , the process ξτ∈ A defined by ξτ:= (1{t≥τ})tis such that τ=τξτ. This gives a natural inclusion of (ξτ t)t, ξτ t=1{t≥τ}into A. We finally observe that, thanks to the definition of τξand the independence of Uwith respect to Wand x0, we have (2.4) Pτξ≤t| Fx0,W t=PU≤ξt| Fx0,W t=Zξt 0 du =ξt, so that ξtcan be interpreted as the (conditional) probability of stopping before time t. With slight abuse of notation, we will evaluate the payoff Jas in (2.1) on randomized stopping times as well. In particular, using (2.4), an application of the tower property and of the independence of Uwith respect to (W, x0)allows to rewrite the payoff functional for any (m, µ)∈ M as (2.5) J(τξ, m, µ) = E"ZT 0 f(s, Xs, ms) (1 −ξs)ds +Z[0,T ] g(Xs,, µ)dξs#. Here, ξ∈ A is the singular control associated to the randomized stopping time τξ, and for a generic continuous process Y:= (Yt)twe have set R[0,T ]Ysdξs:= Y0ξ0+RT 0Ysdξs. 2.3.2. Consistency conditions for randomized strategies. Equation (2.5) provides the expression of the representative player’s payoff functional when she is allowed to play a randomized stopping rule. We now derive a convenient form of the consistency conditions in (2.3) for randomized stopping times τξin terms of the related singular control ξ∈ A. For any ξ∈ A, by using (2.4), we rewrite the first consistency condition in (2.3) as (2.6) PXt∈A, t < τξ=P(Xt∈A, ξt≤U) =EhEh1{Xt∈A,U⩾ξt}| Fx0,W tii =EZ1 0 1{Xt∈A,ξt≤u}du =EZ1 ξt 1{Xt∈A}dq =E1{Xt∈A}(1 −ξt),∀A∈ B(Rd), t ∈[0, T], 7
(2) Jλ(ξ, mξ, µξ)−Jλ(¯ ξ, mξ, µξ)−(Jλ(ξ, m¯ ξ, µ¯ ξ)−Jλ(¯ ξ, m¯ ξ, µ¯ ξ)) ≤0. Notice that Condition 1in Assumption 3.7 is always satisfied when λ > 0. Remark 3.8. We provide here an example in which Condition (2)in Assumption (3.7)is satisfied. For example, one could consider fand gof the form f(t, x, m) = ¯ k(x)¯ ft, ZRd ¯ k(x)mt(dx), g(x, µ) = ¯ ℓ(x)¯ h Z[0,T ]×Rd ¯ ℓ(x)µ(dt, dx)!, with ¯ fnon-increasing in the second argument and ¯ hnon-decreasing. Theorem 3.9. Under Assumptions 2.1 and 3.7, for any λ≥0there exists a unique equilibrium to λ-SC-MFG (2.10). Proof. The proof of this result this theorem is standard and slightly adapted from [8], we include it for completeness. Arguing by contradiction, assume that there exists two distinct mean-field equilibria (ξ, m, µ), (¯ ξ, ¯m, ¯µ). By uniqueness of the optimal control and the definition of equilibrium, we have Jλ(ξ, mξ, µξ)−Jλ(¯ ξ, mξ, µξ)>0and Jλ(¯ ξ, m¯ ξ, µ¯ ξ)−Jλ(ξ, m¯ ξ, µ¯ ξ)>0. Summing up these two inequalities, we obtain Jλ(ξ, mξ, µξ)−Jλ(¯ ξ, mξ, µξ)−(Jλ(ξ, m¯ ξ, µ¯ ξ)−Jλ(¯ ξ, m¯ ξ, µ¯ ξ)) >0, which contradicts the second condition in Assumption 3.7. Therefore, the equilibrium is unique. □ Entropy regularization and the corresponding temperature parameter λare introduced to encourage randomization in a learning environment. It is important to understand the closeness between the equilibria of the entropy-regularized problem (2.10) – which we aim to learn – and those of the original problem (2.3). We discuss this convergence in the following theorem. Theorem 3.10. Under Assumption 2.1, for any λ > 0, let (ˆ ξλ,ˆmλ,ˆµλ)be an equilibrium to the λ-SC-MFG. Then (1) (ˆ ξλ,ˆmλ,ˆµλ)→(ˆ ξ0,ˆm0,ˆµ0)up to subsequence as λ→0, where (ˆ ξ0,ˆm0,ˆµ0)is an equilibrium to the 0-SC-MFG; (2) Under the additional Assumption 3.7, we have (ˆ ξλ,ˆmλ,ˆµλ)→(ˆ ξ0,ˆm0,ˆµ0)as λ→0. Proof. To simplify the notation, we simply write (ξλ, mλ, µλ)instead of (ˆ ξλ,ˆmλ,ˆµλ), for any λ≥0. Fix any sequence (λn)n⊂(0,∞)with λn→0as n→ ∞ and set (ξn, mn, µn) := (ξλn, mλn, µλn). Since (ξn, mn, µn)are assumed to be equilibria, we have (mn, µn) = Γ(Rλ(ξn)), so that (mn, µn)n⊂ Γ(A). By compactness of A×Γ(A)(see Lemma 3.3), we can extract a subsequence (not relabeled) of (ξn, mn, µn)nand a limit point (ξ0, m0, µ0)such that (ξn, mn, µn)→(ξ0, m0, µ0). Thanks to Lemma 2, we have that ξ0∈R0(m0, µ0). Moreover, by repeating the arguments in the proof of Theorem 3.6, we have that Γ(ξ0) = (m0, µ0). This, in turn implies that ξ0∈R0(m0, µ0) = R0(Γ(ξ0)) = R0(ξ0), so that (ξ0, m0, µ0)is an equilibrium of the 0-SC-MFG. This completes the proof of Claim (1). When the additional Assumption 3.7 holds, by Theorem 3.9 we have the uniqueness of the equilibrium of the 0-SC-MFG. Hence, by the previous argument, any sequence (ξλn, mλn, µλn)converges to (ξ0, m0, µ0), thus proving Claim (2). □ 14
4. Fictitious play algorithms In this section, we introduce two novel fictitious play algorithms for computing the unique meanfield equilibrium λ-SC-MFG in the case when λ > 0, and we establish their convergence under different data assumptions. Note that the convergence of iterative numerical schemes – such as fictitious play algorithms under known model parameters – serves as a fundamental foundation for analyzing the convergence of RL algorithms in environments with unknown parameters. In many popular RL methods, such as Q-learning and actor-critic algorithms [60], the known quantities in these schemes (e.g., the value function) are simply replaced by their corresponding estimates, computed from available observations. Algorithm 1: Fictitious play algorithm Data: A number of steps nfor the equilibrium approximation; a control ¯ ξ0∈ A and (¯µ0,¯m0) := Γ(¯ ξ0)∈ Rλ; Result: Approximate MFG Nash equilibrium 1for k= 0,1, . . . , n −1do 2Set ξk+1 := arg maxξ∈A Jλ(ξ, ¯mk,¯µk); 3Set ¯ ξk+1 := k k+1 ¯ ξk+1 k+1ξk+1 =1 k+1 Pk+1 ℓ=1 ξℓ; 4Set (¯µk+1,¯mk+1) := Γ(¯ ξk+1). 5end Observe that, due to the linear dependence of the measure (m, µ)with respect to ξ, we have that (¯µn,¯mn) = 1 nPn k=1(µξk, mξk). In the next two subsections, we establish the convergence of the fictitious play algorithm ( ¯mn,¯µn)n to the equilibria of the λ-SC-MFG, under suitable structural conditions on the data, and for suitable choice of the initialization. 4.1. Fictitious play under the Lasry-Lions condition. We start with the fictitious play algorithm under the Lasry-Lions monotonicity condition. This fictitious-play algorithm is novel in the context of SC-MFGs and also brings new contributions compared to the fictitious-play algorithm developed in [34] for OS-MFGs. In particular, it extends to cases where the profit-functional is non-linear with respect to therandomized stopping, a setting in which the linear programming approach developed in [34] fails. To simplify the presentation, we set up some notation. First, notice that for a given (m, µ)∈Γ(A) and ξ∈ A, by using the identification between controls and associated measures given by (3.1), we can rewrite Jλas follows: Jλ(ξ, m, µ) = ZT 0ZRdf(t, x, mt)mξ t(dx)dt +g(t, x, µ)µξ(dt, dx)+λEZT 0 E(ξt)dt. We now introduce the bilinear form notation ⟨f(m), m′⟩:= ZT 0ZRd f(t, x, mt)m′ t(dx)dt, ⟨g(µ), µ′⟩:= Z[0,T ]×Rd g(t, x, µ)µ′(dt, dx), where (µ, m),(µ′, m′)∈ Pp([0, T]×R)×Mp. Therefore, Jλwrites as follows: Jλ(ξ, m, µ) = ⟨f(m), mξ⟩+⟨g(µ), µξ⟩+λEZT 0 E(ξt)dt. We assume the following conditions hold: 15
Assumption 4.1. There exist constants cf>0and cg>0such that for all t∈[0, T ],x, x′∈Rd, m, m′∈ Psub p(Rd),µ, µ′∈ Pp([0, T]×Rd), f(t, x, m)−f(t, x, m′)≤cf(1 + |x|)ZRd (1 + |z|p)|m−m′|(dz), |f(t, x, m)−f(t, x, m′)−f(t, x′, m) + f(t, x′, m′)| ≤ cf|x−x′|ZRd (1 + |z|p)|m−m′|(dz), |g(t, x, µ)−g(t, x, µ′)| ≤ cg(1 + |x|)Z[0,T ]×Rd (1 + |z|p)|µ−µ′|(ds, dz), |g(t, x, µ)−g(t, x, µ′)−g(t′, x′, µ)+g(t′, x′, µ′)| ≤ cg(|t−t′|+|x−x′|)Z[0,T ]×Rd (1+|z|p)|µ−µ′|(ds, dz), where |m|denotes the total variation measure of m. Before proving the main result on the convergence of the fictitious play algorithm, we first recall from [34] some estimates (on the linear part of the payoff functional, as well as on the admissible measures), which will be needed in the proof of Theorem 4.5 below. Lemma 4.2. (Estimates on distances) For all n≥1, we have the following estimates: (i) There exists a constant C1>0such that |dM 1( ¯mn+1,¯mn)| ≤ C1 n,|d1(¯µn+1,¯µn)| ≤ C1 n. (ii) There exists a constant C2>0(independent of n) such that for all t∈[0, T ]: Z[0,T ]×R (1 + |x|p)|¯µn+1 −¯µn|(dt, dx)≤C2 n,Z[0,T ]×R (1 + |x|p)|¯mn+1 t−¯mn t|(dt, dx)≤C2 n.(4.1) Lemma 4.3 (Estimates on the rewards).There exist constants Cf>0and Cg>0such that for all n≥1 ⟨f( ¯mn+1)−f( ¯mn), mn+2 −mn+1⟩ ≤ Cf ndM 1(mn+1, mn+2), ⟨g(¯µn+1)−g(¯µn), µn+2 −µn+1⟩ ≤ Cg nd1(µn+1, µn+2). Lemma 4.4. (Proximity between two successive best responses) The following convergence results hold: lim n→∞ d1(µn, µn+1)=0,lim n→∞ dM 1(mn, mn+1)=0.(4.2) Proof. We follow the same arguments as in [34] by using the continuous map Γ◦Rλfrom (M, dM 1⊗d1) to (M, dM 1⊗d1).□ We now give the main convergence result of this section. Theorem 4.5. Under Assumptions 2.1,3.7 and 4.1, for any initialization (¯ ξ0,¯m0,¯µ0), the sequence (¯ ξn,¯mn,¯µn)ndefined in Algorithm 1 converges to the unique equilibrium distribution (ˆ ξλ,ˆmλ,ˆµλ). Proof. The proof is organized in two steps. Step 1. We first introduce the exploitability errors (εn)n≥1which are defined as follows: εn:= Jλ(ξn+1,¯mn,¯µn)−Jλ(¯ ξn,¯mn,¯µn) =⟨f( ¯mn), mn+1 −¯mn⟩+⟨g(¯µn), µn+1 −¯µn⟩+λEZT 0 (E(ξn+1 t)− E(¯ ξn t))dt.(4.3) By the optimality of ξn+1, it is easy to observe that εn≥0, for all n≥1. 16
We write εn+1 −εn=ε(1) n+ε(2) n, where ε(1) n:= ⟨f( ¯mn),¯mn⟩+⟨g(¯µn),¯µn⟩−⟨f( ¯mn+1),¯mn+1⟩−⟨g(¯µn+1),¯µn+1⟩ +λEZT 0 (E(¯ ξtn)− E(¯ ξt n+1))dt,(4.4) ε(2) n:= ⟨f( ¯mn+1), mn+2⟩+⟨g(¯µn+1), µn+2⟩−⟨f( ¯mn), mn+1⟩−⟨g(¯µn), µn+1⟩ +λEZT 0 (E(ξn+2 t)− E(ξn+1 t)))dt.(4.5) In the following C > 0denotes a constant (independent of n) that may vary from line to line. We first estimate ε(1) n. First, by the concavity of the entropy functional E(·), we get EZT 0E(¯ ξn t)− E(¯ ξn+1 t)dt=EZT 0E(¯ ξn t)− E n n+ 1 ¯ ξn t+1 n+ 1ξn+1 t)dt ≤1 n+ 1EZT 0E(¯ ξn t)− E(ξn+1 t)dt.(4.6) Then, by using the estimates from Lemma 4.3, we have −1 n+ 1⟨f( ¯mn+1)−f( ¯mn), mn+1 −¯mn⟩ ≤ 1 n+ 1⟨|f( ¯mn+1)−f( ¯mn)|, mn+1 + ¯mn⟩ ≤C n2. We deduce that ⟨f( ¯mn),¯mn⟩−⟨f( ¯mn+1),¯mn+1⟩=⟨f( ¯mn),¯mn⟩ − ⟨f( ¯mn+1),¯mn+1 n+ 1(mn+1 −¯mn)⟩ =⟨f( ¯mn)−f( ¯mn+1),¯mn⟩ − 1 n+ 1⟨f( ¯mn+1), mn+1 −¯mn⟩ ≤ ⟨f( ¯mn)−f( ¯mn+1),¯mn⟩ − 1 n+ 1⟨f( ¯mn), mn+1 −¯mn⟩+C n2. Similar to the above inequality, we also derive ⟨g(¯µn),¯µn⟩−⟨g(¯µn+1),¯µn+1⟩ ≤ ⟨g(¯µn)−g(¯µn+1),¯µn⟩ − 1 n+ 1⟨g(¯µn), µn+1 −¯µn⟩+C n2. Therefore, by combining the last three inequalities, we derive ε(1) n=⟨f( ¯mn),¯mn⟩+⟨g(¯µn),¯µn⟩−⟨f( ¯mn+1),¯mn+1⟩−⟨g(¯µn+1),¯µn+1⟩ +λEZT 0 (E(¯ ξtn)− E(¯ ξt n+1))dt ≤ ⟨f( ¯mn)−f( ¯mn+1),¯mn⟩+⟨g(¯µn)−g(¯µn+1),¯µn⟩ − εn n+ 1 +C n2. We shall now proceed with the estimation of ε(2) n. First notice that, since Jλ(ξn+1,¯mn,¯µn)≥Jλ(ξ, ¯mn,¯µn),∀ξ∈ A,(4.7) 17
it follows that ⟨f( ¯mn), mn+1⟩+⟨g(¯µn), µn+1⟩+λEZT 0 E(ξn+1 t)dt ≥ ⟨f( ¯mn), mn+2⟩+⟨g(¯µn), µn+2⟩+λEZT 0 E(ξn+2 t)dt.(4.8) We therefore have ε(2) n=⟨f( ¯mn+1), mn+2⟩+⟨g(¯µn+1), µn+2⟩+λEZT 0 (E(ξn+2 t) − ⟨f( ¯mn), mn+1⟩−⟨g(¯µn), µn+1⟩ − λEZT 0 E(ξn+1 t))dt ≤ ⟨f( ¯mn+1), mn+2⟩+⟨g(¯µn+1), µn+2⟩−⟨f( ¯mn), mn+2⟩−⟨g(¯µn), µn+2⟩ =⟨f( ¯mn+1)−f( ¯mn), mn+2⟩+⟨g(¯µn+1)−g(¯µn), µn+2⟩ =⟨f( ¯mn+1)−f( ¯mn), mn+1⟩+⟨g(¯µn+1)−g(¯µn), µn+1⟩ +⟨f( ¯mn+1)−f( ¯mn), mn+2 −mn+1⟩+⟨g(¯µn+1)−g(¯µn), µn+2 −µn+1⟩. By using the estimates of Lemma 4.3, we obtain ⟨f( ¯mn+1)−f( ¯mn), mn+2 −mn+1⟩ ≤ Cf ndM 1(mn+1, mn+2), ⟨g(¯µn+1)−g(¯µn), µn+2 −µn+1⟩ ≤ Cg nd1(µn+1, µn+2). Therefore ε(2) n≤ ⟨f( ¯mn+1)−f( ¯mn), mn+1⟩+⟨g(¯µn+1)−g(¯µn), µn+1⟩ +C n(dM 1(mn+1, mn+2) + d1(µn+1, µn+2)). We set δn:= CdM 1(mn+1, mn+2) + d1(µn+1, µn+2) + 1 n, and we derive εn+1 −εn=ε(1) n+ε(2) n≤ ⟨f( ¯mn)−f( ¯mn+1),¯mn⟩+⟨g(¯µn)−g(¯µn+1),¯µn⟩ − εn n+ 1 +⟨f( ¯mn+1)−f( ¯mn), mn+1⟩+⟨g(¯µn+1)−g(¯µn), µn+1⟩+δn n =⟨f( ¯mn+1)−f( ¯mn), mn+1 −¯mn⟩+⟨g(¯µn+1)−g(¯µn), µn+1 −¯µn⟩ − εn n+ 1 +δn n = (n+ 1) ⟨f( ¯mn+1)−f( ¯mn),¯mn+1 −¯mn⟩+⟨g(¯µn+1)−g(¯µn),¯µn+1 −¯µn⟩ −εn n+ 1 +δn n ≤ − εn n+ 1 +δn n, where the last inequality comes from the Lasry–Lions monotonicity condition. Observe that δn→0 by Lemma 4.4. By Lemma 3.1 in [43], we conclude that εn→0as n→ ∞. The result follows. Step 2. We show now that the sequence (¯ ξn,¯mn,¯µn)converges to the unique λ-SC-MFG equilibrium (ˆ ξλ,ˆmλ,ˆµλ).By compactness of the set A × Γ(A), from any subsequence (¯ ξkn,¯mkn,¯µkn) we can subtract a further subsequence (not relabeled) such that (¯ ξkn,¯mkn,¯µkn)converges in the 18
topology τw 2⊗dp⊗Wpto some (¯ ξ, ¯m, ¯µ), where τw 2represents the topology associated to the weak convergence in L2(Ω ×[0, T]). First, by using similar arguments as for the proof of (3.9), we can establish that ( ¯m, ¯µ) = Γ(¯ ξ).(4.9) Then, by optimality of ξkn+1 ·, we get Jλ(ξkn+1 ,¯mkn,¯µkn)≥Jλ(ξ, ¯mkn,¯µkn),for all ξ∈ A.(4.10) By definition of εkn, we derive Jλ(¯ ξkn,¯mkn,¯µkn)≥Jλ(ξ, ¯mkn,¯µkn)−εkn,for all ξ∈ A.(4.11) By using similar arguments as the ones used to establish (3.6), (3.7) and (3.8) and by using the convergence of (εn)to 0when n→ ∞, we get Jλ(¯ ξ, ¯m, ¯µ)≥Jλ(ξ, ¯m, ¯µ),for all ξ∈ A.(4.12) From the above inequality and relation (4.9), we conclude that (¯ ξ, ¯m, ¯µ)is a λ-SC-MFG equilibrium. By uniqueness of the equilibrium, it coincides with (ˆ ξλ,ˆmλ,ˆµλ). Since from any subsequence of (¯ ξn,¯mn,¯µn)we can subtract a further subsequence which converges to the unique λ-SC-equilibrium, we conclude that the whole sequence converges to the unique MFG equilibrium. □ 4.2. Fictitious play in the supermodular case. We now provide convergence of the fictitious play algorithm when the so-called supermodularity condition holds true (see, e.g., [27,29,64]). Such a property naturally appears in many financial and economic applications, as for example in models for bank runs (see [13] and Example 4.7 below). The interested reader is referred to the textbook [63] for further details on supermodular games. We thus enforce the following structural condition: Assumption 4.6 (Supermodularity).The following hold true: (1) For any (m, µ)∈ M,Rλ(m, µ)is a singleton. (2) For each (t, x)∈[0, T ]×Rd, the function f(t, x, ·)is nondecreasing with respect to the measure argument in the following sense: for any two measures m, ¯m∈ Psub p(Rd), if m(A)≤ ¯m(A)for all A∈ B(Rd), then f(t, x, m)≤f(t, x, ¯m). (3) We have g(x, µ) = ˜g(x, ⟨ψ, µ⟩), for some functions ˜g∈ C2(Rd×R)and ψ∈ C1,2([0, T]×Rd) such that (∂t+L)ψ≥0and y7→ L˜g(x, y)is nondecreasing for any x∈Rd. For the sake of illustration, we discuss some key examples. Example 4.7. An example in which the supermodularity condition is satisfied is when the function gdoes not depend on µ(in this case, Condition (3)above is not needed) and f(t, x, m) = f0(t, x, 1− m(Rd)), with f0(t, x, ·)nonincreasing for any t, x. This setting represents supermodular MFGs with equilibrium condition ˆτ∈arg max τ EZτ 0 f0(t, Xt,ˆmt)dt +g(Xτ)and ˆmt=P[ˆτ≤t], which are also referred to as MFGs of timing. The supermodularity condition is typical of models for bank runs [13]. Natural examples are of additive type f0(t, x, y) = f1(x)−f2(y)or of multiplicative type f0(t, x, y) = −f1(x)f2(y)(with f1≥0), for nondecreasing functions f1, f2. 19
Example 4.8. The underlying process evolves as a one dimensional geometric Brownian motion dZt=Zt(b0dt +σ0dWt), Z0=z0>0,P-a.s., with b0∈R,σ0>0, and the payoffs are f(t, z, m) = ZR (z+y)m(dy)and g(t, z, µ) = ZR×[0,T ] (t+s)µ(ds, dy). With respect to the previous notation, the state process becomes Xt:= (t, Zt)so that time-dependence in gis allowed. Before discussing the main result of this subsection, we recall that the controls ξare interpreted as probability measures; in particular, ξtis the probability of stopping before time t. In this sense, if ξt≥¯ ξtfor any t∈[0, T ],P-a.s., then τξ≤τ¯ ξP-a.s., and we say that ξis earlier than ¯ ξ(or that ¯ ξis later than ξ). Given two equilibria (ξ, m, µ)and (¯ ξ, ¯m, ¯µ), we say that the first equilibrium is earlier (resp. later) that the second one if ξis earlier (resp. later) that ¯ ξ. Given λ≥0, the earliest equilibrium (ξλ, mλ, µλ)and the latest equilibrium (ξλ, mλ, µλ)are such that ξλ t≥ξt≥ξλ tfor any t∈[0, T],P-a.s., for any other equilibrium (ξ, m, µ). The following theorem discusses the existence and approximation of the earliest and latest equilibria. Theorem 4.9. Under Assumptions 2.1 and 4.6, the earliest equilibrium (ξλ, mλ, µλ)and the latest equilibrium (ξλ, mλ, µλ)exist. Moreover, the following statements hold true: (1) With initialization ¯ ξ0 t= 1 for any t∈[0, T ], the sequence ¯ ξnis nonincreasing, ¯ ξn+1 ≤¯ ξn, and (¯ ξn,¯mn,¯µn)converges to the earliest equilibrium (ξλ, mλ, µλ). (2) With initialization ¯ ξ0 t= 0 for any t∈[0, T ], the sequence ¯ ξnis nondecreasing, ¯ ξn+1 ≥¯ ξn, and (¯ ξn,¯mn,¯µn)converges to the latest equilibrium (ξλ, mλ, µλ). Proof. We limit ourself to show the claim on the existence and approximation of the earliest equilibrium, as the proof for the latest equilibrium follows by the same argument. The proof is divided into three steps. Step 1. In this step we show the some monotonicity properties of the maps Γand Rλ. We first show anti-monotonicity property of the consistency map Γ. Observe that, if ξt≤ξ′ t,∀t∈ [0, T],P-a.s., from the definition of mξ, mξ′have mξ t(A) = E[1A(Xt)(1 −ξt)] ≥E[1A(Xt)(1 −ξ′ t)] = mξ′ t(A),∀A∈ B(Rd), t ∈[0, T]. Moreover, for ψas in Assumption 4.6, by using integration by parts and that ξT= 1, we have ⟨ψ, µξ⟩:= EZ[0,T ] ψ(t, Xt)dξt=Eψ(T, XT)−ZT 0 ξt(∂t+L)ψ(t, Xt)dt, and the analogous expression can be obtained for ξ′. Thus, since (∂t+L)ψ≥0, we find ⟨ψ, µξ⟩=Eψ(T, XT)−ZT 0 ξt(∂t+L)ψ(t, Xt)dt ≥Eψ(T, XT)−ZT 0 ξ′ t(∂t+L)ψ(t, Xt)dt=⟨ψ, µξ′⟩. To summarize, we have found that (4.13) if ξt≤ξ′ t,∀t, P-a.s., then mξ t(A)≥mξ′ t(A),∀A, t, and ⟨ψ, µξ⟩ ≥ ⟨ψ, µξ′⟩, which is the desired anti-monotonicity of Γ. 20
We next show the anti-monotonicity of the best reply map Rλ; that is, for any m, m′, (4.14) if mt(A)≥m′ t(A),∀A, t, and ⟨ψ, µ⟩ ≥ ⟨ψ, µ′⟩, then Rλ(m, µ)t≤Rλ(m′, µ′)t,∀t, P-a.s. To see this, set ξ:= Rλ(m, µ), ξ′:= Rλ(m′, µ′), and define the processes ξ∧ξ′:= min{ξt, ξ′ t}tand ξ∨ξ′:= max{ξt, ξ′ t}t. For a generic process ζ∈ A, using integration by parts for the integral in dζ, we find Jλ(ζ, m, µ) = E−ZT 0f(t, Xt, mt) + Lg(Xt, µ)ζtdt +ZT 0f(t, Xt, mt) + λE(ζt)dt +g(XT, µ), and, defining ˆ f(t, x, m, µ) = f(t, x, m) + Lg(x, µ), we rewrite the latter expression as Jλ(ζ, m, µ) = E−ZT 0 ˆ f(t, Xt, mt, µ)ζtdt +ZT 0f(t, Xt, mt) + λE(ζt)dt +g(XT, µ). Now, noticing that ξ′ t−ξt∧ξ′ t=ξt∨ξ′ t−ξtand that E(ξt∨ξ′ t)− E(ξ′ t) = E(ξt)− E(ξt∧ξ′ t), we first find Jλ(ξ∨ξ′, m′, µ′)−Jλ(ξ′, m′, µ′) =EZT 0ˆ f(t, Xt, m′ t, µ′)ξ′ t−ξt∨ξ′ t+λ(E(ξt∨ξ′ t)− E(ξ′ t))dt =EZT 0ˆ f(t, Xt, m′ t, µ′)ξt∧ξ′ t−ξt+λ(E(ξt)− E(ξt∧ξ′ t))dt. Thus, if mt(A)≥m′ t(A),∀A, t, and ⟨ψ, µ⟩≥⟨ψ, µ′⟩, from Conditions 2and 3in Assumption 4.6 we have Jλ(ξ∨ξ′, m′, µ′)−Jλ(ξ′, m′, µ′) ≥EZT 0ˆ f(t, Xt, mt, µt)ξ′ t∧ξt−ξt+λ((E(ξt)− E(ξt∧ξ′ t))dt =Jλ(ξ, m, µ)−Jλ(ξ∧ξ′, m, µ). Hence, from the optimality of ξ′for (m′, µ′)and the fact that ξ∨ξ′∈ A, we deduce that 0≥Jλ(ξ∨ξ′, m′, µ′)−Jλ(ξ′, m′, µ′)≥Jλ(ξ, m, µ)−Jλ(ξ∧ξ′, m, µ). This in turn implies that ξ∧ξ′=Rλ(m, µ), and, by uniqueness of the optimizer as in Condition 1 in Assumptoin 4.6, we conclude that ξ∧ξ′=ξ, thus proving (4.14). Step 2. In this step we show the monotonicity of ¯ ξnby an induction argument. To simplify the notation, we will write m≥m′instead of mt(A)≥m′ t(A),∀A, t. Moreover, (mk, µk) := Γ(ξk), so that (due to the linear dependence of the measure (m, µ)with respect to ξ) we have the relation (¯µn,¯mn) = 1 n n X k=1 (µξk, mξk) = 1 n n X k=1 (µk, mk), that will be used several times in the sequel. We proceed with an induction argument. Since ¯ ξ0≡1and ( ¯m0,¯µ0)=(mξ0, µξ0), we have ξ1=Rλ( ¯m0,¯µ0)≤¯ ξ0, (m1, µ1)=(mξ1, µξ1), ¯m1=m1=mξ1≥m¯ ξ0= ¯m0, ⟨ψ, ¯µ1⟩=⟨ψ, µ1⟩=⟨ψ, µξ1⟩≥⟨ψ, µ¯ ξ0⟩=⟨ψ, ¯µ0⟩. 21
Hence, by the monotonicity of Rλin (4.14) and of Γin (4.13), we have ξ2=Rλ( ¯m1,¯µ1)≤Rλ( ¯m0,¯µ0) = ξ1, m2=mξ2≥mξ1=m1, ⟨ψ, µ2⟩=⟨ψ, µξ2⟩≥⟨ψ, µξ1⟩=⟨ψ, µ1⟩, ¯m2=1 2(m2+m1)≥m1= ¯m1, ⟨ψ, ¯µ2⟩=⟨ψ, 1 2(µ2+µ1)⟩≥⟨ψ, µ1⟩≥⟨ψ, ¯µ0⟩. Assume that for a generic nwe have ξn≤ξn−1≤ · · · ≤ ξ1, mn=mξn≥mn−1≥ · · · ≥ m1, ⟨ψ, µn⟩=⟨ψ, µξn⟩≥⟨ψ, µn−1⟩ ≥ · · · ≥ ⟨ψ, µ1⟩, ¯mn=1 n n X k=1 mk≥¯mn−1≥ · · · ≥ ¯m1, ⟨ψ, ¯µn⟩=⟨ψ, 1 n n X k=1 µk⟩≥⟨ψ, ¯µn−1⟩ ≥ · · · ≥ ⟨ψ, ¯µ1⟩. Then, the monotonicity of Rλin (4.14) and of Γin (4.13), again gives ξn+1 =Rλ( ¯mn,¯µn)≤Rλ( ¯mn−1,¯µn−1) = ξn, mn+1 =mξn+1 ≥mξn=mn≥ · · · ≥ m1, ⟨ψ, µn+1⟩=⟨ψ, µξn+1 ⟩≥⟨ψ, µξn⟩=⟨ψ, µn⟩ ≥ · · · ≥ ⟨ψ, µ1⟩. Moreover, from the last two inequalities we deduce that mn+1 ≥¯mn=1 n n X k=1 mkand ⟨ψ, µn+1⟩ ≥ ⟨ψ, ¯µn⟩=⟨ψ, 1 n n X k=1 µk⟩, which allows to conclude that ¯mn+1 =1 n+ 1(mn+1 +n¯mn)≥¯mn, ⟨ψ, ¯µn+1⟩=⟨ψ, 1 n+ 1(µn+1 +n¯µn)⟩ ≥ ⟨ψ, ¯µn⟩, thus completing the induction argument. Finally, the monotonicity of the sequence ¯ ξnfollows from the monotonicity of ξn. Step 3. We next define (thanks to the monotonicity of ξnand ¯ ξn) the limit points ξ:= inf nξn= inf n ¯ ξn, m := mξ,and µ:= µξ, and show that these coincide with the minimal equilibria (ξλ, mλ, µλ). Notice that the sequence (¯ ξn)nconverges strongly. Moreover, by monotonicity, it is easy to see that the consistency map Γis continuous along this sequence. Therefore, we have (m, µ) = (mξ, µξ) = Γ(ξ) = lim nΓ(¯ ξn) = lim n( ¯mn,¯µn), 22
so that the sequence ( ¯mn,¯µn)nconverges to (m, µ)as well. Hence, the continuity of Rλin turn implies that ξ= lim nRλ( ¯mn,¯µn) = Rλ(m, µ), which proves the optimality. This show that (ξ, m, µ)is a mean-field equilibrium. Finally, we show the minimality of (ξ, m, µ). Let (ˆ ξ, ˆm, ˆµ)be another mean-field equilibrium. By definition of ¯ ξ0, we have ˆ ξ≤¯ ξ0. Hence, by monotonicity of Γin (4.13), we have mˆ ξ= ˆm≥m1= mξ1= ¯m1and ⟨ψ, µˆ ξ⟩=⟨ψ, ˆµ⟩≥⟨ψ, µ1⟩=⟨ψ, µξ1⟩=⟨ψ, ¯µ1⟩. Hence, using the monotonicity of Rλso that ˆ ξ=Rλ( ˆm, ˆµ)≤Rλ( ¯m1,¯µ1) = ξ2by monotonicity of Rλ. Proceeding by induction, we conclude that ˆ ξ≤¯ ξn,∀n, so that ˆ ξ≤ξ= infn¯ ξn, thus giving the desired minimality. □ Appendix A. Connection between the λ-singular control MFG problem and a λ-optimal stopping MFG problem In this part, we present the connection between the singular mean-field problem λ-SC-MFG and a new type of OS-MFG problems, that we introduce below and we call λ-OS-MFG (λ-optimal stopping MFG). Let ¯ Tdenote the set of Fx0,W,U -stopping times with values in [0, T ], where the filtration Fx0,W,U has been introduced in Section 2.3. For ease of exposition, we restrict ourselves to the relative entropy E(z) = −zlog z; the analysis for a general entropy function satisfying Assumption 2.3 is analogous. Definition A.1. A triple ( ˆmλ,ˆµλ,ˆ θλ)is said to be an equilibrium to the λ-OS-MFG problem if: ˆ θλ∈arg maxθ∈¯ TEhRθ 0f(t, Xt,ˆmλ t)−λ(1 + log U)dt +g(Xθ,ˆµλ)i ˆmλ t(A) = PhXt∈A, t ≤ˆ θλi, A ∈ B(Rd) ˆµλ(B×[0, t]) = PhXˆ θ∈B, ˆ θλ≤tiB∈ B(Rd). Observe that, for any stopping time θ∈¯ T, since θis a Fx0,W,U -stopping time, then there exists a measurable function, still denoted by θ, such that θ=θ(x0, W, U). We now establish the first relation between the two game problems. Theorem A.2. Let ( ˆmλ,ˆµλ,ˆ ξλ)be an equilibrium to the λ-SC-MFG. Define ˆ θλ:= θλ(·,1−U),(A.1) where θλ(·, u) := inf{t≥0|ˆ ξλ t≥u}.(A.2) Then ( ˆmλ,ˆµλ,ˆ θλ)is an equilibrium to the λ-OS-MFG . Proof. Let ( ˆmλ,ˆµλ,ˆ ξλ)be an λ-SC-MFG equilibrium and let ˆ θλbe given by (A.1). Observe that {t < ˆ θλ}={ˆ ξλ t<1−U}.(A.3) In view of this observation, we obtain: E"Zˆ θλ 0 f(t, Xt,ˆmλ t)dt#=EZT 0 f(t, Xt,ˆmλ t)1{t<ˆ θλ}dt=EZT 0 f(t, Xt,ˆmλ t)P(U < 1−ˆ ξλ t|Fx0,W t)dt =EZT 0 f(t, Xt,ˆmλ t)(1 −ˆ ξλ t)dt23
