scieee AI-readable full text Open interactive document viewer

Mean-field ranking games with diffusion control

Ankirchner, S.,Kazi-Tani, N.,Wendt, J.,Zhou, C.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Ankirchner, S.; Kazi-Tani, N.; Wendt, J.; Zhou, C. Article — Published Version Mean-field ranking games with diffusion control Mathematics and Financial Economics Provided in Cooperation with: Springer Nature Suggested Citation: Ankirchner, S.; Kazi-Tani, N.; Wendt, J.; Zhou, C. (2024) : Mean-field ranking games with diffusion control, Mathematics and Financial Economics, ISSN 1862-9660, Springer, Berlin, Heidelberg, Vol. 18, Iss. 2, pp. 313-331, https://doi.org/10.1007/s11579-024-00354-2 This Version is available at: https://hdl.handle.net/10419/315443 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/ Mathematics and Financial Economics (2024) 18:313–331 https://doi.org/10.1007/s11579-024-00354-2 Mean-field ranking games with diffusion control S. Ankirchner1 ·N. Kazi-Tani2 ·J. Wendt1 ·C. Zhou3 Received: 30 June 2023 / Accepted: 9 January 2024 / Published online: 26 March 2024 © The Author(s) 2024 Abstract We consider a stochastic differential game, where each player continuously controls the diffusion intensity of her own state process. The players must all choose from the same diffusion rate interval [σ1,σ 2], and have individual random time horizons that are independently drawn from the same distribution. The players whose states at their respective time horizons are among the best p∈(0,1)of all terminal states receive a fixed prize. We show that in the mean field version of the game there exists an equilibrium, where the representative player chooses the maximal diffusion rate when the state is below a given threshold, and the minimal rate else. The symmetric n-fold tuple of this threshold strategy is an approximate Nash equilibrium of the n-player game. Finally, we show that the more time a player has at her disposal, the higher her chances of winning. Keywords Diffusion control ·Game ·Rank-based reward ·Mean field limit ·Oscillating Brownian motion Mathematics Subject Classification Primary: 91A15; secondary: 91A06 ·91A10 ·91A16 · 93E20 BS. Ankirchner [email protected] N. Kazi-Tani [email protected] J. Wendt [email protected] C. Zhou [email protected] 1Institute for Mathematics, University of Jena, Ernst-Abbe-Platz 2, 07743 Jena, Germany 2Institut Elie Cartan de Lorraine, Université de Lorraine, UFR MIM, 3 rue Augustin Fresnel, 57073 Metz Cedex 03, France 3Department of Mathematics and Risk Management Institute, National University of Singapore, 10 Lower Kent Ridge Road, 119076 Singapore, Singapore 123 314 Mathematics and Financial Economics (2024) 18:313–331 1 Introduction We consider the following stochastic differential game: each player can control the fluctuation intensity of her own state process, up to an individual random time horizon. The controls of a player are individually defined as a set of progressively measurable processes, with respect to a filtration modeling the player’s information flow, with values in a bounded interval [σ1,σ 2], where 0 <σ 1<σ 2are the same for all. The players whose terminal states are among the highest p∈(0,1)receive a fixed prize, set to be equal to one. The other players do not receive anything. The game models in stylized form competitions where only the best performing agents receive a fixed reward and where every agent can choose between risky and safe actions. The game thus allows to analyze the impact of rank-based rewards on the risk appetite of the competitors. The game has multiple interpretations. We refer to Sect. 6for more details. As usual, we fall back on the concept of Nash equilibria for predicting the players’ behavior. Given the discontinuous rewards, it turns out to be difficult to compute explicit equilibria. Moreover, it seems already difficult to prove existence by abstract means. A way out for games with many players is to fall back on the game’s mean field version and to derive an approximate equilibrium. The equilibrium of the mean field game here consists of a control and the respective state distribution at the terminal time. Under regularity conditions on the rewards and on the state equation coefficients, existence of equilibria in mean field games with diffusion control and common noise is proved in [2] using a relaxed formulation of the problem and the theory of second order BSDEs. Other contributions consider mean field games with diffusion control, such as [8] in a Principal-Agent setting, with applications to optimal energy demand management or [6,7] in the case of extended mean field games with control interactions. In this paper, we avoid second order BSDEs and show, using a direct argument, that there exists an equilibrium with a threshold control that consists in choosing the maximal diffusion rate σ2when the state is below a given threshold, and the minimal rate σ1else. The corresponding equilibrium distribution is the distribution of an oscillating Brownian motion at the terminal time. The game bears similarities with the diffusion control game studied in [1]. In contrast to [1], however, we allow here the agents to be heterogeneously informed. Each player is assumed to observe her own state process, but the assumption on how much the agents know beyond this remains general. Some of the players may know, e.g., the time horizons, and some not. Whether a player knows her own time horizon or the time horizons of the opponents turns out to have no effect on the equilibrium. Indeed, in the many player game the empirical distribution of the time horizons is close to T. Thus knowing Tis sufficient for implementing an approximate equilibrium control. Similarly, it has no effect on the equilibrium whether the agents can observe the state processes of the opponents or not. For implementing the threshold control of the equilibrium each agent needs only to oberve her own process. By defining the state dynamics in terms of solutions to controlled martingale problems and choosing controls of open loop type, we obtain a model framework that allows to cover general information structures, in particular situations where a player has only partial knowledge about the other players’ states and about the individual random time horizons. In the game of [1] the state dynamics are described in terms of stochastic differential equations and the players’ controls are modeled as closed loop controls. In the appropriate mean field version of the game the representative player knows the distribution Tof the time horizon, but the actual time horizon arrives unpredicted. The 123 Mathematics and Financial Economics (2024) 18:313–331 315 threshold describing the equilibrium control strongly depends on the distribution T.Forthe cases where Tis an exponential distribution or a uniform distribution, we characterize the threshold of the equilibrium control as the unique root of a simple equation. The article is organized as follows: in Sect.2, we introduce the game model in more detail. For deriving a candidate for an approximate Nash equilibrium, we study the corresponding mean field game in Sect.3and show that an equilibrium control is given by a threshold control. In Sect.4, we show that the n-tuple consisting of the mean field equilibrium controls is an O(n−1/2)-Nash equilibrium of the n-player game. This means, in the approximate Nash equilibrium players only use information about their own state and choose maximal diffusion intensity below the optimal threshold and minimal diffusion intensity above. All additional information about the opponents’ states or the random times is irrelevant. In Sect.5,we analyze the winning probability of a given player depending on the time horizon Tand in Sect.6, we provide a particular application of our model setting to online competitions. Finally, we consider in Sect. 7an extension of the game to more general reward functions that are continuous, have exponential growth, satisfy a symmetry condition, and are convex below a certain threshold and concave above. We show that the same tuple consisting of the mean field equilibrium controls is also an approximate Nash equilibrium for these reward functions. 2 Game model We describe the players’ states by means of controlled martingale problems. To this end, let R+:= [0,∞),andC(R+,Rn)denote the space of continuous functions f:R+→Rn equipped with the metric d(ω1,ω 2):= ∞  n=1 1 2n supt∈[0,n]|ω1(t)−ω2(t)| 1+supt∈[0,n]|ω1(t)−ω2(t)|,ω 1,ω 2∈C(R+,Rn). Let B(C(R+,Rn))denote the corresponding Borel σ-algebra on C(R+,Rn)and denote by X=(X1,...,Xn)the canonical process on C(R+,Rn), i.e., Xi t(ω) := ωi(t),t≥0,for i=1,...,nand ω∈C(R+,Rn). We refer to [18], Section 1.3, for more details on the construction of this measurable space and its properties. Let Tbe a probability measure on R+, equipped with the usual Borel σ-algebra B(R+), that describes the distribution of the players’ random times. We suppose that Tsatisfies: Assumption 2.1 ∞ 0 1 √tT(dt)<∞. Let BRn +be the Borel σ-algebra on Rn +and define the measure Tn=n i=1Ton Rn +, i.e., Tnis the n-fold product of T.Wedefineτ=(τ1,...,τ n)as canonical map on Rn +, i.e., τ(ω) =(τ1(ω), . . . , τn(ω)) =(ω1,...,ω n)for any ω∈Rn +. Note that τ1,...,τ nare independent and identically distributed under Tnby definition and the law of τiis given by T. We define a common measurable space by setting: (i) := C(R+,Rn)×Rn +, (ii) F:= B(C(R+,Rn))⊗BRn +, i.e., the smallest σ-algebra containing all sets of the form A×Bfor A∈B(C(R+,Rn)),B∈BRn +. 123 316 Mathematics and Financial Economics (2024) 18:313–331 We extend the definitions of Xand τto by setting X(ω1,ω 2):= X(ω1)and τ(ω1,ω 2):= τ(ω2)for (ω1,ω 2)∈. We define on (, F)the filtration (Ft)t≥0by Ft=σ(τ1,...,τ n)∨σ(Xs:0≤s≤t), t≥0. The filtration (Ft)t≥0describes the overall information flow including the information about the values of the time horizons. Moreover, for Player i, we introduce the filtration (Fi t)t≥0 describing the private information of Player i.WeassumethatσXi s:0≤s≤t⊆Fi t⊆Ft for all t≥0. Let Aidenote the set of all (Fi t)t≥0-progressively measurable α:×R+→[σ1,σ 2]. We refer to elements of Aias admissible controls or strategies and write Anfor the product A1×... ×An. We characterize the law of the state processes by means of martingale problems, introduced by Stroock and Varadhan. We refer to the monograph [18] for more details on martingale problems. Definition 2.2 Let α=(α1,...,αn)∈An. Then, a probability measure Pαon (, F)is called a feasible state distribution if (i) Pα◦X−1 0=δ0, (ii) Pα(C(R+,Rn)×B)=Tn(B)for all B∈BRn +, (iii) for all f∈C2 c(Rn,R), i.e., for all twice continuously differentiable f:Rn→Rwith compact support, the process Mf,definedby Mf s:= f(Xs)−f(0)−1 2s 0 n  j=1αj r2 ∂jj f(Xr)dr,s≥0,(1) is an (Ft)t≥0-martingale under Pα. We denote by Q(α) the set of all feasible state distributions Pα. Remark 2.3 The assumptions on the tuple αin the previous definition do not exclude that Q(α) is empty. For a tuple αto be a Nash equilibrium it is necessary, however, that Q(α) =∅ (see Definition 2.8 below). Remark 2.4 Note that condition (ii) implies that each random time τihas distribution T. Moreover, condition (iii) yields that each state process Xiis a local (Ft)t≥0-martingale and Xi,Xjt=t 0(αi s)2ds,if i=j, 0,else, t≥0. The state processes are even true martingales that are square integrable, i.e., EPα(Xi t)2< ∞,t≥0, because the controls are bounded and thus, EPαXi,Xit<∞for all t≥0 (see, e.g., [17], Section II.6, Corollary 3, p.73). Remark 2.5 The definition of the filtrations (Fi t)t≥0is rather general. Each filtration (Fi t)t≥0 can contain information about the other players’ states and random times. Therefore, the game setting covers the cases where players can or cannot observe each other, and have knowledge about the random time horizons. If, e.g., Fi t=Ft, each player can observe the state processes of the opponents and make her strategy depend on the opponents’ state trajectories. Moreover, each player has prior knowledge of the random time horizons. If, however, Fi t=σ(Xi s:0≤s≤t), then each player can only observe her own state process. Neither information about the opponents’ states nor about the random times is available. 123 Mathematics and Financial Economics (2024) 18:313–331 317 Despite this quite general information structure, we show in Theorem 4.1 that an approximate Nash equilibrium is given by a tuple of threshold controls depending only on the position of the single player’s state process. Remark 2.6 The state processes can be equivalently described as weak solutions to an ndimensional SDE (or via a stochastic integral w.r.t. some Brownian motion). Indeed, for some control α∈Anwith Q(α) =∅and P∈Q(α), there exists an n-dimensional Brownian motion Won (, F,(Ft)t≥0,P)such that the state processes X1,...,Xnsatisfy P-a.s. Xi t=t 0 αi sdWi s,t≥0,i=1,...,n,(2) and (, F,(Ft)t≥0,P,X,W)is a weak solution to (2) (see, e.g., [11], Proposition 5.4.6). We use this connection in the mean field game presented in Sect.3, and thus, characterize the state processes via stochastic integrals. Remark 2.7 For all measurable feedback functions a:R+×Rn→[σ1,σ 2]n, there exists a solution Pato the martingale problem (1) with (αs)s≥0=a(s,X1 s,...,Xn s)s≥0.This follows, e.g., from [13], Theorem 2.6.1, and [11], Proposition 5.4.11. If the filtration (Fi t)t≥0 contains the information about all states, i.e. if σ(Xs:0≤s≤t)⊂Fi t, then the control a(s,X1 s,...,Xn s)s≥0is contained in Aiand {Pa}⊆Q(α). This particularly holds for control tuples where each entry is a threshold control, i.e., each entry is given by the feedback function mb(x)=σ2,if x≤b, σ1,if x>b,(3) where b∈R. In this case, the solution to the martingale problem (1) is even unique: Remark 3.1 below implies that the SDE (7) with m=mbhas a unique strong solution and [11], Corollary 5.4.9, then implies that uniqueness for the martingale problem (1) holds. We suppose that each player aims at maximizing the probability of her own state at her terminal random time to be greater than the empirical (1−p)-quantile of all states at the individual random times. More precisely, let μn=1 n n  i=1 δXi τi be the empirical distribution of the players’ states at the terminal times. We define the empirical (1−p)-quantile by q(μn,1−p)=inf{r∈R:μn((−∞,r])≥1−p}.Notethat Xi τi>q(μn,1−p)if and only if the state of Player iis among the best npplayers at the terminal times. It is standard to predict or explain the players’ behavior in terms of (approximate) Nash equilibria, which are here defined as follows. Definition 2.8 Let ε≥0. A tuple α=(α1,...,α n)∈Anwith Q(α) =∅is called ε-Nash equilibrium of the n-player game if for all i∈{1,...,n}and Pα∈Q(α),wehave Pα(Xi τi>q(μn,1−p)) +ε≥sup β∈Ai sup P∈Q(α−i,β) P(Xi τi>q(μn,1−p)), (4) where (α−i,β)=(α1,...,α i−1,β,α i+1,...,α n)and sup ∅=−∞. 123 318 Mathematics and Financial Economics (2024) 18:313–331 Note that for ε=0, the tuple αin Definition 2.8 is a Nash equilibrium in the usual sense. In the case ε>0, the tuple αis also called an approximate Nash equilibrium. We do not assume that the solutions to the martingale problem (1) are unique by limiting the set of controls. Hence, we require that (4) holds for all solutions to the martingale problem (1). We emphasize that for an arbitrary control tuple α∈An, uniqueness of solutions to the martingale problem (1) can fail. For example, if n≥3, then it was shown in [15] that there exists a diffusion coefficient, and hence, an operator such that the corresponding martingale problem does not have a unique solution. Equivalently, there exists a diffusion coefficient σ:Rn→Rn×nthat is uniformly elliptic and such that for the corresponding SDE uniqueness in law fails (see also [5], Example 1.24). Thus, one can interpret (4) as follows: Player ihas no incentive to change her strategy from αto β, no matter which distribution from Q(α−i,β) is chosen. We stress, however, that for the approximate Nash equilibria derived from the corresponding mean field game, uniqueness is always satisfied. We do not compute an exact Nash equilibrium for the n-player game. As in [1], we compute an approximate Nash equilibrium for large games by considering the mean field limit of the game. We show that a mean field equilibrium strategy is given by a threshold control. Remark 2.9 Let T>0andτ1=... =τn=T, i.e., T=δT.SetFi t=σ(Xs:0≤s≤t), t≥0, i=1,...,n, and consider only controls of the form αt=a(t,X1 t,...,Xn t)for some a:R+×Rn→[σ1,σ 2]. Then, the above model is equivalent to the game model presented in the article [1]. 3 Mean field game In this section we describe the mean field version of the game introduced in Sect.2. Let 0 <σ 1<σ 2and (, F,(Ft)t≥0,P)be a complete filtered probability space satisfying the usual conditions. We suppose that (, F,(Ft)t≥0,P)supports an (Ft)t≥0-Brownian motion (Wt)t≥0and an R+-valued random variable τ. We assume that (Wt)t≥0and τare independent and Eτ−1 2<∞,(5) see Assumption 2.1 above. Let ˜ Abe the set of all processes α:×R+→[σ1,σ 2]that are (Ft)t≥0-progressively measurable. Given an agent chooses the control α∈˜ A,thestate process is defined by Xα t:= t 0 αsdWs,t≥0.(6) Remark 3.1 All feedback controls with a feedback function m:R→[σ1,σ 2]of bounded variation are contained in ˜ A. Indeed, the SDE dXt=m(Xt)dB t,X0=0,(7) has a weak solution because of Theorem 2.6.1 in [13], and pathwise uniqueness applies according to results in [16]. Hence, there exists a unique strong solution Xmto (7)(cf. Section 5.3 in [11]) and the control (m(Xm t))t≥0is contained in ˜ A. Let p∈(0,1)and denote by q(μ, 1−p)the (1−p)-quantile of some probability measure μ∈P(R), i.e., for a Borel probability measures μon R,q(μ, 1−p)=inf{r∈ 123 Mathematics and Financial Economics (2024) 18:313–331 319 R:μ((−∞,r])≥1−p}.Ifμ=Law(Xα τ), i.e., μis the law of Xα τfor a control α∈˜ A,we write q(Xα τ,1−p)for q(μ, 1−p). In the mean field game that corresponds to the n-player game described in Sect.2, a single player wants to maximize the probability of being larger than the population (1−p)-quantile over all admissible controls. In the mean field game, we define equilibria in the following sense: Definition 3.2 A tuple (μ∗,α∗)∈P(R)ט Ais called mean field equilibrium if (i) for all α∈˜ A, PXα∗ τ>q(μ∗,1−p)≥P(Xα τ>q(μ∗,1−p)), (ii) it holds μ∗=Law(Xα∗ τ). Lemma 3.3 Let b ∈R.Then P(Xmb τ>b)=max α∈˜ A P(Xα τ>b), (8) where mbis defined in Eq.(3). Proof The result follows from the diffusion control problem studied by McNamara [14]. In more detail, let α∈˜ A. Assume, without loss of generality, that there exists a regular conditional probability Q :R+×F→[0,1]for Fgiven τ. Because τand Ware independent, Wis also a Brownian motion under Q(t,·)for Pτ-a.e. t∈R+. Hence, we see that Q(t,{Xα τ>b})=Q(t,{Xα t>b})≤Q(t,{Xmb t>b}), for Pτ-a.e. t∈R+, using either [14], Remark 8, or [19], Proposition C.5. Finally, P(Xα τ>b)=∞ 0 Q(t,Xα t>b)Pτ(dt)≤∞ 0 Q(t,Xmb t>b)Pτ(dt)=P(Xmb τ>b). We refer to [11], Section 5.3.C, or [10], Section 1.3, for more details on regular conditional probabilities.  Lemma 3.3 implies that the optimal control for (8) is the threshold control mb.Inthe following, we just write Xbfor the state Xmb. The process Xbis a so-called oscillating Brownian motion (OBM) with threshold b, introduced in [12]. In more detail, OBM is defined as follows. Definition 3.4 Let b∈R. We call the solution Xbof the SDE dXt=mb(Xt)dB t,X0=0,(9) oscillating Brownian motion (OBM) with threshold band initial value 0. There exists indeed a unique strong solution to the SDE (9) because of Remark 3.1. For OBMs, one can explicitly calculate the probability density function and the cumulative distribution function. For the reader’s convenience, we recall the following result on OBMs. Proposition 3.5 Let b ∈Rand Xbbe an OBM with threshold b and initial value 0. Then, for all t >0, the random variable Xb thas a probability density function p(t,−b,·−b)with 123 320 Mathematics and Financial Economics (2024) 18:313–331 respect to the Lebesgue measure, where p(t,x,y)= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ 2σ1 σ2(σ1+σ2) 1 √2πte−(x σ1−y σ2)21 2t,if x ≥0,y<0, 2σ2 σ1(σ1+σ2) 1 √2πte−(y σ1−x σ2)21 2t,if x <0,y≥0, 1 σ1√2πte−(y−x)2 2σ2 1t+σ2−σ1 σ1+σ2e−(y+x)2 2σ2 1t,if x ≥0,y≥0, 1 σ2√2πte−(y−x)2 2σ2 2t+σ1−σ2 σ1+σ2e−(y+x)2 2σ2 2t,if x <0,y<0, for all x,y∈R. Moreover, the cumulative distribution function of Xb tis given by Fb t(x)= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ x σ2√t−σ2−σ1 σ1+σ2x−2b σ2√t,if x <b,b≥0, 2σ2 σ1+σ2x−b1−σ1 σ2 σ1√t−σ2−σ1 σ1+σ2,if x ≥b,b≥0, 2σ1 σ1+σ2x−b1−σ2 σ1 σ2√t,if x <b,b<0, x σ1√t−σ2−σ1 σ1+σ22b−x σ1√t,if x ≥b,b<0. (10) The proof of Proposition 3.5 follows either from [12], Theorem 1, or [19], Proposition B.2 and Proposition B.4. For more details on OBMs, we refer to [12]and[19], Appendix B. Using the independence of the OBM Xband the random time τ, one can derive from Proposition 3.5 the cumulative distribution function of Xb τ. Lemma 3.6 Let b ∈Rand Xbbe an OBM with threshold b and initial value 0. Then the cumulative distribution function of the random variable Xb τ, denoted by Fb τ, is given by Fb τ(x)=∞ 0 Fb t(x)Pτ(dt), x∈R. We refer to Proposition B.8 in [19] for the proof of Lemma 3.6. One can show that the functions Fb tand Fb τare Lipschitz continuous in the state variable as well as in the threshold b(see Appendix B in [19] for more details). The standard approach to solve mean field games is to consider mappings from probability distributions to the distributions of optimally controlled states and find their fixed points, the so-called equilibrium measures (see, e.g., [3]and[4]). However, Lemma 3.3 allows to study the distributions of OBMs only, which can be parameterized by the real-valued threshold b∈R. For identifying equilibria, it suffices to show that the function f:R→R,b→ q(Xb τ,1−p), has a unique fixed point. Indeed, if f(b)=b,thenb=q(Xb τ,1−p). Lemma 3.3 further implies that P(Xb τ>q(Xb τ,1−p)) =maxβ∈MP(Xβ τ>q(Xb τ,1−p)); hence, (Law(Xb τ), mb)is an equilibrium. The main result of this section is the following: Theorem 3.7 There exists a unique b∗∈Rsuch that q Xb∗ τ,1−p=b∗. The tuple (Law(Xb∗ τ), mb∗)is a mean field equilibrium. 123 Mathematics and Financial Economics (2024) 18:313–331 327 Notice that (31) implies that the threshold level b∗of Theorem 3.7 is positive. We denote the mean field equilibrium distribution by μ∗. Recall that μ∗is the law of the OBM with threshold b∗at an independent random time with distribution T. Now suppose that nis large and that in the game with nplayers everyone controls their states with the threshold control mb∗. We select one player and assume that her realized time horizon is t. Then the probability for this particular player to be among the best pat the end of the game is approximately given by w(t):= PXb∗ t>qμ∗,1−p=1−Fb∗ t(b∗)=2σ2 σ1+σ21−b∗ σ2√t. We refer to the function was the winning probability. Note that the winning probability is continuous and increasing in t. Moreover, we have lim t↓0w(t)=0. Thus, if the actual time horizon tis small, then the winning probability is close to zero. This is plausible, since the state process starts in zero and is stopped early, and hence attains the positive level b∗with a small probability only. Next observe that lim t→∞w(t)=σ2 σ1+σ2 . Thus, the winning probability is bounded by σ2 σ1+σ2, and the bound is almost attained for large time horizons t. The bound corresponds to the expected average time that an OBM is spending above the threshold in the long run. Indeed, irrespective of the threshold b, one can show that lim t→∞ P(Xb t≥b)=σ2 σ1+σ2 .(32) The bound allows also for a control theoretical interpretation: to this end consider the ergodic control problem with target functional J(α) := lim inf T→∞ 1 TET 0 1{Xα t≥b}dt,α∈˜ A, where ˜ Ais defined as in Sect.3and Xαis defined as in (6). One can show that mbis an optimal control (see Remark 8 in [14]) and hence, using (32), supα∈˜ AJ(α) =σ2 σ1+σ2. 6 An application to online competitions Our game formulation is generic and can correspond to a variety of practical situations. For example, the game applies to managers of mutual funds striving for their funds to be among the best performing. This application is described, for homogeneous managers, in detail in Section 7 of [1]. We here provide an alternative application to online competitions in which, usually, teams have to collaborate in order to solve a problem and provide a solution within a limited time frame. Hackatons are examples of such competitions, but also data science and machine learning related competitions offered on some platforms such as Kaggle, DrivenData, AICrowd etc. During these competitions, a score, based on a given evaluation metric, can be calculated 123 328 Mathematics and Financial Economics (2024) 18:313–331 by each participant and a "public leaderboard" displays the relative ranks during the whole length of the competition. In this context, the private state Xi tof player iis interpreted as her score displayed at time t. The number of teams involved in online competitions can reach several tens of thousands, enough to consider the mean field approximation. Level of risk. We interpret the diffusion control as the possibility to control the level of risk taken. In the online contest’s setting, teams can indeed choose to try and use well established methods, whose robustness is already studied and for which errors can be more easily and quickly corrected. We interpret this as low risk and diffusion coefficient σ1. On the other hand, the choice of diffusion coefficient σ2is interpreted as trying new techniques, for which there is less or no experience. Our results rigorously show in this context that if subtasks are going well, players will play safe, whereas if a given team is poorly performing on the evaluation metric, it has an incentive to play risky and try less common strategies. Observability. Participants to online competitions can submit a solution, and in that case their current score is calculated and displayed to all participants. However, if a team obtains a solution and tests it offline on the provided data set, then the associated score is not visible during the time it is not submitted on the platform. A given team can choose to reveal its solution and score only towards the end of the submission period. Moreover, it is possible to design online competitions with only partial observability: one could easily imagine that teams only observe the best score, or a given quantile of the scores distribution, to assess their relative performance. Our results show that if the number of players is large, observability does not matter, at least for the particular type of discontinuous criteria that we consider, which are common in these competitions, where a fixed cash prize is offered to the best performing team. This result also holds for continuous functions of the rank satisfying the symmetry condition given in Assumption 7.1. Terminal time. We consider two cases. Firstly, the case where the time horizons of all players are constant equal to T∈(0,∞), interpreted as the date at which a final assessment of the evaluation metric is made by the contest organizers. Secondly, the time horizon τiof team iquantifies the resources it can put into the competition, e.g. the number of working hours. For example, τi can be set proportional to the deterministic assessment date and the number of team members. Section5reveals that larger teams have an advantage compared to smaller teams. 7 Extensions In this section, we discuss more general reward criteria for the n-player game. In particular, we consider rewards at the random time horizons that are given by measurable functions g:R×R→Rof the state and the population quantile, instead of the “all-or-nothing” payoff given by the function (x,q)→ 1(q,∞)(x)before. First, we show that the mean field equilibrium control of Sect.3is also an equilibrium for the reward functions g,ifg satisfies a symmetry and convexity condition. Then, we prove that this equilibrium provides an approximate Nash equilibrium of the n-player game. 123 Mathematics and Financial Economics (2024) 18:313–331 329 7.1 Mean field game Assume that we are in the setting of Sect.3. The whole analysis in Sect. 3depends on the optimality of the threshold control mbfor the particular choice of the reward 1(b,∞)(Lemma 3.3). Results of McNamara [14] imply that mbis not only optimal for this reward but also for more general reward functions that are continuous, have exponential growth, and satisfy a convexity condition and a symmetry condition. In more detail, we can generalize Theorem 3.7 to measurable functions g:R×R→Rsatisfying: Assumption 7.1 (i) g(·,q)is continuous and has exponential growth for any q∈R, (ii) g(·,q)is convex on (−∞,q]and concave on [q,∞)for any q∈R, (iii) for all x≥0andq∈Rit holds σ2g(σ1x+q,q)+σ1g(−σ2x+q,q)=(σ1+σ2)g(q,q). Proposition 7.2 Let b∗be given by Theorem 3.7 and g satisfy Assumption 7.1. Then, (Law(Xb∗ τ), mb∗)is also a mean field equilibrium for the reward function g, i.e., EgXb∗ τ,qXb∗ τ,1−p=sup α∈˜ A EgXα τ,qXb∗ τ,1−p. Proof As in Lemma 3.3, one can show for fixed b∈Rthat EgXb τ,b=sup α∈˜ A EgXα τ,b, using either [14], Theorem 6, or [19], Theorem C.5. The assumptions on gguarantee that these theorems apply. Moreover, Theorem 3.7 implies the existence of a unique fixed point b∗of the map b→ qXb τ,1−p. For this fixed point b∗,weseethat EgXb∗ τ,qXb∗ τ,1−p=sup α∈˜ A EgXα τ,qXb∗ τ,1−p, i.e., (Law(Xb∗ τ), mb∗)is an equilibrium for the reward function g. 7.2 Approximate Nash equilibrium in the n-player game Now, in the setting of Sect.2, we show that the n-tuple with each entry equal to the mean field equilibrium strategy provides an approximate Nash equilibrium of the n-player game with reward g. Definition 7.3 Let ε>0. A tuple α=(α1,...,α n)∈Anwith Q(α) =∅is called ε-Nash equilibrium of the n-player game if for all i∈{1,...,n}and Pα∈Q(α) EPαgXi τi,q(μn,1−p)+ε≥sup β∈Ai sup P∈Q(α−i,β) EPgXi τi,q(μn,1−p), where (α−i,β)=(α1,...,α i−1,β,α i+1,...,α n)and sup ∅=−∞. With some additional assumptions on the terminal reward g,wecanshow: Proposition 7.4 Let g satisfy Assumption 7.1. In addition, assume that g is uniformly bounded, continuous, and g(x,·)is monotonically decreasing for all x ∈R.Letα∗∈Anbe defined as in Theorem 4.1. Then, there exists a sequence εn≥0with limn→∞ εn=0such that the control tuple α∗=(α1,∗,...,αn,∗)is an εn-Nash equilibrium of the n-player game. 123 330 Mathematics and Financial Economics (2024) 18:313–331 Proof All details of the proof can be found in Proposition 3.2.25 in [19]. Let i∈{1,...,n}. Moreover, let β∈Aisuch that Q(α−i,∗,β)=∅and choose P∈Q(α∗),˜ P∈Q(α−i,∗,β). Using the empirical quantiles A(n)and D(n)defined in (16)and(18), respectively, we find that ˜ Eg(Xi τi,q(μn,1−p))≤˜ Eg(Xi τi,D(n)), Eg(Xi τi,q(μn,1−p))≥Eg(Xi τi,A(n)), because of the monotonicity of g(x,·). We use the notation Eand ˜ Efor the expectation w.r.t. Pand ˜ P, respectively. Define the function G(q):= E[g(Xi τi,q)]=∞ 0∞ −∞ g(x,q)p(t,−q,x−q)dxP τ(dt), q∈R, where pdenotes the probability density function of the OBM defined in Proposition 3.5. Note that Eg(Xi τi,A(n))=E[G(A(n))], because X1 τ1,...,Xn τnare independent under P. Similar to S, one can show that ˜ Eg(Xi τi,D(n))≤˜ E[G(D(n))]=E[G(D(n))], similar to Step 2 in the proof of Theorem 3.2.16 in [19]. The upper bound follows from Theorem 6 in [14]: the control maximizing the left-hand side is the threshold control with threshold D(n), if conditioned on D(n). We conclude that ˜ Eg(Xi τi,q(μn,1−p))−Eg(Xi τi,q(μn,1−p))≤E[G(D(n))]−E[G(A(n))]. Note that Gis bounded and continuous, and the right-hand side only depends on the distribution of n−1 independent OBMs with threshold b∗. Moreover, A(n)and D(n)converge to b∗in probability (see Lemma 3.2.20 in [19]) and hence, also in distribution. This means lim n→∞|E[G(D(n))]−E[G(A(n))]|=0. Therefore, we can find a sequence (εn)n∈Nwith the desired properties. The sequence (εn)n∈N is independent of i∈{1,...,n},β∈A,P∈Q(α∗),and ˜ P∈Q(α−i,∗,β) because the distributions of A(n)under Pand of D(n)under ˜ Pare unique.  Acknowledgements We thank two anonymous referees for carefully reading the article and for providing many suggestions for improvements. Support from the German Research Foundation through the project AN 1024/5-1 is gratefully acknowledged. Nabil Kazi-Tani’s research is supported by the ANR project DREAMES ANR-21-CE46-0002-03. Funding Open Access funding enabled and organized by Projekt DEAL. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. 123 Mathematics and Financial Economics (2024) 18:313–331 331 References 1. Ankirchner, S., Kazi-Tani, N., Wendt, J., Zhou, C.: Large ranking games with diffusion control. Math. Oper. Res. (2023) (in press) 2. Barrasso, A., Touzi, N.: Controlled diffusion mean field games with common noise and Mckean–Vlasov second order backward sdes. Theory Probab. Appl. 66(4), 613–639 (2022) 3. Carmona, R., Delarue, F.: Probabilistic theory of mean field games with applications. I, volume 83 of Probability Theory and Stochastic Modelling. Springer, Cham. Mean field FBSDEs, control, and games (2018) 4. Carmona, R., Delarue, F.: Probabilistic theory of mean field games with applications. II, volume 84 of Probability Theory and Stochastic Modelling. Springer, Cham. Mean field games with common noise and master equations (2018) 5. Cherny, A., Engelbert, H.-J.: Singular stochastic differential equations. Lecture Notes in Mathematics, vol. 1858. Springer, Berlin (2005) 6. Djete, M.F.: Extended mean field control problem: a propagation of chaos result. Electron. J. Probab. 27, 1–53 (2022) 7. Djete, M.F.: Mean field games of controls: on the convergence of nash equilibria. Ann. Appl. Probab. 33(4), 2824–2862 (2023) 8. Elie, R., Hubert, E., Mastrolia, T., Possamaï, D.: Mean-field moral hazard for optimal energy demand response management. Math. Financ. 31(1), 399–473 (2021) 9. Folland, G.B.: Fourier Analysis and Its Applications. Wadsworth & Brooks/Cole Advanced Books & Software, Pacific Grove (1992) 10. Ikeda, N., Watanabe, S.: Stochastic differential equations and diffusion processes, volume 24 of NorthHolland Mathematical Library, 2 edn. North-Holland Publishing Co., Amsterdam; Kodansha, Ltd., Tokyo (1989) 11. Karatzas, I., Shreve, S. E.: Brownian motion and stochastic calculus, volume 113 of Graduate Texts in Mathematics, 2nd edn. Springer, New York (1991) 12. Keilson, J., Wellner, J.A.: Oscillating Brownian motion. J. Appl. Probab. 15(2), 300–310 (1978) 13. Krylov, N.V.: Controlled diffusion processes, volume 14 of Stochastic Modelling and Applied Probability. Springer, Berlin. Translated from the 1977 Russian original by A. B. Aries, Reprint of the (1980) edition (2009) 14. McNamara, J.M.: Optimal control of the diffusion coefficient of a simple diffusion process. Math. Oper. Res. 8(3), 373–380 (1983) 15. Nadirashvili, N.: Nonuniqueness in the martingale problem and the Dirichlet problem for uniformly elliptic operators. Annali della Scuola Normale Superiore di Pisa. Classe di Scienze. Serie IV 24(3), 537–549 (1997) 16. Nakao, S.: On the pathwise uniqueness of solutions of one-dimensional stochastic differential equations. Osaka J. Math. 9(3), 513–518 (1972) 17. Protter, P.E.: Stochastic integration and differential equations, volume 21 of Stochastic Modelling and Applied Probability, 2 edn. Springer, Berlin, 2005. Version 2.1, Corrected third printing (2005) 18. Stroock, D.W., Varadhan, S.R.S.: Multidimensional diffusion processes. In: Classics in Mathematics. Springer, Berlin. Reprint of the (1997) edition (2006) 19. Wendt, J.: Diffusion control and games. Ph.d thesis, Friedrich-Schiller-Universität Jena (2023) Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. 123