scieee AI-readable full text Open interactive document viewer

Generalized hyperbolic discounting in security games of timing

Merlevede, Jonathan,Johnson, Benjamin,Grossklags, Jens,Holvoet, Tom

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Merlevede, Jonathan; Johnson, Benjamin; Grossklags, Jens; Holvoet, Tom Article Generalized hyperbolic discounting in security games of timing Games Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Merlevede, Jonathan; Johnson, Benjamin; Grossklags, Jens; Holvoet, Tom (2023) : Generalized hyperbolic discounting in security games of timing, Games, ISSN 2073-4336, MDPI, Basel, Vol. 14, Iss. 6, pp. 1-52, https://doi.org/10.3390/g14060074 This Version is available at: https://hdl.handle.net/10419/330067 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Citation: Merlevede, J.; Johnson, B.; Grossklags, J.; Holvoet, T. Generalized Hyperbolic Discounting in Security Games of Timing. Games 2023,14, 74. https://doi.org/ 10.3390/g14060074 Academic Editors: Michael Wooldridge and Ulrich Berger Received: 30 June 2023 Revised: 12 November 2023 Accepted: 26 November 2023 Published: 30 November 2023 Copyright: © 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). games Article Generalized Hyperbolic Discounting in Security Games of Timing Jonathan Merlevede 1, Benjamin Johnson 2, Jens Grossklags 2,* and Tom Holvoet 1 1imec-DistriNet, Department of Computer Science, KU Leuven, 3001 Heverlee, Belgium 2 School of Computation, Information and Technology, Technical University of Munich, 85748 Garching, Germany *Correspondence: jens.gr[email protected] Abstract: In recent years, several high-profile incidents have spurred research into games of timing. A framework emanating from the FlipIt model features two covert agents competing to control a single contested resource. In its basic form, the resource exists forever while generating value at a constant rate. As this research area evolves, attempts to introduce more economically realistic models have led to the application of various forms of economic discounting to the contested resource. This paper investigates the application of a two-parameter economic discounting method, called generalized hyperbolic discounting, and characterizes the game’s Nash equilibrium conditions. We prove that for agents discounting such that accumulated value generated by the resource diverges, equilibrium conditions are identical to those of non-discounting agents. The methodology presented in this paper generalizes the findings of several other studies and may be of independent interest when applying economic discounting to other models. Keywords: game theory; games of timing; discounting 1. Introduction Game theory’s great promise is to predict the evolution of a multi-agent interaction from agents’ preferences. Unfortunately, the particulars of scenarios tend to complicate the derivation of these predictions, and accurately describing a real scenario requires progressively more advanced models of agent behavior. This is markedly true for interactions that persist over time. Research on games of timing dates back to the Cold War period, and various scenarios have by now motivated a sizable body of literature (see, for example, [ 1 ]). One effort is a series of papers centered around the so-called FlipIt model [ 2 ], which attempts to capture the dynamics of a persistent two-agent interaction involving a single contested resource. This resource is an abstract representation of, e.g., knowledge of cryptographic keys or passwords or ownership of cloud resources [3]. FlipIt models capture scenarios taken in part from a series of high-profile (cyber) attacks against supposedly well-protected targets—the Iranian industrial control systems [ 4 ], the U.S. Office of Personnel Management [ 5 , 6 ], large telecommunication providers [ 7 , 8 ], major health care insurance companies [ 9 ], and, recently, critical infrastructure companies in Ukraine [ 10 ]. An essential property shared by these threats is that against them, perfectly effective preventative investments are nearly impossible [ 11 , 12 ]. Mitigation strategies such as in-depth security audits or longer-term investments to reorganize existing IT infrastructures are therefore weighed against the economic incentives of would-be attackers. While this line of work has explored diverse facets of security decision-making, a limitation in most studies on games of timing in general and on FlipIt-like games in particular is that they do not consider economic discounting in the valuation of the contested resource—so although the environment changes over time, the value of the resource and the costs of attacking or defending it do not. A few more recent efforts have attempted to Games 2023,14, 74. https://doi.org/10.3390/g14060074 https://www.mdpi.com/journal/games Games 2023,14, 74 2 of 52 bridge this deficiency by considering the effect of exponential discounting on the valuation and defense costs over time [ 13 – 15 ]. Such studies have found new predictions of optimal defense investments due mainly to the fact that an exponentially discounted investment has a finite cumulative valuation over time. Although mathematically appealing, exponential discounting is known to have low descriptive accuracy in most contexts. It remains to consider a more comprehensive approach to time-based discounting. Our work meets this need by analyzing a FlipIt-like contested resource scenario under a more generic present-focused discounting framework using a two-parameter family of generalized hyperbolic discounting functions. These well-studied functions, introduced by Loewenstein and Prelec in 1992 [ 16 ], smoothly interpolate an agent’s temporal valuation between an arbitrary rate of exponential discounting and a constant valuation, nesting exponential and hyperbolic discounting as special cases. As we cover the entire space of generalized hyperbolic discounting functions, our conclusions subsume those of earlier research efforts involving persistent resource control—including some that do not involve discounting and work on exponential discounting [13,14]. Our primary technical contribution is to show that for a wide range of agent discounting preferences, the agent’s strategic considerations and equilibrium conditions are the same as if they were not discounting. We also provide a complete characterization of player utilities and (Nash) equilibrium conditions for the FlipIt model in which both agents are interacting using the same strategic class out of periodic or exponential (defined in Section 4.5) and when both agents are applying a form of discounting with parameters from the same class out of sub-hyperbolic or super-hyperbolic (defined in Section 3). We show that generalized hyperbolic discounting unlocks the possibility of equilibria in which neither player moves and prove that periodic strategies no longer always strictly dominate the class of so-called renewal strategies, as they do for the model without discounting, a result first stated by van Dijk et al. [2] . Beyond these new discoveries, and because some of the derivations developed for this work can be used to extend results from less generic models, we believe that our framework would likely be helpful in other timing-based games besides FlipIt. We organized the rest of the paper as follows. In Section 2, we discuss related work. Then, we describe our model, which builds on FlipIt, in Section 4. We present our analysis of the model in Section 5, and we discuss some interesting aspects of our findings in Section 6. Finally, we conclude in Section 7. Following the main document are appendices containing formal proofs and derivations to support the claims made in Sections 5and 6. 2. Related Work Many interactions, including security interactions and especially those dealing with Advanced Persistent Threats (APTs), have an important temporal dimension. Studying such interactions necessitates comparing the value of present and future potential gains or losses. It is here that discounting models come into play, offering precise explanations for the valuations of economic resources, including privacy and security, by individuals, institutions, or societies over time. Within the field of economics, the exponential discounted-utility model pioneered by Samuelson [17] and Ramsey [18] has become the predominant approach to discounting. The exponential model posits that the rate at which assets depreciate remains constant, allowing a single discount factor, β , to encapsulate all individuals’ disparate perceptions and behaviors at different times. Specifically, it discounts utility flows with the function e −βt , where t is the horizon of the utility flow and β is a free parameter representing the discounting rate. The exponential model is popular because it is relatively easy to understand and reason about and mathematically tractable. There are also solid theoretical arguments for exponential discounting. Strotz [19] was the first one to derive exponential discounting as the only approach to intertemporal choice available to decision-makers whose choices are always consistent, in the sense that they are never in conflict with themselves at later points Games 2023,14, 74 3 of 52 in time. This theoretical effort was later repeated by others in different settings [ 20 , 21 ]. The consistency property has established exponential discounting as the dominant model that is often considered to be normatively correct [22]. However, we know that the exponential discounting utility model is only very rarely descriptively accurate—it does not do well predicting the behaviors of agents in the real world. Numerous behavioral research works have uncovered facets of human and animal behavior that cannot be captured by the exponential discounting framework [ 23 ]. This includes present-focused preferences, preference reversals, self-control problems [ 16 , 24 ], effects of temptation [ 25 ], psychometric distortions such as subjective time [ 26 ] and probability perception [ 27 ], magnitude effects, and myopic decision-making on account of our limited cognitive faculties [ 28 ]. The search for more descriptive accuracy has resulted in many other discounting models, including present-biased or quasi-hyperbolic discounting [ 29 ] and the hyperbolic and generalized hyperbolic models discussed below. Ericson and Laibson [30] present an excellent overview of theories on intertemporal choice. Within the wealth of existing models, the model studied most by psychologists is the hyperbolic discounting function, first implied by Herrnstein [31] . It is a mathematically uncomplicated one-parameter model that has been shown to greatly outperform exponential discounting in terms of descriptive accuracy and incorporates present-focused preferences and preference reversals. Hyperbolic discounting scales utility with the function ( 1 +αt)−1 , where t is again the horizon of the utility flow, and where α is a free parameter related to the discounting rate. Hyperbolic discounting preserves the proportion between the ratio of future valuations and the corresponding future times—e.g., the relative drop in valuation between times 10 and 20 is (almost) the same as the relative drop in valuation between times 100 and 200. This implements a strictly decreasing discount rate; when compared with exponential discounting, hyperbolic discounting depreciates faster near t= 0 and slower for large t. The hyperbolic model is often extended with an additional parameter β to ( 1 +αt)−β/α . It is then called generalized hyperbolic discounting or hyperboloid discounting. Here, the ratio β/α reflects the nonlinear scaling of the amount and delay [ 32 , 33 ]. Parameter α can also be interpreted as how much the function departs from constant-rate (exponential) discounting [16], with lower values of αcorresponding to more (traditional) “rational” behavior. The generalized hyperbolic discounting function has many desirable properties making it an excellent choice for study. For one, it embeds several other major models: classical exponential discounting, hyperbolic discounting, and constant or no discounting, for α→∞ , α=β , and α→ 0, respectively. Generalized hyperbolic discounting can be shown to have quantitatively higher predictive accuracy beyond that which can be accounted for simply by the addition of an additional parameter to the hyperbolic model [ 33 , 34 ]. The additional parameter allows the modeling of very different agent types. For example, many studies have shown that humans, in many contexts, tend to exhibit behavior in accordance with αβ , while animals such as pigeons exhibit behavior where α≈β ; we refer to Vanderveldt et al. [23] for a review of these studies. Lastly, in addition to expressing delay discounting behavior of agents with present-focused preferences, generalized hyperbolic discounting can be rationalized as a result of risk and probability discounting [ 35 ]. The underlying idea is that the value of a future reward should be discounted because there is a risk that rewards will not be realized. Generalized hyperbolic discounting can arise both when the hazard rate is horizon-dependent and when it is uncertain [36–38]. Our work builds on these behavioral and economic insights to advance the literature on games of timing [1] through the adoption of generalized hyperbolic discounting in the framework of a FlipIt-like game [ 2 , 39 ]. As such, our work contributes also to the overall space of security economics and, in particular, the application of game theory to security and privacy challenges [40,41]. FlipIt is a game motivated by persistent, stealthy, sophisticated attacks by an advanced persistent threat (APT) and models a situation where two players compete for the ownership of a resource generating value over time [ 2 , 39 ]. Apart from the inclusion of discounting, the Games 2023,14, 74 4 of 52 interaction we investigate is exactly that of FlipIt described by van Dijk et al. [2] , although our mathematical description of it is significantly different and more streamlined. The interesting strategic aspects of FlipIt have led to numerous follow-up studies. Most closely related to our work are studies on the impact of exponential discounting in the framework of the FlipIt game [ 13 – 15 ]. These have found new predictions of optimal defense investments due largely to the fact that an exponentially discounted investment has a finite cumulative valuation over time. Throughout the paper, we will discuss the relationship of model variations with generalized hyperbolic discounting, exponential discounting, and no discounting in detail. To the best of our knowledge, there is no comprehensive review article about the FlipIt game. Most of the literature has centered on the study of variations of the game’s structure, such as the types of involved players [ 42 , 43 ], the speed and efficacy of moves [ 42 – 45 ], the game’s time horizon [ 45 , 46 ], the structure of the resource and partial control [ 47 – 49 ], variations on the assumption of perfect stealthiness [ 42 – 44 , 46 ], discretization [ 45 ], the effects of budget constraints [ 46 ], and more sophisticated strategies involving machine learning [ 50 ]. FlipIt can also be embedded as a stage into other games [ 51 ]. More recently, Banik and Bopardikar [52] have formulated a variation of FlipIt as a zero-sum discrete control game, where the defender aims to stabilize a dynamical system. Miura et al. [53] have formulated a FlipIt-like game, augmented by an epidemic model, where defender and attacker compete to control as many computing resources as possible over a set period of time. Although the concept of time forms the very core of the FlipIt game, van Dijk et al. [2] and most follow-up work consider many aspects of the interaction to remain unchanged as time progresses. This includes the value of the resource, as well as the costs to attack it. This presents some conceptual challenges, as nothing is everlasting, and also results in some questionable predictions, such as every resource being attacked by attackers at non-zero rates. In [ 13 , 14 ], Merlevede et al. introduced exponential discounting of future gains and costs, addressing some of these concerns. In [ 15 ], these authors also introduce a novel class of “discounted strategies” to the exponentially discounted game in which players move less frequently as the resource loses value. We are unaware of further studies considering discounting in the context of the FlipIt game. 3. Discounting The term discounting implies that the future value of a measurable quantity is less than its present value. The effect of this value decrease can be formalized by a discount function D(t) . D(t) gives a multiplicative factor expressing the value at future time t> 0 relative to its present value (at time t=0). An intuitive measure of the behavior of the discount function is the discount rate: ρ(t) = −D0(t) D(t). The discount rate indicates how fast the value decreases at any time t≥0. Our model discounts player gains and player costs along a member of the family of generalized hyperbolic discounting functions, introduced by Loewenstein and Prelec [16]1 , which have the form (D(t) = 1 (1+αt)β/α|α>0, β>0). Games 2023,14, 74 5 of 52 Figure 1a illustrates the hyperbolic discount function, while Figure 1b displays the corresponding discount rates. For generalized hyperbolic discounting, discount rates are always strictly decreasing: ρ(t) = β 1+αt. 0246 8 10 0 1 3 1 t D(t) α=106 α=10 α=1/2 α=1/10 exp (a) Generalized hyperbolic discount function 0 2 46 8 10 0 0.2 0.4 0.6 0.8 1 t ρ(t) α=106 α=10 α=1/2 α=1/10 exp ln(3) 4 (b) Generalized hyperbolic discount rate Figure 1. Hyperbolic discount function and corresponding discount rates. Each blue curve corresponds to a different value of α . Parameter β is chosen so that D( 4 ) = 1/3 ( β is increasing in α ). The dashed curve is a true hyperbolic discounting curve with α=β=1/2 . The exponential function crossing the point (4, 1/3)is shown in red. Parameter α determines how much the function resembles the exponential function. For varying α , the family of generalized hyperbolic discounting functions spans a wide variety of discounting behaviors, including exponential discounting for α→ 0, true hyperbolic discounting for α=β, and no discounting for α→∞. lim α→0 1 (1+αt)β/α=e−βt(exponential) (1) 1 (1+αt)β/αα=β =1 1+βt(hyperbolic) (2) lim α→∞ 1 (1+αt)β/α=1 (no discounting). (3) Note that for exponential discounting ( D(t) = e−βt ), the discount rate is equal to a constant value (ρ(t) = β). The case of “true” hyperbolic discounting ( α=β ) corresponds to a boundary condition separating two qualitatively distinct classes of discounting functions. We will refer to generalized hyperbolic discounting with α<β as super-hyperbolic discounting, and to discounting with α≥β as sub-hyperbolic discounting. We include hyperbolic discounting ( α=β ) as a member of the class of sub-hyperbolic discounting functions. The following observations largely explain the behavioral differences between these two classes of functions. 1. For α<β (super-hyperbolic discounting), the area under D(t) ’s curve is finite and given by: Z+∞ τ=0D(τ)dτ=1 β−α. (4) 2. For α≥β (sub-hyperbolic discounting), RT τ=0D(τ)dτ does not converge for T→+∞ . Games 2023,14, 74 6 of 52 4. Model This section introduces our model for stealthy timing-based security games with generalized hyperbolic-discounted costs and resource valuations. We model the same interaction first presented in van Dijk et al. [2] but present a simplified, more streamlined mathematical characterization and include discounting. Our model subsumes models without discounting [ 2 ] and with exponential discounting [ 14 ] as special cases (Table 1). Although adding generalized hyperbolic discounting is a small technical change, it has a large conceptual impact as a model for risk, ephemerality, or bounded rationality. It also necessitates an entirely different mathematical approach to model analysis. Table 1. Discounting models used with FlipIt-like games with their discounting parameters and relation to the other models. Class of Discounting No Discounting [2] Exponential Discounting [14] Generalized Hyperbolic Discounting Discounting parameters ∅{βA,βc A,βD,βc D}2{αD , βD , αc D , βc D , αA , βA , αc A , βc A} Relation to other models n.a. lim λ→ 0 is None lim α→0 is Exponential 4.1. Overview In our two-player game, a defender ( D ) and an attacker ( A ) vie for control over a central resource. To obtain control, either player i∈ {A , D} can choose to pay a fixed instantaneous cost ci to ‘flip’ or immediately assume control of the resource. The last player to execute a move always controls the resource. The controlling player accrues utility at a rate that decreases over time along a generalized hyperbolic function. The cost to execute a move is also time-discounted along a (possibly different) generalized hyperbolic function. Control is stealthy in the sense that neither player knows who controls the resource until the moment that they initiate an instantaneous flip. The remainder of this section formalizes the game. 4.2. Player Strategies For player i∈ {D,A}, define ~ ti= (ti,0,ti,1,ti,2, . . .) to be a strictly increasing sequence of real times at which player i moves. 3 The length of ~ ti can be finite or infinite. A player strategy in this game is defined completely by a probability distribution over a set of possible~ ti. 4.3. Player Control The player control function indicates, for a given pair of move sequences ( ~ tD , ~ tA) , whether the defender or the attacker is deriving utility from the resource at any given moment in time. In the particular case where different players’ moves collide in time, we define the outcome as a no-op, meaning that resource ownership remains unchanged. Thus, removing any such collisions if necessary, we may assume without loss of generality that ~ tD∩~ tA=∅. Let ~ t=~ tD∪~ tA= (t0,t1,t2, . . .) be the strictly increasing sequence of player move times. Then, for any time t≥t0 , we define the latest flip time function by Games 2023,14, 74 7 of 52 LFT:t7→ max{tk∈~ t:tk≤t}. From time t= 0 until the time of the first flip t0 , the defender has control of the resource. The player control function can, therefore, be expressed as PC:t7→    Dif t<t0or LFT(t)∈~ tD Aif LFT(t)∈~ tA.(5) The asymmetry of the player control function shows that the defender has an advantage due to starting the game in control of the resource. We also define a player control indicator function PCi:t7→~ 1PC(t)=i, (6) which can be useful for integration. The player control indicator function tells us, for a given player i∈ {D , A} and time t> 0, whether that player controls the resource at that time. 4.4. Player Utilities Each player’s utility is defined to be the difference between the player’s gains and the player’s costs: ui=Gi−Ci. Both gains and costs are subject to generalized hyperbolic discounting. We define generalized hyperbolic discounting, player gains, and then player costs in the following subsections. 4.4.1. Gains Players achieve gains when in control of the resource. Gains initially accrue value at some rate of V units of value per unit of time. Gains decrease over time according to a player-dependent generalized hyperbolic discount function Di:     [0, +∞[→]0, 1] t7→ 1 (1+αit)βi/αi,(7) where αi and βi are parameters characterizing player i ’s (im)patience. We use reversed brackets to indicate open interval boundaries to avoid confusing intervals for ordered tuples, notably strategy profiles. The discounted gain rate of player i up to time t may be determined by computing the expected weighted integral of PCi(t) up to t , normalized with respect to the total (discounted) value of the resource up to t: Gi(t) = EhRt τ=0PCi(τ)·V·Di(τ)dτi Rt τ=0V·Di(τ)dτ. (8) The expectation is taken over possible game outcomes resulting from the player’s strategies, represented by stochastic process PCi . Normalization allows comparing player gains for different discount rates and interpreting gain as a fraction of total achievable gain. Games 2023,14, 74 8 of 52 The gain of player i is the maximum limit of her discounted gain rate as time t moves to infinity: Gi=lim sup t→∞ Gi(t). (9) Note that Gi(t)∈[ 0, 1 ] for every t , and so Gi∈[ 0, 1 ] as well. When the strategies are restricted to periodic or exponential, as described in Section 4.5, then limt→∞Gi(t) exists for each player i∈ {D , A} , so that we can replace lim sup with limit. If we would further assume that players have the same discounting factors so that (αD , βD) = (αA , βA) , then we would obtain GD+GA=1 by the linearity of expectation. 4.4.2. Costs When players perform a move, this comes at a fixed instantaneous cost of ci> 0. As with gains, we discount costs according to a player-dependent generalized hyperbolic discount function Dc i:     [0, +∞[→]0, 1] t7→ 1 (1+αc it)βc i/αc i ,(10) where αc i and βc i characterize i ’s (im)patience, this time with respect to costs. Note that αc i and βc i do not have to equal αi and βi , meaning that costs may be discounted at different rates from gains. This results in four discounting parameters per player and eight discounting parameters in total. The discounted spending rate of player i up to time t may be determined by computing the expected weighted sum of instantaneous costs of the moves made by i up to time t , normalized with respect to the total (discounted) value of the resource: Ci(t) = Eh∑τ∈ ~ ti,τ≤tciDc i(τ)i Rt τ=0V·Di(τ)dτ. (11) As with gains, the expectation is taken with respect to the distribution used to define ~ ti . Scaling gains and costs by the same factor ( Rt τ=0V·Di(τ)dτ ) makes the normalization operation neutral with respect to the behavior of utility-maximizing players. Finally, since we only deal with normalized costs and gains and since ci and V are both free parameters, we can, without loss of generality, assume that V=1. (12) This assumption allows us to consider the instantaneous cost ci=ci/V as a unitless value, expressing a fraction of the initial rate at which the resource accrues value per unit of time. A player’s (total) cost is defined as the limit of her spending rate as time t moves to infinity: Ci=lim t→∞Ci(t). (13) 4.5. Restricted Strategies The description of all possible player strategies given in Section 4.2 is the most general class of strategies to which we can define a definite outcome according to our discounted FlipIt model. However, to exhibit an effective strategy and facilitate analysis, it is necessary to introduce additional constraints that reduce the number of free parameters in a strategy specification. Each of the two strategy classes we consider in this paper—the class of exponential and the class of periodic strategies—has been studied extensively in prior work. Both classes are described by a single real parameter corresponding to the expected number of moves per unit of time, which we will refer to as flip rate or move rate throughout the Games 2023,14, 74 15 of 52 0 0.5 1 1.5 0 0.5 1 1.5 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 νD νA (a) Defender gain 0 0.5 1 1.5 0 0.5 1 1.5 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 νD νA (b) Attacker gain 0 0.5 1 1.5 0 0.5 1 1.5 νD νA (c) Defender utility 0 0.5 1 1.5 0 0.5 1 1.5 νD νA (d) Attacker utility Figure 3. Contour plots of gains and utilities for periodic play and super-hyperbolic discounting (with αD=αA= 0.5, βD=βA= 1, and cD=cA= 0.3). Warmer colors indicate higher values; the precise numerical values are not important. 5.4. Player Incentives and Best Responses Knowing players’ utility functions, our attention turns to how players behave as they try to optimize their utility. The first step involves player incentives. Definition 4 (Player incentive) . Player i ’s incentive is the partial derivative of her utility to her play rate, expressed in standard mathematical notation as ∂ui ∂νi . Player incentives are closely related to player best responses because the roots of a player’s incentive function point to the local optima of her utility function. A player’s incentive is generally a function of both her own play rate and her opponent’s play rate. Appendix C.1 states expressions for the player incentives for exponential and periodic play. Figures 4and 5display players’ incentive functions for exponential and periodic play. The defender’s incentive to play is low when the attacker moves slowly. She generally has less motivation to move than the attacker, as she is always in control while the resource is the most valuable. Note that the graphs assume that executing a move is free ( cD=cA= 0), causing incentives to always be positive. Increasing the cost of a move does not change the shape of the incentive function but shifts it downwards (by an amount ci ). Other properties of these graphs are discussed in the following sections. Games 2023,14, 74 16 of 52 0 0.5 1 1.5 0 0.5 1 1.5 νD νA (a) Defender incentive 0 0.5 1 1.5 0 0.5 1 1.5 νD νA (b) Attacker incentive Figure 4. Contour plots of incentives for exponential play and super-hyperbolic discounting (with αD=αA= 0.5, βD=βA= 1, and cD=cA= 0). Warmer colors indicate higher values; the precise numerical values are not important. 0 0.5 1 1.5 0 0.5 1 1.5 νD νA (a) Defender incentive 0 0.5 1 1.5 0 0.5 1 1.5 νD νA (b) Attacker incentive Figure 5. Contour plots of incentives for periodic play and super-hyperbolic discounting (with αD=αA= 0.5, βD=βA= 1, and cD=cA= 0). Warmer colors indicate higher values; the precise numerical values are not important. 5.4.1. Directionality of Incentives and Base Incentive Using standard analytical techniques, we can show that incentives are decreasing for both exponential and periodic play (see Appendix C.2). Specifically, for exponential play, a player’s incentive is always strictly decreasing in her own play rate. For periodic play, the direction of a player’s incentive depends on whether she is the faster or slower-moving player. If she is the slower-moving player, her incentive is independent of her play rate, while her incentive is strictly decreasing in her play rate if she is the faster-moving player. We state this more precisely in Lemmas A18 and A22. The independence of incentive is clearly visible in Figure 5as perfectly horizontal and vertical contour lines. As players’ incentives are always decreasing, their incentive is maximal when they are not playing (νi= 0 ) , and the existence of a non-zero best response depends on their incentive when they do not play. This observation motivates the definition of base incentive. Definition 5 (Base incentive) . Player i ’s base incentive when playing against a player who moves at rate ¯ νjis her incentive when not playing (νi=0). Define player i’s base incentive function as: BIi(ν) = ∂ui ∂νiνi=0 νj=ν . (31) Games 2023,14, 74 17 of 52 Note that a player’s base incentive function is a function of her opponent’s play rate. Figure 6displays base incentive functions for exponential play. Graphs of the base incentive for periodic play look very similar. 0 0.5 1 1.5 0 1 2 νA BID αD=0.20 αD=0.40 αD=0.45 αD=0.50 αD=0.55 αD=0.80 (a) Defender base incentive 0 0.5 1 1.5 0 1 2 3 νD BIA αA=0.20 αA=0.40 αA=0.45 αA=0.50 αA=0.55 αA=0.80 (b) Attacker base incentive Figure 6. Base incentive functions of the defender and the attacker for exponential play, ci= 0, βi= 1, and varying αi. Using the concept of base incentive and leveraging what we know of the directionality of the incentive function, we can state the following. Corollary 2 (Best responses for exponential play) . For exponential play, each player has a unique, single-valued best response determined by her base incentive and incentive as follows: •If her base incentive is negative, her best response is not to play. • If her base incentive is strictly positive, then her incentive function has a single root at ν? i> 0, and her best response is to play at rate ν? i. Corollary 3 (Best responses for periodic play) . Player i ’s best response to her opponent playing at rate ¯ νjcan be characterized in terms of her base incentive as follows: •If her base incentive is strictly negative, then her unique best response is not to play. •If her base incentive is zero, then any play rate νi∈[0, ¯ νj]is a best response. • If her base incentive is strictly positive, then her incentive function has a single root at ν? i> 0, and her best response is to play at rate ν? i. Corollaries 2and 3imply that each of the contour lines in Figures 4and 5corresponds to a best-response curve for a specific flip cost ( ci ). The properties of the incentive functions stated above allow finding best responses as the roots of the incentive function using straightforward numerical procedures. 5.4.2. Directionality and Origin of Base Incentive We can analyze the behavior of the base incentive function using standard analytical techniques (see Appendix C.3). The attacker’s base incentive is always strictly decreasing. The behavior of the defender’s base incentive function is discontinuous at νA=0 and depends on the relative size of 2 αD and βD and νA= 0. If 2 αD≥βD , then her base incentive function is strictly decreasing, similar to the attacker’s. If 2 αD<βD , then her base incentive is first strictly increasing, then strictly decreasing. We state this more precisely in Lemma A23. As players’ base incentives are usually maximal when their opponents are barely moving at all, whether or not players choose to participate in the game is related to the sign of the base incentive near the origin. We can again analyze the behavior of base incentives near the origin using standard analytical techniques. For both exponential and periodic play, the base incentive of the defender is equal to −cD at the origin. The value of the right limit of the base incentive of the defender at zero ( νA↓ 0 ) , and the value of the base incentive of the attacker at zero ( νD= 0), depends on the relative size of 2 αi and βi . The Games 2023,14, 74 18 of 52 origin base incentive becomes unboundedly large if 2 αi>βi and converges to specific values for 2αi<βi. We state this precisely in Lemmas A28 and A29. The properties of the base incentive function, together with their values at the origin, allow us to state the following results on best responses to non-participatory players. Corollary 4 (Defender best response to non-participatory attacker) . For exponential and periodic play, the defender’s best response to a non-participatory attacker is not to play. Corollary 5 (Attacker best response to non-participatory defender) . For exponential and periodic play, the attacker’s best response to a non-participatory defender is as follows: •If 2αA≥βA, move at non-zero rates, regardless of cost (ci). •If 2αA<βA, move at non-zero rates iff cA<1 βA−2αAand at the zero rate otherwise. 5.5. Equilibria for Super-Hyperbolic Discounting This section characterizes and presents numerical procedures for finding Nash equilibria when both players discount super-hyperbolically. 4 It starts with an investigation of equilibria in which at least one player does not move for both the periodic and the exponential regimes and to which we will refer as non-participatory equilibria (Section 5.5.1). It then discusses participatory equilibria for the exponential regimes (Section 5.5.2) and the periodic regimes (Section 5.5.3). 5.5.1. Non-Participatory Equilibria Non-participatory equilibria are (Nash) equilibria in which at least one player never moves. The following can be stated as a corollary of Corollaries 4and 5and Lemma A28. Corollary 6. For both exponential and periodic play, characterize the set of non-participatory equilibria as follows: • If 2 αA<βA and cA≥1 2αA−βA , then there is a non-participatory equilibrium in which neither player moves. • Otherwise, there may be an equilibrium in which only the attacker plays if 2 αD<βD or if 2αD=βDand cD≥1 αD. There are no other non-participatory equilibria. Algorithm 1presents a way to numerically determine the set of non-participatory strategy profiles. Algorithm 1 Procedure for finding the set of non-participatory equilibria for exponential and periodic play 1: if 2αA<βAand cA≥1 2αA−βAthen 2: return {(0, 0)}// Neither player moves 3: end if 4: if 2αD=βDand cD≥1 αDthen 5: return {(0, ν? A)}// Only attacker moves 6: end if 7: if 2αD<βDthen // Compute best response of attacker to non-participatory defender 8: ν? A→Bisect ∂uA ∂νAνA=x νD=0 ,xmin =0, xmax =1 cA// Strictly decreasing // Inspect defender’s base incentive at ν? A 9: if BID(ν? A)≤0then 10: return {(0, ν? A)}// Only attacker moves 11: end if 12: end if 13: return ∅// No non-participatory equilibria Games 2023,14, 74 19 of 52 5.5.2. Participatory Equilibria for Exponential Play We now focus on equilibria in which both players move at non-zero rates, beginning with equilibria for the exponential strategy regime. From our characterization of best responses for exponential play in terms of the base incentive function (Corollary 2) and the direction and the roots of this base incentive function (Lemma A23), we can bound the domain in which to look for participatory equilibria (ν? D,ν? A). •ν? D∈] 0, ¯ νD[ , where ¯ νD> 0 is the (unique) root of the attacker’s base incentive function. • If 2 αD≥βD , then ν? A∈] 0, ¯ νA[ , where ¯ νA> 0 is the root of the defender’s base incentive function. • If 2αD<βD, then ν? A∈]¯ ν(1) A,¯ ν(2) A[, where ¯ ν(1) Aand ¯ ν(2) Aare the roots of the defender’s base incentive function with ¯ ν(2) A≥¯ ν(1) A>0. If the mentioned roots do not exist, there is no participatory equilibrium. We can use a numerical procedure to find the participatory roots for exponential play. We find equilibria by searching through the domain ] 0, ¯ νD[ for the stable points ν? D of the function νD7→ BRD(BRA(νD)), equivalently the roots of the function r:   ]0, ¯ νD[→]−∞,νD] νD7→ νD−BRD(BRA(νD)).(32) The strategy profile (ν? D, BRA(ν? D)) is then an equilibrium. We can obtain the same results by looking for stable points of the function νA7→ BRA(BRD(νA)). 5.5.3. Participatory Equilibria for Periodic Play For periodic play, the insight that a players’ incentives are independent of their play rates when they are the slower player allows us to restrict the set of equilibria immediately. The following can be stated as a corollary of Corollary 3. Corollary 7 (Faster play rate in participatory equilibrium).In any participatory equilibrium, the faster-moving player moves at a rate corresponding to a root of the slower player’s base incentive function. When the faster player moves at a rate ¯ νf that is a root of the slower player’s base incentive, any rate νs∈] 0, ¯ νf] is a best response for the slower player. Whether there exists an equilibrium involving ¯ νf depends on whether one of these play rates makes playing ¯ νf a best response for the faster player (¯ νf=BRf(νs)). We analyze the direction of the faster player’s incentive to see if such a νs exists. It turns out that we can define the direction of player incentives in terms of the base incentive function. Lemma 5. For periodic play, the direction of the incentive of the faster player playing at rate ¯ νf with respect to the slower player’s play rate is opposite to the direction of the slower player’s base incentive in ¯ νf: ∂2uf ∂νs∂νfνs≤νf =−d BIs,f(νf) dνf =−BI0 s,f(νf), where BIs,f is the base incentive function of the slower player evaluated using the faster player’s discounting parameters αfand βf. Games 2023,14, 74 20 of 52 Proof. For equal discounting parameters, the sum of player gains is constant (Equation (A43)). For exponential and periodic play, the impact of cost on incentive is constant, so the change of the faster player’s incentive with respect to the slower player’s move rate is opposite to the change of the slower player’s incentive with respect to the faster player’s move rate: ∂ ∂νs  ∂uf ∂νf =∂ ∂νs −∂us ∂νf −cf−cs =−∂ ∂νf∂us ∂νs. The slower player’s incentive equals her base incentive (Lemma A22). This gives us an elegant description of the faster player’s incentive. Lemma 6. For periodic play, the faster player’s incentive is given by ∂uf ∂νf =BIf(νf) + (νf−νs)BI0 s,f(νf). (33) Proof. Lemma 5implies that the faster player’s incentive changes linearly as a function of the slower player’s move rate, with the slope given by −BI0 s,f(νf) . The functions for incentive for slower and faster players are equal to each other when play rates are equal; therefore, when νs=νf, player fhas an incentive of BIf(νf). Corollary 6and Equation (33) reveal that the faster player’s incentive behaves linearly with respect to the slower player’s move rate. We leverage this to determine the play rates by the slower player for which the faster player’s move rate is a best response. We can state the following as a corollary of Corollary 3and Lemma 6. Corollary 8 (Slower play rate in participatory equilibrium for periodic play) . Let ¯ νf> 0be the play rate of the faster-moving player. Move rates for the slower player to which move ¯ νf is a best response for the faster player can be characterized as follows: Ss(¯ νf) =            [0, ¯ νf]if BI0 s,f(¯ νf) = 0and BIf(¯ νf) = 0 ∅if BI0 s,f(¯ νf) = 0and BIf(¯ νf)6=0 BIf(¯ νf) BI0 s,f(¯ νf)−¯ νf∩[0, ¯ νf]otherwise. Corollaries 7and 8and our knowledge of the base incentive function allow us to finally characterize the Nash equilibria for periodic play. Theorem 4 (Participatory equilibria) . The set of equilibria where the defender moves faster than the attacker (ν? D≥ν? A) is given by: (ν? D,ν? A)|ν? D∈ {νD|BIA(νD) = 0},ν? A∈SA(ν? D), where SA is as in Lemma 8. If 2 αA≥βA or cA is <1 βA−2αA , this set has zero or one element(s). Otherwise, it is empty. Proof. Because the attacker’s base incentive is always strictly decreasing ( BI0 A,D< 0), we can disregard the case where BI0 A,D= 0, so SA(ν? D) contains at most one element regardless of the value of ν? D . The set {νD|BIA(νD) = 0 } contains at most one element. This is because the attacker’s base incentive function is strictly decreasing (Lemma A23). Therefore, there is at most one ν? D . It exists if and only if the attacker’s base incentive is strictly positive, which is the case if 2 αA<βA and cA<1/βA−2αA or if 2 αA≥βA (Lemma A29). Whether a corresponding ν? A exists is easily verified by evaluating SA(ν? D) , that is, by evaluatiing BID(ν? D)/ BI0 A,D(ν? D)−ν? D. Games 2023,14, 74 21 of 52 Theorem 5 (Participatory equilibria with faster attacker for periodic play) . The set of equilibria where the attacker moves faster than the defender (ν? A≥ν? D) is given by: (ν? D,ν? A)|ν? A∈ {νA|BID(νA) = 0},ν? D∈SD(ν? A), where SA is as in Lemma 8. This set has zero or one element(s) if 2 αD>βD . If 2 αD≥βD= 0, it has zero or one element(s) if 2 αD>βD and is empty otherwise. If 2 αD<βD , this set can have zero, one, two, or a continuum of elements. Proof. The proof of Theorem 5is identical to that of Theorem 4, except that if 2 αA<βA , then the defender’s base incentive function is first increasing then decreasing (Lemma A23). This results in up to two elements in {νA|BID(νA) = 0 } and introduces the possibility for BI0 D,A(ν∗ A)to be equal to zero, resulting in infinite roots if BIA(νA) = 0. We can derive all periodic Nash equilibria algorithmically. Algorithm 2gives an algorithm for deriving the equilibria with faster defender play. A similar but slightly longer algorithm yields the equilibria for faster attacker play. Algorithm 2 Algorithm for finding the set of participatory equilibria with faster-moving defender for periodic play // Check if the attacker’s base incentive has a root 1: if 2αA<βAand cA≥1 βA−2·αAthen return ∅end if // Determine the defender’s equilibrium play rate 2: ν? D←Bisect BIA(x),xmin =0, xmax =1/cA// Strictly decreasing // Determine attacker’s equilibrium play rate 3: ν? A←BID(ν? D)/ BI0 A,D(ν? D)−ν? D 4: if ν? A6∈ ]0, ν? D]then return ∅end if 5: return {(ν? D,ν? A)} 6. Discussion In this section, we interpret our results further and look closely at how generalized hyperbolic discounting qualitatively impacts player utilities and utility-maximizing player behavior. 6.1. Degenerated Discounting Behavior Our principal result is that complex discounted behavior often degenerates into nondiscounting behavior (Theorem 1). There is a mathematical component to this result and a practical component. We begin with the mathematical component and shed some light on how it was established. Mathematically, our result falls into the space of games involving infinite time horizons. In the original FlipIt paper, a significant part was dedicated to proving that if players are using renewal strategies, then the formalism that defines the game’s outcome in terms of very long-term (until infinity) running averages is valid. The important point was to show that if we define a game with an infinite time horizon, then there is some way to calculate the outcome from a class of strategies that can be specified finitely. Specifically, an initial result was that the formalism works for renewal strategies. Because of this background, earlier versions of our result about sub-hyperbolic discounting assumed that the strategies were renewal. This gave us the additional property that restarting the game at a later point in time yields essentially the same outcome pattern. However, we could not find any counterexample to the discounting equivalence even while violating the renewal assumption. Games 2023,14, 74 22 of 52 One key example function helped to provide insight: PC(t) =    1 if ∃m,m(m+1) 2≤t<(m+1)2 2 0 otherwise. (34) This example satisfies lim t→∞Rt 0PC(τ)dτ t=1 2. (35) However, for every integer m , there is a point in time s such that if we restart the game at time s , then the value of the functional form remains bounded away from its correct limit value for duration m . Well-behaved functions generated from renewal strategies essentially cannot depend on absolute time, so they do not have this property. Nevertheless, for this same example, the discounted limit value is also 1 2. If we perform our windowing construction using this example, we see that the restart time is quadratic in the window size. Iterating on this example, we see very quickly that whenever the restart time always dominates the next window size, then the discounted limit will converge. However, when we try to exemplify a window size that is about the same as the restart time, the original functional form does not converge. This gives us enough information to imagine the current proof structure. The narrative we now follow does not assume that the function is well behaved but rather starts from a worst-case scenario in which the player control function is “optimized” to make the discounted limit deviate from the non-discounted one. In this framing, our tactic is to use the limit definition to construct, for each fixed ε> 0, and as a function of t , a monitor on the function’s tail until the end of time. We use the properties of this monitor to show that whatever the function is doing, there is, at most, a finite amount of time after which it can no longer deviate from the truth of the original limit’s definition by more than 4 ε . From this, we show that discounting (sub-)hyperbolically results in the same outcome as not discounting at all. The main intuition here is that discounting hyperbolically or sub-hyperbolically still assigns an infinite value to future gains and costs. Therefore, whatever happens in the now is not as important as the long-term future, which, independent of any (sub-)hyperbolic discounting parameters, always tends toward the same thing. The source of gravity for this sameness is that, if we wait long enough, the discount rate drops over time for all possible parameter values until it becomes essentially flat. In any case, the result is both moderately philosophical and mathematically precise, which makes it interesting. There are also interesting behaviorally grounded aspects. Experimental research suggests that humans do temporally discount the future but tend to do so sub-hyperbolically [30,56] . 5 We have shown that the behavior of rational actors who do not discount, the behavior of those who discount sub-hyperbolically, and the behavior of those who discount hyperbolically is identical. Additionally, we observed that the behavior of super-hyperbolic discounters ( α<β ) remains similar as long as α and β are close to each other; there is no major jump or discontinuity in behavior. Consequently, the impact of discounting is absent or small for what appears to be the most relevant part of the parameter space. Our result seemingly has some implications that are perhaps counter-intuitive: • All existing research and results on FlipIt-like timing games without discounting carry over to the most commonly observed discounting behavior. This includes the results presented in van Dijk et al. [2] and follow-up work without discounting. • The players’ utilities, incentives, best responses, and the game’s equilibria are not impacted by the fact that the defender starts the game in control of the resource. Differences between defender and attacker emerge only when discounting superhyperbolically. Upon encountering a surprising result, we must evaluate if it represents a significant finding or if it is an unintended artifact of model imperfections. We discern some caveats Games 2023,14, 74 23 of 52 that may temper some of the implications of Theorem 1: one related to commitment and two arguments for super-hyperbolic discounting in our context. Commitment Consider a choice between the following two options: a reward at time T or a somewhat larger reward a fixed duration later. For sufficiently large T , a hyperbolic discounter always prefers to wait for the larger reward, as her discount rate is lower for times further removed from the present. However, unless she somehow commits to her choice, she might choose to claim the smaller reward anyway when the time T actually comes because her discount rate is high for times close to the present. This is known as “preference reversal” or “time-inconsistent preferences”. A utility-maximizing hyperbolic discounter knows that preference reversal can happen and, in the absence of commitments, has to take her own time-inconsistent preferences into account when making decisions. This can change her best response; for example, she could opt for a third option of claiming an even smaller reward sometime before T , even if she would prefer the larger reward after Twhen she opts for it. Our model assumes that the decision on which strategy to execute is made at the start of the game. We, therefore, implicitly assume that players can (and, in fact, must) commit to this strategy and cannot change it later on if their preferences were to change. Setting fixed strategies at the start of the game is common in games of timing where no new information becomes available as the game progresses, but for discounters with timeinconsistent preferences, it takes away any tension that may exist between players’ desired long-run strategies today and what their future selves would choose to do. Since we model instantaneous rewards and costs as constant over time, it appears somewhat unlikely that such tension exists, as agents are not rewarded for patience. Nevertheless, the role of commitment is something to take into account, and it could be interesting to extend the game with multiple decision-making points, e.g., in the style of extensive form games. Organizational Actors Experimental research suggests that individuals tend to discount sub-hyperbolically, but the actors in the modeled interaction typically represent highly sophisticated (adversarial) organizations. Little research exists on appropriate values for α and β in this context. Presumably, organizational actors tend to act more rationally in an economic sense, which might mean they do discount super-hyperbolically—for example, exponentially or close to exponentially (small α). Rationalization of Risk Sub-hyperbolic discounting appears to be a more appropriate model for presentfocused preferences and delay discounting. However, super-hyperbolic discounting can arise naturally as a result of rationalizing risk. We can, instead of discounting the generated value, consider the risk that the resource might stop generating value at some point in time. Assigning a constant probability of disappearance to any fixed-length time interval results in an exponentially discounted model—an extreme form of super-hyperbolic discounting. However, the entire family of (generalized) hyperbolic discount functions can appear if we consider resources that have been around for longer to be less likely to disappear or by introducing uncertainty over the probability of disappearance. 6.2. Short-Horizon and Long-Horizon Discounting For both exponential and periodic play, our description of incentives often resulted in two separate cases based on whether α was more or less than half the value of β . It turns out that the behavioral shift around this point is similarly significant to that between super- and sub-hyperbolic discounters. The result is a partitioning of the parameter space in three zones, with αi>βi being equivalent to that of FlipIt-like games without discounting [ 2 ], Games 2023,14, 74 24 of 52 2 αi<βi resulting in games that exhibit mostly the same characteristics as the exponentially discounted model presented in Merlevede et al. [13 , 14] , and a “new” space for αi<βi< 2 αi exhibiting part of the behavior of exponential discounters but not all (e.g., no decreasing attacker best-response curves or the possibility for three equilibria). In the remainder of this section, we refer to super-hyperbolic discounters with 2 α<β as short-horizon discounters and to super-hyperbolic discounters with 2 α≥β as long-horizon discounters. Figures 7and 8illustrate player behavior through best-response graphs for periodic play and varying costs for defenders and attackers. In all figures, αand βare chosen such that the value of the resource is halved over the course of the first unit of time. 6 We discuss some features of these graphs over the course of the following sections. Figure 7. Best responses of non-discounting ( β/α= 1), long-horizon ( β/α= 2), and short-horizon ( β/α= 6) defender behavior for periodic play and changing defender costs ( cD ). Attacker play rates (νA) are on the horizontal axis; defender play rates (νD) are on the vertical axis. Figure 8. Best responses of non-discounting ( β/α= 1), long-horizon ( β/α= 2), and short-horizon ( β/α= 6) attacker behavior for periodic play and changing attacker costs ( cA ). Defender play rates (νD) are on the horizontal axis; attacker play rates (νA) are on the vertical axis. 6.3. Player Indifference for Periodic Play A remarkable property of periodic strategies is that the incentive of the slower player is independent of her own play rate. This is visible as the vertical lines on Figures 7and 8 . This independence resulted in a set-valued best-response function, an interesting mathematical analysis, and, sometimes, many Nash equilibria. It appears for all values of α and β and could, therefore, also be observed in related work that did not include discounting [ 2 ] or that modeled exponential discounting [ 14 ]. However, where does the player indifference come from, and does it hold up in a real-world setting? To reason about this, we only need to concern ourselves with gains, as the impact of cost on incentives is constant irrespective of discounting parameters (Lemma 4). Careful thought 7 reveals that irrespective of the precise play rate of the slower player, every move she makes remains guaranteed to be successful, resulting in an unchanged incentive. In contrast, for faster periodic or exponential play, moving more frequently increases the probability of moving while already in control. In any real-life scenario, the assumptions made by our model are unlikely to hold exactly. As the qualitative analysis above shows, changing either the assumption of perfect stealthiness or perfect periodic-ness would disrupt the perfectly vertical or horizontal lines on our best-response curves, especially near their edges. However, the indifference is “stable” in the sense that the slope changes only slightly if assumptions are changed only slightly; Nash equilibria induced by indifference do not appear particularly brittle. It is interesting to see that the slower player’s indifference can be observed regardless of Games 2023,14, 74 31 of 52 For such an m, we can evaluate the deviation error of f(0, sm+1)as follows: f(0, sm+1)−v=sm sm+1 (f(0, sm)−v) + t∗ m sm+1 (f(sm,t∗ m)−v). (A10) The first of the two terms on the right has absolute value at most sm sm+1 ·cε, while the second term has absolute value exactly t∗ m sm+1 ·ε≥c(sm+sm+1) sm+1 ·ε=sm sm+1 ·cε+cε. No matter how these numbers are added together with their signs, the absolute value of the sum will be at least cε . We have thus exhibited a fixed constant cε , and infinitely many msatisfying |f(0, sm+1)−v| ≥ cε, which contradicts limt→∞f(0, t) = v. Corollary A1. lim t→∞ t∗ k sk+1 =0 (A11) Proof. lim t→∞ t∗ k sk+sk+1 =0⇒lim t→∞ t∗ k 2sk+1 =0⇒lim t→∞ t∗ k sk+1 =0 Corollary A2. t∗ k=o(sk), or, equivalently, lim t→∞ t∗ k sk =0. (A12) Proof. If not, then ∃c∈(0, 1),t∗ k≥cskfor infinitely many k. For each such k, t∗ k=ct∗ k+ (1−c)t∗ k ≥ct∗ k+ (1−c)csk[assumption t∗ k≥csk] =csk+1−c2sk[definition, Equation (A7)] >csk+1−c2sk+1[skis strictly increasing in k] =c(1−c)sk+1. Therefore, ∃d∈( 0, 1 ) , t∗ k≥dsk+1 for infinitely many k , contradicting the previous Corollary. Having established sufficient structure around the convergence of f(s , t) , we now proceed to the discounted limit. Define the function g(s,t)to be the discounted variant of f(s,t), namely: g(s,t) = EhRs+t τ=sPC(τ)D(τ)dτi Rs+t τ=sD(τ)dτ. (A13) Games 2023,14, 74 32 of 52 Let us define functions a(s , t) and b(s , t) to be the numerator and denominator of g(s,t), respectively. a(s,t) = EZs+t τ=sPC(τ)D(τ)dτ, (A14) b(s,t) = Zs+t τ=sD(τ)dτ. (A15) And finally, using our sequences hskiand ht∗ ki, define for each integer k, ak=a(sk,t∗ k) bk=b(sk,t∗ k). With this notation, we can finally outline our strategy for proving our main result from Equation (14), which, using our recent notation, is equivalent to lim t→∞g(0, t) = v. Given f , v , and ε> 0, we determiniscally construct the sequences hski and ht∗ ki . This construction defines a map t7→ mt , where mt is the unique integer with smt<t≤smt+1 . We thus obtain a well-defined representation of g(0, t)as g(0, t) = a(0, smt) + a(smt,t−smt) b(0, smt) + b(smt,t−smt)=∑mt−1 k=0ak+a(smt,t−smt) ∑mt−1 k=0bk+b(smt,t−smt). We now construct an upper bound for |g( 0, t)−v| (for sufficiently large t ) using three steps. First, we find an integer k∗such that ∀r≥k∗, ∑r k=k∗ak ∑r k=k∗bk −v <2ε. We then construct an integer K∗≥k∗such that ∀r≥K∗, ∑r k=0ak ∑r k=0bk −v <3ε. Finally, we construct an integer m∗≥K∗such that ∀t≥sm∗, ∑mt−1 k=0ak+a(smt,t−smt) ∑mt−1 k=0bk+b(smt,t−smt)−v <4ε. The desired result then follows because |g( 0, t)−v| is exactly the bounded value in the expression from the previous step, and the real number sm∗ in that expression is just a calculated maximum time value t=T∗ , and we can carry out this entire constructive process for an arbitrarily small ε. Assumption A1. Before doing the computations, we need a note about v and ε . These will remain fixed for the entire proof. To avoid unnecessary edge cases for v∈( 0, 1 ) , by convention and without loss of generality, we assume that ε is sufficiently small so as to ensure 0 <v− 4 ε<v+ 4 ε< 1. All of our calculations will be valid under these assumptions. If v= 0, then the calculation involving v−kε may involve some division by zero or multiplication of negative numbers. However, the result that g(something)>v−kε will still hold because g(s , t)∈[ 0, 1 ] from the definitions. The other edge case when v =1seems unproblematic with respect to the calculations in this proof. Games 2023,14, 74 33 of 52 We begin with the computation of k∗. Lemma A6. ∃k∗,∀k≥k∗,v−2ε<g(sk,t∗ k)<v+2ε. (A16) Proof. The discounting function D(τ) = 1 (1+ατ)β/α is decreasing in τ , so we can bound each value of the form g(sk , t∗ k) by replacing the function D(τ) by its constant minimum or maximum in its defined integration interval [sk,sk+1]. We have g(sk,t∗ k)≥D(sk+1)·EhRsk+1 skPC(τ)dτi D(sk)·Rsk+1 sk1 dτ=D(sk+1) D(sk)f(sk,t∗ k)≥D(sk+1) D(sk)(v−ε), g(sk,t∗ k)≤D(sk)·EhRsk+1 skPC(τ)dτi D(sk+1)·Rsk+1 sk1 dτ=D(sk) D(sk+1)f(sk,t∗ k)≤D(sk) D(sk+1)·(v+ε). Computing these ratios, we have D(sk+1) D(sk)= 1 (1+αsk+1)β/α 1 (1+αsk)β/α = 1+αsk 1+αsk+1!β/α =1− αt∗ k 1+αsk+1!β/α D(sk) D(sk+1)= 1 (1+αsk)β/α 1 (1+αsk+1)β/α =1+αsk+1 1+αskβ/α =1+ αt∗ k 1+αsk!β/α Since lim t→∞ t∗ k sk+1 =lim t→∞ t∗ k sk =0, we can find a sufficiently large number k∗ as to make the ratio D(sk+1) D(sk) as close to 1 as we wish, and the same applies to its inverse. Let us choose k∗ so that the term less than 1 is at least (1−ε v−ε)and that the term greater than 1 is at most 1 +ε v+ε. Then, we will have g(sk,t∗ k)≥1−ε v−ε(v−ε) = v−2ε g(sk,t∗ k)≤1+ε v+ε(v+ε) = v+2ε. And so, ∀k≥k∗,v−2ε<g(sk,t∗ k)<v+2ε. (A17) Corollary A3. Given k∗as above, we have ∀r≥k∗,v−2ε<∑r k=k∗ak ∑r k=k∗bk <v+2ε. (A18) Proof. Expressing the result of Lemma A6 into sequence notation, we have ∀k≥k∗,(v−2ε)·bk<ak<bk·(v+2ε). (A19) Since ∑r k=k∗bkis positive for each r≥k∗, we can write v−2ε=∑r k=k∗(v−2ε)·bk ∑r k=k∗bk <∑r k=k∗ak ∑r k=k∗bk <∑r k=k∗bk·(v+2ε) ∑r k=k∗bk =v+2ε. (A20) Games 2023,14, 74 34 of 52 Since ∑bk diverges, the finite number of terms involving ak and bk with k<k∗ cannot significantly affect the ratio of sums after sufficiently many terms. The following Lemma makes this precise. Lemma A7. ∃K∗≥k∗,∀r≥K∗,v−3ε<∑r k=0ak ∑r k=0bk <v+3ε. (A21) Proof. We may write ∑r k=0ak ∑r k=0bk =∑k∗−1 k=0ak+∑r k=k∗ak ∑k∗−1 k=0bk+∑r k=k∗bk Since limk→∞∑r k=k∗bk=∞ , and both ∑k∗−1 k=0ak and ∑k∗−1 k=0ak are finite, we may choose aK∗large enough so that ∀r≥K∗,∑k∗−1 k=0bk ∑r k=k∗bk <εand ∑k∗−1 k=0ak ∑r k=k∗bk <ε v−2ε. For the upper bound, we have ∑r k=0ak ∑r k=0bk =∑k∗−1 k=0ak ∑r k=0bk +∑r k=k∗ak ∑r k=0bk <∑r k=0ak ∑r k=k∗bk +∑r k=k∗ak ∑r k=k∗bk <ε+ (v+2ε) = v+3ε. For the lower bound, we have ∑r k=0ak ∑r k=0bk >∑r k=k∗ak ∑r k=0bk =∑r k=k∗ak ∑r k=k∗bk 1−∑k∗−1 k=0bk ∑r k=0bk!>∑r k=k∗ak ∑r k=k∗bk 1−∑k∗−1 k=0bk ∑r k=k∗bk! >(v−2ε)1−ε v−2ε=v−3ε. The next lemma formalizes a result that is analogous to tk=o(sk) in terms of the discounted variant g . Intuitively, the result for f is that as we go further in time, the potential accumulated mass of f starting at sk and going on for t∗ k is eventually (for all sufficiently large k ) dominated by the potential accumulated mass of f over the totality of time preceding sk . In the discounted version, this dominance should be even greater because the potential value of the current interval is more heavily discounted compared to everything that was accumulated before. Lemma A8. lim k→∞ bk b(0, sk)=0 (A22) Proof. Let δ> 0. Since limk→∞ t∗ k sk= 0, we may choose an integer N such that ∀k≥N , t∗ k sk <δ. Then, for k≥N, we have bk b(0, sk)≤D(sk)t∗ k b(0, sk)<D(sk)skδ b(0, sk)<b(0, sk)δ b(0, sk)=δ. Finally, we have Games 2023,14, 74 35 of 52 Lemma A9. ∃m∗,∀t≥sm∗, ∑mt−1 k=0ak+a(smt,t−smt) ∑mt−1 k=0bk+b(smt,t−smt)−v <4ε. (A23) Proof. Using the previous lemma, choose m∗≥K∗ such that ∀k≥m∗ , bk b(0,sk)<minnε,ε v−3εo . We have the relations: a(smt,t−smt)≤b(smt,t−smt)≤bmt, so then for t≥sm∗, we will have mt≥m∗and thus a(smt,t−smt) b(0, smt)≤b(smt,t−smt) b(0, smt)≤bmt b(0, smt)<minε,ε v−3ε. Using algebra similar to that of Lemma A7, together with the equivalence ∑mt−1 k=0bk= b(0, smt), we have for the upper bound, ∑mt−1 k=0ak+a(smt,t−smt) ∑mt−1 k=0bk+b(smt,t−smt) <∑mt−1 k=0ak ∑mt−1 k=0bk +a(smt,t−smt) b(0, smt) <(v+3ε) + ε=v+4ε. For the lower bound, we have ∑mt−1 k=0ak+a(smt,t−smt) ∑mt−1 k=0bk+b(smt,t−smt) >∑mt−1 k=0ak ∑mt−1 k=0bk 1−b(smt,t−smt) b(0, smt)! >(v−3ε)1−ε v−3ε=v−4ε. Since g(0, t) = ∑mt−1 k=0ak+a(smt,t−smt) ∑mt−1 k=0bk+b(smt,t−smt), this completes the proof that limt→∞g(0, t) = v. Appendix A.2. Costs The proof for costs is largely analogous to the proof for gains, with a few exceptions. Here, we supply a new modified notation and provide a new Lemma A10 summarizing the different properties of the new notation. We modify the statement and proof of Lemma A4 ; additionally, we provide modified proofs for Lemmas A5 and A9. The desired result for costs follows directly from applying the modifications and substituting the new notation into the proof from the previoussubsection. Define fc(s,t) = Eh∑τ∈ ~ ti,s<τ≤s+tcii t, (A24) and assume that for some value vc, lim t→∞fc(s,t) = vc. (A25) The function fc(s , t) has a few properties that are different from the f(s , t) defined for gains, which we summarize in the following lemma. Lemma A10. fc(s,t)is right continuous in t, and for every τ>0, 0≤fc(s,τ)−lim t→τ−fc(s,t)≤ci τ Games 2023,14, 74 36 of 52 Proof. The support of fc(s , t) (without taking an expectation) is a set of functions having the form C(s,t) = ∑τ∈ ~ ti,s<τ≤s+tci t. If we imagine that time starts at s , then this function behaves as kci t , from the time of player i ’s k th flip (or the beginning of time s , if k= 0) until just before her k+ 1st flip. At the k+ 1st flip time τ , there is an upward transition of the value by an amount ci τ , at which point the function continues as the next (k+1)ci t . Therefore, it is right continuous, not left continuous at the flip times, generally decreasing except at the flip times, and for any point τat which it is not continuous, it increases by an amount ci τ. Not all of these properties are preserved by taking weighted averages via an expectation over some members of this function class; however, right continuity, the maximum shift in the value at points of discontinuity, and the direction of the shift are preserved. The statement in the lemma succinctly formalizes these properties. The relation fc(s,t) = t s+tfc(0, s+t)−s+t sfc(0, s)shows that for every s, lim t→∞fc(s,t) = vc. (A26) Using this, for a fixed ε, inductively define sequences sc k= k−1 ∑ m=0 t∗c m(A27) t∗c k=max1, infnt∗∈R:∀t≥t∗,fcsc k,t−vc<εo. (A28) These sequences have the identical basic properties given in Lemma A3. The next Lemma is a modified version of Lemma A4. Lemma A11. If t∗c k>1, then fc(sc k,t∗c k)−vc≥ε−ci t∗c k . Proof. If t∗c k> 1, then t∗c k=inft∗∈R:∀t≥t∗,fcsc k,t−vc<ε . This definition implies limt→t∗c k−fc(sc k,t)−vc≥ε . Then, from Lemma A10, we must have limt→t∗c k−fc(sc k , t)≤ fc(sc k,t∗c k)≤limt→t∗c k−fc(sc k,t) + ci t. Therefore, fc(sc k,t∗c k)−vc≥ε−ci t. The next Lemma is more or less identical in its statement to that of Lemma A5, but it requires a new proof due to the different lower bound, which was derived in Lemma A11. Lemma A12. The newly constructed hsc kiand ht∗c kimust satisfy the following constraint: lim k→∞ t∗c k sc k+sc k+1 =0. (A29) Proof. We only need to modify the previous proof slightly. For the same reasons as before, if limk→∞ t∗c k sc k+sc k+16= 0, there must be infinitely many m satisfying t∗c m≥d·(sc m+sc m+1) for a fixed constant d . We have a revised lower bound of the form |fc(sc m , t∗c m)−vc| ≥ ε−ci t∗c m , and since there are infinitely many mwith t∗c m>maxn1, ci ε2,d·(sc m+sc m+1)o, the lower bound |fc(sc m,t∗c m)−vc| ≥ ε(1−ε) Games 2023,14, 74 37 of 52 holds for infinitely many m . The upper bound was chosen arbitrarily using the limit definition to make the algebra simple, so we can now take mlarge enough so that |fc(0, sc m)−vc|<dε(1−ε). We may thus express f( 0, sc m+1)−vc as the sum of two terms, one of which has absolute value at most sc m sc m+1 ·dε(1−ε), and the other of which has absolute value at least t∗c m sc m+1 ·ε(1−ε). Summing these will give a term having absolute value at least dε( 1 −ε) , and having infinitely many m satisfying |fc( 0, sc m+1)−vc| ≥ D for the same fixed constant D=dε(1−ε)contradicts our main assumption that limt→∞fc(0, t) = vc. The remainder of the proof steps, with one small exception, can be carried out exactly as before using the following new notations. gc(s,t) = Eh∑τ∈ ~ ti,s<τ≤s+tciDc i(τ)i Rs+t τ=sDi c(τ), (A30) ac(s,t) = E ∑ τ∈ ~ ti,s<τ≤s+t ciDc i(τ) , (A31) bc(s,t) = Zs+t τ=sDi c(τ), (A32) ac k=ac(sc k,t∗c k), (A33) bc k=bc(sc k,t∗c k). (A34) The remaining exception is that in the proof of Lemma A9, we wrote two inequalities containing and assuming the relation a(smt,t−smt)≤b(smt,t−smt). This inequality is true for the gains but not necessarily for the costs since it is possible for the flip costs to accrue early on in the interval that starts with smt . However, the only property we required was much weaker, namely that for mt≥m∗, we have ac(sc mt,t−sc mt) bc(0, sc mt)<ε. And for this we could use any number of relations, which do not even depend on vc≤1, for example, ac(sc mt,t−sc mt)≤ac mt<bc mt·(vc+2ε) together with bc k=o(bc(0, sc k)). Therefore, with these specific modifications to the lemmas and their proofs, we can perform the substitution f(),sk,t∗ k,g(),a(),b(),ak,bk7→ fc(),sc k,t∗c k,gc(),ac(),bc(),ac k,bc k, Games 2023,14, 74 38 of 52 and the modified proof for gains shows that we have the desired result for costs also; namely that lim t→∞gc(0, t) = vc. (A35) Appendix B. Player Gains Appendix B.1. Structural Aspects This subsection derives closer-to-analytical expressions for player gains when discounting super-hyperbolically and introduces precise definitions for the defender advantage and player anonymous gain. It presents formulae for player gains in terms of these concepts. The fact that player i ’s calculated resource valuation is finite, together with the restriction of strategies that allow us to replace lim sup with limit, allows us to obtain a much closer-to-analytic expression for the gain of player i(starting from Equation (9)): Gi=lim sup t→∞ EhRt τ=0PCi(τ)·Di(τ)dτi Rt τ=0Di(τ)dτ=lim t→∞ E  Zt τ=0 βi−αi (1+αiτ) βi αi PCi(τ)dτ  , (A36) where PCi is the player control indicator function from Equation (6) that encodes the game’s temporal outcomes, and expectations are over PCi. At this point, it becomes useful to separate the expression for gain according to two time periods—the time before the first flip of the game, and all the time afterward. This allows us to work with two quantities, the first of which has a direct analytic representation, and the second of which has a representing formula that is uniform with respect to the identity of i(A or D): Gi=E  Zt0 τ=0 βi−αi (1+αiτ) βi αi dτ   ~ 1i=D+lim t→∞ E  Zt τ=t0 βi−αi (1+αiτ) βi αi PCi(τ)dτ  . We refer to the first term of this sum, Di , as the defender advantage because its value only accrues toward the defender’s utility. We refer to the second term of the sum, Gi , as the anonymous gain, because the expression is uniform in the identity of i (which is not true of the gain as a whole). Evaluating the defender advantage does not require structural knowledge of the strategy configuration and can be derived by calculus. The term evaluates to Di=~ 1i=D  1−E  1 1+αit0 βi−αi αi   , (A37) which reduces to DD=1−E  1 1+αDt0 βD−αD αD and (A38) DA=0. (A39) Evaluating the anonymous gain requires structural knowledge of the strategies being employed by each player, and this expression is a prominent subject of our investigations in the following subsections. For reference, the anonymous gain transcribes directly to Games 2023,14, 74 39 of 52 Gi=lim t→∞ E  Zt τ=t0 βi−αi (1+αiτ) βi αi PCi(τ)dτ  . (A40) Because Gi=Di+Gi , we can obtain what we need about the gain of players by focusing our attention on the defender advantage and the anonymous gain. Moreover, by extending these two notations only slightly, we may take advantage of additional symmetries of the game formulation. Formally, we define Di,j=~ 1i=D    1−E   1 1+αjt0 βj−αj αj      , and (A41) Gi,j=lim t→∞ E   Zt τ=t0 βj−αj (1+αjτ) βj αj PCi(τ)dτ    . (A42) The first index i of the notation identifies the player whose action strategy is being applied, while the second index j corresponds to the player whose discounting method is being applied. This extended notation subsumes the original via Di=Di,i and Gi=Gi,i . It also allows us to compute some terms from others even when the two players apply different discounting implementations. Such simplifications rely on the normalized gains of two competing players summing to one if the players apply the same discounting parameters. When players discount at different rates, this symmetry is broken, but its usefulness can still be recovered by using the following equation, which uses our extended notation: DD,j+GD,j+GA,j=1. (A43) This works because the discounting of player j is being applied equally to three types of game outcomes—the time before t0 , the time when the defender controls after t0 , and the time when the attacker controls after t0—which together subsume all possible outcomes. Finally, we may also use the property of finite total resource valuation to obtain a closer-to-analytic expression for the costs of player istarting from Equation (13): Ci=lim t→∞ Eh∑τ∈ ~ ti,τ≤tciDc i(τ)i Rt τ=0Di(τ)dτ=lim t→∞ E   ∑ τ∈ ~ ti,τ≤t ci(βi−αi) (1+αc iτ) βc i αc i     . (A44) Throughout the remainder of this section and the following sections, we omit the subscripts to discounting parameters α and β ; which subscript applies should be clear from the context. Appendix B.2. Exponential Play We start with the derivation of anonymous player gains for exponential play. We can obtain the defender’s and the attacker’s gain from these expressions because players’ gains with equal discounting parameters sum to one, as is expressed by Equation (A43). To obtain the anonymous player gain of player i , we derive a formula for the fraction of the total anonymous gain she obtains and a formula for the total anonymous gain. Games 2023,14, 74 40 of 52 Lemma A13. For exponential play, each player i obtains a fraction of the total anonymous gain equal to: νi νi+νj . Proof. At any moment in time, the probability that player i is the next player to move is equal to pi=Z+∞ τi=0νie−νiτiZ+∞ τj=τi νje−νiτjdτjdτi=νi νi+νj . Probability pi is, therefore, also equal to the probability that any flip by either player after time t0is made by player i. Consider the set of all intervals between flips. For each such interval, with probability pi , player i is the player who receives gain over the entire interval. With probability 1 −pi , she receives nothing. Therefore, her expected gain over each interval is pi times the value of the interval. By linearity of expectation, her total expected gain over this set of intervals is then pitimes the total combined value of the intervals. Lemma A14. For exponential play, the time of the first move is exponentially distributed with rate parameter νi+νj. Proof. Let Xibe the time of i’s first flip, and define the random variable Z=min{Xi,Xj} as the time until the first flip by either player. Then FZ(z) = Pr[Z≤z] = Pr[Xi≤zor Xj≤z] =1−Pr[Xi≥zand Xj≥z] =1−Pr[Xi≥z]·Pr[Xj≥z] =1−(1−(1−e−νiz))(1−(1−e−νjz)) =1−e−(νi+νj)z, that is, Zis distributed exponentially with rate parameter νi+νj. Lemma A15 (Anonymous gains for exponential play) . If both players are playing exponentially, then player i’s anonymous gain is given by: Gi=νi νi+νj ·fβi−αi αi ,νi+νj αi. (A45) Proof. From Lemma A13, it follows that proving this theorem means showing that the total anonymous gain is given by G=fβ−α α,νi+νj α. (A46) From Lemma A14, we know that the time of the first move is exponentially distributed with rate parameter νi+νj . The total anonymous gain is defined as the expected total gain obtained after the first move by either player: Games 2023,14, 74 47 of 52 Lemma A24. For exponential play, the attacker’s base incentive function is strictly decreasing. Proof. The attacker’s base incentive function is given by −cA+eνD αA ·Eβ−α ανD αA, the derivative of which is 1 α2 A eνD αEβ−α ανD α−eνD αEβ−2α ανD α!, which is strictly negative as the exponential integral function is strictly decreasing in its order. Lemma A25. For periodic play, the attacker’s base incentive function is strictly decreasing. Proof. We have ∂2uA ∂νD∂νAνA≤νD =∂ ∂νD −cA+hA(νD)−1 2αA−βA!=h0 A(νD), which is strictly negative by Lemma A21. Lemma A26. For periodic play and 2 αD≥βD , the defender’s base incentive function is strictly decreasing for strictly positive attack rates (νA∈]0, +∞[). Proof. Compute the direction and curvature of the defender’s incentive for slower defender play as: ∂2uD ∂νA∂νDνD<νA =−h0 D(νA)−νAh00 D(νA)(A71) = 1−1+α νAα−β α(β−2α)2+(β−α)νA+ν2 A ν2 A (β−2α)(β−3α) ∂3uD ∂ν2 A∂νDνD<νA =2α−β+νA ν4 A1+α νA−β α. (A72) From Equation (A72), we can see that the sign of 2 α−β+νA determines the curvature of the defender’s incentive. Since 2 α≥β , the defender’s incentive is strictly convex for νA> 0. To see that the direction is strictly negative, we compute the limit of the direction for νA→+∞: lim νA→∞ ∂2uD ∂νA∂νDνD<νA (A73) = 1−limνA→+∞1+α νAα−β α(β−2α)2+(β−α)νA+ν2 A ν2 A (β−2α)(β−3α) =1−1·(0+0+1) (β−2α)(β−3α) =0. Games 2023,14, 74 48 of 52 Since the direction is constantly increasing and has a limiting value of zero, it is always strictly negative. Lemma A27. For periodic play and 2 αD<βD , the defender’s base incentive function is first strictly increasing, then strictly decreasing. Proof. Equations (A72) and (A73) show that the direction of the defender’s incentive is negative (and increasing) for large enough νA . Because the curvature of the incentive changes only once and the direction has an asymptotic root at νA→+∞ , its direction has at most one real root and changes sign at most once. The proof, therefore, reduces to showing that the defender’s base incentive is strictly increasing for νA↓0. Looking at Equation (A71), we see that this is equivalent to showing that f=1+α νAα−β α(β−2α)2+ (β−α)νA+ν2 A ν2 A is smaller than one if β> 3 α and greater than one otherwise. We can show that if β> 3 α , then limνA↓0f= 0, and that if β< 3 α , then f becomes unboundedly large for small νA . Appendix C.4. Origin of Base Incentive The following results can be derived using standard analytical techniques and properties of the exponential integral function. We state them without proof. Lemma A28 (Origin of the defender’s base incentive) . For both exponential and periodic play, we can characterize the origin of the defender’s base incentive as follows: •It is equal to −cDfor νA=0. •If 2αD>βD, then it becomes unboundedly large for νA↓0. •If 2αD=βD, then it is equal to 1 αD−cDfor νA↓0. •If 2αD<βD, then it is equal to −cDfor νA↓0. Lemma A29 (Origin of the attacker’s base incentive) . For both exponential and periodic play, we can characterize the origin of the attacker’s base incentive as follows: •If 2αA≥βA, then it becomes unboundedly large for νD↓0. •If 2αA<βA, then it is equal to 1 βA−2·αA−cAfor νD=0. Appendix D. A Renewal Strategy Beating the Periodic Strategy Proof of Theorem 6. We prove this by showing that for every defender and periodic strategy, a renewal strategy exists that outperforms it. Define RPD(ν) as the renewal strategy induced by a Dirac point distribution at 1/ν , resulting in a periodic strategy with the phase fixed to be equal to the period and moving at times (1/ν , 2/ν , . . .) . This strategy outperforms the periodic strategy with rate parameter νfor the defender. To keep the math simple, assume that αD→ 0, that is, assume that the defender discounts by the exponential function e −βDt . The cost of playing the renewal strategy RPD is then: βD +∞ ∑ i=1 cDe−βDi/ν=cDβD eβD/νD−1. This cost is strictly lower than the cost of cDνD of the periodic strategy with rate parameter νD . Because she starts in control of the resource, the total expected gain of the Games 2023,14, 74 49 of 52 defender when playing renewal strategy RPD in response to an attacker playing a periodic strategy with νA≥νDis: βD +∞ ∑ k=0Zk/νD+1/νA τ=k/νD e−βτ dτ=e βD νD−βD νA·eβD/νA−1 eβD/νD−1. We can show that this gain is strictly higher than the gain derived from playing the periodic strategy with rate parameter νD . As its cost is lower and its gain is higher, the renewal strategy RPD yields higher utility than the periodic strategy, at least under the assumed conditions. Although RPD is arithmetic, renewal strategies defined by a narrow (but not “infinitely narrow”) distribution around 1 /νD are non-arithmetic and will yield approximately the same utilities for super-hyperbolic discounters. Notes 1 Similar formulations of the generalized hyperbolic discounting function have been referred to as hyperboloid (e.g., Estle et al. [54] ) or hyperbola-like (e.g., Green and Myerson [55]) discounting functions. 2In Merlevede et al. [13], these parameters are named λDand λAinstead of βDand βA, and ΛDand ΛAinstead of βc Dand βc A. 3A player can move at most once at a particular instance of time and a finite number of times over any finite time interval. 4 We limit ourselves to cases where both players discount super-hyperbolically and do not exhaustively cover scenarios where one player discounts super-hyperbolically and another sub-hyperbolically. As there is no discontinuity in player best responses near α=β , outcomes for mixed scenarios can be observed under the super-hyperbolic discounting by bringing α and β close to each other. 5 This research does primarily focus on income or costs occurring at one particular point in time, not on continuous income and cost streams as is the case here. 6We pick a value for r=β/αand then compute parameters αand βas α=21/r−1 and β=α·r. 7 Define income resulting from a move as follows: ( • ) When a player in control of the resource performs a move, it does not result in income. ( • ) When a player not in control of the resource performs a move, flipping the resource, it results in income equal to the value generated by the resource from the time of the move until the resource flips again. Start by splitting the game into two parts: the part of the game before the attacker’s first move (ante) and the part after the attacker’s first move (post). • (ante) If the attacker is the slower player, this part of the game always yields him precisely zero gain, irrespective of either player’s play rate. If the defender is the slower player, this part of the game does yield her gain, but the expected amount only depends on the duration of ante, which is also independent of her play rate. • (post) The probability density of the slower player flipping at any time is constant and equal to her play rate. Every one of her moves results in an expected income; while difficult to determine this income exactly, it is independent of her own play rate, as it is certain to result in a change of ownership while the value of ownership depends only on the time and the play rate of the faster player. The slower player’s gain is, therefore, proportional to her play rate, and her incentive is independent of her play rate. 8 To be precise, the periodic strategies strictly dominate the class of non-arithmetic or non-lattice renewal strategies. Arithmetic strategies are strategies for which all possible realizations happen at an integer multiple of some real number. 9 The CDF for the time of the first move is 1 −( 1 −t0νf)~ It0≤1/νf·( 1 −tνs)~ It0≤1/νs=t0(νf+νs−t0νfνs)~ It0≤1/νf , where ~ I is the indicator function. The PDF for the time of the first move is its derivative, p(t0) = j(νf+νs− 2 t+ 0 νfνs)~ It0≤1/νf . An expression for the total anonymous gain is, therefore, (β−α)Z+∞ t0=0p(t0)Z+∞ τ=t0 Di(τ)dτdt0=Z1/νf t0=0 νf+νs−2t0νfνs (1+t0α)β−α α , .i f t0. We can confirm that this expression is equal to the sum of Gsand Gfas presented in Lemmas A16 and A17. 10 The probability of the faster player having moved last at any point in time is equal to R1/νf τ=0νf·τ·νsdτ=νs 2νf . The slower player moved last with probability 1 −νs 2νf. 11 The expected duration of the intervals owned by the faster-moving player is R1/νf τ=0( 1 −τ νs)dτ=1 νf−νs 2ν2 f . The intervals of the slower-moving player have a shorter expected length of R1/νf τ=0(1−τ νf)dτ=1 2νf. Games 2023,14, 74 50 of 52 12 https://dlmf.nist.gov/8.19.E15, (accessed on 10 November 2023). References 1. Radzik, T. Results and Problems in Games of Timing; Lecture Notes-Monograph Series; Institute of Mathematical Statistics: Durham, NC, USA, 1996; pp. 269–292. [CrossRef] 2. van Dijk, M.; Juels, A.; Oprea, A.; Rivest, R.L. FlipIt: The Game of “Stealthy Takeover”. J. Cryptol. 2012,26, 655–713. [CrossRef] 3. Pawlick, J.; Farhang, S.; Zhu, Q. Flip the Cloud: Cyber-Physical Signaling Games in the Presence of Advanced Persistent Threats. In Decision and Game Theory for Security; Khouzani, M., Panaousis, E., Theodorakopoulos, G., Eds.; Number 9406 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2015; pp. 289–308. [CrossRef] 4. Kushner, D. The Real Story of Stuxnet. IEEE Spectrum 2013,50, 48–53. [CrossRef] 5. Barrett, D.; Yadron, D.; Paletta, D. U.S. Suspects Hackers in China Breached about 4 Million People’s Records, Officials Say. Available online: http://www.wsj.com/articles/u-s-suspects-hackers-in-china-behind-government-data-breach-sources-say- 1433451888 (accessed on 7 October 2023). 6. Perez, E.; Prokupecz, S. U.S. Data Hack May Be 4 Times Larger Than the Government Originally Said. Available online: http://edition.cnn.com/2015/06/22/politics/opm-hack-18-milliion (accessed on 7 September 2023). 7. Gallagher, R. The Inside Story of How British Spies Hacked Belgium’s Largest Telco. Available online: https://theintercept.com/ 2014/12/13/belgacom-hack-gchq-inside-story/ (accessed on 10 November 2023). 8. Price, R. TalkTalk Hacked: 4 Million Customers Affected, Stock Plummeting, ’Russian Jihadist Hackers’ Claim Responsibility. Available online: http://uk.businessinsider.com/talktalk-hacked-credit-card-details-users-2015-10 (accessed on 10 November 2023). 9. Reuters. Premera Blue Cross Says Data Breach Exposed Medical Data. Available online: http://www.nytimes.com/2015/03/18 /business/premera-blue-cross-says-data-breach-exposed-medical-data.html (accessed on 10 November 2023). 10. ESET. APT Activity Report T2 2022. Available online: https://www.eset.com/int/business/resource-center/reports/eset-apt- activity-report-t2-2022 (accessed on 10 November 2023). 11. Cole, E. (Ed.) Preface. In Advanced Persistent Threat; Syngress: Boston, MA, USA, 2013; pp. xv–xvi. [CrossRef] 12. Nadela, S. Enterprise Security in a Mobile-First, Cloud-First World. Available online: http://news.microsoft.com/security2015/ (accessed on 10 November 2023). 13. Merlevede, J.; Johnson, B.; Grossklags, J.; Holvoet, T. Exponential Discounting in Security Games of Timing. In Proceedings of the Workshop on the Economics of Information Security (WEIS), Boston, MA, USA, 2–3 June 2019. 14. Merlevede, J.; Johnson, B.; Grossklags, J.; Holvoet, T. Exponential Discounting in Security Games of Timing. J. Cybersecur. 2021 ,7, tyaa008. [CrossRef] 15. Merlevede, J.; Johnson, B.; Grossklags, J.; Holvoet, T. Time-Dependent Strategies in Games of Timing. In Decision and Game Theory for Security; Alpcan, T., Vorobeychik, Y., Baras, J.S., Dán, G., Eds.; Number 11836 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2019; pp. 310–330. [CrossRef] 16. Loewenstein, G.; Prelec, D. Anomalies in Intertemporal Choice: Evidence and an Interpretation. Q. J. Econ. 1992 ,107, 573–597. [CrossRef] 17. Samuelson, P. A Note on Measurement of Utility. Rev. Econ. Stud. 1937,4, 155–161. [CrossRef] 18. Ramsey, F.P. A Mathematical Theory of Saving. Econ. J. 1928,38, 543–559. [CrossRef] 19. Strotz, R.H. Myopia and Inconsistency in Dynamic Utility Maximization. Rev. Econ. Stud. 1955,23, 165–180. [CrossRef] 20. Koopmans, T.C. Stationary Ordinal Utility and Impatience. Econometrica 1960,28, 287–309. [CrossRef] 21. Fishburn, P.C.; Rubinstein, A. Time Preference. Int. Econ. Rev. 1982,23, 677–694. [CrossRef] 22. Callender, C. The Normative Standard for Future Discounting. Australas. Philos. Rev. 2021,5, 227–253. [CrossRef] 23. Vanderveldt, A.; Oliveira, L.; Green, L. Delay Discounting: Pigeon, Rat, Human—Does it Matter? J. Exp. Psychol. Anim. Learn. Cogn. 2016,42, 141–162. [CrossRef] 24. O’Donoghue, T.; Rabin, M. Doing It Now or Later. Am. Econ. Rev. 1999,89, 103–124. [CrossRef] 25. Laibson, D. A Cue-theory of Consumption. Q. J. Econ. 2001,116, 81–119. [CrossRef] 26. Zauberman, G.; Kim, B.K.; Malkoc, S.A.; Bettman, J.R. Discounting Time and Time Discounting: Subjective Time Perception and Intertemporal Preferences. J. Mark. Res. 2009,46, 543–556. [CrossRef] 27. Kahneman, D.; Tversky, A. Prospect Theory: An Analysis of Decision under Risk. Econometrica 1979,47, 363–391. [CrossRef] 28. Frederick, S. Cognitive Reflection and Decision Making. J. Econ. Perspect. 2005,19, 25–42. [CrossRef] 29. Laibson, D. Golden Eggs and Hyperbolic Discounting. Q. J. Econ. 1997,112, 443–478. [CrossRef] 30. Ericson, K.M.; Laibson, D. Chapter 1—Intertemporal Choice. In Handbook of Behavioral Economics—Foundations and Applications 2; Elsevier: Amsterdam, The Netherlands, 2019; pp. 1–67. [CrossRef] 31. Herrnstein, R.J. Relative and Absolute Strength of Response as a Function of Frequency of Reinforcement. J. Exp. Anal. Behav. 1961,4, 267–272. [CrossRef] 32. McKerchar, T.L.; Green, L.; Myerson, J. On the Scaling Interpretation of Exponents in Hyperboloid Models of Delay and Probability Discounting. Behav. Process. 2010,84, 440–444. [CrossRef] Games 2023,14, 74 51 of 52 33. Myerson, J.; Green, L. Discounting of Delayed Rewards: Models of Individual Choice. J. Exp. Anal. Behav. 1995 ,64, 263–276. [CrossRef] 34. McKerchar, T.L.; Green, L.; Myerson, J.; Pickford, T.S.; Hill, J.C.; Stout, S.C. A Comparison of Four Models of Delay Discounting in Humans. Behav. Process. 2009,81, 256–259. [CrossRef] [PubMed] 35. Green, L.; Myerson, J.; Vanderveldt, A. Delay and Probability Discounting. In The Wiley Blackwell Handbook of Operant and Classical Conditioning; John Wiley & Sons, Ltd.: Hoboken, NJ, USA, 2014; Chapter 13, pp. 307–337. [CrossRef] 36. Sozou, P.D. On Hyperbolic Discounting and Uncertain Hazard Rates. Proc. R. Soc. B Biol. Sci. 1998,265, 2015–2020. [CrossRef] 37. Azfar, O. Rationalizing Hyperbolic Discounting. J. Econ. Behav. Organ. 1999,38, 245–252. [CrossRef] 38. Fernandez-Villaverde, J.; Mukherji, A. Can We Really Observe Hyperbolic Discounting? 2002. Available online: https: //ssrn.com/abstract=306129 (accessed on 13 July 2023). [CrossRef] 39. Bowers, K.D.; van Dijk, M.; Griffin, R.; Juels, A.; Oprea, A.; Rivest, R.L.; Triandopoulos, N. Defending against the Unknown Enemy: Applying FlipIt to System Security. In Decision and Game Theory for Security; Grossklags, J., Walrand, J., Eds.; Number 7638 in Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2012; pp. 248–263. [CrossRef] 40. Laszka, A.; Felegyhazi, M.; Buttyán, L. A Survey of Interdependent Security Games. ACM Comput. Surv. (CSUR) 2014 ,47. [CrossRef] 41. Manshaei, M.; Zhu, Q.; Alpcan, T.; Bac¸sar, T.; Hubaux, J.P. Game Theory Meets Network Security and Privacy. ACM Comput. Surv. (CSUR) 2013,45, 1–49. [CrossRef] 42. Laszka, A.; Johnson, B.; Grossklags, J. Mitigating Covert Compromises. In Proceedings of the Web and Internet Economics, Cambridge, MA, USA, 11 December 2013; pp. 319–332. [CrossRef] 43. Laszka, A.; Johnson, B.; Grossklags, J. Mitigation of Targeted and Non-Targeted Covert Attacks as a Timing Game. In Decision and Game Theory for Security; Das, S.K., Nita-Rotaru, C., Kantarcioglu, M., Eds.; Number 8252 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2013; pp. 175–191. [CrossRef] 44. Farhang, S.; Grossklags, J. FlipLeakage: A Game-Theoretic Approach to Protect against Stealthy Attackers in the Presence of Information Leakage. In Decision and Game Theory for Security; Zhu, Q., Alpcan, T., Panaousis, E., Emmanouil Tambe, E., Casey, W., Eds.; Number 9996 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2016; pp. 195–214. [CrossRef] 45. Johnson, B.; Laszka, A.; Grossklags, J. Games of Timing for Security in Dynamic Environments. In Decision and Game Theory for Security; Khouzani, M., Panaousis, E., Theodorakopoulos, G., Eds.; Number 9406 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2015; pp. 57–73. [CrossRef] 46. Zhang, M.; Zheng, Z.; Shroff, N.B. A Game Theoretic Model for Defending Against Stealthy Attacks with Limited Resources. In Decision and Game Theory for Security; Khouzani, M., Panaousis, E., Theodorakopoulos, G., Eds.; Number 9406 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2015; pp. 93–112. [CrossRef] 47. Laszka, A.; Horvath, G.; Felegyhazi, M.; Buttyán, L. FlipThem: Modeling Targeted Attacks with FlipIt for Multiple Resources. In Decision and Game Theory for Security; Poovendran, R., Saad, W., Eds.; Number 8840 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2014; pp. 175–194. [CrossRef] 48. Leslie, D.; Sherfield, C.; Smart, N.P. Threshold FlipThem: When the Winner Does Not Need to Take All. In Decision and Game Theory for Security; Khouzani, M., Panaousis, E., Theodorakopoulos, G., Eds.; Number 9406 in Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2015; pp. 74–92. [CrossRef] 49. Hu, P.; Li, H.; Fu, H.; Cansever, D.; Mohapatra, P. Dynamic Defense Strategy against Advanced Persistent Threat with Insiders. In Proceedings of the 2015 IEEE Conference on Computer Communications (INFOCOM), Hong Kong, China, 26 April–1 May 2015; pp. 747–755. [CrossRef] 50. Oakley, L.; Oprea, A. QFlip: An Adaptive Reinforcement Learning Strategy for the FlipIt Security Game. In Decision and Game Theory for Security; Alpcan, T., Vorobeychik, Y., Baras, J.S., Dán, G., Eds.; Number 11836 in Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2019; pp. 364–384. [CrossRef] 51. Zhang, R.; Zhu, Q. FlipIn : A Game-Theoretic Cyber Insurance Framework for Incentive-Compatible Cyber Risk Management of Internet of Things. IEEE Trans. Inf. Forensics Secur. 2020,15, 2026–2041. [CrossRef] 52. Banik, S.; Bopardikar, S.D. FlipDyn: A Game of Resource Takeovers in Dynamical Systems. In Proceedings of the 2022 IEEE 61st Conference on Decision and Control (CDC), Cancun, Mexico, 6–9 December 2022; pp. 2506–2511. [CrossRef] 53. Miura, H.; Kimura, T.; Hirata, K. Modeling of Malware Diffusion with the FlipIt Game. In Proceedings of the 2020 IEEE International Conference on Consumer Electronics—Taiwan (ICCE-Taiwan), Taoyuan, Taiwan, 28–30 September 2020; pp. 1–2. [CrossRef] 54. Estle, S.; Green, L.; Myerson, J. When Immediate Losses are Followed by Delayed Gains: Additive Hyperboloid Discounting Models. Psychon. Bull. Rev. 2019,26, 1418–1425. [CrossRef] 55. Green, L.; Myerson, J. A Discounting Framework for Choice with Delayed and Probabilistic Rewards. Psychol. Bull. 2004 , 130, 769–792. [CrossRef] 56. Frederick, S.; Loewenstein, G.; O’Donoghue, T. Time Discounting and Time Preference: A Critical Review. J. Econ. Lit. 2002 , 40, 351–401. [CrossRef] Games 2023,14, 74 52 of 52 57. Grossklags, J.; Reitter, D. How Task Familiarity and Cognitive Predispositions Impact Behavior in a Security Game of Timing. In Proceedings of the 2014 IEEE 27th Computer Security Foundations Symposium, Vienna, Austria, 19–22 July 2014; pp. 111–122. [CrossRef] 58. Reitter, D.; Grossklags, J. The Positive Impact of Task Familiarity, Risk Propensity, and Need for Cognition on Observed Timing Decisions in a Security Game. Games 2019,10, 49. [CrossRef] Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.