Cooperate without looking in a non-repeated game
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Hilbe, Christian; Hoffman, Moshe; Nowak, Martin A. Article Cooperate without looking in a non-repeated game Games Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Hilbe, Christian; Hoffman, Moshe; Nowak, Martin A. (2015) : Cooperate without looking in a non-repeated game, Games, ISSN 2073-4336, MDPI, Basel, Vol. 6, Iss. 4, pp. 458-472, https://doi.org/10.3390/g6040458 This Version is available at: https://hdl.handle.net/10419/167955 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Games 2015,6, 458-472; doi:10.3390/g6040458 OPEN ACCESS games ISSN 2073-4336 www.mdpi.com/journal/games Article Cooperate without Looking in a Non-Repeated Game Christian Hilbe 1,*, Moshe Hoffman 1and Martin A. Nowak 1,2 1Program for Evolutionary Dynamics, Harvard University, Cambridge, MA 02138, USA; E-Mails: hof[email protected] (M.H.); now[email protected]ard.edu (M.A.N.) 2Department of Organismic and Evolutionary Biology, Department of Mathematics, Harvard University, Cambridge, MA 02138, USA *Author to whom correspondence should be addressed; E-Mail: [email protected]ard.edu; Tel.: +1-617-496-4737. Academic Editor: Ulrich Berger Received: 16 June 2015 / Accepted: 24 September 2015 / Published: 30 September 2015 Abstract: We propose a simple model for why we have more trust in people who cooperate without calculating the associated costs. Intuitively, by not looking at the payoffs, people indicate that they will not be swayed by high temptations to defect, which makes them more attractive as interaction partners. We capture this intuition using a simple four-stage game. In the first stage, nature draws the costs and benefits of cooperation according to a commonly-known distribution. In the second stage, Player 1 chooses whether or not to look at the realized payoffs. In the third stage, Player 2 decides whether to exit or let Player 1 choose whether or not to cooperate in the fourth stage. Using backward induction, we provide a complete characterization for when we expect Player 1 to cooperate without looking. Moreover, we show with numerical simulations how cooperating without looking can emerge through simple evolutionary processes. Keywords: evolutionary game theory; cooperation; emotions; principled behavior JEL classifications: C72; C73; D03; D64
Games 2015,6459 1. Introduction In various instances, the mere act of considering one’s strategic options and gathering information about the possible costs and benefits of an action will be met with distrust (see also [1]). For example, in most romantic relationships, it will be taken as suspicious if one of the partners starts to flirt with another potential mate. Similarly, societies often have a set of taboo topics (like eating beef in India) whose mere consideration is socially sanctioned. In these examples, people are prevented from seeking for detailed information to ensure that they withstand situations in which there would be a high temptation to defect. Herein, we propose a simple model that formalizes this mechanism. Our model is a modified version of the envelope game proposed by Hoffman et al. [1], and we refer to it as the non-repeated envelope game (in contrast with the original model, which requires repeated interactions). The non-repeated envelope game describes a strategic interaction with four stages, involving two players, Player 1 and Player 2 (as illustrated in Figure 1; a more detailed explanation is given in the next section). In the first stage, nature randomly determines the possible payoffs of the game. In the second stage, Player 1 can decide whether she wants to learn the realized payoffs. However, Player 1’s decision is observed by Player 2, who can then decide in the third stage whether to exit the interaction or whether to continue. Only if Player 2 continues, there is a fourth stage in which Player 1 decides whether to cooperate or to defect. We will show that if nature only occasionally chooses a game in which Player 1 is tempted to defect (but these instances of defection are extremely damaging to the co-player), then there is an equilibrium in which Player 1 cooperates without looking. Nature high low Player 1's defection payoff or Player 1 Do not look Look Player 2 Continue Exit Player 1 Cooperate Defect Payoffs Player 1 Player 2 a b cidi 0 0 Stage 1 Stage 2 Stage 3 Stage 4 Figure 1. Schematic representation of the non-repeated envelope game. We consider a game with four stages, involving two players. In the first stage of the game, nature randomly determines Player 1’s payoff from defection, which can be either high (h) or low (l). Nature puts the respective information into an envelope. In the second stage, Player 1 is given the chance to learn the realized payoff by opening the envelope. In the third stage, Player 2 decides whether to exit the interaction, depending on whether Player 1 knows the realized payoffs. If Player 2 continues, there is a fourth stage in which Player 1 decides whether or not to cooperate. We assume that if Player 1 defects, the payoffs of both players, ciand di, depend on the state of nature i∈ {l, h}: when the state is h, then Player 1 has a higher incentive to defect (i.e.,ch> cl). In such an equilibrium, Player 1 may sometimes cooperate, although cooperation happens to be against her own material interests. In this way, our model suggests an alternative explanation for
Games 2015,6460 altruistic acts in anonymous one-shot social dilemmas (as documented, for example, in [2–5]). Previous explanations have suggested that such seemingly irrational behaviors have evolved because there is uncertainty about the possibility of future encounters [6] or that they are a consequence of group selection [7].1Others have emphasized the role of heuristics [13–15]. According to that latter explanation, people develop simple heuristics in order to help them deal with recurring situations without having to employ the more costly mental process of conscious deliberation [16–18]. In contrast, our model does not require intuitive cooperation to be inherently more efficient than a conscious calculation of payoffs. Instead, our model suggests an endogenous incentive to ignore payoff-relevant information: when people cooperate without looking, they can hope to gain their co-player’s trust. 2. Model: The Non-Repeated Envelope Game In the following, let us describe the four stages of the non-repeated envelope game in more detail. In the first stage, nature chooses between two different states Ω = {l, h}, with the two letters land hreferring to the defection payoff of Player 1 (which can be either low or high). Let pdenote the probability that nature chooses the low defection payoff. We assume that pis common knowledge, but the players do not observe directly which state nature has chosen. That is, both players know how often on average there will be a low defection payoff for Player 1, but they do not know whether defection is particularly beneficial for Player 1 in the current situation. In Figure 1, this scenario is represented by a closed envelope, which contains the information about the realized state of nature. In the second stage, Player 1 decides whether or not to learn the state of nature (i.e., whether or not to open the envelope). We suppose that learning the state of nature does not entail any direct costs. However, it may affect the players’ actions in the subsequent course of the game. Specifically, we assume that in the third stage, Player 2 can decide whether or not to continue the interaction, based on whether or not Player 1 has looked at the state of nature in the previous stage. If Player 2 decides to exit the interaction, the game ends, and both players’ payoffs are zero. Only if Player 2 continues, we assume that there is a fourth stage in which Player 1 decides whether to cooperate (C) or to defect (D). Player 1’s cooperation decision may depend on the information she has up to that point (i.e., on what she has learned about the realized payoffs during the second stage of the game). The players’ payoffs are: C(a , b ) D(ci, di)(1) 1In addition, several authors have suggested proximate mechanisms to explain altruistic behaviors in social dilemmas. The corresponding models typically argue that people have social preferences, such as inequity aversion [8,9], a preference for fairness [10] or that people care for efficiency and social welfare [11,12]. While such models can explain certain patterns of human behavior in experimental games, these models do not explain how the presumed social preferences have evolved. Herein, we aim to propose an ultimate mechanism for intuitive cooperation. In particular, we stress that our payoff parameters do not reflect the players’ utilities from a certain outcome, but they directly correspond to the players’ material benefit. Our model then explores under which circumstances individuals learn to ignore payoff-relevant information and cooperate without looking.
Games 2015,6461 The first variable (aor ci) refers to the payoff of Player 1, whereas the second variable (bor di) corresponds to the payoff of Player 2. The values of ciand didepend on the state of nature i∈Ω. Without loss of generality, we can assume that payoffs satisfy: cl< ch di< b for i∈ {l, h}(2) The first inequality cl< chformalizes our interpretation that the state lcorresponds to a situation in which Player 1gets a low payoff from defection. The second inequality di< b, incorporates the assumption that Player 1’s move in the fourth stage can be interpreted as an act of cooperation: Player 2 would always prefer his co-player to choose C. After these four stages, the game is over. For this non-repeated envelope game to be interesting, we impose five further inequalities on the payoff parameters. The first two inequalities, ch> a > cl(3) ensure that Player 1 has no dominant action in Stage 4; in some situations cooperation is in her best interest, whereas in other situations, she is tempted to defect. The next two inequalities, pdl+ (1 −p)dh<0< b (4) prevent Player 2’s exiting decision in the third stage from being trivial. If pdl+ (1 −p)dh>0, Player 2 would continue even if Player 1 always chose D, whereas for b < 0, Player 2 would always exit. Although the final stage of the envelope game involves a decision between cooperation and defection, the primary aim of the model is not merely to explore when people will cooperate. Rather, we will be interested in situations in which Player 1 cooperates without looking, without even paying attention to whether cooperation is in her own interest. In order to allow for such acts of cooperation, we will furthermore make the final assumption that: a > 0(5) With this assumption, we ensure that Player 1 prefers a game that ends with cooperation to a game that ends by Player 2 exiting the interaction. 3. Static Analysis The non-repeated envelope game in Figure 1describes a simple sequential interaction, which we can solve by backward induction. The result is summarized in the following theorem; to allow for a convenient treatment of degenerate cases, the proof uses the tie-breaking rule that Player 1 cooperates when she is indifferent between C and D and that Player 2 continues when he is indifferent between continue and exit.
Games 2015,6462 Theorem 1. Depending on the probability p, backward induction implies the following behaviors along the unique equilibrium path: (ONLYL) If p≥− dh b−dh, Player 1 looks in the second stage, Player 2 continues in the third stage, and in the fourth stage, Player 1 cooperates if and only if her payoff from defection is low. (CWOL) If ch−a ch−cl≤p < −dh b−dh, Player 1 does not look in the second stage, Player 2 continues in the third stage, and Player 1 always cooperates in the fourth stage. (EXIT) If p< ch−a ch−cl, Player 2 exits the interaction in the third stage, independent of Player 1’s decision in the second stage. Proof. The proof is by backward induction. Optimal behavior in Stage 4: Suppose the interaction has reached Stage 4. Then, there are two cases, depending on whether or not Player 1 knows the state of nature: (a) if Player 1 has not looked at the state of nature in the second stage, she should cooperate if and only if pcl+ (1 −p)ch≤a(6) holds, such that cooperation yields at least the expected payoff from defection; (b) if Player 1 knows the state of nature, it follows from Equation (3) that she should cooperate when there is a low defection payoff, whereas she should defect when there is a high defection payoff. Optimal behavior in Stage 3: Again, we distinguish the previous two cases based on Player 1’s information: (a) without knowledge about the state of nature, Player 1 will cooperate in the fourth stage if and only if condition Equation (6) holds; in that case, it pays for Player 2 to continue in the third stage; (b) when Player 1 knows her own state, she will cooperate if and only if her defection payoff happens to be low. That implies that Player 2 continues if and only if pb + (1 −p)dh≥0,(7) that is, if the expected continuation payoff is above the exit payoff. Optimal behavior in Stage 2: Since there is no direct cost for acquiring information, Player 1 would always prefer to have as much information as possible (subject to the restriction that this does not lead Player 2 to exit the interaction). That is, if possible, Player 1 looks at the state of nature, which will be accepted by Player 2 if and only if Equation (7) holds (or, equivalently, when p≥ − dh b−dh). If looking at the payoffs is not feasible, Player 1 will either cooperate without looking (if Equation (6) holds, or, equivalently, p≥ch−a ch−cl), or Player 2 will exit the interaction irrespective of Player 1’s choice in the second stage (if p< ch−a ch−cl). Remark. In the following, let us make a few simple observations that follow immediately from Theorem 1. 1. The two equilibrium scenarios ONLYL and CWOL both allow for some cooperation along the equilibrium path. However, Player 2’s interpretation of an observed cooperative act of Player 1 will be different between these two scenarios. In an ONLYL equilibrium, Player 2 knows that Player 1 only cooperated because cooperation happened to be in Player 1’s own interest. Only in the CWOL equilibrium, Player 2 can be certain that Player 1 cooperates in any case; since Player 1 does not even care to learn the possible gains from defection.
Games 2015,6463 2. In a CWOL equilibrium, there may be instances in which Player 1 cooperates although realized payoffs do not make it profitable for Player 1 to do so. In such instances, an outsider who only observes the fourth stage of the game may interpret Player 1’s behavior as altruistic. Our model suggests that such seemingly altruistic acts occur because Player 1’s cooperation decision is only the final stage of a larger game. When the whole game is considered, Player 1’s cooperation behavior is part of an equilibrium. In fact, Player 1 cooperates in such instances because: (i) Player 2 would not accept a purely opportunistic co-player as an interaction partner (i.e., the general parameters of the game do not allow for an ONLYL equilibrium); and (ii) on average, it pays for Player 1 to engage in such interactions and to occasionally cooperate, even if cooperation may happen to be not in Player 1’s self-interest. 3. There is a strong analogy between the equilibrium conditions in Theorem 1 and the equilibrium conditions for the repeated-game setup considered in Hoffman et al. [1]. For the repeated game model, Hoffman et al. [1] find that Player 2 will exit in the first round of the game if payoffs satisfy p < (ch−˜a)/(ch−cl), with ˜a=a/(1 −w)being the expected payoff from cooperating in every round (wis the constant probability of having another round). If p > (ch−˜a)/(ch−cl), they observe the occurrence of an additional equilibrium, in which Player 1 cooperates without looking. The condition p < −d/(b−d), with d:= dl=dh, is incorporated as a basic assumption, which is introduced exactly to exclude the opportunistic ONLYL equilibria. 4. The equilibrium condition for CWOL, ch−a ch−cl≤p < −dh b−dh, suggests that CWOL can only be sustained if there is sufficient variability in Player 1’s temptation to defect. If it happens too often that Player 1’s payoff of defection is low (i.e., if p > −dh b−dh), then Player 2 would also accept a Player 1 who defects in the few cases in which she is tempted to defect, resulting in an ONLYL equilibrium. On the other hand, if it only happens occasionally that Player 1 has a low defection payoff (i.e., if p < ch−a ch−cl), then Player 1 who does not look at the state of nature would prefer to defect by default, forcing Player 2 to exit the interaction. In particular, the equilibrium conditions for CWOL are easier to satisfy when dhdecreases and when Player 1’s cooperation payoff aapproaches her maximum defection payoff ch. This observation suggests that CWOL is more likely to occur if occasional events of defection can be very harmful to Player 2, while only giving a negligible advantage to Player 1. 5. Whether CWOL can be obtained as an equilibrium also depends on the correlation between Player 1’s defection payoff ciand the resulting payoff dito Player 2. To see this in more detail, let us assume that dland dhtake the following form, dl=λd1+ (1 −λ)d2 dh= (1 −λ)d1+λd2 (8) with d1and d2being payoff parameters, such that d1< d2<0. The parameter 0≤λ≤1 is a measure for the correlation between the payoffs of the two players when Player 1 defects. When λ= 1, then dl=d1and dh=d2, and thus, our assumptions imply dl< dhand cl< ch. In this case, Player 1 gets a comparably low payoff from defection if and only if Player 2’s payoff after defection is comparably low (i.e., if the payoffs of the two players have a positive correlation). On the other hand, if λ= 0, then dl> dhwhile still cl< ch. This implies that when Player 1 is tempted to defect, defection is particularly harmful to Player 2 (i.e., there is a negative correlation
Games 2015,6464 between the players’ payoffs). Because dhis monotonically increasing in λand because −dh b−dh is monotonically decreasing in dh, it follows from Theorem 1 that CWOL is most likely to be an equilibrium when λ→0,i.e., when there is a negative correlation between payoffs. Intuitively, if there is a negative correlation, instances in which Player 1 would be tempted to defect are more harmful to Player 2, and thus, Player 2 has a stronger interest to prevent his co-player from looking. Finally, let us note that our previous conclusions are independent of the assumption that nature chooses among two states only. To see this, suppose there are nstates of nature, Ω = {ω1, . . . , ωn}, and let p= (p1, . . . , pn)be the corresponding probability distribution (such that p1+. . .+pn= 1). If Player 1 cooperates, the payoffs are a(for Player 1) and b(for Player 2), respectively. If Player 1 defects, the players’ payoffs again depend on the state of nature; they are given by ci(for Player 1) and di(for Player 2), with 1≤i≤n. For simplicity, let us assume that the states are ordered such that c1< c2< . . . < cnand such that there is a jwith cj<a<cj+1 with 1≤j < n. That is, in the first jstates, Player 1 has an incentive to cooperate, whereas in the remaining n−jstates, Player 1 would like to defect. Then, we can define pas the probability that nature chooses one of the first jstates, cl and dldenote the expected payoffs in that case, whereas chand dhdenote the expected payoffs given that nature has chosen one of the remaining n−jstates. Assuming that this extended model again satisfies Inequalities (2) and (4), one can take the same proof as in Theorem 1 to show that CWOL is an equilibrium if ch−a ch−cl≤p < −dh b−dh. 4. Evolutionary Simulations In the previous section, we have used backward induction to analyze when people will cooperate without looking. However, backward induction imposes certain assumptions on the players’ rationality and on the players’ trust in their co-players’ rationality [19]. For some applications, such a rationality-based view may appear inappropriate. In the following, we will thus argue that our previous conclusions remain valid even if we consider a simple evolutionary setup, in which individuals do not perform the calculations necessary for backward induction. Instead, they simply imitate others based on the relative success of their strategies [20–23].2 To this end, let us consider two populations of players, a population of individuals who play the non-repeated envelope game in the role of Player 1 and a population of individuals who play the game in the role of Player 2. We will refer to these populations as Population 1 and Population 2, with population sizes N1and N2, respectively. Members of the two populations are randomly matched to play the game, 2Of course, there has been extensive research on the relationship between backward induction and evolutionary equilibrium selection (see, e.g. [24–26]). For example, in Hart [26], it is shown that the backward induction outcome is the unique evolutionarily stable outcome if the stochastic dynamical process of mutation and selection satisfies certain reasonable assumptions. Seen from that angle, the results presented in this section may not come as a surprise. However, to show a correspondence between backward induction and evolutionary equilibrium selection, analytical models typically need to consider appropriate limits (such as the limit of rare mutation or the limit of infinite population sizes). The simulation results presented in this section can therefore be regarded as a robustness check: backward induction gives a reasonable prediction for the evolution of strategies in the envelope game even if populations are of moderate size and mutation rates are bounded away from zero, provided that selection is sufficiently strong to wipe out deleterious mutations.
Games 2015,6465 and they derive a payoff depending on their respective strategies. A strategy for Player 1 is a four-tuple (qL;q0, ql, qh) with qi∈ {0,1}. The first entry qLis the player’s probability to look in the second stage; the other entries q0,qland qhcorrespond to Player 1’s cooperation probability in the fourth stage (given that she does not know the state of the world, she knows the state is lor she knows the state is h, respectively). Overall, players in Population 1 can thus choose among 16 strategies.3Similarly, a strategy for Player 2 is a two-tuple (rL, rN), where ri∈ {0,1}is the probability to exit the interaction in the third stage depending on whether or not Player 1 has looked in the second stage, respectively.4 Players in Population 2 can thus choose among four possible strategies. The current composition of a population can be described by a vector, n1= (n(1) 1, . . . , n(16) 1)and n2= (n(1) 2, . . . , n(4) 2), where n(j) idenotes the number of players in population ithat employ strategy j. In particular, it follows that n(1) 1+. . . +n(16) 1=N1and n(1) 2+. . . +n(4) 2=N2. Initially, we suppose that individuals in Population 1 are non-looking defectors, whereas individuals in Population 2 always exit the interaction. However, the composition of the two populations is allowed to change over time, depending on the relative success of each strategy. Specifically, we consider a simple pairwise comparison process (see also [30–33]). In each time step, some individual ifrom one of the two populations is chosen at random and given the chance to update the strategy. To this end, another individual jis chosen from the same population to act as a potential role model. If expected payoffs of the two players are given by πiand πj(which depend on the current composition of the other population), then we assume that the focal individual iadopts j’s strategy with probability: ρ=1 1 + exp−β(πj−πi)(9) The parameter β≥0is called the strength of selection. It measures how strongly i’s updating decision is affected by relative payoff differences. In the extreme case β= 0, the imitation probability becomes ρ= 1/2(irrespective of the payoffs of the players), and imitation essentially occurs at random. In the other extreme case β→ ∞, the focal player ionly imitates co-players that have at least the payoff of player i(because πj< πiimplies ρ= 0). In addition to these imitation events, we also allow for random strategy exploration (corresponding to spontaneous mutation events in evolutionary biology). Specifically, we assume that in each time step, there is a probability µthat one of the individuals (in one of the two populations) is chosen at random. This individual then randomly switches to a different strategy (with all alternative strategies for 3We note that this parametrization of the strategy space allows for several distinct strategies that are behaviorally indistinguishable. For example, strategies of the form (1; 0, pl, ph)give rise to the same behavior as strategies of the form (1; 1, pl, ph); as Player 1 looks at the payoffs, it is irrelevant what she would do if she had no payoff information. We have chosen this payoff parametrization in order to decrease the bias in the mutation kernel (see also [27]): the number of strategies that prescribe Player 1 to look is the same as the number of strategies that prescribe Player 1 not to look. However, we stress that the simulation results presented in this section are robust; even if we only allowed for one strategy that gives rise to some particular behavior (and excluded all behaviorally-equivalent copies), our qualitative results would not change as long as selection is sufficiently strong. 4As is typical for models in evolutionary game theory, we assume here that individuals can only choose among the pure strategies [20,21,28]. However, an extension to the case of mixed strategies is straightforward (see, e.g., [29]), and since the non-repeated envelope game does not have any mixed equilibria, the results can be expected to be similar.
Games 2015,6472 36. Hardy, C.; Van Vugt, M. Nice guys finish first: The competitive altruism hypothesis. Personal. Soc. Psychol. Bull. 2006,32, 1402–1413. 37. McNamara, J.M.; Barta, Z.; Fromhage, L.; Houston, A. The coevolution of choosiness and cooperation. Nature 2008,451, 189–192. 38. Roberts, G. Competitive altruism: From reciprocity to the handicap principle. Proc. R. Soc. B Biol. Sci. 1998,265, 427–431. 39. Ghang, W.; Nowak, M.A. Indirect reciprocity with optional interactions. J. Theor. Biol. 2015,365, 1–11. 40. Pacheco, J.M.; Traulsen, A.; Nowak, M.A. Active linking in evolutionary games. J. Theor. Biol. 2006,243, 437–443. 41. Sasaki, T.; Uchida, S. The evolution of cooperation by social exclusion. Proc. R. Soc. B Biol. Sci. 2013,280, 20122498, doi:10.1098/rspb.2012.2498. 42. Frank, R.H. Passions Within Reason; W. W. Norton & Company: New York, NY, USA, 1989. 43. Nowak, M.A.; Tarnita, C.E.; Antal, T. Evolutionary dynamics in structured populations. Philos. Trans. R. Soc. B 2010,365, 19–30. 44. Santos, F.C.; Santos, M.D.; Pacheco, J.M. Social diversity promotes the emergence of cooperation in public goods games. Nature 2008,454, 213–216. c 2015 by the authors; licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution license (http://creativecommons.org/licenses/by/4.0/).