scieee AI-readable full text Open interactive document viewer

Private monitoring and communication in the repeated prisoner's dilemma

Awaya, Yu

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Awaya, Yu Article Private monitoring and communication in the repeated prisoner's dilemma Games Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Awaya, Yu (2021) : Private monitoring and communication in the repeated prisoner's dilemma, Games, ISSN 2073-4336, MDPI, Basel, Vol. 12, Iss. 4, pp. 1-10, https://doi.org/10.3390/g12040080 This Version is available at: https://hdl.handle.net/10419/257562 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ games Article Private Monitoring and Communication in the Repeated Prisoner’s Dilemma Yu Awaya   Citation: Awaya, Y. Private Monitoring and Communication in the Repeated Prisoner’s Dilemma. Games 2021,12, 80. https:// doi.org/10.3390/g12040080 Academic Editors: Andrew M. Colman and Ulrich Berger Received: 4 August 2021 Accepted: 11 October 2021 Published: 25 October 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). Department of Economics, University of Rochester and Institute of Social and Economic Research, Osaka University, Harkness 228, Rochester, NY 14627, USA; YuA[email protected] Abstract: This paper provides a model of the repeated prisoner’s dilemma in which cheap-talk communication is necessary in order to achieve cooperative outcomes in a long-term relationship. The model is one of complete information. I consider a continuous time repeated prisoner’s dilemma game where informative signals about another player’s past actions arrive following a Poisson process; actions have to be held fixed for a certain time. I assume that signals are privately observed by players. I consider an environment where signals are noisy, and the correlation of signals is higher if both players cooperate. We show that, provided that players can change their actions arbitrary frequently, there exists an equilibrium with communication that strictly Pareto-dominates all equilibria without communication. Keywords: repeated games; private monitoring; communication 1. Introduction In this paper, I identify an environment in which communication is necessary to achieve cooperation. I consider a model of the repeated prisoner’s dilemma game with complete information and private monitoring (a repeated game in which players can only get noisy signals on other players’ past actions, and moreover, the signals are private information of the player). I borrow the specifications of the model from Abreu, Milgrom and Pearce [ 1 ] (henceforth AMP)—this is a continuous time model in which actions have to be held fixed for a certain period, and a noisy signal arrives following a stochastic process. I depart from AMP, however, by assuming monitoring is private instead of public. It is assumed that the monitoring structure has the following properties. First, private signals are sufficiently noisy. Second, like Aoyagi [2] and Zheng [3] , the degree of their correlation depends on actions, and it is higher when both players cooperate than when one of them defects. Finally, regardless of action pair, the correlation is high enough so that given one player gets a signal, it is more likely that the other player also observes a signal. The main result of the paper is that there exists an equilibrium with communication that strictly Pareto-dominates all equilibria without communication, provided that the period over which actions are held fixed is sufficiently short. In this model, communication not only facilitates coordination, as in Compte [4] and Kandori and Matsushima [5] , but also helps players to find defects. To see this, notice that if the signals disagree across players, that is a sign of a defect. Signals that are too noisy to sustain cooperation if observed individually are informative enough to sustain cooperation when they indicate disagreement. Of course because of private monitoring, players cannot utilize this information. Communication, however, allows players to get information about the other players’ signals. Thus, this paper provides a channel through which communication facilitates cooperation. Note also that in our model, messages are not verifiable. This result suggests that antitrust authorities may have to be careful even when exchanging unverifiable messages. Lately, Awaya and Krishna [6 , 7] and Spector [8] also showed the necessity of communication. In particular, Awaya and Krishna [6] considered Stigler’s 1964 secret price cutting Games 2021,12, 80. https://doi.org/10.3390/g12040080 https://www.mdpi.com/journal/games Games 2021,12, 80 2 of 10 model where firms compete in prices and sales are drawn from a log-normal distribution. Awaya and Krishna [7] considered a general game with finite action and signal space. This paper is a precursor to these. AMP specification helps to state the result in a cleaner manner. In particular, unlike Awaya and Krishna [6 , 7] , this paper does not focus on the limiting case where signals are arbitrarily noisy. A remark is in order. While this paper shows that without communication, cooperation cannot be achieved, this result does not contradict folk theorem (see Sugaya [10] ). Folk theorem fixes other parameters, and works on the limit when players become very patient. In this paper, on the other hand, the patience of players is fixed. More broadly, this paper relates to the recent advances in the dynamic prisoner’s dilemma and social dilemma games that examine when cooperation can be achieved. For example, using the evolutionary approach, Szolnoki and Perc [11] and Danku et al. [ 12 ] pointed to importance of knowing others’ strategies to achieve cooperation. In this study, in contrast, communication was used to learn other’s past signals, not strategies. 2. Preliminaries 2.1. Repeated Game Consider the following prisoner’s dilemma as the stage game with expected payoffs, denoted by G: The time structure is borrowed from AMP. This prisoner’s dilemma game is played continuously over the time interval [ 0, ∞) . The payoffs in the matrix above are flow rates of payoff. The interest rate is denoted by r , and so if the payoff at time t is ut , then the net present value of payoffs for the whole game is Re−rtutdt . Notice r represents patience of players, and the smaller the ris, the more patient the players are. It is assumed that actions have to be held fixed for a certain time interval denoted by ∆∈( 0, ∞) . In other words, players can choose their actions only at periods 0, ∆ , 2 ∆ , . . . . In effect, the “discount factor” is interpreted as e−r∆. Like AMP, I assume that the signals arrive following a stochastic process which depends on the pair of actions. Unlike AMP, however, I assume that the signal is private. In particular, it follows the following process. Let zi∈ {s , o} represent arrival or absence of a signal to player i , where s (resp. o ) denotes arrival (resp. absence). Given an action pair a∈ {C , D} × {C , D} , the probability that each player gets a signal within a time interval ∆ is given by the following figure: s o s1−e−η(a)∆−1 21−e−γ(a)∆1 21−e−γ(a)∆ o1 21−e−γ(a)∆e−η(a)∆−1 21−e−γ(a)∆ where η(a) , γ(a)> 0. It is clear that each element is smaller than 1. Later I will present an assumption that guarantees each element of the figure is strictly positive when ∆→ 0. The time interval ∆ is assumed to be small enough so that the probability that a signal is observed more than once within ∆ is negligible. Finally, Figure 1 is stated in terms of expected payoffs. One interpretation is that the game ends with probability rand players get payoffs only after it ends. C D C1, 1 −l, 1 +g D1+g,−l0, 0 Figure 1. The payoff matrix. where g,l>0. Games 2021,12, 80 3 of 10 Notice that the marginal probability that a private signal arrives is given by Pr(zi=s|a) = 1−e−η(a)∆ In other words, η(a) is the arrival rate of a signal to a player. From this, we can interpret maxa,a0|η(a)−η(a0)| as a measure of noisiness of the signal. The private signal is noisier if maxa,a0|η(a)−η(a0)| is smaller. An extreme case is maxa,a0|η(a)−η(a0)|= 0. In this case, the signal is completely noise. Notice also that Pr(z16=z2|a) = 1−e−γ(a)∆ Hence, γ(a) is an arrival rate of an event in which the arrival of a signal differs among players. If γ(a) is small, then signals are highly correlated in the sense that if one player gets the signal, then the other players will also get the signal, most likely. I will assume that (i) private signals are noisy relative to the patience of players, (ii) the degree of their correlation depends on actions and it is higher when both players cooperate than when one of them defects, and (iii) regardless of action pair, the correlation is high enough so that given a player gets a signal, it is more likely that the other player also observes a signal. Formally, first Assumption 1. max a,a0|η(a)−η(a0)|<min{g,l} 1+g+lr This assumption states the relationship between the noisiness of the private signals and the patience of the players. Roughly, the assumption is easily satisfied when (i) the private signal is noisy and (ii) the players are impatient. Second, Assumption 2. g<min{γ(CD),γ(DC)} − γ(CC) r+γ(CC) This assumption implies that γ(CC)<min{γ(CD) , γ(DC)} , or the disagreement over arrivals happens more frequently when a player defects than when both players cooperate. To see the implication of this assumption, consider a fictitious case in which signals are public in the sense that both players could commonly observe the pair of signals (z1 , z2)∈ {s , o} × {s , o} , instead of their own signal zi . Then the event z16=z2 constitutes the arrival of bad news in terms of AMP—an event which occurs more frequently when a player defects. Lemma 1. lim ∆→0¯ vP(∆) = 1−g/min{γ(CD),γ(DC)} γ(CC)−1 Finally, assume Assumption 3. For any a, 1−1 2 γ(a) η(a)>1/2 η(a)>γ(a) This assumption means that given if a player gets a signal, it is likely that the other player also observes a signal. To see this, just observe that Pr z1=z2=s−Pr {zi=s} ∩ {z−i=o}=e−γ(a)∆−e−η(a)∆>0 Games 2021,12, 80 4 of 10 Additionally, notice that this assumption guarantees that the first element of the figure takes a strictly positive value when ∆→0. Note that similar assumptions—such as that signals based on past actions are highly correlated, and correlation is higher when they cooperate—appear in other contexts. See, for example, Fleckinger [13] and Awaya and Do [14]. I will provide some parametric examples that satisfy all these assumptions. Example 1. Let g =l=1and r =0.05. Additionally, let η(a) = 0.2 for any a and γ(a) = (0.01 if a =CC 0.1 otherwise It is clear that Assumption 1is satisfied, because right-hand side is 0. Assumption 3is clearly satisfied as well. To see Assumption 2, min{γ(CD),γ(DC)} − γ(CC) r+γ(CC)=0.09 0.06 =1.5 >1=g Example 2. Again let g =l=1, r =0.05, and γ(a) = (0.01 if a =CC 0.1 otherwise Now let η(a) =      0.2 if a =CC 0.205 if a =CD or DC 0.21 if a =DD To see that these satisfy Assumption 1, notice maxa,a0|η(a)−η(a0)|=0.01 and min{g,l} 1+g+lr=0.05 3w0.01667 The remaining assumptions follow by the same reason as above. Now I am ready to define histories and strategies. Let An i=∏ ν∈{0,1,...,n} {C,D} be the history of actions that player itook at periods 0, ∆, 2∆, ..., n∆, and Zt i=∏ ν∈{0,1,...,n} {s,o} be the history of signals that player i has observed by period n∆ . Recall that ∆ is assumed to be small enough so that the probability so that the signal is observed twice within ∆ is negligible. A pure strategy σifor player iis a sequence of functions (σ0 i,σ1 i, ...)where σ0 i:∅→ {C,D} and for n∈N, σn i:An−1 i×Zn∆ i→ {C,D} Games 2021,12, 80 5 of 10 where ∅ is the null history (throughout the paper, subscripts indicate players, superscripts indicate periods, and bold letters indicate histories). An implicit assumption here is that a player can condition his action on a signal he observes in a certain instance. This assumption is innocuous because this event happens with probability zero. 2.2. Repeated Game with Communication Next consider an infinite horizon repeated game version of the stage game in which players can communicate. They can send messages in a finite set Mi and let M=M1×M2 . Communication entails no cost so that it is “cheap talk”. All messages are observed publicly and without any error. Assume players send messages only when they decide their actions, that is, at 0, ∆ , 2 ∆ , .... The communication is done by an instance without taking any physical time, and when players decide their actions they can take the message sent at that moment into account. Formally, let Mn i=∏ ν∈{0,1,...,n} Mi be the history of messages that player isend at periods 0, ∆, 2∆, ..., n∆, and let Mn=Mn 1×Mn 2 A pure strategy in this game is a pair of µ∗ i,σ∗ i where µ∗ i=µ∗0 i,µ∗1 i, ... is a sequence of message strategies and σ∗ i=σ∗0 i,σ∗1 i, ...is a sequence of action strategies. Then µ∗0 i:∅→Mi σ∗0 i:M→ {C,D} and for n≥1, µ∗n i:An−1 i×Zn∆ i×Mn−1→Mi σ∗n i:An−1 i×Zn∆ i×Mn→ {C,D} Again it is assumed that a player can condition a message sent, or an action taken, on a signal observed at the instance. 3. The Result The main result of the paper is that communication is necessary to achieve cooperation. Theorem 1. There exists an equilibrium with communication that strictly Pareto-dominates all equilibria without communication, provided that players can change their actions in an arbitrarily small instant. The theorem is established by showing the following two propositions whose proofs appear in the next section. The first proposition states the negative result that the trivial equilibrium is the only equilibrium if communication is impossible. Proposition 1. Suppose communication is impossible. Then there exists a ∆> 0such that for any ∆<∆repetition of mutual defection is the only equilibrium. This proposition is a consequence of an assumption that signal is too noisy. I provide ∆for the examples discussed before. Example 3 (Continuation of Example 1) . Take parameter values provided in Example 1. Then the trivial equilibrium is the only equilibrium for any ∆ . In other words, ∆ can be any value. Games 2021,12, 80 6 of 10 The reason here is clear. As maxa,a0|η(a)−η(a0)|= 0, the signal generated by playing C is exactly the same as playing D . In other words, there is no way to detect defection, even statistically. Hence, it is rational to play the stage-game dominant strategy, D. Example 4 (Continuation of Example 2) . Take parameter value provided in Example 2 and let ∆= 0.01. Then it is shown that the trivial equilibrium is the only equilibrium. The argument is found in the next section. The second proposition states the positive result that a better outcome (than the trivial equilibrium) can be achieved in the equilibrium if communication is possible. Proposition 2. Suppose communication is possible. Let ¯ v(∆) be the supremum of the symmetric equilibrium payoff when players can change their actions at time interval ∆. Then lim ∆→0¯ v(∆)≥1−g min{γ(CD),γ(DC)} γ(CC)−1 >0 The rough idea of the proof is as follows. As is already mentioned, the disagreement of the signal constitutes bad news. Now consider a fictitious public monitoring game in which each players could observe signals of both players. Then, Proposition 5 of AMP applies here, and it is shown there exists a non-trivial equilibrium in the limit as ∆→ 0. Now the remaining task is to show truth telling. It is shown that given an arrival (resp. absence) of a signal to a player, it is more likely that an arrives (resp. absence) of the signal happens to the other player regardless of action pair. This means that truth telling minimizes the probability of an arrival of bad news, which is disagreement over the arrival times of signals. Example 5 (Continuation of Examples 1 and 2) . Given parameter values provided Examples 1 and 2, lim ∆→0¯ v(∆)≥1−1 0.1 0.01 −1 w0.8889 4. Proofs 4.1. Without Communication We first establish that without communication, players cannot cooperate. The intuition is that because signals only contain noisy information about past behavior of the opponent (Assumption 1), deviations are hardly detected and so will be left unpunished. By way of contradiction, suppose there is an equilibrium in which at least one player— say player i —cooperates at some histories. Without loss of generality, suppose player i cooperates in period n∆ following after a history hi . For this to be rational, player i ’s future payoff must non trivially depend on whether −i gets a signal or not during (n∆ , (n+ 1 )∆] . Now let vi(s) be the continuation payoff of player i when player −i gets a private signal in (n∆ , (n+ 1 )∆] and vi(o) be the continuation payoff of player i when player −i does not. Of course, these expected continuation payoffs depend on history hi . As vi(o) and vi(s)are continuation payoffs, they are bounded, and in particular, |vi(o)−vi(s)|≤|vi(o)|+|vi(s)| ≤ 1+g+l Take an action pair a∈ {C , D} × {C , D} . Then the lifetime payoff of player i given a is given by (1−e−r∆)u(a) + e−r∆[e−η(a)∆vi(o) + (1−e−η(a)∆)vi(s)] Games 2021,12, 80 7 of 10 where u(a) is the stage-game payoff given an action pair. Let the value be denoted by V(a , ∆) . Notice for given player −i ’s strategy σ−i , the lifetime payoff when he takes an action aiis given by Eσ−iV(ai,a−i,∆) where the expectation is taking over player −i’ action a−i, following his strategy σ−i. In the following, I show that for any σ−i, lim ∆→0 Eσ−iV(D,a−i,∆)−V(C,a−i,∆) ∆>0; that is, in contrast to the assumption, it is not rational to play C. To see this, first define W(∆) = V(D,a−i,∆)−V(C,a−i,∆) Then, W(∆)=(1−e−r∆)[u(D,a−i)−u(C,a−i)] +e−r∆e−η(D,a−i)∆−e−η(C,a−i)∆(vi(o)−vi(s)) and hence, lim ∆→0 W(∆) ∆=r(u(D,a−i)−u(C,a−i)) + (η(D,a−i)−η(C,a−i))(vi(s)−vi(o)) This value takes a strictly positive value by the following reason. First observe that for any a−i, u(D,a−i)−u(C,a−i)≥min{g,l} Second, (η(D,a−i)−η(C,a−i))(vi(s)−vi(o)) ≤ |η(D,a−i)−η(C,a−i)||vi(s)−vi(o)| ≤max a,a0|η(a)−η(a0)||vi(o)−vi(s)| ≤max a,a0|η(a)−η(a0)|(1+g+l) By combining these inequalities, I have lim ∆→0 W(∆) ∆≥rmin{g,l} − max a,a0|η(a)−η(a0)|(1+g+l) The right-hand side is strictly positive by Assumption 1. Now the claim follows because lim∆→0(V(D,a−i,∆)−V(C,a−i,∆))/∆is bounded away from 0 for any a−i. Now consider the following example again. Example 6 (Continuation of Examples 2 and 4).First notice that W(∆)≥(1−e−r∆)[u(D,a−i)−u(C,a−i)] −e−r∆e−η(D,a−i)∆−e−η(C,a−i)∆|vi(o)−vi(s)| ≥(1−e−r∆)−3e−r∆e−η(D,a−i)∆−e−η(C,a−i)∆ Games 2021,12, 80 8 of 10 where the last inequality follows from the facts that min a−i [u(D,a−i)−u(C,a−i)] = min{g,l}=1 and max|vi(o)−vi(s)|=1+g+l=3 Now because ∆=0.01, e−r∆=.999500125. Also, max a−ie−η(D,a−i)∆−e−η(C,a−i)=e−η(D,D)∆−e−η(C,C)∆=.0000997952 Thus, W(0.01)≥0.000200639 >0 4.2. With Communication Next we establish that players can cooperate if communication is available. Here, communication plays a role to “aggregate” signals. Correlation across signals are higher when they cooperate (Assumption 2). Thus, by aggregating signals and checking correlation, players can learn more about their past behavior. Thus, with communication, deviations are easily detected and punished. The outline of the proof is as follows. First, consider the fictitious case where each player can observe the signals of both players (z1 , z2) . Recall that Assumption 2implies that disagreement in signals z16=z2 is bad news that occurs more likely when a player plays D . This fictitious game is one of the public monitoring games that AMP considered with the interpretation that z16=z2 is bad news, and they established that some degree of cooperation can be achieved. This part is stated in Lemma 2. Note also that this game is equivalent to the original game with communication: each player follows truth telling—that is, sends a message that says the player observes a signal if and only if he actually does—with disagreement m16=m2being bad news. Then the statement is established by showing that players have incentives to report truthfully. This is established in Lemma 3. 4.2.1. Fictitious “Public Monitoring” Game Consider again a fictitious public monitoring game in which each player can observe the signals of both players (z1 , z2) . Let ¯ vP(∆) be the supremum of the symmetric perfect public equilibrium payoff of this fictitious game (following AMP, I use “symmetric” in the sense that the equilibrium specifies symmetric actions after all histories). Now Assumption 2implies that the first case of Proposition 5 of AMP applies here (with the interpretation that z16=z2 is bad news), and hence the following lemma is an immediate corollary of the proposition. Lemma 2. lim ∆→0¯ vP(∆) = 1−g min{γ(CD),γ(DC)} γ(CC)−1. 4.2.2. Truth Telling I complete the proof by showing that players have an incentive to report truthfully. As a disagreement over arrivals of signals constitutes bad news, each player tries to minimize the probability of disagreement when he sends a message. Now I will establish that truth telling minimizes the probability of the disagreement given that the other player also tells the truth, regardless of action pair.