scieee AI-readable full text Open interactive document viewer

Cooperation through image scoring: A replication

Russell, Yvan I.,Stoilova, Yana,Dosoftei, Aura-Adriana

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Russell, Yvan I.; Stoilova, Yana; Dosoftei, Aura-Adriana Article Cooperation through image scoring: A replication Games Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Russell, Yvan I.; Stoilova, Yana; Dosoftei, Aura-Adriana (2020) : Cooperation through image scoring: A replication, Games, ISSN 2073-4336, MDPI, Basel, Vol. 11, Iss. 4, pp. 1-15, https://doi.org/10.3390/g11040058 This Version is available at: https://hdl.handle.net/10419/257476 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ games Article Cooperation through Image Scoring: A Replication Yvan I. Russell *, Yana Stoilova and Aura-Adriana Dosoftei Department of Psychology, Middlesex University, London NW4 4BT, UK; y[email protected] (Y.S.); [email protected] (A.-A.D.) *Correspondence: yvanr[email protected] Received: 11 August 2020; Accepted: 25 November 2020; Published: 30 November 2020   Abstract: “Image scoring” is a type of social evaluation, originally used in agent-based models, where the reputation of another is numerically assessed. This phenomenon has been studied in both theoretical models and real-life psychology experiments (using human participants). The latter are aimed to create conditions in the laboratory where image scoring can be elicited. One influential paper is that of Wedekind and Milinski (2000), WM. Our paper is a replication of that study, deliberately employing very similar methodology to the original. Accordingly, we had six groups of ten participants play an economic game. In each round, each player was randomly paired with another player whose identity was unknown. The participant was given a binary choice of either (1) donating money to that person, or (2) not donating money. In each round, the player was passively exposed to information about the past generosity of the other player. In our study, we successfully replicated the central result of WM. Participants in our replication gave significantly more money to partners with higher image scores (more generous reputations) than those with lower image scores (less generous reputations). This paper also provides a critical review of the methodology of WM and the study of image scoring. Keywords: cooperation; reputation; psychology; replication 1. Introduction “Reputation” refers to knowledge about the typical behavior of an observed individual [ 1 – 3 ]. A numerical index of reputation, called an “image score”, was created in theoretical studies (agent-based modelling) that aimed to investigate the evolution of cooperation [ 4 ]. The image score scale ranged from –5 to +5 (including a neutral zero), allowing agents to accumulate a negative score for selfishness and a positive score for generosity. The motivation behind research in this area was the search for solutions to the “tragedy of the commons”, a common phenomenon whereby a useful or necessary public good is diminished (perhaps fatally) because of overconsumption by self-interested parties [ 5 ]. The evolution of cooperation (that which would solve the tragedy) requires a complex and multi-faceted explanation [ 6 ]. Indirect reciprocity [ 7 , 8 ] is just one proposed solution to the tragedy of the commons, expressed in positive form as “A observed B help C, therefore A helps B” [ 2 ]: generosity becomes self-serving because the generosity itself might be rewarded [ 1 , 7 ]. A substantial number of agent-based models have been published [ 8 ] that explore the conditions that allow indirect reciprocity, notably that of Nowak and Sigmund [ 4 ], who found that a discriminating “image scoring” strategy (give only to those who have given to others) was an evolutionarily stable strategy in contrast to undiscriminating/selfish strategies. Following the Nowak and Sigmund model [ 4 ], Wedekind and Milinski [ 9 ] (hereafter, WM) published a short paper in Science titled “cooperation by image scoring in humans”. This study converted the computer simulation into a psychology experiment, using actual human participants (instead of “agents” in a model). The results of this human study appeared to provide support for the conclusions of the Nowak and Sigmund model [ 4 ] because the human participants appeared to Games 2020,11, 58; doi:10.3390/g11040058 www.mdpi.com/journal/games Games 2020,11, 58 2 of 15 act in accordance to the predictions of the model. In the twenty years since WM [ 9 ] was published, the general methodology of the human studies have become known as the indirect reciprocity game (IRG). In our new study, we adopt the basic design of the IRG. In doing so, we model our study design most directly on that of WM. Below, we describe the IRG as we implement it in our current study. The IRG entails that a group of participants are brought into a group and given a sum of money. Then, the experimenter proceeds to engineer pairwise encounters between participants (a different pairing in each round). Upon pairing, the donor and receiver proceed to participate in a version of the dictator game [ 10 ], where the donor makes a unilateral decision about whether to donate money to the receiver. Participants never choose their own role (donor or receiver). The role of a player is never fixed, however, because it can change from round to round (often, the role is randomly scheduled by the computer program). Crucial to the design of the IRG is that the possibility of direct reciprocity is removed. Players cannot take revenge, nor reward a good deed. There are two ways that direct reciprocity is prevented. One way is to tell players that they will never be identifiably paired with the same player in reversed roles. Another way is that the identities of the other players are hidden. Even though participants are often sitting in the same room together, they do not know with whom they are being paired in a given round. Because direct reciprocity is not possible, players are consequently guided toward the perception that they are playing a string of one-shot games. Donors know nothing about the receiver except for a numerical measure of image score. The image score is an index of a donor’s generosity in previous rounds. Players who choose to donate to another player are awarded one point. Players who decline to donate lose one point. Just like in theoretical models (e.g., [ 4 ]), a high score connotes generosity and a low score connotes selfishness. At the end of the game, participants are paid according to how much money is in their account. Participants had been allocated a sum at the beginning of the study. During the game, money was very likely added to the account (due to being generous) or subtracted from the account (due to being selfish). IRG is also called a “donor game” by some authors. In the following, some figures from WM are specifically mentioned. To avoid confusion with figures in our current paper, the figures from WM are prefixed by “WM-” (e.g., WM-Figure 2 refers to Figure 2 in their paper). In the WM [ 9 ] version of the IRG, they tested seventy-nine human participants. Their main result was that reputation played a significant role in the participants’ decisions about whom to reward. Participants who received more money had higher image scores (for more details see Table 1and discussion below). This result was interpreted as a demonstration of the importance of reputation as an explanatory factor in why we cooperate. The WM paper [ 9 ] can justifiably be described as influential, because, to date, it has been cited more than 900 times. We carefully combed over these citations and found that the vast majority of the citations were brief and uncritical, used as theoretical background in literature reviews (e.g., in Ref. [11]). Among the 900+citations, we found only nine papers [12–20] that we considered to be replications. All of them successfully replicated WM. Table 1is a summary of these replications, compared against the original WM paper (Ref. [ 9 ]) and compared against our own replication (bottom row). Five of the replications were co-authored by the same people who did the original study [ 12 – 16 ] and four of the replications were from other groups [ 17 – 21 ]. There were two criteria that we used for including papers in Table 1. The first was that the replication directly cited WM [ 9 ]. The second criterion was that the replication adopted the basic design of the IRG (as described in the previous paragraph). This excludes empirical studies of indirect reciprocity which did not use the IRG (e.g., studies which used the Prisoner’s Dilemma game, PDG, instead). As shown in Table 1, the replications of WM should probably be called “partial replications”, or “conceptual replications”. None of those studies used exactly the same methodology, nor the same analysis. This was true even for those with the same authors as the original [ 12 – 16 ]. Looking at the aims, methods, and results columns, it is clear that all replications were using the basic IRG design as a core procedure, but adding new conditions, additional new games (e.g., alternating IRGs with non-IRGs), and new hypotheses (e.g., Ref. [ 19 ] used the IRG paradigm to study strategic reputation building, something that WM had not investigated). In Table 1(third and fourth columns), one can also see great variation in the Games 2020,11, 58 3 of 15 parameters. The mean sample size (refs. [ 12 – 20 ]; n=10 studies) was 128.8 participants (std. dev. =55.09), median 114, range 79–228. The mean group size (students playing together; n=134 groups) was 8.45 participants (std. dev. =3.18), median 7.00, range 4–16. The mean number of rounds of the IRG (n=10 studies) was approximately 32 rounds (std. dev. ≈ 31.2) (it was not possible to calculate this precisely due to randomized game-endpoints in Ref. [ 18 ]). Also shown in Table 1(fourth column) is the length of the history of interactions that allows the donor to evaluate the image score of the receiver. The history length increases as the game proceeds. At the beginning of the game, the history length is zero. As shown in Table 1, the studies took two main approaches. Some studies [ 14 , 16 – 20 ] censored the history length to a maximum number of rounds. In the studies that did this, the mean maximum history length was approximately five rounds. Other studies [ 12 , 13 , 15 ] appeared to maintain a history length through the duration of the game. Therefore, a study like Ref. [ 12 ], for example, had 16 rounds, allowing a history length of 0–15 rounds (caveat: some of the papers did not mention history length, and therefore the inference was made that the history consisted of all rounds minus one). Research using the specific WM design lasted only a few years. Most of the studies in Table 1 were conducted around the early 2000s (this includes refs. [ 17 – 19 ], which were also conducted around the early 2000s despite being published years later). Only Ref. [ 20 ] was conducted much later than the rest. There have been more recent (post-2015) developments on the study of reputation and indirect reciprocity, but these have branched out into new and different methods (we review more recent papers in our Discussion below). In our laboratory, we chose to replicate the original WM [ 9 ] paper. We did this because all of the prior replications [ 12 – 20 ] lack the same analysis as WM. The main claim of WM is based on a somewhat unusual means of transforming the data. The raw image score of each player was converted into a “deviation” score. This measured how much higher or lower that player’s image score was compared to the group average on the given round. WM justified this approach as useful “to correct for group and round effects” (Ref. [ 9 ], p. 851). No further explanation is given, but presumably this refers to group-specific and round-specific confounds that influence the rate of giving. Following this transformation, there was another unusual aspect, concerning the analysis which supports the key result. The unusual part is not the statistical procedure (a standard repeated measures ANOVA was used), but the way that the data were set up. The focus was on the donor’s perspective. During the course of the game, the donor had a number of opportunities to donate money. In each round, the question for the donor was always a YES or NO. In a given round, the donor is paired with a receiver, and the receiver’s image score is shown. The donor’s choice is whether or not to donate money to that receiver. The theoretical expectation in WM [ 9 ] was that positive image scores will be rewarded (YES) and negative image scores punished (NO). However, in the game, the donor was free to violate this expectation. It was completely possible to unjustly decide YES for a negative image score and NO for a positive one. Whatever the case, WM splits the donor’s data into two columns: mean image score for (1) recipients of all YES decisions, and (2) all NO decisions. We believe that they did it this way because it allowed the binary YES/NO choice to be applied directly to the image score as encountered by the donor in a given round. WM did not report the descriptive statistics for the YES column and the NO columns, but the repeated measures ANOVA showed that those who benefited from YES decisions had significantly higher image scores than those who suffered NO decisions. WM’s approach, which focused on “type of donors’ decision” (Ref. [ 9 ], p. 851) was quite different from that in the replications that followed. None of the replications [ 12 – 20 ] did it that way. Yet, all of them were inspired by WM and copied their basic IRG paradigm. In our replication of WM below, we did our best to adhere to the specifics of the original analysis. We felt it was important to confirm that the original is replicable. Games 2020,11, 58 4 of 15 Table 1. List of replications, plus original and current study. Abbreviations: WM =Wedekind and Milinski; IRG =indirect reciprocity game; PGG =public goods game; PDG =prisoner’s dilemma game; CAG =competitive altruism game; UNICEF =United Nations Children’s Fund. Ref. Author(s) (Year) Overall n (Number of Groups × Group Sizes) Rounds (Shown History) Aims and Methods Result [9] WM (2000) (original study) 79 (7 ×10, 1×9) 6 (0–5) Human experiment to investigate effects shown in Ref. [4]. Each round gave the participant one opportunity to donate and two opportunities to receive. Independent variable was a corrected version of receiver image score, measured not as raw value but as deviations from the group mean in a given round. “. . . the image score of the receivers who were given money... was on average higher than the score of those who got nothing” (Ref. [9], p. 851). Compared to more generous players, less generous players were more discriminating, giving more to recipients with higher image scores. [12]Milinski et al. (2001) 161 (23 ×7) 16 0–15) IRG was used as control group in study about the usefulness of additional, “standing strategy” information (e.g., when a player had declined to donate in an instance where recipient was not deserving). Statistical unit was group, not participant. Image scoring was successful (analogous results to Ref. [9]), but the standing strategy was not successful. Analysis based on donations to “NO” player (confederate who never donated). [13]Milinski et al. (2002) 72 (12 ×7) 16 (0–15) The donor, after making the yes/no decision to donate, was given an additional question of whether to donate money to UNICEF. Reputational information was shown for both decisions. Each group had confederates: “always yes” and “always no” players. Both types of generosity (giving to other players/giving to UNICEF) tended to be rewarded, including in results of a mock-election at the end of the game to vote for other players for student council. [14]Milinski et al. (2002) 114 (19 ×6) 20 (0–7) Information about the generosity of other players was derived from a PGG which alternated with an IRG in first sixteen rounds. In the first treatment, the games alternated. In the second treatment, eight PG games were followed by eight IRG. Last four rounds consisted of PGG only. Players who were more generous in the PGG received more money in IR game. Final PGG showed very high cooperation if players uncertain about whether future IRGs would occur. Showed that concern for constant reputation monitoring increased cooperation. [15] Wedekind and Brathwaite (2002) 114 (6 ×9, 6×10) 24 1 (0–23) Investigated the relation between direct and indirect reciprocity. Each group played three games: PDG, IRG, then PDG again. Players who were more generous in the IRG received more money (as in Ref. [ 9 ]) but also in the subsequent PDG. Games 2020,11, 58 5 of 15 Table 1. Cont. Ref. Author(s) (Year) Overall n (Number of Groups × Group Sizes) Rounds (Shown History) Aims and Methods Result [16] Semmann, Krambeck, and Milinski (2005) 228 (19 ×12) 16 2 (0–5) In a similar design to Ref. [14], information about the generosity of other players was derived from a PGG which alternated with an IRG. Statistical unit was group, not participant, in most analyses. Players who were more generous in the PGG tended to receive more money in the IRG, showing that reputational information transfers between games and groups. [17] Bolton, Katok, and Ockenfels (2006) 192 (16 ×12) 14 (0–1) Had similar aims to Refs. [12,18]. Manipulated cost (high/low) and type of information. For the latter, they distinguished between first-order (image score) and second-order information (cf. standing strategy). Contrary to Ref. [12], giving was highest in response to second-order information (compared to first-order and a no-information control condition). Giving is higher in the low-cost condition. [18]Seinen and Schram (2006) 168 (12 ×14) 90+ (0–6) Had similar aims to Refs. [12,17]. Manipulated cost of giving (high/low) and information (information about past generosity in high/low cost condition, no-information about past generosity in high cost condition only. Number of rounds were designed to be unpredictable (min. 90, avg. 99). Dependent variable was the fraction of helping behavior (from 0–1). More donations occurred in the information condition and when the cost was low. Players with best image score tended to receive more money. Individual strategies were partitioned into six categories. [19] Engelmann and Fischbacher (2009) 80 (5 ×16) 80 (0–5) A study of strategic reputation building, comparing “public” (image score seen by all) and “private” (image score not seen) conditions. Half the participants had public scores in the first 40 rounds and private scores in the last 40 rounds (and vice versa for the other half). Image scoring was successful. Contributions were higher when image scores were public compared to private, suggesting that participants altered their behavior in response to being observed. Analysis of individual results showed a mix of apparent strategies among players. [20]Sylwester and Roberts (2013) 3 80 (20 ×4) 30 (6+) A study of reputation building in both direct and indirect contexts. PGG was alternated with a IRG (for half the sample) and alternated with a CAG (for the other half) (there were two types of alternating scheme: “one-shot” and “iterated”). In the CAG, players could choose each other in advance for a directly reciprocal game. Analysis based on groups rather than individuals. Image score was derived from the PGG (rather than in IRG play). The IRG was successful, but was not the main focus of the paper. Generosity was higher overall (including in PGG) for CAG than IRG. Participant showed an immediate reaction to the introduction of reputational incentives (causing them to contribute more). Games 2020,11, 58 6 of 15 Table 1. Cont. Ref. Author(s) (Year) Overall n (Number of Groups × Group Sizes) Rounds (Shown History) Aims and Methods Result - Russell, Stoilova, and Dosoftei (current paper) 60 (6 ×10) 12 (0–5) Replication of WM [9], using the same horizon (0–5 rounds) of reputational information and the same group sizes, but smaller overall sample. Unlike in Ref. [9], but like Ref. [19], not every player played in every round. However, on average, the randomization allowed six recipient rounds per player, as in Ref. [9]. Image scores were displayed only to donor in a given pairing (like private condition in Ref. [19], but unlike the displays of most other refs). The main result of WM [9] was replicated, using the analysis where image score was measured as deviations from the group mean in a given round, contrasting those who received a high amount versus a low amount of money. See main article for details of methodology and further analyses. 1. This excludes PDG rounds played before and after the IRG rounds. 2. This includes six PGG rounds alternating with IRG rounds (included because PGG contributed to reputation). 3. Earlier versions of this study are reported in Ref. [21]. Games 2020,11, 58 7 of 15 We know that some readers may question the value of replicating such an old paper. Our response is that, in the context of the “replication crisis” in psychology [ 22 ], replications can be considered as opportunities to look forward as well as backward [ 23 ]: replications can lead to the development of new methods for paradigms that perhaps need to be rethought. Furthermore, each new replication adds a data point to quantitative analyses of the success/failure rates of replications [ 22 ]. Finally, we can also point to equivalent replication programmes in other areas of psychology. The 1963 Milgram [ 24 ] study of obedience is a good example. That classic paper has been replicated umpteen times as a means of probing the nature of obedience as thoroughly as possible [ 25 ]. Replications of Milgram [ 24 ] continue to the present day, utilizing the most modern methodologies and theoretical perspectives [ 26 ]. We think that the WM study [ 9 ]—like the Milgram study [ 24 ]—deserves to be explored further, allowing us, in future, to probe the effects of reputation as thoroughly as possible. 2. Results In our replication, the main result of WM [ 9 ] was replicated. Data files and other material are viewable in the Supplementary Material (Documents S1–S12). In summary, we found that players provided YES decisions preferentially to receivers who had higher image scores in contrast to the receivers who received NO decisions. Our methods differed in several small ways from WM ( see Methods for details ), but produced directly comparable results. Figure 1is modelled directly on WM-Figure 2. As shown, there were six groups of players. To conduct an equivalent analysis to WM (as described in our introduction), we measured image score as individual deviations from the mean image score for the group and round that a player is in. Thus, if the group mean for in a given round were +0.2, but an individual’s raw image score were − 1 in that round, then that individual’s adjusted image score is − 1.2. The overall mean for adjusted image score was 0.00 (std. dev. =2.90). We then calculated two other variables: image scores of individuals who (1) received and (2) did not receive a donation in a given round (i.e., YES-donate, NO-donate). The first step of this calculation was to identify the recipient in every round and determine their adjusted image score at the beginning of the round (i.e., the image score perceived by the donor before making a choice). In the data file (Document S9) where the rows were participants, the image score of the recipient would be inputted into the row of the donor. For example, if participant 47 was the donor in round 10 and participant 46 was the receiver, then participant 46 0 s image score was inputted into participant 47 0 s row as the image score of the receiver from the point of view of participant 47 in that round. Thereafter, for every player, two variables were created that were the mean adjusted image scores of all recipients to which a player had (1) donated, or (2) not donated. The overall mean adjusted image score for recipients (from the point of view of the donor) was 0.92 (std. dev. =0.97; range –2.00 to +2.90; n=58) for those who received donations, and − 0.46 (std. dev. =1.47; range -4.20 to +2.10; n=42) for those who were declined donations. Because this was a comparison of two separate variables, the analysis below included only those players who were seen to make both decisions (YES-donate; NO-decline donation) during their play. Players who were all-yes (n=8) and all-no (n=3) were therefore excluded. A repeated measures ANOVA (the same analysis as WM) was then conducted with donors as replicates using the two aforementioned variables (if the donor gives or not) as the within-subjects variables and with groups of individuals as between-subjects factors. The effect of image score in giving/not giving was significant, F(1, 34) =6.563, p=0.015, η2p =0.1618, but there was no significant effect of group, F(5, 34) =0.346, p=0.881, η2p=0.0484, and no significant interaction, F(5, 34) =1.093, p=0.382, η2p=0.1384. Games 2020,11, 58 8 of 15 Games 2020, 11, x FOR PEER REVIEW 8 of 15 Figure 1. Image score of receivers when money given (black bars) or not (white bars). We decided to perform an alternative analysis, to confirm that the WM result was not an artefact of their choice of analysis. WM had partitioned receiver’s image score into two separate variables, according to whether they were (1) YES versus (2) NO decisions. In contrast, we decided to merge receiver’s image score into a single variable. We constructed a logistic GLMM (generalized linear mixed model) using donor trial as repeated measures and a binomial distribution with a logit link. The donor trial referred to every instance that a player was assigned the donor role by the computer program (these varied across participants; see Methods). The opportunity to donate was randomly assigned, not fixed to a particular round. For example, looking at the data file, we see that participant one had five opportunities to donate and these occurred in rounds 2, 5, 7, and 8. Player two, in contrast, had four opportunities to donate and these occurred in rounds 3, 4, 10, and 12. All players had a different pattern in this regard. That is why the donor trial is differentiated from the game round. There was a significant positive correlation between receiver image score (original) and the donor trial, rho = 0.214, p < 0.001, which shows the accumulative nature of the image score (the same correlation was not significant using the adjusted image score). The binary target (dependent variable) was the yes-no decision about whether or not to donate money to the recipient in a given round (hereafter called “decision”). This new analysis required the construction of a new data file where each row in an SPSS file recorded a player’s decision in a given round. This created a file with 330 decisions made by sixty players. A player’s group membership and ID number was recorded in columns, allowing a data structure where participants were nested within their testing groups. The random effect was therefore the testing group (n = 6). In the datafile, we made two exclusions. One is that we excluded all decisions from the first round. We did this because it was too early to accumulate an image score in the first round. Second, we excluded all-YES and all-NO players (leaving n = 49). The number of exclusions was therefore the same as in the repeated measures ANOVA reported above. The results below were not significant unless these exclusions were applied. The fixed effect (independent variable) was the image score of the recipient in a given round. We ran the test using both versions of the receiver’s image score (original and adjusted). The result was significant for the adjusted image score only, BIC (Bayesian Information Criterion) = 895.783, fixed effects F = 4.767, df = 192, p = 0.030, fixed co-efficient = 0.178 ± 0.082, t = 2.183, p = 0.030. The co-variance parameters for each donor trial (from second to ninth opportunity) were: Figure 1. Image score of receivers when money given (black bars) or not (white bars). We decided to perform an alternative analysis, to confirm that the WM result was not an artefact of their choice of analysis. WM had partitioned receiver’s image score into two separate variables, according to whether they were (1) YES versus (2) NO decisions. In contrast, we decided to merge receiver’s image score into a single variable. We constructed a logistic GLMM (generalized linear mixed model) using donor trial as repeated measures and a binomial distribution with a logit link. The donor trial referred to every instance that a player was assigned the donor role by the computer program (these varied across participants; see Methods). The opportunity to donate was randomly assigned, not fixed to a particular round. For example, looking at the data file, we see that participant one had five opportunities to donate and these occurred in rounds 2, 5, 7, and 8. Player two, in contrast, had four opportunities to donate and these occurred in rounds 3, 4, 10, and 12. All players had a different pattern in this regard. That is why the donor trial is differentiated from the game round. There was a significant positive correlation between receiver image score (original) and the donor trial, rho =0.214, p<0.001, which shows the accumulative nature of the image score (the same correlation was not significant using the adjusted image score). The binary target (dependent variable) was the yes-no decision about whether or not to donate money to the recipient in a given round (hereafter called “decision”). This new analysis required the construction of a new data file where each row in an SPSS file recorded a player’s decision in a given round. This created a file with 330 decisions made by sixty players. A player’s group membership and ID number was recorded in columns, allowing a data structure where participants were nested within their testing groups. The random effect was therefore the testing group (n=6). In the datafile, we made two exclusions. One is that we excluded all decisions from the first round. We did this because it was too early to accumulate an image score in the first round. Second, we excluded all-YES and all-NO players (leaving n=49). The number of exclusions was therefore the same as in the repeated measures ANOVA reported above. The results below were not significant unless these exclusions were applied. The fixed effect (independent variable) was the image score of the recipient in a given round. We ran the test using both versions of the receiver’s image score (original and adjusted). The result was significant for the adjusted image score only, BIC (Bayesian Information Criterion) =895.783, fixed effects F=4.767, df =192, p=0.030, fixed co-efficient =0.178 ± 0.082, t=2.183, p=0.030. The co-variance parameters Games 2020,11, 58 15 of 15 29. Christensen, L.B. Experimental Methodology, 5th ed.; Allyn & Bacon: Boston, MA, USA, 1991. 30. Schram, A. Artificiality: The tension between internal and external validity in economic experiments. J. Econ. Methodol. 2005,12, 225–237. [CrossRef] 31. Okada, I.; Yamamoto, H.; Sato, Y.; Uchida, S.; Sasaki, T. Experimental evidence of selective inattention in reputation-based cooperation. Sci. Rep. 2018,8, 14813. [CrossRef] 32. Binmore, K. Why do people cooperate? Politics Philos. Econ. 2006,5, 81–96. [CrossRef] 33. Ledyard, J.O. Public goods: A survey of experimental research. In The Handbook of Experimental Economics; Kagel, J.H., Roth, A.E., Eds.; Princeton University Press: Princeton, NJ, USA, 1995; pp. 110–194. 34. Ule, A.; Schram, A.; Riedl, A.; Cason, T.N. Indirect punishment and generosity towards strangers. Science 2009,326, 1701–1704. [CrossRef] 35. Swakman, V.; Molleman, L.; Ule, A.; Egas, M. Reputation-based cooperation: Empirical evidence for behavioural strategies. Evol. Hum. Behav. 2015,37, 230–235. [CrossRef] 36. Camera, G.; Casari, M. Monitoring institutions in indefinitely repeated games. Exp. Econ. 2018 ,21, 673–691. [CrossRef] 37. Kamei, K.; Nesterov, A. Endogenous Monitoring through Gossiping in An Infinitely Repeated Prisoner’s Dilemma Game: Experimental Evidence; SSRN Working Paper; Durham University Business School: Durham, UK, 2020; pp. 1–53. 38. Leimar, O.; Hammerstein, P. Evolution of cooperation through indirect reciprocity. Proc. R. Soc. B Biol. Sci. 2001,268, 745–753. [CrossRef] [PubMed] 39. Hilbe, C.; Schmid, L.; Tkadlec, J.; Chatterjee, K.; Nowak, M.A. Indirect reciprocity with private, noisy, and incomplete information. Proc. Natl. Acad. Sci. USA 2018,115, 12241–12246. [CrossRef] [PubMed] 40. Duca, S.; Nax, H.N. Groups and scores: The decline of cooperation. J. R. Soc. Interface 2018 ,15. [CrossRef] [PubMed] Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. © 2020 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).