scieee AI-readable full text Open interactive document viewer

Adverse Selection, Heterogeneous Beliefs, and Evolutionary Learning

Buchen, Clemens,Palermo, Alberto

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Buchen, Clemens; Palermo, Alberto Article — Published Version Adverse Selection, Heterogeneous Beliefs, and Evolutionary Learning Dynamic Games and Applications Provided in Cooperation with: Springer Nature Suggested Citation: Buchen, Clemens; Palermo, Alberto (2021) : Adverse Selection, Heterogeneous Beliefs, and Evolutionary Learning, Dynamic Games and Applications, ISSN 2153-0793, Springer US, New York, NY, Vol. 12, Iss. 2, pp. 343-362, https://doi.org/10.1007/s13235-021-00396-x This Version is available at: https://hdl.handle.net/10419/287051 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Dynamic Games and Applications (2022) 12:343–362 https://doi.org/10.1007/s13235-021-00396-x Adverse Selection, Heterogeneous Beliefs, and Evolutionary Learning Clemens Buchen1 ·Alberto Palermo2 Accepted: 5 July 2021 / Published online: 21 July 2021 © The Author(s) 2021 Abstract We relax the common assumption of homogeneous beliefs in principal-agent relationships with adverse selection. Principals are competitors in the product market and write contracts also on the base of an expected aggregate. The model is a version of a cobweb model. In an evolutionary learning set-up, which is imitative, principals can have different beliefs about the distribution of agents’ types in the population. The resulting nonlinear dynamic system is studied. Convergence to a uniform belief depends on the relative size of the bias in beliefs. Keywords Evolutionary game theory ·Imitation equilibrium ·Heterogeneous beliefs · Adverse selection ·Cobweb model JEL Classification C61 ·C73 ·D82 ·D83 ·E32 1 Introduction Usually in mechanism design it is assumed that players have a subjective probability distribution over a set of possible elements or outcomes, which represents information privately known to other players. More specifically, in a principal–agent relationship with adverse selection the principal does not know the type of agent that she is matched with, but the distribution of types is common knowledge. Given a belief about this distribution, princiWe are grateful to seminar participants at the University of Tartu, the University of Marburg, EBS Business School, and the EARIE conference in Munich with a previous version. We thank an associate editor and two anonymous referees for very helpful comments and remarks, which helped us greatly improve the exposition of the paper. BClemens Buchen [email protected] Alberto Palermo [email protected] 1WHU-Otto Beisheim School of Management, Burgplatz 2, 56179 Vallendar, Germany 2Institute for Labour Law and Industrial Relations in the European Union (IAAEU), Trier University, Behringstr. 21, 54296 Trier, Germany 344 Dynamic Games and Applications (2022) 12:343–362 pals write court-enforceable contracts and agents self-select. Against this backdrop of the standard model, we introduce a bias on the part of the principals concerning their beliefs. We model an aggregative game where principals are firms in an competitive market. Each principal is randomly matched with an agent who, in exerting effort, generates an output. Principals offer contracts based on the expected aggregate quantity and their belief over the distribution of types in the economy. Whereas the payoff of the agent depends on his (privately known) cost, the principals’ payoff is affected by the realized aggregate output and their beliefs. We then study imitation equilibria in this market characterized by adverse selection and heterogeneous beliefs. The overarching question in the set-up then becomes: What are possible long-run equilibria of the beliefs principals hold? To this end, we formulate conditions under which biased beliefs can persist. The aggregate of all individual firm decisions has an externality effect on all market participants. Building on that notion, we are interested in the way that a bias can affect that externality and, in a feedback effect, how the externality affects the bias. This means that, on the one hand, and this is to be expected, firms acting on biased beliefs influence the aggregate quantity in the market, because their individual output decisions are changed by the bias. On the other hand, however, convergence toward a beliefs equilibrium depends on the market quantity and the realized profits. There is a long-term effect on market fundamentals, i.e., price, quantity, utilities and labor contracts. We show that the magnitude of the bias is decisive for the long-run outcomes. As intuition suggests, a large bias would be eradicated by market forces, whereas a modest degree of bias can persist. The reason is that for a range of biases the net of all externalities affecting an individual firm is positive. Imitation in a game-theoretic setting is developed by Björnerstedt and Weibull [5], VegaRedondo [22] and Schlag [18], where individuals imitate those strategies that offer higher profits. Apesteguia et al. [2] synthesize these approaches and test the theory with an experiment. Selten and Ostmann [20] study an imitation equilibrium in which higher profits also determine who will be imitated. In addition they introduce the notion of the reference group, which comprises all other players any individual would at all consider imitating. The precise definition of reference group is made in each case on the basis of the problem at hand. For example in a spatial sense as in Selten and Apesteguia [19], where firms imitate only neighboring firms. Or, as in Rothschild and Stiglitz [15] and Ania et al. [1], the reference group for principals—while not so named—includes contracts that are similar enough to one’s own contract. We contribute to this literature on imitation. We assume (and w.l.o.g.) that some of the principals are optimistic about the distribution.1This means that they believe that the distribution of types is more favorable than it actually is.2As a result, the profits are different depending on the belief a principal holds. Hence, we study a polymorphic population characterized by unbiased and optimistic principals. To the best of our knowledge, our paper is the first to introduce a bias for the uninformed market side in an evolutionary set-up.3In our model, learning takes place by imitation of other principals’ beliefs. That is, information 1The analysis is carried out with the assumption of an optimistic bias. A pessimistic bias would lead to a specular model with specular results. 2This is different from Arifovic and Karaivanov [3], who start from an adverse selection model evolving over time assuming that principals are unable to solve the correct maximization problem. 3There is a related literature on biases of agents, for example with respect to the perception of their own ability ([11]), their own and others’ ability ([17]) or the success probability of a project and the agent’s contribution to the success ([6,24]). Dynamic Games and Applications (2022) 12:343–362 345 about the beliefs is shared by word-of-mouth communication as in Banerjee and Fudenberg [4]. Alternatively, beliefs can be inferred if the menus of contracts are observable as in Ania et al. [1]. In putting the model squarely in the tradition of imitation, we effectively assume that principals are memoryless about past matches. Instead, we could assume that the lifespan is short. Alternatively, if we assumed principals able to revise beliefs as Bayesian updaters, then, with an appeal to the law of large numbers, the answer would be straightforward. As more information accumulates over time, Bayesian updating with perfect memory would ultimately lead to all principals having the same unbiased belief. However, in this scenario, the information requirements are strong, because principals have to live a sufficiently long time to collect enough data points. We imagine a continuum between a perfect Bayesian updater at the one end and imitation only at the other. A Bayesian approach requires that firms potentially live forever and that (in the long run) perfectly recover the underlying distribution of types. In assuming memorylessness, we are at the one end of the spectrum where the evolution of beliefs is not straightforward. Fundamentally, we are interested in studying potential outcomes if economic actors do not or are unable to accumulate enough observations to recover the true distribution. In this context, the assumption of memorylessness is a mathematical convenience. In our set-up, the probability to change beliefs depends on the matching, the propensity to switch, and the payoff difference between principals. We partially follow Selten and Ostmann’s [20] notion of reference groups. There it is stated that individuals tend to compare with similar others (i.e., membership of the same reference group). In this sense, we introduce a mechanism that describes the willingness of individuals to compare themselves with others. To be concrete, consider a model with two types of agents. There is one group including the principals matched with a high-cost agent and the second with those matched with a low-cost agent. This lets us define a propensity to compare for each principal, which expresses the willingness to compare with a randomly chosen different principal. Hence, we do not restrict the comparison to a given member of a reference group, but rather we conceive of a probability representing the willingness of a given principal to compare herself with others. Whereas this assumption represents an extension of the notion of reference groups in economics, it is acommonviewinthesocial comparison theory, which is commonplace in other disciplines. Starting with Festinger [9], psychologists point out that individuals tend to carry out “social comparisons” preferably (but not exclusively) with similar others. The latter point suggests that our principals treat information gained from different reference groups differently. A precursor to our approach in experimental economics is the work by Todt [21]. The subjects in these early experiments tended to be more open to imitation if the situation of the other party was perceived to be more similar to their own. Principals offer contracts that stipulate the production of a given quantity of a homogeneous good. Aside from a belief about the distribution of types, each principal also forms an expectation about the aggregate quantity produced in the market. We focus our attention on naive adaptive expectations, which implies that each principal writes a contract assuming that the aggregate quantity in a given period is the same as in the period before. The rationale is the following. Principals in our model do not know the salient characteristics of the market, simply because they are not aware that there are different beliefs. Rational expectations about the quantity would run contrary to this view of the role of principals and therefore are not useful in this respect.4Given the assumption about naive expectations, our model is 4There is a complementary behavioral approach to our set-up in Esponda [7] and Frick et al. [10]. They go in the direction of aggregation under misperception where actors would attempt to estimate an aggregate 346 Dynamic Games and Applications (2022) 12:343–362 then a version of a cobweb model where, traditionally, fluctuations arise due to a disconnect between the time quantities are chosen and prices are realized. We focus on just two possible beliefs, but obviously one could imagine a population in which each principal holds her own prior belief about the distribution. Then, the switching process described above should sooner or later lead to a situation in which the polymorphism of the population consists of either two remaining strategies or a stable configuration with more than two beliefs present. Whereas the latter case would require an analysis of the conditions under which a configuration of multiple beliefs can coexist, the former is a study of convergence toward a unique belief. We focus on this case keeping in mind that this is a “reduced” problem, because we start the analysis at a moment in time in which a potentially large number of beliefs has already been eliminated from the population and where the only polymorphism consists of two beliefs.5 The paper proceeds as follows: the next section introduces the interaction of the different groups in the model. The resulting dynamic system is studied in Sect. 3. Finally, we offer some concluding remarks in Sect. 4. All proofs are relegated to Appendix. 2 Population Interactions In Sect. 2.1, we describe the interaction between principals and agents in the stage game of the model. Then, in Sect. 2.2, we set up and describe the interaction among principals with different beliefs. This entails defining a mechanism that allows principals to modify their beliefs and therefore change their contract offers in the stage game. Hence, in the next two subsections we derive the nonlinear map which governs the evolution of the population composition and of the quantity. 2.1 Stage Game There are two large populations of principals and agents of equal size. Each principal wants to delegate a task to an agent in order to produce a quantity q. Agents are heterogeneous with regard to their ability to produce the quantity. They have a linear cost function defined as C(q,θ)=θq.Asisstandard,weassumethatθ∈θ, θ,whereθ>θand we denote θ =θ−θ. The proportion of agents with marginal cost θis v∈(0,1). In the next step, the stage game will be defined. Principals are heterogeneous in that they hold different beliefs about the distribution of agents’ abilities. In particular, we assume that some principals believe that the proportion of low-cost agents in the population is larger than it really is: Assumption 1 Each principal has a belief φabout the prevalence of low-cost agents, with φ∈{ρ,v}where ρ>v. We will sometimes refer to those biased principals as optimistic. The agent’s production provides a benefit to a principal i, which is measured by a function S(qi t,˜qt),whereqi tis the quantity produced by the agent working for principal iin period t (footnote 4 continued) ignoring the effect of biased choices. In our behavioral approach, evolution is not driven by an incomplete estimation but by imitation. 5Including the belief which turns out to be the true one (and not assuming a situation characterized by two biases) is less restrictive and comes from the aim of showing convergence toward a biased belief, and a possible coexistence of beliefs. A situation with two biases would mean, therefore, assuming that such convergence has already happened. However, a model with only biased beliefs would not substantially change the dynamic. Dynamic Games and Applications (2022) 12:343–362 347 and ˜qtis a sufficient statistic of the aggregate quantity in the market. The precise definition of ˜qtwill be given below. Timing Time is discrete. In a generic period t, the fraction of principals who write contracts on the basis of belief vis denoted by αt. The timing of the game follows the timing of the cobweb model where there is a lag between output decisions and realizations. In our model, this disconnect is between the contracting stage and the observability of outcomes. This means that the principal designs a contract in tfor a quantity that is only observed in t+1. The payment is conditioned on the observed quantity as well, contracted in t, but paid only when the quantity is observed. Technically speaking, the contracting stage takes place in a period t, when principals offer menus of contracts for the next period based on the beliefs as in Assumption 1. The functional form ([23, 231]) of the benefit that the principal expects to gain in t+1is: Sqt+1,Et(˜qt+1)=βqt+1−(qt+1)2 2+δqt+1Et(˜qt+1) In t, each principal defines a mechanism qt+1(θ),w t+1(θ)which entails a transfer wt+1 for each observed quantity in t+1. We assume βto be a positive constant and δ∈(−1,0) which is a measure of the degree of substitutability between principals’ outputs. Et(˜qt+1) denotes the expectation a principal forms in tabout the value of ˜qt+1in t+1, when all production is carried out. We make the following assumption about this expectation. Assumption 2 In each period t, each principal has a naive expectation about the aggregate quantity in the market: Et(˜qt+1)=˜qt. Assumption 2is the simplest way of modeling adaptive expectations compared to the alternative of Bayesian learning. The timing for the contracting-production stage in a flow period t,t+1 can be summarized as follows: 1. (Period t) Each agent realizes his type. 2. (Period t) Principals write contracts according to their beliefs about the distribution of types and according to naive expectations about the aggregate quantity ˜qt+1. 3. (Period t) Each agent is randomly matched with a principal and the agent decides whether to accept the contract or not. 4. (Period t+1) Contracts are executed and outcomes are realized and observed: profits and payments to agents are realized. Contracts In each period t, principals write contracts which entail a rent for each quantity observed in t+1. The quantities contracted in tand observed in t+1 are indicated by qt+1qt+1(θ)for the low-cost type and qt+1qt+1(θ) for the high-cost type. We will use either of the two notations where convenient. In addition, a similar notation will hold for the transfers wt+1(θ). We restrict our analysis to direct revelation mechanisms that are truthful. This can be done because the agent’s rent is only a function of his principal’s contract and of the aggregate quantity in the market in the previous period. The rent is U(qt+1(θ),w t+1(θ),θ)=wt+1(θ)−θqt+1(θ). Moreover, we assume that agents are protected in every state of the world by limited liability on the rent. Formally, each principal maximizes expected profits given the usual incentive (IC) and participation constraints (PC): 348 Dynamic Games and Applications (2022) 12:343–362 max {qt+1,wt+1,qt+1,wt+1}φSqt+1,Et(˜qt+1)−wt+1+(1−φ)Sqt+1,Et(˜qt+1)−wt+1 s.t wt+1(θ)−θqt+1(θ)≥0∀θ∈θ, θ(PCs) wt+1θ−θqt+1θ≥wt+1θ−θqt+1θ IC θ wt+1θ−θqt+1θ≥wt+1θ−θqt+1θ IC θ Recall that principals in this model use the aggregate quantity from the previous period as expectation for the next. That is, with regard to the quantity, the model we present is a cobweb model, because given the timing and the specification of the benefit function it is mathematically equivalent to a model where principals form expectations about the price instead of the quantity with a linear demand function.6 Given the standard nature of the maximization problem the following proposition is straightforward. Proposition 1 Given different beliefs and the same naive expectations about ˜qt+1, the quantities for the low-cost types are equal, or qv t+1=qρ t+1, whereas for the high-cost type we have qv t+1>qρ t+1. The rent U(·,θ) for the high-cost type is equal to zero for both types of principals, whereas for the low-cost type the rent U(·,θ)is higher with a v-principal than a ρ-principal. For the low-cost agent both types of principals stipulate the same, first-best quantity. However, the ρ-principal offers a smaller rent, because she mistakenly believes that there are more low-cost agents than there really are. For the high-cost type both contracts offer the same (zero) rent, but the v-principal stipulates a bigger quantity. This is so because the odds of being matched with a low-cost agent appear too large for the ρ-principal. Given that the quantity for the high-cost type is decreasing in the odds of being matched with one, the quantity of the optimistic principal is set too low. Quantities We denote the expected quantity over the different types for a principal with belief φby Eθqφ t+1(θ)=vqφ t+1+(1−v)qφ t+1and with ˜qt=iqt,idi the aggregate quantity in the market, where iis an indicator of the principals in the population. Given the different proportions of principals with different beliefs, an informal appeal to the law of large numbers allows us to write the aggregate quantity as: ˜qt+1=i qt+1,idi =αtEθqv t+1(θ)+(1−αt)Eθqρ t+1(θ)(1) Profits For contracts stipulated in t, the realized profits in t+1 for each θand for a given belief are functions πt+1˜qt+1,˜qt,φ,θand πt+1˜qt+1,˜qt,φ,θ. The expected profits are 6This equivalent model can be summarized as follows: 1. Each principal maximizes: max {qt+1,wt+1,qt+1,wt+1}φEtPt+1qt+1−q2 t+1 2−wt+1+(1−φ)EtPt+1qt+1−q2 t+1 2−wt+1 under (ICs) and (PCs), and therefore a linear supply function is obtained. 2. Principals have naive expectations about the price: EtPt+1=Pt. 3. The demand is linear: Qt+1=A−BP t+1. 4. Market clears: the prices are computed on the demand function. The connection to our model is established for β=A Band δ=−1 B. Our choice of the interval for δ∈(−1,0) defines a standard stable cobweb model (in the absence of any kind of heterogeneity of expectations about any variable). See Hommes [12] for an overview and a recent reapprecitation of the cobweb model. Dynamic Games and Applications (2022) 12:343–362 349 Eθπt+1(˜qt+1,˜qt,φ,v). The presence of ˜qtcomes from the fact that each contracted quantity qt+1(θ)is a function of ˜qt(Assumption 2). Proposition 2 Given different beliefs and the same naive expectations about the total quantity, for any realization θ∈θ, θ, the realized profits πφ t+1(·,θ)are such that: πρ t+1>π v t+1with πρ t+1−πv t+1=(θ)2ρ−v (1−ρ)( 1−v)(2) πv t+1πρ t+1with πv t+1−πρ t+1=θ ρ−v (1−ρ)( 1−v)δ(˜qt+1−˜qt)+1 2θ v 1−v+ρ 1−ρ (3) Parts of the results in Proposition 2are a direct result from Proposition 1. Due to the fact the unbiased v-principal pays a higher rent for the low-cost agent, but produces the same quantity as the biased ρ-principal, profits must be smaller. This can be seen from Eq. (2). Further, as can be seen in Eq. (3), the difference in profits for the high-cost agent depends on the change of the quantity ˜qfrom one period to the next. The v-principal makes a larger profit than the ρ-principal if the quantity decreases or is constant from one period to the next. The reverse is true if the change is positive and large enough. To summarize, the basic stage game defines an aggregative game in which the profit of principals in a particular period depends on the belief about the distribution of types, the specific match and the behavior of all other principals, which affects the aggregate quantity in the market. 2.2 Evolutionary Learning by Imitation We use a proportional imitation rule to model the replica equation ([18]). For that purpose we define the conditional switch rate, which is the probability that at the end of a period a principal changes beliefs. To do that, we periodically allow some principals at the end of a period to observe the profit of a second principal. For each principal two scenarios are possible. Either she meets a principal from the same reference group, who got matched with the same type of agent, or from a different reference group, i.e., a principal who got matched with a different type of agent. The propensity to compare is a measure of how open a principal is toward comparing her situation to a different principal. In what follows we use the following assumption. Assumption 3 The propensity to compare is equal to zero if the principals come from different reference groups. Assumption 3immensely simplifies the following exposition and analysis, and the intuitions are not hidden behind the algebra. We shall relax this assumption in Sect. 3.4. Given the proportion of low-cost types vand the proportion αtof principals using v, P(φ ¬φ) =αt(1−αt)is the probability that a principal with a belief φmeets a principal with the different belief. Since we assume that matching between principals and different types of agents is type-independent, the probability that two principals were matched with a low-cost agent is simply (v)2and the probability that both were matched with a high-cost agent is given by (1−v)2. Hence, we have the probabilities that two principals with different beliefs and in the same reference group meet: γvρ t=P(vρ)v2=αt(1−αt)v2 350 Dynamic Games and Applications (2022) 12:343–362 γρv t=P(ρv)( 1−v)2=αt(1−αt)(1−v)2 In words, γvρ tis the probability that a v-principal would consider switching to belief ρ, with a similar interpretation of γρv t. The probability of switching to the other strategy is linearly dependent on the payoff difference. Formally, it is the product ·πφ t+1−π¬φ t+1, where >0 is chosen to scale the payoff difference in such a way that it can be used as a probability. We find three mechanisms to justify why principals infer whether or not they come from the same reference group. Either because of a mechanism based on word-of-mouth communication as in Banerjee and Fudenberg [4], or because the contracts and profits are observed as in Ania et al. [1], or simply because the contracting quantities are observed, because principals with low-cost agents obtain equal quantities. We assume that principals are memoryless about past plays or past switches and that learning takes place by imitation of other principals’ beliefs. This assumption is based on the following considerations. Information about the beliefs is shared by word-of-mouth communication (see above). Alternatively, beliefs can be inferred if the menus of contracts are observable. Putting the pieces together, the dynamic over time is described by the following equation: αt+1=αt+γρv tπv t+1−πρ t+1−γvρ tπρ t+1−πv t+1 The equation should be read as follows. The fraction of v-principals in a period is equal to the fraction in the previous period plus all ρ-principals who switch to vminus all v-principals who switch to ρ. From Proposition 2we know that the term πv t+1−πρ t+1can be positive or negative depending on the magnitudes and direction of fluctuations of the quantity in the market. If the term is negative the direction of proportional imitation is reversed, which means that the v-principal switches to ρwith the given probability. The resulting equation is equivalent.7Substituting the specific switch rates defined above, we arrive at the discrete change of αfrom one period to the next: αt+1=αt+αt(1−αt)(1−v)2πv t+1−πρ t+1−v2πρ t+1−πv t+1 (4) 2.3 Overview To recap, the model aims at combining insights from cobweb models and the problem of asymmetric information, in particular, adverse selection outcomes. The three main assumptions we make reflect this basic goal. Certainly, one could imagine alternatives to Assumption 1. If one assumes that all principals have the same belief (effectively, ρ=v), then Eq. (4) simply disappears and the model reverts back to a basic cobweb. If, instead of optimism, one assumed pessimism (ρ<v), all of the results presented below would be symmetrically reversed. Assumption 2is integral to the cobweb model reflecting a version of adaptive expectations. As discussed, rational expectations would go against the spirit of boundedly rational principals. We use naive expectations, which are the simplest version of adaptive expectations taking only one preceding period into account. For an overview of the role of expectations see Evans and Honkapohja [8]. 7To see this, write the dynamic for πv t+1< πρ t+1as αt+1=αt−γρv t{[πρ t+1−πv t+1]} − γvρ t{[πρ t+1− πv t+1]}, which is equivalent. Dynamic Games and Applications (2022) 12:343–362 357 biased beliefs have on aggregate market outcomes, on the other hand, raises new questions to study in competitive markets. Funding Open Access funding enabled and organized by Projekt DEAL. Declarations Conflict of interest The authors have no relevant financial or non-financial interests to disclose. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. A Appendix A.1 Proof of Proposition 1 As is standard (see, e.g., Laffont and Martimort [14]), the participation constraint of the low-cost type is implied by PC(θ)andIC(θ). The incentive constraint of the high-cost type is slack at the optimum. Moreover, the other two are binding constraints. Then, using the binding constraints to substitute wages in the objective function, the maximization problem reads as follows: max {qt+1,qt+1}φ{S[qt+1,Et(˜qt+1)]−θqt+1}−φθqt+1 +(1−φ){Sqt+1,Et(˜qt+1)−θqt+1}. (A.1) The quantities at the optimum are defined implicitly by: S q(·)=θ,S q(·)=θ+φ 1−φθ. Substituting for Et(˜qt+1)=˜qtand using the specific functional form for S(·), we obtain that in any generic time the quantities set by principal are: qv t+1=qρ t+1=β+δ˜qt−θ(A.2) qφ t+1=β+δ˜qt−θ−φ 1−φθ (A.3) From ρ>v, follows: qv t+1=qρ t+1>qv t+1>qρ t+1. The binding PC(θ)clarifies that the high-cost types realize a zero rent independently from the principal they are matched with. Conversely, from the binding IC(θ),wehavethatthe rent of the low-cost types (rentφ t+1(θ)) differs according to the principals’ belief. It holds: rentφ t+1(θ)=θqφ t+1,(A.4) and therefore rentv t+1(θ)>rentρ t+1(θ). 358 Dynamic Games and Applications (2022) 12:343–362 A.2 Proof of Proposition 2 Recall that principals design contracts based on the belief Et(˜qt+1)=˜qt;meaningthatina generic time tcontracts are defined on the basis of quantities as in (A.2)and(A.3). Hence, their choices about quantities in a time t+1 are based on the belief about the aggregate quantity, which in our set-up equals the quantity one period before ˜qt.However,int+1the payoff is affected by the realization of the aggregate quantity ˜qt+1which is described by (1). To compute the differences in payoffs, it is useful computing the difference in quantities for the high-cost type. From Eq. (A.3) we obtain: qv t+1−qρ t+1=θ ρ 1−ρ−v 1−v.Wehave that for a match with a low-cost type the quantity for both principals is equal. Hence, the surpluses are equal and the only difference is in the paid informational rent. It follows: πρ t+1−πv t+1=θqv t+1−θqρ t+1=(θ)2ρ−v (1−ρ)(1−v),(A.5) which is Eq. (2) in the paper. Conversely, for a match with a high-cost type πv t+1−πρ t+1 =(β −θ+δ˜qt+1)qv t+1−(qv t+1)2 2−(β −θ+δ˜qt)qρ t+1+(qρ t+1)2 2 =(β −θ+δ˜qt+1)(qv t+1−qρ t+1)−(qv t+1+qρ t+1)(qv t+1−qρ t+1) 2 =(qv t+1−qρ t+1)(β −θ+δ˜qt+1)−(qv t+1+qρ t+1) 2 =(qv t+1−qρ t+1)(β −θ+δ˜qt+1)−1 2(2β+2δ˜qt−2θ−v 1−vθ −ρ 1−ρθ) =θ ρ−v (1−ρ)( 1−v)δ(˜qt+1−˜qt)+1 2θ v 1−v+ρ 1−ρ, (A.6) which is Eq. (3) in the text. A.3 Derivation of the Nonlinear Map We start by computing the equation describing the evolution of the aggregate quantity over time. From (1), we know: ˜qt+1=αtEθ[qv t+1(θ)]+(1−αt)Eθ[qρ t+1(θ)](A.7) Using (A.2)and(A.3), we compute Eθ[qv t+1(θ)]and Eθ[qρ t+1(θ)], where the expectation is w.r.t. the true realization of the variable θ(i.e., the distribution for which it holds Pr(θ = θ)=v). Eθ[qv t+1(θ)]=vqv t+1+(1−v)qv t+1 =v(β +δ˜qt−θ) +(1−v)(β +δ˜qt−θ−v 1−vθ) =β−θ+δ˜qt (A.8) Dynamic Games and Applications (2022) 12:343–362 359 and Eθ[qρ t+1(θ)]=vqρ t+1+(1−v)qρ t+1 =v(β +δ˜qt−θ) +(1−v)(β +δ˜qt−θ−ρ 1−ρθ) =β−θ+δ˜qt+vθ −ρ(1−v) 1−ρθ =β−θ+δ˜qt−c, (A.9) with c=θ ρ−v 1−ρas defined in the text. Using (A.8)and(A.9)in(A.7), it is immediate to obtain: ˜qt+1=β−θ+δ˜qt−(1−αt)c,(A.10) which is Eq. (6) in the paper. Subtracting ˜qtto both sides of this equation, we obtain: ˜qt+1−˜qt=β−θ+(δ −1)˜qt−(1−αt)c(A.11) Substituting (A.11)inEq.(A.6) to eliminate its dependence on ˜qt+1gives the difference in realized payoffs: πv t+1−πρ t+1=aθδ β−θ+(δ −1)˜qt−(1−αt)c+b(A.12) Then, using both differences in realized payoffs (A.5 and A.12) in the replica equation (4), we obtain (5). A.4 Proof of Proposition 3 The proof involves a simple inspection of the eigenvalues of the Jacobians. The two eigenvalues are δand 1 ±ak. Then, for k=0 one eigenvalue crosses the unit circle for both points (fold bifurcation). Moreover, for kchanging sign one point has both eigenvalues smaller than one, whereas the other becomes a saddle point exchanging stability (transcritical bifurcation). This implies that for k<0 the point X0is a stable hyperbolic steady state, which corresponds to the situation depicted in Fig. 1a. Accordingly, the reverse case of k>0is shown in Fig. 1b, in which X1is stable. A.5 Proof of Lemma 2 The following proof works on the basis of the center manifold theorem. The theorem claims that whenever the system is close enough to a steady state, the stable and unstable manifolds are tangent to the respective stable and unstable eigenvectors of the linearized system (see, e.g., Kuznetsov [13] page 157). Given the theorem, the proof can be formulated as follows. Given the Jacobians in the steady state, the two eigenvalues are δand 1±ak.Since|δ|<1, the corresponding eigenvectors are the stable ones, they are invariant and correspond to the vertical line in α=0andα=1. Then, for an it is sufficient to define a rescaling of θ such that ak=ak. The rest follows from the center manifold theorem. A.6 Proof of Theorem 1 The proof is based on the results of Proposition 3, and therefore, it requires to identify the stable and unstable fixed points. As seen, the stability of the steady states depends on the 360 Dynamic Games and Applications (2022) 12:343–362 sign of k. Whenever the proportion of low-cost agents is greater than half of the population (v>1 2), kcan be greater than, smaller than or equal to zero. Recall that k=0 is satisfied for ρ=ρcand that the sign of kdepends on the relation between ρand ρc.Ifρ>ρ cthen k>0 and, therefore, from Proposition 3the fixed point X1is a sink and X0is a saddle node; the opposite is true for ρ<ρ c, which implies that k<0. Hence, given that there are only two fixed points, one stable and the other unstable, for any initial state the population converges to the stable one. It remains to prove that for v≤1 2for every ρ>vit follows ρ>ρ cand, therefore, that X1is the sink. In fact, solving k=0 to obtain ρc,wehave: ρc(v) =v(3v−1) 4v2−3v+1. The function ρc(v) in the interval v∈[0,1 2]has a unique minimum, and therefore, it is U-shaped. Moreover, it holds ρc(v =0)=ρc(v =1 3)=0. Hence, for v∈[0,1 3)it holds that ρc<0 and therefore every ρ>0>ρ c. Conversely, for v∈(1 3,1 2]it holds that ρc(v) is increasing, but ρc(v =1 2)=1 2. Hence, the function ρc(v) is always below the straight line vand therefore every ρ>vit also such that ρ>ρ c. A.7 Proof of Theorem 2 The aim is to show that there can exist a limit-two cycle. Hence, we will proceed in computing the second iterate for the equations describing the evolution of the aggregate quantity and the fraction α. Then, we will show that the conditions (δrelatively small or θ relatively large) as in the theorem ensure the existence of the limit-two cycle. To simplify the algebra, let R≡a(1−v)2δ 1+δcand S≡ak.FromEq.(6), recursively, we compute the second iterate, i.e., ˜qt+2as only dependent on ˜qt. With q(2)we denote the solution imposing ˜qt+2=˜qtwhich is equal to: q(2)=β−θ 1−δ−(1−αt)δ 1−δ2c−(1−αt+1)1 1−δ2c.(A.13) Inserting (A.13)in(5), and simplifying using the expressions for Rand S, we can write: αt+1[1+αt(1−αt)R]=αt[1+αt(1−αt)R]+αt(1−αt)S(A.14) The same relationship holds for αt+1,αt+2: αt+21+αt+1(1−αt+1)R=αt+11+αt+1(1−αt+1)R+αt+1(1−αt+1)S (A.15) To simplify the algebra (and, more importantly, the subsequent analysis) even further, we will use the following substitution: H:= 1+αt(1−αt)R With this last equation, it is helpful to rewrite (A.14)as: αt+1=αt+αt(1−αt)S H=αt H+(1−αt)S H To reduce the amount of computations, we will write: αt+1=αt L Hwith L:= H+(1−αt)S(A.16) Dynamic Games and Applications (2022) 12:343–362 361 Substituting (A.16)in(A.15): αt+21+αt HL1−αt HLR =αt HL1+αt HL1−αt HLR+αt HL1−αt HLS We denote with α(2)the steady state of the second iterate, i.e., α(2)≡αt+2=αt. Hence, dividing the previous equation by α(2), we can write: 1+α(2) HL1−α(2) HLR =L H1+α(2) HL1−α(2) HLR+L H1−α(2) HLS Adding and subtracting in the last bracket L/Hand collecting common factors: 1+α(2) HLH−α(2) HLR−LS H=L2 H1−α(2) HS This last expression can be simplified, obtaining: (1−α(2))HL2S+(H−1)LS(H−α(2)L)−(1−α(2))HLS2=−(1−α(2))H2S−→ LHL +(H−1)(H−α(2)L) 1−α(2)−HS=−H2 Using the expression for Land simplifying: (H+(1−α)S)(2H2−2HαS−H+αS)=−H2 Factoring and recalling the expression for H,wewrite: H+1−α(2)SH−α(2)S[2H−1]=−H2(A.17) H=1+α(2)1−α(2)R(A.18) Equations (A.17)and(A.18) determine α(2),(A.14) determines αt+1and (A.13) determines q(2). Equation (A.13) has an unique solution for q(2). The solutions for α(2)are not easily obtainable, and they can be in the set of complex numbers. Hence, in what follows, we discuss the conditions leading to real solutions. Equations (A.17)and(A.18) define a polynomial of degree 6 for α(2). Observe that δ→−1 implies R→−∞. It follows that (independently of θ)Hfrom (A.18) can be sufficiently negative to allow for the LHS of (A.17)tobenegative and therefore potentially ensure real solutions for α(2). Conversely, suppose δis sufficiently large; we show that also θ sufficiently large ensures the existence of a solution. Equations (A.17)and(A.18) describe a function of α(2),say, f(α(2)).Wehavetoprovethat f(α(2))=0 is possible. With this aim, observe that limα(2)→0f(α(2))=2+Sand limα(2)→1f(α(2))= 2−S, implying that there is at least one solution whenever |S|>2. Notice that the sign of kdoes not depend on θ, and it is straightforward to see that S≡ak∝(θ)2. It follows that the value of θ can be chosen large enough to ensure |S|>2. 362 Dynamic Games and Applications (2022) 12:343–362 References 1. Ania AB, Tröger T, Wambach A (2002) An evolutionary analysis of insurance markets with adverse selection. Games Econ Behav 40(2):153–184 2. Apesteguia J, Huck S, Oechssler J (2007) Imitation-theory and experimental evidence. J Econ Theory 136(1):217–235 3. Arifovic J, Karaivanov A (2010) Social learning in a model of adverse selection. In: Industrial organization, trade and social interaction: essays in Honour of B. Curtis Eaton. University of Toronto Press, Toronto 4. Banerjee A, Fudenberg D (2004) Word-of-mouth learning. Games Econ Behav 46(1):1–22 5. Björnerstedt J, Weibull JW (1996) Nash equilibrium and evolution by imitation. In: Arrow KJ, Colombatto E, Perlman M, Schmidt C (eds) The rational foundations of economic behavior. Macmillan, Houndmills, pp 155–171 6. de la Rosa LE (2011) Overconfidence and moral hazard. Games Econ Behav 73(2):429–451 7. Esponda I (2008) Behavioral equilibrium in economies with adverse selection. Am Econ Rev 98(4):1269– 91 8. Evans GW, Honkapohja S (2001) Learning and expectations in macroeconomics. Princeton University Press, Princeton 9. Festinger L (1954) A theory of social comparison processes. Hum Relat 7(2):117–140 10. Frick M, Iijima R, Ishii Y (2020) Stability and robustness in misspecified learning models. Cowles Foundation Discussion Paper, 2235 11. Gervais S, Goldstein I (2007) The positive effects of biased self-perceptions in firms. Rev Finance 11(3):453–496 12. Hommes CM (2013) Behavioral rationality and heterogeneous expectations in complex economic systems. Cambridge University Press, Cambridge 13. Kuznetsov YA (1998) Elements of applied bifurcation theory. Springer, New York 14. Laffont J-J, Martimort D (2002) The theory of incentives: the principal-agent model. Princeton University Press, Princeton 15. Rothschild M, Stiglitz JE (1976) Equilibrium in competitive insurance markets: an essay on the economics of imperfect information. Q J Econ 90(4):629–649 16. Sandholm WH (2011) Population games and evolutionary dynamics. MIT Press, Cambridge 17. Santos-Pinto L (2008) Positive self-image and incentives in organisations. Econ J 118(531):1315–1332 18. Schlag KH (1998) Why imitate, and if so, how? J Econ Theory 78(1):130–156 19. Selten R, Apesteguia J (2005) Experimentally observed imitation and cooperation in price competition on the circle. Games Econ Behav 51(1):171–192 20. Selten R, Ostmann A (2001) Imitation equilibrium. Homo oeconomicus 18(1):111–149 21. Todt H (1972) Pragmatic decisions on an experimental market. In: Sauermann H (ed) Contributions to experimental economics, pp 608–634 22. Vega-Redondo F (1997) The evolution of Walrasian behavior. Econometrica 65(2):375–384 23. Vives X (2001) Oligopoly pricing: old ideas and new tools. MIT Press, Cambridge 24. Wang J, Zhuang X, Yang J, Sheng J (2014) The effects of optimism bias in teams. Appl Econ 46(32):3980– 3994 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.