Modeling the formation of R&D alliances: An agentbased model with empirical validation
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Tomasello, Mario Vincenzo; Burkholz, Rebekka; Schweitzer, Frank Working Paper Modeling the formation of R&D alliances: An agentbased model with empirical validation Economics Discussion Papers, No. 2017-107 Provided in Cooperation with: Kiel Institute for the World Economy – Leibniz Center for Research on Global Economic Challenges Suggested Citation: Tomasello, Mario Vincenzo; Burkholz, Rebekka; Schweitzer, Frank (2017) : Modeling the formation of R&D alliances: An agentbased model with empirical validation, Economics Discussion Papers, No. 2017-107, Kiel Institute for the World Economy (IfW), Kiel This Version is available at: https://hdl.handle.net/10419/172009 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Received December 4, 2017 Accepted as Economics Discussion Paper December 5, 2017 Published December 6, 2017 © Author(s) 2017. Licensed under the Creative Commons License - Attribution 4.0 International (CC BY 4.0) Discussion Paper No. 2017-107 | December 06, 2017 | http://www.economics-ejournal.org/economics/discussionpapers/2017-107 Modeling the formation of R&D alliances: an agentbased model with empirical validation Mario V. Tomasello, Rebekka Burkholz, and Frank Schweitzer Abstract The authors develop an agent-based model to reproduce the size distribution of R&D alliances of firms. Agents are uniformly selected to initiate an alliance and to invite collaboration partners. These decide about acceptance based on an individual threshold that is compared with the utility expected from joining the current alliance. The benefit of alliances results from the fitness of the agents involved. Fitness is obtained from an empirical distribution of agent's activities. The cost of an alliance reflects its coordination effort. Two free parameters ac and a1 scale the costs and the individual threshold. If initiators receive R rejections of invitations, the alliance formation stops and another initiator is selected. The three free parameters (ac; a1; R) are calibrated against a large scale data set of about 15,000 firms engaging in about 15,000 R&D alliances over 26 years. For the validation of the model the authors compare the empirical size distribution with the theoretical one, using confidence bands, to find a very good agreement. As an asset of our agent-based model, they provide an analytical solution that allows to reduce the simulation effort considerably. The analytical solution applies to general forms of the utility of alliances. Hence, the model can be extended to other cases of alliance formation. While no information about the initiators of an alliance is available, the results indicate that mostly firms with high fitness are able to attract newcomers and to establish larger alliances. (Published in Special Issue Agent-based modelling and complexity economics) JEL L14 Keywords R&D network; alliance; collaboration; agent Authors Mario V. Tomasello, Chair of Systems Design, ETH Zurich, Department of Management, Technology and Economics, Zurich, Switzerland Rebekka Burkholz, Chair of Systems Design, ETH Zurich, Department of Management, Technology and Economics, Zurich, Switzerland Frank Schweitzer, Chair of Systems Design, ETH Zurich, Department of Management, Technology and Economics, Zurich, Switzerland, [email protected] Citation Mario V. Tomasello, Rebekka Burkholz, and Frank Schweitzer (2017). Modeling the formation of R&D alliances: an agent-based model with empirical validation. Economics Discussion Papers, No 2017-107, Kiel Institute for the World Economy. http://www.economics-ejournal.org/economics/discussionpapers/2017-107
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation 1 Introduction 25 Collaboration can be widely observed in dierent social and economic systems, where agents 26 strive to reach a common goal. Scientists collaborate to write joint publications (Katz and 27 Martin, 1997), rms collaborate to le joint patents (Hoang and Rothaermel, 2005; Kim and 28 Song, 2007), and software developers collaborate to create joint software products (Bitzer and 29 Geishecker, 2010; Lakhani and Wolf, 2003). To explain collaboration, economic research has tra30 ditionally focused on dierent aspects of labor division (Durkheim, 2014) and productivity of 31 teams (Scholtes et al. , 2016). However, in the wake of technology-driven economic growth, the 32 question how to boost collaboration to foster knowledge transfer and innovation has become more 33 important (Frenz and Ietto-Gillies, 2009). 34 The current research about the dynamics of R&D networks can be seen as a major contribution to 35 better understand how rms collaborate in patenting activities. In this network representation, 36 nodes depict the economic agents, i.e. the rms, and links between nodes their collaboration. 37 Specically, rms formally declare this collaboration in publicly announced alliances , which can 38 involve more than two partners. So, it makes sense to ask how the size of alliances , i.e. the 39 number of partners involved , can be explained by means of an agent based model, which is the 40 aim of the current paper. 41 To address this questions, we can build on a number of empirical studies about R&D alliances. 42 It was shown that, because rms are involved in dierent alliances at the same time, their 43 collaboration results in a large network component, in which even rms not directly collaborating 44 are still connected through other rms (See Figure 1). At the same time, a large number of small 45 rm alliances exist that are not connected to the rest of the network. These co-existing sub46 networks are called components in the following. 47 The formation of a strongly connected component can be seen as an emergent property of the 48 economic network because it is not planned top down, but emerges during the process of alliance 49 formation, if (some) rms become engaged in more than one alliance. Once such a strongly 50 connected component exists, it greatly enhances the transfer of knowledge and the diusion of 51 innovations even between distant rms, so it is benecial from a policy perspective. 52 (Tomasello et al. , 2014, 2017) have already proposed an agent-based model that is able to re53 produce most of the properties of the observed R&D network. These properties include (i) the 54 distribution of component sizes, i.e. the number of components of a given size plus the size of the 55 largest connected component, (ii) the distribution of local clustering coecients, i.e. the fraction 56 of rms in a component that form triads (closed triangles) in their collaboration, (iii) the distri57 bution of the lengths of shortest paths that connect any two rms in the network, and (iv) the 58 distribution of degrees, i.e. the number of partners of a rm. 59 This agent-based model, while successfully reproducing network features along dierent dimen60 sions, takes two empirical distributions as an input: (a) the distribution of agent's activities , 61 2/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation i.e. their propensity to engage into a collaboration, and (b) the distribution of alliance sizes , 62 i.e. the number of partners involved in an alliance. The latter has been investigated empirically 63 (see Hagedoorn, 2002; Tomasello et al. , 2016). Remarkably, one nds a broad and right-skewed 64 distribution of alliance sizes (see Figure 2). The same distribution was found even for dierent 65 industrial sectors (including manufacturing, research, nancial and service sectors), that exhibit 66 substantial dierences otherwise. 67 Previous modeling attempts in this eld have, to the best of our knowledge, limited themselves 68 only to general features of the R&D network, such as the degree distribution or small world prop69 erties. With the current paper, we want to move the agent-based modeling one important step 70 forward, by explaining the distribution of alliance sizes as an emergent feature of an underlying 71 agent-based model instead of taking it from observations. This requires us to explicitly model 72 how agents form alliances, which implies to consider why agents form alliances. But given the 73 empirical work on alliance sizes, we have some ground truth to later judge the performance of 74 our agent-based model in reproducing the distribution of alliance sizes. 75 2 Empirical ndings 76 2.1 The network of R&D alliances 77 The dataset. We build our empirical R&D network using the SDC Platinum database, 1 that 78 reports approximately 672,000 publicly announced alliances in all countries, from 1984 to 2009, 79 with a granularity of 1 day, between several kinds of economic actors (including manufacturing 80 rms, investors, banks and universities) for which we commonly use the term rm in the 81 following. We then select all the alliances characterized by the R&D ag; after applying this 82 lter, a total of N= 14,829 alliances, connecting n= 14,561 rms, are listed in the dataset. 83 An R&D alliance is dened as an declared partnership between two or more rms. This can 84 range from formal joint ventures to more informal research agreements, specically aimed at 85 research and development purposes. Note that we do not have any information about the rm 86 that initiated the alliance, nor about the sequence in which rms joined an alliance. 87 The analysis of the data set, as well as all the network analyses and plots, are done by means of 88 the R software for statistical computing. 2 89 Reconstructing the collaboration network. In the present study, we investigate the R&D 90 network aggregated over all years and all industrial sectors, which has to be reconstructed from 91 the data set. Firms are represented as nodes in the network and R&D alliances as undirected 92 links between nodes. Isolate nodes, i.e. rms not taking part in any R&D partnership, are simply 93 excluded from our network representation. 94 1 http://thomsonreuters.com/sdc-platinum/ 2 http://www.r-project.org/ 3/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation Figure 1: Visual representation of the R&D network that we analyze in this study the size of the nodes encodes their tness. We have used the igraph package (Csardi and Nepusz, 2006) and the Fruchterman-Reingold layout algorithm (Fruchterman and Reingold, 1991), which minimizes the number of crossing links. When an R&D alliance involves more than two rms, we assume that all the corresponding 95 nodes are connected in pairs, forming a fully connected clique. A standard two-partner alliance 96 is then a fully connected clique of size 2. The choice of the fully connected clique rather than 97 less interconnected network architectures derives from the fact that alliances of more than two 98 partners, although representing only a minority, require great coordination and resources. There99 fore, they have to be associated with a higher number of links in the corresponding collaboration 100 networks. By following this procedure, the 14,829 R&D alliances listed in the dataset result in a 101 total of 21,572 links. The resulting network is shown in Figure 1. 102 Distribution of alliance sizes A salient feature of the R&D alliances in the SDC dataset is 103 the variable number of partners they involve. Most of the collaborations (93%) are stipulated 104 between two partners, the remaining ones involve three or more partners. In the following, we 105 denote by s the size of the alliance, whereas n indicates the number of rms and N the number 106 of alliances. 107 We report the empirical distribution of the alliance size, pe s(s) in the R&D network in Figure 2. 108 As clearly visible, it spans one order of magnitude and is right-skewed. It should be noted that an 109 4/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation identication of the functional form of the distribution (e.g., power-law, exponential, log-normal 110 and so on) is outside of the scope of this study. Our aim instead is to develop an agent-based 111 model to reproduce this distribution, as described below. 112 2 5 10 1 10 100 1000 10000 Firms per alliance Counts Figure 2: Histogram of the empirical alliance size distribution pe s(s) measured on the R&D network. This distribution is later used to evaluate the outcome of our agent-based model. 2.2 Dening a tness measure for agents 113 Our agent-based model requires an attribute, called tness, that is assigned to each agent. It 114 describes how attractive an agent is for the other agents, to form an alliance. To keep the model 115 as a general as possible, we decide to proxy tness by a measure which is not system specic, such 116 as the operational value of a rm. We choose the so-called agents' activity (Perra et al. , 2012), 117 which has been already successfully used on various data sets, such as online microblogging, actor 118 networks, R&D and co-authorship networks (Tomasello et al. , 2014). The empirical activity η∆t i,t 119 of an agent i at time t , over a time window ∆t , is dened as the number of alliances n∆t i,t that 120 involve agent i in the time window ∆t ending at time t , divided by the total number of alliances 121 N∆t t involving any agent in the network during the same time period: 122 η∆t i,t =n∆t i,t N∆t t . (1) It was found that activity distributions in most collaboration networks are right skewed and 123 dispersed over several orders of magnitude, as in many other social and technological systems 124 (Barabasi, 2005; Pastor-Satorras et al. , 2001). This is conrmed also for the case of R&D net125 works, where the empirical activity values range from low 0.002 to the maximum value of 1 . 126 Applying this to the tness of agents, this means that the agent with the highest tness has a 127 value of 2-3 orders of magnitude larger than the agents with the lowest tness. Indeed, the vast 128 5/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation majority of the agents has a tness equal to the minimum value, which is also the median value, 129 and the average tness is only slightly higher than that. Only one agent has a tness equal to 1 130 (the highest possible value). 131 Contrary to most network indicators that display strong variability and dependence on time, 132 especially in R&D networks (see Tomasello et al. , 2016), activity is a stable attribute that can 133 be assigned to rms to eectively estimate their propensity to engage in new collaborations, as 134 well as their attractiveness to potential new collaborators. Empirical activities are robust with 135 respect to (a) the time t at which they are measured, (b) the length of the selected time window 136 ∆t , (c) the sectoral classication of rms or authors, as shown by Tomasello et al. (2014). Such 137 a stability makes activity a perfect empirical proxy for our tness attribute. 138 Given the robustness with respect to the time window, we decide to compute the tness values 139 using the longest possible window, i.e. the entire observation period, therefore ηi≡η∆t=26years i,t=2009 . 140 This considers the full information from the data set and results in activities ηi that are always 141 strictly greater than 0 because, by denition, all rms in our network must be involved in at 142 least 1 alliance. In Figure 3 we report the empirical distribution of activity, i.e. of tness, pe η(η) , 143 for the analyzed R&D network. Further, in Figure 1, we have used the empirical tness values 144 of agents to scale their size in the collaboration network. Agents with higher tness obviously 145 form the core of the empirical R&D network. 146 ● ● ● ● ● ●● ● ● ● ● ● ● ● ●● ● ● ● ● ● ● ● ● ●● ● ● ●● ● ●● ● ● ●● ● ● ● ● ● ● ● ● ● ●● ● ● ● ● ●●● ● ● ●●●●●●● ● ● ●●●●●● ● ●●●●●● ● ●●●●●●●●●●●● ● 0.002 0.005 0.020 0.050 0.200 0.500 1 10 100 1000 10000 Fitness Counts Figure 3: Histogram of agents' empirical tness distribution, pe η(η) , measured on the R&D network. This distribution is later used as an input for our agent-based computer simulations. 6/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation 3 The modeling approach 147 3.1 Agent-based model of alliance formation 148 In the following, we develop an agent-based model to reproduce the observed size distribution of 149 consortia, shown in Figure 2. This distribution is the result of a dynamic process in which agents 150 decide to initiate or to join an alliance, i.e. it can only be understood by modeling the growth of 151 the collaboration network. 152 Fitness of agents and initiation of an alliance. Our model is a considerable extension 153 of the network tness model rst proposed by Bianconi and Barabási (2001). Each agent i is 154 assigned a tness ηi which is xed and independent of time. The values for the tness are obtained 155 from the empirical distribution pe η(η) , shown in Figure 3. 156 In our model, all n agents can become active with a uniform probability , which is chosen to be 157 1/n , independent of their tness. The sampling occurs with replacement, i.e. agents can also 158 be chosen more than once to become active. Activity means here that an agent initiates a new 159 alliance; hence we refer to her as the initiator. We do not have empirical information about the 160 agent that initiated an alliance, hence the assumption of a uniform probability for the activation 161 is reasonable. 162 For our simulations, we choose a discrete time t which measures the time to form an alliance . 163 I.e. each time a new initiator is selected we start with t= 0 and the maximum time for alliance 164 formation is denoted as T . The newly created alliance can grow only if new collaborators join. 165 This process is reected in two steps: (i) the initiator invites new collaboration partners, one per 166 time step, (ii) the invitees accept or reject to join the alliance. 167 Utility of consortia. A number of agents form an alliance C(st) of size st which can change over time as new agents join the alliance. There can be many consortia of dierent sizes coexisting over time. The utility function, ut , of the alliance combines the benets, bt , and the costs, ct , of the collaboration of the st agents, i.e. both bt(st) and ct(st) depend on the current alliance size, st . Further, the benets should be a monotonous function of the tness values of the currently involved agents, i.e. bt(..., ηi, ηj, ...) , whereas the costs should reect the coordination eort of the alliance and thus should be a monotonous function of the size of the alliance. For simplicity, we assume linear dependencies for the monotonous functions, i.e. the utility of an alliance is dened as ut=bt−ct;bt= st X m=1 ηm;ct=ac·st−1 (2) where the parameter ac allows to scale costs against benets. We note that the costs scale with 168 the number of alliance partners rather than with the number of their possible connections , which 169 7/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation would be quadratic, i.e. st(st−1)/2 , if an alliance is seen as a fully connected clique. The latter 170 would account for a superlinear increase in the coordination eort between partners which sets 171 strong limitations to larger consortia. Here, instead we assume some sort of administration cost 172 based on the number of parties. 173 Invitation of alliance partners. As the alliance grows, its utility will change in a non174 monotonous manner. Precisely, according to Equation 2, the utility will grow only if the tness 175 ηj of the new alliance member j is larger than the scaling constant ac . To ensure that this 176 condition is met, an initiator preferably invites agents with a high tness. Precisely, similar 177 to the tness model of Bianconi and Barabási (2001), the initiator i chooses potential alliance 178 partners j , one at a time, with a probability proportional to their tness ηj . 179 Dierent from the mentioned model, it is however left to the agents to decide whether they want 180 to accept this invitation, i.e. to join the alliance. The initiator repeats the invitation procedure 181 until a number R of invited partners refuse the invitation to the alliance. I.e. R is a parameter 182 of our agent-based model. If the current number of rejections , rt , reaches R , we assume that the 183 alliance is fully formed and stops to grow in size. At the same time the initiator loses its active 184 status. If the initiator receives R rejections already from the rst R selected partners, then no 185 alliance is formed. 186 Formation of collaboration links The second step in the formation of the alliance is the 187 decision of the invitee to accept, or to not accept, the invitation. An agent j decides to join an 188 alliance C at time t+ 1 if the utility of the alliance, ut is larger than a certain threshold, uthr j , 189 which is assumed to be heterogeneous across agents. 190 Specically, we argue that the threshold of an agent to join an alliance increases with her tness. 191 The rationale behind this is that agents with a high tness are very attractive for initiators of 192 consortia and thus receive invitations very often. On the other hand, because of scarce resources, 193 agents cannot simply accept all invitations, they have to be selective. Therefore, the higher the 194 own tness and the attractiveness for consortia, the higher the threshold to accept an invitation. 195 Conversely, agents with a low tness are not invited very often for an alliance, therefore they will 196 be more inclined to accept invitations, i.e. their threshold is lower because of the lower tness. 197 Hence, it is reasonable to argue that uthr j=alηj , i.e. uthr j is simply proportional to the tness, 198 where al is a parameter of the model, to be determined later. 199 This results in the following condition for agent j to join the alliance C in the next time step j∈ Ct+1 if ut≥uthr j⇒ st X m=1 ηm−ac[st−1] ≥al·ηj (3) Implications. Our agent-based model builds on an interesting tension between the attractive200 ness and the willingness to become an alliance member, which is a novel point in the discussion 201 8/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation Our model builds on heterogeneous agents characterized by one individual parameter, their tness 357 ηi . The distribution of tness values is proxied by the empirical activity distribution, Figure 3, 358 as the only model input. Activity describes how often an agent was engaged in an alliance during 359 the observation period, which is 26 years in our case. This can be seen as an indication of the 360 agent's attractiveness for other agents to collaborate with, and tness should be interpreted in 361 the same manner. We have shown that the tness distribution obtained this way is right skewed 362 and very broad. 363 Further, our agent-based model uses three free parameters that need to be determined in compar364 ison with empirical data. The calibration process is based on a maximum likelihood estimation 365 that returns those parameter values that match best the target, which is the empirical distribu366 tion of alliance sizes. 367 It is interesting to note that only two of these parameters, the scaling factors ac for the cost of 368 the consortium and al for the individual threshold to accept an invitation obtain a stable value 369 in the maximum likelihood estimation, whereas the third parameter R , the number of rejections 370 to stop forming an alliance does not reach a denite value. Instead, we observe that equally 371 good likelihoods are obtained for a larger range of R between 1 and 20. Hence, our model works 372 without assuming a specic value of R . In other words, R can vary across time, industrial sectors 373 or even alliances without questioning the validity of our model. 374 For our model validation we used a high number of rejections, ˆ R= 20 . This is denitely realistic 375 for a system such as the global, inter-sectoral R&D network that we analyze. Here, rms have to 376 search for their partners among a huge number of potential candidates, making the establishment 377 of an R&D alliance potentially costly and risky. We argue that this leads to a very long and 378 cautious selection process, from the side of both the initiator and the invited rm. Therefore, 379 rms have to be willing to accept a high number of rejections, if they want to gain access to 380 external knowledge and eventually establish R&D collaborations with other rms. 381 Regarding the other two parameters, ˆac= 0.04 and ˆal= 2 , we note from Equation (5) that 382 actually their ratio matters, as it determines the range of tness values [0, η?] for which agents 383 join an alliance. The denite value of ˆac should be interpreted as rather large. I.e. when multiplied 384 with the size of the alliance, the cost in Equation (2) is rather high in comparison with the benet 385 of the alliance, which is the sum of the tness values of the agents. This has two consequences. 386 First, it restricts the maximum size of an alliance to values below 20. Second, it restricts the 387 maximum number of alliances with sizes larger than 2, because most agents in the system have a 388 rather low tness and are thus not able to overcome the considerable cost of forming an alliance. 389 To illustrate this, an alliance of two agents with median tness values exhibits a benet of 0.004 390 and a cost of 0.04; or an alliance of four agents with median-tness agents exhibits a benet of 391 0.008 and a cost of 0.12, i.e. almost an order of magnitude larger. 392 This reects the intention of our agent-based model. Agents with high tness (typically incumbent 393 rms) are the ones that are most likely to receive an invitation. At the same time, they will most 394 often refuse the invitation, if an alliance consists of only agents with medium or low tness nodes 395 15/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation (typically mid-size rms or startups). This leads to the high value of rejections obtained from 396 the maximum likelihood estimation. But if agents with a high tness initiate an alliance, agents 397 with medium or low tness are likely to accept this invitation. On the other hand, agents with 398 low tness are not very selective to refuse any invitation because their threshold utility uthr i is 399 rather low, also as a consequence of the small value of ˆal . 400 The good match of our agent-based model with the empirical observations allows us to draw 401 some conclusions about the formation of real R&D alliances, for which no data is available. 402 As we have seen, alliances are more likely initiated by an incumbent rm of high tness which 403 directs its interest toward a mid-size company or a startup. At the same time, the bottleneck 404 in establishing new alliances is probably on the initiator's side, which has to take rejections and 405 keep looking for new partners until it nds the right one. 406 Our agent-based model was developed to reproduce the empirical distribution of alliance sizes . 407 One could be interested to know whether this model, using the parameter from the maximum 408 likelihood estimation, is also able to reproduce other features of the observed topology of the 409 R&D network. This is not the aim of the paper, but we can comment at least on the degree 410 distribution which was analyzed already by Tomasello et al. (2014). Degree refers to the number 411 of collaboration partners of an agent, not to the number of alliances the agent is involved. As such, 412 degree is not independent of the size of an alliance, and indeed the empirical degree distribution 413 was also shown to be right skewed and very broad. 414 However, we argue that the degree distribution cannot simply be obtained from our agent-based 415 model because this does not take degree-degree correlations into account. Assortativity reects the 416 tendency of agents with high degree to form alliances with other agents with high degree, whereas 417 dissortativity would indicate that agents with high degree have the tendency to form alliances 418 with agents of low degree. Such degree-degree correlations have been detected by (Tomasello 419 et al. , 2016) both for sectoral R&D networks and for the aggregated R&D network used in this 420 paper. They play a role in particular for agents with high degree. Therefore, we can assume that 421 our agent-based model will be able to reproduce the right skewed and broad degree distribution, 422 but becomes increasingly worse in the range of larger degrees. 423 To conclude, our agent-based model provides a considerable step forward in identifying the real 424 mechanisms for alliance formation (Ahuja, 2000). In particular, with the distribution of alliance 425 sizes we are able to reproduce a feature that has received some attention in the existing literature, 426 but never a conclusive explanation. Our model can be used for stochastic agent-based simulations, 427 it also provides an analytical solution that considerably reduces the computational eort. We 428 emphasize that it is rather rare to obtain an analytic description of an agent based model. Our 429 derivations also apply to cases with dierent cost functions and are thus quite general. This 430 should inspire further agent based modelling approaches. 431 Our agent-based model is fully calibrated and validated against real data from the global interrm 432 R&D network. It shows the emergence of a broad, right-skewed distribution of alliance sizes, 433 taking into account a heterogeneous tness distribution of agents. On the methodological side, 434 16/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation our study provides an approach to infer the correct parameter values for the agent-based model, to 435 interpret them and check their consistency with reality. Like for any agent-based model approach, 436 we cannot conclude that our model is the only one able to explain and reproduce the alliance 437 size distribution. However, the very good match with reality is a clear sign of plausibility for the 438 set of agent rules that we propose, thus providing us with new insights into the micro dynamics 439 of alliance formation. 440 Acknowledgements 441 M.V.T. and F.S. acknowledge nancial support from the Swiss National Science Foundation, 442 through grant 100014_126865, R&D Network Life Cycles. M.V.T. and F.S. acknowledge nan443 cial support from the Seed Project SP-RC 01-15 Performance and resilience of collaboration 444 networks, granted by the ETH Risk Center of ETH Zurich. The authors thank Ryan Murphy 445 for comments on an early version of this paper. 446 References 447 Ahuja, G. (2000). The duality of collaboration: Inducements and opportunities in the formation 448 of interrm linkages. Strategic management journal 21(3) , 317343. 449 Barabasi, A.-L. (2005). The origin of bursts and heavy tails in human dynamics. Nature 450 435(7039) , 207211. 451 Bianconi, G.; Barabási, A. (2001). Competition and multiscaling in evolving networks. EPL 452 (Europhysics Letters) 54 , 436. 453 Bitzer, J.; Geishecker, I. (2010). Who contributes voluntarily to OSS? An investigation among 454 German IT employees. Research Policy 39(1) , 165172. 455 Csardi, G.; Nepusz, T. (2006). The igraph software package for complex network research. 456 InterJournal, Complex Systems 1695(5) . 457 Durkheim, E. (2014). The division of labor in society . Simon and Schuster. 458 Frenz, M.; Ietto-Gillies, G. (2009). The impact on innovation performance of dierent sources 459 of knowledge: Evidence from the UK Community Innovation Survey. Research Policy 38(7) , 460 11251135. 461 Fruchterman, T.; Reingold, E. (1991). Graph Drawing by Force-directed Placement. Software462 Practice and Experience 21(11) , 11291164. 463 Hagedoorn, J. (2002). Inter-rm R&D partnerships: an overview of major trends and patterns 464 since 1960. Research policy 31(4) , 477492. 465 17/18
M. V. Tomasello, R. Burkholz, F. Schweitzer: Modeling the formation of R&D alliances: An agent-based model with empirical validation Hoang, H.; Rothaermel, F. T. (2005). The eect of general and partner-specic alliance experience 466 on joint R&D project performance. Academy of Management Journal 48(2) , 332345. 467 Katz, J.; Martin, B. R. (1997). What is research collaboration? Research Policy 26(1) , 118. 468 Kim, C.; Song, J. (2007). Creating new technology through alliances: An empirical investigation 469 of joint patents. Technovation 27 , 461470. 470 Lakhani, K.; Wolf, R. G. (2003). Why Hackers Do What They Do: Understanding Motivation 471 and Eort in Free/Open Source Software Projects. SSRN Electronic Journal . 472 Pastor-Satorras, R.; Vazquez, A.; Vespignani, A. (2001). Dynamical and Correlation Properties 473 of the Internet. Physical Review Letters 87 . 474 Perra, N.; Goncalves, B.; Pastor-Satorras, R.; Vespignani, A. (2012). Activity driven modeling 475 of time varying networks. Scientic Reports 2 , 469. 476 Scholtes, I.; Mavrodiev, P.; Schweitzer, F. (2016). From Aristotle to Ringelmann: A large-scale 477 analysis of team productivity and coordination in Open Source Software projects. Empirical 478 Software Engineering 21(2) , 642683. 479 Tomasello, M. V.; Napoletano, M.; Garas, A.; Schweitzer, F. (2016). The Rise and Fall of R&D 480 Networks. ICC - Industrial and Corporate Change 26(4) , 617646. 481 Tomasello, M. V.; Perra, N.; Tessone, C. J.; Karsai, M.; Schweitzer, F. (2014). The Role of 482 Endogenous and Exogenous Mechanisms in the Formation of R&D Networks. Scientic Reports 483 4 , 5679. 484 Tomasello, M. V.; Vaccario, G.; Schweitzer, F. (2017). Data-driven modeling of collaboration 485 networks: A cross-domain analysis. EPJ Data Science 6 , 22. 486 18/18
Please note: You are most sincerely encouraged to participate in the open assessment of this discussion paper. You can do so by either recommending the paper or by posting your comments. Please go to: http://www.economics-ejournal.org/economics/discussionpapers/2017-107 The Editor © Author(s) 2017. Licensed under the Creative Commons License - Attribution 4.0 International (CC BY 4.0).