Full text
The network structure of success: Evidence from an empirical study of European patents Alex Stivala1Alessandro Lomi1,2 1Universit`a della Svizzera italiana, Lugano, Switzerland 2The University of Exeter Business School, Exeter, UK Fifth Annual Australian Social Network Analysis Conference (ASNAC 2020). On-line conference hosted by the University of Western Australia. November 26-27, 2020. 1 / 35
Introduction IOne measure of the “success” of a patent is the number of citations it receives from other patents. IThese are known as “forward citations”, and is just the in-degree in the citation network. IInnovation involves the combination of knowledge in different ways. IBut not all possible combinations of knowledge are equally likely to succeed. So what factors contribute to success? IWe will use the ideas of categorical contrast and niche width (Hannan et al., 2007; Kov´acs and Hannan, 2010, 2015), as well as a new measure of technology class boundary crossing, to try to answer this question. IWe will use both negative binomial regression and ERGM, as appropriate, to test hypotheses. 2 / 35
Contrast (Kov´acs and Hannan, 2010) IThe contrast of a category captures the idea of sharpness or fuzziness of category boundaries: IA category has high contrast (sharp boundaries) if it is seldom assigned low or moderate levels of category membership. IA category has lower contrast (fuzzier boundaries) as partial membership is more common. IA technology class that is seldom assigned together with other classes to a patent has high contrast. IA technology class that is frequently assigned together with other classes to a patent has low contrast. IContrast is defined as the average grade-of-membership (GoM) in a category, for those with nonzero GoM. IWhen the category membership is binary (as in patent technology classes), then for each patent GoM is just 0 if the patent does not have that class, and 1/Kpwhen it does, where Kpis the number of categories assigned to patent p. 3 / 35
Niche width (Hannan et al., 2007; Kov´acs and Hannan, 2010) INiche width captures the idea of breadth: IA patent with high niche width spans many categories (technology classes); it is generalist. IA patent with a single technology class has a niche width of 0; it is specialized. IThe niche width of a patent is the Simpson diversity index of the GoM vector. IEquivalently, 1 −Hwhere His the Herfindahl concentration index. IFor binary memberships as used here, niche width is just 1−1/Kp. 4 / 35
Assigned technology classes or cited technology classes? IPatents are assigned technology classes by the patent office. IIn our data, multiple classes can be assigned. ISo GoM can be defined in two ways: IBy the set of technology classes assigned to a patent. IBy the set of technology classes assigned to the patents cited by a patent. IThe latter is claimed to better capture the combination of knowledge by a patent (Gruber et al., 2013; Ferguson and Carnabuci, 2017). IWe will use both. IWhen niche width is defined by classes of cited patents, it is the same as the “originality” of Trajtenberg et al. (1997); Hall et al. (2001). 5 / 35
Class crossing ratio I INiche width is a monotonic function of the number of technology classes, so it captures just breadth and not diversity as such. IWe define the class crossing ratio to capture a particular idea of diversity or “boundary crossing”: IConsider each citation as an arc between each of the classes in the citing patent to each of the classes in the cited patents. IThe class crossing ratio is the ratio of the number of these virtual arcs which join different classes, to the total number of virtual arcs. ISo class crossing ratio is high when a patent cites patents that have different technology classes than those it is assigned itself. 6 / 35
Class crossing ratio II IThis is conceptually different from the typicality measure of Ferguson and Carnabuci (2017) which measures similarity among sets of technology classes assigned to the cited patents only, with a Jaccard index. IIt is also different from Jaccard similarity between classes of citing patent and union of classes of cited patents. 7 / 35
Class crossing ratio illustration 1 abc a c d b X Y Z Class crossing ratio = 9/12 = 0.75 J(X,Y∪Z)=3/4=0.75 8 / 35
Class crossing ratio illustration 2 abc a c da X Y Z Class crossing ratio = 9/12 = 0.75 J(X,Y∪Z)=2/4=0.5 9 / 35
Summary statistics of IPC sections IPC Section Description N A Human necessities 405804 B Performing operations; transporting 497492 C Chemistry; metallurgy 464874 D Textiles; paper 54695 E Fixed constructions 78438 F Mechanical engineering; lighting; heating ... 227017 G Physics 477022 H Electricity 438685 Y General ... 0 Note that a patent need not be assigned to only a single section; the sections are not mutually exclusive. 16 / 35
Summary statistics of the patent citation network Description N Components Giant Mean Density component degree EPO (full) 4903886 746741 3789545 2.30 0.0000002 EPO (subgraph) 1933231 1119794 673306 1.15 0.0000003 Description Reciprocity Clustering Assortativity coefficient coefficient EPO (full) 0.0005 0.03125 0.08300 EPO (subgraph) 0.0025 0.07862 0.13231 The “full” network is the network containing not only patents in the data set, but also nodes representing patents outside the data set, but which are cited by a patent in the data set. The “subgraph” network is the network induced by only those nodes in the data set itself. 17 / 35
Distribution of contrast values of technology classes Contrast Density 0.3 0.4 0.5 0.6 0.7 0.8 01234 The highest value of contrast (0.812) is for A43 (footwear), and the lowest value (0.250) is for C99 (chemistry; metallurgy). 18 / 35
Distribution of maximum contrast value of patents Max. contrast Density 0.3 0.4 0.5 0.6 0.7 0.8 0 2 4 6 8 10 19 / 35
Distribution of class crossing ratio of patents Class crossing ratio Density 0.0 0.2 0.4 0.6 0.8 1.0 0123456 The class crossing ratio of a patent is the number of backward citations that represent a direct citation from a class assigned to the patent, to a different class in the cited patent, divided by the total number of possible class citations (to both the same or different classes). 20 / 35
Distribution of technology class Jaccard similarity Jaccard similarity between technology classses and directly cited technology classes Density 0.0 0.2 0.4 0.6 0.8 1.0 0 2 4 6 8 Distribution of the Jaccard similarity between the sets of technology classes assigned to a patent, and the union of the sets of technology classes assigned to the backward citations (directly cited patents) of the patent. N= 650511, median = 0.667, mean = 0.674, sd = 0.307. 21 / 35
Negative binomial models, citations as response variable I Model 1 Model 2 Model 3 App. Year [base 1978] −0.12 (0.00)∗∗∗ −0.12 (0.00)∗∗∗ −0.12 (0.00)∗∗∗ Section A −0.24 (0.01)∗∗∗ −0.31 (0.01)∗∗∗ −0.31 (0.01)∗∗∗ Section B 0.04 (0.00)∗∗∗ −0.05 (0.01)∗∗∗ −0.04 (0.01)∗∗∗ Section C 0.25 (0.00)∗∗∗ 0.15 (0.01)∗∗∗ 0.14 (0.01)∗∗∗ Section D 0.07 (0.01)∗∗∗ −0.00 (0.01) −0.01 (0.01) Section E −0.39 (0.01)∗∗∗ −0.46 (0.01)∗∗∗ −0.45 (0.01)∗∗∗ Section F −0.07 (0.01)∗∗∗ −0.15 (0.01)∗∗∗ −0.14 (0.01)∗∗∗ Section G 0.19 (0.00)∗∗∗ 0.11 (0.01)∗∗∗ 0.10 (0.01)∗∗∗ Section H 0.17 (0.01)∗∗∗ 0.09 (0.01)∗∗∗ 0.09 (0.01)∗∗∗ Pub. Language German −0.29 (0.00)∗∗∗ −0.29 (0.00)∗∗∗ −0.31 (0.00)∗∗∗ Pub. Language French −0.31 (0.01)∗∗∗ −0.31 (0.01)∗∗∗ −0.32 (0.01)∗∗∗ Backward citations (all) 0.17 (0.00)∗∗∗ 0.17 (0.00)∗∗∗ 0.17 (0.00)∗∗∗ Max. contrast −2.36 (0.44)∗∗∗ −2.70 (0.44)∗∗∗ −2.68 (0.44)∗∗∗ Max. contrast23.45 (0.34)∗∗∗ 3.65 (0.34)∗∗∗ 3.61 (0.35)∗∗∗ Niche width 0.22 (0.01)∗∗∗ 0.23 (0.01)∗∗∗ Appplicant Switzerland −0.05 (0.02)∗∗ Inventor Switzerland −0.07 (0.03)∗∗ Appplicant Switzerland×Inventor Switzerland 0.27 (0.03)∗∗∗ Cited max. contrast Cited max. contrast2 Cited niche width AIC 3331171.47 3330604.41 3248519.42 BIC 3331371.01 3330816.43 3248768.46 Log Likelihood −1665569.73 −1665285.20 −1624239.71 Deviance 1181391.34 1181445.64 1157693.65 Num. obs. 1927639 1927639 1889616 22 / 35
Negative binomial models, citations as response variable II Model 4 Model 5 Model 6 App. Year [base 1978] −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ Section A −0.14 (0.01)∗∗∗ −0.14 (0.01)∗∗∗ −0.14 (0.01)∗∗∗ Section B −0.00 (0.01) −0.01 (0.01) −0.00 (0.01) Section C 0.11 (0.01)∗∗∗ 0.10 (0.01)∗∗∗ 0.10 (0.01)∗∗∗ Section D 0.01 (0.01) 0.01 (0.01) 0.01 (0.01) Section E −0.35 (0.01)∗∗∗ −0.35 (0.01)∗∗∗ −0.35 (0.01)∗∗∗ Section F −0.04 (0.01)∗∗∗ −0.05 (0.01)∗∗∗ −0.04 (0.01)∗∗∗ Section G 0.06 (0.01)∗∗∗ 0.06 (0.01)∗∗∗ 0.05 (0.01)∗∗∗ Section H −0.03 (0.01)∗∗∗ −0.03 (0.01)∗∗∗ −0.03 (0.01)∗∗∗ Pub. Language German −0.34 (0.01)∗∗∗ −0.34 (0.01)∗∗∗ −0.36 (0.01)∗∗∗ Pub. Language French −0.34 (0.01)∗∗∗ −0.34 (0.01)∗∗∗ −0.34 (0.01)∗∗∗ Backward citations (all) 0.04 (0.00)∗∗∗ 0.04 (0.00)∗∗∗ 0.04 (0.00)∗∗∗ Max. contrast −2.48 (0.72)∗∗∗ −2.63 (0.71)∗∗∗ −2.63 (0.72)∗∗∗ Max. contrast22.88 (0.57)∗∗∗ 3.21 (0.57)∗∗∗ 3.19 (0.58)∗∗∗ Niche width 0.23 (0.01)∗∗∗ 0.18 (0.01)∗∗∗ 0.18 (0.01)∗∗∗ Appplicant Switzerland −0.07 (0.02)∗∗ Inventor Switzerland −0.05 (0.03) Appplicant Switzerland×Inventor Switzerland 0.23 (0.04)∗∗∗ Cited max. contrast 0.01 (0.75) −0.02 (0.75) 0.01 (0.76) Cited max. contrast20.78 (0.59) 0.55 (0.59) 0.53 (0.60) Cited niche width 0.11 (0.01)∗∗∗ 0.11 (0.01)∗∗∗ AIC 1615185.10 1615025.57 1579868.81 BIC 1615401.42 1615253.28 1580130.28 Log Likelihood −807573.55 −807492.79 −789911.40 Deviance 548718.44 548738.71 538346.76 Num. obs. 650434 650434 639387 23 / 35
Negative binomial models with secondary contrast I Model 1 Model 2 Model 3 App. Year [base 1978] −0.10 (0.00)∗∗∗ −0.10 (0.00)∗∗∗ −0.10 (0.00)∗∗∗ Section A −0.19 (0.01)∗∗∗ −0.21 (0.01)∗∗∗ −0.21 (0.01)∗∗∗ Section B 0.02 (0.01)∗∗ −0.00 (0.01) −0.00 (0.01) Section C 0.11 (0.01)∗∗∗ 0.07 (0.01)∗∗∗ 0.07 (0.01)∗∗∗ Section D 0.07 (0.02)∗∗∗ 0.04 (0.02)∗0.04 (0.02)∗ Section E −0.33 (0.02)∗∗∗ −0.34 (0.02)∗∗∗ −0.34 (0.02)∗∗∗ Section F 0.03 (0.01)∗∗ 0.01 (0.01) 0.01 (0.01) Section G 0.08 (0.01)∗∗∗ 0.06 (0.01)∗∗∗ 0.05 (0.01)∗∗∗ Section H −0.02 (0.01) −0.04 (0.01)∗∗∗ −0.04 (0.01)∗∗∗ Pub. Language German −0.27 (0.01)∗∗∗ −0.27 (0.01)∗∗∗ −0.29 (0.01)∗∗∗ Pub. Language French −0.26 (0.01)∗∗∗ −0.26 (0.01)∗∗∗ −0.27 (0.01)∗∗∗ Backward citations (subgraph) 0.16 (0.00)∗∗∗ 0.16 (0.00)∗∗∗ 0.16 (0.00)∗∗∗ Max. contrast −1.00 (0.97) −1.43 (0.98) −1.67 (0.99) Max. contrast22.39 (0.76)∗∗ 2.73 (0.76)∗∗∗ 2.90 (0.77)∗∗∗ Class crossing ratio 3.01 (0.25)∗∗∗ 2.39 (0.26)∗∗∗ 2.41 (0.27)∗∗∗ Class crossing ratio2−2.34 (0.19)∗∗∗ −2.00 (0.19)∗∗∗ −2.01 (0.19)∗∗∗ Secondary contrast −5.73 (0.77)∗∗∗ −6.03 (0.77)∗∗∗ −5.77 (0.78)∗∗∗ Secondary contrast25.14 (0.67)∗∗∗ 5.25 (0.67)∗∗∗ 5.03 (0.68)∗∗∗ Niche width 0.49 (0.05)∗∗∗ 0.50 (0.05)∗∗∗ Appplicant Switzerland −0.04 (0.03) Inventor Switzerland −0.08 (0.05) Appplicant Switzerland×Inventor Switzerland 0.23 (0.06)∗∗∗ Cited max. contrast Cited max. contrast2 Cited secondary contrast Cited secondary contrast2 Cited niche width AIC 761913.27 761804.30 745025.12 BIC 762124.45 762026.05 745278.11 Log Likelihood −380936.63 −380881.15 −372488.56 Deviance 251830.34 251837.63 247074.61 Num. obs. 284767 284767 279728 24 / 35
Negative binomial models with secondary contrast II Model 4 Model 5 Model 6 App. Year [base 1978] −0.10 (0.00)∗∗∗ −0.10 (0.00)∗∗∗ −0.10 (0.00)∗∗∗ Section A −0.21 (0.01)∗∗∗ −0.21 (0.01)∗∗∗ −0.21 (0.01)∗∗∗ Section B 0.00 (0.01) 0.00 (0.01) 0.00 (0.01) Section C 0.07 (0.01)∗∗∗ 0.07 (0.01)∗∗∗ 0.06 (0.01)∗∗∗ Section D 0.02 (0.02) 0.02 (0.02) 0.02 (0.02) Section E −0.34 (0.02)∗∗∗ −0.35 (0.02)∗∗∗ −0.35 (0.02)∗∗∗ Section F 0.02 (0.01) 0.02 (0.01) 0.03 (0.01)∗ Section G 0.06 (0.01)∗∗∗ 0.06 (0.01)∗∗∗ 0.06 (0.01)∗∗∗ Section H −0.05 (0.01)∗∗∗ −0.05 (0.01)∗∗∗ −0.05 (0.01)∗∗∗ Pub. Language German −0.26 (0.01)∗∗∗ −0.26 (0.01)∗∗∗ −0.28 (0.01)∗∗∗ Pub. Language French −0.26 (0.02)∗∗∗ −0.26 (0.02)∗∗∗ −0.26 (0.02)∗∗∗ Backward citations (subgraph) 0.14 (0.00)∗∗∗ 0.14 (0.00)∗∗∗ 0.14 (0.00)∗∗∗ Max. contrast −0.45 (1.42) −0.22 (1.43) −0.54 (1.45) Max. contrast21.61 (1.12) 1.49 (1.13) 1.75 (1.14) Class crossing ratio 1.76 (0.32)∗∗∗ 1.36 (0.32)∗∗∗ 1.35 (0.33)∗∗∗ Class crossing ratio2−1.66 (0.23)∗∗∗ −1.46 (0.24)∗∗∗ −1.46 (0.24)∗∗∗ Secondary contrast −4.76 (0.95)∗∗∗ −4.99 (0.95)∗∗∗ −4.85 (0.96)∗∗∗ Secondary contrast24.13 (0.83)∗∗∗ 4.46 (0.83)∗∗∗ 4.35 (0.84)∗∗∗ Niche width 0.64 (0.06)∗∗∗ 0.62 (0.06)∗∗∗ 0.63 (0.06)∗∗∗ Appplicant Switzerland −0.04 (0.04) Inventor Switzerland −0.07 (0.06) Appplicant Switzerland×Inventor Switzerland 0.25 (0.07)∗∗∗ Cited max. contrast −3.32 (1.54)∗−3.63 (1.55)∗−3.44 (1.57)∗ Cited max. contrast23.09 (1.20)∗3.29 (1.21)∗∗ 3.13 (1.23)∗ Cited secondary contrast −1.47 (1.07) −1.52 (1.07) −1.76 (1.08) Cited secondary contrast21.31 (0.92) 1.15 (0.92) 1.35 (0.93) Cited niche width 0.31 (0.05)∗∗∗ 0.31 (0.05)∗∗∗ AIC 597762.41 597712.34 584577.96 BIC 598019.71 597979.92 584875.90 Log Likelihood −298856.21 −298830.17 −292259.98 Deviance 195603.94 195596.40 191914.04 Num. obs. 217890 217890 214014 25 / 35
Results for hypotheses II H2 Success increases with maximum contrast of technology classes. IPartly confirmed: there is a quadratic relationship between success and max. contrast, with success decreasing with maximum contrast up to a point, but increasing thereafter. IThis applies for both maximum contrast of a patent’s classes, and of maximum contrast of its cited patents’ classes. IThe ERGM models also show a similar pattern with the Receiver effect on max. contrast. H3 But spanning high contrast categories makes success less likely. IPartly confirmed: there is a quadratic relationship between success and secondary contrast, with success decreasing with secondary contrast only up to a point, after which it increases. IThere is a similar pattern in the ERGM for the Receiver effect for secondary contrast. H4 Patents with high maximum contrast are unlikely to cite other patents with high maximum contrast. 32 / 35
Results for hypotheses III IContradicted: In the ERGM model the effect for heterophily (AbsDiff) on max. contrast is negative and significant. IDiffSign is not significant. IIt seems that, contrary to H4, there is significant homophily on max. contrast. IIs this a poor test of H4, as it is confounded by patents citing patents with the same technology class? IPositive significant Jaccard similarity of technology class sets in all models in which it is included (unsurprising: patents cite other patents in the same technology classes). INote ERGM parameter estimation does not converge well with both Jaccard similarity of technology classes and the AbsDiff effect for max. contrast included. H5 (Geographical knowledge spillover): citations are more likely to be geographically localized. IConfirmed: The effect for Jaccard similarity is positive and significant for both applicant countries and inventor countries in all ERGM models. 33 / 35
Acknowledgments IThis work was funded by Swiss National Science Foundation NRP 75 Big Data project 167326 “The Global Structure of Knowledge Networks: Data, Models and Empirical Results”. IWe thank Mr Manajit Chakraborty and Prof. Fabio Crestani for assisting with access to patent data. IWe used the high performance computing cluster at the Institute of Computational Science, Universit`a della Svizzera italiana, for all data processing and statistical computations. 34 / 35
Unpublished work IThis is unpublished work (as of November 2020). IDetails including methods and references are in the “hidden bonus slides” after this one. II will make these slides available on my website: Ihttps://sites.google.com/site/alexdstivala/home/ conferences 35 / 35
Hidden bonus slides 1 / 27
CPC technology sections AHuman necessities BPerforming operations; transporting CChemistry; metallurgy DTextiles; paper EFixed constructions FMechanical engineering; lighting; heating; weapons; blasting engines or pumps GPhysics HElectricity YGeneral tagging of new technological developments ... https: //www.epo.org/searching-for-patents/helpful-resources/first-time-here/classification/cpc.html 2 / 27
Jaccard similarity The Jaccard similarity 0 ≤J(A,B)≤1 between two sets is the size of their intersection over the size of their union: J(A,B) = |A∩B| |A∪B| If |A∪B|= 0 i.e. Aand Bare both empty, then define J(A,B) = 1. 3 / 27
Class crossing ratio example IAssume patent X has classes a,b,c and it cites patent Y with classes a,c,d and patent Z with class b only IWe consider the total of 3 ×3+3×1 = 12 virtual ties (a–a, a–c, a–d, b–a, b–c, ... , c–b) IOf these 12 virtual ties 9 are “boundary crossing” (a–c, a–d, b–a, ..., but not a–a, c–c, b–b, ...) ISo we would give it a boundary crossing score of 9/12 = 0.75 I(In R we can do this using the vector outer product.) INote that this is like a kind of generalized E-I index (Krackhardt and Stern, 1988) IAlthough it is in [0,1] not [−1,+1] — to make it more like E-I index we would have the numerator as (mismatching - matching) not just mismatching, applicable to sets of categories on nodes, rather than just a simple nodal categorical variable. 4 / 27
Contrast share IContrast share is the ratio of the maximum contrast of assigned categories to their sum (Kov´acs and Hannan, 2010). IIn our data, contrast share is highly inversely correlated with niche width, so we use only niche width. 5 / 27
Summary statistics of publication languages Language N English 1355416 German 435373 French 141397 NA 1045 6 / 27
Linear correlation between cited niche width and class crossing ratio of patents 13 / 27
Linear correlation between class crossing ratio and Jaccard similarity between technology classes and union of directly cited technology classes 14 / 27
Methods I IPower law and log-normal distributions were fitted using the methods of Clauset et al. (2009) implemented in the poweRlaw package (Gillespie, 2015). INegative binomial regression models were estimated using the MASS (Venables and Ripley, 2002) and formatted with the texreg (Leifeld, 2013) packages in R (R Core Team, 2016). Robust standard errors (Hinkley, 1977; MacKinnon and White, 1985) were estimated with the sandwich (Zeileis, 2004, 2006) and lmtest (Zeileis and Hothorn, 2002) packages in R. Residual diagnostics from the DHARMa R package (Hartig, 2019). IERGM models were estimated with EstimNetDirected (Byshkin et al., 2018; Borisenko et al., 2020; Stivala et al., 2019b). 15 / 27
Methods II IThe ERGM DiffSign parameter to control for citation temporal direction was introduced by Graham et al. (2018); McLevey et al. (2018) and also used in Stivala et al. (2019a). IIn the full 4.9 million node network, only 1.9 million nodes represent patents in the data set. The remaining 3 million nodes (61% of the nodes) represent patents cited by one of those in the data set, but for which we have no data. IAn ERGM model with NA for all values on those 3 million nodes does not converge (unlike the 3.7 million node NBER patent citation network where only 27% of the nodes have no data in Stivala et al. (2019a)). ISo conditional estimation based on snowball sampling structure (Pattison et al., 2013; Stivala et al., 2016) was used. The 1.9 million nodes (39%) with data are treated as wave 0 (seeds) and the remaining 3 million nodes treated as wave 1, and estimation is conditional on this structure. 16 / 27
Negative binomial models with class crossing ratio I Model 1 Model 2 Model 3 App. Year [base 1978] −0.13 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ Section A −0.20 (0.01)∗∗∗ −0.06 (0.01)∗∗∗ −0.13 (0.01)∗∗∗ Section B 0.12 (0.00)∗∗∗ 0.11 (0.01)∗∗∗ 0.03 (0.01)∗∗ Section C 0.06 (0.00)∗∗∗ 0.14 (0.01)∗∗∗ 0.04 (0.01)∗∗∗ Section D 0.08 (0.01)∗∗∗ 0.10 (0.01)∗∗∗ 0.02 (0.01) Section E −0.21 (0.01)∗∗∗ −0.21 (0.01)∗∗∗ −0.28 (0.01)∗∗∗ Section F 0.09 (0.01)∗∗∗ 0.09 (0.01)∗∗∗ 0.00 (0.01) Section G 0.13 (0.00)∗∗∗ 0.13 (0.01)∗∗∗ 0.05 (0.01)∗∗∗ Section H 0.14 (0.01)∗∗∗ 0.06 (0.01)∗∗∗ −0.02 (0.01)∗ Pub. Language German −0.25 (0.00)∗∗∗ −0.33 (0.01)∗∗∗ −0.33 (0.01)∗∗∗ Pub. Language French −0.27 (0.01)∗∗∗ −0.33 (0.01)∗∗∗ −0.33 (0.01)∗∗∗ Backward citations (subgraph) 0.43 (0.00)∗∗∗ 0.16 (0.00)∗∗∗ 0.17 (0.00)∗∗∗ Max. contrast −1.74 (0.43)∗∗∗ −3.34 (0.56)∗∗∗ −4.04 (0.56)∗∗∗ Max. contrast22.67 (0.34)∗∗∗ 4.01 (0.44)∗∗∗ 4.43 (0.44)∗∗∗ Class crossing ratio 0.33 (0.02)∗∗∗ 0.18 (0.03)∗∗∗ Class crossing ratio2−0.48 (0.03)∗∗∗ −0.42 (0.03)∗∗∗ Niche width 0.30 (0.02)∗∗∗ Cited max. contrast Cited max. contrast2 Cited niche width Appplicant Switzerland Inventor Switzerland Appplicant Switzerland×Inventor Switzerland AIC 3318050.97 1610355.84 1609898.86 BIC 3318250.52 1610560.78 1610115.18 Log Likelihood −1659009.49 −805159.92 −804930.43 Deviance 1199294.95 549422.32 549407.66 Num. obs. 1927639 650434 650434 17 / 27
Negative binomial models with class crossing ratio II Model 4 Model 5 Model 6 App. Year [base 1978] −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ Section A −0.13 (0.01)∗∗∗ −0.13 (0.01)∗∗∗ −0.13 (0.01)∗∗∗ Section B 0.03 (0.01)∗∗∗ 0.02 (0.01)∗∗ 0.03 (0.01)∗∗∗ Section C 0.04 (0.01)∗∗∗ 0.03 (0.01)∗∗∗ 0.03 (0.01)∗∗ Section D 0.02 (0.01) 0.02 (0.01) 0.01 (0.01) Section E −0.27 (0.01)∗∗∗ −0.28 (0.01)∗∗∗ −0.28 (0.01)∗∗∗ Section F 0.01 (0.01) 0.00 (0.01) 0.01 (0.01) Section G 0.05 (0.01)∗∗∗ 0.05 (0.01)∗∗∗ 0.05 (0.01)∗∗∗ Section H −0.02 (0.01)∗−0.02 (0.01)∗−0.02 (0.01)∗ Pub. Language German −0.33 (0.01)∗∗∗ −0.33 (0.01)∗∗∗ −0.34 (0.01)∗∗∗ Pub. Language French −0.33 (0.01)∗∗∗ −0.33 (0.01)∗∗∗ −0.33 (0.01)∗∗∗ Backward citations (subgraph) 0.16 (0.00)∗∗∗ 0.16 (0.00)∗∗∗ 0.16 (0.00)∗∗∗ Max. contrast −2.81 (0.73)∗∗∗ −3.18 (0.73)∗∗∗ −3.21 (0.74)∗∗∗ Max. contrast23.14 (0.58)∗∗∗ 3.59 (0.59)∗∗∗ 3.60 (0.59)∗∗∗ Class crossing ratio 0.14 (0.03)∗∗∗ −0.11 (0.03)∗∗∗ −0.11 (0.03)∗∗∗ Class crossing ratio2−0.41 (0.03)∗∗∗ −0.30 (0.03)∗∗∗ −0.30 (0.03)∗∗∗ Niche width 0.34 (0.02)∗∗∗ 0.38 (0.02)∗∗∗ 0.38 (0.02)∗∗∗ Cited max. contrast −0.69 (0.76) −0.98 (0.77) −0.91 (0.78) Cited max. contrast21.01 (0.60) 1.00 (0.61) 0.94 (0.62) Cited niche width 0.20 (0.01)∗∗∗ 0.20 (0.01)∗∗∗ Appplicant Switzerland −0.06 (0.02)∗∗ Inventor Switzerland −0.04 (0.03) Appplicant Switzerland×Inventor Switzerland 0.21 (0.04)∗∗∗ AIC 1609786.42 1609545.75 1574445.58 BIC 1610025.52 1609796.23 1574729.79 Log Likelihood −804872.21 −804750.87 −787197.79 Deviance 549418.03 549427.11 539036.16 Num. obs. 650434 650434 639387 18 / 27
Negative binomial models using cited contrast only I Model 1 Model 2 Model 3 App. Year [base 1978] −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ Section A 0.04 (0.01)∗∗∗ −0.01 (0.01) −0.01 (0.01) Section B 0.15 (0.01)∗∗∗ 0.13 (0.01)∗∗∗ 0.14 (0.01)∗∗∗ Section C 0.07 (0.01)∗∗∗ 0.12 (0.01)∗∗∗ 0.13 (0.01)∗∗∗ Section D 0.10 (0.01)∗∗∗ 0.07 (0.02)∗∗∗ 0.08 (0.02)∗∗∗ Section E −0.04 (0.01)∗∗ −0.14 (0.02)∗∗∗ −0.14 (0.02)∗∗∗ Section F 0.11 (0.01)∗∗∗ 0.13 (0.01)∗∗∗ 0.13 (0.01)∗∗∗ Section G 0.17 (0.01)∗∗∗ 0.17 (0.01)∗∗∗ 0.17 (0.01)∗∗∗ Section H 0.21 (0.01)∗∗∗ 0.14 (0.01)∗∗∗ 0.14 (0.01)∗∗∗ Pub. Language German −0.34 (0.01)∗∗∗ −0.32 (0.01)∗∗∗ −0.32 (0.01)∗∗∗ Pub. Language French −0.33 (0.01)∗∗∗ −0.32 (0.01)∗∗∗ −0.32 (0.01)∗∗∗ Backward citations (subgraph) 0.17 (0.00)∗∗∗ 0.15 (0.00)∗∗∗ 0.15 (0.00)∗∗∗ Class crossing ratio 0.32 (0.02)∗∗∗ Class crossing ratio2−0.49 (0.03)∗∗∗ Cited max. contrast −1.22 (0.96) −1.05 (0.96) Cited max. contrast21.94 (0.74)∗∗ 1.80 (0.74)∗ Cited secondary contrast −3.88 (0.76)∗∗∗ −3.77 (0.76)∗∗∗ Cited secondary contrast23.25 (0.65)∗∗∗ 3.24 (0.65)∗∗∗ Cited niche width −0.12 (0.03)∗∗∗ Appplicant Switzerland Inventor Switzerland Appplicant Switzerland×Inventor Switzerland AIC 1611861.29 964173.87 964153.35 BIC 1612043.45 964368.85 964359.16 Log Likelihood −805914.64 −482068.94 −482057.67 Deviance 549299.64 322525.69 322527.82 Num. obs. 650434 373983 373983 19 / 27
Negative binomial models using cited contrast only II Model 4 Model 5 App. Year [base 1978] −0.11 (0.00)∗∗∗ −0.11 (0.00)∗∗∗ Section A −0.01 (0.01) 0.01 (0.01) Section B 0.14 (0.01)∗∗∗ 0.17 (0.01)∗∗∗ Section C 0.12 (0.01)∗∗∗ 0.14 (0.01)∗∗∗ Section D 0.07 (0.02)∗∗∗ 0.10 (0.02)∗∗∗ Section E −0.14 (0.02)∗∗∗ −0.11 (0.02)∗∗∗ Section F 0.13 (0.01)∗∗∗ 0.16 (0.01)∗∗∗ Section G 0.17 (0.01)∗∗∗ 0.19 (0.01)∗∗∗ Section H 0.14 (0.01)∗∗∗ 0.17 (0.01)∗∗∗ Pub. Language German −0.33 (0.01)∗∗∗ −0.33 (0.01)∗∗∗ Pub. Language French −0.32 (0.01)∗∗∗ −0.32 (0.01)∗∗∗ Backward citations (subgraph) 0.15 (0.00)∗∗∗ 0.15 (0.00)∗∗∗ Class crossing ratio 0.40 (0.11)∗∗∗ Class crossing ratio2−0.58 (0.09)∗∗∗ Cited max. contrast −0.97 (0.97) −1.01 (0.97) Cited max. contrast21.74 (0.75)∗1.73 (0.75)∗ Cited secondary contrast −3.86 (0.77)∗∗∗ −4.03 (0.78)∗∗∗ Cited secondary contrast23.30 (0.66)∗∗∗ 3.43 (0.66)∗∗∗ Cited niche width −0.12 (0.03)∗∗∗ 0.10 (0.04)∗∗ Appplicant Switzerland −0.06 (0.03) −0.05 (0.03) Inventor Switzerland −0.05 (0.04) −0.05 (0.04) Appplicant Switzerland×Inventor Switzerland 0.23 (0.06)∗∗∗ 0.23 (0.06)∗∗∗ AIC 943423.07 943090.97 BIC 943661.00 943350.52 Log Likelihood −471689.54 −471521.49 Deviance 316546.11 316510.52 Num. obs. 367615 367532 20 / 27
ERGM results, 1 933 231 node network I Effect Model 1 Model 2 Model 3 Arc −13.638 (−13.896,−13.380) −13.932 (−14.224,−13.639) −13.417 (−13.703,−13.131) Isolates −0.182 (−0.253,−0.111) 0.046 (−0.009,0.101) 0.087 (0.023,0.151) Sink −0.763 (−0.848,−0.679) −0.486 (−0.541,−0.430) −0.490 (−0.559,−0.421) Source −0.225 (−0.290,−0.159) −0.223 (−0.269,−0.176) −0.222 (−0.285,−0.160) Popularity spread (AinS) 0.784 (0.697,0.870) 0.757 (0.684,0.831) 0.775 (0.685,0.865) Activity spread (AoutS) 1.238 (1.096,1.381) 0.841 (0.744,0.937) 0.847 (0.728,0.966) Two-path (A2P-T) −0.003 (−0.016,0.010) −0.023 (−0.041,−0.005) −0.029 (−0.046,−0.012) Shared popularity (A2P-D) −0.213 (−0.246,−0.180) −0.119 (−0.146,−0.091) −0.120 (−0.149,−0.092) Shared activity (A2P-U) 0.074 (0.055,0.092) 0.062 (0.047,0.078) 0.057 (0.038,0.076) Sender App. Year [base 1978] 0.454 (0.442,0.465) 0.417 (0.402,0.432) 0.449 (0.431,0.466) Receiver App. Year [base 1978] −0.523 (−0.540,−0.505) −0.505 (−0.525,−0.486) −0.532 (−0.554,−0.509) DiffSign App. Year 1.872 (1.741,2.003) 2.032 (1.916,2.148) 2.050 (1.937,2.164) AbsDiff App. Year −0.625 (−0.650,−0.599) −0.600 (−0.624,−0.576) −0.629 (−0.659,−0.600) Jaccard similarity Applicant countries 0.756 (0.582,0.931) 0.808 (0.646,0.970) 0.786 (0.615,0.957) Jaccard similarity Inventor countries 0.586 (0.432,0.739) 0.573 (0.443,0.702) 0.551 (0.399,0.704) Jaccard similarity Sections 3.837 (3.518,4.156) 1.501 (1.360,1.643) 1.402 (1.269,1.535) Matching Pub. Language 0.102 (0.050,0.154) 0.044 (0.004,0.083) −0.025 (−0.061,0.011) 21 / 27
ERGM results, 1 933 231 node network II Sender Max. contrast −1.409 (−1.596,−1.221) −0.975 (−1.383,−0.567) −3.547 (−3.849,−3.245) Sender Max. contrast2−0.788 (−0.946,−0.630) −1.375 (−1.762,−0.988) 0.668 (0.490,0.847) Receiver Max. contrast −6.515 (−6.802,−6.229) −5.204 (−5.433,−4.975) −8.099 (−8.373,−7.825) Receiver Max. contrast25.169 (4.917,5.420) 3.303 (3.108,3.497) 5.067 (4.788,5.346) Jaccard similarity Classes — 4.563 (4.308,4.817) 5.802 (5.523,6.080) DiffSign Max. contrast 0.008 (−0.001,0.018) — — AbsDiff Max. contrast −15.999 (−17.996,−14.002) — — Sender Niche width — — 1.487 (1.424,1.551) Receiver Niche width — — 1.978 (1.798,2.159) Sender Secondary contrast — — — Sender Secondary contrast2— — — Receiver Secondary contrast — — — Receiver Secondary contrast2— — — Converged runs 20 20 20 Total runs 20 20 20 22 / 27