Patent policy, patent pools, and the accumulation of claims in sequential innovation
Abstract
We present a dynamic model where the accumulation of patents generates an increasing number of claims on sequential innovation. We compare innovation activity under three regimes -patents, no-patents, and patent pools- and find that none of them can reach the first best. We find that the first best can be reached through a decentralized tax-subsidy mechanism, by which innovators receive a subsidy when they innovate, and are taxed with subsequent innovations. This finding implies that optimal transfers work in the exact opposite way as traditional patents. Finally, we consider patents of finite duration and determine the optimal patent length.
Full text
PATENT POLICY, PATENT POOLS, AND THE ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION GAST´ ON LLANES†AND STEFANO TRENTO‡ Abstract. We present a dynamic model where the accumulation of patents generates an increasing number of claims on sequential innovation. We compare innovation activity under three regimes –patents, no-patents, and patent pools– and find that none of them can reach the first best. We find that the first best can be reached through a decentralized tax-subsidy mechanism, by which innovators receive a subsidy when they innovate, and are taxed with subsequent innovations. This finding implies that optimal transfers work in the exact opposite way as traditional patents. Finally, we consider patents of finite duration and determine the optimal patent length. Keywords: Sequential Innovation, Patent Policy, Patent Pools, Anticommons, Double Marginalization, Complementary Monopoly (JEL L13, O31, O34). Date: November 10, 2010. †Escuela de Administraci´on, Pontificia Universidad Cat´olica de Chile, [email protected]. ‡Department of Economics and Economic History, Universitat Aut`onoma de Barcelona, stefano.tren[email protected]. We are grateful to Michele Boldrin for his guidance and advice. We thank Antonio Cabrales, Antonio Ciccone, Marco Celentani, Andres Erosa, Bel´en Jerez, Gerard Llobet, Xavier Vives, and participants of seminars at Harvard Business School, Universidad Carlos III de Madrid, Universitat Aut`onoma de Barcelona, and Center for Applied Economics of Universidad de Chile for useful comments and suggestions. All remaining errors are our responsibility. We gratefully acknowledge financial support from the Ministry of Education of Spain (Llanes, FPU grant AP2003-2204), the Ministry of Science and Technology of Spain (Trento, grant SEJ2006-00538), and the Comunidad Aut´onoma de Madrid (Trento). Stefano Trento is also affiliated to MOVE and Barcelona Graduate School of Economics. 1
2 LLANES AND TRENTO 1. Introduction Knowledge builds upon previous knowledge. This is true for most innovations nowadays, especially in high-tech industries such as molecular biology, plant biotechnology, semiconductors, and software. In some cases, the innovation consists of an improvement of an older version of the same good. In other cases, the research leading to the discovery of the new good depends on the access to research tools, techniques and inputs that were previous innovations themselves. The sequential nature of innovation introduces the issue of how to divide the revenues from the chain of inventions among the different innovators. Suppose two innovations may be introduced sequentially. If the first innovator receives a patent, she may obtain a claim over part of the second innovator’s revenues. Then, the policy maker faces an important trade-off: if the patent covering the first innovation is strong, the second innovation may become unprofitable, but if that patent is weak, it may provide low incentives to introduce the first innovation. The literature on sequential innovation, pioneered by Scotchmer (1991), has studied this problem in depth. Usually, this literature has analyzed the optimal division of profits between two sequential innovators. But what happens when a continuous sequence of innovations exists, each building on all previous inventions? Recent research has suggested the possibility that the accumulation of claims on sequential innovations may generate a tragedy of the anticommons (Heller, 1998; Heller and Eisenberg, 1998). When too many agents have exclusion rights over the use of a common resource, this resource tends to be underutilized, in clear duality with the tragedy of the commons in which too many agents hold rights of use and the resource tends to be overused. In our case, the anticommons could arise if too many patent holders have exclusive claims on separate components of the state-of-theart technology, creating an obstacle for future research. However, the anticommons hypothesis has not yet been studied formally in a dynamic model with endogenous innovation. In particular, patents produce claims on subsequent innovations that may more than compensate for the negative effect of having to pay licensing fees to previous innovators. Several interesting questions arise: what is the net effect of patents on innovation incentives? How should policy parameters be set to maximize social welfare?
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 3 Two streams of literature provide partial answers to these questions.1 The literature on sequential innovation is mainly concerned with substitute innovations, where later innovations are applications or improvements of earlier innovations. This setting allows little room for studying the effects of the accumulation of claims. On the other hand, the literatures of complementary monopoly, patent thickets, and patent pools study the problem of the accumulation of complementary patents, but from a static point of view. As Shapiro (2001) states: “The generic problem inherent in the patent thicket is well understood as a matter of economic theory, at least in its static version.” The main contribution of our paper is to develop a dynamic model to study how the accumulation of complementary claims affects innovation incentives. We relate the literature on sequential innovation to the static literatures of complementary monopoly and patent pools. Extending the analysis of complementary monopoly to a dynamic framework allows us to gain relevant insights into the emergence of patent thickets and the net effect of different patent policy regimes on innovation at different stages of industry maturity. We present a dynamic model to study the division of profits between sequential innovators when each innovation builds on several prior inventions. An infinite sequence of innovations n= 1,2, . . . exists, where innovation ncannot be introduced until innovation n−1 has been introduced. Each innovation has a commercial value (the profit it generates as a final good), which is random and private information of the innovator, and requires a deterministic cost of R&D to be developed. Our model provides a good description of the innovation process in several industries. For example, in the software industry, the first programs were written from scratch and therefore built on little prior knowledge. As more and more programs were developed, they progressively became more dependent on technologies the first programs had introduced. According to Garfinkel et al (1991), modern software programs contain thousands of previously developed mathematical algorithms and techniques. Similar examples can be found in other hightech industries. Formally, our model is a multi-stage game in discrete time with an uncertain end. Interestingly, the probability of reaching the next period is determined endogenously. The equilibrium concept we use is subgame perfect equilibrium with Markovian strategies (Markov perfect equilibrium). 1Read section 1.1 for more details.
4 LLANES AND TRENTO In the first sections of the paper, we study equilibrium dynamics in three scenarios: patents, no patents, and patent pools. With patents, innovation becomes harder and harder with more complex innovations. The probability of innovation goes to 0 as n→ ∞. The probability of innovation is higher than in the static case, but not high enough to prevent the tragedy of the anticommons. Therefore, we show that complementary monopoly inefficiencies, originally explored by Cournot (1838), also extend to a dynamic framework where we remove the bound on the social value complementary monopolists share, as we explain in section 1.1. Without patents, the probability of innovation is constant and depends on the degree of appropriability of the innovation’s commercial value in the final goods sector. When patents protect ideas, the formation of a patent pool increases the probability of innovation for all innovations. Interestingly, the probability of innovation with a pool is constant and higher than it would be in the static case. This result strengthens the findings of Shapiro (2001), Lerner and Tirole (2004), and Llanes and Trento (2009) for static models. We find that pools are dynamically unstable: the temptation to remain outside the pool increases as the sequence of innovations advances, which means early innovators have more incentives than later innovators to enter the pool. The design of a mechanism to solve the pool instability problem, along the lines of Brenner (2009), is beyond the scope of this paper. However, we find that a scheme in which each innovator buys all patent rights from the preceding innovator, instead of paying only for the permission to use the idea, can replicate the patent pool outcome. The complete sale of patent rights will therefore generate higher innovation than licensing. An alternative scheme, leading to the same innovation outcome, is to allow subsequent competition between the licensee and the original licensor. This alternative scheme removes the monopoly power of all but the last patent, eliminating the anticommons effect. We study the optimal innovation policy that maximizes the expected welfare of the sequence of innovations and find that innovation is suboptimal in the three policy regimes. In the no-patents regime, there is a dynamic externality: innovators do not consider how their decisions impact the technological possibilities of future innovators. In the two other policy regimes, the inefficiency stems from asymmetric information and market power: patent holders do not know the exact value of the innovation, but they know its probability distribution. The asymmetric information generates a downward-sloping expected demand for
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 5 the use of ideas, and patent holders’ market power results in a price for old ideas above their marginal cost. We also show that the first best can be reached by decentralizing the innovation decision and implementing a tax-subsidy scheme. Surprisingly, this result holds even if the government and previous innovators do not know the value of the innovation. With respect to the timing of the optimal transfers, we find that innovators should receive a subsidy to innovate, and then be taxed with subsequent innovations. Therefore, optimal transfers work in the exact opposite way as patents, which require an innovator to pay previous inventors upfront and then be compensated by following innovators. These findings extend the results of Erkal and Scotchmer (2007), who show that the scarcity of ideas has a non-trivial impact on the optimal reward to the innovator. In this sense, the two papers are complementary. While Erkal and Scotchmer show that the scarcity of ideas affects the size of the optimal reward, we show that it also affects the timing of this reward. Finally, we find that symmetric information over the value of the innovation leads to the first best. We then turn to the analysis of the optimal patent length. With respect to the optimal patent length, we find that short patents maximize the probability of innovation, because in our model, absent any form of price collusion or agreement between patent holders, reducing patent length is the only way to reduce the complementary monopoly problem. This finding complements the findings of previous papers that focus on the final goods sector and substitute innovations. Another important finding is that the main results of the literature of complementary monopoly extend to a dynamic framework with endogenous innovation. Also, patent pools composed of complementary patents are welfare improving with endogenous innovation. This finding partially answers a question Lerner and Tirole (2004) posed regarding the desirability of patent pools when their ex-ante effect on innovation activity is taken into account. 1.1. Related Literature. We extend the literature on sequential innovation by analyzing the case in which patents generate cumulative claims on subsequent innovations. This extension is important because it allows us to study the emergence of patent thickets. Many papers of sequential innovations, such as Green and Scotchmer (1995), Chang (1995), and Scotchmer (1996), analyze the optimal distribution of profits between two sequential innovators. If the first innovation has low commercial value (basic research for instance), granting the first innovator a strong patent is optimal. This is not necessarily true in
6 LLANES AND TRENTO our model, where the accumulation of patents generates a problem of anticommons. Abstracting from transaction costs, or the possibility that one or more patent holders refuse to license their ideas thereby blocking innovation, the tragedy of the anticommons is similar to a complementary monopoly problem, first analyzed by Cournot (1838). Cournot modeled a competitive producer of brass who has to use copper and zinc as inputs in production, and showed that, when two different monopolists sell the inputs, the total cost of producing brass is higher than when the same monopolist sells both inputs. Sonnenschein (1968) showed that complementary monopoly is the dual of a duopoly model with quantity competition and homogeneous goods, and Bergstrom (1978) generalized this result to a general number nof inputs and any degree of complementarity among them. Chari and Jones (2000) showed that the market outcome in a complementary monopoly setting is increasingly inefficient as the number of agents increases. Recently Shapiro (2001) and Lerner and Tirole (2004) applied the instruments of complementary monopoly to the analysis of patent pools. Their results reinforce the results on complementary monopoly: patent pools (or, equivalently, a single monopolist owning all the patented inputs) reduce the cost of innovation when patents are complements and increase the cost when patents are substitutes. Boldrin and Levine (2005) and Llanes and Trento (2009) also made use of complementary monopoly to show that, as the number of complementary patents increases, the probability that a future innovation will be profitable goes to zero. All of these papers, although they make important contributions, present static models. In other words, the first innovation has been invented already, so patents and patent pools only affect the profitability of introducing a second innovation. This structure introduces an important asymmetry between previous and future innovations that our dynamic model eliminates. We believe that adding a dynamic dimension is an important step towards a better understanding of the mechanism of anticommons in sequential innovation. In particular, one would expect the complementary monopoly problem to be weaker in a dynamic context for two reasons: first, in the static model there is a limit on the revenues input producers share; we eliminate this limit by extending the analysis to a dynamic framework with potentially infinite innovations. Second, setting high license fees increases the probability that the innovation chain, or a particular research line, will come to a halt: a patent holder would then have an incentive to moderate the license fee to be able to reap part
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 7 of the revenues of further innovations. We find that, in spite of these two effects, the complementary monopoly problem is so strong that innovation eventually becomes unprofitable. Our paper is related to O’Donoghue et al (1998) and Hopenhayn et al (2006), who also present models of cumulative innovation. However, Hopenhayn et al (2006) have no accumulation of claims since only one patent is valid at any given time. O’Donoghue et al (1998), on the other hand, have accumulation of claims but rule out complementary monopoly, as bargaining among patent holders is efficient by assumption. Also, in both papers, innovations are substitutes: the introduction of a new product automatically implies the disappearance of old versions from the market. The substitutability between innovation introduces a natural limit on appropriability and produces an important trade-off, because granting a patent to the first innovator limits what can be offered to the second innovator. In our paper, innovations are complementary and do not compete with each other in the final goods sector. This setting eliminates the appropriability problem. In such a model, one would expect a patent system to perform well. However, the opposite happens: granting too many patent rights on sequential innovations produces a complementary monopoly problem that hampers innovation. Finally, our paper is also related to Menezes and Pitchford (2004), who present a dynamic model of anticommons. Menezes and Pitchford model the case of a buyer who has to combine complementary assets from two sellers. Sellers may have an incentive to avoid entering into negotiations with the buyer because they may get a higher share of total surplus by negotiating after the buyer agrees with the other seller. Holdout occurs if at least one seller is not present in the first round of negotiations. The authors show that complementarity is a necessary condition for holdout, and also that a rise in complementarity leads to an increase in the possibility of holdout. In our case, the accumulation of claims may lead to increasing delays in the agreement between current and past innovators, furthering the welfare loss caused by complementary monopoly. 2. The model We study a model with an infinite sequence of innovations n= 1,2, . . . Each innovation cannot be introduced until all previous innovations have been introduced. This innovation process reflects the fact that earlier innovations do not have a solid background upon which
8 LLANES AND TRENTO to build, while further innovations become more and more indebted to previous ones as the market matures. At each stage, an innovator gets an idea of how to develop a particular innovation. If the innovator decides to perform the innovation, the game continues and, in the following stage, another innovator will get an idea for the next innovation. If the innovator decides not to introduce the innovation, two things may happen: (i) with probability φ, the game continues and in the following stage another innovator tries to perform the failed innovation, and (ii) with probability 1 −φ, the game ends and no other innovations are possible. The parameter φ∈[0,1] represents the degree of scarcity of ideas. If ideas are more scarce (lower φ), each idea is more difficult to substitute, and another innovator is less likely to have a different approach to implement a failed innovation. Let n, j represent the jth innovator trying to introduce innovation n (j−1 innovators have already tried to introduce innovation nwithout success). At the beginning of the stage, the innovator gets an idea with random value vn,j, which she may develop by incurring in a deterministic R&D cost of ε. vn,j represents the revenues obtained by selling the new product in the final-goods market. To concentrate on the effects of patents on innovation activity, we will assume the innovator is a perfect price discriminator in the final-goods market, which means the private value of the innovation is equal to the social surplus the new product generates. The value of the innovation is private information of the innovator. Patent holders only know vn,j is drawn from a uniform distribution between 0 and 1, with cumulative distribution function F(vn,j) = vn,j. The innovator’s decision on whether to perform the innovation will depend not only on vn,j and ε, but also on the licensing revenues and cost that may arise depending on the particular patent regime under analysis. Given that at each stage the innovator will perform the innovation with a certain probability, the game is a multi-stage game with uncertain end, in which the probability that the game continues is determined endogenously. 3. Innovation with patents In this case, patents with infinite length and breadth protect ideas (we will relax these assumptions in section 10), which means each innovator has to pay license fees to all previous inventors (patent holders), in case she wants to introduce the innovation. The cost of innovation
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 9 is the sum of the cost of R&D and the licensing fees paid to previous innovators. A patent of infinite length will also protect the new idea, which means the innovator can request licensing fees from all subsequent innovators. The total revenues of the innovation equal the commercial value of the innovation plus future licensing revenues. The timing of the game within each stage is the following: (i) The n−1 patent holders set licensing fees pi n,j, (ii) Nature extracts a value for vn,j from distribution F(vn,j), (iii) the innovator decides whether to innovate (In,j = 1) or not (In,j = 0). If the revenues from the innovation are higher than the cost, innovator n, j will introduce the innovation, and in the next stage, innovator n+ 1,1 will try to introduce innovation n+ 1. If revenues are lower than cost, innovator n, j will not introduce the innovation, and in the following stage (reached with probability φ), innovator n, j + 1 will try to introduce innovation nbased on a different approach. This innovator n, j + 1 will face the same n−1 patent holders and will have a new draw for the value of innovation, vn,j+1. Let Ji n,j be the expected future licensing revenues of patent holder iat trial jof innovation n, given that stage n, j has been reached. Expressed in a recursive way, Ji n,j =Prn,j (pi n,j +β Ji n+1,1) + (1 −Prn,j)φ β Ji n,j+1, where Prn,j is the probability that innovation nis introduced at trial j, given that n−1 prior innovations have been introduced and that j−1 trials to introduce innovation nhave already failed. With probability Prn,j, the patent holder gets the price pi n,j plus the continuation value of the first trial of the next innovation, Ji n+1,1, discounted by a factor β∈[0,1]. With probability (1 −Prn,j)φ, the innovation is not introduced but the game continues, in which case the patent holder gets the continuation value corresponding to the next trial of the current innovation, Ji n,j+1, discounted by the factor β.βcan be interpreted both as the discount factor or, for a fixed discount factor, as the time between innovations: lower values of βimply ideas arrive less frequently. The innovator’s payoff is In,j(vn,j +βJn n+1,1−cn,j −ε), where cn,j = Pn−1 i=1 pi n,j is the sum of licensing fees paid to previous innovators. We will focus on Markov strategies. A strategy for player ispecifies an action conditioned on the state, where actions are prices and the state is simply n, j. The equilibrium concept is Markov perfect equilibrium, which implies future prices will be determined by a Nash equilibrium in the subsequent games. Thus players understand that
16 LLANES AND TRENTO which does not depend on n. If, on the other hand, innovator n−1 decides to join the pool with the n−2 previous innovators, her expected revenue will be Jn=1−φβ −√1−φβp1−β+β(1 −φ)ε2 (n−1)(1 −φβ)β2(1 −φ)2, which is decreasing in n. This result is due to the fact that the patent pool maximizes joint profits, thus keeping the total cost of innovation constant. This constant amount must be divided among an increasing number of insiders; therefore, the expected revenue of an insider is decreasing in nand converges to 0 as n→ ∞. Therefore some innovator n0>2 with incentives to deviate by remaining outside the pool always exists. Figure 2 illustrates this finding and shows the gains from deviating from the pool as a function of n, for ε= 0.1, φ= 0.5, and β= 0.95. The gains become positive after innovator 3, which means the fourth innovator would gain by remaining outside the pool. Figure 2. Gains from not joining the patent pool. ε= 0.1, φ= 0.5, and β= 0.95 Patent pools can improve innovation activity, but are dynamically unstable. Early innovators have more incentives to enter the pool than subsequent innovators. Brenner (2009) finds an elegant mechanism to solve the instability problem for socially desirable patent pools in a static model. We leave the design of an equivalent mechanism in the context of a dynamic model for future research. Without such a mechanism, however, patent pools are likely to be unstable. This instability might explain why governments sometimes have to enforce the creation of patent pools, as the U.S. government did in the radio and aircraft industry, for example.
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 17 9. Socially optimal innovation The relevant measure of welfare is the expected social value generated by the sequence of innovations. The social value of an innovation is equal to the increase in consumer surplus minus the cost of the resources spent in R&D. Therefore, when considering trial jof innovation n, the social value generated is vn,j −εif the innovation is performed, and 0 otherwise. Consider the decision of performing innovation n, j. If the value of the innovation is greater than the cost, obviously the innovation should be performed. However, the social planner could still decide to perform an innovation with negative social value, because in the opposite case, the sequence of innovations will stop with a probability of 1 −φ. The decision will depend, therefore, on a comparison between the current cost of performing an innovation with negative social value and the expected future benefits of continuing with the chain of innovations. Let Wm,k be the expected social welfare from stage m, k onwards. Once we know the realization of vn,j, expected welfare is vn,j −ε+ β Wn+1,j if the innovation is performed, and β φ Wn,j+1 if the innovation is not performed. Therefore, the innovation should be performed if vn,j −ε+β Wn+1,j ≥β φ Wn,j+1. Proposition 1 shows the socially optimal innovation policy. Proposition 1 (Socially optimal innovation).To maximize expected social welfare, innovation n, j should be performed if and only if vn,j ≥ v∗, where v∗=(0if ε≤β 2 1−φ 1−β φ , β−1+√1−β φ √1−2β(1−(1−φ)ε−φ/2) β(1−φ)if ε > β 2 1−φ 1−β φ . Proof. Given the assumptions of the model, innovations vn,j and vm,k are equivalent for any n, j, m, k. It follows that Wn,j =Wm,k =W, and the optimal decision is time-invariant: a value v∗∈[0, ε] exists such that innovation n, j should be performed if and only if vn,j ≥v∗. By definition, v∗solves v∗−ε+β W =β φ W. Therefore, we need to determine the value of W. In particular, Wm,k is given by Wm,k =Pr(vm,k ≥v∗) (E(vm,k −ε/vn,j ≥v∗) + β Wm+1,k) + (1 −Pr(vm,k ≥v∗)) β φ Wm,k+1. Imposing Wm,k =Wm+1,k =Wm,k+1 =W, and solving for W, we get W=1−v∗ 1−β(1 −(1 −φ)v∗)1 + v∗ 2−ε.
18 LLANES AND TRENTO Substituting this result into v∗−ε+β W =β φ W, and solving for v∗, we get the optimal policy stated in the proposition. Proposition 1 implies that innovation will be suboptimal in the three cases studied above. There are three reasons why this is so: dynamic externalities, market power, and asymmetric information. The dynamic externality is best described by analyzing the no-patents case. Without patents, the innovator will perform the innovation when vn≥ε/θ. Given that v∗≤ε, the innovator may decide not to perform a socially desirable innovation, even if θ= 1, because she ignores the effect of her decision on the technological possibilities of future innovators. This effect is well known in the literature of sequential innovation (Scotchmer, 1991; Hopenhayn et al, 2006) and is similar to the one found in the literature of moral hazard in teams (e.g., Holmstrom, 1982), where each agent internalizes only his reward from the effort exerted. The solution to the first problem would require intertemporal transfers. In section 9.1, we show that the first best can be reached by decentralizing the innovation decision and implementing a tax-subsidy scheme. Surprisingly, this result holds even if information is asymmetric, that is, if neither the government nor previous innovators know the value of the innovation. In the patents and patent-pool cases, the inefficiency arises from a different source: market power and asymmetric information. Because patent holders care about the stream of future licensing revenues they will lose if the sequence of innovations stops, they internalize the dynamic externality. However, asymmetric information implies a downward-sloping expected demand for old innovations, and market power implies inefficient pricing of patents, which leads to suboptimal innovation. As the number of holders of rights on innovation increases, the inefficiency due to market power increases (because of the complementary monopoly), which is why the patent-pools case is more efficient than the patents case. In order to show the importance of the asymmetric information assumption, in Section 9.2, we show that under symmetric information there exists an equilibrium that reaches the first best.2This means that without asymmetric information, the dynamic externality could be perfectly internalized, reaching the first best. 2A continuum of equilibria exists under symmetric information. Some of these equilibria do not reach the first best, but the important fact is that some of them do.
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 19 9.1. Optimal transfers. In this section, we show that the first best can be reached by decentralizing the innovation decision and implementing a tax-subsidy scheme. This is a surprising finding, because it does not require the government to know the value of the innovation in order to be implemented. In addition, we find that the optimal timing of the tax-subsidy scheme is as follows: innovators should receive a subsidy if they innovate, and then be taxed when the following innovation is performed. The structure of transfers is the following: if innovator n, j decides to innovate, she will have to pay a transfer tnto the innovator who successfully performed innovation n−1, but she will also have the right to receive a transfer tn+1 from the innovator who performs innovation n+ 1. Therefore, given vn,j, the innovator will innovate if vn,j −ε− tn+β Jn+1,1≥0, where (1) Jn+1,1=Prn+1,1tn+1 + (1 −Prn+1,1)β φ Jn+1,2. Proposition 2 shows the optimal intertemporal transfer that implements the first best. Proposition 2. The optimal transfer is constant and equal to t∗=(v∗−ε) (1 −φ β v∗) 1−β(1 −v∗+φ v∗). t∗≤0for any value of the parameters, and t∗<0if and only if ε > 0, φ β < 1. Proof. The problems of innovators n, j and m, k are equivalent for the social planner for any n, j, m, k, which means tn=tn+1 =t, and Jn,j =Jm,k =Jfor any n, j, m, k. Given transfers, the probability of innovation is Pr = 1−ε−t+β J. We want to make this probability equal to the optimal probability, which is Pr∗= 1 −v∗. The optimal transfer then solves v∗=ε+t−β J. On the other hand, from equation (1), we get J= (1−v∗)t/(1−φ β v∗). Substituting the latter expression into the former, and solving for t, we get the optimal transfer stated in the proposition. Finally, t∗≤0 because v∗≤εfrom Proposition 1. Note that v∗has a kink when ε=β 2 1−φ 1−β φ , so t∗will also have a kink at that point. An interesting feature of the optimal transfer is that it is negative. Therefore, the innovator should receive a subsidy to innovate, and be taxed with the following innovation. Most importantly, optimal transfers work in the exact opposite way as patents, which require an innovator to pay previous inventors upfront and then be compensated by following innovators.
20 LLANES AND TRENTO 9.2. Symmetric information. To analyze the reasons for inefficiency in the different cases, in this section we study the effects of removing the asymmetric information assumption while keeping the market power assumption (i.e., past innovators are price-setters, whereas the current innovator is a price taker). At each stage, a value of vn,j is drawn from F(vn,j), and previous innovators set the levels of licensing fees the innovator will pay, just as in the basic model. The difference is that now, previous innovators learn the realization of vn,j, and use this information when setting their licensing fees. In equilibrium, previous innovators will set a level of fees that will leave the innovator indifferent between innovating or not. Otherwise, one of the previous innovators could raise her fee without affecting the innovation decision, thereby raising her profits. Therefore, any sequence of prices such that n−1 X i=1 pi n,j =vn,j −ε+β Jn n+1,1 is an equilibrium. Consider an equilibrium in which each innovator pays a fee only to the previous innovator: pn−1 n,j =vn,j −ε+β Jn n+1,1. For the remainder of this section, let pn,j =pn−1 n,j and Jn,j =Jn−1 n,j . If vn,j −ε+β Jn+1,1<0, innovator n−1 will not allow innovator n, j to innovate (she can do this by setting any price above vn,j −ε+β Jn+1,1). In case vn,j −ε+β Jn+1,1≥0, on the other hand, innovator n−1 will allow innovation n, j only if pn,j ≥β φ Jn,j+1. Proposition 3 shows that in equilibrium, innovations will be performed only if they are socially desirable. Proposition 3 (Symmetric information).Under symmetric information, the equilibrium in which each innovator only pays a licensing fee to the previous innovator is socially optimal. Proof. The problem at trial jis the same as the problem at trial j+ 1, which means the lowest value of vn,j that innovator n−1 will tolerate is constant. Let ˆvindicate this value. ˆvsolves ˆv−ε+β Jn+1,1=β φ Jn,j+1, and Jn,j+1 solves Jn,j+1 =Pr E(pn,j+1/vn,j+1 ≥ˆv) + (1 −Pr)β φ Jn,j+2, where Pr =Pr(vn,j+1 ≥ˆv) = 1 −ˆv, and E(pn,j+1/vn,j+1 ≥ˆv) = 1+ˆv 2−ε+Jn+1,1. The problem for a different nand/or jis equivalent,
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 21 so Jn,j =Jm,k =Jfor all n, j, m, k. Using this result in the above equations, we obtain ˆv=ε−β(1 −φ β)J J= (1 −ˆv)1 + ˆv 2+β φ J+ ˆv φ β J. Solving this system of equations for ˆvand J, we get that ˆv=v∗. 9.3. Static versus dynamic incentives. Previous models of complementary monopoly, sequential innovation, and patent pools were static (Shapiro, 2001; Lerner and Tirole, 2004; Boldrin and Levine, 2005; Llanes and Trento, 2009). Looking at what changes when we add the dynamic dimension is interesting. To see what happens in the static case, assume only one innovation is under consideration. The innovation uses n−1 old ideas, which have already been invented. If the innovation is performed, the innovator obtains a value vfrom a uniform distribution between 0 and 1, and incurs in a cost εof R&D. The probability of innovation is Pr = 1−ε−cn, with patents or patent pool and Pr = 1 −ε/φ without patents. With patents, the patent holder’s problem is to maximize Pr pi. As a result, the equilibrium price is 1−ε nand the probability of innovation is 1−ε n. We have shown that in the dynamic model, the probability of innovation is 1−ε n+βJn+1 −n−1 nφβJn, with Jn, Jn+1 >0. These extra terms arise because the innovator gets licensing revenues from future innovators. Dynamic incentives imply a higher probability of innovation, but the increase is not enough to prevent the probability of innovation from converging to 0 as n→ ∞. A patent pool would consider cross-price effects, which would lead to a price of 1−ε 2 (n−1) and a probability of innovation of 1−ε 2. The probability of the corresponding dynamic model is 1−ε 2+n 2βJn+1 −n−1 2φβJn, with Jn, Jn+1 >0. In this case, the extra terms arise due to not only the future licensing revenues of the innovator, but also to the pool’s concern with keeping the future licensing revenues of current patent holders. With respect to the no-patents case, the profit-maximizing decision is the same as in the dynamic case. Innovators will therefore perform the innovation if φ vn≥ε, which leads to a probability of Pr = 1 −ε/φ. However, in the dynamic case, innovation is suboptimal even when φ= 1, which contrasts with the static case, where innovation is socially optimal because no intertemporal link between innovations exists and therefore neither does any externality.
22 LLANES AND TRENTO 10. Finite patents We have seen that with patents of infinite length, innovation is stifled as nincreases due to the complementary monopoly problem. In this section, we ask whether appropriately reducing the length of patents can prevent this problem.3 For simplicity, we assume φ= 0 and β= 1, which is the most favorable case for patents: ideas are scarce and the discount factor is small, so, in principle, society could largely benefit if innovators with a low value for their inventions could get additional revenues from charging other innovators. Each stage corresponds to one period and only one innovation is attempted at each period. If the innovator decides to introduce the innovation, she obtains a patent for Lperiods. The innovator, therefore, has to pay licensing fees for Lprevious innovations, but she also charges licenses to Lfuture innovators. The main difficulty of the present analysis is that, unlike in the previous sections, the identity of the patent holders matters. The price and future expected licensing revenues will be different for different patent holders, depending on how long her patent lasts. The innovator will introduce the innovation if the revenues from innovation are larger than the cost: vn+ n+L X m=n+1 pn m m Y k=n+1 Prk≥ n−1 X i=n−L pi n+ε, which means the probability of innovation is Prn= 1 + n+L X m=n+1 pn m m Y k=n+1 Prk− n−1 X i=n−L pi n−ε. The Lcurrent patent holders differ in their objective functions. Let Ji nbe the future expected revenues of patent holder iat stage n, given that stage nhas been reached. Then Ji n=Prn(pi n+Ji n+1). The patent holder charging a license for the last time is patent holder n−L, so Jn−L n+1 = 0. The patent of n−L+ 1, on the other hand, will 3We have also analyzed the effects of reducing the breadth of patents. For example, suppose new inventions may infringe on old patents with certain probability. Within our framework, the effects of reducing breadth are similar to the effects of reducing patent length: a lower breadth implies that the innovator will have to pay fewer licensing fees, but it also means fewer future inventions will infringe on her patent.
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 23 last for one more period, so Jn−L+1 n+1 =Prn+1 pn−L+1 n+1 . In this way, we can construct the future expected revenues of the Lpatent holders. The profit maximization problem is max pi n Ji n=Prn(pi n+Ji n+1). The first-order condition is −pi n−Ji n+1 +Prn= 0, so pi n+Ji n+1 =Prn and Ji n=Pr2 nfor all i, which also implies pn−L n=Prn. We are interested in stationary equilibria, which means Prn=Pr for all n. Stationarity, together with the first order condition, implies pi n=Pr(1 −Pr) for i≥n−L. Substituting the equilibrium prices into the probability of innovation, we get P r = 1 + n+L X m=n+1 pn m m Y k=n+1 P rk− n−1 X i=n−L pi n−ε = 1 + L−1 X m=1 P r(1 −P r)P rm+P rP rL−(L−1)P r(1 −P r)−P r −ε. Solving for Pr, we get: Pr =L+ 1 −p(L−1)2+ 4Lε 2L, which is the stationary equilibrium probability of innovation. Figure 3 shows the probability of innovation as a function of the patent length for ε= 0.2. We can see the probability of innovation decreases with L, which means patents hurt more than benefit the innovator, because the innovator has to pay licenses to the patent holders. Future licensing revenues are uncertain, however, as they depend on future innovations being performed. Also note that Pr →0 when L→ ∞ and Pr →1−εwhen L→0, which corresponds to the previously analyzed patents and no-patents cases (with θ= 1). 10.1. Revenues depend on patent length. We have assumed that the revenues from selling the new product in the final goods market are independent of patent length. In this subsection, we analyze what happens when we relax this assumption. Assume the revenues of the innovator are ψ(L)vn, with ψ0(L)≥0, ψ00(L)≤0,limL→0ψ(L) = ψand limL→∞ ψ(L) = 1. Here, ψis the fraction of social surplus the innovator would appropriate without any patent protection due to trade secrets or first-mover advantages.
24 LLANES AND TRENTO Figure 3. Probability of innovation and patent length. In this case, the innovator will innovate if ψ(L)vn+ n+L X m=n+1 pn m m Y k=n+1 Prk≥ε+ n−1 X i=n−L pi n. Applying a procedure similar to that in the previous case, we obtain the probability of innovation in the stationary equilibrium: Pr =L+ 1 −p(L−1)2+ 4Lε/ψ(L) 2L. The effect of patent length on the probability of innovation depends on the functional form of ψ(L). Let ψ(L) = 1 −1−ψ (L+1)γ, where γmeasures the speed at which revenues grow when Lincreases. Figure 4a shows that when ψis more concave (γ= 1), the probability of innovation first increases and then decreases with patent length. The optimal length is positive and finite (in this case L= 1). Figure 4b shows that, for a lower degree of concavity of ψ(L), completely removing patents is optimal. Therefore, the results do not change significantly when the revenues in the final-goods sector depend on patent length. In this model, short patents therefore perform better than long patents. O’Donoghue et al (1998) find that when patent breadth is infinite, which is always the case in this model, long patents stimulate innovation activity. These apparently different results arise because we are looking for solutions to different problems. We analyze the effect of patent policy on the complementary monopoly problem, disregarding its effect on the final-goods market. O’Donoghue et al (1998) do exactly the opposite. Therefore, our findings are not opposed to theirs, but rather are complementary.
ACCUMULATION OF CLAIMS IN SEQUENTIAL INNOVATION 25 (a) ψ= 0.2, ε = 0.1, γ = 1. (b) ψ= 0.2, ε = 0.1, γ = 0.1. Figure 4. Probability of innovation as a function of patent length. 11. Conclusion In this paper, we build a dynamic model where the accumulation of patents generates an increasing number of claims on cumulative innovation. The model is intended to reproduce the central feature of innovation activity in high-tech industries: new products are more complex than old products, because they build on a larger stock of previously accumulated knowledge. We study the policy that maximizes expected social welfare and compare it with the outcome of three patent-policy regimes: patents, patent pools, and no patents. We find that, even abstracting from the monopolistic inefficiencies of patents, none of these policies attains the optimum. With patents, the innovator has to pay an increasing number of license fees to previous innovator. Asymmetric information on the value of the innovation and uncoordinated market power of licensors create an anticommons effect that reduces the incentives to innovate as innovation becomes more complex. The anticommons effect is weaker than in the static case, but it is still strong enough to drive the probability of innovation to zero as the number of licenses grows. Enforcing a patent pool solves the lack of coordination but not the asymmetricinformation problem. As a result, the outcome of patent pools is more desirable but still does not achieve the first best. Eliminating patent protection solves the two problems but introduce a non-internalized externality: previous innovations set the foundations for future innovations. Therefore the social cost of one innovation may be higher than its instantaneous social value (the social value the innovation creates per se), and yet the innovation may be socially desirable because it