A new method for measuring tail exponents of firm size distributions
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Fujimoto, Shouji; Ishikawa, Atushi; Mizuno, Takayuki; Watanabe, Tsutomu Working Paper A new method for measuring tail exponents of firm size distributions Economics Discussion Papers, No. 2011-29 Provided in Cooperation with: Kiel Institute for the World Economy – Leibniz Center for Research on Global Economic Challenges Suggested Citation: Fujimoto, Shouji; Ishikawa, Atushi; Mizuno, Takayuki; Watanabe, Tsutomu (2011) : A new method for measuring tail exponents of firm size distributions, Economics Discussion Papers, No. 2011-29, Kiel Institute for the World Economy (IfW), Kiel This Version is available at: https://hdl.handle.net/10419/48826 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by-nc/2.0/de/deed.en
A New Method for Measuring Tail Exponents of Firm Size Distributions Shouji Fujimoto Kanazawa Gakuin University Atushi Ishikawa Kanazawa Gakuin University Takayuki Mizuno University of Tsukuba Tsutomu Watanabe Hitotsubashi University and University of Tokyo Abstract We propose a new method for estimating the power-law exponent of a firm size variable, such as annual sales. Our focus is on how to empirically identify a range in which a firm size variable follows a power-law distribution. As is well known, a firm size variable follows a power-law distribution only beyond some threshold. On the other hand, in almost all empirical exercises, the right end part of a distribution deviates from a power-law due to finite size effect. We modify the method proposed by Malevergne et al. (2011) so that we can identify both of the lower and the upper thresholds and then estimate the power-law exponent using observations only in the range defined by the two thresholds. We apply this new method to various firm size variables, including annual sales, the number of workers, and tangible fixed assets for firms in more than thirty countries. Paper submitted to the special issue New Approaches in Quantitative Modeling of Financial Markets JEL C16, C18, D20, E23 Keywords Econophysics; power-law distributions; power-law exponents; firm size variables; finite size effect Correspondence Shouji Fujimoto, Kanazawa Gakuin University, The Department of Informatics and Business, 2-6-36 Naga-machi, Ishikawa, Kanazawa, Japan; e-mail: [email protected]; [email protected] Discussion Paper No. 2011-29 | July 29, 2011 | http://www.economics-ejournal.org/economics/discussionpapers/2011-29 © Author(s) 2011. Licensed under a Creative Commons License - Attribution-NonCommercial 2.0 Germany
conomics Discussion Paper Introduction Power-law distributions are frequently observed in social phenomena (e.g., Pareto (1897); Newman (2005); Clauset et al. (2009)). One of the most famous examples in Economics is the fact that personal income follows a power-law, which was first found by Pareto (1897) about a century ago, and thus referred to as Pareto distribution. Specifically, the probability that personal income xis above x0is given by P>(x)∝x− µ for x>x0(1) where µ is referred to as a Pareto exponent or a power-law exponent. As for the variables related to firm behavior, it is well known that there are several variables that follow a power-law, including firm sales for a particular period (e.g., annual sales), the number of workers employed by a firm, and the amount of fixed assets, like machinery equipments, held by a firm. The fact that the firm size variables mentioned above follow power-law distributions implies that the behavior of these variables at the aggregate level is dominantly affected by a very small number of firms that are extremely large in their size. The purpose of this paper is to propose a new method for estimating the powerlaw exponent of a distribution. Our special focus is on how to empirically determine a range in which a variable follows a power-law distribution. On the one hand, as shown in equation (1), a variable follows a power-law distribution only when it exceeds some threshold, for example x0in (1); the variable deviates from a power-law below that threshold. Thus we need to empirically specify where such a threshold exists. On the other hand, in almost all empirical exercises, the right end part of a distribution deviates from a power-law due to the limited number of observations. It is often the case that the right end part of a distribution exhibits a much quicker decay than implied by a power-law due to such a finite size effect. We need to eliminate that part of a distribution before estimating a power-law exponent. Our strategy is to empirically specify the range of a variable, which is defined by a lower threshold x0and an upper threshold x1, and then estimate a power-law exponent using only observations only in that range. Our method is based on the one proposed by Malevergne et al. (2011).1 Malevergne et al. (2011) propose to test the null hypothesis that, beyond some threshold, the upper tail of a distribution is characterized by a power law distribution against the alternative that the upper tail follows a lognormal beyond the same threshold. It is important to note that their intention was to detect a lower threshold x0by conducting this test, and that they did not pay any particular attention to the presence of an upper threshold x1. However, as we will show later, in 1See, for example, Hisano and Mizuno (2011) for an application of their method. www.economics-ejournal.org 2
conomics Discussion Paper applying this method to firm size variables, one often encounters a situation that the threshold detected by this method is not x0but x1. Needless to say, this failure leads to an imprecise estimate of a power-law exponent. In our method, we first apply the test by Malevergne et al. (2011) to detect a upper threshold, x1. We then repeat the test, but we “thin out” observations before conducting the second round test. Specifically, we discard observations above x1, which is detected by the first round test, and similarly we thin out observations below x1. Then we apply the test to the thinned out set of observations to detect x0.The rest of this paper is organized as follows. In Section 1, we will provide detailed explanation on our new method. In Section 2, we will apply the new method to firm size variables, including annual sales, the number of workers, and tangible fixed assets for firms in more than thirty countries. Section 3 concludes the paper. 1 Methodology Let us start by showing the empirical distributions for tangible fixed assets, which is denoted by K, the number of workers, L, and annual sales, Y. The cumulative distributions for these three variables for Japanese firms are shown in figure 1 with horizontal and vertical axes being in logarithm. We see that dots are on a straight line in each of the three figures, indicating that each of the distributions is a powerlaw. However, dots deviate from a straight line when the firm size variables take very small or very large values. In other words, K,L, and Yfollow power-law distributions only with some range. That is, P>(K)∝K− µ Kfor K0<K<K1,(2) P>(L)∝L− µ Lfor L0<L<L1,(3) P>(Y)∝Y− µ Yfor Y0<Y<Y1.(4) The main issue of this paper is how to estimate the range in which dots are on a straight line; namely, [K0,K1],[L0,L1], and [Y0,Y1]. Our method is based on Malevergne et al. (2011), which propose a method to identify the boundary between a power-law and a lognormal. Consider a case described by equation (1). For each value of x, they test the null hypothesis that xfollows a power-law distribution beyond that value against the alternative that x follows a lognormal distribution beyond the same value. They start this test for the maximum value of x, and repeat the test for the second largest, the third largest values, and so on, until the null is finally rejected. Note that their test is equivalent to testing the null that the upper tail of the log of xfollows an exponential distribution against the alternative that the www.economics-ejournal.org 3
conomics Discussion Paper 10−6 10−4 10−2 100 100102104106108 P>(K) P/AK (inthousandUSdollars ) 2004 2005 2006 2007 2008 2009 (a) CDFs of tangible fixed assets in 2004–2009 10−6 10−4 10−2 100 100102104106 P>(L) TheNumberofEmployeeL 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 (b) CDFs of the number of workers in 2000–2009 10−6 10−4 10−2 100 100102104106108 P>(Y) SalesY (inthousandUSdollars ) 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 (c) CDFs of annual sales in 2000–2009 Figure 1: Cumulative Distribution Functions of Firm Size Variables for Japanese Firms. log of xfollows a truncated normal distribution. For this transformed test, del Castillo and Puig (1999) have shown that the clipped empirical coefficient of variation provides the uniformly most powerful unbiased test. Specifically, let us consider a random variable zwhich follows a truncated normal distribution with truncation occurring at z=A. The probability density function P(z)is given by P(z; α , β ) = exp[− α (z−A)− β (z−A)2]/NC( α , β )(5) where NC( α , β )represents a scaling value, and it is defined by NC( α , β ) = √ π β exp( α 2 4 β )[1−Φ( α √2 β )] (6) where Φ(·)is the CDF of a standard normal distribution. Note that it can be shown by using asymptotic expansion that P(z; α , β )→ α exp[− α (z−A)] as ( α , β )→ ( α ,0). Suppose there are nobservations for z(namely, z1,z2,···,zn). The log likelihood is given by l( θ ) = l( α , β ) = − α n ∑ i=1 (zi−A)− β n ∑ i=1 (zi−A)2−nlogNC( α , β )(7) and the maximum likelihood estimate for θ = ( α , β )is characterized by − γ h( γ )+ γ 2+1/2 (h( γ )− γ )2−1=c2(8) where γ and h( γ )are defined by γ ≡ α /2√ β ;h( γ )≡exp(− γ 2) 2√ π (1−Φ(√2 γ )) (9) www.economics-ejournal.org 4
conomics Discussion Paper and c2is the square of the coefficient of variation for (xi−A), which is defined by c2=⟨(x−A)2⟩−⟨x−A⟩2 ⟨x−A⟩2=⟨(x−A)2⟩ ⟨x−A⟩2−1 (10) where ⟨·⟩ represents the sample mean. For a give value of A, one can calculate c2from the data, and then obtain a maximum likelihood estimate of γ from (8), which is denoted by ˆ γ . Note that the expression on the left hand side of (8) is monotonically increasing with respect to γ , so that one can obtain a solution just by applying a simple method like the Newton-Raphson method. If zfollows an exponential distribution rather than a truncated normal distribution, β in equation (5) is equal to zero, and the log likelihood is given by l( θ ) = l( α ,0) = − α n ∑ i=1 (xi−A)−nlog( α )(11) The maximum likelihood estimate for θ is given by ˜ θ = ( ˜ α ,0) = (1/⟨x−A⟩,0). Then the null hypothesis that zis exponentially distributed can be tested against the alternative that zfollows a truncated normal distribution by conducting a likelihood ratio test, in which the likelihood ratio is given by W=2(l(ˆ θ )−l(˜ θ )) (12) The random variable zis more likely to follow a truncated normal distribution if the value of Wis above zero, and it is more likely to be exponentially distributed if Wis below zero. Specifically, it is known that the asymptotic distribution of W around W=0 is a 50-50 mixture of a χ 2distribution with a degree of freedom of one, and a constant zero (see Self and Liang (1987) and Geyer (1994)). Therefore, the asymptotic distribution of Wis given by W(ˆ γ ) = 0 if cis greater than unity, and W(ˆ γ ) = n[2log{2h(ˆ γ )(h(ˆ γ )−ˆ γ )}+2ˆ γ 2−2ˆ γ h(ˆ γ )+1](13) if cis less than unity. del Castillo and Puig (1999) adopts a more precise approximation to Wby using W∗=W(ˆ γ )+2L(ˆ γ )+L2(ˆ γ )/W(ˆ γ )(14) where L(·)is defined by L(ˆ γ ) = 1 2log[2ˆ γ 3h(ˆ γ )−4ˆ γ 2h(ˆ γ )2+ˆ γ h(ˆ γ )(2h(ˆ γ )2+3)−3h(ˆ γ )2+1 4(h(ˆ γ )−ˆ γ )2(W(ˆ γ )/n)](15) In sum, the procedure proposed by del Castillo and Puig (1999) and Malevergne et al. (2011) is as follows. www.economics-ejournal.org 5
conomics Discussion Paper 1. Pick up the largest nobservations and take log. Set the threshold Aequal to the log of the value for the largest observation. 2. Compute ˆ γ by solving (8). 3. Compute W∗and p-value associated with it by inserting the value of ˆ γ into (14). 4. Repeat this procedure for n=1,2,3,... until the p-value associated with W∗is sufficiently large to reject the null hypothesis. Let us show how the method proposed by Malevergne et al. (2011) works by applying it to the distribution for the number of workers employed by Japanese firms in 2004. The black dots in Figure 2 represent empirical CDF produced using actual observations. There are two vertical dotted lines in the figure, but the right one represents the threshold identified by the procedure proposed by Malevergne et al. (2011), which corresponds to the 17th largest observation with the value (i.e., the number of workers) of 84,899. If their method works well, this result indicates that the number of workers follows a power-law beyond this threshold, but a lognormal below it. However, as one can clearly see from the figure, the back dots are on a straight line even below this threshold, implying that their method fails to detect a correct threshold. This failure happens because the right end part of the distribution decays quicker than the other part of the distribution due to the limited number of observations. The possibility of such a finite size effect is not seriously considered in Malevergne et al. (2011). It is important to note that this particular case is not an exception, but in fact we encounter similar failures quite often in estimating the power-law exponents of firm size distributions. To cope with this problem, we propose to modify their procedure in the following way. Basically what we will do is to “thin out” observations so as to minimize the extent to which one suffers from the finite size effect. Specifically, after detecting the 17th largest observation as a (wrong) threshold, we discard 16 observations above it. We also discard the 18th, 19th, 20th,..., and 33rd largest observations, the 35th, 36th, 37th,..., and 50th largest observations, and so on. By repeating this procedure, we end up with a thinned out set of observations which consist of the 17th largest observation, the 34th largest observation, the 51st largest observation, and so on. These thinned out observations are indicated by grey circles in Figure 2. Then we apply again the method by Malevergne et al. (2011), but this time not to the original set of observations but to the thinned out set of observations. This second round test identifies a new threshold, which is represented by another vertical dotted line in Figure 2. This corresponds to the 24701st largest among the original set of observations and the 1453rd largest among the thinned out set of www.economics-ejournal.org 6
conomics Discussion Paper Figure 2: Cumulative Distribution Function for the Number of Workers Employed by Japanese Firms in 2004. Black dots represent the original set of observations, while grey dots represent the thinned-out set of observations. The two vertical dotted lines indicate the upper and lower thresholds, which are estimated using the method described in the text. The power-law exponent is estimated using only observations within the range defined by these two thresholds. observations. The number of workers corresponding to this second threshold is 60, which is substantially lower than the number corresponding to the first threshold. We see from the figure that dots, both black and grey, are on a straight line in the range indicated by the two dotted vertical lines. To see how our method works, consider a size-rank equation of the form lnr=const− µ lns(16) where srepresents a firm size, ris the rank associated with it, and µ is a power-law exponent. We assume that this size-rank equation holds for r∈[r0,r1]. We know the value of r0from the first round test (r0is 17th in the above example). Let s0 www.economics-ejournal.org 7
conomics Discussion Paper represents the size associated with the rank r0. Equation (16) implies that ln(r r0)=− µ ln(s s0)(17) holds for r=r0,2r0,3r0,4r0,... as far as ris smaller than r1. Thus we can estimate a power-law exponent µ using a thinned out set of observations {r0,2r0,3r0,4r0,...}. Note that discarding only observations with higher ranks than r0does not work, because, in this case, the rank in the new set of observations is r−r0, rather than r/r0in equation (17), and the log of r−r0does not depend linearly on the log of s. The procedure we propose is summarized as follows. 1. Apply the method proposed by Malevergne et al. (2011) to the original observations to detect an observation (we refer to this as k-th largest observation), above which the CDF is steeper than the other part due to finite size effect. 2. Create a new (thinned out) set of observations, consisting of the k-th largest observation, the 2k-th largest observation, the 3k-th largest observations, and so on. 3. Apply the method proposed by Malevergne et al. (2011) to the thinned out set of observations to detect a new threshold (we refer to this as K-th largest observations). 4. Estimate the slope of a straight line within the range defined by the value associated with the k-th largest observation and the value associated with the K-th largest observation. 2 Empirical Results In this section we apply the new method to firm size variables, including annual sales, the number of workers, and tangible fixed assets for firms in more than thirty countries.2The data comes from ORBIS provided by Bureau van Dijk, which contains B/S and P/L information for more than 60 million firms all over the world. The sample period is 1999 to 2009.3 2There is a long list of papers that investigate various aspects of firm size distributions, including Stanley et al. (1995), Okuyama et al. (1999), Ramsden and Kiss-Haypál (2000), Mizuno et al. (2006), Axtell (2001), Gaffeo et al. (2003), Fujiwara et al. (2004), and Zhang et al. (2009). 3More detailed information on the dataset employed in this paper is available at http://www.bvdinfo.com/Home.aspx. www.economics-ejournal.org 8
conomics Discussion Paper (a) Black line represents f(x)in (21) computed using a built-in error function; Gray line represents the same function f(x), but it is computed using an equation obtained from asymptotic expansion equation (up to xto the 25th power). (b) Black line represents L(x)in (15) computed using a built-in error function; Gray line represents the same function L(x), but it is computed using an equation obtained from asymptotic expansion equation (up to xto the 25th power). Figure 6: Comparison between the result from built-in error function and the result from asymptotic expansion see that the built-in error function is able to return a precise outcome up to x=6, but unable to do so for the values greater than that due to underflow. To fix this problem, we use a built-in error function for up to x=4, but use an equation obtained from asymptotic expansion for x>4. Turning to a function L(·)in equation (15), we compare in Figure 6b the result obtained from a built-in error function and the result obtained using asymptotic expansion up to xto the 25th power. Again we see that the built-in error function fails to return a precise outcome for xgreater than 4. More importantly, there is a discontinuous jump around at x=4, which cannot be completely eliminated even if we increase the order of expansion. To fix this, we use the built-in function for x<4 and use an equation obtained from asymptotic expansion for x>5, and adopt a linear extrapolation between the two. Also we set an approximate value of L(x)for x>15 at zero since it can be shown analytically that L(x)→ −0 as x→∞. Finally, a similar problem occurs for L(x)2/W(x)in equation (14). We use the built-in function for x<4 and use an equation obtained from asymptotic expansion for x>6, and adopt a linear extrapolation between the two. We also set www.economics-ejournal.org 15
conomics Discussion Paper an approximate value of L(x)2/W(x)at 64/9 for x>10 since it can be shown analytically that L(x)2/W(x)→64/9 as x→∞. www.economics-ejournal.org 16
Please note: You are most sincerely encouraged to participate in the open assessment of this discussion paper. You can do so by either recommending the paper or by posting your comments. Please go to: http://www.economics-ejournal.org/economics/discussionpapers/2011-29 The Editor © Author(s) 2011. Licensed under a Creative Commons License - Attribution-NonCommercial 2.0 Germany