Distinguishing log-concavity from heavy tails
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Asmussen, Søren; Lehtomaa, Jaakko Article Distinguishing log-concavity from heavy tails Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Asmussen, Søren; Lehtomaa, Jaakko (2017) : Distinguishing log-concavity from heavy tails, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 5, Iss. 1, pp. 1-14, https://doi.org/10.3390/risks5010010 This Version is available at: https://hdl.handle.net/10419/167911 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Article Distinguishing Log-Concavity from Heavy Tails Søren Asmussen * and Jaakko Lehtomaa Department of Mathematics, Aarhus University, Ny Munkegade 118, DK-8000 Aarhus C, Denmark; [email protected] *Correspondence: [email protected]; Tel.: +45-2066-3866 Academic Editor: Qihe Tang Received: 14 November 2016; Accepted: 17 January 2017; Published: 7 February 2017 Abstract: Well-behaved densities are typically log-convex with heavy tails and log-concave with light ones. We discuss a benchmark for distinguishing between the two cases, based on the observation that large values of a sum X1+X2occur as result of a single big jump with heavy tails whereas X1,X2are of equal order of magnitude in the light-tailed case. The method is based on the ratio |X1−X2|/(X1+X2), for which sharp asymptotic results are presented as well as a visual tool for distinguishing between the two cases. The study supplements modern non-parametric density estimation methods where log-concavity plays a main role, as well as heavy-tailed diagnostics such as the mean excess plot. Keywords: heavy-tailed; log-concave; mean excess function; principle of a single big jump MSC: 60E05; 60G50; 62G20 1. Introduction General interest towards non-parametric thinking has increased over the last few years. One example is density estimation under shape constraints instead of requiring the membership of a parametric family. Here, a particularly robust alternative to parametric tests is provided by searching for the best fitting log-concave density. Another example is the mean excess plot that aims at distinguishing light and heavy tails. Throughout the paper, we consider i.i.d. random variables X,X1,X2, . . . >0 with common distribution Fhaving density fand tail F(x) = P(X>x). Then, Xis (right) heavy-tailed if EesX =∞ for all s>0 and light-tailed otherwise. The density fis log-concave, if f(x) = eφ(x), where φis a concave function. If φis convex, then fis log-convex. This paper aims to illustrate that light-tailed asymptotic behaviour is associated with log-concave densities. Likewise, log-convexity seems to be connected to heavy-tailed behaviour. One can use the connection to assess potential heavy-tailedness by searching for patterns that are typically present among distributions with log-concave or log-convex densities. Log-concavity is a widely studied topic in its own right [1,2]. There also exists substantial literature regarding its connections to probability theory and statistics [3,4]. Several papers concentrate on the statistical estimation of density functions assuming log-concavity [5,6]. This is due to the fact that log-concavity provides desirable statistical properties for estimators. For instance, maximum likelihood estimation becomes applicable and the estimate is unique. The topic is discussed in detail in the beginning of [7]. Unfortunately, much less emphasis seems to be put on verification of the log-concavity property itself. Specifically, it seems to be relatively little studied if it is feasible that the sample be generated by a log-concave distribution. See, for example [8,9]. A distribution with a log-concave density fis necessarily light-tailed. In contrast, fis log-convex in the tail in the standard examples of heavy tails such as regular variation, the lognormal distribution and Weibull case F(x) = e−xαwith α<1. An important class of heavy-tailed distributions are the Risks 2017,5, 10; doi:10.3390/risks5010010 www.mdpi.com/journal/risks
Risks 2017,5, 10 2 of 14 subexponential ones defined by P(X1+X2>d)∼2F(d). The intuition underlying this definition is the principle of a single big jump: X1+X2is large if one of X1,X2is large, whereas the other remains typical. This then motivates R=|X1−X2| X1+X2 (1) being close to 1. In contrast, the folklore is that X1,X2contribute equally to X1+X2with light tails. We are not aware of general rigorous formulations of this principle, but it is easily verified in explicit examples like a gamma or normal F(see further below) and, for a large number of summands rather than just 2, it is supported by conditioned limit theorems (see e.g., ([10] (VI.5))). However, it was recently shown in [4] that these properties of Rhold in greater generality and that asymptotic properties of the corresponding conditioned random variable Yd=RX1+X2>d(2) provide a sharp borderline between log-convexity and log-concavity. In this paper, we provide a wider perspective in terms of both sharper and more general limit results and of the usefulness for visual statistical data exploration. To this end, we propose a feature based nonparametric test. It can be used as a visual aid in identification of log-concavity or heavy-tailed behaviour. It complements earlier ways to detect signs of heavy-tailedness such as the mean excess plot [11]. Further tests based on probabilistic features have been previously utilised in e.g., [12–14]. 2. Background A property holds eventually, if there exists a number y0so that the property holds in the set [y0,∞). Standard asymptotic notation is used for limiting statements. These and basic properties of regularly varying functions with parameter α, denoted RV(α), can be recalled from e.g., [15]. We note that the principle of a single big jump relates to the fact that joint distributions of independent random variables concentrate probability mass to different regions. For example, a distribution with tail function F(x) = e−xαsatisfies F(x) = o(F(x/2)2) for α>1 and F(x/2)2=o(F(x)) for α∈(0, 1), as x→∞. We refer to [16–19] for related work in this direction. It is shown in Lemma 1.2 of [4] that log-concavity or log-concavity of the density is closely related to the occurrence of the principle of a single big jump. A further observation in this direction is the following lemma. It states that contour lines of joint densities of independent variables behave differently for log-concave and log-convex densities, and thereby leads naturally to different concentrations of probability mass of joint densities (recall that a contour line corresponding to a value p∈Rof joint density f:R2→Ris the set of points in the plane defined as {(x,y)∈R2:f(x,y) = p}). Lemma 1. Suppose X1and X2are i.i.d. unbounded non-negative random variables. Assume further that they have a common twice differentiable density function fof the form f(x) = e−h(x), where his a strictly increasing function. If fis log-concave (log-convex), then, for any fixed p∈(0, e−h(0)), there exists a convex (concave) function ψpdefining a contour line of fX1,X2corresponding to psuch that fX1,X2(x,ψp(x)) = pfor all x∈[0, h−1(−log p−h(0))].
Risks 2017,5, 10 3 of 14 Lemma 1implies that log-convex and log-concave densities cause maximal points of joint densities to accumulate into different regions in the plane. Log-convex densities tend to put probability mass near the axis, while log-concave densities have a tendency to concentrate mass near the graph of the identity function. The exponential density is the limiting case where all contour lines are straight lines. More generally, for fα(x) = Cαe−xα, where Ca>0 is an integration constant, the contour lines are circles for α=2, straight lines for α=1, and parabolas for α=1/2. 3. Theoretical Results The emphasis of the paper is on the mathematical formulation of the connection between log-convexity and the principle of a single big jump. However, some additional theoretical results are provided concerning convergence rates of the conditional ratio defined in Equation (3). These rates, or estimates for the rates, are obtained in some standard distribution classes. Their proofs are mainly based on sharp asymptotics of subexponential distributions obtained in [20–22]. Recall that some main classes of such distributions are RV(α), meaning regularly varying ones, where F(x) = L(x)/xα with α>0 and L(·)are slowly varying, Weibull tails with F(x) = e−xαfor some α∈(0, 1), and lognormal tails which are close to the case γ=2 of F(x) = e−logγxfor x≥1 and some γ>1; we refer in the following to this class as lognormal type tails. 3.1. Convergence Properties Define the function g:(0, ∞)→[0, 1]by gX(d) = g(d) = E|X1−X2| X1+X2X1+X2>d. (3) It can be viewed as a generalisation of the function fZdconsidered in [4] and has the same interpretation as in the case with densities: if both X1and X2contribute equally to the sum X1+X2, then gshould eventually obtain values close to 0; similarly, if only one of the variables tends to be the same magnitude as the whole sum, then gis close to 1 for large d. Note also that gis scale independent in the sense that gaX(d) = gX(d/a)for all a>0. Due to this property, two or more samples can be standardised to have, say, equal means in order to obtain graphs on the same scale. In Proposition 1, sharp asymptotic forms of gare exhibited in some classes of distributions. Proposition 1. The following convergence rates hold for gdefined in Equation (3). 1. Let Xbe RV(α)with α>1, eventually decreasing density f. Then, g(d) = 1−c d+o(1/d)where c=2αE[X] α+1. 2. Let Xbe Weibull distributed with α<1. Then, g(d) = 1−o(dα−1). (4) 3. Let Xbe of lognormal type. Then, g(d) = 1−o(logγ−1x/x). (5) Remark 1. In the case of Weibull and lognormal distributions, the implication is that g(d)converges to 1 at a larger rate than their associated hazard rates tend to zero. In addition, inspection of the proof shows lim inf d→∞d|g(d)−1|>0.
Risks 2017,5, 10 4 of 14 This implies that the actual convergence rate can not be substantially larger than in the regularly varying case, where the leading term is explicitly identified. The light-tailed case appears to be more difficult to study than the heavy-tailed case. Difficulty arises mainly from the lack of good asymptotical approximations for probabilities of the form P(X1+X2>d)when P(X1>d)decays much faster than e−d. Interestingly, the full asymptotic form of gcan be recovered in the special case of the normal distribution if we allow Xto obtain negative values. Proposition 2. Suppose that Xis normally distributed with E[X] = 0 and Var[X] = 1/√2. Then, g(d) = c d+o(1/d),where c =E[|X1−X2|]. (6) The following theorem can be used to assess if a sample is coming from a source with log-concave density. It can be seen as a natural continuation as well as a generalisation to [4]. Theorem 1. Assume the density fis twice differentiable and eventually log-concave. Then, lim sup d→∞ g(d)≤1 2. (7) Similarly, if fis eventually log-convex, then lim inf d→∞g(d)≥1 2. (8) 4. Statistical Application: Visual Test Suppose (X1,Y2),(X2,Y2), . . . is a sequence of i.i.d. vectors whose components are also i.i.d. One can formulate the empirical counterpart of Quantity (3) by setting ˆ g(d,n) = n ∑ k=1 Rk1(Xk+Yk>d) n ∑ k=1 1(Xk+Yk>d) , (9) where Rk=|Xk−Yk| Xk+Yk , and 1(A)is the indicator function of the event A. Remark 2. Equation (9) requires as input a two-dimensional sequence of random variables. One can form such a sequence from a real valued i.i.d. source Z1,Z2, . . . , ZNusing any pairing of the Zi. Obvious examples are to take Xk=Z2k−1,Yk=Z2kto take the set {(Xk,Yk)}as all pairings of the Zk or as a randomly sampled subset of these N(N−1)/2 pairings. If the data is truly i.i.d, this should not have any effect on the outcome. 4.1. Examples and Applications A graph of ˆ g(d,n)as function of dcan be used to determine if the data support the density being log-concave or light-tailed behaviour. According to Theorem 1, the graph should then stay below 1/2. Figures 1–4illustrate such graphs using experimental data.
Risks 2017,5, 10 5 of 14 The test method is visual. A similar idea has been used at least in the classical mean excess plot, where one visually assesses if the tail excess in the sample points is increasing at the level, as is the case for heavy tails. (a) (b) (c) Figure 1. Graphs of ˆ g(d,n)for n=10, 000 for Gamma distributed random variables with shapes 0.2, 1 and 5 in figures (a), (b) and (c), respectively. All variables are standardised to have mean 3. (a) (b) (c) Figure 2. Graphs of ˆ g(d,n)for n=10, 000 for Weibull distributed random variables with shapes 0.2, 1 and 5 in figures (a), (b) and (c), respectively. All variables are standardised to have mean 3. Figure 3. Graph of ˆ g(d,n)from a classical set of Danish fire insurance data that can be obtained for instance from data set ‘danish’ in the R package [23]. The data is scaled to have mean 1. The sample is traditionally used to illustrate how heavy-tailed data behaves. A similar set of data was previously used in [24]. The graph supports the usual finding that the data set is heavy-tailed.
Risks 2017,5, 10 6 of 14 Figure 4. The graphs of multiple versions of ˆ g(d,n)based on a dataset obtained from Hansjörg Albrecher (private communication) and related to occurrences of floods in a particular area. The data is scaled to have mean 1. The sample size is n=39. Bivariate vectors (X1,Y1), . . . , (X19,Y19)were sampled several times randomly without replacement from the original data. The overall appearance of the paths points to the data being heavyrather than light-tailed. 4.2. Finer Diagnostics The idea of plotting ˆ g(d,n)as a function of dwas introduced as a graphical test for distinguishing between heavy-tailed and log-concave light tailed distributions. It seems reasonable to ask if the plot can be used for finer diagnostics, in particular to further specify the tail behaviour of Fwhen Fwas found to be heavy-tailed. Such an idea would be based on the rate of convergence of ˆ g(d,n)to 1. To gain some preliminary insight, we simulated R=5000 =5·103and R=5·106i.i.d pairs of r.v.’s X1,X2from an Fwhich was either RV(1.5), lognormal(0,1) or Weibull(0.4). The results are in Figure 5as plots of d1−ˆ g(d,n), with three runs in each subfigure. A first conclusion is that a sample size of R=5000 is grossly insufficient for drawing conclusions about the way in which ˆ g(d,n)approaches 1—random fluctuations take over long before a systematic trend is apparent. The sample size R=5·106is presumably unrealistic in most cases, but even for this, the picture is only clear in the RV case. Here, d1−ˆ g(d,n)seems to have a limit c, as it should be, and the plot is in good agreement with the value 2.4 =2αE[X]/(α+1)of cpredicted by Proposition 1. Whether a limit exists in the lognormal or Weibull case is less clear. The results of Proposition 1 are less definite here, but, actually, a heuristic argument suggests that the limit cshould exist and be 2E[X]. To this end, let X(1)<X(2)be the order statistics. According to subexponential theory (see, in particular, [25]), X(1),X(2)are asymptotically independent given X1+X2>d, with X(1)having asymptotic distribution Fand X(2)being of the form d+e(d)Ewith e(d), and Eas in the proof of Proposition 1. For large d, this gives |X1−X2| X1+X2∼X(2)−X(1) X(2)∼1−X(1)/X(2) 1+X(1)/X(2)∼1−2X(1)/X(2). (10) In the lognormal or Weibull case, one has e(d) = o(d)and so d+e(d)E∼d. Taking expectations gives the conjecture.
Risks 2017,5, 10 7 of 14 0 50 100 150 200 250 300 350 400 450 500 0 1 2 3 4 5 6 7 8 0 50 100 150 200 250 300 350 400 450 500 0 1 2 3 4 5 6 7 8 0 50 100 150 0 1 2 3 4 5 6 7 8 9 10 0 50 100 150 0 1 2 3 4 5 6 7 8 9 10 0 50 100 150 200 250 300 0 5 10 15 0 50 100 150 200 250 300 0 5 10 15 Figure 5. d1−ˆ g(d,n). Pareto in the first row, lognormal in the second, and Weibull in the last. R=5000 (left), R=5×106(right). For the Weibull, 2E[X] = 6.65, and the R=5·106part of Figure 5is rather inconclusive concerning the conjecture. We did one more run with R=5·107over a range of parameters (the variance σ2of log Xfor the lognormal and βfor the Weibull). All X1,X2were normalized to have mean 1 so that the conjecture would assert convergence to c=2. This is not seen in the results in Figure 6. Large values of σ2and small values of βcould appear to give convergence, but not to 2.
Risks 2017,5, 10 8 of 14 0 50 100 150 0 1 2 3 4 5 6 7 8 9 10 2=1/4 2=1/2 2=1 2=2 2=4 0 50 100 150 200 250 300 0 1 2 3 4 5 6 7 8 9 10 =0.1 =0.3 =0.5 =0.7 =0.9 Figure 6. R=5×107. Lognormal (left), Weibull (right). It should be noted that the heuristics give the correct result in the RV case. Namely, here we can take e(d) = dand P(E>x) = 1/(1+x)α. This easily gives E1/(1+E) = α/(α+1)so that Quantity (10) is approximately 1 −2αE[X]/d(α+1), as rigorously verified in Proposition 1. The overall conclusion is that the finer diagnostic value of the method is quite limited, and restricted to RV and sample sizes which may be unrealistically large in many contexts. 5. Proofs Proof of Lemma 1.Suppose his concave and p∈(0, 1). The contour line corresponding to value pis formed as the set of points (x,y)that satisfy fX1,X2(x,y) = p, or equivalently h(x) + h(y) = −log p. (11) For any such pair (x,y)one can solve Equation (11) for yto obtain y=h−1(−log p−h(x)). (12) Firstly, h−1is convex as the inverse of an increasing concave function. Secondly, the composition of an increasing convex function and a convex function remains convex. Thus, as a function of x, Expression (12) defines a convex function when x∈[0, h−1(−log p−h(0))]. Thus, one can define ψp(x) = h−1(−log p−h(x)). If his convex, the proof is analogous. The following technical lemma is needed in the proof of Proposition 1. It applies to Pareto, Weibull and lognormal type distributions. Indeed, condition (13) follows from Proposition 1.2; (ii) of [21] and further needed assumptions are easily verified apart from strong subexponentiality, which is known to hold in the mentioned examples. Lemma 2. Suppose X1and X2are non-negative i.i.d. variables with a common density f, where the hazard rate r(d) = f(d)/F(d)is eventually decreasing with r(d) = o(1). Assume further that P(X1+X2>d)−2P(X1>d)∼2E[X]f(d). (13) Then, 2P(X1>d) + 2f(d)E[X] P(X1+X2>d)=1+o(r(d)). (14) If in addition F(d/2)2=o(F(d)), then P(X1≤d/2, X2≤d,X1+X2>d) P(X1>d)=E[X]r(d) + o(r(d)). (15)