scieee AI-readable full text Open interactive document viewer

Extreme points of Lorenz and ROC curves with applications to inequality analysis

Baíllo, Amparo,Cárcamo Urtiaga, Javier,Mora Corral, Carlos

Abstract

A. Baíllo and J. Cárcamo are supported by the Spanish MCyT grant PID2019-109387GB-I00. C. Mora-Corral is supported by the Spanish MCyT grant MTM2017-85934-C3-2-P.

Full text

J. Math. Anal. Appl. 514 (2022) 126335 Contents lists available at ScienceDirect Journal of Mathematical Analysis and Applications www.elsevier.com/locate/jmaa Regular Articles Extreme points of Lorenz and ROC curves with applications to inequality analysis Amparo Baíllo a, Javier Cárcamo b,∗, Carlos Mora-Corral a aDepartamento de Matemáticas, Universidad Autónoma de Madrid, 28049 Madrid, Spain bDepartamento de Matemáticas, Universidad del País Vasco, Aptdo. 644, 48080 Bilbao, Spain a r t i c l e i n f o a b s t r a c t Article history: Received 12 November 2021 Available online 16 May 2022 Submitted by S. Geiss Keywords: Extreme points Gini index Lorenz curve Lorenz ordering Inequality ROC curve We find the extreme points of the set of convex functions  :[0, 1] →[0, 1] with a fixed area and (0) =0, (1) =1. This collection is formed by Lorenz curves with a given value of their Gini index. The analogous set of concave functions can be viewed as Receiver Operating Characteristic (ROC) curves. These functions are extensively used in economics (inequality and risk analysis) and machine learning (evaluation of the performance of binary classifiers). We also compute the maximal L1-distance between two Lorenz (or ROC) curves with specified Gini coefficients. This result allows us to introduce a bidimensional index to compare two of such curves, in a more informative and insightful manner than with the usual unidimensional measures considered in the literature (Gini index or area under the ROC curve). The analysis of real income microdata illustrates the practical use of this proposed index in statistical inference. © 2022 The Authors. Published by Elsevier Inc. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). 1. Introduction Given a non-empty, compact and convex set of a locally convex space, the collection of its extreme points plays a prominent role in optimization and mathematical programming. By Bauer’s maximum principle (see, e.g., Phelps [25, Proposition 16.6] or Aliprantis and Border [1, 7.69]) any convex, upper-semicontinuous functional defined on this set attains its maximum at an extreme point. In other words, we can search for maximizers of such functionals within the set of extreme points. The relevance of extreme points can be also comprehended through the Krein–Milman theorem (see, e.g., Simon [30, Theorem 8.14]), which is a central result in convex analysis. This theorem affirms that the original set is actually the closed convex hull of its extreme points. Therefore, we can retrieve the entire convex set by knowing the (usually much smaller) subset of extreme points. The practical application of these powerful results goes through the *Corresponding author. E-mail address: [email protected] (J. Cárcamo). https://doi.org/10.1016/j.jmaa.2022.126335 0022-247X/© 2022 The Authors. Published by Elsevier Inc. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). 2A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 explicit computation of the extreme points of the convex set under study, which is usually a difficult task in infinite-dimensional spaces. As a matter of fact, one of the aims of this work is to find the extremal points of some infinite-dimensional collections of functions. Specifically, we define L={:[0,1] →[0,1] : convex and (0) = 0,(1) = 1},(1) R={r:[0,1] →[0,1] : rconcave and r(0) = 0,r(1) = 1},(2) and we consider La={∈L:=(1−a)/2},a∈[0,1],(3) Ra={r∈R:r=a},a∈[1/2,1],(4) where  ·is the usual norm in L1=L1([0, 1]). We will provide a detailed description of the main properties of Laand Ra, which are actually compact and convex sets in the space L1; see Proposition 1. Our main motivation lies in the fact that the curves in Laand Raappear repeatedly in many applied disciplines. On the one hand, an important interpretation of the functions in Lis as Lorenz curves of positive and integrable random variables. In addition, Lais the collection of Lorenz curves with Gini index a. In economics, Lorenz curves are extensively used to provide a graphical representation of the distributions of income or wealth in populations, while the Gini index is perhaps the most prominent inequality measure. On the other hand, Rcan be viewed as the set of Receiver Operating Characteristic (ROC) curves. ROC curves are employed to evaluate the quality of classifiers in probability forecasts (see for example Fawcett [15]) and appear in many scientific disciplines in which classification of binary outcomes is relevant (medical diagnostic, credit scoring, classification of financial transactions, and so on). The Gini coefficient of a ROC curve is defined as twice its area minus 1. However, the area under the ROC curve (AUC scoring) is used more frequently as a measure of the global accuracy of the underlying classifier; see Section 2for details. A second goal of this work is to quantify how “far” two of the curves in (1)(or (2)) can be from one another. Specifically, given two Lorenz (or ROC) curves with fixed Gini indices, we want to compute the maximal L1-distance between them. In other words, for a, b ∈[0, 1], we are interested in computing the distance between the sets Laand Lb(or Raand Rb, for a, b ∈[0, 1/2]). This question turns out to be an infinite-dimensional convex maximization problem with two linear constraints. We will solve this optimization problem by a careful analysis of the distance between the extreme points of the considered sets. This approach also allows us to identify the maximizers of the underlying functional. This paper is structured as follows: Section 2reviews the main interpretations of the elements in Laand Ra. We recall the definition of the Lorenz curve and the Gini index of an integrable variable, as well as the main elements to describe ROC curves. In Section 3, we determine the set of extreme points of Laand Ra. Section 4is devoted to the computation of the maximal L1-distance between two of these sets. We also identify extremal curves, that is, functions for which this maximal distance is attained, and show their connection with stochastic orderings. As an application, in Section 5we introduce a bidimensional inequality index to compare different characteristics of two Lorenz or ROC curves and enumerate its main properties. In the case of Lorenz curves, this new index measures simultaneously inequality and dissimilarity, while for ROC curves, it indicates accuracy and dissimilarity as well. We consider the empirical version of this inequality index to be used in practice. We also propose to combine a bootstrap approach with a non-parametric set estimation technique to obtain a confidence region for the population index. The procedure can be further applied to test simple hypotheses related to the index. These techniques have been implemented in the software R and are illustrated via the analysis of some real income microdata samples from Spain. Finally, the proofs of the main results are collected in Section 6. A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 3 2. Lorenz and ROC curves and the Gini index In this section we recall the definitions of Lorenz and ROC curves and the Gini index, as well as their connections with the sets La(in (3)) and Ra(in (4)). Also, we briefly describe the scientific fields in which these concepts are widely used. 2.1. Lorenz curves and the Gini index in economics Let Xbe a positive random variable with finite mean μ >0and cumulative distribution function F(x) =P(X≤x), for x ≥0. The Lorenz curve of the variable X(or of the distribution F) is (t)= 1 μ t 0 F−1(x)dx, 0≤t≤1,(5) where F−1(x)=inf{y≥0:F(y)≥x}(6) (0 <x <1) is the quantile function of X, that is, the generalized inverse of F. If Xmeasures income in a population, for each value t ∈[0, 1], the function in (5)gives us the (normalized) total income accumulated by the proportion tof the poorest in that population. Note that F−1is nondecreasing, μ =1 0F−1(x) dxand (t) =F−1(t)/μ a.e. t ∈(0, 1). In particular, is continuous except perhaps at the point 1and has positive second derivative  a.e. Moreover, as the quantile function in (6) characterizes the probability distribution, determines the distribution of the underlying variable up to a (positive) scale transformation. Explicit analytic expressions for the Lorenz curves of the usual parametric distributions can be found in Kleiber and Kotz [20, Section 2.1.2]. It is easy to check that the set Lin (1)is the closure (with respect to the pointwise convergence) of the set of Lorenz curves of positive and integrable random variables with strictly positive expectation. For simplicity, we will refer to Las the class of Lorenz curves. By convexity, for every  ∈Lit holds that pi ≤≤pe,(7) where pe(t)=t(0 ≤t≤1) and pi(t)=0,if 0 ≤t<1, 1,if t=1.(8) Fig. 1shows a graphical representation of the inequalities in (7). The function pe is called the perfect equality curve as it corresponds to the Lorenz curve of a Dirac delta measure, i.e., the probability measure corresponding to a population in which all individuals have equal (and positive) incomes. Additionally, pi is the perfect inequality curve because it can be viewed as the limit (when the total number of individuals tends to infinity) of Lorenz curves in finite populations where only one person accumulates all the wealth. Note that the function pi defined in (8)(see also Fig. 1), which is not a proper Lorenz curve, belongs to L. In practice, it is very common to synthesize the information of the Lorenz curve in a single numerical value that quantifies income inequality. Different characteristics, functionals and values of the Lorenz curve are employed to construct those inequality indices; see Arnold and Sarabia [4]. The most popular inequality measure derived from the Lorenz curve is the Gini index. This index has a vast number of interesting 4A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 Fig. 1. A Lorenz curve together with the perfect equality and inequality curves. interpretations and representations; see Yitzhaki and Schechtman [35, Chapter 2]. One possible way to define it is the following: G()=2 1 0 (t−(t)) dt=1−2.(9) Therefore, we have that for a ∈[0, 1], the set Ladefined in (3)is precisely La={∈L:G()=a},(10) the collection of Lorenz curves with Gini index a. Observe that G()= −pe pe −pi.(11) The denominator in (11)equals 1/2 (the maximum L1-distance between Lorenz curves) and acts as a normalizing constant so that 0 ≤G(X) ≤1. Graphically, G()is twice the shaded area in Fig. 1. The Gini index has many convenient properties: it is scale-free (because the Lorenz curve is itself invariant under positive scaling); it can be computed whenever the considered random variable is integrable (so finite second moment is not necessary); it is normalized so that it takes values between 0 (perfect equality) and 1 (perfect inequality); it has a simple and effective interpretation (small values of this index amount to fair income distributions, whereas high values indicate unequal distributions). Another important instrument to compare distributions according to inequality is the so-called Lorenz ordering. Let X1and X2be two variables with Lorenz curves 1and 2, respectively. It is said that X1is less than or equal to X2in the Lorenz order, written X1≤LX2, if 1(t) ≥2(t), for all t ∈[0, 1]. In this case, we have that pe ≥1≥2, where pe is the perfect equality curve defined in (8). In other words, income is distributed in a more equitable manner in X1than in X2. 2.2. ROC curves in machine learning Binary supervised classification is one of the main statistical techniques in machine learning; see Hastie et al. [17]. In this context, we want to classify an object in one of two groups labelled 0 and 1. We observe a random vector (X, Y), where Xis the predictor, usually a multidimensional (or functional) variable, and A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 5 Y∈{0, 1}indicates the group membership. Usual procedures in machine learning combine the information in Xto construct a score or marker Sto predict Y. This score is typically an estimate of the posterior probability P(Y=1|X=x)or some increasing function of this quantity. We can assume that members of the class {Y=0}have often smaller values of the score; if not, we can interchange the labels. Then, higher values of Sprovide stronger evidence in favour of the event {Y=1}. We denote by Fi(t) =P(S≤t|Y=i), i =0, 1and t ∈R, the conditional distribution functions of the score in the groups. Any reasonable x ∈Rcan be used as a cut-off to obtain a classification rule by assigning {Y=1}whenever {S>x}, and {Y=0}if {S≤x}. In practice, the selected value for the cut-off point xtypically depends on the prior probability of the groups and misclassification costs. The probability of detection of the classifier generated by the threshold xis PD(x)=P(S>x|Y=1)=1−F1(x) and the probability of false detection is PFD(x)=P(S>x|Y=0)=1−F0(x). The functions PD(x)and PFD(x)are also known as hit rate and false alarm rate, respectively. In this context, ROC curves are commonly used to evaluate the predictive ability of binary classifiers. Formally, the ROC curve is the parametric curve in [0, 1]2given by {(PFD(x), PD(x)) :x ∈R}. If Fi, i =0, 1are continuous and strictly increasing, the ROC curve is the graph of the function r(t)=1−F1(F−1 0(1 −t)),t∈(0,1), with r(0) =0and r(1) =1. We observe that a good classifier should have a high probability of detection and low probability of false detection. Therefore, classifiers with ROC curves close to the constant 1 are preferible. A perfect classifier has the ROC curve r(0) =0and r(x) =1, x ∈(0, 1], while a random classifier has a ROC curve on the diagonal of [0, 1]2. Therefore, the ROC curve gives information about the precision of a binary classifier. Further, ROC curves are concave (see Lloyd [22]) and the area under the ROC curve (AUC) is used to evaluate the performance of the classifier. In particular, Rain (4)is the set of ROC curves with AUC scoring a. At this point it should be commented that within the machine learning community ROC curves are considered convex, as they are viewed from the line {(x, 1) :x ∈(0, 1]}. We follow here the usual terminology in mathematics. Finally, if two classifiers have ordered ROC curves, the above curve corresponds to the classifier that is uniformly better than the other one; that is, it gives better results for each cut-off point xthat is selected to carry out the classification procedure. 3. Extreme points of Laand Ra The sets Ladefined in (3)(see also (10)) and Rain (4)are clearly convex. The following result asserts that they are also compact in L1. Proposition 1. For each a ∈[0, 1] and b ∈[1/2, 1], the sets Laand Rbare compact in L1. As Laand Rbare convex and compact, their extreme points acquire special relevance. We recall that an extreme of a convex set is a point that cannot be expressed as a proper convex combination of other points within the set. Formally, given a convex set C, x ∈Cis an extreme point of Cif x =tx1+(1 −t)x2, for some t ∈(0, 1) and x1, x2∈C, implies that x1=x2. In the following we denote by Ext(C)the set of extreme points of C. 6A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 Fig. 2. The functions a x1, with a =0.5and x1=0.25 (left panel), ma x2, with a =0.5and x2=0.7 (central panel) and na x1,x2, with a =0.5and x1=0.3, x2=0.9 (right panel). The next theorem determines the set of extreme points of La. Theorem 1. For a ∈[0, 1], we have that Ext(La)=a x1:x1∈[0,a]∪ma x2:x2∈(a, 1)∪na x1,x2:x1∈(0,a),x 2∈(a, 1), where a x1, ma x2and na x1,x2are the piecewise affine functions of Lasuch that a x1:⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ 0→ 0 x1→ 0 1−→ 1−a 1−x1, ma x2:⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ 0→ 0 x2→ x2−a 1→ 1, na x1,x2:⎧ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎩ 0→ 0 x1→ 0 x2→ x2−a 1−x1 1→ 1 (12) (with the notation a x1(1−) = limt↑1a x1(t)and the convention 1 1=pi in (8)). Theorem 1summarizes the information of La(an infinite-dimensional collection) in the set of its extreme points, which has only dimension 2. Furthermore, as Lais compact in L1, it identifies all possible maximizers of convex and continuous functionals. To prove Theorem 1(Section 6) we first show that twice differentiation determines an affine isomorphism between Laand the set of non-negative measures on (0, 1) with some restrictions. Afterwards, we identify those combinations of delta measures that are extreme points. In Fig. 2we have depicted various extreme points of La, with a =0.5. The probabilistic and economic meaning of some of these Lorenz curves is described in Section 4. Remark 1. Let us consider the map T(x)=1−(1 −x),∈L,x∈[0,1].(13) Observe that, for a ∈[0, 1], Tdefines a natural bijection between Laand R(1+a)/2whose inverse is itself. Further, for t ∈[0, 1] and 1, 2∈L, it holds that T(t1+(1−t)2)=tT(1)+(1−t)T(2). Therefore, Tpreserves extreme points and the set Ext(Ra)is directly obtained from Theorem 1: Ext(Ra)={T():∈Ext(L2a−1)},for a∈[1/2,1]. A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 7 4. Maximum distance between Lorenz or ROC curves Given two Lorenz (or ROC) curves with fixed Gini indices, in this section we quantify how “far” they can be from one another. We only consider the problem for Lorenz curves, as the corresponding question for ROC curves is analogous (see Remark 3). We also show another property of those curves for which the maximal distance is attained, which is connected with stochastic orders. 4.1. Computation of the maximum distance We are interested in computing the value d(La, Lb)(a, b ∈[0, 1]), for a suitable metric don L ×L, where Lis defined in (1), and Laand Lbas in (3)(see also (10)). Theorem 1is extremely useful for this purpose. If dis defined through a norm, dis a convex and continuous functional on (the convex set) La×Lb. Therefore, as long as Laand Lbare compact, by Bauer’s maximum principle, the supremum of don La×Lbis attained in Ext(La×Lb) =Ext(La) ×Ext(Lb). Thus, thanks to Theorem 1, we reduce the calculation of d(La, Lb) to a finite-dimensional problem. The exact computation of d(La, Lb) will eventually depend on the particular choice of the metric d. We note that the Gini coefficient itself is defined in terms of a (normalized) L1-distance between Lorenz curves; see formula (11). This is indeed a sensible and convenient choice to measure dissimilarities between Lorenz curves (and their associated probability distributions). The L1distance between Lorenz curves has also been used in Zheng [36] related to almost stochastic dominance of Leshno and Levy [21]. Explicitly, endow the set Lin (1)(or, analogously, the set Rin (2)) with the Lorenz distance dL(1, 2)= 1−2 pe −pi=21−2=2 1 0|1−2|, 1, 2∈L.(14) Remark 2. Let us consider X1, X2two positive and integrable random variables with positive expectation and Lorenz curves 1and 2, respectively. We can define d(X1, X2) =dL(1, 2). We have that dis actually a pseudo-metric (on the space of positive and integrable random variables) because d(X1, X2) =0holds if and only if X1=st cX2, where c >0is a constant and ‘=st’ stands for stochastic equality. By (7), the diameter of Lwith respect to the metric dLis diam(L)=sup{dL(1, 2):1, 2∈L}=dL(pe, pi)=1. We further observe that Lain (10)is the set of  ∈Lsuch that dL(, pe) =a. For any fixed a, b ∈[0, 1], La and Lbare compact sets in L1(see Proposition 1). Therefore, from Theorem 1, the maximum M(a, b)=max{dL(1, 2):1∈L aand 2∈L b}(15) is attained at Ext(La) ×Ext(Lb). Note that M(a, b)= max 1,2∈L 21−2subject to 1 0 1=1−a 2and 1 0 2=1−b 2. Therefore, the computation of (15), which in principle is an infinite-dimensional convex maximization problem with two linear constraints, is reduced to a finite-dimensional optimization problem. 8A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 Fig. 3. The functions − 0.4(left panel) and + 0.5(right panel). Definition 1. We say that the pair (1, 2) ∈L ×Lis extremal if dL(1, 2)=M(G(1),G(2)). The pair of probability distributions associated to an extremal pair of Lorenz curves will also be called extremal distributions. For notational convenience, we rename the functions a aand a 0in (12)as − aand + a, respectively. In other words, for 0 ≤a ≤1, − a, + a∈L aare defined as − a(t)=max0,t−a 1−aand + a(t)=(1 −a)t, if 0 ≤t<1, 1,if t=1 (16) (with the agreement that − 1≡pi defined in (8)). These two functions will play an essential role in the rest of the section. In Fig. 3we display two of these functions. The following theorem, which is the main theoretical result of this section, provides an explicit expression for M(a, b)and shows that this maximum distance is precisely attained at functions of the form (16). The computation of M(a, b), which begins at Theorem 1and is collected in Section 6, reveals that this issue is more delicate and complex than expected. Theorem 2. For 0 ≤a, b ≤1, let M(a, b)be as in (15). We have that M(a, b)=(1 −a)b2+(1−b)a2 a+b−ab (17) (the value M(0, 0) =0is taken by continuity). Moreover, (− a, + b)and (+ a, − b)are pairs of extremal Lorenz curves within the set La×Lb. Theorem 2asserts that the maximum distance in (15)is attained at the pairs (− a, + b)and (+ a, − b). Hence, the associated probability distributions (unique up to positive scale transformations) are extremal. The function − ais the Lorenz curve of a population in which a proportion aof the people have 0income and the rest, a proportion 1 −a, have equal and positive income. In other words, (up to positive scale transformations) − ais the Lorenz curve of a variable X− awith Bernoulli (1 −a) distribution, that is, P(X− a=0) =aand P(X− a=1) =1 −a. On the other hand, + bis not a proper Lorenz curve, but it can be expressed as the limit (as ngoes to infinity) of Lorenz curves of populations with nindividuals where n −1 of them fairly share a proportion (1 −b)of the wealth and there is only one “lucky person” who accumulates A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 9 the rest of the total wealth (the proportion b). Hence, + bcan be obtained as the limit (as ngoes to infinity) of the Lorenz curves of a sequence of random variables X+ b(n)taking the values (1 −b)n/(n −1) (with probability 1 −1/n) and bn (with probability 1/n). The minimum value of dL(1, 2)for 1∈L aand 2∈L bis easily seen to be |b−a|(see Lemma 1in Section 6). With this and Theorem 2we can see the range of values of the distance dL(1, 2). Corollary 1. For 0 ≤a, b ≤1, we have that {dL(1, 2): 1∈L aand 2∈L b}=[|b−a|,M(a, b)] , where M(a, b)is given in (17). Proof. The minimum and maximum of dL(, m)among  ∈L aand m ∈L bare |b−a|and M(a, b), as calculated in Lemma 1and Theorem 2, respectively. Now, the set {dL(, m) : ∈L aand m ∈L b}is compact and connected as it is the image by the continuous function dLof the compact and connected set La×Lb(in the space L1). The result follows.  Theorem 2also allows us to compute the maximum distance between Lorenz curves with a given difference of their Gini indices. Corollary 2. For −1 ≤c ≤1, let us consider M∗(c)=max{M(a, b):a, b ∈[0,1] and b−a=c}. We have that M∗(c)=M(ac,a c+c)with ac=(4−c−8+c2)/2 (18) and M∗(c)=8−8+(c2+8) 3/2 c2+4 .(19) Definition 2. We say that the pair (1, 2) ∈L 2is super-extremal if dL(1, 2)=M∗(G(1)−G(2)). The associated pairs of probability distributions will be also called super-extremal distributions. Obviously, each super-extremal pair is extremal because it always holds that M(a, b)≤M∗(a−b),for 0 ≤a, b ≤1. However, from Theorem 2and for any 0 ≤c ≤1, among all the pairs (− a, + a+c)and (− a+c, + a)(with a ∈[0, 1 −c]) of extreme Lorenz curves with a value cfor the difference of their Gini indices there are only two super-extremal curves. Namely, the pairs corresponding to a =acin (18). Observe that M∗(0) is the maximum possible distance between Lorenz curves with equal Gini indices. By Theorem 2and Corollary 2, we have that the maximum distance between two income distributions both with Gini indices equal to ais 16 A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 Fig. 8. The inequality index ˆ Ifor Spanish equivalised disposable income corresponding to year 2008 (X1) and each of the years in the span 2009–2020 (X2). We finally give some insight on the surprising value of ˆ Ifor 2020, that would indicate a reduction in Spanish income inequality with respect to 2008, in spite of the economic effects of the Covid-19 pandemic on the world economy and, particularly, on the Spanish one (see, e.g., IMF [19]). As explained by the INE in a press release, the LCS collects information from a sample of households regarding their living conditions at the time of the interview (fourth trimester of the year) as well as their income in the previous year. Thus, the impact of the pandemic on the LCS-2020 is only partially reflected. 5.4. A confidence region for the inequality index I(1, 2) In this section, we employ a standard bootstrap scheme (see Efron [12]), a widely used statistical technique based on plug-in estimation and resampling, combined with a non-parametric set estimation technique to obtain a confidence region for the bidimensional relative inequality index Iin (23). Let xj1, ..., xjnjdenote the observations from population Xjfor j=1, 2. Based on the samples we construct a confidence region for the population index I(1, 2)at the confidence level 1 −αin the following way. First, we extract Bbootstrap samples from each of the original samples {xji}nj i=1, j=1, 2, and we compute the corresponding bootstrap version of the empirical inequality index: Original samples Bootstrap samples Bootstrapped indices x11,...,x 1n1 x21,...,x 2n2 −→ −→ x∗b 11,...,x ∗b 1n1 x∗b 21,...,x ∗b 2n2−→ ˆ I∗b,b=1,...,B Bootstrap usually gives good results with commonly large sample sizes available in inequality analysis and machine learning settings. After obtaining the bootstrap sample ˆ I∗1, ..., ˆ I∗Bof empirical inequality indices, we use the local convex hull (LoCoH) (also called k-nearest neighbour convex hull in the literature of set estimation, see Getz and Wilmers [16]) to construct a confidence region for I(1, 2). The LoCoH provides results that adapt well to the bootstrap sample and to the shape of the region Δ where Itakes values. The construction of the LoCoH is as follows. For a fixed integer k>0, we construct the convex hull of each ˆ I∗band its k−1nearest neighbours. Then these hulls are ordered according to their area, from smallest to largest. The LoCoH is the polygonal region that results of progressively taking the union of the hulls A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 17 Table 1 Values of kand p(k)for Spanish income data in 2008 and 2019 (in bold the values of k and p(k)such that this proportion is closest to 95%). k100 200 300 400 500 600 700 800 900 p(k0.05) 0.923 0.932 0.935 0.936 0.942 0.933 0.938 0.934 0.938 from the smallest upwards, until a proportion 1 −αof ˆ I∗b’s is included in the region. We use the notation ˆ S(k) = LoCoHk(ˆ I∗1, ..., ˆ I∗B)to denote the resulting confidence region. Selecting an adequate value for the number of neighbours kis a relevant matter, as it critically affects the shape of the LoCoH, but no automatized selection procedures have yet been considered in the literature. We propose to use a leave-one-out scheme to select the “optimal” kamong the values in a grid k1, ..., kM. The idea is to determine, for each b =1, ..., B, the LoCoH ˆ S(b)(k) = LoCoHk(ˆ I∗β, β=1, ..., B, β=b) based on the sample of bootstrapped inequality indices from which the b-th one has been removed. Then, we compute the proportion of times that ˆ S(b)(k)contains the left-out ˆ I∗b p(k)= B  b=1 1ˆ S(b)(k)(ˆ I∗b)/B. Our proposal for choosing the number of neighbours in the LoCoH is kα=argmin k∈{k1,...,kM}|p(k)−(1 −α)|.(26) We use the procedure described above to compute a bootstrap confidence region for the inequality index Icorresponding to two of the samples considered in Section 5.3. We have chosen the equivalised disposable income of Spain in 2008 (n1= 12987) and in 2019 (n2= 15861). The Gini index in this case takes the value 32.9% for both years, and the empirical bidimensional inequality index is ˆ I=(−10−4, 0.006). As the components of the index are not both close to the origin (0,0), we can conclude that the distribution of income was not exactly the same in 2008 and 2019. Indeed, this is confirmed by the two empirical Lorenz curves and their difference (re-scaled by the maximum absolute difference to improve the visualization) plotted in Fig. 9: in 2019 the poorest half of the Spanish population had a smaller cumulative proportion of income than in 2008. Regarding the proximity of the two Lorenz curves in Fig. 9(a) and the low values of the two components of ˆ I, let us note that Lorenz curves for the income of the same country in two different years are never radically different. We have extracted B= 1000 bootstrap samples from each of the data sets corresponding to 2008 and 2019 and computed the resulting empirical indices ˆ I∗b, b =1, ..., 1000. We have determined the bootstrap LoCoH confidence region ˆ S(k0.95), at the confidence level of 95%, for the relative inequality index Icomparing 2008 and 2019. The value of k0.95 was selected from the grid (100, 200, ..., 900) via the leave-one-out procedure described before and according to (26)with 1 −α=0.95. Table 1displays the proportion p(k)of times that the left-out index ˆ I∗bis contained in the LoCoH ˆ S(b)(k) constructed with the remaining 999 indices. The number k0.95 = 500 of nearest neighbours yielding the proportion p(k)nearest to 0.95 results in the LoCoH ˆ S(500) of Fig. 10. Remark the contrast with Fig. 11, where we have plotted the bootstrap inequality indices corresponding to years 2008 and 2016 in Spain and the confidence region constructed using the LoCoH with k0.95 = 600. 5.5. Hypothesis tests for the inequality index I(1, 2) Once determined a confidence region for I(1, 2)at a level (1 −α)(as explained in the previous section), we can use it to carry out the statistical test (with significance level α) 18 A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 Fig. 9. (a) Empirical Lorenz curves, ˆ 1and ˆ 2, for the equivalised household income in Spain in 2008 and 2019 respectively; (b) Difference ˆ 1−ˆ 2re-scaled by ˆ 1−ˆ 2∞. Fig. 10. Confidence region for the bidimensional relative inequality index Iof Spain in 2008 and 2019. The crossed point in white is the empirical index ˆ I. H0:I(1, 2)=I0versus H1:I(1, 2)=I0, where I0is a known fixed value in the region Δgiven in (24). The procedure to test the simple hypothesis H0:I=I0follows by the duality between confidence regions and hypothesis tests: we reject H0at level α whenever I0does not belong to the confidence region for I(1, 2)at a level (1 −α). For instance, the confidence region at level 95% in Fig. 10 does not intersect the left diagonal L2, which means that there is evidence to reject that the relative inequality index Icomparing 2008 and 2019 lies on any point of the left diagonal. In other words, we can affirm (with significance level 5%) that 2019 did not distribute income more fairly than 2008. However, for the same years, the confidence region ˆ S(500) intersects the right diagonal L1, so we cannot reject that the 2019 Lorenz curve is below that of 2008. The situation is even more clear for the years 2008 and 2016 (see Fig. 11). A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 19 Fig. 11. Confidence region for the bidimensional relative inequality index Iof Spain in 2008 and 2016. The crossed point in white is the empirical index ˆ I. 6. Proofs: extremal points and maximum distance Here we collect the proofs of Proposition 1and Theorems 1and 2. First, we enumerate some regularity properties of the functions in the set Ldefined in (1)that will be useful throughout this section. 6.1. Regularity properties and compactness of La We start with a slight change in the definition of the functions in the set Lof (1). Given  ∈Lwe redefine the value of at 1as (1) =sup [0,1) . This redefinition is motivated by the fact that, as shown in the following proposition, functions in Lbecome continuous in [0, 1]. In addition, the convexity of Land Laremains true. With this definition, Lbecomes the set of convex  :[0, 1] →[0, 1] such that (0) =0and (1) =sup [0,1) . In the following proposition we denote by W1,1(0, 1) the Sobolev space W1,1in the interval (0, 1), which is equivalent to the set of absolutely continuous functions in [0, 1]; it is endowed with the norm W1,1(0,1) =+, where is the distributional derivative of , which coincides a.e. with the derivative of . For α∈(0, 1) we denote by W1,∞(0, α)the Sobolev space W1,∞in the interval (0, α), which is equivalent to the set of Lipschitz continuous functions in [0, α]; it is endowed with the norm W1,∞(0,α)=sup (0,α)||+ ess sup (0,α)||. See, e.g., Brezis [6, Chapter 8] or Evans and Gariepy [14, Chapter 4] for the definition and properties of these spaces. Proposition 5. Let a ∈[0, 1] and  ∈L a. Then (a) The function is non-decreasing, absolutely continuous in [0, 1] and Lipschitz in [0, α]for each α∈(0, 1). Moreover, W1,1(0,1) ≤1+1−a 2, 20 A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 and for each α∈(0, 1), W1,∞(0,α)≤1+ 1−a (1 −α)2. (b) The function is locally of bounded variation, the right derivative (x+)exists for all x ∈[0, 1) and is non-decreasing. Moreover, (0+) ≥0. (c)  is a non-negative Radon measure. Proof. Convex functions are locally Lipschitz (see Evans and Gariepy [14, Theorem 6.3.1] or Simon [30, Theorem 1.19]), have a first derivative locally of bounded variation (see Evans and Gariepy [14, Theorem 6.3.3]) and have a second derivative in the sense of distributions (see Evans and Gariepy [14, Theorem 6.3.2] or Simon [30, Theorem 1.29]), which in fact is a non-negative Radon measure. Further, the right derivative (x+), exists for all x ∈[0, 1) and is non-decreasing (see Simon [30, Theorem 1.26]). As (0) =0and  ≥0, we necessarily have that (0+) ≥0, so (x+) ≥0for all x ∈[0, 1). By the version of the fundamental theorem of calculus for convex functions (see Simon [30, Theorem 1.28]), is non-decreasing. In addition, the derivative of exists a.e. and coincides a.e. with the right derivative, so ≥0a.e. In particular, = 1 0 (t)dt=(1) −(0) ≤1. We conclude that W1,1(0,1) ≤1 +1−a 2. On the other hand, we observe that the affine function s :[0, 1] →Rgiven by s(x) =(α+)(x −α) +(α) is a supporting line of at the point (α, (α)). By convexity, we hence have that s ≤and then, 1−a 2=≥ 1 α (t)dt≥ 1 α s(t)dt=(α+)(1 −α)2 2+(α)(1 −α)≥(α+)(1 −α)2 2. We conclude that ess sup (0,α) ≤(α+)≤1−a (1 −α)2. Consequently, W1,∞(0,α)≤1 +1−a (1−α)2. The set Lais clearly convex. We are now ready to prove Proposition 1. ProofofProposition1in Section 3(Compactness of Laand Rb). Let {n}n∈Nbe a sequence in La. By Proposition 5, {n}n∈Nis bounded in W1,1(0, 1), so by the Rellich–Kondrachov theorem (see Brezis [6, Theorem 8.8]), there exists a subsequence (not relabelled) and an  ∈L1such that n→in L1as n →∞. This also implies that G() =a. On the other hand, for each α∈(0, 1) we have by Proposition 5(a) that {n}n∈Nis bounded in W1,∞(0, α), so by the Ascoli–Arzelà theorem (see Brezis [6, Theorems 4.25 and 8.8]), for a further subsequence, n→uniformly in [0, α]as n →∞. In particular, (0) =0and 0 ≤ ≤1in [0, α]. As the pointwise limit of convex function is a convex function, we obtain that is convex in [0, α]. Therefore, 0 ≤ ≤1in [0, 1) and is convex in [0, 1). We redefine (1) as (1) =sup [0,1), so that becomes continuous in [0, 1]. We also obtain that 0 ≤ ≤1in [0, 1] and is convex in [0, 1]. Therefore,  ∈L aand the proof is finished.  A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 21 6.2. Proof of Theorem 1in Section 3(extreme points of La) We will use an alternative description of the elements in Lain terms of positive measures concentrated on the interval (0, 1). The main idea is based on the following fact: any curve  ∈L ais univocally determined by its second derivative, , together with the conditions (0) =0and G() =a(or, equivalently,  =1−a 2). Given a ∈[0, 1], we denote by Mathe set of non-negative Radon measures μconcentrated on the interval (0, 1) and such that 1 0 (1 −s)2dμ(s)≤1−aand 1 0 s(1 −s)dμ(s)≤a. (27) Proposition 6. For a ∈[0, 1], the map Ta:La→M adefined by Ta() = is an affine isomorphism with inverse T−1 a:Ma→L agiven by T−1 aμ(x)=⎡ ⎣1−a− 1 0 (1 −s)2dμ(s)⎤ ⎦x+ x 0 (x−s)dμ(s),x∈[0,1].(28) Proof. First we see that the map Tais well defined. Given  ∈L a, we have from Proposition 5that  is a non-negative Radon measure. As is locally of bounded variation, (t)=(0+)+ t 0 d(s),a.e. t∈(0,1). As is locally Lipschitz and (0) =0, for x ∈[0, 1], we have that (x)= x 0 (t)dt= x 0 ⎡ ⎣(0+)+ t 0 d(s)⎤ ⎦dt =(0+)x+ x 0 (x−s)d(s), (29) where for the last equality we have used Fubini’s theorem. Integrating in x ∈(0, 1) equality (29)(and by Fubini’s theorem again) we obtain the restriction 1−a 2==(0+) 2+1 2 1 0 (1 −s)2d(s).(30) As (0+) ≥0, from (30)we directly obtain the first inequality of (27). On the other hand, (29)and (30) show that (x)=⎡ ⎣1−a− 1 0 (1 −s)2d(s)⎤ ⎦x+ x 0 (x−s)d(s),x∈[0,1].(31) Hence, imposing (1) ≤1we have the second inequality of (27). Now, for μ ∈M a, we define as in the right-hand side of (28)and we will check that  ∈L a. First, (0) =0and 22 A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 1 0 (x)dx=1 2⎡ ⎣1−a− 1 0 (1 −s)2dμ(s)⎤ ⎦+ 1 0 1 s (x−s)dxdμ(s)=1−a 2. Further, thanks to the second inequality of (27), we obtain that (1) = 1 −a+ 1 0 s(1 −s)dμ(s)≤1. By Leibniz integral rule (differentiation under the integral sign), it can also be checked that (x)=1−a− 1 0 (1 −s)2dμ(s)+ x 0 dμ(s),a.e. x∈(0,1) (32) and hence  =μas measures. (33) As μis positive, from (32)we have that is essentially non-decreasing, is convex and (x)≥1−a− 1 0 (1 −s)2dμ(s)≥0,a.e. x∈(0,1), by the first inequality of (27), so is non-decreasing. In particular,  ≥0. This shows that  ∈L a. Finally, we prove that the maps Taand (28)are mutually inverse. Given  ∈L a, if we apply first Taand then (28)we get back thanks to (31). Conversely, given μ ∈M a, if we apply first (28)and then Tawe recover μby (33). Since Tais affine, the proof is concluded.  Next we calculate Ext(Ma). We denote by δxthe Dirac measure at x ∈[0, 1]. Proposition 7. For a ∈[0, 1], we have that Ext(Ma)={0}∪1−a (1 −x1)2δx1:x1∈(0,a]∪a x2(1 −x2)δx2:x2∈(a, 1) ∪x2−a (1 −x1)(x2−x1)δx1+a−x1 (x2−x1)(1 −x2)δx2:x1∈(0,a),x 2∈(a, 1). Proof. The proof is divided into several smaller results. Step 1: The null measure μ ≡0 ∈Ext(Ma). This is direct as all the measures in Maare non-negative. Step 2: For all x1∈(0, a], the measure μ =1−a (1−x1)2δx1∈Ext(Ma). Clearly, μ ∈M a. Assume that μ =t1μ1+t2μ2, for some t1, t2>0with t1+t2=1and μ1, μ2∈M a. Then μi=βiδx1and, due to (27), βi≥0,β i(1 −x1)2≤1−a, βix1(1 −x1)≤a, i =1,2. A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 23 Thus, 1−a (1 −x1)2=t1β1+t2β2≤t1 1−a (1 −x1)2+t2 1−a (1 −x1)2=1−a (1 −x1)2, so, necessarily, β1=β2=1−a (1−x1)2, and, hence, μ1=μ2. Therefore, μ ∈Ext(Ma). Step 3: For all x2∈(a, 1), the measure a x2(1−x2)δx2∈Ext(Ma). The proof is similar to the one of the previous step and it is therefore omitted. Step 4: If for some x1∈(0, a]and α∈R \{0, 1−a (1−x1)2}, μ =αδx1∈M a, then μ /∈Ext(Ma). The fact μ ∈M aimplies that 0 <α< 1−a (1−x1)2. Therefore, for ε >0small enough, we have that μ ±εδx1∈M asince both are positive measures and the restrictions (27)are satisfied; indeed, 1 0 (1 −s)2d(μ±εδx1)(s)=(α±ε)(1 −x1)2<1−a and 1 0 s(1 −s)d(μ±εδx1)=(α±ε)x1(1 −x1)<1−a 1−x1 x1≤a. Finally, we can write μ =1 2(μ +εδx1) +1 2(μ −εδx1), and, hence, μ /∈Ext(Ma). Step 5: If for some x2∈(a, 1) and α∈R \{0, a x2(1−x2)}, μ =αδx2∈M a, then μ /∈Ext(Ma). The proof is similar to that of Step 4 and it is left to the reader. Step 6: For all x1∈(0, a)and x2∈(a, 1), μ =x2−a (1−x1)(x2−x1)δx1+a−x1 (x2−x1)(1−x2)δx2∈Ext(Ma). It is immediate to check that 1 0 (1 −s)2dμ(s)=1−aand 1 0 s(1 −s)dμ(s)=a. Therefore, μ ∈M a. Moreover, if μ =t1μ1+t2μ2for some t1, t2>0with t1+t2=1and μ1, μ2∈M a, then 1 0 (1 −s)2dμi(s)=1−aand 1 0 s(1 −s)dμi(s)=a, i =1,2. Furthermore, as μi=2 j=1 βijδxjfor some βij ≥0(for i, j=1, 2), we have that 2  j=1 βij(1 −xj)2=1−aand 2  j=1 βijxj(1 −xj)=a, i =1,2, and, hence, 24 A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 βi1=x2−a (1 −x1)(x2−x1),β i2=a−x1 (x2−x1)(1 −x2),i=1,2. Therefore, μ1=μ2. Consequently, μ ∈Ext(Ma). Step 7: If for some α1, α2>0and 0 <x 1<x 2<1with α1=x2−a (1 −x1)(x2−x1)or α2=a−x1 (x2−x1)(1 −x2) the measure μ =2 i=1 αiδxi∈M a, then μ /∈Ext(Ma). By (27), we have that 2  i=1 αi(1 −xi)2≤1−aand 2  i=1 αixi(1 −xi)≤a. (34) If both inequalities in (34)were equalities, we necessarily have that α1=x2−a (1 −x1)(x2−x1)and α2=a−x1 (x2−x1)(1 −x2), against our assumption. Therefore, at least one of the two inequalities of (34)is strict. If 2 i=1 αi(1 −xi)2< 1 −a, then we consider the signed measure defined by μ0=x2(1 −x2)δx1−x1(1 −x1)δx2. Then, it is straightforward to check that, for small enough ε >0, μ ±εμ0∈M aand μ=1 2(μ+εμ0)+1 2(μ−εμ0).(35) Therefore, μ /∈Ext(Ma). If, instead, 2 i=1 αixi(1 −xi)2<a, we then consider the signed measure μ0=(1−x2)2δx1−(1 −x1)2δx2. Again, we have that, for small enough ε >0, μ ±εμ0∈M aand equality (35)holds. We conclude that μ /∈Ext(Ma). Step 8: If μ ∈M ais supported in more than two points, then μ /∈Ext(Ma). In this case, there exist Borel disjoint sets Ai⊂(0, 1) such that μ|Ai=0, for i =1, 2, 3. Let (α1, α2, α3) ∈ R3\{(0, 0, 0)}be such that 3  i=1 αi Ai (1 −s)2dμ(s)=0 and 3  i=1 αi Ai s(1 −s)2dμ(s)=0. We consider the signed measure μ0=3 i=1 αiμ|Ai. For ε >0small enough, define μ+and μ−as μ±= μ ±εμ0. Then μ =1 2μ++1 2μ−with μ±=μ. Moreover, μ±are positive measures since A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 25 μ±=μ|Ac 1∩Ac 2∩Ac 3+ 3  i=1 (1 ±εαi)μ|Ai, where Acstands for the complement of the set Ain (0, 1). In fact, μ±∈M asince 1 0 (1 −s)2dμ±(s)= 1 0 (1 −s)2dμ(s)≤1−aand 1 0 s(1 −s)dμ±(s)= 1 0 s(1 −s)dμ(s)≤a. Therefore, we conclude that μ /∈Ext(Ma). The eight steps above complete the proof.  Step 8 of the previous proof is related to the works by Winkler [33]and Pinelis [26], where they analyze the set of extreme points of subset of measures defined through some inequalities. Proof of Theorem 1in Section 3(Extreme points of La). By Proposition 6, we have the equality Ext(La)=T−1 a(Ext(Ma)). Now, we can use Proposition 7to determine the set Ext(La). By (28)and Proposition 7, we obtain three families of extreme curves in La. First, for x1∈(0, a], let a x1=T−1 a1−a (1−x1)2δx1and a 0=T−1 a(0). More explicitly, we have that, for x1∈[0, a] a x1(x)= 1−a (1 −x1)2max{0,x−x1},x∈[0,1]. Second, for x2∈(a, 1), we set ma x2=T−1 aa x2(1−x2)δx2. We obtain that ma x2(x)= 1 x2(x2−a)x+a 1−x2 max{0,x−x2},x∈[0,1]. Finally, for x1∈(0, a)and x2∈(a, 1), let na x1,x2=T−1 ax2−a (1−x1)(x2−x1)δx1+a−x1 (x2−x1)(1−x2)δx2. In this case we have that na x1,x2(x)= 1 x2−x1x2−a 1−x1 max{0,x−x1}+a−x1 1−x2 max{0,x−x2},x∈[0,1]. These curves admit the characterization as piecewise affine functions given in (12). Therefore, the proof of Theorem 1is complete.  6.3. Proof of Theorem 2in Section 4(maximal distance) Here we compute the exact value of M(a, b)in (15). The proof of this theorem is long and we have divided it into several results. It is based on following proposition. Proposition 8. For a, b ∈[0, 1], let M(a, b)be defined in (15). We have that M(a, b)=max{dL(1, 2):1∈Ext(La)and 2∈Ext(Lb)}.(36) 32 A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 A 1−√1−a, 1+b 2!=4√1−a−1b+a(b−2) −4√1−a+b2+4 b, which is seen to be less than the value of dL(− a, + b) (computed in Lemma 3), thanks to the restriction a +b2<2b. After having checked the value of Aat the critical points in the interior of region (44), we analyze the value of Aon the boundary. Describing the boundary of region (44)is cumbersome since it involves several cases, according to the values of a, b. In any case, the boundary is clearly contained in the set (x1,x 2)∈[0,a]×[b, 1] : x1=0orx1=aor x2=bor x2=1 or (1 −a)(x2−x1)−(x2−b)(1 −x1)+x1(1 −x1)(x2−b)=0 . When (1 −a)(x2−x1) −(x2−b)(1 −x1) +x1(1 −x1)(x2−b) =0the curves a x1and mb x2do not have a proper crossing, so, by Lemma 1, A =|b −a|, which does not release a maximum. Therefore, we are led to the maximization of A(x1, x2)in the set (x1,x 2)∈[0,a]×[b, 1] : x1=0orx1=aor x2=bor x2=1  The value of Awhen x1=0is A(0,x 2)=a+b−2ab a+b−ax2 , which is decreasing in x2, so the maximum is attained at x2=band equals A(0,b)=a+b−2ab a+b−ab =(1 −a)b2+(1−b)a2 a+b−ab =dL(− a, + b). The value of Awhen x1=ais A(a, x2)=a+b−2ab ax2+b−ab which is increasing in x2, so the maximum is attained at x2=1and equals A(a, 1) = (1 −a)b2+(1−b)a2 a+b−ab =dL(− a, + b). The value of Awhen x2=bis A(x1,b)=a2b−a2+ab2−b2+(−4ab +2a+2b)x1+(a+b−2) x2 1 ab −a−b+2x1−x2 1 , which is decreasing in x1, so the maximum is attained at x1=0and equals A(0,b)=(1 −a)b2+(1−b)a2 a+b−ab =dL(− a, + b). The value of Awhen x2=1is A(x1,1) = a2b−a2+ab −b2+a+2b2−3abx1+a−b2x2 1+(b−1) x3 1 ab −a−b+2x1−x2 1 , A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 33 which is increasing in x1. Therefore, the maximum is attained at x1=aand equals A(a, 1) = (1 −a)b2+(1−b)a2 a+b−ab . This concludes the proof.  We finally observe that the proof of Theorem 2directly follows from Lemmas 3, 5, 6, and 7. Acknowledgments A. Baíllo and J. Cárcamo are supported by the Spanish MCyT grant PID2019-109387GB-I00. C. MoraCorral is supported by the Spanish MCyT grant MTM2017-85934-C3-2-P. References [1] C.D. Aliprantis, K.C. Border, Infinite Dimensional Analysis: A Hitchhiker’s Guide, Springer, 2006. [2] L.E. Alonso, C.J. Fernández, R. Ibáñez, “I think the middle class is disappearing”: crisis perceptions and consumption patterns in Spain, Int. J. Consum. Stud. 41 (2017) 389–396. [3] F. Andreoli, Robust inference for inverse stochastic dominance, J. Bus. Econ. Stat. 36 (1) (2018) 146–159. [4] B.C. Arnold, J.M. Sarabia, Majorization and the Lorenz Order with Applications in Applied Mathematics and Economics, Springer, 2018. [5] P. Blavier, How did the great recession affect income inequality in Spain?, Econ. Sociol. Eur. Electron. Newsl. 19 (1) (2017) 7–14. [6] H. Brezis, Functional Analysis, Sobolev Spaces and Partial Differential Equations, Springer, 2011. [7] J. de la Cal, J. Cárcamo, Inverse stochastic dominance, majorization, and mean order statistics, J. Appl. Probab. 47 (1) (2010) 277–292. [8] F. Cowell, Measuring Inequality, Oxford University Press, 2011. [9] D. Dentcheva, A. Ruszczyński, Inverse stochastic dominance constraints and rank dependent expected utility theory, Math. Program. 108 (2) (2006) 297–311. [10] F.G. De Maio, Income inequality measures, J. Epidemiol. Community Health 61 (2007) 849–852. [11] The Economist, Lessons from Spain’s recovery after the euro crisis, Available: www .economist .com /leaders /2018 /07 /26 / lessons -from -spains -recovery -after -the -euro -crisis, 26 Jul. 2018. [12] B. Efron, Bootstrap methods: another look at the jackknife, Ann. Stat. 7 (1979) 1–26. [13] I.I. Eliazar, I.M. Sokolov, Measuring statistical evenness: a panoramic overview, Phys. A, Stat. Mech. Appl. 391 (4) (2012) 1323–1353. [14] L.C. Evans, R.F. Gariepy, Measure Theory and Fine Properties of Functions, CRC Press, 1992. [15] T. Fawcett, An introduction to ROC analysis, Pattern Recognit. Lett. 27 (8) (2006) 861–874. [16] W.M. Getz, C.C. Wilmers, A local nearest-neighbor convex-hull construction of home ranges and utilization distributions, Ecography 27 (2004) 489–505. [17] T. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning: Data Mining, Inference and Prediction, second edition, Springer, New York, 2009. [18] International Monetary Fund, Article IV Consultation-Press Release; Staff Report; and Statement by the Executive Director for Spain, Country Report No. 17/319, 6 Oct. 2017. [19] International Monetary Fund, Article IV Consultation-Press Release; Staff Report; and Statement by the Executive Director for Spain, Country Report No. 2020/298, 13 Nov. 2020. [20] C. Kleiber, S. Kotz, Statistical Size Distributions in Economics and Actuarial Sciences, vol. 470, John Wiley & Sons, 2003. [21] M. Leshno, H. Levy, Preferred by “all” and preferred by “most” decision makers: almost stochastic dominance, Manag. Sci. 48 (8) (2002) 1074–1085. [22] C.J. Lloyd, Estimation of a convex ROC curve, Stat. Probab. Lett. 59 (1) (2002) 99–111. [23] P. Muliere, M. Scarsini, A note on stochastic dominance and inequality measures, J. Econ. Theory 49 (2) (1989) 314–323. [24] F.D. Parker, Integrals of inverse functions, Am. Math. Mon. 62 (6) (1955) 439–440. [25] R.R. Phelps, Lectures on Choquet’s Theorem, second edition, Lecture Notes in Mathematics, vol. 1757, Springer, 2001. [26] I. Pinelis, On the extreme points of moments sets, Math. Methods Oper. Res. 83 (3) (2016) 325–349. [27] L.A. Prendergast, R.G. Staudte, A simple and effective inequality measure, Am. Stat. 72 (4) (2018) 328–343. [28] M. Shaked, J.G. Shanthikumar, Stochastic Orders, Springer, 2006. [29] A.F. Shorrocks, J.E. Foster, Transfer sensitive inequality measures, Rev. Econ. Stud. 54 (3) (1987) 485–497. [30] B. Simon, Convexity, Cambridge University Press, 2011. [31] S.S. Vallender, Calculation of the Wasserstein distance between probability distributions on the line, Theory Probab. Appl. 18 (4) (1974) 784–786. 34 A. Baíllo et al. / J. Math. Anal. Appl. 514 (2022) 126335 [32] C. Villani, Optimal Transport: Old and New, Springer, 2009. [33] G. Winkler, Extreme points of moment sets, Math. Oper. Res. 13 (4) (1988) 581–587. [34] Wolfram Research, Mathematica, Inc., Version 9.0, Champaign, IL, 2012. [35] S. Yitzhaki, E. Schechtman, The Gini Methodology: A Primer on a Statistical Methodology, Springer Science & Business Media, 2012. [36] B. Zheng, Almost Lorenz dominance, Soc. Choice Welf. 51 (1) (2018) 51–63. [37] C. Zoli, Inverse stochastic dominance, inequality measurement and Gini indices, J. Econ. 77 (1) (2002) 119–161.