POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Budny, Katarzyna Article POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Budny, Katarzyna (2019) : POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, NY, Vol. 20, Iss. 3, pp. 155-170, https://doi.org/10.21307/stattrans-2019-029 This Version is available at: https://hdl.handle.net/10419/207949 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/
STATISTICS IN TRANSITION new series, September 2019 155 STATISTICS IN TRANSITION new series, September 2019 Vol. 20, No. 3, pp. 155–170, DOI 10.21307/stattrans-2019-029 Submitted – 04.03.2019; Paper ready for publication – 14.05.2019 POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE Katarzyna Budny 1 ABSTRACT In the paper some multivariate power generalizations of Chebyshev’s inequality and their improvements will be presented with extension to a random vector with singular covariance matrix. Moreover, for these generalizations, the cases of the multivariate normal and the multivariate t distributions will be considered. Additionally, some financial application will be presented. Key words: multivariate Chebyshev’s inequality, Mahalanobis distance, multivariate normal distribution, multivariate t distribution. 1. Introduction Chebyshev’s inequality yields a bound on the probability of a univariate random variable taking values close to the mean expressed by its variance. Pearson (1919) proposed its univariate power generalization presenting bounds by the central moments of a random variable of even orders. Theorem 1.1. (Pearson, 1919). If we take a random variable R: with finite central moments of 2s order s2 , then for all 0 ss s EP 22 2 . (1.1) There also exist multivariate generalizations of Chebyshev’s inequality (see, e.g. Olkin and Pratt, 1958, Marshall and Olkin, 1960, Osiewalski and Tatar, 1999). In the paper we present one of those providing upper bounds on the probability that the Mahalanobis distance of a random vector from its mean is greater or equal than the fixed value. These bounds will be given by the power transformations and will constitute the multivariate extension of (1.1). There are many applications of the Mahalanobis distance in statistical analysis. In particular, this is used in classification methods and in cluster analysis. The multivariate power generalization of Chebyshev’s inequality presented below can be exploited to detect outliers. 1 Department of Mathematics, Cracow University of Economics, Poland. E-mail: [email protected]. ORCID ID: https://orcid.org/0000-0002-3683-0327.
156 K. Budny: Power generalization of Chebyshev’s… 2. Multivariate power generalization of Chebyshev’s inequality We begin by recalling the inequality which is given by the measure of multivariate kurtosis. Theorem 2.1. (Mardia, 1970) Let n R:X be a random vector with nonsingular covariance matrix and finite fourth–order moments. Then, for any 0 the following inequality holds 2 ,2 1 X XXXX n TEEP , (2.1) where 2 1 ,2 XXXXX EEE T n is Mardia’s kurtosis of a random vector (Mardia, 1970). Chen (2007, 2011) proposed a tight upper bound (see Navarro, 2014) in the case of a random vector for which only mean and covariance matrix are known. Theorem 2.2. (Chen, 2007, Chen, 2011) Assume that n R:X is a random vector with positive covariance matrix . Then, for all 0 we get n EEP T XXXX 1 . (2.2) Budny (2014) obtained the multivariate power generalization of Chebyshev’s inequality. Theorem 2.3. (Budny, 2014) Suppose that n R:X is a random vector with nonsingular covariance matrix . Let us consider any 0s such that s T ns EEEI XXXXX 1 , exists. Then, for all 0 s ns TI EEP X XXXX , 1 . (2.3) Remark 2.1. (Budny, 2014) Observe that theorems 2.1 and 2.2 can be considered as the special cases of theorem 2.3. Taking 1s we get (2.2) and for 2s we obtain (2.1). Budny (2016), following Navarro (2016), extended (2.3) to the case of a random vector with singular covariance matrix by using the spectral decomposition.
STATISTICS IN TRANSITION new series, September 2019 157 Assume that n R:X is a random vector with covariance matrix , mrank , nm ,...,1 . Let T PP be a spectral decomposition of a covariance matrix, i.e. P is an orthogonal matrix such that n TT IPPPP and 0,...,0,,...,diag 1m is the diagonal matrix with the ordered eigenvalues 0...... 11 nmm . Hence, the Moore-Penrose generalized inverse matrix of is of the form T PCP , where 0,...,0,,...,diag 11 1 m C . Let us consider any 0s such that s T ms EEEI XXXXX , exists. Theorem 2.4. (Budny, 2016) Under the above assumptions, for any 0 , we have s ms TI EEP X XXXX , . (2.4) We will denote by S the set of all 0s such that X ms I, exists. Let us define, for fixed 0 , the function RSBd : : s ms I sBd X , , Ss . It is easily seen that for Sss 21, if 21 ss , then the following conditions are equivalent: 21 sBdsBd 12 1 2 1 , ,ss ms ms I I X X and 21 sBdsBd 12 1 2 1 , , 0ss ms ms I I X X . Summarizing, we get following remark.
158 K. Budny: Power generalization of Chebyshev’s… Remark 2.2. For Sss 21, if 21 ss , then the upper bound 2 sBd of XXXX EEP T is better than 1 sBd for all 12 1 2 1 , ,ss ms ms I I X X . On the contrary, the upper bound 1 sBd is better than 2 sBd for all 12 1 2 1 , , ,0 ss ms ms I I X X . In particular, if we consider 1 1s , 2 2s and nrank , then the upper bound 2 sBd is better than 1 sBd for all n nX ,2 . Conversely, the upper bound 1 sBd is better than 2 sBd for all n nX ,2 ,0 . 3. The case of the multivariate normal distribution Budny (2016) proposed the form of the multivariate power generalization of Chebyshev’s inequality for a normally distributed random vector for all 0\Ns . In the next theorem we extend this result to the case of any real 0s . Theorem 3.1. Let n R:X be a normally distributed random vector with mean and covariance matrix , , n N~X . Suppose that mrank , nm ,...,1 . Then, for all 0 and 0s we obtain XX T P 2 2 2 m s m s . (3.1) Proof: The proof is similar to that presented for theorem 3.1 in Budny (2016). A slight change is that we consider sth uncorrected moment (sth moment about zero) of a chi-square distribution with m degrees of freedom for any real 0s (not only for 0\Ns ). Hence, for 0s : 2 2 2 ,m s m I s ms X (Johnson, Kotz and Balakrishnan, 1994, p. 420) and it completes the proof.
STATISTICS IN TRANSITION new series, September 2019 159 Remark 3.1. For 0\Ns the inequality (3.1) takes the following form s Tsmmm P 12...2 XX (Budny, 2016). Remark 3.2. On account of remark 2.2, if 21 ss , then the upper bound 2 sBd is better than 1 sBd for all 12 1 12 22 2ss s m s m . Conversely, the upper bound 1 sBd is better than 2 sBd for all 12 1 12 22 2,0 ss s m s m . Particularly if we take 1 1s and 2 2s , then the upper bound 2 sBd is better than 1 sBd for all 2 m . Conversely, the upper bound 1 sBd is better than 2 sBd for all 2,0 m . Example 3.1. Let us consider normally distributed random vector n R:X with mean and covariance matrix , , n N~X . Assume that 3rank m . A random variable XX T has a chi-square distribution with m degrees of freedom (Kotz, Balakrishnan and Johnson, 2000, p. 110, Budny, 2016), hence we know the exact value of P . From remark 3.2 for 1 1s and 2 2s we get that the upper bound 2 sBd is better than 1 sBd for all 52 m and the upper bound 1 sBd is better than 2 sBd for all 5,0 (see Figure 3.1).
160 K. Budny: Power generalization of Chebyshev’s… Figure 3.1. The upper bounds ( 1s , 2s ) and exact value of P for ,~ n NX , rank 3 . In turn Figure 3.2 shows the upper bounds of XX 1 T P for various values of s .
STATISTICS IN TRANSITION new series, September 2019 161 Figure 3.2. The upper bounds ( 5.0s , 1s , 2s , 4s ) and exact value of P for ,~ n NX , rank 3 . 4. The case of the multivariate t distribution A n variate random vector n R:X is said to have multivariate t distribution with degrees of freedom , mean and nonsingular covariance matrix R 2 , 2 , denoted by nt ,,R , if its joint probability density function (pdf) is given by 2/ 1 2/1 2/ 1 1 2 πν 2n T nxRx R n xf n Rx . If nt ,,~ RX , then the random variable n T XRX 1 has a central F-distribution with n , degrees of freedom, ,~ nF (Lin, 1972).
162 K. Budny: Power generalization of Chebyshev’s… It follows that for any 2 s we get 22 22 n ss n n E s s (4.1) (Johnson, Kotz and Balakrishnan, 1995, p. 349). The power generalization of Chebyshev’s inequality for multivariate t distribution is established by our next theorem. Theorem 4.1. Assume that nt ,,~ RX , rank nR . Then, for any 0 the inequality 3.2 takes the following form 22 22 2 1 n ss n P s TXX (4.2) for any 0s such that 2 s . Proof: We first observe that 11 2 R . From this it is obvious that ss s ns EnI 2 ,X . (4.3) Substituting (4.1) into (4.3) yields 22 22 2 , n ss n Is ns X . (4.4) This establishes the inequality (4.2). Remark 4.1. For 0\Ns , 2 s from 4.4 we get svvv snnnv Is ns 2...42 12...22 , X . (4.5) Hence, the inequality (4.2) is of the form svvv snnn P s T 2...42 12...22 XX 1 .
STATISTICS IN TRANSITION new series, September 2019 169 6. Applications in finance Let us take a random vector t r of n assets returns on a specific day t with mean (sample mean vector of historical returns) and covariance matrix (sample covariance matrix of historical returns). Kritzman and Li (2010) propose to use the Mahalanobis distance as a measure of financial turbulence, which is understood as occurrence of unusual multivariate financial data. They defined (the so-called “the turbulence index”) turbulence for a particular time t as: t T tt rrd 1 . In the examples presented in section 3 and 4, for any 0 , we know the exact value of t T tt rrdP 1 . In the general case, this probability may not be easy to compute and if we are able to calculate the upper bounds 2.5 , then we can estimate the exact value of P. Other financial applications of the Mahalanobis distance were presented by Stöckl and Hanke (2014). Acknowledgement The publication was financed from the funds granted to the Faculty of Finance and Law at Cracow University of Economics, within the framework of the subsidy for the maintenance of research potential. REFERENCES BUDNY, K., (2014). A generalization of Chebyshev's inequality for Hilbert-spacevalued random elements. Statistics and Probability Letters, 88, pp. 62–65. BUDNY, K., (2016). An extension of the multivariate Chebyshev’s inequality to a random vector with a singular covariance matrix, Communications in Statistics – Theory and Methods, 45 (17), pp. 5220–5223. CHEN, X., (2007). A new generalization of Chebyshev inequality for random vectors. Available at: <https://arxiv.org/abs/0707.0805>[Accessed 5 July 2007]. CHEN, X., (2011). A new generalization of Chebyshev inequality for random vectors. Available at: <https://arxiv.org/abs/0707.0805v2>[Accessed 24 June 2011]. JOHNSON, N.L., KOTZ, S., BALAKRISHNAN, N., (1994). Continuous univariate distribution. Vol. 1, 2nd ed. John Wiley & Sons Inc. JOHNSON, N.L., KOTZ, S., BALAKRISHNAN, N., (1995). Continuous univariate distribution, Vol. 2, 2nd ed. John Wiley & Sons Inc.
170 K. Budny: Power generalization of Chebyshev’s… KRITZMAN, M., Li, Y., (2010). Skulls, financial turbulence, and risk management. Financial Analysts Journal, 66 (5), pp. 30–41. KOTZ, S., BALAKRISHNAN, N., JOHNSON, N.L., (2000). Continuous multivariate distribution, Vol. 1: Models and applications, 2nd ed. John Wiley & Sons Inc. LIN, P., (1972). Some characterizations of the multivariate t distribution, Journal of Multivariate Analysis, 2, pp. 339–344. LOPERFIDO, N., (2014). A probability inequality related to Mardia’s kurtosis. In: C. Perna, M. Sibillo (eds.). Mathematical and statistical methods of actuarial science and finance, Springer: Springer International Publishing Switzerland 201. pp. 129–132. MARDIA, K.V., (1970). Measures of multivariate skewness and kurtosis with applications, Biometrika, 57 (3), pp. 519–530. MARSHALL, A., OLKIN, I., (1960). Multivariate Chebyshev inequalities, The Annals of Mathematical Statistics, 31, pp. 1001–1014. NAVARRO, J., (2014). Can the bounds in the multivariate Chebyshev inequality be attained? Statistics and Probability Letters, 91, pp. 1–5. NAVARRO, J., (2016). Avery simple proof of the multivariate Chebyshev’s inequality. Communications in Statistics – Theory and Methods, 45 (12), pp. 3458–3463. OLKIN, I., PRATT, J.W., (1958). A multivariate Tchebycheff inequality. The Annals of Mathematical Statistics, 29, pp. 226–234. OSIEWALSKI, J., TATAR, J., (1999). Multivariate Chebyshev inequality based on a new definition of moments of a random vector, Przegląd Statystyczny (Stat. Rev.), 2, pp. 257–260. PEARSON, K., (1919). On generalised Tchebycheff theorems in the mathematical theory of statistics, Biometrika,12(3–4), pp. 284–296. STÖCKL, S., HANKE, M., (2014). Financial applications of the Mahalanobis distance, Applied Economics and Finance, 1 (2), pp. 78–84.