scieee AI-readable full text Open interactive document viewer

KERNEL ESTIMATION OF CUMULATIVE DISTRIBUTION FUNCTION OF A RANDOM VARIABLE WITH BOUNDED SUPPORT

Baszczyńska, Aleksandra

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Baszczyńska, Aleksandra Article KERNEL ESTIMATION OF CUMULATIVE DISTRIBUTION FUNCTION OF A RANDOM VARIABLE WITH BOUNDED SUPPORT Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Baszczyńska, Aleksandra (2016) : KERNEL ESTIMATION OF CUMULATIVE DISTRIBUTION FUNCTION OF A RANDOM VARIABLE WITH BOUNDED SUPPORT, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, NY, Vol. 17, Iss. 3, pp. 541-556, https://doi.org/10.21307/stattrans-2016-037 This Version is available at: https://hdl.handle.net/10419/207829 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ STATISTICS IN TRANSITION new series, September 2016 541 STATISTICS IN TRANSITION new series, September 2016 Vol. 17, No. 3, pp. 541–556 KERNEL ESTIMATION OF CUMULATIVE DISTRIBUTION FUNCTION OF A RANDOM VARIABLE WITH BOUNDED SUPPORT Aleksandra Baszczyńska 1 ABSTRACT In the paper methods of reducing the so-called boundary effects, which appear in the estimation of certain functional characteristics of a random variable with bounded support, are discussed. The methods of the cumulative distribution function estimation, in particular the kernel method, as well as the phenomenon of increased bias estimation in boundary region are presented. Using simulation methods, the properties of the modified kernel estimator of the distribution function are investigated and an attempt to compare the classical and the modified estimators is made. Key words: boundary effects, cumulative distribution function, kernel method, bounded support. 1. Introduction Nonparametric methods are becoming increasingly popular in statistical analysis of economic problems. In most cases, this is caused by the lack of information, especially historical data, about the economic variable being analysed. Smoothing methods concerning functions, such as density or cumulative distribution, play a special role in a nonparametric analysis of economic phenomena. Knowledge of density function or cumulative distribution function, or their estimates, allows one to characterize the random variable more completely. Estimation of functional characteristics of random variables can be carried out using kernel methods. The properties of the classical kernel methods are satisfactory, but when the support of the variable is bounded, kernel estimates may suffer from boundary effects. Therefore, the so-called boundary correction is needed in kernel estimation. 1 Department of Statistical Methods, University of Łódź. E-mail: alba[email protected]. 542 A. Baszczyńska: Kernel estimation of cumulative … Kernel estimator of cumulative distribution function has to be modified when the support of the variable is defined as   ,a ,   b, or   ba, . Such a situation is frequently observed in an economic analysis, for example, when data are considered only on the positive real line (e.g.: arable land, energy use, CO2 emission, external debts stocks, current account balance, total reserves, etc.). Near zero, the classical kernel distribution function estimator is poor because of its considerable bias. The bias comes from the behaviour of kernel estimator, which has no knowledge of the boundary and assigns probability on the negative real line. A range of boundary correction methods for kernel distribution function estimator is present in the literature. They are addressed mainly to boundary kernels (Tenreiro, 2013; Tenreiro, 2015) and reflection method (Koláček, Karunamuni, 2009; Koláček, Karunamuni, 2012). In Section 2 we introduce the kernel method, which for the first time was implemented in density estimation in the late 1950s. The properties of the kernel density estimator, as well as the modifications, are presented taking into account the boundary effects reduction of classical kernel density estimator. In Section 3 some selected methods of distribution function estimation are presented, including the kernel method. Some methods of choosing the smoothing parameter of kernel method and properties of estimator are shown, and methods of boundary correction are used in the case of cumulative distribution function estimation. In Section 4 the results of a simulation study are given and an attempt to compare the considered estimators is made. In addition, the comparison of the values of smoothing parameters is presented. The simulations and the plots were carried out using MATLAB software. The aim of the paper is to give a detailed presentation of methods of the modified kernel distribution function estimation and to compare the considered methods. The simulation shows that when boundary correction is used in kernel estimation of distribution function, the estimator has better properties. 2. Kernel method The kernel method originated from the idea of Rosenblatt and Parzen dedicated to density estimation. The Rosenblatt-Parzen kernel density estimator is as follows (cf.: Härdle, 1994; Wand, Jones, 1995; Silverman 1996; Domański et al., 2014):          n in i n nh Xx K nh xf 1 1 )( ˆ , (1) where: n XXX ,...,, 21 is the random sample from the population with unknown density function   xf ; n is the sample size; n h is the smoothing parameter, which controls the smoothness of the estimator ( 0lim   h n ,   nh n lim ). STATISTICS IN TRANSITION new series, September 2016 543 Throughout this paper the notation n hh  will be used.   uK is the weighting function called the kernel function. When   uK is symmetric and unimodal function and the following conditions are fulfilled:                              ,0 ,0 ,1 2 2  duuKu duuuK duuK (2) the kernel function is called the second order kernel function (or classical kernel function). The most frequently used Gaussian kernel function         2 2 1 exp 2 1uuK  is a function belonging to this group, although its support is unbounded. It stands in contrast to other kernel functions fulfilling conditions (2), like functions presented in Table 1, for which support is bounded. The indicator function   1uI is defined as follows:   11 uI for 1u ,   01 uI for 1u . Table 1. Kernel functions Kernel function   uK Uniform   1 2 1uI Triangle     11  uIu Epanechnikov     11 4 32 uIu Quartic     11 16 15 2 2 uIu Triweight     11 32 35 3 2 uIu Cosine   1 2 cos 4      uIu  544 A. Baszczyńska: Kernel estimation of cumulative … Higher order kernel functions (the order of the kernel is the order of the first nonzero moment) can be used, especially in reducing the mean squared error of the estimator. But the higher order kernels properties may sometimes be unacceptable because they may result in taking negative values for the density function estimators. When the support of random variable is, for example, left-bounded (support of random variable is   ,0 ), the properties of the estimator (1) may differ in boundary region   h,0 and in inner region   ,h (cf.: Jones, 1993; Jones, Foster, 1996; Li, Racine, 2007). The estimator (1) is not consistent in boundary region. As a result, the support of the kernel density estimator may differ from the support of the random variable and the estimator may be non-zero for negative values of random variable. Moreover, this situation may appear when the kernel function has unbounded as well as bounded support. Removing boundary effects can be done in various ways. The best known and most often used method is the reflection method, which is characterized by both simplicity and best properties. Assuming that the support of random variable is   ,0 , the reflection kernel density estimator, based on reflecting data about zero, has the following form (cf. Kulczycki, 2005):                    n i ii nR h Xx K h Xx K nh xf 1 1 )( ˆ . (3) The estimator (3) is a consistent estimator of unknown density function f . Moreover, it integrates to unity and for x close to zero the bias is of order )(hO . The analysis of the properties of this estimator is presented in Baszczyńska (2015), among others. 3. Distribution function estimation Let n XXX ,...,, 21 denote independent random variables with a density function f and a cumulative distribution function F . One can estimate the cumulative distribution function (CDF) by:          n i ixn XI n xF 1 , 1 ˆ , (4) where A I is the indicator function of the set A: 1)( xIA for Ax , 0)( xIA for Ax . The empirical distribution function defined by (4) is not smooth, at each point nn xXxXxX  ,...,,2211 it jumps by n 1 . STATISTICS IN TRANSITION new series, September 2016 545 The smoothed version of the empirical distribution estimator is the Nadaraya kernel estimator of CDF:               n i h Xx n i i i h Xx W n dyyK n xF 1 1 11 )( ˆ , (5) where h is a smoothing parameter such as 0lim   h n ,   nh n lim and    x dttKxW 1 )()( . Assuming that function 0)( xK is a unimodal, symmetric kernel function of the second order with support   1,1 (examples of these kernels are presented in Table 1), the properties of function )(xW are the following: 0)( xW for   1,x , 1)( xW for    ,1x ,    1 1 1 1 21)()( dxxWdxxW , (6)      1 12 1 )( dxxKxW ,             1 1 2 1 1 )(1 2 1 )( dxxWdxxKxxW . Function )(xW is a cumulative distribution function because )(xK is a probability density function. For example, when the kernel function is Epanechnikov kernel, the function )(xW has the form:              .1for1 ,1for 2 1 4 3 4 1 ,1for0 3 x xxx x xW 546 A. Baszczyńska: Kernel estimation of cumulative … Assuming additionally that )(xF is twice continuously differentiable, the mean integrated squared error (MISE) of kernel distribution estimator (5) is as follows:                             ,1 1 ˆˆ 44 21 2 h n h ohc n h cdxxFxF n dxxFxFExFMISE (7) where:    1 1 2 1)(1 dttWc , (8)          dttFc 2 2 2 2 24  , (9)   s F denotes the sth derivative of the cumulative distribution function. Kernel distribution estimator (5) is a consistent estimator of the distribution function. The expectation value, bias and variance are, respectively:             2 2 22 2 1 ˆhohxFxFxFE   ,           2 2 22 2 1 ˆhohxFxFB   ,                            n h odttWxhf n xFxF n xFD 1 1 22 )(1 1 1 1 ˆ . The method of choosing the value of the smoothing parameter in kernel estimation of the cumulative distribution function is of crucial interest, as it is in kernel estimation of the density function. Some procedures used frequently in CDF estimation are presented in Table 2. STATISTICS IN TRANSITION new series, September 2016 547 Table 2. Methods of choosing the smoothing parameter in kernel estimation of the cumulative distribution function Method Smoothing parameter Cross-validation, CV   hCVh n Hh CV  minarg ˆ ,               n i iix dxhxFXI n hCV 1 2 ,, ˆ 1 )( ,   hxF i, ˆ is a kernel estimator based on the sample with i X deleted Maximal smoothing principle, MSP 2 3 1 3 1 2 2 1ˆ 7 15 7            n c hMS Plug-in, PI 3 1 3 1 2 21 1 ˆ           n c hPI  ,               n ji ji k kg XX L gn g 1, 2 2 1 ˆ  , g is an initial smoothing parameter,   k L2 is the 2kth derivative of the initial kernel function L Iteration, IM            n ji ji jIT jijIT jIT h XX nc h h 1, ,1 , 1, 4 , ,...1,0j ,      uWWWWKWWKKu  2 , gf  denotes convolution When the random variable has bounded support (without loss of generality one can take   ,0 ), as in the case of the kernel density estimation, the properties of the kernel distribution function get poorer, in comparison with the situation when the support is unbounded. For x in boundary region   hx ,0 , let hxc / , 10  c , the expectation value and the variance of estimator (5) are the following: 548 A. Baszczyńska: Kernel estimation of cumulative …               ,)()( 2 0 )(0 ˆ 2 11 2 12 1 hodtttWdttWc c fh dttWhfxFxFE cc c B                                  .)(0 1 1 1 ˆ 1 22 hocdttWhf n xFxF n xFD c B           It is to note that in boundary region the estimator is not consistent, but variance is of the same order. The reflection kernel distribution estimator has the form (cf. Horovà et al., 2012):                    n i ii Rh Xx W h Xx W n xF 1 1 )( ˆ . (10) The generalized reflection kernel distribution estimator, improving the bias of the estimator and holding onto low variance, is the following (cf. Karunamuni, Alberts, 2005; Karunamuni, Zhang, 2008):                        n i ii GR h Xgx W h Xgx W n xF 1 21 1 )( ˆ , (11) where 1 g and 2 g are cubic polynomials with such coefficients that the bias of the estimator is of order   2 hO . In boundary region the expectation value and variance of the estimator (11) are, respectively:                             , )(00)(00)()(2 2 0 ˆ 2 1 1 2 2 2 1 1 2 12 ho dttWtcgfdttWtcgfdtttWdttWc c fh xF xFE c cc c c GR                                          hodttWdtctWtWdttWhf n xFxF n xFD ccc GR             1 2 11 22 )()2()(2)(0 1 1 1 ˆ . STATISTICS IN TRANSITION new series, September 2016 555 In the procedure of kernel estimation of cumulative distribution function, two kernel functions: Gaussian and Epanechnikov functions, influence the kernel estimator in a very similar way. The application of these kernel functions is connected with almost the same values of smoothing parameters. It can indicate that Gaussian and Epanechnikov kernels have similar smoothing properties, although they are characterized by different support. When quartic kernel is used, the smoothing parameter is smaller in comparison with other kernel functions. For samples from populations with small shape parameter, the kernel distribution estimator with smaller smoothing parameters was used. The bigger the shape parameter, the bigger the smoothing parameter in kernel estimation. In general, using Silverman’s reference rule ensures smaller values of smoothing parameter. When the shape parameter of population distribution is small, the iterative method is rather poor, the smoothing parameter is unacceptably big, which is denoted by a grey spot in Table 3. 5. Conclusion The kernel method is an intuitive, simple and useful procedure, especially in density and distribution function estimation. When the support of the random variable is bounded, this procedure needs modification. The modified kernel distribution function estimator ensures that the estimator is consistent, even in boundary region, and the support of the estimator is the same as the support of the random variable being analysed. In kernel method two parameters should be predetermined: kernel function and smoothing parameter. Quartic kernel function indicates higher values of smoothing parameter. Silverman’s reference rule, though based on the assumption that the population distribution is normal, gives smaller values of the smoothing parameter. REFERENCES BASZCZYŃSKA, A., (2015). Bias Reduction in Kernel Estimator of Density Function in Boundary Region, Quantitative Methods in Economics, XVI, 1. DOMAŃSKI, C., PEKASIEWICZ, D., BASZCZYŃSKA, A., WITASZCZYK, A., (2014). Testy statystyczne w procesie podejmowania decyzji [Statistical Tests in the Decision Making Process], Wydawnictwo Uniwersytetu Łódzkiego, Łódź. HÄRDLE, W., (1994). Applied Nonparametric Regression, Cambridge University Press, Cambridge. LI, Q., RACINE, J. S., (2007). Nonparametric Econometrics. Theory and Practice, Princeton University Press, Princeton and Oxford. 556 A. Baszczyńska: Kernel estimation of cumulative … JONES, M. C., (1993). Simple Boundary Correction for Kernel Density Estimation, Statistics and Computing, 3, pp. 135–146. JONES, M. C., FOSTER, P. J., (1996). A Simple Nonnegative Boundary Correction Method for Kernel Density Estimation, Statistica Sinica, 6, pp. 1005–1013. KARUNAMUNI, R. J., ALBERTS, T., (2005). On Boundary Correction in Kernel Density Estimation, Statistical Methodology, 2, pp. 191–212. KARUNAMUNI, R. J., ZHANG, S., (2008). Some Improvements on a Boundary Corrected Kernel Density Estimator, Statistics and Probability Letters, 78, pp. 497–507. KOLÁČEK, J., KARUNAMUNI, R. J., (2009). On Boundary Correction in Kernel Estimation of ROC Curves, Australian Journal of Statistics, 38, pp. 17–32. KOLÁČEK, J., KARUNAMUNI, R. J., (2012). A Generalized Reflection Method for Kernel Distribution and Hazard Function Estimation, Journal of Applied Probability and Statistics, 6, pp. 73–85. KULCZYCKI, P., (2005). Estymatory jądrowe w analizie systemowej [Kernel Estimators in Systems Analysis], Wydawnictwa Naukowo-Techniczne, Warszawa. HOROVÀ, I., KOLÁČEK, J., ZELINKA, J., (2012). Kernel Smoothing in MATLAB. Theory and Practice of Kernel Smoothing, World Scientific, New Jersey. SILVERMAN, B.W., (1996). Density Estimation for Statistics and Data Analysis, Chapman and Hall, London. TENREIRO, C., (2013). Boundary Kernels for Distribution Function Estimation, REVSTAT Statistical Journal, 11, 2, pp. 169–190. TENREIRO, C., (2015). A Note on Boundary Kernels for Distribution Function Estimation, http://arxiv.org/abs/1501.04206 [14.11.2015]. WAND, M. P., JONES, M.C., (1995). Kernel Smoothing, Chapman and Hall, London.