scieee AI-readable full text Open interactive document viewer

Tests in nonparametric regression based on the error distribution

Pardo Fernández, Juan Carlos

Full text

Universidade de Santiago de Compostela Departamento de Estat´ıstica e Investigaci´on Operativa Tests in Nonparametric Regression Based on the Error Distribution Juan Carlos Pardo Fern´andez Tests in Nonparametric Regression Based on the Error Distribution Juan Carlos Pardo Fern´andez Realizado o acto p´ublico de defensa e mantemento desta tese doutoral o d´ıa 1 de decembro de 2005 na Facultade de Matem´aticas da Universidade de Santiago de Compostela, perante o tribunal formado por Presidente Dr. D. Ricardo Cao Abad Vogais Dr. D. No¨el Veraverbeke Dr. D. Fr´ed´eric Ferraty Dr. D. Jacobo de U˜na ´ Alvarez Secretario Dr. D. C´esar Andr´es S´anchez Sellero, sendo os directores Dr. D. Wenceslao Gonz´alez Manteiga e Dra. Dna. Ingrid Van Keilegom, obtivo a m´axima cualificaci´on de sobresaliente cum laude. Ded´ıcolles este traballo aos meus pais e irmaus Agradecementos Quero comezar esta p´axina expresando o meu agradecemento aos profesores Wenceslao Gonz´alez Manteiga e Ingrid Van Keilegom polo labor de direcci´on deste traballo. I am very grateful to Ingrid for inviting me to visit the Institut de Statistique at Louvain-la-Neuve and accepting me as a doctorate student. I have learned a lot from both of them. Moitas grazas, danke je. Gustar´ıame tam´en darlles as grazas aos compa˜neiros do Departamento de Estat´ıstica e Investigaci´on Operativa da Universidade de Vigo que me apoiaron ao longo dos ´ultimos anos. Desexo mencionar expresamente a Jacobo e a Gustavo polo interese mostrado cara ao colectivo dos profesores visitantes. Con moitos deles ´uneme, ademais dunha relaci´on profesional, unha grande amizade: Alberto e Javier (tantas horas compartidas f´ora e dentro do despacho), Mar´ıa, Leticia, Amalia, Silvia. Agrad´ezolles a todos os bos momentos. Une grande partie de cette th`ese a ´et´e r´ealis´ee pendant plusieurs stages de recherche dans l’Institut de Statistique de l’Universit´e catholique de Louvain (Belgique). Je voudrais remercier tout le personnel de l’Institut (mes coll`egues jeunes chercheurs, professeurs et personnel administratif) pour le formidable accueil. Ils ont fait que le temps que j’ai pass´e `a Louvain-la-Neuve ait ´et´e tr`es agr´eable. Un grand merci `a Bianca, Anouar, Agust´ın et Oana pour leur amiti´e et pour m’avoir aid´e plusiuers fois avec mes probl`emes de logement. Finalmente, fago constar que este traballo foi financiado polo Ministerio de Ciencia y Tecnolog´ıa (proxecto BFM2002-03213, que incl´ue fondos FEDER), pola Secretar´ıa Xeral de Investigaci´on da Xunta de Galicia a trav´es da convocatoria de axudas para estad´ıas en centros de investigaci´on e polo Vicerreitorado de Investigaci´on da Universidade de Vigo. Tor´es, xullo de 2005 Juan Carlos Pardo Fern´andez Introduction 5 under the null hypothesis. These estimated residuals must be considered as censored data because they depend on the censored observations Zij. The comparison of the distributions of the residuals is now carried out with the Kaplan-Meier estimator (Kaplan and Meier, 1958), which is the equivalent of the empirical distribution function in the censored model. Theoretical results and simulations are also included. The problem of testing for the equality of nonparametric regression curves has been treated in the literature during the last fifteen years, mainly in more restrictive situations. Some authors, such as Young and Bowman (1995) and Dette and Neumeyer (2001), named this problem ‘Nonparametric Analysis of Covariance’. The first references about comparison of regression curves assume some restrictions on the model. Fixed design and homoscedastic errors are common assumptions in several references. H¨ardle and Marron (1990), Hall and Hart (1990), King, Hart and Wehrly (1991), and Delgado (1993) assume equal design points, equal sample sizes and homoscedastic errors. H¨ardle and Marron (1990) introduce a semiparametric method to test the hypothesis of equality of two regression curves by parameterizing the difference between them. Hall and Hart (1990) work with bootstrap methods. Delgado (1993) presents an approach based on the empirical process of the difference of the responses, which does not require smoothing at all. Young and Bowman (1995) use the classical idea of analysis of covariance in order to develop a method to compare more than two curves, assuming homoscedastic and normal distributed errors. The same authors in Bowman and Young (1996) present a graphical method to check the equality of regression curves. Kulasekera (1995) and Kulasekera and Wang (1997) propose a test applicable in the case of unequal design points, but it still requires homoscedastic errors and the impact of the choice of the smoothing parameter is important and cannot be avoided in an easy way. Munk and Dette (1998) consider the problem of testing for the equality of two curves with heteroscedastic errors and fixed design. The testing procedure is based on the estimation of the L2-distance of the difference between the two curves. Scheike (2000) establishes a test by comparing two cumulative regression functions 6 Chapter 1 and extends the ideas of Delgado (1993) to different and random designed covariates. Dette and Neumeyer (2001) describe several methods based on smoothing techniques to compare more than two curves in a nonparametric framework with fixed design. Kulasekera and Wang (2001) apply the ideas of the likelihood ratio tests to compare two regression curves. Finally, Neumeyer and Dette (2003) use an empirical process approach in a complete nonparametric, heteroscedastic and random designed setup to compare two regression curves. Their test enjoys the capability of detecting alternatives converging to the null hypothesis at a rate n−1/2, which is of interest. A different approach consists of testing the equality of two regression curves against more specified alternatives. The contributions of Hall, Huber and Speckman (1997), Koul and Schick (1997, 2003), Neumeyer and Dette (2005) and Speckman, Chin, Hewett and Bertelson (2003) deal with one-sided alternatives, which is an important particular situation. In this case the null hypothesis is the equality of the curves and the alternative says that one curve is strictly bigger than the other one. Ferreira and Stute (2004) also work with the ideas of Delgado (1993) in a time series context, and Vilar-Fern´andez and Gonz´alez-Manteiga (2004) consider dependent errors. As can easily be seen, most of these references are devoted to testing for the equality of only two regression curves. Their extensions to more than two curves are not straightforward. In practical situations the problem of testing for the equality of more than two regression curves can arise very easily. If the comparison is performed pairwise, a correction in the level of the tests must be done, and consequently the power can be affected. This motivates the implementation of general procedures for more than two curves, as has been done in this thesis. To the best of our knowledge, the problem of testing for the equality of nonparametric regression curves with censored data has not been treated in the literature. In Chapter 2 we also propose a method to test for the equality of the distributions of the errors in kregression models. Consider the same statistical setup such as in the comparison of regression curves problem. We test for the hypothesis H0ε:Fε1=Fε2=. . . =Fεk Introduction 7 versus the general alternative of any difference between these distributions. The equality of the distributions of the errors is a relatively common assumption in some references about comparison of regression curves, although it is not necessary for our method. Parametric regression models are appealing in many situations. They describe the relationship between the response variable and the covariate in a simple way and usually allow for interpretability of the parameters (for instance in linear regression). Nevertheless, if the parametric model fails then the conclusions will be erroneous. This motivates the development of specific goodness-of-fit tests to check the validity of parametric models. Given model (1.1) and a parametric family of regression functions M={mθ;θ∈Θ}, with Θ ⊂Rp, we are interested in testing for the null hypothesis H0:m∈ M versus the general alternative H1:m /∈ M. In Chapter 4 we study a goodness-of-fit test for parametric models when the response variable is subject to right censoring. The goodness-of-fit procedure consists of comparing two estimators of the distribution of the errors. Let (Xi, Zi,∆i), i= 1, . . . , n, be an i.i.d. sample of (X, Z, ∆). Assume that ˆ θis an estimator of the parameter under the null hypothesis and ˆmis a nonparametric estimator of the regression function. We compare the distribution of the residuals estimated in a parametric way Zi−mˆ θ(Xi) ˆσj(Xi) with the distribution os the residuals estimated in a completely nonparametric way Zi−ˆm(Xi) ˆσ(Xi). 8 Chapter 1 These residuals are considered as censored data and we compare the Kaplan-Meier estimators of their distributions via Kolmogorov-Smirnov and Cram´er-von Mises type statistics. Asymptotic results will be stated. The critical values of the tests are obtained by a bootstrap procedure, which is studied by means of simulations. Several goodness-of-fit tests for complete data have been proposed in the literature during the last two decades. Dette and Munk (1998), Dette (1999) and Dette, Munk and Wagner (2000) estimate the minimum L2-distance between mand the parametric family or other related quantities. H¨ardle and Mammen (1993) consider the L2-distance of the difference between the parametric and nonparametric estimators of the regression function. Stute (1997) and Stute, Gonz´alez-Manteiga and Presedo-Quidimil (1998) proposed a different approach based on the marked empirical process of the integrated regression function and its bootstrap approximation. Ramil-Novo and Gonz´alez-Manteiga (2000) studied methods to fit polynomial models. More recently, Dette and Hetzler (2004) studied tests based on processes indexed by the bandwidths. For censored data Stute, Gonz´alez-Manteiga and S´anchez-Sellero (2000) studied a test based on the marked empirical process of the integrated regression function. They assume independence between the response Yand the censoring variable Cand focus on the conditional mean. The notation is self-contained in each chapter. 1.2 Nonparametric estimation in regression In this section we briefly describe some estimation techniques in nonparametric regression. Let (X, Y ) verify the regression model Y=m(X) + σ(X)ε, where m(x) = E(Y|X=x) is the conditional mean, σ2(x) = V ar(Y|X=x) is the conditional variance and the error εis independent of X. Let (Xi, Yi), i= 1, . . . , n, be nindependent observations of (X, Y ). The objective of nonparametric Introduction 9 regression is to give an estimate of the regression curve mbased on the sample without assuming any parametric model. Smoothing techniques estimate the value of the regression curve at a certain point xby averaging locally the values of the response corresponding to the values of the covariate which are close to x. A smooth estimator of m(x) is ˆm(x) = n X i=1 Wi(x, h)Yi,(1.2) where Wi(x) is a sequence of weights depending on xand a smoothing parameter, h. The smoothing parameter plays a fundamental role in any smoothing procedure because it controls the importance of a particular response according to the proximity to xof the corresponding covariate. The choice of the weights sequence leads to different smoothing methods. In this piece of research we will work with kernel estimators. Particularly, NadarayaWatson weights, introduced simultaneously by Nadaraya (1964) and Watson (1964), will be used to estimate the regression function m Wi(x, h) = K((x−Xi)/h) Pn i0=1 K((x−Xi0)/h),(1.3) where h > 0 is the smoothing parameter or bandwidth and Kis a kernel function (normally, a symmetric density). The Nadaraya-Watson estimator of the variance function is ˆσ2(x) = n X i=1 Wi(x, h)Y2 i−ˆm2(x). As mentioned above, the bandwidth hplays a fundamental role in kernel estimation. If his too large the estimation ˆm(x) is biased and the resulting curve oversmoothed. If his too small, then the resulting estimator has large variance and it is undersmoothed. There are several mechanisms to choose the bandwidth parameter in an optimal way: cross-validation, plug-in methods, bootstrap techniques. On the other hand, the choice of the kernel function Kdoes not represent a big impact on the estimation. For further discussion about the choice of hand Kand other topics in kernel estimation see e.g. H¨ardle (1990), Wand and Jones (1994) or H¨ardle, M¨uller, Sperlich and Werwatz (2004). 10 Chapter 1 The local polynomial methods, studied in Fan and Gijbels (1996), constitute an extension of the Nadaraya-Watson estimator. Alternative choices of the weights sequence in (1.2) include Gasser and M¨uller weights and k-nearest neighbor weights (see H¨ardle, 1990). Other smoothing techniques are spline methods and orthogonal series estimators (see Hart, 1997). Akritas and Van Keilegom (2001) studied the properties of the following nonparametric estimator of the distribution of the errors ˆ Fε(y) = 1 n n X i=1 IµYi−ˆm(Xi) ˆσ(Xi)≤y¶,(1.4) which is the empirical distribution of the estimated residuals, where ˆmand ˆσare the Nadaraya-Watson estimators of mand σ. Assume now that the response variable Yis subject to right censoring. This means that there exists a random variable C, independent of Ygiven X, such that instead of Ywe observe the pair (Z, ∆), where Z= min{Y, C}and ∆ = I(Y≤C). We work with the following definitions for the regression and variance functions m(x) = Z1 0 F−1(s|x)J(s)ds, (1.5) σ2(x) = Z1 0 F−1(s|x)2J(s)ds −m2(x),(1.6) where F−1(s|x) = inf{t;F(t|x)≥s}is the conditional quantile function of Ygiven X=xand Jis a given score function satisfying R1 0J(s)ds = 1. In order to estimate the regression and the variance functions nonparametrically, we need a nonparametric estimator of the conditional distribution of Y given X. Let (Xi, Zi,∆i), i= 1, . . . , n, be a sample of independent observations of (X, Z, ∆). The nonparametric estimator of the conditional distribution with censored data was introduced by Beran (1981) ˆ F(y|x) = 1 −Y Zi≤y,∆i=1 µ1−Wi(x, h) Pn l=1 I(Zl≥Zi)Wl(x, h)¶, where Wi(x, hn) are the Nadaraya-Watson type weights given in (1.3). This estimator was studied, among others, by Dabrowska (1989), Gonz´alez-Manteiga and Introduction 11 Cadarso-Su´arez (1994) and Van Keilegom and Veraverbeke (1996, 1997). In absence of censoring ˆ F(y|x) reduces to the estimator of the conditional distribution proposed by Stone (1977). The right tails of the Beran estimator may be inconsistent to estimate the conditional distribution due to the censoring structure. This motivates the redefinition of the regression and variance functions to include the score function J. From equations (1.5) and (1.6), it is clear that once we have an estimator of the conditional distribution function, we can immediately construct an estimator of the regression function. Nonparametric estimators of the regression and variance functions are given by ˆm(x) = Z1 0 ˆ F−1(s|x)J(s)ds, and ˆσ2(x) = Z1 0 ˆ F−1(s|x)2J(s)ds −ˆm2(x). The residuals of the regression model can now be estimated by ˆ Ei=Zi−ˆm(Xi) ˆσ(Xi), considered as censored quantities since they are measured with respect to the observable variables Ziand not with respect to the true (and not observable) responses Yi. Van Keilegom and Akritas (1999) proposed estimating the distribution of the errors by the Kaplan-Meier estimator of the censored sample ( ˆ Ei,∆i), i= 1, . . . , n, ˆ Fε(y) = 1 −Y ˆ Ei≤y,∆i=1 Ã1−1 Pn l=1 I(ˆ El≥ˆ Ei)!. Obviously, this estimator of the distribution of the errors reduces to the one given in (1.4) in absence of censoring. Chapter 2 Testing for the equality of k regression curves In this chapter we introduce a procedure to test the hypothesis of equality of the kregression functions in a completely nonparametric framework. The test is based on the comparison of two estimators of the distribution of the errors in each population. Kolmogorov-Smirnov and Cram´er-von Mises type statistics are considered and their asymptotic distributions are obtained. The proposed tests can detect local alternatives converging to the null hypothesis at the rate n−1/2. We describe a bootstrap procedure in order to approximate the critical values, and present the results of a simulation study, in which the behavior of the tests for small and moderate sample sizes is studied. Finally, we include an application to a real data set. As a by-product, a method to test the equality of the distribution of the errors in several regression models is obtained and developed theoretically and from a practical point of view. This chapter is mainly based on Pardo-Fern´andez, Van Keilegom and Gonz´alezManteiga (2005a). 13 14 Chapter 2 2.1 Introduction. Statistical model The comparison of two or more groups is an important problem in statistical inference. This comparison can be performed through the regression curves in a nonparametric context. Let (Xj, Yj) be kindependent random vectors, and assume that they satisfy the following nonparametric regression models, for j= 1, . . . , k, Yj=mj(Xj) + σj(Xj)εj,(2.1) where the error variable εj, with distribution Fεj, is independent of Xj, mj(Xj) = E(Yj|Xj) is the unknown regression function and σ2 j(Xj) = V ar(Yj|Xj) is the conditional variance function. By construction E(εj) = 0 and V ar(εj) = 1. Suppose that the covariates Xjhave common support RX. Let (Xij, Yij), i= 1, ..., nj, be an i.i.d. sample from the distribution of (Xj, Yj), for j= 1, . . . , k, and denote n=Pk j=1 nj. We are interested in testing the null hypothesis of equality of the regression functions H0:m1=m2=··· =mk(2.2) versus the general alternative Ha:mi6=mjfor some i, j ∈ {1,2, . . . , k}. The idea of the testing procedure proposed in this thesis is to compare, in each population, the empirical distribution functions of the residuals with the same distribution function estimated under the null hypothesis (2.2). More precisely, let us fix one population, say j. Let Yij −ˆmj(Xij) ˆσj(Xij) Testing for the equality of kregression curves 21 2.3.1 Asymptotic results under the null hypothesis In this section we state the asymptotic results related to the testing procedure under the null hypothesis. Theorem 2.2 gives a representation for the difference of the two estimators of the error distribution in each population ˆ Fεj0(y)−ˆ Fεj(y) as a sum of i.i.d. random variables plus a negligible term, which will be used in Theorem 2.3 to state the weak convergence of the k-dimensional process ˆ W(y). Finally, the asymptotic distributions of the test statistics are obtained in Corollary 2.1. Theorem 2.2 Assume (A1)-(A4). Then, under the null hypothesis H0, for j= 1, . . . , k ˆ Fεj0(y)−ˆ Fεj(y) =fεj(y) k X l=1 pl(1 nl nl X i=1 Yil −m(Xil) σj(Xil)µfXj(Xil) fmix(Xil)−I(l=j) pj¶)+oP(n−1/2), uniformly in y. Theorem 2.3 Assume (A1)-(A4). Then, under the null hypothesis H0, the kdimensional process ˆ W(y)=(ˆ W1(y), . . . , ˆ Wk(y))tconverges weakly to W(y) = (fε1(y)W1, . . . , fεk(y)Wk)t, where, W1, . . . , Wkare normal random variables with mean zero and covariance structure Cov(Wj, Wj0) = p1/2 jp1/2 j0× × k X l=1 plE·σ2 l(Xl) σj(Xl)σj0(Xl)µfXj(Xl) fmix(Xl)−I(l=j) pj¶µfXj0(Xl) fmix(Xl)−I(l=j0) pj0¶¸. Corollary 2.1 Assume (A1)-(A4). Then, under the null hypothesis H0, T1 KS d → k X j=1 |Wj|sup y|fεj(y)|, T1 CM d → k X j=1 W2 jZf2 εj(y)dFεj(y), T2 KS d →sup y|W(y)|, T2 CM d →ZW2(y)dFε(y), 22 Chapter 2 where W(y) = Pk j=1 p1/2 jfεj(y)Wjand Fε(y) = Pk j=1 pjFεj(y). 2.3.2 Asymptotic results under local alternatives Consider now the limiting behavior of the test statistics under the following local alternatives: Hl.a. :mj=m0+n−1/2rj, where the functions rjsatisfy (A5) (i) rjis twice continuously differentiable, for j= 1, . . . , k. (ii) V ar[rj(Xl)] <∞, for j= 1, . . . , k and l= 1, . . . , k. In Theorem 2.4 we state the weak convergence of ˆ W(y) under the alternative hypothesis Hl.a. and in Corollary 2.2 we give the corresponding asymptotic distributions of the test statistics. Theorem 2.4 Assume (A1)-(A5). Then, under the alternative hypothesis Hl.a., the k-dimensional process ˆ W(y) = ( ˆ W1(y), . . . , ˆ Wk(y))tconverges weakly to W(y)+ D(y), where W(y)is defined in Theorem 2.3 and D(y) = (p1/2 1fε1(y)d1, . . . , p1/2 kfεk(y)dk)t, with dj=E·R(Xj)−rj(Xj) σj(Xj)¸, and R(u) = Pk j=1 pj fXj(u) fmix(u)rj(u). Corollary 2.2 Assume (A1)-(A5). Then, under the alternative hypothesis Hl.a., T1 KS d → k X j=1 |Wj+p1/2 jdj|sup y|fεj(y)|, T1 CM d → k X j=1 (Wj+p1/2 jdj)2Zf2 εj(y)dFεj(y), T2 KS d →sup y|W(y) + d(y)|, T2 CM d →Z(W(y) + d(y))2dFε(y), Testing for the equality of kregression curves 23 where d(y) = Pk j=1 pjfεj(y)dj, the random variables Wjare defined in the statement of Theorem 2.3, and W(y)and Fε(y)are defined in the statement of Corollary 2.1. We can analyze in detail the effect of the local alternatives if we consider the simpler situation of two regression curves where one of the curves is fixed and the other one varies with n. The null hypothesis is H0:m1=m2and the alternative Hl.a. :m2=m1+n−1/2r. In this situation d1=p2E·fX2(X1) fmix(X1) r(X1) σ1(X1)¸ and d2=−p1E·fX1(X2) fmix(X2) r(X2) σ2(X2)¸, and these values may be zero in some cases. Nevertheless, there are important situations with consistency against alternatives converging to the null hypothesis at a rate n−1/2, such as the one-sided alternatives (when ris a positive function). And, of course, the testing procedure is universally consistent in the sense of Theorem 2.1. 2.4 Bootstrap approximation To apply this testing procedure in practice the asymptotic distribution of the test statistics can be used so as to obtain the critical values of the test. These asymptotic distributions, given in Corollary 2.1, can be estimated by plugging in estimators for pj, m, σj, Fεj,fεj,fXjand fmix. Alternatively, one can use a bootstrap procedure in order to approximate the distributions of the test statistics under the null hypothesis. We will now consider this second option in detail. First, for j= 1, . . . , k and i= 1, ..., nj, estimate the residuals in a nonparametric way, using each sample separately, that is Yij −ˆmj(Xij) ˆσj(Xij). 24 Chapter 2 These residuals are then standardized in order to have mean zero and variance one. Let ˜ Fεjbe the empirical distribution of these standardized residuals obtained from each sample. We propose a smooth bootstrap of the residuals. Note that the asymptotic representation given in Theorem 2.2 involves the density of the errors fεj. This suggests that a smoothed version of the bootstrap of the residuals must be used. In the bootstrap of the residuals the samples are drawn from the empirical distribution, while in the smooth bootstrap the resamples are drawn from an estimation of the corresponding density. See Freedman (1981) for the bootstrap of the residuals and, e.g., Davison and Hinkley (1997) or Silverman and Young (1987) for the smoothing in the bootstrap. The bootstrap procedure can be described in the following steps. For fixed B and for b= 1, ..., B, 1. For j= 1, . . . , k, let {ε∗ ij,b, i = 1, . . . , nj}be an i.i.d. sample from the distribution of (1 −a2 j)1/2Vj+ajZ, where Vjhas distribution ˜ Fεjand Zis, e.g., a standard normal random variable. The constants aj, which determine the amount of smoothing in the bootstrap, are related to the sample size in each sample. 2. For j= 1, . . . , k, define new responses under the null hypothesis Y∗ ij,b by (i= 1, . . . , nj) Y∗ ij,b = ˆm(Xij) + ˆσj(Xij)ε∗ ij,b. 3. Let T1∗ KS,b,T1∗ CM,b,T2∗ KS,b and T2∗ CM,b be the test statistics obtained from the bootstrap samples {(Xij, Y ∗ ij,b), i = 1, . . . , nj},j= 1, . . . , k. Since in step 2 the bootstrap resamples are constructed under the null hypothesis of equal regression functions, this mechanism approximates the distribution of the test statistics under the null hypothesis. If we denote T1∗ KS,(b)for the order statistics of the values T1∗ KS,1, . . . , T1∗ KS,B obtained in step 3, and analogously for T1∗ CM,(b),T2∗ KS,(b)and T2∗ CM,(b), then T1∗ KS,([(1−α)B]),T1∗ CM,([(1−α)B]),T2∗ KS,([(1−α)B]) and T2∗ CM,([(1−α)B]) approximate the (1 −α)-quantiles of the distribution of T1 KS,T1 CM , T2 KS and T2 CM under the null hypothesis respectively. Testing for the equality of kregression curves 25 2.5 Simulation study This section is devoted to studying the practical behavior of the bootstrap procedure by means of simulations. In the first part we restrict our study to the comparison of two regression curves in order to be able to compare our method with others in the literature. Particularly, we will compare our method with the procedure developed by Neumeyer and Dette (2003), which is based on a marked empirical process approach. The comparison is carried out for a selection of the models considered in the simulation section of Neumeyer and Dette’s paper (except for model (ii), which is not considered in their paper). We consider the following models: (i)m1(x) = 1; m2(x) = 1 (ii)m1(x) = x;m2(x) = x (iii)m1(x) = sin(2πx); m2(x) = sin(2πx) (iv)m1(x) = exp(x); m2(x) = exp(x) (v)m1(x) = 1; m2(x) = 1 + x (vi)m1(x) = exp(x); m2(x) = exp(x) + x (vii)m1(x) = sin(2πx); m2(x) = sin(2πx) + x (viii)m1(x) = 1; m2(x) = 1 + sin(2πx) Clearly, models (i)−(iv) correspond to the null hypothesis and models (v)− (viii) to the alternative hypothesis. In each case, we consider a homoscedastic and a heteroscedastic situation. In the homoscedastic case the variance functions are (as in Neumeyer and Dette, 2003) σ2 1(x) = 0.25 and σ2 2(x) = 0.50,(2.10) and in the heteroscedastic case the variance functions are σ2 1(x) = σ2 2(x) = ex R1 0etdt.(2.11) The distribution of ε1and ε2is the standard normal distribution. Other simulations have been carried out with other distributions for the errors and similar 26 Chapter 2 results were obtained. In all cases the covariates X1and X2have uniform distribution on [0,1]. For the kernel function needed in the estimation of the regression and variance curves, we choose the Epanechnikov kernel K(u) = 0.75(1 −u2)I(|u|<1). In the theoretical results we work with only one bandwidth hnand we have found in simulations a better approximation of the level when the same bandwidth is used to estimate the common regression curve and the regression curves in each population, especially for ‘oscillating’ functions, as in model (iii). This can be explained as follows: when using different bandwidths the estimation of the regression curve in one population can be oversmoothed with respect to the estimation of the common regression function, and then the test could detect different curves when they are really the same. More precisely, we consider a bandwidth of the form h=cn−3/10, which verifies the regularity conditions given in Section 2.2. Cases c= 0.5 and c= 1 are presented. In other situations the value of cmust be adapted to the support of the regressor variables. Concerning the amount of smoothing we apply in the bootstrap, we recommend different constants ajdepending on the sample sizes nj. We work in all cases with aj= 2n−3/10 j. When it is considered as a bandwidth, ajchosen in this way is a small bandwidth to estimate the density of the errors (a standard normal in our simulations). Tables 2.1-2.4 register the proportion of rejections in 1000 trials for sample sizes (n1, n2) = (50,50),(100,50) and (100,100) and B= 200 bootstrap replications. The significance levels are α= 0.05 and α= 0.10. Tables 2.1 and 2.3 show that the level is well approximated in most cases and for both bandwidth choices. The approximation is better for the tests based on the statistics T1 KS and T1 CM . The tests based on T2 KS and T2 CM seem to be somewhat conservative. The behavior of the power (Tables 2.2 and 2.4) of the tests based on T1 KS and T1 CM is good for models (v), (vi) and (vii). On the other hand, the tests based on T2 KS and T2 CM give good power for model (viii). In both cases the obtained power is better for the larger sample sizes. The Cram´er-von Mises test gives better power than the Kolmogorov-Smirnov test in most situations. Also note Testing for the equality of kregression curves 27 that the choice of the bandwidth has little impact on the rejection probabilities. In Neumeyer and Dette (2003) only the homoscedastic case is considered. For a comparison with their simulations we found that our procedure based on T1 KS and T1 CM yields better or similar results for the power in most cases for models (v), (vi) and (vii), whereas for model (viii) we obtained better results with the tests based on T2 KS and T2 CM . Our method is valid for more than two curves. We explore now the behavior of the testing procedure in a three regression curves setup. We consider the following models (ix)m1(x) = x;m2(x) = x;m3(x) = x (x)m1(x) = x;m2(x) = x+ 0.25; m3(x) = x+ 0.5 (xi)m1(x) = x;m2(x) = 0.5; m3(x) = 1 −x Model (ix) corresponds to the null hypothesis and models (x) and (xi) correspond to the alternative. The variance functions are σ2 1(x) = σ2 2(x) = σ2 3(x) = 0.5.(2.12) As in the previous simulated models, the covariates are uniformly distributed in [0,1] and the errors are distributed as a standard normal. The choice of the kernel, the bandwidth and the amount of smoothing in the bootstrap is the same as in the previous simulations: Kis the Epanechnikov kernel, h=cn−3/10 (cases c= 0.5 and c= 1 are displayed) and aj= 2n−3/10 j. The obtained results for both versions of the test statistics, with samples sizes (50,50,50), (100,50,50), (100,100,50) and (100,100,100) and significance levels α= 0.05 and α= 0.10 are shown in Table 2.5. The tests based on T1 KS and T1 CM approximate the level well, while the tests based on T2 KS and T2 CM seem to be a bit conservative, as happened in the two-curve examples. In all cases the power increases with the sample sizes. The first version of the test statistics produces better power in model (x) and the second version gives better power in model (xi). 28 Chapter 2 Table 2.1: Rejection probabilities under the null hypothesis –models (i) to (iv)– of the tests based on T1 KS and T1 CM . The models are homoscedastic, with variances given in (2.10), and heteroscedastic, with variances given in (2.11). c= 0.5c= 1 T1 KS T1 CM T1 KS T1 CM (n1, n2)α: 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 Homoscedastic models (50,50) (i) 0.048 0.095 0.050 0.108 0.051 0.093 0.052 0.105 (ii) 0.053 0.101 0.053 0.106 0.051 0.101 0.054 0.102 (iii) 0.066 0.102 0.058 0.107 0.064 0.110 0.071 0.122 (iv) 0.058 0.110 0.057 0.105 0.057 0.102 0.055 0.108 (100,50) (i) 0.063 0.111 0.057 0.103 0.055 0.099 0.055 0.107 (ii) 0.055 0.103 0.061 0.108 0.048 0.100 0.050 0.106 (iii) 0.060 0.111 0.065 0.118 0.067 0.117 0.076 0.127 (iv) 0.056 0.100 0.062 0.110 0.048 0.097 0.050 0.111 (100,100) (i) 0.054 0.100 0.053 0.100 0.052 0.088 0.050 0.106 (ii) 0.056 0.097 0.051 0.099 0.054 0.088 0.053 0.107 (iii) 0.049 0.094 0.050 0.102 0.058 0.110 0.061 0.110 (iv) 0.055 0.096 0.053 0.099 0.052 0.095 0.059 0.104 Heteroscedastic models (50,50) (i) 0.050 0.090 0.050 0.098 0.052 0.097 0.050 0.095 (ii) 0.053 0.095 0.052 0.100 0.047 0.095 0.053 0.103 (iii) 0.055 0.106 0.060 0.105 0.049 0.096 0.052 0.103 (iv) 0.054 0.098 0.053 0.100 0.042 0.090 0.054 0.094 (100,50) (i) 0.059 0.102 0.054 0.104 0.054 0.099 0.052 0.114 (ii) 0.050 0.105 0.053 0.103 0.048 0.096 0.053 0.102 (iii) 0.058 0.109 0.060 0.099 0.060 0.116 0.063 0.121 (iv) 0.049 0.102 0.055 0.099 0.050 0.099 0.053 0.102 (100,100) (i) 0.055 0.097 0.049 0.102 0.050 0.096 0.056 0.099 (ii) 0.057 0.098 0.049 0.101 0.052 0.097 0.056 0.100 (iii) 0.048 0.094 0.046 0.098 0.054 0.098 0.056 0.107 (iv) 0.055 0.096 0.048 0.102 0.052 0.093 0.056 0.104 Testing for the equality of kregression curves 29 Table 2.2: Rejection probabilities under the alternative hypothesis –models (v) to (viii)– of the tests based on T1 KS and T1 CM . The models are homoscedastic, with variances given in (2.10), and heteroscedastic, with variances given in (2.11). c= 0.5c= 1 T1 KS T1 CM T1 KS T1 CM (n1, n2)α: 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 Homoscedastic models (50,50) (v) 0.939 0.972 0.960 0.986 0.948 0.978 0.972 0.986 (vi) 0.943 0.970 0.965 0.983 0.950 0.973 0.969 0.986 (vii) 0.950 0.966 0.966 0.977 0.945 0.972 0.963 0.977 (viii) 0.245 0.431 0.204 0.398 0.213 0.370 0.158 0.312 (100,50) (v) 0.983 0.995 0.992 0.997 0.983 0.994 0.992 0.998 (vi) 0.988 0.994 0.993 0.998 0.983 0.990 0.991 0.995 (vii) 0.986 0.993 0.990 0.998 0.977 0.986 0.984 0.994 (viii) 0.315 0.480 0.239 0.431 0.286 0.436 0.180 0.342 (100,100) (v) 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 (vi) 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 (vii) 1.000 1.000 1.000 1.000 0.999 1.000 1.000 1.000 (viii) 0.488 0.702 0.428 0.688 0.430 0.647 0.324 0.562 Heteroscedastic models (50,50) (v) 0.537 0.673 0.613 0.725 0.596 0.717 0.638 0.762 (vi) 0.543 0.670 0.613 0.727 0.583 0.712 0.643 0.753 (vii) 0.564 0.666 0.612 0.719 0.586 0.701 0.626 0.748 (viii) 0.118 0.201 0.100 0.185 0.122 0.199 0.086 0.167 (100,50) (v) 0.725 0.821 0.791 0.863 0.760 0.854 0.816 0.886 (vi) 0.728 0.817 0.796 0.867 0.740 0.847 0.815 0.881 (vii) 0.727 0.820 0.793 0.862 0.749 0.844 0.810 0.870 (viii) 0.164 0.269 0.133 0.248 0.178 0.302 0.136 0.240 (100,100) (v) 0.898 0.945 0.917 0.954 0.899 0.952 0.928 0.962 (vi) 0.890 0.943 0.915 0.954 0.912 0.953 0.929 0.962 (vii) 0.888 0.939 0.920 0.953 0.890 0.951 0.920 0.958 (viii) 0.208 0.331 0.151 0.282 0.213 0.341 0.136 0.261 30 Chapter 2 Table 2.3: Rejection probabilities under the null hypothesis –models (i) to (iv)– of the tests based on T2 KS and T2 CM . The models are homoscedastic, with variances given in (2.10), and heteroscedastic, with variances given in (2.11). c= 0.5c= 1 T2 KS T2 CM T2 KS T2 CM (n1, n2)α: 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 Homoscedastic models (50,50) (i) 0.038 0.080 0.041 0.091 0.056 0.104 0.049 0.101 (ii) 0.047 0.084 0.045 0.101 0.049 0.095 0.055 0.100 (iii) 0.052 0.097 0.049 0.098 0.047 0.097 0.066 0.121 (iv) 0.052 0.091 0.050 0.100 0.048 0.092 0.055 0.096 (100,50) (i) 0.046 0.092 0.045 0.081 0.048 0.088 0.050 0.105 (ii) 0.043 0.080 0.048 0.078 0.054 0.093 0.055 0.101 (iii) 0.046 0.096 0.048 0.101 0.067 0.130 0.073 0.130 (iv) 0.045 0.092 0.042 0.084 0.059 0.099 0.055 0.110 (100,100) (i) 0.030 0.081 0.034 0.062 0.039 0.090 0.044 0.087 (ii) 0.033 0.073 0.033 0.064 0.044 0.077 0.044 0.092 (iii) 0.044 0.089 0.046 0.080 0.045 0.091 0.060 0.104 (iv) 0.039 0.070 0.032 0.065 0.044 0.086 0.049 0.095 Heteroscedastic models (50,50) (i) 0.046 0.092 0.043 0.087 0.044 0.095 0.057 0.094 (ii) 0.047 0.084 0.043 0.089 0.043 0.095 0.051 0.088 (iii) 0.050 0.087 0.045 0.080 0.044 0.088 0.047 0.087 (iv) 0.040 0.082 0.044 0.091 0.038 0.086 0.046 0.089 (100,50) (i) 0.040 0.073 0.043 0.075 0.042 0.085 0.049 0.087 (ii) 0.036 0.075 0.039 0.076 0.032 0.083 0.038 0.085 (iii) 0.044 0.080 0.036 0.083 0.055 0.101 0.041 0.090 (iv) 0.036 0.080 0.039 0.079 0.039 0.079 0.046 0.082 (100,100) (i) 0.041 0.079 0.034 0.057 0.039 0.067 0.043 0.076 (ii) 0.042 0.072 0.030 0.060 0.034 0.079 0.039 0.073 (iii) 0.038 0.075 0.033 0.063 0.038 0.093 0.051 0.101 (iv) 0.035 0.066 0.032 0.063 0.038 0.083 0.044 0.074 Testing for the equality of kregression curves 37 We consider the following test statistics: a Kolmogorov-Smirnov type statistic SKS = k X j=1 sup y|ˆ Uj(y)|, and a Cram´er-von Mises type statistic SCM = k X j=1 Zˆ U2 j(y)dˆ Fε(y). The null hypothesis (2.13) is rejected when the obtained value of the test statistic SKS or SCM is larger than a certain critical value. 2.7.1 Asymptotic results In the following theorem and corollary we state the weak convergence of the process ˆ U(y) and give the asymptotic distributions of the test statistics SKS and SCM . The regularity assumptions are basically the same we listed in the beginning of Section 2.3, except that now the covariates do not need to have common support (assumption A1-i). The proofs are included in Section 2.8. Theorem 2.5 Assume (A1)-(A4). Then, under the null hypothesis H0ε, the kdimensional process ˆ U(y) = ( ˆ U1(y), . . . , ˆ Uk(y))tconverges weakly to a centered k-dimensional Gaussian process U(y) = (U1(y), . . . , Uk(y))twith covariance structure given by Cov(Uj(y), Uj0(y0)) = k X l=1 p1/2 jp1/2 j0µ1−I(l=j) pj¶µ1−I(l=j0) pj0¶E(ϕl(Xl, Yl, y)ϕl(Xl, Yl, y0)), where, for j= 1, . . . , k, ϕj(u, v, y) =Iµv−mj(u) σj(u)≤y¶−Fε(y) +yfε(y)(v−mj(u))2−σ2 j(u) 2σ2 j(u)+fε(y)v−mj(u) σj(u). 38 Chapter 2 Corollary 2.3 Assume (A1)-(A4). Then, under the null hypothesis H0ε, SKS d → k X j=1 sup |Uj(y)|, SCM d → k X j=1 ZU2 j(y)dFε(y). 2.7.2 Bootstrap, simulations and data analysis Bootstrap approximation. In practical applications, the critical values of the test statistics SKS and SCM can be approximated by the bootstrap procedure described below. Let ˜ Fεbe the standardized version of the empirical distribution function of the joint sample of estimated residuals considered in (2.14). For fixed Band for b= 1, ..., B, 1. For j= 1, . . . , k, let {ε∗ ij,b, i = 1, . . . , nj}be an i.i.d. sample from the distribution of (1 −a2 j)1/2V+ajZ, where Vhas distribution ˜ Fεand Zis, e.g., a standard normal random variable. 2. For j= 1, . . . , k, and i= 1, . . . , nj, define new responses under the null hypothesis Y∗ ij,b by Y∗ ij,b = ˆmj(Xij) + ˆσj(Xij)ε∗ ij,b 3. Let S∗ KS,b and S∗ CM,b be the test statistics obtained from the bootstrap samples {(Xij, Y ∗ ij,b), i = 1, . . . , nj},j= 1, . . . , k. In step 2 the bootstrap resamples are constructed under the null hypothesis of equal distribution for the errors since we draw residuals from the combined sample of estimated residuals in all populations. This mechanism approximates the distribution of the test statistics under the null hypothesis. If we denote S∗ KS,(b)(respectively, S∗ CM,(b)) for the order statistics of the values S∗ KS,1, . . . , S∗ KS,B obtained in step 3 (respectively, S∗ CM,1, . . . , S∗ CM,B) then S∗ KS,([(1−α)B]) (respectively, S∗ CM,([(1−α)B])) approximates the (1 −α)-quantiles of the distribution of SKS (respectively, SCM ) under the null hypothesis. Testing for the equality of kregression curves 39 Table 2.6: Rejection probabilities under models (i) to (iii) of the tests based on SKS and SCM . c= 0.5c= 1 SKS SCM SKS SCM (n1, n2)α: 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 (50,50) (i) 0.061 0.092 0.067 0.115 0.050 0.113 0.053 0.123 (ii) 0.613 0.727 0.775 0.841 0.627 0.747 0.794 0.859 (iii) 0.088 0.165 0.143 0.239 0.106 0.183 0.179 0.290 (100,50) (i) 0.059 0.105 0.060 0.101 0.057 0.100 0.052 0.102 (ii) 0.721 0.812 0.837 0.906 0.733 0.832 0.870 0.920 (iii) 0.146 0.231 0.218 0.321 0.147 0.242 0.230 0.350 (100,100) (i) 0.054 0.096 0.062 0.107 0.061 0.103 0.056 0.107 (ii) 0.919 0.961 0.977 0.989 0.931 0.964 0.983 0.993 (iii) 0.225 0.353 0.377 0.495 0.265 0.401 0.446 0.579 Simulations. We show a simulated example with two populations. The models for the distributions of the errors are (i)ε1, ε2∼N(0,1) (ii)ε1∼N(0,1), ε2∼Exponential(1) −1 (iii)ε1∼N(0,1), ε2∼Uniform[−1/√12,1/√12] Model (i) corresponds to the null hypothesis, and models (ii) and (iii) correspond to the alternative hypothesis. In all cases the covariates follow a uniform distribution on [0,1]. The responses are given by the regression functions m1(x) = m2(x) = xand the variance functions σ1(x) = σ2(x) = 0.50. In this case we estimate every needed curve to construct the residuals form each sample separately. From a practical point of view, we recommend to choose different bandwidths according to the sample sizes in each population. More precisely, we consider bandwidths hj=cn−3/10 jto estimate mjand σjin each population. Cases c= 0.5 and c= 1 are displayed. The amount of smoothing in the bootstrap was chosen to be aj= 2n−3/10 j. We work with the Epanechnikov kernel K(u) = 40 Chapter 2 0.75(1 −u2)I(|u|<1). Table 2.6 shows the proportion of rejections in 1000 trials for sample sizes (n1, n2) = (50,50),(100,50) and (100,100) and B= 200 bootstrap replications. The significance levels are α= 0.05 and α= 0.10. The level is well approximated for both bandwidth choices and the power increases with the sample sizes. The Cram´er-von Mises test gives better power than the Kolmogorov-Smirnov test in most situations. As in the comparison of regression curves problem, the choice of the bandwidth has little impact on the rejection probabilities. Real data analysis. We have also applied this testing procedure to the data described in Section 2.6. We have tested the null hypothesis of equal distribution of the residuals in the three populations. The results confirm this hypothesis. The test was carried out for different values of cranging from 0.5 to 1.5. The p-values were calculated from 1000 bootstrap replications. For the KolmogorovSmirnov type statistic SKS the p-values were between 0.33 and 0.65 and for the Cram´er-von Mises type statistic SCM the p-values were between 0.21 and 0.36. 2.8 Proofs Proof of Theorem 2.1. 1st part. Assume Fεj0(y) = Fεj(y). This implies that the first and the second moment of these distributions are the same. From the first moment we have that EµYj−m(Xj) σj(Xj)¶=EµYj−mj(Xj) σj(Xj)¶= 0. The second moment of Fεj0(y) can be written as V ar µYj−m(Xj) σj(Xj)¶ =E"µYj−m(Xj) σj(Xj)¶2#=E"µYj−mj(Xj) + mj(Xj)−m(Xj) σj(Xj)¶2# Testing for the equality of kregression curves 41 =V ar µYj−mj(Xj) σj(Xj)¶+ 2Eµ(Yj−mj(Xj))(mj(Xj)−m(Xj)) σ2 j(Xj)¶ +Eµ(mj(Xj)−m(Xj))2 σ2 j(Xj)¶. We are assuming that V ar µYj−m(Xj) σj(Xj)¶=V ar µYj−mj(Xj) σj(Xj)¶, and clearly 2Eµ(Yj−mj(Xj))(mj(Xj)−m(Xj)) σ2 j(Xj)¶= 0. Hence Eµ(mj(Xj)−m(Xj))2 σ2 j(Xj)¶= 0, and this implies mj(x) = m(x), for all j= 1, . . . , k, and for all x∈RX, except for a set of points of probability zero. The continuity of the functions mjallows extending the equality to all x∈RX. 2nd part. Assume Fε0(y) = Fε(y). As in the previous part, if we compute the first moment we obtain k X j=1 pjEµYj−m(Xj) σj(Xj)¶= k X j=1 pjEµYj−mj(Xj) σj(Xj)¶= 0. And from the second moment V ar Ãk X j=1 pj Yj−m(Xj) σj(Xj)!= k X j=1 p2 jE"µYj−m(Xj) σj(Xj)¶2# = k X j=1 p2 jV ar µYj−mj(Xj) σj(Xj)¶+ k X j=1 p2 jEµ(mj(Xj)−m(Xj))2 σ2 j(Xj)¶ =V ar Ãk X j=1 pj Yj−mj(Xj) σj(Xj)!+ k X j=1 p2 jEµ(mj(Xj)−m(Xj))2 σ2 j(Xj)¶. Then it holds that k X j=1 p2 jEµ(mj(Xj)−m(Xj))2 σ2 j(Xj)¶= 0. 42 Chapter 2 This and the continuity of the functions mjimplies mj(x) = m(x) for all j= 1, . . . , k and for all x∈RX. The converse implications are trivial. Before proving the main results in Sections 2.3 and 2.7, we state the following three auxiliary lemmas. Lemma 2.1 Assume (A1)-(A3). Then, under the null hypothesis H0, for any j= 1, . . . , k, Zˆm(x)−m(x) σj(x)fXj(x)dx =1 n k X l=1 nl X i=1 Yil −m(Xil) σj(Xil) fXj(Xil) fmix(Xil)+oP(n−1/2). Proof. Consider the kernel estimator of the density of the mixture fmix(x) ˆ fmix(x) = 1 nhn k X l=1 nl X i=1 Kµx−Xil hn¶. It is well known that (see, e.g., Wand and Jones, 1995) ˆ fmix(x) = fmix(x)+O(h2 n)+ OP((nhn)−1/2), uniformly in x. By assumption (A2-ii) we have that O(h2 n) = O((nhn)−1/2), hence ˆ fmix(x)−fmix(x) = OP((nhn)−1/2) and ˆ fmix(x)f−1 mix(x)−1 = OP((nhn)−1/2). Similarly ˆm(x)−m(x) = OP((nhn)−1/2). Then we obtain ˆm(x)−m(x) = 1 nhnˆ fmix(x) k X l=1 nl X i=1 Kµx−Xil hn¶(Yil −m(x)) =1 nhnfmix(x) k X l=1 nl X i=1 Kµx−Xil hn¶(Yil −m(x)) + OP((nhn)−1)), uniformly in x. Taking this and condition (A2-ii) into account, the integral becomes Zˆm(x)−m(x) σj(x)fXj(x)dx =1 nhn k X l=1 nl X i=1 ZK((x−Xil)h−1 n)(Yil −m(x)) σj(x) fXj(x) fmix(x)dx +oP(n−1/2). Testing for the equality of kregression curves 43 Denote L(x) = (Yil −m(x))fXj(x)(fmix(x)σj(x))−1. Using the change of variable u= (x−Xil)h−1 n, a Taylor expansion of second order of Laround Xil and assumption (A3), we have that ZK((x−Xil)h−1 n)(Yil −m(x)) σj(x) fXj(x) fmix(x)dx =hnZK(u)L(Xil +hnu)du =hnL(Xil)ZK(u)du +h2 nL0(Xil)ZuK(u)du +OP(h3 n) =hnL(Xil) + OP(h3 n). Then, using assumption (A2), Zˆm(x)−m(x) σj(x)fXj(x)dx =1 n k X l=1 nl X i=1 Yil −m(Xil) σj(Xil) fXj(Xil) fmix(Xil)+oP(n−1/2). Lemma 2.2 Assume (A1)-(A3). Then, for any j= 1, . . . , k, Zˆmj(x)−mj(x) σj(x)fXj(x)dx =1 nj nj X i=1 Yij −mj(Xij) σj(Xij)+oP(n−1/2 j). Proof. Similar to the proof of the previous lemma. Lemma 2.3 Assume (A1)-(A3). Then, for any j= 1, . . . , k Zˆσj(x)−σj(x) σj(x)fXj(x)dx =1 nj nj X i=1 (Yij −mj(Xij)2−σ2 j(Xij) 2σ2 j(Xij)+oP(n−1/2 j). Proof. Write ˆσj(x)−σj(x) = ˆσ2 j(x)−σ2 j(x) 2σj(x)−(ˆσj(x)−σj(x))2 2σj(x). From Proposition 3 in Akritas and Van Keilegom (2001) the second term is of order OP((njhn)−1), and hence we can write ˆσj(x)−σj(x) = ˆσ2 j(x)−σ2 j(x) 2σj(x)+OP((njhn)−1) 44 Chapter 2 uniformly in x, and if we denote ˜σ2 j(x) = Pnj i=1 W(j) ij (x, hn)(Yij −mj(x))2, ˆσ2 j(x) = nj X i=1 W(j) ij (x, hn)Y2 ij −ˆm2 j(x) = nj X i=1 W(j) ij (x, hn)(Y2 ij −ˆm2 j(x)) = nj X i=1 W(j) ij (x, hn)(Yij −mj(x))2−(mj(x)−ˆmj(x))2 = ˜σ2 j(x) + OP((njhn)−1), because ˆmj(x)−mj(x) = OP((njhn)−1/2). Since σ2 j(x) can be considered as the regression function of the variable (Yij −mj(x))2, we have that ˜σ2 j(x)−σ2 j(x) = OP((njhn)−1/2). Now we can write ˜σ2 j(x)−σ2 j(x) = 1 njhnˆ fXj(x) nj X i=1 Kµx−Xij hn¶((Yij −mj(x))2−σ2 j(x)) =1 njhnfXj(x) nj X i=1 Kµx−Xij hn¶((Yij −mj(x))2−σ2 j(x)) +Ã1− ˆ fXj(x) fXj(x)!(˜σ2 j(x)−σ2 j(x)) =1 njhnfXj(x) nj X i=1 Kµx−Xij hn¶((Yij −mj(x))2−σ2 j(x)) +OP((njhn)−1). Using (A2-ii) write Zˆσj(x)−σj(x) σj(x)fXj(x)dx =Z˜σ2 j(x)−σ2 j(x) 2σ2 j(x)fXj(x)dx +OP((njhn)−1) =1 njhn nj X i=1 ZK((x−Xij)h−1 n)((Yij −mj(x))2−σ2 j(x)) 2σ2 j(x)dx +oP(n−1/2 j). Testing for the equality of kregression curves 45 Using a Taylor expansion of second order we obtain Zˆσj(x)−σj(x) σj(x)fXj(x)dx =1 nj nj X i=1 (Yij −mj(Xij))2−σ2 j(Xij) 2σ2 j(Xij)+OP(h2 n) + oP(n−1/2 j) =1 nj nj X i=1 (Yij −mj(Xij))2−σ2 j(Xij) 2σ2 j(Xij)+oP(n−1/2 j), and this concludes the proof of the lemma. Proof of Theorem 2.2. Write ˆ Fεj0(y)−ˆ Fεj(y) = ( ˆ Fεj0(y)−Fεj(y)) −(ˆ Fεj(y)−Fεj(y)).(2.15) First we will study the asymptotic behavior of ˆ Fεj0(y)−Fεj(y). We will use some results and proofs from Akritas and Van Keilegom (2001). These authors assume that the functions m,mj, and σjare L-functionals depending on a certain score function J. In our case these functionals are the conditional mean and variance, that correspond to J≡1. This choice of Jis not covered by the results in Akritas and Van Keilegom (2001). However, it is easy to check that the results in that paper can be extended to J≡1. From the proof of Theorem 1 in Akritas and Van Keilegom (2001) we have that ˆ Fεj0(y)−Fεj(y) = 1 nj nj X i=1 IµYij −m(Xij) σj(Xij)≤y¶−Fεj(y) (2.16) +fεj(y)Zy(ˆσj(x)−σj(x)) + ˆm(x)−m(x) σj(x)fXj(x)dx +Rnj(y), where supy|Rnj(y)|=oP(n−1/2 j). Note that in Akritas and Van Keilegom (2001) the estimation of the distribution of the residuals is considered from one sample. This means that, with our notation, the error εij is estimated by (Yij − ˆmj(Xij))/ˆσj(Xij). However, the decomposition given in (2.16) remains valid when the errors are estimated with ˆm, because their Lemma 1 holds in that case. 46 Chapter 2 The integral on the right side above can be decomposed as follows: Zy(ˆσj(x)−σj(x)) + ˆm(x)−m(x) σj(x)fXj(x)dx =yZˆσj(x)−σj(x) σj(x)fXj(x)dx +Zˆm(x)−m(x) σj(x)fXj(x)dx. Now, using Lemma 2.1, Lemma 2.3 and the fact that under the null hypothesis m=m1=··· =mk, ˆ Fεj0(y)−Fεj(y) = 1 nj nj X i=1 IµYij −m(Xij) σj(Xij)≤y¶−Fεj(y) (2.17) +yfεj(y)1 nj nj X i=1 (Yij −m(Xij))2−σ2 j(Xij) 2σ2 j(Xij) +fεj(y)1 n k X l=1 nl X i=1 Yil −m(Xil) σj(Xil) fXj(Xil) fmix(Xil)+oP(n−1/2) uniformly in y. Note that nj=O(n) because of condition (A2-i). Again from the proof of Theorem 1 in Akritas and Van Keilegom (2001) ˆ Fεj(y)−Fεj(y) = 1 nj nj X i=1 IµYij −m(Xij) σj(Xij)≤y¶−Fεj(y) +fεj(y)Zy(ˆσj(x)−σj(x)) + ˆmj(x)−m(x) σj(x)fXj(x)dx +oP(n−1/2 j), and using Lemma 2.2 and Lemma 2.3 ˆ Fεj(y)−Fεj(y) = 1 nj nj X i=1 IµYij −m(Xij) σj(Xij)≤y¶−Fεj(y) (2.18) +yfεj(y)1 nj nj X i=1 (Yij −m(Xij))2−σ2 j(Xij) 2σ2 j(Xij) +fεj(y)1 nj nj X i=1 Yij −m(Xij) σj(Xij)+oP(n−1/2 j) uniformly in y. Testing for the equality of kregression curves 53 where dn2j(x) = 1 and dn1j(x) = (R(x)−rj(x))σ−1 j(x). Clearly, these functions are in the conditions of the proof of Lemma 1 of the aforementioned paper. Using their notation dn2j∈˜ C1+δ 2(RX) for all δ > 0 and P(n−1/2dn1j∈C1+δ 1(RX)) →1 as n→ ∞ for 0 < δ ≤1, since dn1jis twice continuously differentiable on the compact set RXby assumption (A1). Hence taking into account (2.24) and (2.22), we can write 1 nj nj X i=1 IµYij −mn(Xij) σj(Xij)≤y¶(2.25) =1 nj nj X i=1 IµYij −mjn(Xij) σj(Xij)−n−1/2R(Xij)−rj(Xij) σj(Xij)≤y¶ =1 nj nj X i=1 IµYij −mjn(Xij) σj(Xij)≤y¶+n−1/2fεj(y)E·R(Xj)−rj(Xj) σj(Xj)¸+oP(n−1/2). As in the proof of Theorem 2.3 we have ˆ Fεj(y)−Fεj(y) = 1 nj nj X i=1 IµYij −mjn(Xij) σj(Xij)≤y¶−Fεj(y) (2.26) +fεj(y)Zy(ˆσj(x)−σj(x)) σj(x)fXj(x)dx +fεj(y)Zˆmj(x)−mjn(x) σj(x)fXj(x)dx +oP(n−1/2). By combining (2.23), (2.25) and (2.26) we obtain ˆ Fεj0(y)−ˆ Fεj(y) =fεj(y)Zˆm(x)−mn(x) σj(x)fXj(x)dx −fεj(y)Zˆmj(x)−mjn(x) σj(x)fXj(x)dx +n−1/2fεj(y)E·R(Xj)−rj(Xj) σj(Xj)¸+oP(n−1/2), and following a similar procedure as in Lemmas 2.2 and 2.3, we can write ˆ Fεj0(y)−ˆ Fεj(y) (2.27) =fεj(y) k X l=1 pl(1 nl nl X i=1 Yil −mln(Xil) σj(Xil)µfXj(Xil) fmix(Xil)−I(l=j) pj¶) +n−1/2fεj(y)E·R(Xj)−rj(Xj) σj(Xj)¸+oP(n−1/2). 54 Chapter 2 Note that (Yl−mln(Xl))/σj(Xil) = εl, which does not depend on n. The leading term of the representation given in (2.27) under Hl.a. is the same as the leading term of the representation given in Theorem 2.3 under the null hypothesis. Therefore their limit distributions are the same. This concludes the proof. Proof of Corollary 2.2. Similar to the proof of Corollary 2.1, only by taking into account the weak convergence of ˆ W(y) under Hl.a. given in Theorem 2.4. Proof of Theorem 2.5. We use the representation given in (2.18) ˆ Fεj(y)−Fε(y) = 1 nj nj X i=1 ϕj(Xij, Yij, y) + oP(n−1/2 j), uniformly in y, where ϕj(u, v, y) =Iµv−mj(u) σj(u)≤y¶−Fε(y) +yfε(y)(v−mj(u))2−σ2 j(u) 2σ2 j(u)+fε(y)v−mj(u) σj(u). Note that ˆ Fε(y)−Fε(y) = k X l=1 nl n(ˆ Fεl(y)−Fε(y)). By simply writing ˆ Fε(y)−ˆ Fεj(y) = ( ˆ Fε(y)−Fε(y)) −(ˆ Fεj(y)−Fε(y)), we have ˆ Fε(y)−ˆ Fεj(y) = k X l=1 ³nl n−I(l=j)´1 nl nl X i=1 ϕl(Xil, Yil, y) + oP(n−1/2). To proof the weak convergence of the k-dimensional process we will make use of the Cram´er-Wold device, as we did in the proof of Theorem 2.3. Let ˆ V(y) = Pn j=1 ajˆ Uj(y) be a linear combination of the components of the multidimensional process. It suffices to show the weak convergence of ˆ V(y). It is easy to prove that ˆ V(y) = k X j=1 (p1/2 jA−aj)1 n1/2 j nj X i=1 ϕj(Xij, Yij, y) + oP(1), Testing for the equality of kregression curves 55 where A=Pk l=1 p1/2 lal. The leading term of the process ˆ V(y) consists of a sum of kindependent processes (multiplied by a constant, which will not influence the weak convergence). Now write ˆ V(y) = k X j=1 (p1/2 jA−aj)ˆ Vj(y) + oP(1), where ˆ Vj(y) is the process Gj-indexed by the class of functions Gj={(u, v)→ϕj(u, v, y),−∞ < y < ∞}. The class of functions Gjcan be written as Gj=Gj1+Gj2+Gj3+Gj4, where Gj1=½(u, v)→Iµv−mj(u) σj(u)≤y¶,−∞ < y < ∞¾, Gj2={(u, v)→ −Fε(y),−∞ < y < ∞}, Gj3=½(u, v)→yfε(y)(v−mj(u))2−σ2 j(u) 2σ2 j(u),−∞ < y < ∞¾, Gj4=½(u, v)→fε(y)v−mj(u) σj(u),−∞ < y < ∞¾. The classes Gj3and Gj4factorize in a part not depending on yand a bounded function of y. Also the class Gj2can be treated in the same way, since it does not depend on (u, v). Let Mbe such that supy{|yfε(y)|,|fε(y)|,|Fε(y)|} < M. Hence, for l= 2,3,4, N[ ](δ, Gjl, L2(P)) ≤2Mδ−1if δ < 2Mand N[ ](δ, Gjl, L2(P)) = 1 if δ > 2M. The functions in class Gj1are decreasing in (v−mj(u))/σj(u). Then, by Theorem 2.7.5 in van der Vaart and Wellner (1996), N[ ](δ, Gj1, L2(P)) = O(exp(Kδ−1)) for some constant K > 0. Obviously, since the functions in Gj1are bounded with values in [0,1], if δ > 1 then N[ ](δ, Gj1, L2(P)) = 1. These features of the bracketing numbers of the classes involved in the class Gjensure the weak convergence of the process ˆ Vj(y), as we explained in the proof 56 Chapter 2 of Theorem 2.3. Therefore we can conclude that our process of interest ˆ V(y) is weakly convergent, and by the Cram´er-Wold device so it is ˆ U(y). The covariance structure given in the statement of the Theorem can be obtained immediately. Proof of Corollary 2.3. The proof of the convergence of the test statistics SKS and SCM can be obtained by following the same arguments as in the proof of Corollary 2.1. Chapter 3 Comparison of regression curves with censored responses In this chapter we consider the problem of comparison of regression curves when the response variables are censored. The test is based on a comparison of Kaplan-Meier estimators of the distribution of the censored residuals. As in the previous chapter, Kolmogorov-Smirnov and Cram´er-von Mises type statistics are considered. Some asymptotic results are proved: weak convergence of the process of interest, convergence of the test statistics and behavior of the process under local alternatives. We also describe a bootstrap procedure in order to approximate the critical values of the test. A simulation study and an application to a real data set conclude the chapter. This chapter is based on Pardo-Fern´andez and Van Keilegom (2005) and it extends the ideas and results obtained in Chapter 2. 3.1 Motivation and statistical model Regression models are used for describing the relationship between a response and a covariate. In the field of survival analysis it can be useful to allow for censoring in the response variable. For instance, we can consider a model where the survival time (for patients having a certain disease) is the response variable and the age 57 58 Chapter 3 is the covariate. If we can distinguish two or more groups in the population (by gender, treated patients and non-treated patients, etc.), we may be interested in testing for the equality of the corresponding regression curves. This kind of test allows us to check whether the effect of the covariate over the variable of interest is the same in all groups. As pointed out in Fan and Gijbels (1994), when the response variable is censored the usual tools of regression (scatter plots, residuals plots, etc.) are not directly applicable to check, at least visually, the shape of the regression curves. This motivates the development of analytic tools in censored regression. In this context, the statistical model can be described as follows. Let (Xj, Yj), j= 1, . . . , k, be independent random vectors, where Yjrepresents a certain response variable associated with the covariate Xj. Suppose that the covariates have common support RX. Assume that, for j= 1, . . . , k, the response variable Yj is subject to random right censoring. This means that there exists a censoring variable Cj, independent of Yjgiven Xj, such that we can observe Zj= min{Yj, Cj} and the indicator of censoring ∆j=I(Yj≤Cj). For j= 1, . . . , k, assume that the following non-parametric regression models hold Yj=mj(Xj) + σj(Xj)εj, where the error variable εjis independent of Xj,mjis an unknown conditional location function mj(x) = Z1 0 F−1 j(s|x)J(s)ds (3.1) and σjis an unknown conditional scale function representing possible heteroscedasticity σ2 j(x) = Z1 0 F−1 j(s|x)2J(s)ds −m2 j(x),(3.2) where Fj(·|x) is the conditional distribution of Yjgiven the value xof the covariate Xj,F−1 j(s|x) = inf{t;Fj(t|x)≥s}is the corresponding quantile function and J(s) is a score function satisfying R1 0J(s)ds = 1. Denote Fεj(y) = P(εj≤y) for the distribution of the error εjin population j. By the definitions (3.1) and (3.2), R1 0F−1 εj(s)J(s)ds = 0 and R1 0F−1 εj(s)2J(s)ds = 1. Comparison of regression curves with censored responses 59 The choice of the score function Jleads to different location and scale functions. In particular if J(s) = I(0 ≤s≤1) then mj(x) = E(Yj|Xj=x) is the conditional mean function and σ2 j(x) = V ar(Yj|Xj=x) is the conditional variance function. However, it may happen that this choice of Jis not appropriate because of the inconsistency of the estimator of the conditional distribution Fj(·|x) in the right tail due to the censoring (see Section 1.2). A useful choice is J(s) = 1 q−pI(p≤s≤q) for some 0 ≤p < q ≤1, which leads to trimmed means and trimmed variances. The conditional median or other conditional quantiles can be seen as limits of trimmed means. The samples are (Xij, Zij,∆ij), i= 1, ..., nj, from the distribution of (Xj, Zj,∆j), for j= 1, . . . , k. Denote n=Pk j=1 nj. We are interested in testing the null hypothesis of equality between the location (regression) functions H0:m1=m2=··· =mk,(3.3) versus the alternative Ha:mi6=mjfor some i, j ∈ {1, . . . , k}. When the distributions of the errors and the variance functions are the same in all the groups (we do not assume so, but it is an interesting situation), if the null hypothesis holds for a particular definition of the location function, that is for a particular choice of J, then it holds for all possible location functions. However, in a general situation with different variances or different error distributions, H0 can be true for a particular choice of the functions mjand false for another one. In Chapter 2 a mechanism of comparison of regression curves for complete data was developed via the estimation of the distribution of the errors of the regression models. The idea of the testing procedure is to compare two estimators of the distribution of the errors in each population. In this chapter we will extend that methodology to the situation where the response variable may be censored. Now, 60 Chapter 3 because of the censoring in the response variable, we will consider Zij −ˆmj(Xij) ˆσj(Xij) and Zij −ˆm(Xij) ˆσj(Xij) to estimate the censored residuals, and we will substitute the empirical distribution by the Kaplan-Meier estimator of the distribution under random censoring (Kaplan and Meier, 1958). The chapter is organized as follows. In Section 3.2 we will explain the testing procedure. In Section 3.3 we will state the main asymptotic results. A bootstrap procedure to approximate the critical points of the test is described in Section 3.4 and a simulation study is presented in Section 3.5. Finally, we include an application to real data in Section 3.6. The proofs of the main results are deferred to Section 3.7. 3.2 Testing procedure The testing procedure is based on the comparison of two non-parametric estimators of the distribution of the residuals Fεjin each population. This involves nonparametric estimation of the location and scale functions. All these estimators will be constructed using the estimator of the conditional distribution function Fj(·|x) when the response is censored introduced by Beran (1981): ˆ Fj(y|x) = 1 −Y Zij ≤y,∆ij =1 Ã1−W(j) ij (x, hn) Pnj l=1 I(Zlj ≥Zij)W(j) lj (x, hn)!, where W(j) ij (x, hn) = K((x−Xij)/hn) Pnj l=1 K((x−Xlj)/hn) are Nadaraya-Watson type weights, Kis a known kernel and hnis an appropriate bandwidth sequence. Comparison of regression curves with censored responses 61 Now consider the following estimator of the location function for each sample, for j= 1, . . . , k, ˆmj(x) = Z1 0 ˆ F−1 j(s|x)J(s)ds, and an estimator of the common location function under the null hypothesis (which we will denote by m) taking into account all the samples ˆm(x) = k X j=1 nj n ˆ fXj(x) ˆ fmix(x)ˆmj(x), where, ˆ fXj(x) = 1 njhn nj X i=1 Kµx−Xij hn¶ is the kernel estimator of the density fXjof Xj, and ˆ fmix(x) = k X j0=1 nj0 nˆ fXj0(x). Note that ˆ fXj(x) can be computed in the usual way because the covariates do not suffer from censoring. The estimator of the scale function σjfrom each sample is ˆσ2 j(x) = Z1 0 ˆ F−1 j(s|x)2J(s)ds −ˆm2 j(x). The score function Jwill be chosen so that ˆmj(x) and ˆσ2 j(x) are consistent, even in the case of the tails of the Beran estimator not being consistent. Compute the estimators of the censored residuals in each sample ˆ Eij =Zij −ˆmj(Xij) ˆσj(Xij) for i= 1, . . . , nj,j= 1, . . . , k, and estimate the distribution of the residuals from the censored sample ( ˆ Eij,∆ij) using the Kaplan-Meier estimator ˆ Fεj(y) = 1 −Y ˆ Eij ≤y,∆ij =1 Ã1−1 Pnj l=1 I(ˆ Elj ≥ˆ Eij)!.(3.4) 62 Chapter 3 If the null hypothesis is true, we can estimate the residuals in each sample using the estimator of the common regression function ˆm, that is ˆ Eij0=Zij −ˆm(Xij) ˆσj(Xij) for i= 1, . . . , nj,j= 1, . . . , k, and estimate the corresponding distribution from the censored sample ( ˆ Eij0,∆ij) ˆ Fεj0(y) = 1 −Y ˆ Eij0≤y,∆ij =1 Ã1−1 Pnj l=1 I(ˆ Elj0≥ˆ Eij0)!.(3.5) Under the null hypothesis, both ˆ Fεjand ˆ Fεj0are estimators of Fεj. The fact that there exists some difference between these two estimators of the distribution of the errors gives evidence for the inequality of the location functions. This idea is formalized theoretically in the following theorem. Note that ˆm(x) estimates consistently m(x) = Pk j=1 pj fXj(x) fmix(x)mj(x), where fmix(x) = Pk j=1 pjfXj(x) is the mixture of the densities of the covariates, provided that nj/n →pj>0. Let Fεj(y) = PµYj−mj(Xj) σj(Xj)≤y¶ and Fεj0(y) = PµYj−m(Xj) σj(Xj)≤y¶ be the theoretical versions (without estimated curves) of the distributions considered in (3.4) and (3.5). Theorem 3.1 Assume that mjis continuous, j= 1, . . . , k, and the moments of order νof the distributions Fεj(y)and Fεj0(y)exist for all ν∈N. Then Fεj(y) = Fεj0(y),−∞ < y < ∞,j= 1, . . . , k, if and only if m1(x) = . . . =mk(x)for all x∈RX. The equivalence given in the previous result is a theoretical justification of the proposed testing procedure. Its proof can be found in Section 3.7. Comparison of regression curves with censored responses 69 pothesis given in Corollary 3.1 are very complicated. Here we consider a bootstrap procedure based on the censored residuals to approximate the critical values. First, for j= 1, . . . , k and i= 1, . . . , nj, estimate the censored residuals in a non parametric way, using each sample separately ˆ Eij =Zij −ˆmj(Xij) ˆσj(Xij). From the censored sample of estimated residuals {(ˆ Eij,∆ij), i = 1, . . . , nj}compute the Kaplan-Meier estimator ˆ Fεjand standardize these residuals in order to verify the initial assumption of having location function 0 and scale function 1. The standardized residuals are ˜ Eij = ( ˆ Eij −λ1j)/λ2j, where λ1j=Rˆ F−1 εj(s)J(s)ds and λ2j= (Rˆ F−1 εj(s)2J(s)ds −λ2 1j)1/2. For resampling the censored residuals we use the ‘naive bootstrap’ described in Efron (1981) and studied in Akritas (1986). Different approaches of smooth bootstrap for censored data were considered in Gonz´alez-Manteiga, Cao and Marron (1996). The bootstrap procedure we propose consists of the following steps. For fixed Band for b= 1, . . . , B, 1. For each j= 1, . . . , k and i= 1, . . . , nj: •Let Y∗ ij,b = ˆm(Xij)+ ˆσj(Xij)ε∗ ij,b, where ε∗ ij,b =Vij,b+ajSij,b,Vij,b is drawn from ˆ Fεj(standardized), and Sij,b is a random variable with mean zero and variance one. •Select at random a C∗ ij,b from a smoothed version of ˆ Gj(·|Xij), which is the Beran estimator of Gj(·|Xij) obtained by replacing ∆ij by 1 −∆ij in the expression for ˆ Fj(·|Xij). •Let Z∗ ij,b = min{Y∗ ij,b, C∗ ij,b}and ∆∗ ij,b =I(Y∗ ij,b ≤C∗ ij,b). 2. The bootstrap samples are, for j= 1, . . . , k,{(Xij, Z∗ ij,b,∆∗ ij,b), i = 1, . . . , nj}. 3. Let T∗ KS,b and T∗ CM,b be the test statistics obtained from the bootstrap samples. 70 Chapter 3 If we denote T∗ KS,(b)the b-th order statistic of the values T∗ KS,1, . . . , T∗ KS,B obtained in step 3 (analogously for T∗ CM,(b)), then T∗ KS,([(1−α)B]) and T∗ CM,([(1−α)B]) approximate the (1 −α)-quantiles of the distribution of TKS and TCM under the null hypothesis respectively. 3.5 Simulation study In this section we present some simulations in order to study the practical behavior of the proposed bootstrap procedure. We restrict ourselves to two-sample situations (k= 2). More precisely, we consider the following models: (i)m1(x) = x;m2(x) = x (ii)m1(x) = exp(x); m2(x) = exp(x) (iii)m1(x) = x;m2(x) = x+ 1 (iv)m1(x) = exp(x); m2(x) = exp(x) + x Clearly, models (i) and (ii) correspond to the null hypothesis and models (iii) and (iv) to the alternative hypothesis. In each case we consider a homoscedastic and a heteroscedastic situation. In the homoscedastic case the variances are σ2 1(x) = 0.25 and σ2 2(x) = 0.50, (3.8) while in the heteroscedastic case the variance functions are σ2 1(x) = ex R1 0etdt and σ2 2(x) = e2x R1 0e2tdt.(3.9) Note that in the heteroscedastic case the variances are larger than in the homoscedastic case. The censoring variables are Cj=mj(Xj) + σj(Xj)ρj, where ρjhas survival function 1 −Fρ(y) = (1 −Fε(y))β. This mechanism of censoring can be seen as a ‘conditional Koziol-Green model’ (see Koziol and Green, 1976) and it allows us to have the same amount of censoring over all the support of the covariates. The expected proportion of censored data is (1 + β)−1. In the tables we consider β= 1/3 (25% of censoring) and β= 1 (50% of censoring). Comparison of regression curves with censored responses 71 In the theoretical results we have used only one bandwidth. As in Section 2.5, we have found that the bandwidth has not a big impact on the results of the tests, but it is recommendable to use the same bandwidth to estimate mand mj. The variance functions could be estimated with different bandwidths. In these simulations we use a bandwidth of the form h=cn−3/10 to estimate m,mjand σj, for j= 1,2. The bandwidths chosen in this way verify the regularity conditions assumed in the theoretical results. In the tables the cases c= 1 and c= 1.5 are shown. This will allow us to check the sensitivity of the test to the change of the bandwidth. For the kernel needed to calculate the weights that appear in the Beran estimator, we choose the Epanechnikov kernel K(u) = 0.75(1 −u2)I(|u|<1). In Tables 3.1 and 3.2 the distribution of the errors is Exponential, transformed such that R1 0F−1 εj(s)J(s)ds = 0 and R1 0F−1 εj(s)2J(s)ds = 1 and the covariates are uniformly distributed in [0,1]. The regression and variance functions are those corresponding to expressions (3.1) and (3.2) with the choice J(s)=0.75−1I(0 ≤ s≤0.75) for the score function. For the test statistics in (3.6) and (3.7) we take as the threshold Tthe value corresponding to the quantile 75% of the combined sample of the estimated residuals under the null hypothesis. Note that all these choices are reasonable for the models and censoring mechanisms we have considered. We work with aj=n−3/10 jin the smooth bootstrap. Table 3.1 shows the proportion of rejections in 1000 trials for sample sizes (n1, n2) = (50,50), (100,50) and (100,100), and when the expected amount of censored data is 25%. Table 3.2 shows the proportion of rejections in 1000 trials for sample sizes (n1, n2) = (100,100), (200,100) and (200,200) when the expected amount of censored data is 50%. In all cases we worked with B= 200 bootstrap replications and significance levels α= 0.05 and α= 0.10. Larger samples sizes for models with 50% of censored data are justified by the difficulty of those models. The approximation of the level –models (i) and (ii)– is good in most cases. The results for models (iii) and (iv) show that the tests gain power as the sample sizes increase. In almost all cases the test based on TCM gives better results than the test based on TKS, and we also observe that the choice of the bandwidth has little impact on the rejection probabilities. 72 Chapter 3 Table 3.1: Rejection probabilities under models (i) to (iv) of the tests based on TKS and TCM when the expected amount of censored data is 25%. The models are homoscedastic, with variances given in (3.8), and heteroscedastic, with variances given in (3.9). c= 1 c= 1.5 TKS TCM TKS TCM (n1, n2)α: 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 Homoscedastic models (50,50) (i) 0.049 0.084 0.048 0.093 0.044 0.095 0.051 0.091 (ii) 0.042 0.089 0.048 0.094 0.044 0.090 0.051 0.094 (iii) 0.984 0.993 0.989 0.996 0.984 0.994 0.989 0.994 (iv) 0.505 0.629 0.550 0.674 0.489 0.643 0.560 0.667 (100,50) (i) 0.059 0.112 0.057 0.104 0.061 0.098 0.067 0.098 (ii) 0.067 0.107 0.058 0.102 0.055 0.096 0.058 0.100 (iii) 0.999 0.999 0.998 0.999 1.000 1.000 1.000 1.000 (iv) 0.557 0.724 0.541 0.690 0.580 0.720 0.568 0.699 (100,100) (i) 0.055 0.109 0.059 0.103 0.056 0.109 0.058 0.102 (ii) 0.060 0.117 0.062 0.103 0.057 0.105 0.058 0.110 (iii) 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 (iv) 0.863 0.924 0.888 0.935 0.862 0.923 0.879 0.932 Heteroscedastic models (50,50) (i) 0.045 0.090 0.044 0.094 0.044 0.085 0.046 0.090 (ii) 0.042 0.085 0.043 0.089 0.045 0.090 0.048 0.089 (iii) 0.771 0.849 0.792 0.864 0.763 0.853 0.777 0.866 (iv) 0.243 0.360 0.256 0.357 0.245 0.346 0.257 0.356 (100,50) (i) 0.048 0.087 0.044 0.086 0.051 0.102 0.048 0.092 (ii) 0.050 0.086 0.046 0.088 0.048 0.089 0.049 0.091 (iii) 0.909 0.955 0.880 0.939 0.913 0.956 0.892 0.947 (iv) 0.236 0.343 0.191 0.283 0.246 0.363 0.221 0.313 (100,100) (i) 0.066 0.112 0.057 0.116 0.059 0.117 0.059 0.111 (ii) 0.061 0.111 0.059 0.112 0.061 0.109 0.056 0.110 (iii) 0.976 0.987 0.973 0.989 0.971 0.985 0.974 0.988 (iv) 0.472 0.582 0.476 0.583 0.465 0.584 0.474 0.581 Comparison of regression curves with censored responses 73 Table 3.2: Rejection probabilities under models (i) to (iv) of the tests based on TKS and TCM when the expected amount of censored data is 50%. The models are homoscedastic, with variances given in (3.8), and heteroscedastic, with variances given in (3.9). c= 1 c= 1.5 TKS TCM TKS TCM (n1, n2)α: 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 Homoscedastic models (100,100) (i) 0.052 0.093 0.055 0.093 0.042 0.081 0.049 0.083 (ii) 0.051 0.083 0.050 0.099 0.037 0.075 0.046 0.087 (iii) 0.995 0.997 0.995 0.997 0.995 0.997 0.994 0.996 (iv) 0.751 0.855 0.813 0.903 0.723 0.838 0.792 0.888 (200,100) (i) 0.072 0.126 0.078 0.137 0.055 0.108 0.057 0.116 (ii) 0.071 0.119 0.074 0.139 0.064 0.111 0.064 0.114 (iii) 1.000 1.000 0.999 1.000 0.999 1.000 1.000 1.000 (iv) 0.808 0.903 0.841 0.910 0.851 0.922 0.866 0.933 (200,200) (i) 0.053 0.101 0.059 0.105 0.052 0.090 0.057 0.098 (ii) 0.055 0.094 0.059 0.103 0.055 0.092 0.058 0.098 (iii) 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 (iv) 0.980 0.992 0.988 0.993 0.982 0.991 0.986 0.993 Heteroscedastic models (100,100) (i) 0.032 0.067 0.042 0.089 0.038 0.077 0.043 0.089 (ii) 0.035 0.070 0.041 0.081 0.039 0.071 0.045 0.086 (iii) 0.932 0.964 0.942 0.972 0.905 0.953 0.926 0.966 (iv) 0.387 0.520 0.441 0.578 0.370 0.495 0.430 0.565 (200,100) (i) 0.050 0.103 0.065 0.125 0.045 0.098 0.055 0.104 (ii) 0.053 0.091 0.062 0.122 0.046 0.092 0.052 0.110 (iii) 0.987 0.995 0.991 0.997 0.990 0.997 0.990 0.998 (iv) 0.381 0.500 0.366 0.478 0.381 0.513 0.381 0.512 (200,200) (i) 0.044 0.081 0.049 0.090 0.041 0.082 0.048 0.082 (ii) 0.046 0.086 0.049 0.089 0.046 0.086 0.046 0.089 (iii) 0.998 0.999 0.998 1.000 0.998 1.000 0.998 1.000 (iv) 0.688 0.787 0.722 0.812 0.669 0.786 0.722 0.800 74 Chapter 3 3.6 Application to real data We illustrate our testing procedure with an application to the Small Cell Lung Cancer Data. The data set is available in Ying, Jung and Wei (1995) and consists of lifetimes of patients suffering from small cell lung cancer. The patients were divided into two groups which followed two different treatments (Group A and Group B). The first group consisted of 62 patients (15 censored) and the second group consisted of 59 patients (8 censored). We considered the base 10 log of the survival time (in days) as response variable and the age as covariate. The support of the covariate was transformed into the interval [0,1]. We worked with different values for the bandwidth needed in the estimation, ranging from 0.15 to 0.40. We have performed the test of equality of the regression curves of the two curves using as score function J(s)=0.75−1I(0 ≤s≤0.75) and J(s) = 0.50−1I(0.25 ≤ s≤0.75). The second choice of the function Jproduces curves closer to the conditional median. The obtained results are very similar. The p-values are obtained from 1000 bootstrap replications. Figure 3.1 shows the estimated curves, using h= 0.30 as a bandwidth. When testing for the equality of the curves, the null hypothesis is clearly rejected in all cases, with p-values smaller than 0.02 for the statistic TKM and smaller than 0.005 for TCM . However, it seems reasonable to suppose that the regression curves differ only by a shift (see Figure 3.1). A test to check that can be performed by transforming the response variables in Z0 ij =Zij −tj, for j= 1,2 and i= 1, . . . , nj, where tj=n−1P2 l=1 Pnl i=1 ˆmj(Xil). In this case the p-values are larger than 0.55 for the statistic TKM and larger than 0.67 for TCM . All these results are summarized in Figures 3.2 and 3.3, which show graphs of the p-values versus the bandwidth. Comparison of regression curves with censored responses 75 age log (survival time) 0.0 0.2 0.4 0.6 0.8 1.0 1.5 2.0 2.5 3.0 3.5 age log (survival time) 0.0 0.2 0.4 0.6 0.8 1.0 1.5 2.0 2.5 3.0 3.5 Figure 3.1: Scatter plot of ‘log10(survival time)’ versus ‘age’ (rescaled to [0,1]) and estimated regression curves of Group A (solid line, +for uncensored data, ¤for censored data) and Group B (dashed line, ×for uncensored data, 4for censored data), with J(s)=0.75−1I(0 ≤s≤0.75) (top) and J(s) = 0.50−1I(0.25 ≤s≤ 0.75) (bottom). 76 Chapter 3 h p-value 0.15 0.20 0.25 0.30 0.35 0.40 0.0 0.01 0.02 0.03 0.04 0.05 0.06 h p-value 0.15 0.20 0.25 0.30 0.35 0.40 0.0 0.01 0.02 0.03 0.04 0.05 0.06 Figure 3.2: Graphs of the p-values as function of the bandwidth hwhen testing for the equality of the regression curves with the test statistics TKS (line with circles) and TCM (line with crosses). The location curves correspond to J(s) = 0.75−1I(0 ≤ s≤0.75) (left) and J(s) = 0.50−1I(0.25 ≤s≤0.75) (right). The solid horizontal line corresponds to a p-value of 0.05. h p-value 0.15 0.20 0.25 0.30 0.35 0.40 0.0 0.2 0.4 0.6 0.8 1.0 h p-value 0.15 0.20 0.25 0.30 0.35 0.40 0.0 0.2 0.4 0.6 0.8 1.0 Figure 3.3: Graphs of the p-values as function of the bandwidth hwhen testing for constant difference between the regression curves with the test statistics TKS (line with circles) and TCM (line with crosses). The location curves correspond to J(s)=0.75−1I(0 ≤s≤0.75) (left) and J(s)=0.50−1I(0.25 ≤s≤0.75) (right). The solid horizontal line corresponds to a p-value of 0.05. Comparison of regression curves with censored responses 77 3.7 Auxiliary results and proofs Proof of Theorem 3.1. Assume that Fεj(y) = Fεj0(y), for −∞ < y < ∞. We write PµYj−m(Xj) σj(Xj)≤y¶=PµYj−mj(Xj) σj(Xj)+mj(Xj)−m(Xj) σj(Xj)≤y¶, for all y, or equivalently Pµexp ½Yj−m(Xj) σj(Xj)¾≤y¶ =Pµexp ½Yj−mj(Xj) σj(Xj)¾exp ½mj(Xj)−m(Xj) σj(Xj)¾≤y¶, for all y. Since (Yj−mj(Xj))/σj(Xj) and Xjare independent, it follows that E"µexp ½Yj−m(Xj) σj(Xj)¾¶2ν# =E"µexp ½Yj−mj(Xj) σj(Xj)¾¶2ν#E"µexp ½mj(Xj)−m(Xj) σj(Xj)¾¶2ν#, for all ν. Then E"µexp ½mj(Xj)−m(Xj) σj(Xj)¾¶2ν#= 1, for all ν. Carleman’s condition (see e.g. Feller, 1966) ensures that Pµexp ½mj(Xj)−m(Xj) σj(Xj)¾= 1¶= 1 or Pµmj(Xj)−m(Xj) σj(Xj)= 0¶= 1, and this clearly implies that mj(x) = m(x) for all j= 1, . . . , k and for all x∈RX, except for a set of points of probability zero. The continuity allows us to extend the equality of the regression curves to the whole support of the covariates. The converse implication is trivial. First we set four auxiliary lemmas, and then we prove the main results. 78 Chapter 3 Lemma 3.1 Assume (A1)-(A5) and Hej(y|x)satisfy (A6). Then under the null hypothesis H0, for j= 1, . . . , k, ˆ Hej0(y)−Hej(y) =1 nj nj X i=1 I(Eij ≤y)−Hej(y)−1 nj nj X i=1 yhej(y|Xij)ζj(Zij,∆ij|Xij) −1 n k X l=1 nl X i=1 hej(y|Xil)fXj(Xil) fmix(Xil) σl(Xil) σj(Xil)ηl(Zil,∆il|Xil) + oP(n−1/2), uniformly in −∞ < y ≤T. Proof. From the proof of Proposition A.2 in Van Keilegom and Akritas (1999), we have that ˆ Hej0(y)−Hej(y) = 1 nj nj X i=1 I(Eij ≤y)−Hej(y) (3.10) +Zhej(y|x)ˆm(x)−m(x) σj(x)fXj(x)dx +Zyhej(y|x)ˆσj(x)−σj(x) σj(x)fXj(x)dx +oP(n−1/2 j), uniformly in −∞ < y ≤T. The last term is oP(n−1/2 j) because of the uniform consistency of ˆmand ˆσj. The consistency of ˆσjis given in Proposition 4.5 in Van Keilegom and Akritas (1999). The consistency of ˆmcan be obtained using the consistency of ˆml(also given in Proposition 4.5 in Van Keilegom and Akritas, 1999), the consistency of ˆ fXland ˆ fmix and taking into account the relation ˆm(x)−m(x) = k X l=1 nl n ˆ fXl(x) ˆ fmix(x)( ˆml(x)−m(x)) = k X l=1 nl n fXl(x) fmix(x)( ˆml(x)−m(x)) + oP(n−1/2), uniformly in x. First using Proposition 4.8 in Van Keilegom and Akritas (1999) ˆm(x)−m(x) =−1 nhn 1 fmix(x) k X l=1 nl X i=1 σl(x)Kµx−Xil hn¶ηl(Zil,∆il|x) + oP(n−1/2), Comparison of regression curves with censored responses 85 considerations. Under Hl.a., the estimator of the common regression curve ˆm estimates mn(x) = m0(x) + n−1/2R(x), where R(x) = Pk l=1 pl fXl(x) fmix(x)rl(x), and ˆmj(x) estimates mjn(x) = m0(x)+n−1/2rj(x). The censored residuals with respect to mnare Ej0=Zj−mn(Xj) σj(Xj), and with respect to mjn the residuals are Ej=Zj−mjn(Xj) σj(Xj). We have the relation Ej0=Zj−mn(Xj) σj(Xj)=Ej+mjn(Xj)−mn(Xj) σj(Xj)=Ej−n−1/2R(Xj)−rj(Xj) σj(Xj). Note also that now ˆ Fεj0(y) estimates Fεj0(y) = P((Yj−mn(Xj))/σj(Xj)≤y). We denote Hej0(y) = P(Ej0≤y), Hej01(y) = P(Ej0≤y, ∆j= 1), Hej0(y|x) = P(Ej0≤y|Xj=x), Hej01(y|x) = P(Ej0≤y, ∆j= 1|Xj=x). If we define E0 j= (Z0 j−m0(Xj))/σj(Xj) and denote H0 ej(y|x) = P(E0 j≤ y|Xj=x) and H0 ej1(y|x) = P(E0 j≤y, ∆j= 1|Xj=x), then it is easy to check that Hej(y|x) = H0 ej(y|x) and Hej1(y|x) = H0 ej1(y|x), which do not depend on n. Note that the independence of Yjand Cjgiven Xj=ximplies the independence of Y0 jand C0 jgiven Xj=x. Lemma 3.5 Assume (A1)-(A5) and Hej(y|x)satisfy (A6) and (AR) holds. Then, under the alternative hypothesis Hl.a., for any j= 1, . . . , k, ˆ Hej0(y)−Hej0(y) =1 nj nj X i=1 I(Eij ≤y)−Hej(y)−1 nj nj X i=1 yhej(y|Xij)ζj(Zij,∆ij|Xij) −1 n k X l=1 nl X i=1 hej(y|Xil)fXj(Xil) fmix(Xil) σl(Xil) σj(Xil)ηl(Zil,∆il|Xil) + oP(n−1/2 j), uniformly in −∞ < y ≤T. 86 Chapter 3 Proof. We follow the proof of Proposition A.2 in Van Keilegom and Akritas (1999) and write ˆ Hej0(y)−Hej0(y) = n−1 j nj X i=1 I(Eij0≤y)−Hej0(y) (3.13) +ZHej0µyˆσj(x) + ˆm(x)−mn(x) σj(x)¯¯¯¯ x¶fXj(x)dx −ZHej0(y|x)fXj(x)dx +oP(n−1/2), uniformly in −∞ < y ≤T. Note that the remainder term in (3.13) is oP(n−1/2) provided that ˆm−mnsatisfies Propositions 4.5, 4.6 and 4.7 in Van Keilegom and Akritas (1999). These propositions can be shown to hold true under standard lines of proof. We will analyze in detail each term of the expression above. First Hej0(y) = ZHej0(y|x)fXj(x)dx (3.14) =ZHej(y|x)fXj(x)dx +Zhej(y|x)x)mn(x)−mjn(x) σj(x)fXj(x)dx+ =Hej(y) + n−1/2E·hej(y|Xj)R(Xj)−rj(Xj) σj(Xj)¸+o(n−1/2), and ZHej0µyˆσj(x) + ˆm(x)−mn(x) σj(x)¯¯¯¯ x¶fXj(x)dx (3.15) =ZHej(y|x)fXj(x)dx +Zhej(y|x)yˆσj(x) + ˆm(x)−mjn(x)−yσj(x) σj(x)fXj(x)dx +o(n−1/2) =Hej(y) + Zhej(y|x)y(ˆσj(x)−σj(x)) + ˆm(x)−mn(x) σj(x)fXj(x)dx +n−1/2E·hej(y|Xj)R(Xj)−rj(Xj) σj(Xj)¸+o(n−1/2). An application of the proof of Lemma A.1 in Van Keilegom and Akritas (1999) Comparison of regression curves with censored responses 87 yields sup y¯¯¯¯¯ n−1 j nj X i=1 ½IµEij −n−1/2R(Xij)−rj(Xij) σj(Xij)≤y¶ −PµEj−n−1/2R(Xj)−rj(Xj) σj(Xj)≤y¶−I(Eij ≤y) + P(Ej≤y)¾¯¯¯¯ =oP(n−1/2 j). Considering the following probability as a function of yand using a Taylor expansion we obtain PµEj−n−1/2R(Xj)−rj(Xj) σj(Xj)≤y¶ =ZPµEj−n−1/2R(Xj)−rj(Xj) σj(Xj)≤y¯¯¯¯ Xj=x¶fXj(x)dx =ZP(Ej≤y|Xj=x)fXj(x)dx +n−1/2E·hej(y|Xj)R(Xj)−rj(Xj) σj(Xj)¸ +o(n−1/2) =P(Ej≤y) + n−1/2E·hej(y|Xj)R(Xj)−rj(Xj) σj(Xj)¸+o(n−1/2), and hence n−1 j nj X i=1 I(Eij0≤y) = n−1 j nj X i=1 IµEij −n−1/2R(Xij)−rj(Xij) σj(Xij)≤y¶(3.16) =n−1 j nj X i=1 I(Eij ≤y) + n−1/2E·hej(y|Xj)R(Xj)−rj(Xj) σj(Xj)¸+oP(n−1/2). Substituting (3.14), (3.15) and (3.16) in (3.13), we obtain ˆ Hej0(y)−Hej0(y) = n−1 j nj X i=1 I(Eij ≤y)−Hej(y) +Zhej(y|x)y(ˆσj(x)−σj(x)) + ˆm(x)−mn(x) σj(x)fXj(x)dx +oP(n−1/2). Since ˆm(x)−mn(x) = Pk l=1 nl n fXl(x) fmix(x)( ˆml(x)−mln(x)) + oP(n−1/2), the integral on the right-hand side of the expression above can be handled in a similar way as in the proof of Lemma 3.1. 88 Chapter 3 Lemma 3.6 Assume (A1)-(A5) and Hej1(y|x)satisfy (A6) and (AR) holds. Then, under the alternative hypothesis Hl.a., for any j= 1, . . . , k, ˆ Hej10(y)−Hej10(y) =1 nj nj X i=1 I(Eij ≤y, ∆ij = 1) −Hej1(y)−1 nj nj X i=1 yhej1(y|Xij)ζj(Zij,∆ij|Xij) −1 n k X l=1 nl X i=1 hej1(y|Xil)fXj(Xil) fmix(Xil) σl(Xil) σj(Xil)ηl(Zil,∆il|Xil) + oP(n−1/2 j), uniformly in −∞ < y ≤T. Proof. Similar to the proof of Lemma 3.5. Proof of Theorem 3.4. First we write Zy −∞ d(ˆ Hej10(s)−Hej10(s)) 1−Hej0(s)=Zy −∞ d(ˆ Hej10(s)−Hej10(s)) 1−Hej(s)(3.17) +Zy −∞ µ1 1−Hej0(s)−1 1−Hej(s)¶d(ˆ Hej10(s)−Hej10(s)). The proof of Corollary A.5 in Van Keilegom and Akritas (1999) can be adapted here to show that sup −∞<y≤T¯¯¯¯Zy −∞ µ1 1−Hej0(s)−1 1−Hej(s)¶d(ˆ Hej10(s)−Hej10(s))¯¯¯¯ =oP(n−1/2). (3.18) Indeed, equation (3.14) in the proof of Lemma 3.5 establishes that Hej0(y)− Hej(y) = O(n−1/2). Note that this order is not stochastic and better than the equivalent one needed in the above-mentioned proof. It suffices to follow the same steps to obtain (3.18). Hence the last term of the expression (3.17) is oP(n−1/2), and we obtain Zy −∞ d(ˆ Hej10(s)−Hej10(s)) 1−Hej0(s)=Zy −∞ d(ˆ Hej10(s)−Hej10(s)) 1−Hej(s)+oP(n−1/2).(3.19) Similarly to equation (3.14), it holds that Hej10(y) = ZHej1µyσj(x) + mn(x)−mjn(x) σj(x)¯¯¯¯ x¶fXj(x)dx, Comparison of regression curves with censored responses 89 and taking derivatives and a Taylor expansion of hej1around y hej10(y) = Zhej1µyσj(x) + mn(x)−mjn(x) σj(x)¯¯¯¯ x¶fXj(x)dx =hej1(y) + O(n−1/2). It follows that hej10(s) (1 −Hej0(s))2=hej1(s) (1 −Hej(s))2+hej10(s)µ1 (1 −Hej0(s))2−1 (1 −Hej(s))2¶ +1 (1 −Hej(s))2(hej10(s)−hej1(s)) =hej1(s) (1 −Hej(s))2+O(n−1/2), and since ˆ Hej0(y)−Hej0(y) = OP(n−1/2), we obtain Zy −∞ ˆ Hej0(s)−Hej0(s) (1 −Hej0(s))2dHej10(s) = Zy −∞ ˆ Hej0(s)−Hej0(s) (1 −Hej(s))2dHej1(s) + oP(n−1/2). (3.20) Using (3.19) and (3.20), as in the proof of Theorem 3.2, we have that ˆ Fεj0(y)−Fεj0(y) = (1 −Fεj0(y)) "Zy −∞ ˆ Hej0(s)−Hej0(s) (1 −Hej(s))2dHej1(s) + Zy −∞ d(ˆ Hej10(s)−Hej10(s)) 1−Hej(s)# +oP(n−1/2). From the proof of Theorem 3.2, we also have ˆ Fεj(y)−Fεj(y) = (1 −Fεj(y)) "Zy −∞ ˆ Hej(s)−Hej(s) (1 −Hej(s))2dHej1(s) + Zy −∞ d(ˆ Hej1(s)−Hej1(s)) 1−Hej(s)# +oP(n−1/2 j). 90 Chapter 3 Now write Fεj0(y) = PµYj−mn(Xj) σj(Xj)≤y¶ =PµYj−mjn(Xj) σj(Xj)−n−1/2R(Xj)−rj(Xj) σj(Xj)≤y¶ =ZPµYj−mjn(Xj) σj(Xj)−n−1/2R(Xj)−rj(Xj) σj(Xj)≤y¯¯¯¯ Xj=x¶fXj(x)dx. If we consider the probability inside the integral as a function of yand apply a Taylor expansion, we obtain Fεj0(y) = Fεj(y) + n−1/2fεj(y)E·R(Xj)−rj(Xj) σj(Xj)¸+o(n−1/2).(3.21) Straightforward calculations lead to ηj(Zj,∆j|Xj) = η0 j(Z0 j,∆j|Xj). Following the same steps as in the proof of Theorem 3.2, using Lemmas 3.5 and 3.6, and taking into account that Fεj0(y) = Fεj(y) + O(n−1/2), we obtain the expressions ˆ Fεj0(y)−Fεj0(y) (3.22) =1 nj nj X i=1 ξej(Eij,∆ij, y)−1 nj nj X i=1 (1 −Fεj(y))ζj(Zij,∆ij|Xij)γj2(y|Xij) −1 n k X l=1 nl X i=1 (1 −Fεj(y)) fXj(Xil) fmix(Xil) σl(x) σj(Xil)η0 l(Z0 il,∆il|Xil)γj1(y|Xil) + oP(n−1/2), and ˆ Fεj(y)−Fεj(y) (3.23) =1 nj nj X i=1 ξej(Eij,∆ij, y)−1 nj nj X i=1 (1 −Fεj(y))ζj(Zij,∆ij|Xij)γj2(y|Xij) −1 nj nj X i=1 (1 −Fεj(y))η0 j(Z0 ij,∆ij|Xij)γj1(y|Xij) + oP(n−1/2 j). Finally, by combining expressions (3.21), (3.22) and (3.23) we obtain the representation given in the statement of the Theorem. The leading term of the obtained representation does not depend on n, because the functions η0 jare defined in terms of distributions of random variables which do not depend on nand the functions Comparison of regression curves with censored responses 91 γj1are defined in terms of distributions of residuals Ej, which do not depend on neither. Proof of Theorem 3.5. The leading term of the representation given in Theorem 3.4 when working under Hl.a. n1/2 k X l=1 pl(n−1 l nl X i=1 ψ0 jl(Xil, Z0 il,∆il, y)) equals the leading term of the representation given under H0in Theorem 3.2 n1/2 k X l=1 pl(n−1 l nl X i=1 ψjl(Xil, Zil,∆il, y)), where m0in the first expression above plays the role of min the second one. Hence the asymptotic behavior is the same and the weak convergence follows immediately. Proof of Corollary 3.2. The convergence of the test statistics under the alternative hypothesis Hl.a. can be obtained in the same way as the proof of Corollary 3.1, simply by taking into account the weak convergence of the process established in Theorem 3.5. Chapter 4 Goodness-of-fit tests for parametric models in censored regression In this chapter we introduce a goodness-of-fit test for parametric regression models when the response variable is right censored. The test is based on the comparison of a parametric estimator and a nonparametric estimator of the distribution of the errors. Kolmogorov-Smirnov and Cram´er-von Mises type statistics are proposed. A bootstrap mechanism is used to approximate the critical values of the test. Some simulations are included and a real data set is analyzed. This chapter is based on Pardo-Fern´andez, Van Keilegom and Gonz´alez-Manteiga (2005b). 4.1 Introduction and statistical model As explained in Chapter 1, in many practical situations parametric regression models are appealing because they describe the relationship between the response and the covariate in a simple way and allow for interpretability of the parameters. Nevertheless, if the parametric model fails then the conclusions will be erroneous. Any parametric analysis should be accompanied by a test to check its validity and 93 94 Chapter 4 avoid misspecification and wrong conclusions. This motivates the development of specific goodness-of-fit tests for parametric models in regression. In the context of censored data the statistical model can be described as follows. Let (X, Y ) be a random vector, where Yrepresents a certain response variable associated with the covariate X. Assume that the response variable Yis subject to random right censoring. This means that there exists a censoring variable C, independent of Ygiven X, such that we observe Z= min{Y, C}and the indicator of censoring ∆ = I(Y≤C). Consider the following non-parametric regression model: Y=m(X) + σ(X)ε. The error variable εis independent of X,mis an unknown conditional location function m(x) = Z1 0 F−1(s|x)J(s)ds and σis an unknown conditional scale function representing possible heteroscedasticity σ2(x) = Z1 0 F−1(s|x)2J(s)ds −m2(x), where F(·|x) is the conditional distribution of Ygiven the value xof the covariate X,F−1(s|x) = inf{y;F(y|x)≥s}is the corresponding quantile function and J(s) is a score function satisfying R1 0J(s)ds = 1 (in general, for any distribution function Fwe denote F−1(s) = inf{y;F(y)≥s}for the corresponding quantile function and τF= inf{y;F(y) = 1}). Let Fεbe the distribution of the error ε. By construction R1 0F−1 ε(s)J(s)ds = 0 and R1 0F−1 ε(s)2J(s)ds = 1. The sample consists of nindependent replications (Xi, Zi,∆i), i= 1, . . . , n, from the distribution of (X, Z, ∆). We recall here that the choice of the function Jleads to different location and scale functions. An interesting choice is J(s)=(q−p)−1I(p≤s≤q), for some 0≤p < q ≤1, which leads to trimmed means and trimmed variances. Given a particular parametric class of regression functions M={mθ;θ∈Θ}, Goodness-of-fit tests in censored regression 101 (ii) mθ(x) is twice continuously differentiable with respect to θfor all x∈RX. Assuming (A5) and the representation given in (4.4), a Taylor expansion of mθ(u) as function of θaround θ0leads to mˆ θ(u)−mθ0(u) = n−1 n X i=1 ϕθ0(u, Xi, Zi,∆i) + oP(n−1/2), where ϕθ0(u, x, z, δ) = Ã∂mθ(u) ∂θ1¯¯¯¯θ=θ0 ,..., ∂mθ(u) ∂θp¯¯¯¯θ=θ0!~ω(x, z, δ).(4.5) Theorem 4.2 Assume (A1)-(A5). Then, under the null hypothesis H0, ˆ Fε0(y)−ˆ Fε(y) = (1 −Fε(y))n−1 n X i=1 ψθ0(Xi, Zi,∆i, y) + oP(n−1/2), uniformly in −∞ < y ≤T, where ψθ0(x, z, δ, y) = Zσ−1(u)ϕθ0(u, x, z, δ)fX(u)γ0(y|u)du +γ0(y|x)η(z, δ|x). Theorem 4.3 Assume (A1)-(A5). Then, under the null hypothesis H0, the process ˆ W(y) = n1/2(ˆ Fε0(y)−ˆ Fε(y)),−∞ < y ≤T, converges weakly to a centered Gaussian process W(y)with covariance function Cov(W(y), W(y0)) = (1 −Fε(y))(1 −Fε(y0))E(ψθ0(X, Z, ∆, y), ψθ0(X, Z, ∆, y0)). Corollary 4.1 Assume (A1)-(A5). Then, under the null hypothesis H0, TKS d →sup −∞<y<T |W(y)|, TCM d →ZT −∞ W2(y)dFε(y). 102 Chapter 4 4.4 Bootstrap approximation We propose a bootstrap procedure in order to approximate the critical values of the test in a practical situation. The resampling procedure is based on a smoothed version of the ‘naive bootstrap’ described in Efron (1981) and studied in Akritas (1986). For i= 1, . . . , n, estimate the censored residuals in a non parametric way ˆ Ei=Zi−ˆm(Xi) ˆσ(Xi). From the censored sample of estimated residuals ( ˆ Ei,∆i), i= 1, . . . , n, compute the Kaplan-Meier estimator ˆ Fεand standardize these residuals in order to verify the initial assumption of having location function 0 and scale function 1 (if λ1= Rˆ F−1 ε(s)J(s)ds and λ2= (Rˆ F−1 ε(s)2J(s)ds−λ2 1)1/2then the standardized residuals are ˜ Ei= ( ˆ Ei−λ1)/λ2). The bootstrap procedure we propose consists of the following steps. For fixed Band for b= 1, . . . , B, 1. For i= 1, . . . , n: •Let Y∗ i,b =mˆ θ(Xi) + ˆσ(Xi)ε∗ i,b, where ε∗ i,b =Vi,b +aSi,b,Vi,b is drawn from ˆ Fε(standardized), Si,b is a random variable with mean zero and variance one to introduce a small perturbation in the residuals (the perturbation is controlled by the constant a). Note that the bootstrap responses follow the null hypothesis by construction. •Select at random C∗ i,b from a smoothed version of ˆ G(·|Xi), which is the Beran estimator of the conditional distribution G(·|Xi) of the censoring variable obtained by replacing ∆iby 1−∆iin the expression of ˆ F(·|Xi). •Let Z∗ i,b = min{Y∗ i,b, C∗ i,b}and ∆∗ i,b =I(Y∗ i,b ≤C∗ i,b). 2. The bootstrap sample is {(Xi, Z∗ i,b,∆∗ i,b), i = 1, . . . , n}. 3. Let T∗ KS,b and T∗ CM,b be the test statistics calculated with the bootstrap sample. Goodness-of-fit tests in censored regression 103 Let T∗ KS,(b)be the b-th order statistic of T∗ KS,1, . . . , T∗ KS,B, and analogously for T∗ CM,(b). Then T∗ KS,([(1−α)B]) and T∗ CM,([(1−α)B]) approximate the (1 −α)-quantiles of the distribution of TKS and TCM under the null hypothesis respectively. 4.5 Some simulations We present some simulation results in order to study the finite-sample behavior of the goodness-of-fit test and the bootstrap approximation of the critical values. The regression and variance functions are those corresponding to the choice J(s) = 0.75−1I(0 ≤s≤0.75). We consider the following regression functions: (i)m(x) = x (ii)m(x) = x+ 0.6(x−0.5) (iii)m(x) = x+ 2x2 (iv)m(x) = x+ 0.5 sin(4πx) The variance function is σ2(x) = 0.5. The covariate is uniformly distributed in the interval [0,1] and the error is exponentially distributed, transformed such that R1 0F−1 ε(s)J(s)ds = 0 and R1 0F−1 ε(s)2J(s)ds = 1. The censoring variable is C=m(X) + σ(X)ρ, where ρis independent of εand has survival function 1−Fρ(y) = (1 −Fε(y))β, with β= 1/3 (25% of censoring) and β= 1 (50% of censoring). We will test for two different null hypotheses: a complete linear model H0:m(x) = θ1+θ2x (in this case models (i) and (ii) correspond to the null hypothesis and (iii) and (iv) correspond to the alternative hypothesis) and a linear model through the origin H0:m(x) = θx (in this case model (i) corresponds to the null hypothesis and models (ii)-(iv) correspond to the alternative hypothesis). The parameter is estimated by using the method proposed by Akritas (1996) for polynomial models described in Section 4.2. 104 Chapter 4 Table 4.1 shows the rejection probabilities in 1000 trials for sample sizes n= 100 and n= 200 and significance levels α= 0.05 and α= 0.10. We choose the Epanechnikov kernel K(u) = 0.75(1 −u2)I(|u|<1) to calculate the weights that appear in the Beran estimator. The bandwidth was chosen of the form h=cn−3/10 and the cases c= 0.75 and c= 1 are displayed. In the bootstrap we use B= 200 replications, a=n−3/10 and Si,b is a standard normal. The threshold Tin the definition of the test statistics was chosen to be the largest observed value of the sample of censored residuals estimated under the null hypothesis. The level is well approximated in most cases, although this approximation gets worse when the data are heavily censored (50% for censoring). The behavior of the power is as expected: it increases with the sample size and it decreases with the amount of censored data. Model (iii) is very difficult to distinguish from a linear model, especially when the amount of censored data is 50%. We believe that the choice of the threshold Tmay have an impact on the power. When the data are heavily censored, the Kaplan-Meier estimators have large jumps in the right tail of the distribution. This may produce large values of the test statistics even under the null hypothesis. In Table 4.2 we have repeated the same simulations by using as a threshold the quantile of order ˆ F−1 ε0(ˆ Fε0(+∞)−0.10) to avoid this problem. Now the level behaves reasonably well and the power is better than in the previous table, especially when the amount of censoring is 50% and the sample size is 200. 4.6 Data analysis We illustrate the proposed goodness-of-fit test on a data set concerning 90 male patients suffering from larynx cancer, diagnosed and treated during the period 1970-1978 in the Netherlands. More details about this data set can be found in Kardaun (1983). The variable of interest is the time between first treatment and death. At the end of the study 40 patients were alive (their survival times are censored). Heuchenne and Van Keilegom (2005) suggested a linear model to explain the relationship between the log of the age of the patient at diagnosis as Goodness-of-fit tests in censored regression 105 Table 4.1: Rejection probabilities under models (i) to (iv) of the tests based on TKS and TCM . The threshold Tis the largest observed value of the sample of censored residuals estimated under the null hypothesis. c= 0.75 c= 1 TKS TCM TKS TCM % Cens. nModel 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 H0:m(x) = θ1+θ2x 25% 100 (i) 0.054 0.098 0.049 0.097 0.043 0.086 0.042 0.096 (ii) 0.044 0.097 0.048 0.105 0.039 0.080 0.041 0.094 (iii) 0.124 0.202 0.130 0.220 0.099 0.165 0.122 0.210 (iv) 0.290 0.423 0.307 0.472 0.201 0.331 0.214 0.389 200 (i) 0.052 0.108 0.047 0.102 0.050 0.114 0.053 0.110 (ii) 0.044 0.110 0.047 0.106 0.038 0.096 0.051 0.106 (iii) 0.260 0.380 0.263 0.375 0.236 0.365 0.267 0.387 (iv) 0.727 0.838 0.760 0.877 0.604 0.761 0.723 0.848 50% 100 (i) 0.039 0.118 0.065 0.135 0.020 0.080 0.036 0.115 (ii) 0.038 0.103 0.055 0.136 0.030 0.086 0.042 0.116 (iii) 0.038 0.106 0.049 0.117 0.036 0.084 0.040 0.097 (iv) 0.066 0.174 0.070 0.194 0.050 0.124 0.046 0.126 200 (i) 0.027 0.110 0.055 0.125 0.021 0.070 0.045 0.118 (ii) 0.029 0.098 0.050 0.118 0.020 0.081 0.045 0.112 (iii) 0.040 0.108 0.063 0.137 0.023 0.087 0.045 0.114 (iv) 0.155 0.310 0.247 0.435 0.080 0.207 0.156 0.327 H0:m(x) = θx 25% 100 (i) 0.056 0.112 0.056 0.107 0.057 0.108 0.049 0.098 (ii) 0.311 0.427 0.356 0.496 0.283 0.382 0.336 0.466 (iii) 0.439 0.571 0.471 0.601 0.343 0.456 0.438 0.552 (iv) 0.256 0.398 0.148 0.293 0.225 0.385 0.132 0.279 200 (i) 0.065 0.118 0.071 0.121 0.061 0.114 0.060 0.117 (ii) 0.556 0.667 0.596 0.726 0.531 0.636 0.593 0.709 (iii) 0.749 0.827 0.743 0.824 0.688 0.765 0.732 0.807 (iv) 0.703 0.850 0.661 0.865 0.707 0.845 0.693 0.864 50% 100 (i) 0.041 0.103 0.060 0.152 0.030 0.078 0.056 0.117 (ii) 0.098 0.195 0.176 0.321 0.061 0.148 0.134 0.277 (iii) 0.130 0.242 0.215 0.391 0.061 0.154 0.176 0.331 (iv) 0.127 0.259 0.079 0.169 0.098 0.213 0.065 0.137 200 (i) 0.044 0.112 0.059 0.141 0.027 0.088 0.053 0.118 (ii) 0.166 0.295 0.335 0.499 0.135 0.248 0.307 0.471 (iii) 0.280 0.426 0.426 0.590 0.190 0.324 0.394 0.561 (iv) 0.327 0.527 0.356 0.564 0.274 0.457 0.359 0.551 106 Chapter 4 Table 4.2: Rejection probabilities under models (i) to (iv) of the tests based on TKS and TCM . The threshold Tis the quantile of order ˆ F−1 ε0(ˆ Fε0(+∞)−0.10) of the sample of censored residuals estimated under the null hypothesis. c= 0.75 c= 1 TKS TCM TKS TCM % Cens. nModel 0.050 0.100 0.050 0.100 0.050 0.100 0.050 0.100 H0:m(x) = θ1+θ2x 25% 100 (i) 0.063 0.110 0.051 0.107 0.053 0.108 0.054 0.108 (ii) 0.056 0.112 0.052 0.116 0.052 0.092 0.046 0.100 (iii) 0.142 0.226 0.152 0.241 0.127 0.206 0.152 0.241 (iv) 0.340 0.491 0.378 0.542 0.263 0.411 0.290 0.466 200 (i) 0.062 0.128 0.055 0.114 0.061 0.131 0.061 0.115 (ii) 0.057 0.130 0.051 0.115 0.061 0.114 0.058 0.109 (iii) 0.302 0.410 0.277 0.385 0.301 0.431 0.309 0.438 (iv) 0.768 0.868 0.789 0.898 0.704 0.823 0.782 0.886 50% 100 (i) 0.032 0.060 0.050 0.111 0.019 0.044 0.037 0.112 (ii) 0.034 0.064 0.052 0.114 0.029 0.055 0.043 0.090 (iii) 0.064 0.116 0.089 0.146 0.051 0.090 0.059 0.126 (iv) 0.192 0.333 0.196 0.361 0.112 0.230 0.106 0.258 200 (i) 0.034 0.071 0.055 0.109 0.036 0.072 0.050 0.101 (ii) 0.034 0.075 0.054 0.112 0.032 0.059 0.042 0.097 (iii) 0.114 0.204 0.125 0.226 0.090 0.172 0.120 0.196 (iv) 0.500 0.669 0.611 0.757 0.392 0.561 0.526 0.670 H0:m(x) = θx 25% 100 (i) 0.062 0.119 0.059 0.108 0.063 0.119 0.049 0.104 (ii) 0.321 0.447 0.370 0.513 0.299 0.406 0.352 0.484 (iii) 0.462 0.592 0.496 0.617 0.365 0.479 0.455 0.576 (iv) 0.266 0.421 0.166 0.317 0.250 0.416 0.158 0.319 200 (i) 0.069 0.123 0.070 0.122 0.066 0.118 0.062 0.118 (ii) 0.559 0.680 0.606 0.734 0.547 0.644 0.606 0.722 (iii) 0.766 0.840 0.753 0.833 0.709 0.787 0.740 0.817 (iv) 0.713 0.869 0.677 0.877 0.735 0.858 0.714 0.887 50% 100 (i) 0.049 0.093 0.072 0.140 0.037 0.072 0.054 0.111 (ii) 0.185 0.274 0.298 0.418 0.154 0.230 0.262 0.392 (iii) 0.254 0.369 0.365 0.501 0.157 0.256 0.326 0.466 (iv) 0.228 0.373 0.142 0.254 0.185 0.326 0.111 0.196 200 (i) 0.060 0.098 0.067 0.128 0.043 0.084 0.051 0.112 (ii) 0.393 0.519 0.514 0.654 0.336 0.461 0.496 0.638 (iii) 0.550 0.670 0.628 0.749 0.451 0.580 0.603 0.719 (iv) 0.612 0.782 0.594 0.777 0.598 0.747 0.614 0.779 Goodness-of-fit tests in censored regression 107 a covariate and the log of the survival time as response. These authors work with the conditional mean. Figure 4.1 shows the data and regression curves estimated nonparametrically with score functions J(s) = 0.50−1I(0.25 ≤s≤0.75) and J(s) = 0.75−1I(0 ≤s≤ 0.75). We believe that the choice of the score function is not crucial here since the data seem to be homoscedastic. The two estimated curves are almost parallel. We have applied our test to verify the claimed linear model with both choices of the function J. The obtained results were very similar, so we will only discuss the results corresponding to J(s) = 0.50−1I(0.25 ≤s≤0.75). We have performed the test over a wide range of bandwidths (from 0.15 to 0.35) and we have calculated the p-values based on 1000 bootstrap replications. The threshold Twas taken to be the quantile of order ˆ F−1 ε0(ˆ Fε0(+∞)−0.10). The Kolmogorov-Smirnov type statistic TKS produced p-values between 0.29 and 0.99. On the other hand, the Cram´er-von Mises type statistic TCM gave p-values between 0.27 and 0.95. The hypothesis of linearity can then be clearly accepted. Heuchenne and Van Keilegom (2005) also gave bootstrap confidence intervals for the parameters of the linear regression. The interval corresponding to the slope of the regression line contains zero. Hence it is reasonable to test for a constant model instead of the complete linear model. We have applied the goodness-of-fit test to check the constant model and the obtained p-values were between 0.12 and 0.73 for TKS and between 0.08 and 0.80 for TCM . It seems that the constant model can also be accepted in this example. All these results are summarized in Figure 4.2. 108 Chapter 4 log (age) log (survival time) 3.6 3.8 4.0 4.2 4.4 4.6 -3 -2 -1 0 1 2 3 Figure 4.1: Scatter plot of ‘log(survival time)’ versus ‘log(age)’ (crosses for uncensored data and circles for censored data) and estimated regression curves with J(s) = 0.50−1I(0.25 ≤s≤0.75) (solid line) and J(s) = 0.75−1I(0 ≤s≤0.75) (dashed line). h p-value 0.15 0.20 0.25 0.30 0.35 0.0 0.2 0.4 0.6 0.8 1.0 h p-value 0.15 0.20 0.25 0.30 0.35 0.0 0.2 0.4 0.6 0.8 1.0 Figure 4.2: Graphs of the p-values as function of the bandwidth hwhen testing for a linear model (left) and for a constant model (right) with the test statistics TKS (line with circles) and TCM (line with crosses). The score function is J(s) = 0.50−1I(0.25 ≤s≤0.75). The solid horizontal line corresponds to a p-value of 0.05. Goodness-of-fit tests in censored regression 109 4.7 Proofs Proof of Theorem 4.1. The direct implication is trivial. On the other hand, assume that there exists a θ0such that Fε0(y) = Fε(y). We can write PµY−mθ0(X) σ(X)≤y¶=PµY−m(X) σ(X)+m(X)−mθ0(X) σ(X)≤y¶, or equivalently Pµexp ½Y−mθ0(X) σ(X)¾≤y¶ =Pµexp ½Y−m(X) σ(X)¾exp ½m(X)−mθ0(X) σ(X)¾≤y¶, for all y. The residuals (Y−m(X))/σ(X) and the covariate Xare independent, hence the moments of the distributions above verify the relation E"µexp ½Y−mθ0(X) σ(X)¾¶2ν# =E"µexp ½Y−m(X) σ(X)¾¶2ν#E"µexp ½m(X)−mθ0(X) σ(X)¾¶2ν#, and hence E"µexp ½m(X)−mθ0(X) σ(X)¾¶2ν#= 1, for all ν∈N. Carleman’s condition (see e.g. Feller, 1966) ensures that Pµexp ½m(X)−mθ0(X) σ(X)¾= 1¶= 1 or Pµm(X)−mθ0(X) σ(X)= 0¶= 1. This and the continuity of mimplies the equality of m(x) and mθ0(x) for all x∈RX. Proof of Theorem 4.2. Since we are working under the null hypothesis, there exists θ0such that m=mθ0. From the proof of Proposition A.2 in Van Keilegom 110 Chapter 4 and Akritas (1999), we have that ˆ He0(y)−He0(y) =1 n n X i=1 I(Ei0≤y)−He0(y) (4.6) +Zhe0(y|x)mˆ θ(x)−mθ0(x) σ(x)fX(x)dx +Zyhe0(y|x)ˆσ(x)−σ(x) σ(x)fX(x)dx +oP(n−1/2), uniformly in −∞ < y ≤T. The last term is oP(n−1/2) because of the uniform consistency of mˆ θand ˆσ. The consistency of ˆσis given by Proposition 4.5 in Van Keilegom and Akritas (1999), and the consistency of mˆ θcan be obtained in a similar way. Define the class of functions MΘ(RX) = {x→(mθ(x)−m(x))/σ(x), θ ∈Θ}. Firstly, this class verifies P((mˆ θ(x)−m(x))/σ(x)∈MΘ(RX)) →1 as n→ ∞ because of the consistency of the parameter estimate. Secondly, the bracketing number N[ ](λ2, MΘ(RX), L2(P)) = O(λ−2p) for any λ > 0 because of the compactness of the parametric space Θ. This bracketing number is smaller than for the class C1+δ 1(RX) defined in Lemma A.1 in Van Keilegom and Akritas (1999). Then we can replace the class C1+δ 1(RX) by the class MΘ(RX) in that Lemma and this justifies expression (4.6). Using the expression (4.5), the first integral in (4.6) can be written as Zhe0(y|x)mˆ θ(x)−mθ0(x) σ(x)fX(x)dx =n−1 n X i=1 Zhe(y|u)σ−1(u)ϕθ0(u, Xi, Zi,∆i)fX(u)du +oP(n−1/2). From Proposition 4.9 in Van Keilegom and Akritas (1999) and a Taylor expansion, the second integral in (4.6) becomes Zyhe0(y|x)ˆσ(x)−σ(x) σ(x)fX(x)dx =−n−1 n X i=1 yhe0(y|Xi)ζ(Zi,∆i|Xi) + oP(n−1/2). References 117 Hall, P., Huber, C. and Speckman, P. L. (1997). Covariate-matched one-sided tests for the difference between functional means. J. Amer. Statist. Assoc. 92, 1074-1083. H¨ardle, W. (1990). Applied Nonparametric Regression. Cambridge University Press, Cambridge. H¨ardle, W. and Mammen, E. (1993). Comparing nonparametric versus parametric regression fits. Ann. Statist. 21, 1926-1947. H¨ardle, W. and Marron, J. S. (1990). Semiparametric comparison of regression curves. Ann. Statist. 18, 63-89. H¨ardle, W., M¨uller, M., Sperlich, S. and Werwatz, A. (2004). Nonparametric and Semiparametric Models. Springer, New York. Hart, J. D. (1997). Nonparametric Smoothing and Lack-of-fit Tests. Springer, New York. Heuchenne, C. and Van Keilegom, I. (2005). Polynomial regression with censored data based on preliminary nonparametric estimation. Conditionally accepted by Ann. Inst. Statist. Math. Kaplan, E. L. and Meier, P. (1958). Nonparametric estimation from incomplete observations. J. Amer. Statist. Assoc. 53, 457-481. Kardaun, O. (1983). Statistical survival analysis of male larynx-cancer patients – a case study. Statist. Neerlandica 37, 103-125. King, E. C., Hart, J D. and Wehrly, T. E. (1991). Testing the equality of two regression curves using linear smoothers. Statist. Probab. Lett. 12, 239-247. Koul, H. L. and Schick, A. (1997). Testing for the equality of two nonparametric regression curves. J. Statist. Plann. Inference 65, 293-314. Koul, H. L. and Schick, A. (2003). Testing for superiority among two regression curves. J. Statist. Plann. Inference 117, 15-33. Koziol, J. A. and Green, S. B. (1976). A Cram´er-von Mises statistic for randomly censored data. Biometrika 63, 465-474. 118 References Kulasekera, K. B. (1995). Comparison of regression curves using quasi-residuals. J. Amer. Statist. Assoc. 90, 1085-1093. Kulasekera, K. B. and Wang, J. (1997). Smoothing parameter selection for power optimality in testing of regression curves. J. Amer. Statist. Assoc. 92, 500-511. Kulasekera, K. B. and Wang, J. (2001). A test of equality of regression curves using Gˆateaux scores. Aus. N. Z. J. Stat. 43, 89-99. Mora, J. and Neumeyer, N. (2005). The two-sample problem with regression errors: an empirical process approach. Manuscript. Munk, A. and Dette, H. (1998). Nonparametric comparison of several regression functions: exact and asymptotic theory. Ann. Statist. 26, 2339-2368. Nadaraya, E. A. (1964). On estimating regression. Theory Probab. Appl. 9, 141-142. Neumeyer, N. and Dette, H. (2003). Nonparametric comparison of regression curves: an empirical process approach. Ann. Statist. 31, 880-920. Neumeyer, N. and Dette, H. (2005). A note on one-sided nonparametric analysis of covariance by ranking residuals. Math. Methods Statist. 14, 80-104. Pardo-Fern´andez, J. C. (2005). Comparison of error distributions in nonparametric regression. Manuscript. Pardo-Fern´andez, J. C. and Van Keilegom, I. (2005). Comparison of regression curves with censored responses. Conditionally accepted by Scand. J. Statist. Pardo-Fern´andez, J. C., Van Keilegom, I. and Gonz´alez-Manteiga, W. (2005a). Testing for the equality of kregression curves. Conditionally accepted by Statist. Sinica. Pardo-Fern´andez, J. C., Van Keilegom, I. and Gonz´alez-Manteiga, W. (2005b). Goodness-of-fit tests for parametric models in censored regression. Manuscript. Ramil-Novo, L. A. and Gonz´alez-Manteiga, W. (2000). F tests and regression analysis of variance based on smoothing spline estimators. Statist. Sinica 10, 819-837. References 119 Rao, C. R. (1965). Linear Statistical Inference and its Applications. Wiley, New York. Scheike, T. H. (2000). Comparison of non-parametric regression functions through their cumulatives. Statist. Probab. Lett. 46, 21-32. Seber, G. A. F. (1977). Linear Regression Analysis. Wiley, New York. Serfling, R. J. (1980). Approximation Theorems of Mathematical Statistics. Wiley, New York. Silverman, B. W. and Young, G. A. (1987). The bootstrap: To smooth or not to smooth? Biometrika 74, 469-479. Speckman, P. L., Chin, J. E., Hewett, J. E. and Bertelson, S. E. (2003). A onesided test adjusting for covariates by ranking residuals following smoothing. Manuscript. Stone, C. J. (1977). Consistent nonparametric regression. Ann. Statist. 5, 595645. Student (1908). The probable error of a mean. Biometrika 6, 1-25. Stute, W. (1997). Nonparametric model checks for regression. Ann. Statist. 25, 613-641. Stute, W. (1999). Nonlinear censored regression. Statist. Sinica 9, 1089-1102. Stute, W., Gonz´alez-Manteiga, W. and Presedo-Quindimil, M. (1998). Bootstrap approximations in model checks for regression. J. Amer. Statist. Assoc. 93, 141-149. Stute, W., Gonz´alez-Manteiga, W. and S´anchez-Sellero, C. (2000). Nonparametric model checks in censored regression. Comm. Statist. Theory Methods 29, 16111629. van der Vaart, A. W. and Wellner, J. A. (1996). Weak Convergence and Empirical Processes. Springer, New York. Van Keilegom, I. and Akritas, M. G. (1999). Transfer of tail information in censored regression models. Ann. Statist. 27, 1745-1784. 120 References Van Keilegom, I., Gonz´alez-Manteiga, W. and S´anchez-Sellero, C. (2005). Goodness-of-fit tests in parametric regression based on the estimation of the error distribution. Conditionally accepted by Test. Van Keilegom, I. and Veraverbeke, N. (1996). Uniform strong convergence results for the conditional Kaplan-Meier estimator and its quantiles. Commun. Statist. Theory Methods 25, 2251-2265. Van Keilegom, I. and Veraverbeke, N. (1997). Estimation and bootstrap with censored data in fixed design nonparametric regression. Ann. Inst. Statist. Math. 49, 467-491. Vilar-Fern´andez, J. M. and Gonz´alez-Manteiga, W. (2004). Nonparametric comparison of curves with dependent errors. Statistics 38, 81-99. Wand, M. P. and Jones, M. C. (1995). Kernel Smoothing. Chapman and Hall, London. Watson, G. S. (1964). Smooth regression analysis. Sankhy¯a Series A 26, 359-372. Ying, Z., Jung, S. H. and Wei, J. (1995). Survival analysis with median regression models. J. Amer. Statist. Assoc. 90, 178-184. Young, S. and Bowman, A. (1995). Nonparametric analysis of covariance. Biometrics 51, 920-931. Summary in Galician / Resumo en galego A regresi´on constit´ue un problema fundamental na estat´ıstica. Un modelo de regresi´on describe a relaci´on entre unha variable explicativa ou covariable Xe unha variable resposta Y. Nun contexto non param´etrico, a relaci´on entre YeX p´odese expresar como Y=m(X) + σ(X)ε, onde m´e unha funci´on de regresi´on suave, σ´e a funci´on de varianza, que representa posible heterocedasticidade, e ε´e o erro do modelo de regresi´on, que asumimos independente da covariable X. Observaremos datos do vector aleatorio (X, Y ). Se a variable resposta Yrepresenta un tempo, con frecuencia ocorre que os datos est´an incompletos debido a diversas raz´ons. Unha fonte importante de incompletitude ´e a censura. Isto significa que existe unha variable aleatoria C, chamada variable de censura, que pode ocultar a nosa variable de interese porque s´o podemos observar Ycando o seu valor ´e menor ou igual ca C. Dun xeito m´ais formal e nun contexto condicional, observamos o vector aleatorio (X, Z, ∆), onde Z= min{Y, C}´e o tempo observado asociado ´a covariable Xe ∆ = I(Y≤C) ´e o indicador de censura (1 representa un dato non censurado e 0 representa un dato censurado). Asumimos que a covariable non est´a censurada e que as variables resposta Ye de censura Cson independentes dada a covariable X. Habitualmente as funci´ons de regresi´on e de varianza son a media e a varianza condicionais respectivamente, pero tam´en se poden considerar outros funcionais. 121 122 Resumo en galego Se a funci´on de regresi´on ´e a media condicional ent´on m(x) = ZydF(y|x), onde F(·|x) ´e a funci´on de distribuci´on condicional da variable resposta Ydado o valor xda covariable X. Esta ´ultima expresi´on tam´en se pode escribir como m(x) = Z1 0 F−1(s|x)ds onde F−1(s|x) = inf{t;F(t|x)≥s}´e a correspondente funci´on cuantil condicional. Para estimar mabonda con dispo˜ner dun estimador da funci´on de distribuci´on condicional e incorporalo ´a expresi´on anterior. Cando a variable resposta ´e censurada pode ocorrer que as colas dereitas do estimador da funci´on de distribuci´on condicional deixen de ser consistentes debido ao propio mecanismo que xera a censura. Neste caso resulta ´util redefinir as funci´ons de regresi´on e de varianza do modelo de regresi´on como m(x) = Z1 0 F−1(s|x)J(s)ds e σ2(x) = Z1 0 F−1(s|x)2J(s)ds −m2(x), onde J´e unha funci´on tal que R1 0J(s)ds = 1. A elecci´on da funci´on Jcond´ucenos a distintas funci´ons condicionais de localizaci´on e de escala (que tam´en denominamos, nun senso m´ais amplo, funci´ons de regresi´on e de varianza). En particular, se J(s) = I(0 ≤s≤1) obtemos de novo a media e varianza condicionais. Unha elecci´on interesante ´e J(s) = (q−p)−1I(p≤s≤q), que corresponde a medias e varianzas recortadas. A mediana condicional ou outros cuant´ıs condicionais p´odense ver como l´ımites de medias recortadas. O obxectivo principal desta tese consiste en desenvolver contrastes de hip´oteses sobre a funci´on de regresi´on men diversos ´ambitos, tanto para datos completos coma para datos censurados. Os contrastes propostos est´an baseados na funci´on de distribuci´on dos erros do modelo de regresi´on, que ven dada por Fε(y) = P(ε≤y) = PµY−m(X) σ(X)≤y¶. Summary in Galician 123 A estimaci´on desta funci´on de distribuci´on implica a estimaci´on da funci´on de regresi´on e da funci´on de varianza condicional. No Cap´ıtulo 1 facemos unha breve introduci´on sobre os novos m´etodos desenvolvidos neste traballo e dos resultados obtidos. Realizamos tam´en unha revisi´on bibliogr´afica sobre comparaci´on de curvas de regresi´on e sobre contrastes de bondade de axuste para modelos param´etricos de regresi´on. Finalmente inclu´ımos un breve repaso das t´ecnicas de estimaci´on non param´etrica na regresi´on, tanto con datos completos coma con datos censurados. No Cap´ıtulo 2 propo˜nemos un novo m´etodo de comparaci´on de curvas de regresi´on desde un punto de vista totalmente non param´etrico. A comparaci´on de distintos grupos de individuos ou poboaci´ons ´e un problema com´un na estat´ıstica. Nun contexto incondicional (sen presenza de covariables), procedementos xa cl´asicos como o t-test ou a An´alise da Covarianza est´an relacionados con este problema. A existencia de covariables asociadas ´a variable de interese leva o problema a un contexto condicional. Agora o obxectivo consiste en comparar as correspondentes curvas de regresi´on para saber se o efecto das covariables sobre as variables resposta ´e o mesmo nos distintos grupos de individuos. Se as curvas de regresi´on est´an contidas nun modelo param´etrico ent´on o problema red´ucese ´a comparaci´on dos correspondentes par´ametros. Sen embargo, en moitas aplicaci´ons pr´acticas non conv´en asumir ning´un modelo param´etrico e a´ında as´ı a comparaci´on das curvas de regresi´on segue a ser un problema de interese. O problema de comparaci´on de curvas de regresi´on foi tratado amplamente na literatura ao longo dos ´ultimos quince anos, pero principalmente en situaci´ons m´ais restrictivas c´a considerada nesta tese. A maior´ıa das referencias existentes est´an dedicadas a contrastar a igualdade de soamente d´uas curvas de regresi´on e as s´uas extensi´ons para comparar m´ais de d´uas curvas non son doadas. En aplicaci´ons pr´acticas o problema da comparaci´on de m´ais de d´uas curvas pode aparecer de forma natural. Se a comparaci´on se fixese par a par ent´on tam´en se deber´ıa aplicar unha correcci´on no nivel e como consecuencia a potencia pode verse afectada. Isto motiva a implementaci´on de procedementos xerais para contrastar a 124 Resumo en galego igualdade de m´ais de d´uas curvas de regresi´on, tal e como facemos neste traballo. O modelo estat´ıstico considerado nesta tese ´e o que se describe a continuaci´on. Sexan kvectores aleatorios independentes (Xj, Yj) que satisfan os seguintes modelos de regresi´on non param´etricos, para j= 1, . . . , k, Yj=mj(Xj) + σj(Xj)εj, onde as covariables te˜nen soporte com´un, e o erro εj´e independente de Xje ten funci´on de distribuci´on Fεj. Sexa (Xij, Yij), i= 1, ..., nj, unha mostra de observaci´ons independentes e identicamente distribu´ıdas do vector aleatorio (Xj, Yj), para j= 1, . . . , k. Estamos interesados en contrastar a hip´otese nula de igualdade das funci´ons de regresi´on H0:m1=m2=··· =mk fronte ´a alternativa xeral Ha:mi6=mjpara alg´un i, j ∈ {1,2, . . . , k}. A idea do procedemento de contraste proposto e estudado nesta tese consiste en comparar en cada poboaci´on a distribuci´on emp´ırica dos residuos coa distribuci´on emp´ırica dos residuos estimados supo˜nendo que a hip´otese nula ´e certa. Fixemos unha poboaci´on, digamos j. Sexan ˆmje ˆσjestimadores non param´etricos da funci´on de regresi´on mje da funci´on de varianza σj(constru´ıdos coa mostra correspondente ´a poboaci´on j), e sexa ˆmun estimador da funci´on de regresi´on com´un baixo a hip´otese nula, que denotamos por m(constru´ıdo utilizando todas as mostras de maneira conxunta). Ent´on a cantidade (Yij −ˆmj(Xij))/ˆσj(Xij) estimar´a o erro εij e a cantidade (Yij −ˆm(Xij))/ˆσj(Xij) estimar´a o mesmo valor cando se aume que a hip´otese nula ´e verdadeira. Constru´ımos as distribuci´ons emp´ıricas correspondentes a estes residuos estimados ˆ Fεj(y) = 1 nj nj X i=1 IµYij −ˆmj(Xij) ˆσj(Xij)≤y¶, e ˆ Fεj0(y) = 1 nj nj X i=1 IµYij −ˆm(Xij) ˆσj(Xij)≤y¶. Summary in Galician 125 Cando a hip´otese nula de igualdade de curvas de regresi´on ´e verdadeira, as d´uas funci´ons de distribuci´on emp´ıricas anteriores estiman a correspondente funci´on de distribuci´on dos erros Fεjna poboaci´on j. Sen embargo, se a hip´otese nula non ´e certa, estas funci´ons emp´ıricas estiman distribuci´ons diferentes porque os residuos est´an medidos con respecto a curvas distintas (a verdadeira funci´on de regresi´on na poboaci´on je a funci´on de regresi´on com´un baixo a hip´otese nula). Polo tanto calquera diferenza entre estas d´uas funci´ons de distribuci´on emp´ıricas d´a evidencia en contra da hip´otese nula de igualdade das curvas de regresi´on. P´odese demostrar que a igualdade das curvas de regresi´on queda caracterizada pola igualdade das funci´ons de distribuci´on que se estiman coas distribuci´ons emp´ıricas anteriores. A comparaci´on l´evase a cabo a trav´es de estat´ısticos de contraste de tipo Kolmogorov-Smirnov e Cram´er-von Mises definidos sobre o proceso emp´ırico kdimensional ˆ W(y) = ( ˆ W1(y), . . . , ˆ Wk(y))t, onde ˆ Wj(y) = n1/2 j(ˆ Fεj0(y)−ˆ Fεj(y)) para j= 1, . . . , k. Os resultados te´oricos obtidos incl´uen unha representaci´on case segura para a diferenza entre os dous estimadores da distribuci´on dos erros ˆ Fεj0(y)−ˆ Fεj(y), a converxencia feble do proceso multidimensional ˆ W(y) e a converxencia en distribuci´on dos estat´ısticos de contraste. Ademais tam´en probamos que este procedemento de contraste pode detectar alternativas locais converxendo ´a hip´otese nula a unha taxa n−1/2, t´ıpica en procedementos param´etricos. Nas demostraci´ons destes resultados asint´oticos empr´egase, ademais de t´ecnicas com´uns en estat´ıstica non param´etrica, a teor´ıa dos bracketing numbers e a ferramenta de Cram´er-Wold. As distribuci´ons asint´oticas dos estat´ısticos de contraste resultan ser bastante complicadas. Como soluci´on para aproximar os valores cr´ıticos do contraste propo˜nemos un mecanismo de remostraxe bootstrap. A representaci´on case segura obtida para a diferenza das distribuci´ons emp´ıricas dos residuos incl´ue na s´ua expresi´on a densidade dos erros do modelo de regresi´on. Isto suxire que se debe empregar un m´etodo de remostraxe baseado nos residuos suavizados. Na construcci´on das mostras bootstrap util´ızase a funci´on de regresi´on estimada de maneira conxunta con todas as observaci´ons. Desta maneira as variables resposta verificar´an a 126 Resumo en galego hip´otese nula e os cuant´ıs dos estat´ısticos de contraste obtidos coas mostras bootstrap aproximar´an os correspondentes cuant´ıs das distribuci´ons dos estat´ısticos de contraste baixo a hip´otese nula. Nun estudo de simulaci´on amosamos o compartamento pr´actico do m´etodo de contraste e da aproximaci´on bootstrap dos valores cr´ıticos. Consideramos escenarios de simulaci´on con d´uas e tres curvas de regresi´on. Os resultados obtidos amosan un comportamento satisfactorio na maior´ıa dos casos considerados. A aproximaci´on do nivel ´e boa e a potencia mellora cando aumenta o tama˜no das mostras. Ademais obs´ervase que a elecci´on do par´ametro de suavizado que se precisa para a estimaci´on dos residuos (na estimaci´on das funci´ons de regresi´on e de varianza) non ten unha repercusi´on importante nas probabilidades de rexeitamento. Para ilustrar a metodolox´ıa proposta neste cap´ıtulo realizamos unha aplicaci´on a datos reais extra´ıdos do arquivo de datos do Journal of Applied Econometrics. Consid´eranse a relaci´on entre o gasto total e o gasto en comida dun conxunto de familias holandesas. Como covariable tomamos o logaritmo do gasto total mensual e como variable resposta consideramos o logaritmo do gasto mensual en comida. Realizamos o contraste de igualdade das correspondentes curvas de regresi´on distinguindo as familias polo n´umero de persoas que as compo˜nen. Tam´en no Cap´ıtulo 2 se estuda un m´etodo para comparar as distribuci´ons dos erros dos modelos de regresi´on en distintos grupos. Esta ´e unha suposici´on relativamente com´un nalg´uns m´etodos de comparaci´on de curvas de regresi´on existentes na literatura, a´ında que non ´e necesario para o noso. Sexa Fεja distribuci´on do erro εjna poboaci´on j. Agora queremos contrastar a hip´otese nula de igualdade das funci´ons de distribuci´on dos erros H0ε:Fε1=Fε2=. . . =Fεk fronte ´a alternativa xeral de que existe algunha diferenza entre estas distribuci´ons. Denotamos por Fεa funci´on de distribuci´on com´un cando se verifica a hip´otese