Upper bounds for the Euclidean distances between the BLUPs
Full text
Open Access. © 2018 Augustyn Markiewicz and Simo Puntanen, published by De Gruyter. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 License. Spec. Matrices 2018; 6:249–261 Research Article Open Access Special Issue on Linear Algebra and its Applications (ICLAA2017) Augustyn Markiewicz and Simo Puntanen* Upper bounds for the Euclidean distances between the BLUPs https://doi.org/10.1515/spma-2018-0020 Received December 19, 2017; accepted May 16, 2018 Abstract: In this article we consider the general linear model {y,Xβ,V}, where yis the observable random vector with expectation Xβand covariance matrix V. Our interest is on predicting the unobservable random vector y∗, which comes from y∗=X∗β+ε∗, where the expectation of y∗is X∗βand the covariance matrix of y∗ is known as well as the cross-covariance matrix between y∗and y. The random vector y∗can be considered as a kind of unknown future value. We introduce upper bounds for the Euclidean distances between the BLUPs, the best linear unbiased predictors, when the prediction is based on the original model and when it is based on the transformed model {Fy,FXβ,FVF}. We also show how the upper bounds are related to the concept of linear sufficiency, and we apply our results into the mixed linear model. Keywords: Best linear unbiased estimator, best linear unbiased predictor, Euclidean norm, linear sufficiency, transformed linear model. 1 Introduction Let us begin with some words about the notation. The symbol Rm×ndenotes the set of m×nreal matrices, while A,A−,A+,C(A), and C(A)⊥, denote, respectively, the transpose, a generalized inverse, the (unique) Moore– Penrose inverse, the column space, and the orthogonal complement of the column space of the matrix A.By (A:B) we denote the partitioned matrix with Aa×band Bc×das submatrices, where a=c. The symbol A⊥ stands for any matrix satisfying C(A⊥)=C(A)⊥. Furthermore, we will use PA=AA+=A(AA)−Ato denote the orthogonal projector (with respect to the standard inner product) onto the column space C(A), and QA=I−PA, where Irefers to the identity matrix of conformable dimension. In particular, we use notation M=In−PX, where Xn×prefers to the model matrix, see (1.1). One convenient choice for X⊥is obviously M; convenience follows from the symmetry and idempotence of M. We will consider the general linear model y=Xβ+ε, or shortly the triplet M={y,Xβ,V},(1.1) where X∈Rn×pis a known model matrix, the vector yis an observable n-dimensional random vector (socalled response vector), β∈Rpis vector of unknown parameters, and εis an unobservable vector of random errors with expectation E(ε)=0, and covariance matrix cov(ε)=V. Often the covariance matrix is of the type cov(ε)=σ2V, where σ2is an unknown nonzero constant. However, in our considerations σ2has no role and hence we omit it. The nonnegative definite matrix Vis known and can be singular. Let y∗denote a q× 1 unobservable random vector containing new future observations. The new observations are assumed to be generated from y∗=X∗β+ε∗, (1.2) Augustyn Markiewicz: Poznań University of Life Sciences, Poland, E-mail: [email protected] *Corresponding Author: Simo Puntanen: University of Tampere, Finland, E-mail: simo.[email protected] Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
250 | Augustyn Markiewicz and Simo Puntanen where X∗is a known q×pmatrix, β∈Rpis the same vector of fixed but unknown parameters as in M, and ε∗is a q-dimensional random error vector with E(ε∗)=0. We will also use the notation μ=Xβ,μ∗=X∗β. The covariance matrix of y∗and the cross-covariance matrix between yand y∗are assumed to be known and thus we have Ey y∗=μ μ∗=X X∗β,cov y y∗=covε ε∗=VV 12 V21 V22=Γ∈NNDn+q, (1.3) where NNDn+qrefers to the set of (n+q)×(n+q) nonnegative definite matrices. This setup can be denoted shortly as M∗=y y∗,X X∗β,VV 12 V21 V22.(1.4) We are particularly interested in predicting the unobservable y∗on the basis of the observable y. While doing this, we look for linear predictors of the type By, where B∈Rq×n, that would “conveniently” utilize the knowledge of (1.4). It is noteworthy that, literally taken, M∗in (1.4) is not a standard linear model—this is due to the fact that the y∗-part is not a proper response variable which in usual linear model is assumed to be observable. To clarify the situation, we will call M∗“the linear model with new future observations”. One of the first articles to consider the setup M∗was Goldberger [9, 1962], who assumed Γto be positive definite and y∗a scalar so that y∗=x ∗β+ε∗. Goldberger called x∗the vector “prediction regressors” and ε∗ the “prediction disturbance”. Premultiplying the model Mby an f×nmatrix Fyields the transformed model Fy =FXβ+Fε, or shortly Mt={Fy,FXβ,FVF}.(1.5) Suppose we wish to do the prediction using the transformed model Mt. Corresponding to M∗, we then have the following transformed setup: Mt∗=Fy y∗,FX X∗β,FVFFV12 V21FV22 .(1.6) We shall concentrate on the linear unbiased estimators, LUEs, and predictors, LUPs, and hence we need the concept of estimability. For example, X∗βis estimable under Mif there exists a matrix Bsuch that E(By)= X∗βfor all β∈Rp. Such a matrix B∈Rq×nexists only when C(X ∗)⊂C(X). The LUE By is the best linear unbiased estimator, BLUE, of estimable X∗βif By has the smallest covariance matrix in the Löwner sense among all linear unbiased estimators of X∗β: cov(By)≤ Lcov(B#y) for all B#:B#X=X∗,(1.7) that is, cov(B#y)−cov(By) is nonnegative definite for all B#:B#X=X∗. Correspondingly, the random vector y∗is called predictable under M∗if there exists a matrix Dsuch that the expected prediction error is zero, i.e., E(y∗−Dy)=0for all β∈Rp. Then Dy is a linear unbiased predictor (LUP) of y∗. Such a matrix D∈Rq×nexists if and only if C(X ∗)⊂C(X), that is, X∗βis estimable under M. Thus y∗is predictable under M∗if and only if X∗βis estimable. Now a LUP Dy is the best linear unbiased predictor, BLUP, for y∗, if we have the Löwner ordering cov(y∗−Dy)≤ Lcov(y∗−D#y) for all D#:D#X=X∗.(1.8) The Lemma 1.1 below provides so-called fundamental BLUEand BLUP-equations. For the BLUP, see, e.g., Christensen [6, p. 294], and Isotalo & Puntanen [18, p. 1015], and for the BLUE, Drygas [7, p. 55], Rao [28, p. 282], and Puntanen et al. [27, Th. 10]. For the reviews of the BLUP-properties, see, Robinson [31] and Haslett & Puntanen [14]. Lemma 1.1. Consider the linear model with new observations defined as M∗in (1.4), where C(X ∗)⊂C(X), i.e., y∗is predictable. Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
Upper bounds for the Euclidean distances between the BLUPs | 251 (a) The linear predictor Ay is the BLUP for y∗if and only if A∈Rq×nsatisfies the equation A(X:VX⊥)=(X∗:V21X⊥) . (1.9) (b) The linear estimator By is the BLUE of μ∗=X∗βif and only if B∈Rq×nsatisfies the equation B(X:VX⊥)=(X∗:0) . (1.10) In particular, Cy is the BLUE for μ=Xβif and only if C∈Rn×nsatisfies the equation C(X:VX⊥)=(X:0) . (1.11) (c) The linear predictor Dy is the BLUP for ε∗if and only if D∈Rq×nsatisfies the equation D(X:VX⊥)=(0:V21X⊥) . (1.12) We will use the following short notations: ˜ y∗= BLUP(y∗|M∗), ˜ μ∗= BLUE(μ∗|M∗), ˜ ε∗= BLUP(˜ ε|M∗) . (1.13) Notice that obviously BLUE(μ∗|M∗) = BLUE(μ∗|M). Lemma 2.2.4 of Rao & Mitra [30] appears very useful for our considerations. It says that for nonnull matrices Aand Cthe following holds: AB−C=AB+C⇐⇒ C(C)⊂C(B)& C(A)⊂C(B) . (1.14) In other words, (1.14) characerizes when the matrix product AB−Cis invariant with respect to the choice of B−. In particular, we observe the following: AB−B=AB+B⇐⇒ C(A)⊂C(B) . (1.15) One well-known solution for Cin (1.11) is PX;W−:= X(XW−X)−XW−, (1.16) where Wis a matrix belonging to the set of nonnegative definite matrices defined as W=W∈NNDn:W=V+XUUX,C(W)=C(X:V).(1.17) In (1.17) the matrix U(having prows) can be chosen arbitrarily subject to the condition C(W)=C(X:V). One obvious choice is Ipand we can choose U=0if C(X)⊂C(V). We could replace Wwith a set of matrices of the type W=V+XUX, where C(W)=C(X:V), U∈Rp×p, and thus Wwould not necessarily be nonnegative definite. However, to simplify our considerations, we will use (1.17) to define the set W. In view of (1.14), the matrix X(XW−X)−Xis invariant for any choices of the generalized inverses involved and the same concerns PX;W−yfor y∈C(X:V). For a review of the properties of W, see, e.g., Puntanen et al. [27, Sec. 12.3]. We assume the model Mto be consistent in the sense that the observed value of ylies in C(X:V) with probability 1. Hence we assume that under the model M, y∈C(X:V)=C(X:VX⊥)=C(X:VM)=C(X)⊕C(VM) , (1.18) where ⊕refers to the direct sum. For the equality C(X:V)=C(X:VM), see, e.g., Rao [29, Lemma 2.1]. The corresponding consistency as in (1.18) is assumed in all models that we will deal with. Let Aand Bbe m×nmatrices. Then, in the consistent linear model M, the estimators Ay and By are said to be equal with probability 1 if Ay =By for all y∈C(X:V) , (1.19) which will be a crucial property in our considerations. For the equality of two estimators, see, e.g., Groß & Trenkler [11]. Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
252 | Augustyn Markiewicz and Simo Puntanen As for the structure of this article, our main goal is to introduce upper bounds for the Euclidean distances between the best linear unbiased predictors (BLUPs) of y∗when the prediction is based on the original model M∗, defined in (1.4), and when it is based on the transformed model Mt∗, defined in (1.6). Corresponding considerations are made for the BLUPs of ε∗. Our attempt has been to make the paper self-readable so that the necessary background concepts, starting from BLUP and BLUE, have been presented to a reasonable amount. In Section 2 we introduce, review and comment on some presentations for the BLUPs under the original and the transformed model. Because the concept of linear sufficiency is strongly connected with the transformed model, we devote Section 3 for this topic. In Section 5 we apply our results into the mixed linear model which is a special case of the model with new future observations. Our considerations are rather mathematical and for this paper, we have no practical statistical applications in mind. 2 Representations for the BLUPs under the original and the transformed model If B1y= BLUE(X∗β) and B2y= BLUP(ε∗) under M∗, then, in view of Lemma 1.1, we have B1 B2(X:VM)=X∗0 0V 21M.(2.1) Premultiplying (2.1) by (Iq:Iq) leads to (B1+B2)(X:VM)=(X∗:V21M) , (2.2) and thereby (B1+B2)y= BLUP(y∗), i.e., BLUP(y∗) = BLUE(X∗β) + BLUP(ε∗) , or shortly, ˜ y∗=˜ μ∗+˜ ε∗.(2.3) Recall that equations of the type (2.3) hold “with probability 1”, that is, they hold for all y∈C(X:V). Now one solution for B2satisfying B2(X:VM)=(0:V21M)isB2=AM, where Asatisfies AMVM = V21M. Thus one expression for the BLUP of ε∗is BLUP(ε∗)=V21M(MVM)−My =V21 ˙ My ,(2.4) where we have denoted ˙ M=M(MVM)−M.(2.5) The matrix ˙ Mappears to be very useful in this context. As noted by Isotalo et al. [17, p. 1439], the matrix ˙ Mis not necessarily unique with respect to the choice of the generalized inverse (MVM)−. It is unique if and only if C(M)⊂C(MV), which further is equivalent to Rn=C(X:V). However, in view of (1.14) and the fact that y∈C(X:V), the expression V21 ˙ My is invariant with respect to the choice of (MVM)−. For the Moore–Penrose inversewehave M(MVM)+M=(MVM)+M=M(MVM)+=(MVM)+.(2.6) Hence in light of (1.14) and (2.6), BLUP(ε∗)=V21M(MVM)−My =V21(MVM)+yfor all y∈C(X:V). (2.7) For further properties of ˙ M, see Isotalo et al. [17] and Puntanen et al. [27, Ch. 17]. It is also of interest to substitute y=Xa +Vb (for some vectors aand b) into (2.4) and obtain BLUP(ε∗)=V21M(MVM)−M(Xa +Vb) =V21V+1/2PV1/2MV+1/2Vb =V21V+1/2PV1/2M0V+1/2Vb Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
Upper bounds for the Euclidean distances between the BLUPs | 253 =V21M0(M 0VM0)−M 0y,(2.8) where M0is any matrix satisfying C(M0)=C(M); here V+1/2 is the nonnegative definite square root of V+, and thereby V+1/2V1/2 =V1/2V+1/2 =PV. Observe that V21PV=V21 because C(V12)⊂C(V) in light of the nonnegative definiteness of Γin (1.3). In view of the identity, see Haslett et al. [13, Sec. 2], X(XW−X)−XW+=PW−VM(MVM)−MPW,(2.9) the residual of the BLUE(μ|M) can be written as y− BLUE(μ|M)=(In−G)y=VM(MVM)−My for all y∈C(W) , (2.10) and the BLUP(ε∗) can be expressed, for example, as follows: BLUP(ε∗)=V21M(MVM)−My =V21W−(In−G)y=V21V−(In−G)y, (2.11) where W∈W,y∈C(W), and G=X(XW−X)−XW−=PX;W−. For our purposes we assume that the parametric function μ∗=X∗βis estimable under Mas well as under Mt, which happens if and only if C(X ∗)⊂C(X)∩C(XF)=C(XF) so that X∗=LFX for some matrix L∈Rq×f,μ∗=X∗β=LFXβ=LFμ. (2.12) In other words, estimability under Mtimplies the estimability under M. The parametric function μ=Xβis of course always estimable under Mwhile under Mtit is estimable whenever C(X)=C(XF) , i.e., rank(X) = rank(FX) . (2.13) In light of (2.12), one expression for BLUE(X∗β) under Mis BLUE(X∗β|M) = BLUE(LFXβ|M)=LF BLUE(Xβ|M)=LFGy ,(2.14) where G=X(XW−X)−XW−=PX;W−. If Xβis estimable under Mt, then, in view of part (b) of Lemma 1.1, the statistic BFy is the BLUE for Xβ under Mtif and only if Bsatisfies B(FX :FVFQFX)=(X:0) . (2.15) Thus, see Kala et al. [19, Sec. 6] and Markiewicz & Puntanen [24, Sec. 3], the BLUE of Xβunder Mthas, for example, the representation BLUE(Xβ|Mt) = BLUE(μ|Mt)=˜ μt=Gty, (2.16) where Gt=X[XF(FWF)−FX]−XF(FWF)−F.(2.17) Correspondingly, BLUE(X∗β|Mt) = BLUE(LFXβ|Mt)=LFGty=˜ μt. (2.18) The BLUP of ε∗under Mt∗is CFy if and only if Csatisfies C(FX :FVFQFX)=(0:V21FQFX) , (2.19) that is, C=AQFX, where Asatisfies AQFXFVFQFX =V21FQFX . Thus one expression for BLUP of ε∗under Mt∗is BLUP(ε∗|Mt∗)=V21FQFX(QFXFVFQFX)−QFXFy .(2.20) According to Markiewicz & Puntanen [24, Sec. 2], we have C(FQFX)=C(F)∩C(M), MFQFX =FQFX .(2.21) Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
254 | Augustyn Markiewicz and Simo Puntanen Denoting N=PFQFX =PC(F)∩C(M), (2.22) and proceeding as in (2.8) we get the following expression: BLUP(ε∗|Mt∗)=V21N(NVN)−Ny.(2.23) Thus, see also Isotalo et al. [16, Sec. 4], the BLUP(y∗) under M∗can be written as BLUP(y∗|M∗) = BLUE(μ∗|M)+V21V−[y− BLUE(μ|M)] =LFGy +V21V−(In−G)y =LFGy +V21M(MVM)−My = BLUE(μ∗|M) + BLUP(ε∗|M∗) , (2.24) or shortly, ˜ y∗=˜ μ∗+˜ ε∗, and BLUP(y∗|Mt∗) = BLUE(μ∗|Mt)+V21F(FVF)−F[y− BLUE(μ|Mt)] =LFGty+V21F(FVF)−F(In−Gt)y =LFGty+V21N(NVN)−Ny = BLUE(μ∗|Mt) + BLUP(ε∗|Mt∗) , (2.25) or shortly, ˜ yt∗=˜ μt∗+˜ εt∗. Recall that N=PFQFX =PC(F)∩C(M)and that Nhas properties C(N)=C(FQFX)=C(F)∩C(M), N=MN =NM =N2.(2.26) Notice that the use of term BLUE(Xβ|Mt), as in the first two expressions in (2.25), requires, of course, that Xβis estimable under the transformed model Mt. The use of other expressions in (2.25) does not require this assumption; the estimability of X∗βunder Mtis only needed. We observe that the random vectors ˜ μ∗and ˜ ε∗are uncorrelated and the corresponding property holds also for ˜ μt∗and ˜ εt∗.Hencewehave cov(˜ y∗)=cov( ˜ μ∗)+cov( ˜ ε∗), cov( ˜ yt∗)=cov( ˜ μt∗)+cov( ˜ εt∗) . (2.27) Nowwehave˜ ε∗=V21M(MVM)−My, and ˜ εt∗=V21N(NVN)−Ny, with covariance matrices cov(˜ ε∗)=V21M(MVM)−MV12 , cov(˜ εt∗)=V21N(NVN)−NV12 .(2.28) Straightforward calculation shows that cov(˜ ε∗,˜ εt∗)=cov( ˜ εt∗) , and cov(˜ ε∗−˜ εt∗)=cov( ˜ ε∗)−cov( ˜ εt∗) , (2.29) and thereby we have the Löwner ordering cov(˜ ε∗)≥ Lcov(˜ εt∗). It is worth noting that for ˜ μ∗and ˜ μt∗we have the reverse Löwner ordering cov(˜ μ∗)≤ Lcov(˜ μt∗). 3 Conditions for linear sufficiency Consider the model M={y,Xβ,V}and let Fbe an f×nmatrix. Then Fy is called linearly sufficient (sometimes called BLUE-sufficient) for estimable X∗β, where X∗∈Rq×p, if there exists a matrix Aq×fsuch that AFy is the BLUE for X∗β. We use the notation Fy ∈S(X∗β) do indicate that Fy is linearly sufficient for X∗β. Let y∗be predictable under the model M∗, i.e., C(X ∗)⊂C(X). Then Fy is called linearly (prediction) sufficient (BLUP-sufficient) for y∗if there exists a matrix Aq×fsuch that AFy is the BLUP for y∗; that is, there exists Asuch that AF(X:VM)=(X∗:V21M) . (3.1) Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
Upper bounds for the Euclidean distances between the BLUPs | 255 If we want to emphasize that we are dealing with the prediction, we could talk about linear prediction sufficiency. We use the short notation Fy ∈S(y∗). The transformed model Mthas very strong connection with the concept of linear sufficiency. The equality of BLUEs under the original model and the transformed model can be characterized via the linear sufficiency and correspondingly the equality of the BLUPs. For the following Lemma 3.1 collects some useful results. For (a)–(c), see, e.g., Baksalary & Kala [3, 4], Drygas [8], Tian & Puntanen [32, Th. 2.8], and Kala et al. [20, Th. 2], and for (d)–(f), see Isotalo & Puntanen [18], and Isotalo et al. [16]. Lemma 3.1. Let μ∗=X∗βbe estimable under Mt(and thereby under M). Then the following statements are equivalent: (a) Fy is linearly sufficient for μ∗=X∗β, i.e., Fy ∈S(X∗β), (b) BLUE(X∗β|M∗) = BLUE(X∗β|Mt∗), or shortly, ˜ μ∗=˜ μt∗with probability 1, (c) cov(˜ μ∗)=cov( ˜ μt∗). Moreover, the following statements are equivalent: (d) Fy is linearly sufficient for y∗, i.e., Fy ∈S(y∗), (e) BLUP(y∗|M∗) = BLUP(y∗|Mt∗), or shortly, ˜ y∗=˜ yt∗with probability 1, (f) cov(˜ y∗−˜ yt∗)=0. The following lemma gives some BLUP-sufficiency properties of Fy for ε∗; see Isotalo et al. [16]. Lemma 3.2. The following statements are equivalent: (a) Fy is linearly sufficient for ε∗, i.e., Fy ∈S(ε∗), (b) C(MV12)⊂C(MVFQFX), (c) BLUP(ε∗|M∗) = BLUP(ε∗|Mt∗), or shortly, ˜ ε∗=˜ εt∗with probability 1, (d) cov(˜ ε∗)=cov( ˜ εt∗), (e) C(V12)⊂C(VN :X)=C(VFQFX :X), where N=PFQFX , (f) V21M=V21N(NVN)−NVM. Details of Lemma 3.2 are proved in Markiewicz & Puntanen [25] but let us take a brief look at the claim (f), which will be needed later on. To confirm (f), we can start from (c): BLUP(ε∗|M∗) = BLUP(ε∗|Mt∗) with probability 1, (3.2) i.e., V21M(MVM)−My =V21N(NVN)−Ny for all y∈C(X:VM) , (3.3) where N=PFQFX and Nhas properties like in (2.26). Choosing y∈C(X) yields zeros on both sides of (3.3). For y∈C(VM) the left-hand side of (3.3) becomes V21M(MVM)−MVM =V21M,(3.4) where we have used (1.15). Hence (3.3) can be expressed as V21M=V21N(NVN)−NVM =V21MN(NVN)−NVM := V21ME ,(3.5) where E=N(NVN)−NVM ∈Rn×n. 4 Some upper bounds for the Euclidean distance between the BLUPs In this section we provide new results giving upper bounds for the Euclidean norms of differences BLUP(ε∗|M∗) − BLUP(ε∗|Mt∗) and BLUP(y∗|M∗)−BLUP(y∗|Mt∗) . (4.1) Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
256 | Augustyn Markiewicz and Simo Puntanen For this purpose, let ch1(·) denote the largest eigenvalue of the matrix argument and let the matrix norm be defined as A2=ch1(AA). In particular, a2 2=aa, where a∈Rn. Denoting ˜ ε∗= BLUP(ε∗|M∗), ˜ εt∗= BLUP(εt∗|Mt∗), and N=PFQFX =MN ,wehave ˜ ε∗−˜ εt∗=V21[M(MVM)−M−N(NVN)−N]y=S0y, (4.2) where S0=V21[M(MVM)−M−N(NVN)−N] . (4.3) Putting y=Xa +VMb and denoting E=N(NVN)−NVM, gives, on account of (3.5), ˜ ε∗−˜ εt∗2 2=S0(Xa +VMb)2 2 =V21M[In−N(NVN)−NVM]Mb2 2 =V21M(In−E)Mb2 2 := SMb2 2 ≤S2 2Mb2 2 =ch 1(SS)bMb ,(4.4) where S=S0VM =V21M(In−E) . (4.5) The inequality in (4.4) follows from the consistency and multiplicativity of the matrix norm ·2; see, for example, Ben-Israel & Greville [5, pp. 19–20]. In view of part (f) of Lemma 3.2, we know that S=0if and only if Fy ∈S(ε∗). The scalar bMb is obviously zero if and only if b∈C(X). Thus we have proved the following theorem. Theorem 4.1. Consider the model M∗. Then for all y=Xa +VMb, ˜ ε∗−˜ εt∗2 2=BLUP(ε∗|M∗)−BLUP(ε∗|Mt∗)2 2 ≤ch 1(SS)bMb := α1,(4.6) where S=V21M(In−E)∈Rq×n,E=N(NVN)−NVM ∈Rn×n.(4.7) If b/ ∈C(X), then the upper bound α1in (4.6) is equal to zero if and only if S=0, i.e., Fy is linearly sufficient for ε∗. An alternative upper bound can be found out as follows: ˜ ε∗−˜ εt∗2 2=S0My2 2≤S02 2yMy =ch 1(S0S 0)yMy =ch 1(S0S 0)bMVMVMb := α2, (4.8) where S0=V21[M(MVM)−M−N(NVN)−N]. It is easy to conclude that yMy = 0 for all y∈C(X:V) if and only if VM =0, or, equivalently, C(V)⊂C(X). Groß [10, p. 317] calls a model with property VM =0adegenerated model. Remark 4.1. As one of the referees pointed out, the upper bounds in (4.6) and (4.8) depend on vector y, which is an arbitrary vector in Rnbelonging to the column space C(X:V). If VM =0, then α1=α2= 0, but of course the situation is somewhat pathological as yMy = 0 for all y∈C(X:V). If M∗is not a degenerated model, then the upper bound α2in (4.8) is equal to zero if and only if S0=V21[M(MVM)−M−N(NVN)−N]=0,(4.9) or, equivalently, S0S 0=V21[M(MVM)−M−N(NVN)−N]2V12 =0. (4.10) Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM
Upper bounds for the Euclidean distances between the BLUPs | 257 Of course, (4.9) implies V21[M(MVM)−M−N(NVN)−N]V12 = cov(˜ ε∗)−cov( ˜ εt∗)=0, (4.11) as well as S0VM =V21[M(MVM)−M−N(NVN)−N]VM =S=0, (4.12) which both are necessary and sufficient conditions for Fy being linearly sufficient for ε∗. Interestingly, but somewhat unwishfully, the linear sufficiency, i.e., (4.12), does not imply that α2= 0; exception for this is the case when C(X:V)=Rn. It remains an open question which upper bound α1or α2is sharper. Let us take a look at the Euclidean distance between the BLUEs of μ∗=X∗βin the original and the transformed model. As in Kala et al. [19, Sec. 6], we can observe that GtG=Gand hence (Gt−G)y=(Gt−GtG)y=Gt(In−G)y=GtVM(MVM)−My for all y∈C(W) , (4.13) where we have used (2.10). Then, for all y∈C(W), and μ∗=LFXβ,wehave ˜ μ∗−˜ μt∗2 2=LF(Gt−G)y2 2 =LFGtVM(MVM)−My2 2 ≤LFGtVM2 2(MVM)+2 2My2 2 =R2 2(MVM)+2 2My2 2 =a b2yMy := γ2,(4.14) where R=LFGtVM, the scalar ais the largest eigenvalue of RR, and bis the smallest nonzero eigenvalue of MVM. Moreover, if Mis not a degenerated model then γ2is zero if and only if Fy is linearly sufficient for X∗β. An alternative upper bound for ˜ μ∗−˜ μt∗2 2can be obtained by substituting y=Xa +VMb into (4.14). This yields ˜ μ∗−˜ μt∗2 2=LFGtVM(MVM)−MVMb2 2 =LFGtVMb2 2 ≤LFGtVM2 2bMb =R2 2bMb =abMb := γ1, (4.15) where a=ch 1(RR). Similarly as for α1and α2in (4.6) and (4.8), the question whether γ1is a sharper upper bound than γ2remains open. The BLUPs of y∗in the original and the transformed model, respectively, are BLUP(y∗|M∗)=LFGy +V21M(MVM)−My , or shortly, ˜ y∗=˜ μ∗+˜ ε∗, (4.16a) BLUP(y∗|Mt∗)=LFGty+V21N(NVN)−Ny , or shortly, ˜ yt∗=˜ μt∗+˜ εt∗. (4.16b) Putting y=Xa +VMb and using earlier notation, gives, in light of the triangle inequality, ˜ y∗−˜ yt∗2=BLUP(y∗|M∗)−BLUP(y∗|Mt∗)2 =(˜ μ∗−˜ μt∗)+( ˜ ε∗−˜ εt∗)2 ≤˜ μ∗−˜ μt∗2+˜ ε∗−˜ εt∗2 =RMb2+SMb2 =ch1(RR)bMb +ch1(SS)bMb =γ1+α1:= α. (4.17) We can now write the following theorem. Brought to you by | Tampere University Library Authenticated Download Date | 6/26/18 1:39 PM