scieee AI-readable full text Open interactive document viewer

Panel data estimation for correlated random coefficients models

Hsiao, Cheng,Li, Qi,Liang, Zhongwen,Xie, Wei

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Hsiao, Cheng; Li, Qi; Liang, Zhongwen; Xie, Wei Article Panel data estimation for correlated random coefficients models Econometrics Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Hsiao, Cheng; Li, Qi; Liang, Zhongwen; Xie, Wei (2019) : Panel data estimation for correlated random coefficients models, Econometrics, ISSN 2225-1146, MDPI, Basel, Vol. 7, Iss. 1, pp. 1-18, https://doi.org/10.3390/econometrics7010007 This Version is available at: https://hdl.handle.net/10419/247507 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ econometrics Article Panel Data Estimation for Correlated Random Coefficients Models Cheng Hsiao 1,2,*,†, Qi Li 3,†, Zhongwen Liang 4,† and Wei Xie 1,† 1Department of Economics, University of Southern California, Los Angeles, CA 90089, USA; [email protected] 2Department of Quantitative Finance, NTHU and WISE, Xiamen University, Xiamen 361005, China 3Department of Economics, Texas A&M University, College Station, TX 77843, USA; [email protected] 4Department of Economics, University at Albany, SUNY, Albany, NY 12222, USA; [email protected] *Correspondence: [email protected]; Tel.: +1-213-740-2103 † These authors contributed equally to this work. Received: 30 January 2018; Accepted: 23 January 2019; Published: 1 February 2019   Abstract: This paper considers methods of estimating a static correlated random coefficient model with panel data. We mainly focus on comparing two approaches of estimating unconditional mean of the coefficients for the correlated random coefficients models, the group mean estimator and the generalized least squares estimator. For the group mean estimator, we show that it achieves Chamberlain (1992) semi-parametric efficiency bound asymptotically. For the generalized least squares estimator, we show that when T is large, a generalized least squares estimator that ignores the correlation between the individual coefficients and regressors is asymptotically equivalent to the group mean estimator. In addition, we give conditions where the standard within estimator of the mean of the coefficients is consistent. Moreover, with additional assumptions on the known correlation pattern, we derive the asymptotic properties of panel least squares estimators. Simulations are used to examine the finite sample performances of different estimators. Keywords: panel data; correlated random coefficients; efficiency bound JEL Classification: C13; C33 1. Introduction One useful tool for reducing real-world details for econometric modeling is through “suitable” aggregations of micro data. For aggregation not to distort the fundamental behavioral relationships between the micro data and aggregate data, certain “homogeneity” conditions must hold between the micro units (e.g., Hsiao et al. 2005;Pesaran 2003;Stoker 1993;Theil 1954). However, the “homogeneity” assumption is often rejected by empirical investigators (e.g., Kuh 1963 ; Hsiao and Tahmiscioglu 1997 ). On the other hand, most policy makers are only interested in the average relationships of the population, not the individual relationship. Random coefficients formulation can be a useful tool to accommodate the “heterogeneity” among micro units and policy makers’ desire to find the average relationship (e.g., Hsiao et al. 1993). Standard random coefficients models assume the variation of coefficients are independent of the variation of regressors (e.g., Hsiao 1996;Hsiao and Pesaran 2008). In recent years, a great deal of attention has been devoted to the correlated random coefficients model. For instance, in the human capital literature, let the dependent variable y denote the logarithm of earnings and the explanatory variable x denote the years of schooling; the coefficient β denotes the rate of return. It is possible that the return to schooling declines with the level of schooling. It is also plausible that Econometrics 2019,7, 7; doi:10.3390/econometrics7010007 www.mdpi.com/journal/econometrics Econometrics 2019,7, 7 2 of 18 there are unmeasured ability or motivational factors that affect the return to schooling and are also correlated with the level of schooling (e.g., Card 1995;Heckman and Vytlacil 1998; Heckman et al. 2006 ; Heckman et al. 2010 ). Particularly, Heckman and Vytlacil (1998) propose an instrumental variable method for the population mean of slope coefficients but not the intercept in the cross sectional correlated random coefficients model. They require the existence of both instrumental variables for the regressors and random coefficients. Many people have worked on correlated random coefficients panel data models. For instance, Chamberlain (1992) showed how to apply his general result on the semiparametric efficiency bound to a random coefficients panel model as an example. The model considered in Chamberlain (1992) also allows for time varying parameters, which is more general than our model. However, the expression of efficient bound obtained using Chamberlain (1992)’s formulas are different from the expression obtained using the direct derivation. We show, in this paper, that they are, indeed, exactly the same. Due to the inclusion of the time varying parameters, Chamberlain (1992) requires the number of time periods T is greater than the number of random coefficients K . Otherwise, the information matrix of the time varying coefficients is singular. Graham and Powell (2012) further considered the situation when T=K , and proposed a novel “irregular” method that leads to consistent estimation. Their approach assumes the existence of panel data with two subpopulations, where one corresponds to units whose regressor values do not change across periods and the other changes across periods. Arellano and Bonhomme (2012) discuss the identification of the distribution of random coefficients conditional on the values of the regressors, extending the idea from Chamberlain (1992). Chernozhukov et al. (2013) consider more general nonseparable panel models that allow for correlated random coefficients model as a special case. In this paper, we consider the parametric identification and estimation of the unconditional mean of the random coefficients using panel data when the regularity conditions hold. Two approaches are considered; the approach of ignoring the correlations between the coefficients and regressors and the approach of explicitly modeling the correlations between the coefficients and regressors. The rest of the paper is organized as follows. We discuss the estimation of the unconditional mean of the random coefficients with panel data in Sections 2and 3. Section 2considers the approach without explicitly modeling the pattern of correlations. Section 3considers the approach with explicit assumption about the correlations between the coefficients and regressors. Section 4provides Monte Carlo results of the different estimators in a finite sample. Concluding remarks are in Section 5. 2. Panel Parametric Approaches without Explicit Assumption about the Correlations between Coefficients and Regressors When only cross-sectional data are available, the identification conditions of average effects for a correlated random coefficients model require the existence of instrumental variables, which are very stringent and may not be satisfied for many data sets. However, when panel data are available, it is possible to obtain a consistent estimator of the population mean of random coefficients without the existence of instrumental variables. Suppose there are T time series observations of (yit , xit)t=1,...,T for each individual i . Let yi and xi be T× 1 vector and T×K matrix with typical row elements given by yit and x> it = (xit,1 , . . . , xit,K) , respectively, for i=1, 2, . . . , N. Also, let βi= (βi1, . . . , βiK)>. We have yi=xiβi+ui,i=1, . . . , N. (1) Let ui= (ui1 , . . . , uiT)> , and we assume ui is iid across i , with E(ui|xi) = 0 and E(uiu> i|xi) = Σxi (a T×Tmatrix). We assume that βiis iid with mean βand variance Var(βi) = ∆. Then we can write βi=β+αi, (2) Econometrics 2019,7, 7 3 of 18 where E(αi) = E(βi−β) = 0, and Cov(βi,βj) = E(αiα> j) = (∆, if i=j, 0, if i6=j.(3) Substituting βi=β+αiinto (1) yields yi=xiβ+xiαi+ui=xiβ+vi, (4) where vi=ui+xiαi. The standard random coefficients model assumes that αi is a random draw from a population with E(αi|xi) = 0. Then E(vi|xi) = E(xiαi+ui|xi) = 0, (5) and E(viv> i|xi) = xi∆x> i+σ2 u. (6) Therefore, a consistent estimator of β can be obtained by simply regressing Y on X , where Y and X are of dimensions N× 1 and N×K , respectively. An efficient estimator of β can be obtained by applying the generalized least squares estimator (GLS) (or feasible GLS) (e.g., Hsiao 2003, chp. 6; Swamy 1970). When E(αi|xi) = 0 is violated, which is very common in practice, there exist the correlations between the coefficients and regressors, which is our main focus in the paper. We discuss different conditions and estimations in the following subsections. 2.1. Group Mean Estimator In this subsection we impose the following mild conditional moment restriction: E(ui|xi) = 0. (7) Note that (7) is weaker than E(ui|xi , βi) = 0 as we do not require that αi and ui are orthogonal with each other. Equation (7) implies the following unconditional moment condition E((x> ixi)−1x> iui) = 0. (8) When (x> ixi) is invertible (which requires that T≥K ), then from (1) one obtains (x> ixi)−1x> iui= (x> ixi)−1x> iyi−βi. Taking expectation yields the unconditional moment condition, E[(x> ixi)−1x> iyi−β] = 0. (9) Moment condition (9) leads to the estimator of βgiven by ˆ βGM =1 N N ∑ i=1 ˆ βi. (10) where ˆ βi= (x> ixi)−1x> iyi . Estimator (10) is the group mean (GM) estimator of Pesaran and Smith (1995)orHsiao et al. (1999). Under certain regularity conditions, we show that the GM estimator achieves the semiparametric efficiency bound derived in Chamberlain (1992). Note that (αi=βi−β) ˆ βi−β=αi+ (x> ixi)−1x> iui. (11) Econometrics 2019,7, 7 4 of 18 Then Var(ˆ βi) = Var(αi) + E[(x> ixi)−1x> iΣxixi(x> ixi)−1] + E[αiu> ixi(x> ixi)−1] + E[(x> ixi)−1x> iuiα> i] ≡Ω.(12) Particularly, in the uncorrelated case, we impose the restriction that E(ui|xi,αi) = 0. (13) Then the covariance term in (12) drops out. Moreover, if we further impose the conditional homoskedastic error assumption: Var(ui|xi) = Var(ui) = σ2 uIT. (14) Then Var(ˆ βi)is simplified to ∆+σ2 uE[(x> ixi)−1], (15) where ∆=Var(αi). The following proposition describes the asymptotic behavior of ˆ βGM. Proposition 1. If E(ui|xi) = 0and T ≥K, then (i) The group mean estimator defined in (10) is √N -consistent and asymptotically normally distributed, specifically, we have √N(ˆ βGM −β)d →N(0, Ω), (16) where Ωis defined in (12). (ii) ˆ βGM is semiparametrically efficient. (iii) If conditions (13) and (14) also hold, then the asymptotic variance Ω can be simplified to Ω=∆+ σ2 uE[(x> ixi)−1]. Proof. (i) ˆ βGM =β+1 N∑N i=1[αi+ (x> ixi)−1x> iui] . Hence √N(ˆ βGM −β) = 1 √N∑N i=1wi , where wi= αi+ (x> ixi)−1x> iui is i.i.d. with mean zero and finite variance Ω . Proposition 1(i) follows from the Lindeberg’s central limit theorem. (iii) follows from (i), (13) and (14) directly. We postpone the proof for (ii) to the Appendix A. Remark 1. Note that (16) holds without imposing any restriction on the correlations between xi and αi , and ui and αi . The random coefficient αi can be correlated with both xi and ui with arbitrary correlation patterns. Also, since xi can contain a constant (an intercept), the conventional fixed effects model is included in the correlated random coefficient model as a special case. 2.2. Generalized Least Squares Estimator In this subsection we consider a generalized least squares (GLS) estimator of β under the assumption that Cov(βi , xi) = 0 and compare the relative efficiency of the group mean estimator and the GLS estimator. Under the assumption that E(αi|xi) = 0. (17) and the assumptions of (13) and (14), i.e., E(ui|xi , αi) = 0 and Var(ui|xi) = σ2 uIT , then the best linear unbiased estimator (BLUE) of βis the generalized least squares estimator (e.g., Hsiao 2003, chp. 6): Econometrics 2019,7, 7 5 of 18 ˆ βGLS = N ∑ i=1 x> i(xi∆x> i+σ2 uIT)−1xi!−1 N ∑ i=1 x> i(xi∆x> i+σ2 uIT)−1yi! = N ∑ i=1 Wiˆ βi, (18) where Wi=(N ∑ i=1 [∆+σ2 u(x> ixi)−1]−1)−1 [∆+σ2 u(x> ixi)−1]−1 is a positive definite weight matrix satisfying ∑N i=1Wi=1. Contrary to the group mean estimator (10) that takes the simple average of the individual least squares estimator, ˆ βi, the GLS estimator takes the weighted average of ˆ βi. By noting that yi=xiβi+ui=xiβ+ei , where ei=xiαi+ui , we define Y(NT)×1= (y> 1 , ..., y> N)> , X(NT)×K= (x> 1 , ..., x> N)> and e(NT)×1= (e> 1 , ..., e> N)> . Then we have Var(e) = Ω(NT)×(NT)=Block − diag(Σi)is a block-diagonal matrix with the ith block diagonal element Σi=xi∆x> i+σ2 uIT. Hence, ˆ βGLS = (X>Ω−1X)−1X>Ω−1Y = N ∑ i=1 x> i(xi∆x> i+σ2 uIT)−1xi!−1 N ∑ i=1 x> i(xi∆x> i+σ2 uIT)−1yi! =β+ N−1N ∑ i=1 x> i(xi∆x> i+σ2 uIT)−1xi!−1 N−1N ∑ i=1 x> i(xi∆x> i+σ2 uIT)−1ei!. (19) Then by the law of large numbers and the central limit theorem arguments and by noting that Var(ei|xi) = xi∆x> i+σ2 uIT, we obtain √N(ˆ βGLS −β)d →AN(0, A−1) = N(0, A), (20) where A={E[x> i(xi∆x> i+σ2 uIT)−1xi]}−1. (21) Clearly, ˆ βGLS is not feasible. A feasible GLS estimator is given with ∆and σ2 ureplaced by ˆ ∆=N−1N ∑ i=1 (ˆ βi−ˆ βGM)( ˆ βi−ˆ βGM)>−ˆ σ2 u N N ∑ i=1 (x> ixi)−1, ˆ σ2 u=N−1(T−K)−1N ∑ i=1 T ∑ t=1 (yit −x> it ˆ βi)2, where ˆ βi is given in (10). The consistency of ˆ ∆ and ˆ σ2 u can be proved similarly as that of ˆ ∆∗ and ˜ σ2 u in the Appendix A. Remark 2. Note that without condition that E(αi|xi) = 0of (17), ˆ βGM is still a root-N consistent estimator for β as shown in Proposition 1while ˆ βGLS becomes inconsistent when T is finite due to E(ei|xi) = xiE(αi|xi)6= 0. However, when T is large, by noting that x> ixi/T=E(xitx> it ) + Op( 1 /√T) under the strong mixing condition and the conditions given in Theorem 24.6 in Davidson (1994), the weight matrix Wi is close to a constant matrix: Econometrics 2019,7, 7 6 of 18 Wi=(1 N N ∑ i=1 [∆+1 Tσ2 u(x> ixi/T)−1]−1)−11 N[∆+1 Tσ2 u(x> ixi/T)−1]−1 =(1 N N ∑ i=1 [∆+1 Tσ2 u[E(xitx> it )]−1]−1)−11 N[∆+1 Tσ2 u[E(xitx> it )]−1]−1+Op1 NT3/2  ≡1 N+Op1 NT . (22) It is easy to see that ˆ βGLS =N−1∑N i=1ˆ βi+Op((NT)−1∑N i=1ˆ βi) is a consistent estimate for β=E(βi) . The next proposition compares the relative efficiency of ˆ βGLS and ˆ βGM by comparing their asymptotic variances: Avar(√Nˆ βGLS) = {E[x> i(xi∆x> i+σ2 uIT)−1xi]}−1 and Avar(√Nˆ βGM) = Ω= ∆+σ2 uE[(x> ixi)−1]. Proposition 2. Assuming that T is small (but still T≥K ) and that conditions (13), (14) and (17) hold. Then Avar(√Nˆ βGLS)≤Avar(√Nˆ βGM). The proof of Proposition 2is given in the Appendix A. Proposition 2says that, under some additional assumptions, ˆ βGLS is asymptotically more efficient than ˆ βGM . This is in no contradiction with Proposition 1(ii) because the result of Proposition 1does not require any of the conditions (13), (14) and (17) to hold. With additional conditions, ˆ βGM is no longer a semiparametric efficient estimator of β . However, these additional conditions, especially condition (17), are quite restrictive. It was shown by Hsiao et al. (1999) that when T is large, ˆ βGLS becomes a consistent estimator for βwithout needing the restrictive condition (17). Proposition 3. Under conditions (13) and (14), if both N , T→∞ and N1/2/T→ 0, √N(ˆ βGLS −ˆ βGM) = op(1). In other words, if both N and T are large and if limN,T→∞(N1/2/T)→ 0, contrary to the case of only cross-sectional data are available, one can ignore the issue of possible correlations between αi and xi (i.e., we allow for E(αi|xi)6= 0) and simply treat the model as if βi and xi are uncorrelated and apply the conventional GLS (e.g., Hsiao 2003, eq. (6.2.6)). 2.3. Within Estimator If T<K , neither the GM, nor the GLS can be implemented. However, the standard within estimator can still yield a consistent estimator of β in certain cases. Let ¯ yi·=1 T∑tyit and ¯ xi·=1 T∑txit . The within estimator (or fixed effects estimator) first takes the deviation of each observation from its time series mean, then regress (yit −¯ yi·)on (xit −¯ xi·)(e.g., Hsiao 2003, chp. 3)). Model (1) leads to (yit −¯ yi·) = (xit −¯ xi·)>β+ (xit −¯ xi·)>αi+ (uit −¯ ui·),i=1, . . . , N,t=1, . . . , T, (23) where ¯ ui·=T−1∑tuit. The fixed effects (FE) estimator of βis the least squares estimator of (23). Econometrics 2019,7, 7 7 of 18 ˆ βFE ="∑ i ∑ t (xit −¯ xi·)(xit −¯ xi·)>#−1"∑ i ∑ t (xit −¯ xi·)(yit −¯ yi·)# =β+"∑ i ∑ t (xit −¯ xi·)(xit −¯ xi·)>#−1 ×"∑ i ∑ t (xit −¯ xi·)(xit −¯ xi·)>αi+∑ i ∑ t (xit −¯ xi·)(uit −¯ ui·)#. (24) In general, (24) is inconsistent. However, if the data generating process of xit takes the form: xit =µi+∑ j Bjei,t−j,∑kBjk<∞, (25) where µiis iid with mean aand variance Σµand eit is iid across both iand over twith E(eit|αi) = E(eis|αi)≡difor all t,s=1, . . . , T. (26) Then (23) is consistent. To see this, note that under (25) and (26), we have E(xit|αi) = E(µi|αi) + ∑ j BjE(ei,t−j|αi) = δi+di∑ j Bj≡µ∗ i, (27) where δi=E(µi|αi)and di=E(ei,t−j|αi). Let xit =E(xit|αi) + ηit ≡µ∗ i+ηit, (28) where ηit =xit −µ∗ i= (µi−E(µi|αi)) + ∑jBj(ei,t−j−E(ei,t−j|αi)) . Then xit −¯ xi·=ηit −¯ ηi· , where ¯ mi·=1 T∑T t=1mit ( m can be x or η ). Also, from E(ηit −¯ ηi·|αi) = 0, we know that E(xit −¯ xi·|αi)≡ E(ηit −¯ ηi·|αi) = 0. If in addition the following conditional homoskedastic error assumption holds: E[(xit −¯ xi·)(xit −¯ xi·)>|αi] = E[(ηit −¯ ηi·)(ηit −¯ ηi·)>|αi] = C, (29) where C=E[(ηit −¯ ηi·)(ηit −¯ ηi·)>]is a K×Knonsingular constant matrix. Then 1 NT ∑ i ∑ t (xit −¯ xi·)(xit −¯ xi·)>αi p →C E[αi] = 0 (30) as N→∞. Therefore Proposition 4. Under (25) and (29), the conventional fixed effects estimator is √N -consistent and asymptotically normally distributed as N→∞ . The asymptotic covariance matrix of (24) can be approximated using Newey-West heteroscedasticity-autocorrelation consistent formula. When (xit , αi) has a joint elliptical distribution, the conditional homoscedasticity of E[(xit − ¯ xi·)(xit −¯ xi·)>|αi] = Calso holds (e.g., Fang and Zhang 1990;Gupta et al. 1993). Therefore, Proposition 5. When (xit , αi) are jointly elliptically distributed, or conditional homoscedasticity of (xit −¯ xi·) of (29) holds, the FE estimator (24) is √N-consistent and asymptotically normally distributed. Another case where the fixed effects estimator can be consistent is that (xit , αi) are jointly symmetrically distributed. Since xit −¯ xi· has mean equal to zero, (xit −¯ xi· , αi) will be symmetrically distributed around ( 0, 0 ) , then 1 NT ∑i∑t(xit −¯ xi·)(xit −¯ xi·)>αi p → 0 even though xit has mean different from zero. We have Econometrics 2019,7, 7 8 of 18 Proposition 6. Under (25), (26) and if (xit , αi) are symmetrically distributed, the fixed effects estimator (24) is √N consistent and asymptotically normally distributed. Wooldridge (2005) also discussed conditions for the validity of the fixed effects estimator. Although the conventional FE estimator (24) can yield a consistent estimator of β , if xit contains time-invariant variables, the mean effects of time-invariant variables cannot be identified by the conventional fixed effects estimator. Moreover, the FE estimator only makes use of the within (group) variation. Since, in general, the between group variation is much larger than within group variation, the FE estimator could also mean a loss of efficiency. 3. Panel Least Squares or Generalized Least Squares Estimator If αiis correlated with xi, i.e., E(αi|xi)6=0, we can re-write (1) as yit =x> it β+x> it E(αi|xi) + vit, (31) where vit =uit +x> it wi , wi=αi−E(αi|xi) . Equation (31) is no longer a linear function of xit . For instance, suppose E(αi|xi) = a+Bvec(x> i), (32) as assumed by Mundlak (1978). Noting that E(αi) = E[E(αi|xi)] = a+BE(vec(x> i)) = 0, (33) which implies that a=−BE(vec(x> i)). Hence, (32) can be written as E(αi|xi) = B(vec(x> i)−E(vec(x> i))). (34) Equation (31) then becomes yit =x> it β+x> it Bvec(x> i−E(x> i)) + vit,i=1, . . . , N,t=1, . . . , T. (35) Let vi= (vi1 , . . . , viT)> . Then E(vi|xi) = 0 and E(viv> i) = E(xi∆∗x> i) + σ2 uIT where ∆∗= E(wiw> i) = E(wiw> i|xi) . Therefore, the least squares or the generalized least squares estimator of β is √Nconsistent provided 1 N N ∑ i=1 E x> ixix> i(xi⊗(vec((xi−¯ x·)>))>) (x> i⊗vec((xi−¯ x·)>))xi(x> i⊗vec((xi−¯ x·)>))(xi⊗(vec((xi−¯ x·)>))>)!(36) is a full rank matrix, where ¯ x·=N−1∑N i=1xi. Similar reasoning can be applied if E(αi|xi)is a higher order polynomial of xi, say, E(αi|xi) = a+Bvec(x> i) + Cvec(x> i⊗x> i). (37) Then from E[E(αi|xi)] = 0, we get a=−BE(vec(x> i)) −CE[vec(x> i⊗x> i)], it follows that E(αi|xi) = B[vec(x> i−E(x> i))] + C[vec(x> i⊗x> i−E(x> i⊗x> i))]. (38) Substituting (38) into (31) we have yit =x> it β+ (x> it ⊗(vec((xi−E(xi))>))>)vec(B>) + (x> it ⊗(vec([(xi⊗xi)−E(xi⊗xi)]>))>)vec(C>) + vit,(39) Econometrics 2019,7, 7 15 of 18 Proof of Proposition 7. First, we show that ˜ σ2 u and ˆ ∆∗ are consistent estimators of σ2 u and E(∆∗) , respectively. Let ˜ uit =yit −x> it ˆ βGM, then it can be shown that ˜ σ2 GM = (NT)−1N ∑ i=1 T ∑ t=1 ˜ u2 it = (NT)−1N ∑ i=1 T ∑ t=1 [x> it (βi−ˆ βGM) + uit]2 = (NT)−1N ∑ i=1 T ∑ t=1 [x> it (βi−ˆ βGM)(βi−ˆ βGM)>xit +u2 it] + op(1) = (NT)−1N ∑ i=1 T ∑ t=1 [x> it (βi−β)(βi−β)>xit +u2 it] + op(1)(A4) = (T)−1T ∑ t=1 E[x> it (βi−β)(βi−β)>xit] + σ2 u+op(1) = (T)−1T ∑ t=1 E{x> it E[(βi−β)(βi−β)>|X]xit}+σ2 u+op(1) = (T)−1T ∑ t=1 E{x> it ∆xit}+σ2 u+op(1) where the first op( 1 ) term comes from E(ui|xi) = 0, and others used ˆ βGM =β+Op(N−1/2) and also the law of large numbers. Furthermore, we have ˜ VGM =N−1N ∑ i=1 (ˆ βi−ˆ βGM)( ˆ βi−ˆ βGM)> =N−1N ∑ i=1 [βi+ (x> ixi)−1x> iui−ˆ βGM][βi+ (x> ixi)−1x> iui−ˆ βGM]> =N−1N ∑ i=1 [βi−β+ (x0 ixi)−1x> iui][βi−β+ (x> ixi)−1x> iui]>+op(1)(A5) =N−1N ∑ i=1{(βi−β)(βi−β)>+ (x> ixi)−1x> iu2 ixi(x> ixi)−1] + op(1) =E[(βi−β)(βi−β)>] + E[(x> ixi)−1x> iE(u2 i|X)xi(x> ixi)−1] + op(1) =∆+σ2 uE[(x> ixi)−1] + op(1) = Var[E(αi|xi)] + E[∆∗] + σ2 uE[(x> ixi)−1] + op(1). Combining (A4) and (A5), we have ˜ σ2 u p →σ2 u,ˆ ∆∗p →E(∆∗). Econometrics 2019,7, 7 16 of 18 Next, we look at ˆ βM,PLS. We have ˆ βM,PLS = N ∑ i=1 x> iˆ Σ−1 ixi− N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1 × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 ixi!−1 N ∑ i=1 x> iˆ Σ−1 iyi− N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 iyi! =β+ N ∑ i=1 x> iˆ Σ−1 ixi− N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1 × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 ixi!−1 N ∑ i=1 x> iˆ Σ−1 i[(xi⊗(vec((xi−E(xi))>))>)vec(B>)] − N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1 (A6) × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i[(xi⊗(vec((xi−E(xi))>))>)vec(B>)]! + N ∑ i=1 x> iˆ Σ−1 ixi− N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1 × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 ixi!−1 N ∑ i=1 x> iˆ Σ−1 ivi− N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 ivi! =β+B vec(¯ x> ·−E(x> i)) + N ∑ i=1 x> iˆ Σ−1 ixi− N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1 × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 ixi!−1 N ∑ i=1 x> iˆ Σ−1 ivi− N ∑ i=1 x> iˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>) × N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 i(xi⊗(vec((xi−¯ x·)>))>)!−1N ∑ i=1 (xi⊗(vec((xi−¯ x·)>))>)>ˆ Σ−1 ivi!. Therefore, VM,PLS =lim N→∞Var(√Nˆ βM,GLS) = Var[E(αi|xi)] + E[x> iΣ−1 ixi]−E[x> iΣ−1 i(xi⊗(vec((xi−Exi)>))>)] (A7) ×E[(xi⊗(vec((xi−Exi)>))>)>Σ−1 i(xi⊗(vec((xi−Exi)>))>)]−1E[(xi⊗(vec((xi−Exi)>))>)>Σ−1 ixi]!−1 , where Σi=xi∆∗x> i+σ2 uIT,∆∗=E(wiw> i)and wi=αi−E(αi|xi), which implies √N(ˆ βM,PLS −β)d →N(0, VM,PLS). This completes the proof of Proposition 7. Moreover, let M∗ i= (x> iΣ−1 ixi)−1= ( x> i(xi∆∗x> i+σ2 uIT)−1xi)−1 . Similar as that in the proof of Proposition 2(i) above, we have M∗ i=∆∗+σ2 u(x> ixi)−1. Econometrics 2019,7, 7 17 of 18 References Arellano, Manuel, and Stéphane Bonhomme. 2012. Identifying distributional characteristics in random coefficient panel data models. Review of Economic Studies 79: 987–1020. [CrossRef] Card, David. 1995. Using geographic variation in college proximity to estimate the return to schooling. In Aspects of Labour Market Behavior: Essays in Honour of John Vanderkamp. Edited by Christofides Loizos Nicolaou, E. Kenneth Grant and Robert Swidinsky. Toronto: University of Toronto Press, pp. 201–22. Chamberlain, Gary. 1992. Efficiency bounds for semiparametric regression. Econometrica 60: 567–96. [CrossRef] Chernozhukov, Victor, Iván Fernández-Val, Jinyong Hahn, and Whitney Newey. 2013. Average and quantile effects in nonseparable panel models. Econometrica 81: 535–80. Davidson, James. 1994. Stochastic Limit Theory: An Introduction for Econometricians. New York: Oxford University Press. Fang, Kai-Tai, and Yao-Ting Zhang. 1990. Generalized Multivariate Analysis. New York: Springer. Graham, Bryan S., and James L. Powell. 2012. Identification and estimation of average partial effects in ‘irregular’ correlated random coefficient panel data models. Econometrica 80: 2105–52. Gupta, Arjun K., Tamas Varga, and Taras Bodnar. 1993. Elliptically Contoured Models in Statistics. Norwell: Kluwer Academic Publishers. Heckman, James J., Daniel Schmierer, and Sergio Urzua. 2010. Testing the correlated random coefficient model. Journal of Econometrics 158: 177–203. [CrossRef] [PubMed] Heckman, James J., Sergio Urzua, and Edward Vytlacil. 2006. Understanding instrumental variables in models with essential heterogeneity. Review of Economics and Statistics 88: 389–432. [CrossRef] Heckman, James, and Edward Vytlacil. 1998. Instrumental variables methods for the correlated random coefficient model: Estimating the average rate of return to schooling when the return is correlated with schooling. Journal of Human Resources 33: 974–87. [CrossRef] Hsiao, Cheng. 1996. Random coefficients models. In The Econometrics of Panel Data: A Handbook of the Theory with Applications. Edited by L´szl Mátyás and Patrick Sevestre. Berlin: Springer, vol. 33, pp. 77–99. Hsiao, Cheng. 2003. Analysis of Panel Data, 2nd ed. Cambridge: Cambridge University Press. Hsiao, Cheng, Trent W. Appelbe, and Christopher R. Dineen. 1993. A general framework for panel data analysis—With an application to Canadian customer dialed long distance service. Journal of Econometrics 59: 63–86. [CrossRef] Hsiao, Cheng, M. Hashem Pesaran, and A. Kamil Tahmiscioglu. 1999. Bayes estimation of short-run coefficients in dynamic panel data models. In Analysis of Panels and Limited Dependent Variables Models. Edited by Cheng Hsiao, M. Hashem Pesaran, Kajal Lahiri and Lung-Fei Lee. Berlin: Springer, pp. 185–213. Hsiao, Cheng, and M. Hashem Pesaran. 2008. Random coefficient models. In The Econometrics of Panel Data: Fundamentals and Recent Developments in Theory and Practice. Edited by L´szl Mátyás and Patrick Sevestre. Berlin: Springer, vol. 46, pp. 185–213. Hsiao, Cheng, Yan Shen, and Hiroshi Fujiki. 2005. Aggregate versus disaggregate data analysis—A paradox in the estimation of money demand function of Japan under the low interest rate policy. Journal of Applied Econometrics 20: 579–601. [CrossRef] Hsiao, Cheng, and A. Kamil Tahmiscioglu. 1997. A panel analysis of liquidity constraints and firm investment. Journal of the American Statistical Association 92: 455–65. [CrossRef] Kuh, Edwin. 1963. Capital Stock Growth: A Micro-Econometric Approach. Amsterdam: North-Holland. Mundlak, Yair. 1978. On the pooling of time series and cross section data. Econometrica 46: 69–85. [CrossRef] Pesaran, M. Hashem. 2003. Aggregation of linear dynamic models: An application of life-cycle consumption models under habit formation. Economic Modelling 20: 227–435. Pesaran, M. Hashem, and Ron Smith. 1995. Estimation of long-run relationships from dynamic heterogenous panels. Journal of Econometrics 68: 79–114. [CrossRef] Stoker, Thomas M. 1993. Empirical approaches to the problem of aggregation over individuals. Journal of Economic Literature 31: 1827–74. Swamy, Paravastu AVB. 1970. Efficient inference in a random coefficient regression model. Econometrica 38: 311–23. [CrossRef] Econometrics 2019,7, 7 18 of 18 Theil, Henri. 1954. Linear Aggregation of Economic Relations. Amsterdam: North Holland. Wooldridge, Jeffrey M. 2005. Fixed-effects related estimators for correlated random-coefficient and treatment-effect panel data models. The Review of Economics and Statistics 87: 385–90. [CrossRef] c 2019 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).