Semiparametric efficiency in nonlinear LATE models
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Hong, Han; Nekipelov, Denis Article Semiparametric efficiency in nonlinear LATE models Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Hong, Han; Nekipelov, Denis (2010) : Semiparametric efficiency in nonlinear LATE models, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 1, Iss. 2, pp. 279-304, https://doi.org/10.3982/QE43 This Version is available at: https://hdl.handle.net/10419/150312 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/3.0/
Quantitative Economics 1 (2010), 279–304 1759-7331/20100279 Semiparametric efficiency in nonlinear LATE models Han Hong Departments of Economics, Stanford University Denis Nekipelov Departments of Economics, University of California at Berkeley In this paper we study semiparametric efficiency for the estimation of a finitedimensional parameter defined by generalized moment conditions under the local instrumental variable assumptions. These parameters identify treatment effects on the set of compliers under the monotonicity assumption. The distributions of covariates, the treatment dummy, and the binary instrument are not specified in a parametric form, making the model semiparametric. We derive the semiparametric efficiency bounds for both conditional models and unconditional models. We also develop multistep semiparametric efficient estimators that achieve the semiparametric efficiency bound. To illustrate the efficiency gains from using the optimal semiparametric weights, we design a Monte Carlo study. It demonstrates that our semiparametric estimator performs well in nonlinear models. Keywords. Semiparametric efficiency bound, local treatment effect, FTP, child achievement, unemployment benefits. JEL classification. C25, C26, C31. 1. Introduction Semiparametric efficiency is an important issue in the estimation of treatment effect models and models with endogenous regressors; see, for example, Chernozhukov and Hansen (2005) and Newey (1990a), among others. Under the strong ignorability assumption, Hahn (1998) and Hirano, Imbens, and Ridder (2003) derived the semiparametric efficiency bound and developed semiparametric efficient estimators for the averaged treatment effect and the averaged treatment effect on the treated. Firpo (2003) extended their analyses to quantile treatment effects. Han Hong: [email protected] Denis Nekipelov: [email protected] The authors acknowledge generous research supports from the NSF and excellent research assistance from Tim Armstrong. The usual disclaimer applies. We would like to thank Alberto Abadie, Shakeeb Khan, Guido Imbens, Bernard Salanié, Ed Vytlacil, conference participants in the 2007 NASM meeting in Duke University, and seminar participants at Columbia University, University of Pennsylvania, Harvard, and MIT for helpful comments. The authors have made use of the FTP data without representing positions by the State of Florida and MDRC. Copyright ©2010 Han Hong and Denis Nekipelov. Licensed under the Creative Commons Attribution- NonCommercial License 3.0. Available at http://www.qeconomics.org. DOI: 10.3982/QE43
280 Hong and Nekipelov Quantitative Economics 1 (2010) An alternative approach to address the endogeneity problem is based on the local instrumental variable (LIV) method. The baseline model for this method has a dummy endogenous regressor and a dummy instrument variable. Under the LIV assumption, the instrumental variable weakly changes the endogenous regressor in one direction. Abadie (2003) showed that the entire distributional causal effect is identified for the complier population where the endogenous regressor changes from 0 to 1 as the instrumental variable changes from 0 to 1, and proposed linear and nonlinear conditional local average treatment effects (LATE) models as an extension of Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996) for the case when the instrument is valid conditional on a vector of covariates, X. In the context of quantile regression, conditional LATE models were first applied in Abadie, Angrist, and Imbens (2002). In contrast to the strong ignorability assumption, semiparametric efficiency under the LIV assumption has not been subject to careful studies. An exception is Frolich (2007), who derived the efficiency bound for the average treatment effect for compliers and showed that the propensity score, properly defined in the LIV context, does not affect the efficiency bound. Henderson, Millimet, Parmeter, and Wang (2006) applied the estimator for the averaged treatment effect on compliers to a fertility analysis. We emphasize that in this paper, we do not develop new models of treatment effects for compliers. The generality that we consider is solely aimed at encompassing the existing conditional and unconditional models of treatment effects for compliers. We make several theoretical contributions in this paper. First we derive the semiparametric efficiency bound for both unconditional and conditional versions of the nonlinear treatment effect parameters, particularly in the context of a general nonlinear conditional mean treatment effect model developed in Abadie (2003). We illustrate the specialization of the efficiency bounds to the quantile and linear treatment effect parameters of Abadie, Angrist, and Imbens (2002) and Abadie (2003). Our semiparametric efficiency calculations include both conditional models and unconditional models, which characterize different treatment effect parameters. The unconditional efficiency bounds include as a special case the mean parameter of Frolich (2007), and also include the treatment effect on the treated compliers, which is related to the average treatment effect of the treated (ATT) when endogeneity is absent. Our results also generalize the efficiency calculations in Hahn (1998) and Chen, Hong, and Tarozzi (2008).Second,we show that the semiparametric efficiency bounds for the treatment effect of treated compliers are different when the propensity score is unknown, is known, or is correctly specified parametrically. We also simplify the structure of the efficiency analysis compared to the existing literature. In addition, we develop semiparametric estimators that achieve the theoretical efficiency bounds. Efficient estimators are developed for both conditional and unconditional models. In the conditional case, we identify a member among a class of estimators admissible under the structure given in Abadie (2003) that achieves the efficiency bound. The structure of the model allows us to make use of the binary instrument feature of a conditional moment model and to reduce the problem of finding semiparametric efficiency bound to the moment-based framework as in Newey (1990b),Bickel, Klaassen, Ritov, and Wellner (1993),andRobins and Rotnitzky (1995). For unconditional
Quantitative Economics 1 (2010) Semiparametric efficiency 281 models, we described efficient estimators for both the treatment effect of compliers and the treatment effect of treated compliers for cases when the propensity score is unknown, known, and parametrically specified. In general, efficiency can be achieved by choosing the instrument functions optimally in the propensity weighting framework of Abadie (2003) or in a conditional expectation projection framework. We demonstrate the efficiency gain from using the optimal weights in a set of Monte Carlo experiments. Section 2develops the semiparametric efficiency results for the complier treatment effect model in Abadie (2003). Section 3develops efficient estimators that achieve the semiparametric efficiency bound, and explicitly quantifies the amount of efficiency improvement over existing methods. This section also gives regularity assumptions that validate the proposed semiparametric efficient estimator. Each of these two sections also discusses extensions to parameters that are defined unconditionally. Section 4re- ports the results from a simulation exercise. Finally Section 5concludes. The Appendix is provided as Supplemental Material (Hong and Nekipelov (2010)): it contains the mathematical proofs and also an application of the efficient estimator to the Florida Transition Program which was offered as an alternative to the existing state welfare system in Florida. 2. Semiparametric efficiency bound 2.1 Local treatment effect parameters The local (complier) treatment effect model (see Imbens and Angrist (1994) and Abadie (2003), for example) is defined through a random vector (Y1Y0)∈R2, a vector of binary variables (D1D0)∈{01}×{01}, a binary instrument Z∈{01}, and a vector of covariates X∈X⊂Rk. The following assumptions are used by these authors to describe the distributions of the variables under consideration: Assumption 1. Almost everywhere in X, (i) (Y1Y0D1D0)⊥Z|X, (ii) E[D1|X]=E[D0|X], (iii) Pr(Z =1|X)∈(01), (iv) Pr(D1≥D0|X)=1. Under these four assumptions, in particular the last assumption (monotonicity), the data directly identify the differences between the cohort that would have been treated for both values of the instrument (always-takers) and the cohort that would not have been treated under any circumstances (never-takers). The combination of the alwaystaker cohort and the never-taker cohort indirectly recovers the compliers, which is the cohort that change behavior when the instrument changes. The variables in the model, Y1,Y0,D1,andD0are not always completely observable. Only the following transformed variables are observed: Y=Y1D+Y0(1−D) and D=D0+Z(D1−D0)
282 Hong and Nekipelov Quantitative Economics 1 (2010) Note that in this setup, covariates Xcan be endogenous. The binary treatment Dis endogenous by construction. On the other hand, as we will see later, the presence of the “switching” dummy Zallows us to recover the effect of the treatment for some subpopulation of the sample. In this way, provided the assumption of conditional independence of Z,andYiand Di(i=01)givenX, we can use variable Zas an instrument for the endogenous treatment. Due to Assumption 1(i), the conditional probabilities of the observable binary variable Dcan be written as Pi(X) =P(D =1|Z=i X) =E[Di|X],i=01,wherethesecond equalities follow from the conditional independence Assumption 1(i). Also define Q(X) =E[Z|X]. Consequently, the conditional probability of the binary treatment dgiven the instrument in terms of the probabilities of treatment dummies can be expressed as P(D=1|Z=zX =x) =F(z x) =P1(x)z +P0(x)(1−z) Taking expectation over Zgiven Xproduces the conditional probability of dgiven only X:P(x) =P1(x)Q(x) +P0(x)(1−Q(x)) The objects of interest that can be identified under Assumption 1are the distributions of the outcomes Y1and Y0given D1>D 0(implying that D1=1and D0=0): for j=01,f(yj|D1>D 0X =x). The subpopulation for which D1>D 0is usually referred to as compliers, for whom random selection into treatment affects the treatment dummy monotonically. Under the monotonicity Assumption 1(iv), the distributions of compliers can be expressed in terms of the observed conditional distributions: f∗∗(y|x d =1) ≡f(y1=y|D1>D 0x) (1) =P1(x) P1(x) −P0(x)f(y|d=1z=1x)−P0(x) P1(x) −P0(x)f(y|d=1z=0x) To see this relation, note that under the monotonicity assumption, P0(x) is the proportion of always-takers (D0=D1=1) conditional on xwhile P1(x) is the sum of alwaystakers and compliers. f(y1|d=1z =1x) gives the distribution of y1conditional on being either an always-taker or a complier and the covariate x.f(y1|d=1z =0x) gives the distribution of y1conditional on being just an always-taker and x. Therefore, P1(x)f(y1|d=1z =1x) can be written as a linear combination for the known distribution of always-takers and the unknown distribution for compliers. Similarly, one can write the joint distribution of y0and the event of being either a never-taker or a complier, which is (1−P0(x))f (y0|d=0z=0x), as a linear combination of the distributions for never-takers and compliers. Hence f∗∗(y|X=x d =0) ≡f(Y0=y|D1>D 0x) (2) =1−P0(x) P1(x) −P0(x)f(y|d=0Z=0x)−1−P1(x) P1(x) −P0(x)f(y|d=0Z=1x)
Quantitative Economics 1 (2010) Semiparametric efficiency 283 The semiparametric model that we consider incorporates the linear quantile regression model of Abadie, Angrist, and Imbens (2002) to a parameter vector βdetermined by a conditional moment equation, ∀xand ∀d, ϕ(βx d) =E[g(ydxβ)|xdD1>D 0]=0(3) for some parametric function g(·).1The conditional expectation for d=10is taken with respect to the corresponding conditional density f∗∗(y|xd). Two direct applications of this general definition are the mean treatment effect of Imbens and Angrist (1994) and the quantile treatment effect of Abadie, Angrist, and Imbens (2002). The mean treatment effect model corresponds to a moment condition g(ydxβ) =y−β1d−(1−d)β0−β 2x The quantile treatment effect model characterizes the difference in conditional distributions of potential outcomes y1and y0for compliers through a linear specification of the conditional quantile functions: Qτ(y|x dD1>D 0)=β0d+β 1x The corresponding moment function that defines the quantile treatment effect (QTE) parameter is, therefore, g(ydxβ) =1(y ≤β0d+β 1x) −τ These models can be extended to allow for a semiparametric component in the conditional moment function. For μ(x) being a nonparametric function of x, we may consider estimating μ(x) and βsimultaneously in the moment function: g(y dx μ(x) β) For example, the parametric mean treatment effect model can be generalized to a semiparametric partial linear model: g(y dx μ(x) β) =y−β1d−(1−d)β0−μ(x) In the rest of the paper, we derive semiparametric efficiency bounds for the parameter vector βand develop a semiparametric procedure that achieves the efficiency bound. This framework can be extended to derive the semiparametric efficiency bound for a nonparametric component in the specification of the conditional moment equations. 2.2 Efficiency bound for treatment effect parameters We will use the arguments of Newey (1990a) and Severini and Tripathi (2001) to construct the efficiency bounds for the system of conditional moments. More specifically, given a set of instrument functions of the covariates x, the conditional moments are first transformed into a system of unconditional moments. Then choosing the instrument functions optimally will produce the semiparametric efficiency bound of the conditional moment model.2 1The class of conditional models of treatment effects for compliers was developed in Abadie (2003). 2We note that the finiteness of the efficiency bound may involve some strong conditions on the absolute integrability of the inverse propensity score function, noted in Khan and Tamer (2009). When these conditions are not satisfied, one cannot estimate the treatment effect estimator at a parametric rate and the efficiency bound becomes meaningless. One, however, may use the efficient procedure outlined in Khan and Nekipelov (2010) even in such a nonregular case.
284 Hong and Nekipelov Quantitative Economics 1 (2010) Theorem 1. Under Assumption 1,the semiparametric efficiency bound for a k-dimen- sional parameter βthat characterizes the subsample of compliers in (3)can be expressed as: V(β)=E(P1(x) −P0(x))2E∂ϕ(βdx) ∂β ζ(xd)x ׯ Ω(x)−1Eζ(xd)∂ϕ(βdx) ∂β x−1 Denote ωdz(x) =V(g(ydxβ)|dzx) and γdz(x) =E(g(ydxβ)|dzx).We can then express the elements of the matrix ¯ Ω(x) in the manner ¯ Ω11(x) =P1(x)ω11(x) Q(x) +P0(x)ω10(x) 1−Q(x) +γ2 11(x)P1(x)P(x) P0(x)Q(x)(1−Q(x))1−P1(x)P0(x) P(x) ¯ Ω22(x) =(1−P1(x))ω01(x) Q(x) +(1−P0(x))ω00(x) 1−Q(x) +γ2 00(x)(1−P0(x))(1−P(x)) Q(x)(1−Q(x))(1−P1(x)) 1−(1−P0(x))(1−P1(x)) 1−P(x) and ¯ Ω21(x) =¯ Ω12(x) =P1(x)(1−P0(x)) Q(x)(1−Q(x)) γ11(x)γ00(x) In this theorem, we have also used the notation ζ(dx) =d P(x)1−d 1−P(x) This structure of the variance bound shows several visible features. First of all, the semiparametric efficiency bound will grow if the fraction of compliers P1(x) −P0(x) in the sample decreases. Moreover, the efficiency bound will be higher if the binary instrument is taking one of the values most of the time, in which case Q(x) is closer to 0or 1. In addition, the proof and the estimation section show that the structure of the variance reflects the optimal instrument function as M(x)ζ(xd),where M(x) =E∂ϕ(dxβ) ∂β ζ(xd)x¯ Ω(x)−1diagP(x) Q(x)1−P(x) 1−Q(x) 2.3 Unconditional parameters Often times researchers can be mainly interested in parameters that are defined unconditionally. For example, under the unconfoundedness assumption where the latent outcome is conditionally independent of the treatment status given exogenous covariates,
Quantitative Economics 1 (2010) Semiparametric efficiency 285 the semiparametric efficiency literature has focused on the average treatment effect and the average treatment effect on the treated, both of which are defined unconditionally with respect to the exogenous covariates X. Under the unconfoundedness assumption, one can also specify a model where the average treatment effect or effect on the treated conditional on each covariate is constant or a known parametric function of the covariates, similar to the analysis in the previous section and in Abadie, Angrist, and Imbens (2002). However, most of the literature has focused on analyzing the average treatment effect or effect on the treated without requiring that this effect is a constant conditional on every value of the exogenous covariate. When Xis not a constant, the conditional model and the unconditional model imply very different parameters of interest. For example, the semiparametric efficiency bound for an average treatment effect that is assumed to be constant across all possible values of covariate Xis tighter than that for the average treatment effect defined unconditionally with respect to the covariates X. This section investigates efficient estimators for unconditionally defined treatment effect parameters under the LIV monotonicity assumption. 2.4 Semiparametric efficiency of unconditional mean treatment effects This section will restrict attention to mean effect parameters to illustrate the ideas. However, the results are readily extendsible to general moment conditions in Section 3.6. Specifically, we consider the average treatment effect on compliers (ATEC) β≡β1−β0= E(Y1−Y0|D1>D 0)and the average treatment effect on the treated compliers (ATTC) γ≡γ1−γ0=E(Y1−Y0|d=1D1>D 0). These parameters reduce to the usual notation of average treatment effect (ATE) and effect on the treated (ATT) under strong ignorability when P(D1>D 0)=1. The efficiency bound for ATEC was derived by Frolich (2007), although we develop a simplified derivation. Our results for ATTC are new and are applicable when the propensity score Q(x) is unknown, known, or parametrically specified. The first theorem considers unknown propensity scores. Theorem 2. The semiparametric efficient bound for βis given by the variance of the efficient influence function 1 P(D1>D 0)z Q(x)(y −E(Y|Z=1x))+E(Y|Z=1x) −1−z 1−Q(x)(y −E(Y|Z=0x))−E(Y|Z=0x) −z Q(x)(d −E(D|Z=1x))+E(D|Z=1x) −1−z 1−Q(x)(d −E(D|Z=0x))−E(D|Z=0x)β
286 Hong and Nekipelov Quantitative Economics 1 (2010) while the semiparametric efficiency bound for γis given by the variance of the efficient influence function 1 P(d =1D1>D 0)y−1−z 1−Q(x)(y −E(Y|Z=0x))−E(Y|Z=0x) −d−1−z 1−Q(x)(d −E(D|Z=0x))−E(D|Z=0x)γ Obviously, under the strong ignorability assumption when Z=DP(D1>D 0)=1, both of these reduce to the corresponding influence functions derived in Hahn (1998). In fact, the only difference (other than the factor outside the bracket) is in the coefficient in front of βand γ, which become 1and zunder strong ignorability. The literature has also been concerned with the semiparametric efficiency when the so-called propensity score, in our case Q(x), is either known or parametrically specified. We will still leave P1(x) −P0(x) nonparametrically specified, even though cases when this is known or parametrically specified can be analyzed too. From the proof of Theorem 2, it is clear that Q(x) does not even enter the moment conditions that define the parameters β. (See equations (15) and (17) in Appendix C). Consequently, any knowledge of Q(x) will have no impact on the efficiency bound for β. Such knowledge, however, will improve the efficiency bound for γ, as described in the following theorem. Theorem 3. When the propensity score Q(x;α) is correctly specified up to a finitedimensional parameter α,the semiparametric efficiency bound for γis the variance of the efficient influence function 1 P(D=1D1>D 0)z(y −E(Y|Z1=1x))+Q(x)E(Y |Z1=1x) −1−z 1−Q(x)Q(x)[y−E(Y|Z=0x)]−Q(x)E(Y |Z=0x) −z(d −E(D|Z1=1x))+Q(x)E(D|Z1=1x) −1−z 1−Q(x)Q(x)[d−E(D|Z=0x)]−Q(x)E(D|Z=0x)γ +Proj(z −Q(x))κ(x)|Sα(z;x) In the above expression we have used the definition κ(x) =E(Y|Z=1x)−E(Y|Z=0x) −(E(D =1|Z=1x)−E(D =1|Z=0x))γ and the efficient influence function of the parametric propensity score model Sα(z;x) =z−Q(x) Q(x)(1−Q(x)) ∂Q ∂α (xα)
Quantitative Economics 1 (2010) Semiparametric efficiency 293 There also exists an alternative estimator that relies on direct estimation of the conditional expectation E[g(Y DXβ)|D X =xD1>D 0]for each candidate parameter β instead of on reweighting the moment conditions using the inverse of ˆ Q(x).Todescribe this estimator, begin with rewriting the identification condition (5)as E(˜ g|D=dD1>D 0X=x) =d P1(x) −P0(x)E(D˜ g|Z=1x)−d P1(x) −P0(x)E(D˜ g|Z=0x) +(1−d) P1(x) −P0(x)E((1−D) ˜ g|Z=0x) −(1−P1(x)) P1(x) −P0(x)E((1−D) ˜ g|Z=1x) For a given instrument matrix M(x), this suggests estimating βby equating to zero the sample analog 1 N N k=1 φk(β) =1 N N k=1dk ˆ P1(xk)−ˆ P0(xk)ˆ E(dk˜ g|Z=1xk) −dk ˆ P1(xk)−ˆ P0(xk)ˆ E(dk˜ g|Z=0xk) (9) +1−dk ˆ P1(xk)−ˆ P0(xk)ˆ E((1−dk)˜ g|Z=0xk) −1−dk ˆ P1(xk)−ˆ P0(xk)ˆ E((1−dk)˜ g|Z=1xk) where each of the conditional expectation terms are estimated nonparametrically at every given parameter value β.Forexample, ˆ E(dk˜ g|Z=1xk)=ˆ Q(xk)−1ˆ E(dkzk˜ g|xk) (10) Both conditional expectations can be estimated using a variety of nonparametric regression methods such as sieve expansion or kernel smoothing. It is easy to show that the asymptotic linear influence function that corresponds to the moment condition 1 NN k=1φk(β) for a given M(x) including the optimal one coincides with the semiparametric efficient function. First of all, similar to before, estimating P1(x) −P0(x) has no impact on the asymptotic variance due to the conditional nature of the moment restrictions. Using the representation theorem of Newey (1994),wecan, for example, expand the first component as 1 √N N k=1 dk ˆ P1(xk)−ˆ P0(xk) ˆ E(dkzk˜ g|xk) ˆ Q(xk) =1 √N N k=1P(xk)dkzk˜ g (P1(xk)−P1(xk))Q(xk)
294 Hong and Nekipelov Quantitative Economics 1 (2010) −P(xk)P1(xk)E( ˜ g|dk=1zk=1xk) (P1(xk)−P1(xk))Q(xk)(zk−Q(xk)) +dk−P(x) (P1(x) −P0(x))E(dk˜ g|zk=1xk)+op(1) Similar calculations can be applied to the other three terms. When summing these four components, we note that the last terms in each of the components cancel out due to the implications of the conditional moment restrictions that E(dk˜ g|zk=1xk)=E(dk˜ g|zk=0xk)and E((1−dk)˜ g|zk=1xk)=E((1− dk)˜ g|zk=0xk) Therefore, it is easy to check that the sum of the four influence functions is identical to the semiparametric efficient influence function when the instrument is chosen optimally. For the sake of brevity, we omit the regularity conditions for the conditional expectation projection estimator. To summarize, the implementation of this estimation method is a two step procedure, each step of which involves a profiled semiparametric estimator. In the first step, for an initial arbitrary choice of the instrument matrix M(x) and for each trial parameter β, the moment condition in each of the terms in (10) in the moment condition (9) is estimated nonparametrically to form the moment condition (9). The near zero of this moment gives an initial estimate of β. In second step, the same procedure is repeated using a consistent estimate of the efficient instrument matrix M(x) which depends on the initial estimate of βfollowing the procedure outlined in Section 3.2. 3.5 Efficient estimation of unconditional parameters It is easy to show that an efficient estimator can be derived from the principle of conditional expectation projection that follows the identification condition. Consider first the case of the average treatment effect on compliers (ATEC) β=E[Y1−Y0|D1>D 0]. Combining equations for the means of the distributions of treated and nontreated observations for compliers, we obtain the unconditional moment equation Eβ(P1(x) −P0(x)) −(E[y|z=1x]−E[y|z=0x])=0 The efficient semiparametric estimator is obtained from the sample analog of this moment equation and takes the form ˆ β=1 N N k=1 (ˆ P1(xk)−ˆ P0(xk))−11 N N k=1 (ˆ E[yk|zk=1xk]− ˆ E[yk|zk=0xk]) Conditional expectations in this expression can be estimated nonparametrically by kernel- or sieve-based methods. Semiparametric efficiency of this estimator can be established by the same projection arguments that we used before to establish efficiency of the estimator for the conditional moment-based model. Similarly to the ATEC, we can estimate the average treatment effect for the treated (ATTC) as γ=E[Y1−Y0|d=1D1>D 0]. The ATTC can be written in terms of the unconditional moment equation Eγ(P1(x) −P0(x)) −(E[Y|Z=1x]−E[Y|Z=0x])Q(x)=0
Quantitative Economics 1 (2010) Semiparametric efficiency 295 By the same principle as the ATEC, we express the efficient estimator as an empirical analog ˆγ=1 N N k=1ˆ Q(xk)( ˆ P1(xk)−ˆ P0(xk))−1 ×1 N N k=1ˆ Q(xk)( ˆ E[yk|zk=1xk]− ˆ E[yk|zk=0xk]) Using the projection argument of Newey (1994), we can easily verify that this estimator achieves the semiparametric efficiency bound when each of the conditional expectations and conditional probabilities above are estimated nonparametrically using either kernel- or sieve-based methods. If the Q(x) is specified as a parametric function or is a known function Qα(x),then the efficient estimator for γbecomes ˆγ=1 N N k=1ˆ Qˆα(xk)( ˆ P1(xk)−ˆ P0(xk))−1 ×1 N N k=1 Qˆα(xk)( ˆ E[yk|zk=1xk]− ˆ E[yk|zk=0xk]) where ˆαis the parametric maximum likelihood estimator (MLE), or the known α0if Q(x) is fully known. When the propensity score Q(x) is entirely unknown, an alternative efficient estimator can be developed using the inverse propensity score weighting approach of Abadie (2003).WhenQ(x) is known or parametrically specified, however, Hahn (1998) and Hirano, Imbens, and Ridder (2003) showed that efficient estimators based on inverse propensity score weighting typically require combining a nonparametric estimate of ˆ Q(x) with the known or parametrically estimated Q(x). A detailed comparison between the conditional expectation projection approach and the inverse propensity weighting approach is provided in Chen, Hong, and Tarozzi (2008), but they maintained the unconfoundedness assumption and did not investigate endogeneity. To summarize, a recipe for empirically implementing the efficient estimators for β and γonly requires summing over the data a function of the nonparametrically estimated choice probabilities ˆ Q(x),ˆ P1(x),and ˆ P0(x), and nonparametric estimates of the conditional expectations ˆ E[yk|zk=1xk]and ˆ E[yk|zk=0xk].Thisisaverystraightforward procedure that does not involve any nonlinear or numerical optimization procedures. 3.6 General separable unconditional model for compliers The treatment effect models considered in this section have a straightforward generalization to the separable conditional moment restrictions expressed in terms of unobservable outcome variables. Consider a problem where a finite-dimensional parameter
296 Hong and Nekipelov Quantitative Economics 1 (2010) β∈Rkis given by the following unconditional moment equation described in terms of unobservable variables Y1and Y0: ϕ(β) =E[g1(Y1xβ)−g0(Y0xβ)|D1>D 0]=0(11) In particular, when g1(Y1β)=Y1−βand g0(Y0β)=Y0+β, parameter βdefines the average treatment effect for compliers. On the other hand, g1(Y1β)=1(Y1≤β1)−τ and g0(Y0β)=1(Y0≤β0)+τdefine a complier analog of the average quantile treatment effect parameter proposed in Firpo (2003). Note that we can represent this moment equation for compliers in terms of distributions for the entire population. Using the Bayes’s rule, we find that this equation is equivalent to E(P1(x) −P0(x))Q(x)E[g1(Y1xβ)|d=1D1>D 0x] −(1−Q(x))E[g0(Y0xβ)|d=0D1>D 0x]=0 which can be redefined in terms of only observable variables in the form E(P1(x) −P0(x))Q(x)d P(x) +(1−Q(x))(1−d) 1−P(x) ×E[dg1(yxβ)−(1−d)g0(yxβ)|dxD1>D 0]=0 This equation in general defines an overidentified system of moments for β. Using a constant matrix A(which we can then choose optimally), we can transform this vector of moments into an exactly identified system. The Jacobi matrix Jfor this system given A is computed in the standard way. The following theorem describes the structure of the efficient influence function for this model. Theorem 6. In the model given by the general moment condition (11), the efficient influence function,which corresponds to finite-dimensional parameter β,can be expressed as (ydxz) =−J−1Az−Q(x) 1−Q(x)dg1(yxβ)+(1−d)g0(yxβ) −1 Q(x)(1−Q(x))E[(1−d)g0(yxβ)|z=1x] +Q(x)E[dg1(yxβ)|z=0x] =−J−1Aφ(ydxz) The structure of the efficient influence function in this case is similar to that in the ATE model which we considered earlier in this section. We can further choose the ma-
Quantitative Economics 1 (2010) Semiparametric efficiency 297 trix Asuch that it minimizes the variance of the efficient influence function. In particular, given that the Jacobi matrix can be expressed as J=A∂ϕ(β) ∂β the semiparametric efficiency bound for this model when Ais chosen optimally takes the form V(β)=(ϕ(β) ∂β E[φ(ydxz)φ(ydxz)]ϕ(β) ∂β)−1. An optimally weighted generalized method of moments (GMM) estimator based on the nonparametrically estimated moment condition 1 N N k=1 (ˆ P1(xk)−ˆ P0(xk))−1 ×1 N N k=1 (ˆ E[g1(ykxkβ)|zk=1xk]− ˆ E[g0(ykxkβ)|zk=0xk])=0 can easily be shown to achieve the efficiency bound derived in Theorem 6. Similarly, it is immediate to develop semiparametric efficiency bounds for a nonlinear treatment effect parameter for treated compliers,definedas ϕ(γ) =E[g1(Y1xγ)−g0(Y0xγ)|d=1D1>D 0]=0 In addition, an optimally weighted GMM estimator based on the nonparametric estimated moment condition 1 N N k=1ˆ Qˆα(xk)( ˆ P1(xk)−ˆ P0(xk))−1 ×1 N N k=1 Qˆα(xk) ׈ E[g1(ykxkγ)|zk=1xk]− ˆ E[g0(ykxkγ)|zk=0xk]=0 where ˆ Qα(xk)can be nonparametrically estimated, parametrically estimated, or the known propensity, can easily be shown to achieve the required corresponding semiparametric efficiency bound when the propensity score is unknown, parametrically specified, or known. 4. Numerical simulations In this section we report the results from a Monte Carlo study to illustrate the finite sample properties of the proposed estimators and the numerical efficiency comparisons with existing estimators. The design of the Monte Carlo study is motivated by the empirical illustration in Appendix E.
298 Hong and Nekipelov Quantitative Economics 1 (2010) 4.1 The structure of the data-generating process To analyze the performance of our semiparametric estimator, we designed an experiment where the outcome variable Ydepends on the endogenous regressor Xin a nonlinear fashion. We construct the data in which a binary instrumental variable Zdepends on X, but is conditional on Xpotential outcomes and the treatments are independent from Z. The observable outcome is generated from these variables. The data-generating mechanism for the simulation is characterized by the vector of potential outcomes (Y1Y0)and the vector of potential treatments (D1D0). Their distributions depend on the vector of covariates X. In light of the empirical illustration, the observable outcome for compliers has a Poisson distribution whose mean depends on the covariates and the parameters. We first consider the benchmark model. In Monte Carlo simulation experiments, we consider alternative parameter values and measure parameter differences in relation to the benchmark values. The sequential sampling scheme in the benchmark model is given by the following steps. First, we generate potential treatments as D1=1(γ0+xγ1+δ+v≥0)and D0= 1(γ0+xγ1+v≥0),whereγ0=−05,γ1=1,δ=1,xis generated from the uniform distribution, and vis standard normal. Second, we generate potential outcomes based on Poisson distributions. We first generate four independent Poisson random variables ξ1∼Poisson(exp(α +xβ)),ξ2∼Poisson(exp(xβ)),ξ3∼Poisson(λ11),and ξ4∼Poisson(λ00),whereα=1,β=05,λ11 =2,andλ00 =1. Denote e=(11).Then we construct the potential outcomes as Y1 Y0=ξ1 ξ2+ξ3e1{D1=1D0=1}+ξ4e1{D1=0D0=0} Note that this structure assures that for compliers (D1=1and D0=0), two potential outcomes are independent and Yi∼Poisson(exp(αi +xβ)) for i=01.Foralwaystakers (D1=D0=1), the potential outcomes have covariance λ11, and for never-takers (D1=D0=0), the potential outcomes have covariance λ00. Finally, the instrument is generated as an independent Bernoulli random variable Z∼Bernoulli((x)). Given the latent variables generated above, we compute the observable treatments and outcomes as D=D1Z+D0(1−Z) and Y=Y1D+Y0(1−D). This structure of the data-generating process guarantees that for compliers, the outcome will be independent from the observable treatment Dand the mean of the treatment outcome can be computed from the pair of Poisson random variables (ξ1ξ2). As a result, the datagenerating process for the Monte Carlo experiment is characterized by conditional moments E[Y−exp(αD +Xβ)|D1>D 0D=dX =x]=0 This model fits into our general conditional moment framework in the LATE context. To analyze the performance of our estimation procedure for parameters αand β,we designed a series of experiments.
Quantitative Economics 1 (2010) Semiparametric efficiency 299 4.2 Experiment 1: Basic comparison with alternative procedures The analysis of our estimation procedure begins with a comparison between our efficient two-stage estimator and the alternative existing estimator. We estimate this model using both the original method of Abadie (2003) and our more efficient estimator. Three parameters are estimated: β0=0is the coefficient on the constant term of the covariate, β1=1is the slope coefficient on the uniform regressor, and α=05is the coefficient on the treatment dummy. The simulation results across 1000 simulations are summarized in Table 1. We separately investigate the impact of the error in Monte Carlo sampling on our results by considering different Monte Carlo sample sizes. We analyze the differences in the mean-squared error of the estimated treatment effect αfor different numbers of Monte Carlo samples. We study the impact of the Monte Carlo sampling error by repeating the simulation results for 1000 Monte Carlo replications (which we use throughout our analysis) 100 times for the baseline parameters of the model. By decomposing the mean-squared error into a between Monte Carlo component and a within Monte Carlo Table 1. Simulation Summary. Parameter Mean Bias Median Bias Std Deviation Mean Squared Errors Sample Size 1000 Abadie et al. estimator β000956 01120 02176 00565 β100559 00527 01880 00385 α−03825 −03186 03254 02504 Semiparametric efficient estimator β000053 −01045 01587 00254 β100493 00540 01605 00282 α−00365 −00286 01672 00288 Sample Size 2000 Abadie et al. estimator β001126 01209 01435 00332 β100452 00384 01271 00182 α−03727 −03510 02205 01875 Semiparametric efficient estimator β000113 0079 01132 00129 β100434 00384 01141 00149 α−00186 −00254 01188 00144 Sample Size 4000 Abadie et al. estimator β000974 01050 00998 00194 β100534 00502 00893 00108 α−03441 −03325 01353 01367 Semiparametric efficient estimator β0−00009 00035 00787 00062 β100464 00453 00787 00083 α−00029 −00069 00824 00068
300 Hong and Nekipelov Quantitative Economics 1 (2010) component, we find that the mean-squared error for the estimated treatment effect associated with sampling error constitutes only 16.3% of the overall mean-squared error across a total of 100,000 Monte Carlo simulations. This leads us to the conclusion that even though the impact of the sampling error is visible, it does not substantially interfere with the results of the Monte Carlo experiment. 4.3 Experiment 2: Proportional reduction of noncompliers In this experiment, we study the robustness of the estimator against the endogeneity of the dependent variable. We proportionally decrease the variances of ξ3and ξ4,which are responsible for the dependence between the binary regressors and the outcome. In the limiting case where these variances λ00 and λ11 are equal to zero, the latent outcomes are completely independent. In Table 2, we document its effect on the mean-squared error of the estimate for the coefficient of the endogenous dummy variable. The tabulated data are obtained from 1000 Monte Carlo replications. The table shows that a reduction in the correlation between the treatment outcomes (Y1and Y0) leads to a smaller variance of the estimated treatment effect. Moreover, an increase in the sample size results in a decrease of the mean-squared error. The numbers in the table are computed from the Monte Carlo sample where we trimmed away the top and bottom 1% of observations to avoid including the cases where the distance minimization algorithm did not converge. Table 2. Simulation Summary for Decreasing Correlation Between Treatment Outcomes. λ11 =2K,λ00 =KSample Sizes K=250 350 450 600 1000 1 0.2716 0.2307 0.1865 0.1782 0.1282 0.9474 0.2774 0.2242 0.1873 0.1640 0.1306 0.8947 0.2791 0.2155 0.1792 0.1642 0.1249 0.8421 0.2578 0.2065 0.1900 0.1667 0.1207 0.7895 0.2509 0.2174 0.1829 0.1519 0.1138 0.7368 0.2453 0.2079 0.1808 0.1577 0.1209 0.6842 0.2528 0.2201 0.1883 0.1544 0.1171 0.6316 0.2504 0.1941 0.1732 0.1621 0.1139 0.5789 0.2463 0.1938 0.1663 0.1566 0.1137 0.5263 0.2506 0.1938 0.1684 0.1457 0.1124 0.4737 0.2305 0.2054 0.1746 0.1554 0.1140 0.4211 0.2169 0.1914 0.1600 0.1463 0.1173 0.3684 0.2187 0.1868 0.1589 0.1495 0.1151 0.3158 0.2135 0.1847 0.1625 0.1431 0.1096 0.2632 0.2116 0.1762 0.1514 0.1408 0.1133 0.2105 0.2181 0.1804 0.1596 0.1383 0.1125 0.1579 0.2149 0.1839 0.1551 0.1372 0.1067 0.1053 0.2074 0.1729 0.1642 0.1385 0.1035 0.0526 0.2137 0.1730 0.1660 0.1305 0.1030 0.0000 0.2050 0.1714 0.1528 0.1312 0.1053
Quantitative Economics 1 (2010) Semiparametric efficiency 301 4.4 Experiment 3: Smaller treatment effect In this experiment we study the sensitivity of our estimator with respect to the value of the treatment effect. We vary the coefficient αof the dummy endogenous variable, keeping the remaining components of the model the same. This exercise illustrates the robustness of our estimation method with respect to the magnitude of the treatment effect parameter of interest relative to the remaining components of the model. The results in Table 3show that, in principle, our procedure gives stable mean-squared errors across different choices of the treatment effect parameters and the sample sizes. Similar to the previous experiment, the numbers in the table are computed from the trimmed Monte Carlo sample (removing top and bottom 1% quantiles) to avoid including the cases where the distance minimization algorithm did not converge. One can see, however, from Table 3that reduction of the actual treatment effect leads to an increase in the distribution range for the estimated treatment effect. 4.5 Experiment 4: Choice of the bandwidth In this experiment, we study the sensitivity of the estimator with respect to the choice of the bandwidth parameter. We use the same structure of the bandwidth for the estimation of all nonparametric components of the model, including the conditional probability of treatment selection Z=0and the conditional probability of treatment D=0.We choose the bandwidth as hn=4σexp(K)n−1/3,whereσis the unconditional variance of Table 3. Simulation Summary for Reduction of the Treatment Effect. λ11 =05KSample Sizes K=250 350 450 600 1000 1 0.2716 0.2276 0.1903 0.1617 0.1347 0.9474 0.2728 0.2287 0.1943 0.1690 0.1319 0.8947 0.2871 0.2383 0.1999 0.1681 0.1266 0.8421 0.2899 0.2332 0.2064 0.1661 0.1288 0.7895 0.2793 0.2358 0.1888 0.1679 0.1216 0.7368 0.2845 0.2315 0.2060 0.1688 0.1223 0.6842 0.2924 0.2371 0.2067 0.1745 0.1251 0.6316 0.2922 0.2172 0.2040 0.1743 0.1271 0.5789 0.2929 0.2314 0.1900 0.1756 0.1322 0.5263 0.2921 0.2321 0.1959 0.1621 0.1299 0.4737 0.2931 0.2291 0.2027 0.1684 0.1313 0.4211 0.2820 0.2118 0.1988 0.1573 0.1234 0.3684 0.2894 0.2224 0.1904 0.1695 0.1326 0.3158 0.2481 0.2101 0.1881 0.1538 0.1259 0.2632 0.2558 0.2200 0.1782 0.1637 0.1255 0.2105 0.2700 0.1825 0.1550 0.1250 0.1117 0.1579 0.2480 0.1886 0.1513 0.1127 0.0970 0.1053 0.2577 0.1613 0.1305 0.1155 0.0779 0.0526 0.2596 0.1687 0.1309 0.1109 0.0796 0.0000 0.2013 0.1532 0.1339 0.1110 0.0868
302 Hong and Nekipelov Quantitative Economics 1 (2010) Table 4. Simulation Summary for Various Bandwidth Choices. h=4σexp(K)n−1/3Sample Sizes K=250 350 450 600 1000 1 0.2860 0.2488 0.2210 0.1854 0.1565 0.9474 0.3022 0.2436 0.2210 0.1900 0.1528 0.8947 0.3002 0.2392 0.2152 0.2000 0.1570 0.8421 0.2669 0.2392 0.2202 0.2016 0.1535 0.7895 0.3215 0.2492 0.2266 0.1961 0.1540 0.7368 0.2868 0.2609 0.2183 0.2045 0.1492 0.6842 0.3027 0.2403 0.2236 0.1922 0.1581 0.6316 0.2926 0.2482 0.2275 0.1857 0.1601 0.5789 0.2825 0.2457 0.2194 0.1950 0.1582 0.5263 0.2749 0.2475 0.2154 0.1972 0.1502 0.4737 0.2863 0.2437 0.2172 0.1972 0.1555 0.4211 0.2895 0.2524 0.2231 0.1990 0.1553 0.3684 0.3074 0.2566 0.2043 0.1795 0.1650 0.3158 0.2802 0.2478 0.2221 0.1857 0.1540 0.2632 0.2809 0.2362 0.2189 0.1969 0.1571 0.2105 0.2841 0.2419 0.2135 0.2008 0.1533 0.1579 0.2857 0.2430 0.2179 0.1876 0.1542 0.1053 0.2937 0.2331 0.2209 0.1824 0.1520 0.0526 0.2759 0.2315 0.2069 0.1996 0.1528 0.0000 0.3066 0.2393 0.2245 0.1875 0.1495 the binary variable (Zor D), nis the sample size, and Kis the constant of choice which we vary from 0 to 1. The results in Table 4demonstrate the mean squared errors across the Monte Carlo simulations. As one can see from the table, the mean-squared error remains stable across all different bandwidth choices. It is especially visible for sample size 1000. This confirms our theoretical results that if the regularity conditions are satisfied, the choice of the estimation procedure for non-parametric components of the model should not have a large impact on the estimated treatment effect parameter. 5. Conclusion In this paper, we derive the semiparametric efficiency bound for the estimation of a finite-dimensional parameter defined by generalized moment conditions under the local instrumental variable assumptions of Imbens and Angrist (1994) and Abadie, Angrist, and Imbens (2002). These parameters identify the treatment effect on the set of compliers under the monotonicity assumption. The moment equation characterizes the parametrized moment of the outcome distribution given a set of covariates and the treatment dummy. The distributions of covariates, the treatment dummy, and the binary instrument are not specified in a parametric form, making the model semiparametric. We also develop multistep semiparametric efficient estimators that achieve the semiparametric efficiency bound. The results of the Monte Carlo simulations demonstrate good performance of the semiparametric efficient estimator for finite samples.