Heterogeneous treatment effects with mismeasured endogenous treatment
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Ura, Takuya Article Heterogeneous treatment effects with mismeasured endogenous treatment Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Ura, Takuya (2018) : Heterogeneous treatment effects with mismeasured endogenous treatment, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 9, Iss. 3, pp. 1335-1370, https://doi.org/10.3982/QE886 This Version is available at: https://hdl.handle.net/10419/217130 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 9 (2018), 1335–1370 1759-7331/20181335 Heterogeneous treatment effects with mismeasured endogenous treatment Takuya Ura Department of Economics, University of California, Davis This paper studies the identifying power of an instrumental variable in the nonparametric heterogeneous treatment effect framework when a binary treatment is mismeasured and endogenous. Using a binary instrumental variable, I characterize the sharp identified set for the local average treatment effect under the exclusion restriction of an instrument and the deterministic monotonicity of the true treatment in the instrument. Even allowing for general measurement error (e.g., the measurement error is endogenous), it is still possible to obtain finite bounds on the local average treatment effect. Notably, the Wald estimand is an upper bound on the local average treatment effect, but it is not the sharp bound in general. I also provide a confidence interval for the local average treatment effect with uniformly asymptotically valid size control. Furthermore, I demonstrate that the identification strategy of this paper offers a new use of repeated measurements for tightening the identified set. Keywords. Local average treatment effect, instrumental variable, nonclassical measurement error, endogenous measurement error, partial identification. JEL classification. C21, C26. 1. Introduction Treatment effect analyses often entail a measurement error problem as well as an endogeneity problem. For example, Black, Sanders, and Taylor (2003) documented a substantial measurement error in educational attainments in the 1990 U.S. Census. At the same time, educational attainments are endogenous as treatment variables in return- to-schooling analyses, because unobserved individual ability affects both schooling decisionsandwages(Card (2001)). The econometric literature, however, has offered only a few solutions for addressing the two problems at the same time. An instrumental variable is a standard technique for correcting endogeneity and measurement error (e.g., Angrist and Krueger (2001)). To the best of my knowledge, however, no existing research has explicitly investigated the identifying power of an instrumental variable for Takuya Ura : [email protected] First version: November, 2015. I would like to thank Federico A. Bugni, V. Joseph Hotz, Shakeeb Khan, and Matthew A. Masten for their guidance and encouragement. I am also grateful to the Editor, four anonymous referees, Luis E. Candelaria, Xian Jiang, Marc Henry, Ju Hyun Kim, Arthur Lewbel, Jia Li, Arnaud Maurel, Marjorie B. McElroy, Ismael Mourifié, Naoki Wakamori, Yichong Zhang, and participants at various seminars and conferences. ©2018 The Author. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE886
1336 Takuya Ura Quantitative Economics 9 (2018) the heterogeneous treatment effect when the treatment is both mismeasured and endogenous.1 I consider a mismeasured treatment in the framework of Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996), and focus on the local average treatment effect as a parameter of interest. My analysis studies the identifying power of a binary instrumental variable under the following two assumptions: (i) the instrument affects the outcome and the measured treatment only through the true treatment (the exclusion restriction of an instrument), and (ii) the instrument weakly increases the true treatment (the deterministic monotonicity of the true treatment in the instrument). These assumptions are an extension of Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996) into the framework with mismeasured treatment. The local average treatment effect is the average treatment effect for the compliers, that is, the subpopulation whose true treatment status is strictly affected by an instrument. Focusing on the local average treatment effect is meaningful for a few reasons. First, the local average treatment effect has been a widely-used parameter to investigate the heterogeneous treatment effect with endogeneity. My analysis offers a tool for a robustness check to those who have already investigated the local average treatment effect. Second, the local average treatment effect can be used to extrapolate to the average treatment effect or other parameters of interest. Imbens (2010) emphasizes the utility of reporting the local average treatment effect in addition to the other parameters of interest because the extrapolation often requires additional assumptions and can be less credible than the local average treatment effect. The mismeasured treatment prevents the local average treatment effect from being point-identified. As in Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996), the local average treatment effect is the ratio of the intent-to-treat effect over the size of compliers.2Since the measured treatment is not the true treatment, however, the size of compliers is not identified and, therefore, the local average treatment effect is not identified. The underidentification for the local average treatment effect is a consequence of the underidentification for the size of compliers; if I assumed no measurement error, I could compute the size of compliers based on the measured treatment and, therefore, the local average treatment effect would be the Wald estimand.3 I take a worst case scenario approach against the measurement error and allow for a general form of measurement error. The only assumption concerning the measurement error in this paper is its independence of the instrumental variable. (Section 3.4 1Many existing methods, including Mahajan (2006) and Lewbel (2007), allow for the treatment effect to be heterogeneous due to observed variables. In this paper, I focus on the heterogeneity due to unobserved variables by considering the local average treatment effect framework. 2The intent-to-treat effect is defined as the mean difference of the outcome between the two groups defined by the instrument. The size of compliers is the probability of being a complier, and it is the mean difference of the true treatment (Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996)). 3The Wald estimand in this paper is defined as the ratio of the intent-to-treat effect over the mean difference of the measured treatment between the two groups defined by the instrument. Note that the Wald estimand is identified because it uses the measured treatment, but it is not the local average treatment effect because it does not use the true treatment.
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1337 | 0 ITT Wald (a) When the intent-to-treat effect is positive 0 | Wald ITT (b) When the intent-to-treat effect is zero | 0 ITT Wald (c) When the intent-to-treat effect is negative Figure 1. Identified set for the local average treatment effect. ITT is the intent-to-treat effect and Wald is the Wald estimand. The thick line is the identified set for the local average treatment effect. Note that the identified set is {0}when the intent-to-treat effect is zero. dispenses with this assumption and shows that it is still possible to bound the local average treatment effect.) The following properties of measurement error are practically and theoretically relevant. First, the measurement error is nonclassical; that is, it can be dependent on the true treatment. The measurement error for a discrete variable is always nonclassical. It is because the measurement error cannot be negative (positive) when the true variable takes the lowest (highest) value. Second, I allow the measurement error to be endogenous (or differential); that is, the measured treatment can be dependent on the outcome conditional on the true treatment. For example, as Black, Sanders, and Taylor (2003) argued, the measurement error for educational attainment depends on the familiarity with the educational system in the U.S., and immigrants may have a higher rate of measurement error. At the same time, the familiarity with the U.S. educational system can be related to the English language skills, which can affect the labor market outcomes. Bound, Brown, and Mathiowetz (2001) also argue that measurement error is likely to be endogenous in some empirical applications. (In Appendix C in the Supplementary Material (Ura (2018)), I explore the identifying power of the exogeneity assumption on the measurement error. The additional assumption yields a tighter sharp identified set, but I still cannot point identify the local average treatment effect in general.) Third, there is no assumption concerning the marginal distribution of the measurement error. It is not necessary to assume anything about the accuracy of the measurement. In the presence of measurement error, I derive the identified set for the local average treatment effect (Theorem 4). Figure 1describes the relationship among the identified set for the local average treatment effect, the intent-to-treat effect, and the Wald estimand. First, the intent-to-treat effect has the same sign as the local average treatment
1338 Takuya Ura Quantitative Economics 9 (2018) effect. Figure 1has three subfigures according to the sign of the intent-to-treat effect: (a) positive, (b) zero, and (c) negative. Second, the intent-to-treat effect is the sharp lower bound on the local average treatment effect in absolute value. Third, the Wald estimand is an upper bound on the local average treatment effect in absolute value. The Wald estimand is the probability limit of the instrumental variable estimator in my framework, which ignores the measurement error but controls only for the endogeneity. This point implies that an upper bound on the local average treatment effect is obtained by ignoring the measurement error. Frazis and Loewenstein (2003) obtain a similar result in the homogeneous treatment effect model. Last, but most importantly, the sharp upper bound in absolute value can be smaller than the Wald estimand. It is a potential cost of ignoring the measurement error and using the Wald estimand. Even for analyzing only an upper bound on the local average treatment effect, it is recommended to take the measurement error into account, which can yield a smaller upper bound than the Wald estimand. Section 3.1 investigates when the Wald estimand coincides with the sharp upper bound. I extend the identification analysis to incorporate covariates other than the treatment variable. In this setting, the instrumental variable satisfies the exclusion restriction after conditioning covariates. Based on the insights from Abadie (2003) and Frölich (2007), I show that the identification strategy of this paper works in the presence of covariates. I construct a confidence interval for the local average treatment effect. To construct the confidence interval, first, I approximate the identified set by discretizing the support of the outcome where the discretization becomes finer as the sample size increases. The approximation for the identified set resembles many moment inequalities in Chernozhukov, Chetverikov, and Kato (2014), who consider a finite but divergent number of moment inequalities. I apply a bootstrap method in Chernozhukov, Chetverikov, and Kato (2014) to construct a confidence interval with uniformly asymptotically valid asymptotic size control. The confidence interval also rejects parameter values which do not belong to the sharp identified set. An empirical application and a Monte Carlo simulation demonstrate a finite sample property of the proposed inference method. The empirical exercise is based on Abadie (2003), who studies the effects of 401(k) participation on financial savings, and I consider a misclassification of the 401(k) participation.4 As an extension, I consider the dependence between the instrument and the measurement error. In this case, there is no assumption on the measurement error, and the measured treatment has no information on the local average treatment effect. Even without using the measured treatment, however, I can still apply the same identification strategy and obtain finite (but less tight) bounds on the local average treatment effect. Moreover, I offer a new use of repeated measurements as additional sources for identification. The existing practice of repeated measurements uses one of them as an instrumental variable, as in Hausman, Newey, Ichimura, and Powell (1991), Hausman, Newey, 4The pension type is subject to a measurement error. See, for example, Gustman, Steinmeier, and Tabatabai (2008) for the pension-type misclassification in the Health and Retirement Study.
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1339 and Powell (1995), Mahajan (2006),andHu (2008).5However, when the true treatment is endogenous, the repeated measurements are likely to be endogenous and are not good candidates for an instrumental variable. My identification strategy demonstrates that those variables are useful for bounding the local average treatment effect in the presence of measurement error, even if none of the repeated measurements are valid instrumental variables. The remainder of this paper is organized as follows. Section 1.1 explains several empirical examples motivating mismeasured endogenous treatments and Section 1.2 reviews the related econometric literature. Section 2introduces mismeasured treatments in the framework of Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996). Section 3constructs the identified set for the local average treatment effect. I also discuss two extensions. One extension describes how repeated measurements tighten the identified set, and the other dispenses with independence between the instrument and the measurement error. Section 4proposes an inference procedure for the local average treatment effect. Section 5conducts an empirical illustrations. Section 6concludes. The Appendix collects proofs, remarks, and Monte Carlo simulations. 1.1 Examples for mismeasured endogenous treatments I introduce several empirical examples in which binary treatments can be both endogenous and mismeasured at the same time. The first example is the return to schooling, in which the outcome is wages, and the treatment is educational attainment, for example, whether a person has completed college or not. Unobserved individual ability affects both the schooling decision and wage determination, which leads to the endogeneity of educational attainment (e.g., Card (2001)). Moreover, survey datasets record educational attainments based on the interviewee’s answers, and these self-reported educational attainments are subject to measurement error. For example, Black, Sanders, and Taylor (2003) estimated that the 1990 Decennial Census had a 177% false positive rate of reporting a doctoral degree. The second example is labor supply response to welfare program participation, in which the outcome is employment status and the treatment is welfare program participation. Self-reported welfare program participation in survey datasets can be mismeasured (Hernandez and Pudney (2007)). The psychological cost for welfare program participation, welfare stigma, affects job search behavior and welfare program participation simultaneously; that is, welfare stigma may discourage individuals from participating in a welfare program, and, at the same time, affect an individual’s effort in the labor market. Moreover, the welfare stigma gives welfare recipients some incentive not to reveal their participation status to the survey, which causes endogenous measurement error in that the unobserved individual heterogeneity affects both the measurement error and the outcome. The third example is the effect of a job training program on wages. As it is similar to the return to schooling, unobserved individual ability plays a crucial role in this example. Self-reported completion of job training program is also subject to measurement 5It is worthwhile to mention that Lewbel (2007) allows for a certain form of the endogeneity in a repeated measurement, under which a repeated measurement still satisfies some exclusion restriction.
1340 Takuya Ura Quantitative Economics 9 (2018) error (Bollinger (1996)). Frazis and Loewenstein (2003) develop a methodology for evaluating a homogeneous treatment effect with mismeasured endogenous treatment, and apply their methodology to evaluate the effect of a job training program on wages. The last example is the effect of maternal drug use on infant birth weight. Kaestner, Joyce, and Wehbeh (1996) estimate that a mother tends to underreport her drug use, but at the same time, she tends to report it correctly if she is a heavy user. When the degree of drug addiction cannot be observed, it becomes an individual unobserved heterogeneity which affects infant birth weight and the measurement in addition to the drug use. 1.2 Literature review Here, I summarize the related econometric literature. Mahajan (2006),Lewbel (2007), and Hu (2008) used an instrumental variable to correct for measurement error in a binary (or discrete) treatment in the homogeneous treatment effect framework and they achieve nonparametric point identification of the average treatment effect. They assume that the true treatment is exogenous, whereas I allow it to be endogenous. Finite mixture models are related to my analysis. I consider the unobserved binary treatment, whereas finite mixture models deal with unobserved type. Henry, Kitamura, and Salanié (2014) and Henry, Jochmans, and Salanié (2015) are the most closely related. They investigate the identification problem in finite mixture models, by using the exclusion restriction in which an instrumental variable only affects the mixing distribution of a type without affecting the component distribution (i.e., the conditional distribution given the type). If I applied their approach directly to my framework, their exclusion restriction would imply conditional independence between the instrumental variable and the outcome given the true treatment. Instead of applying the approaches in Henry, Kitamura, and Salanié (2014) and Henry, Jochmans, and Salanié (2015),Iuseadifferent exclusion restriction in which the instrumental variable does not affect the outcome or the measured treatment directly. A few papers have applied an instrumental variable to a mismeasured binary regressor in the homogenous treatment effect framework. They include Aigner (1973),Kane, Rouse, and Staiger (1999),Bollinger (1996),Black, Berger, and Scott (2000),Frazis and Loewenstein (2003),andDiTraglia and García-Jimeno (2015).Frazis and Loewenstein (2003) and DiTraglia and García-Jimeno (2015) are the most closely related among them, since they allow for endogeneity. Here, I allow for heterogeneous treatment effects, and I contribute to the heterogeneous treatment effect literature by investigating the consequences of the measurement errors in the treatment. Kreider and Pepper (2007),Molinari (2008),Imai and Yamamoto (2010),andKreider, Pepper, Gundersen, and Jolliffe (2012) applied a partial identification strategy for the average treatment effect to the mismeasured binary regressor problem by utilizing the knowledge of the marginal distribution for the true treatment. Those papers use auxiliary datasets to obtain the marginal distribution for the true treatment. The framework in Kreider et al. (2012) is the most closely related to this paper, in that they allow for both treatment endogeneity and endogenous measurement error. My instrumental variable approach can be an an alternative strategy to deal with mismeasured endogenous treatment. It is worthwhile because, as mentioned in Schennach (2013), the availability of an
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1341 auxiliary dataset is limited in empirical research. Furthermore, it is not always the case that the results from auxiliary datasets is transported into the primary dataset (Carroll, Ruppert, Stefanski, and Crainiceanu (2012,p.10)), Some papers investigate mismeasured endogenous continuous variables, instead of binary variables. Amemiya (1985),Hsiao (1989),Lewbel (1998),Song, Schennach, and White (2015) considered nonlinear models with mismeasured continuous explanatory variables. The continuity of the treatment is crucial for their analysis, because they assume classical measurement error. The treatment in my analysis is binary and, therefore, the measurement error is nonclassical. Hu, Shiu, and Woutersen (2015) considered mismeasured endogenous continuous variables in single index models. However, their approach depends on taking derivatives of the conditional expectations with respect to the continuous variable. It is not clear if it can be extended to binary variables. Song (2015) considered the semiparametric model when endogenous continuous variables are subject to nonclassical measurement error. He assumes conditional independence between the instrumental variable and the outcome given the true treatment, which would impose some structure on the outcome equation when a treatment is binary. Instead I propose an identification strategy without assuming any structure on the outcome equation. Chalak (2017) investigated the consequences of measurement error in the instrumental variable instead of the treatment. He assumed that the treatment is perfectly observed, whereas I allow for it to be measured with error. Since I assume that the instrumental variable is perfectly observed, my analysis is not overlapped with Chalak (2017). Manski (2003), Blundell, Gosling, Ichimura, and Meghir (2007), and Kitagawa (2010) have similar identification strategy in the context of sample selection models. These papers also use the exclusion restriction of the instrumental variable for their partial identification results. Particularly, Kitagawa (2010) derived the integrated envelope from the exclusion restriction, which is similar to the total variation distance in my analysis because both of them are characterized as a supremum over the set of the partitions. First and the most importantly, I consider mismeasurement of the treatment, whereas the sample selection model considers truncation of the outcome. It is not straightforward to apply their methodologies in sample selection models into mismeasured treatment problem. Second, I offer an inference method with uniform size control, but Kitagawa (2010) derived only point-wise size control. Last, Blundell et al. (2007) and Kitagawa (2010) used their result for specification test, but I cannot use it to carry out a specification test because the sharp identified set of my analysis is always non-empty. Finally, Calvi, Lewbel, and Tommasi (2017) and Yanagi (2017) have recently discussed identification issues of the local average treatment effect in the presence of a measurement error in the treatment variable. They are built on results in the previous draft of this paper to derive novel and important results when there are additional variables in a dataset: multiple measurements of the true treatment variable (Calvi, Lewbel, and Tommasi (2017)) or multiple instrumental variables (Yanagi (2017)). In contrast, the results of this paper are valid without these additional variables and only requires the assumptions in Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996).
1342 Takuya Ura Quantitative Economics 9 (2018) outcome Ytrue treatment T∗ measured treatment T instrument Z Figure 2. Graphical representation of dependencies among variables. 2. Local average treatment effect framework with misclassification My analysis considers a mismeasured treatment in the framework of Imbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996). The objective is to evaluate the causal effect of a binary treatment T∗∈{01}on an outcome Y,whereT∗=0represents the control group and T∗=1represents the treatment group. To deal with endogeneity of T∗, I use a binary instrumental variable Z∈{01}which shifts T∗exogenously without any direct effect on Y. The treatment T∗of interest is not directly observed, and instead there is a binary measurement T∈{01}for T∗.Iputthe∗symbol on T∗to emphasize that the true treatment T∗is unobserved. I allow Yto be discrete, continuous or mixed; Yis only required to have some known dominating finite measure μYon the real line. For example, μYcan be the Lebesgue measure or the counting measure. Let Ybe the support for the random variable Yand T={01}be the support for T. To describe the data generating process, I consider the counterfactual variables. T∗ z is the counterfactual true treatment when Z=z.Yt∗is the counterfactual outcome when T∗=t∗.Tt∗is the counterfactual measured treatment when T∗=t∗. The individual treatment effect is Y1−Y0. It is not directly observed; Y0and Y1cannot be observed at the same time. Only YT∗is observable. Using the notation, the observed variables (YTZ)are generated by the three equations: T=TT∗(1) Y=YT∗(2) T∗=T∗ Z(3) Figure 2graphically describes the relationship among the instrument Z,the(unobserved) true treatment T∗, the measured treatment T, and the outcome Y.(1)isthe measurement equation, which is the arrow from T∗to Tin Figure 2.T−T∗is the measurement error; T−T∗=1represents a false positive and T−T∗=−1represents a false negative. Equations (2)and(3)arethesameasImbens and Angrist (1994) and Angrist, Imbens, and Rubin (1996).Equation(2) is the outcome equation, which is the arrow from T∗to Yin Figure 2.Equation(3) is the treatment assignment equation, which is the arrow from Zto T∗in Figure 2. A potentially nonzero correlation between (Y0Y1) and (T∗ 0T∗ 1)causes an endogeneity problem.
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1349 The identified set in Theorem 9is weakly smaller than the identified set in Theorem 4. The total variation distance TV(RYT ) in Theorem 9is weakly larger than that in Theorem 4, because, using the triangle inequality, TV(RYT) =1 2 t=01(f(RYT )|Z=1−f(RYT)|Z=0)(ryt)dμR(r) dμY(y) ≥1 2 t=01(f(RYT )|Z=1−f(RYT)|Z=0)(ryt)dμR(r) dμY(y) =1 2 t=01(f(YT )|Z=1−f(YT)|Z=0)(yt)dμY(y) =TV(YT) and the strict inequality holds unless the sign of (f(RYT)|Z=1−f(RYT )|Z=0)(ryt) is constant in rfor every (yt). Therefore, it is possible to test whether the repeated measurement Rhas additional information, by testing whether the sign of (f(RYT )|Z=1− f(RYT)|Z=0)(ryt)is constant in r. 3.4 Dependence between measurement error and instrumental variable It is still possible to apply the same identification strategy and obtain finite (but less tight) bounds on the local average treatment effect, even without the independence between the instrumental variable and the measurement error. (Assumption 1(i) implies that Zis independent of Tt∗for each t∗=01.) Instead Assumption 1is weakened to allow for the measurement error Tt∗to be correlated with the instrumental variable Z. Assumption 10. (i) Zis independent of (Yt∗T∗ 0T∗ 1)for each t∗=01. (ii) T∗ 1≥T∗ 0al- most surely. (iii) 0<P(Z=0)<1. Theorem 11 shows that the above observations characterize the identified set for the local average treatment effect under Assumption 10. Theorem 11. Suppose that Assumption 10 holds,and consider an arbitrary data distribution Pof (YTZ).The identified set ΘI(P) for the local average treatment effect is characterized as follows:ΘI(P) =Θif TVY=0;otherwise, ΘI(P) =⎧ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎩ EP[Y|Z]EP[Y|Z] TVYif EP[Y|Z]>0 {0}if EP[Y|Z]=0 EP[Y|Z] TVY EP[Y|Z]if EP[Y|Z]<0 The difference from Theorem 4is that Theorem 11 does not depend on the measured treatment T. Although it is observed in the dataset, Tdoes not have any information on the local average treatment effect because Assumption 10 does not restrict T.When TVY>0, there are nontrivial upper and lower bounds on the local average treatment effect even without using the measured treatment T.
1350 Takuya Ura Quantitative Economics 9 (2018) 4. Inference Based on the sharp identified set in the presence of covariates (Theorem 7), this section constructs a confidence interval for the local average treatment effect based on an i.i.d. sample {Wi:1≤i≤n}of W=(YTZV). The confidence interval described below controls the asymptotic size uniformly over a class of data generating processes, and rejects all the fixed alternatives. The identified set in Theorem 7is characterized by moment inequalities as follows. Lemma 12. Let Pbe an arbitrary data distribution of W=(YTZV).Under Assumption 6,ΘI(P) is the set of θ∈Θin which EP−Z−π(V ) π(V )1−π(V )sgn(θ)Y ≤0(6) EPZ−π(V ) π(V )1−π(V )sgn(θ)Y −|θ|≤0(7) EPZ−π(V ) π(V )1−π(V )|θ|h(YT V ) −sgn(θ)Y≤0for all h∈H(8) where π(V ) =P(Z =1|V),His the set of measurable functions on Y×T×Vtaking a value in {−0505}and sgn(x) ≡1{x≥0}−1{x<0}. The first condition in (6) states that the local average treatment effect θhas the same sign as the intent-to-treat effect EPEP[Y|ZV ]=EPYZ−EP[Z|V] EP[Z|V]EP[1−Z|V] The second condition in (7)is|θ|≥|EP[EP[Y|ZV ]]|. The last condition in (8) corresponds to |θ|≤EPEP[Y|ZV ] EP[TV(YT)|V]=EPEP[Y|ZV ] sup h∈H EPEPh(YTV ) |ZV where I use TV(YT )|V=suph∈HEP[h(YTV ) |ZV ]. The derivations are found in the proof of Lemma 12 in Appendix B. Iconstructa(1−α)-confidence interval for the local average treatment effect θ with treating πas a nuisance parameter for given α∈(005). I assume that a (1−δ)- confidence interval Cπn(δ) for πis available for researchers for given δ∈(0α).Given Cπn(δ),Iconstructthe(1−α−δ)-confidence interval Cθn(α +δ) for the local average treatment effect as Cθn(α +δ) = π∈Cπn(δ)θ∈Θ:T(θπ)≤c(αθπ)
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1351 where T(θπ) and c(αθπ) are defined below using the bootstrap-based testing (Chernozhukov, Chetverikov, and Kato (2014)). The number of the moment inequalities in Lemma 12 can be finite or infinite, which determines whether some of the existing methods can be applied directly to the inference on the local average treatment effect. When (Y V ) has finite supports and, therefore, His finite, the sharp identified set is characterized by a finite number of inequalities and, therefore, I can apply inference methods based on unconditional moment inequalities. To the best of my knowledge, however, inference for the local average treatment effect in my framework does not fall directly into the existing moment inequality models when either Yor Vis continuous. When either Yor Vis continuous, the sharp identified set is characterized by an uncountably infinite number of inequalities. In the current literature on the partially identified parameters, an infinite number of moment inequalities are mainly considered in the context of conditional moment inequalities. The identified set in this paper is not characterized by conditional moment inequalities. I consider a sequence of finite sets Hnwhich converges to Has a sample size increases. (The convergence is formally defined in Assumption 14, and an example for Hn appears after Assumption 14.) Note that, when His finite, Hncan be equal to H.IfH is replaced with Hnin Lemma 12, the number of the moment inequalities becomes finite. At the same time, as Hnapproaches to H, the approximation error from using Hn converges to zero, and the number of the inequalities can be increasing, particularly diverging to the infinity when Hincludes infinite elements. The approximated identified set is characterized by a finite number of the following moment inequalities: EP−Z−π(V ) π(V )1−π(V )sgn(θ)Y ≤0(9) EPZ−π(V ) π(V )1−π(V )sgn(θ)Y −|θ|≤0(10) EPZ−π(V ) π(V )1−π(V )|θ|h(YTV ) −sgn(θ)Y≤0for all h∈Hn(11) Denote by pnthe resulting number of moment inequalities, that is, the number of elements in Hnplus 2. For the size α∈(005), I construct a test statistic T(θπ) and a critical value c(αθπ) via the two-step multiplier bootstrap in Chernozhukov, Chetverikov, and Kato (2014), where Appendix Adescribes the procedures. Chernozhukov, Chetverikov, and Kato (2014) studies the testing problem for moment inequality models in which the number of the moment inequalities is finite but growing. Since the number of the moment inequalities in (9)–(11) is finite but growing, their results are applicable to construct a confidence interval based on (9)–(11). Assumption 13. Given positive constants C2and η,the class of data generating processes,denoted by P0,and the parameter spaces Θ×Πsatisfy the following properties:(i)max{EP[Y3]2/3EP[Y4]1/2}<C 2, (ii) Θ⊂Ris bounded, (iii) The random vari-
1352 Takuya Ura Quantitative Economics 9 (2018) able inside EPin (9)–(11)has a nonzero variance for every j=1pnand every θ∈Θ, (iv) liminfn→∞infP∈P0P(π ∈Cπn(δ)) ≥1−δ,and (v) η<π(V)<1−ηfor every π∈Π. The first assumption (i) is a regularity condition about the moments of Y.Thesecond assumption (ii) requires researchers to know ex ante upper and lower bounds on the parameter. The third assumption (iii) guarantees that the test statistic is well-defined. The fourth assumption (iv) is that the confidence interval for πcontrols the size uniformly over P0. The last assumption (v) is that the propensity score π(v) =P(Z =1|V= v) is bounded away from zero and one. In this paper, I assume that {Hn}satisfies the following conditions. Assumption 14. (i) Hn⊂Hn+1. (ii) The convergence sup h∈H EPZ−π(V ) π(V )1−π(V )h(YT V )−max h∈Hn EPZ−π(V ) π(V )1−π(V )h(YTV )→0(12) holds uniformly over π∈Πand P∈P0. (iii) The number of elements in Hnsatisfies log7/2(pnn) ≤C1n1/2−c1and log1/2pn≤C1n1/2−c1(13) for some c1∈(01/2)and C1>0. An example of Hnis obtained by discretizing Y×T×V. Partition Y×T×Vby In1InKn, in which the intervals Ink and the grid size Kndepend on the sample size n.Lethnj be a generic function of Y×T×Vinto {−0505}that is constant over Ink for every 1≤k≤Kn.LetHn={hn1hn2Kn}be the set of all such functions. Lemma 15 shows that this construction of Hnsatisfies equation (12) under conditions on Ink and f(YT)|Z=z. The conditions in Lemma 15 guarantee that the approximation error from the discretization vanishes as the sample size nincreases. Lemma 15. Assumption 14 holds if (i) the partition In+11In+1Kn+1is a refinement of the partition In1InKn, (ii) pn=2Kn+2satisfies (13), (iii) there is a positive constant D1such that Ink is a subset of some open ball with radius D1/Knin Y×T×V,and (iv) the density function f(YT)|Z=zV is Hölder continuous in (ytv)with the Hölder constant D0 and exponent d. Theorem 16 shows asymptotic properties of the confidence interval Cθn(α +δ).The first result (i) is the uniform asymptotic size control and the second result (ii) is the consistency against all the fixed alternatives. Theorem 16. Construct T(θπ) and c(αθπ) via the two-step multiplier bootstrap as in Appendix A.Suppose that Assumptions 13 and 14 hold.(i)The confidence interval controls the asymptotic size uniformly: liminf n→∞ inf P∈P0θ∈ΘI(P) Pθ∈Cθn(α +δ)≥1−α−δ (ii) If equation (12)holds,the confidence interval excludes all the fixed alternatives: lim n→∞Pθ∈Cθn(α +δ)=0for every (θP) ∈Θ×P0with θ/∈ΘI(P)
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1353 5. Empirical illustrations This section studies the effects of the 401(k) participation on financial savings using the inference method in Section 4. I introduce a measurement error problem to the analysis of Poterba, Venti, and Wise (1995) and Abadie (2003), which investigate the local average treatment effect using the eligibility for a 401(k) program. The robustness to misclassification is empirically relevant because the retirement pension plan type is subject to a measurement error in survey datasets. Using the Health and Retirement Study, for example, Gustman, Steinmeier, and Tabatabai (2008) estimated that around one-fourth of the survey respondents misclassified their pension plan type. The dataset in my analysis is from the Survey of Income and Program Participation (SIPP) of 1991. I follow the data construction in Abadie (2003). The sample consists of households in which at least one person is in employment, which has no income from self-employment, and whose annual family income is between $10,000 and $200,000. The resulting sample size is 9275. The outcome variable Yis the net amount of financial assets in dollars, and the measured treatment variable Tis the self-reported participation in a 401(k) program. The 401(k) participation can be endogenous because participants in a 401(k) program might be more informed or plan more about retirement savings than nonparticipants. To control for the endogeneity problem, I use the 401(k) eligibility, which is the indicator of whether an employer offers a 401(k) program, as an instrumental variable Z. I also use the participation in an individual retirement account (IRA) as the variable Rin Section 3.3. The control variables Vincludes constant, family income, age, age squared, marital status, and family size. The summary statistics for (YTZR) are in Table 1. To implement the proposed method, I impose Assumption 6in this empirical exercise. Assumption 6(i) requires that, given the covariates V, the 401(k) eligibility Zis exogenous for the self-report Tas well as the saving Yand the true 401(k) participation T∗. Poterba, Venti, and Wise (1995) justified the exogeneity of the 401(k) eligibility given the income level, based on that an employer determines the 401(k) eligibility wheres an employee determines saving and pension plan. Using the same reasoning, the self-report T is an employee’s choice and then the 401(k) eligibility Zis considered to be exogenous for the self-report T. Assumption 6(ii) holds because a 401(k) plan is not available unless an employee is eligible. Assumption 6(iii) requires that the 401(k) eligibility is not deterministic given the covariates V. Given that the 401(k) eligibility is an employer’s Table 1. Summary statistics for Y,T,Z,andR. Mean Standard Deviation Y: Family net financial assets 19,071 63,963 T: 401(k) participation 02762 04472 Z: 401(k) eligibility 03921 04883 R: IRA participation 02543 04355
1354 Takuya Ura Quantitative Economics 9 (2018) decision and the covariates Vare employee’s demographic variables, this condition is likely to hold. To construct the confidence interval Cπn(δ) for π, I use the linear probability model for the regression of the instrumental variable Zon the control variables V. I estimate βπin E[Z|V]=π(V ) =Vβπusing OLS and compute the confidence interval using a nonparametric bootstrap. To compare the proposed method with the exiting methods, I compute the Wald estimator, $16,290,witha95% bootstrapped confidence interval [597627,611].(TheWald estimator is the sample analogue of E[Z−π(V ) π(V )(1−π(V ))Y]/E[Z−π(V ) π(V )(1−π(V ))T]with πbeing estimated in the linear probability model.) The estimated intent-to-treat effect is $10,981 with a 95% bootstrapped confidence interval [416918,558]. For the construction of In1InKn, I construct a partition only over Y×Tto make the number of Knreasonable.7Using an integer Ln, I discretize the outcome variables Yusing {yτ:τ=01/Ln(Ln−1)/Ln1},whereyτis the unconditional τ-quantiles of Y. Define Kn=2Lnand, for every l=1Ln, define Inl ={(yt) :y(l−1)/Ln<y≤ yl/Lnt =0}and Inl+Ln={(yt) :y(l−1)/Ln<y≤yl/Lnt =1}.IuseLn=1234in the following analysis.8Although the data-driven choice of Lnis beyond the scope of this paper, a possible rule-of-thumb choice is the largest Lnfor which each Ink has at least a certain number of observations, for example, 30. In this empirical exercise, the corresponding choice is Ln=3. Table 2shows the confidence intervals for the local average treatment effect, where I use 2000 draws for the bootstraps and set β=01% for the moment selection and δ=1% for the estimation of βπ. The confidence intervals for the local average treatment effect are qualitatively similar to the confidence interval for the Wald estimand. The confidence intervals do not shrink as Lnincreases from 1to 4. It is possibly because the data generation process does not violate the conditions in (5) to a large extent and, therefore, the Wald estimand is close to the sharp upper bound for the local average treatment effect. Table 2. 90% and 95% confidence intervals for the local average treatment effect for different Ln.(BasedonTheorem7). Ln90% CI 95% CI 1[5743 25,287][4465 27,415] 2[5748 25,891][4461 28,062] 3[5707 26,081][4443 28,163] 4[5713 26,122][4430 28,296] 7Otherwise the number of the observations in some Ink’s can be small. Note that the resulting approximated identified set is still weakly tighter than the identified set based on the Wald estimand. Particularly when Ln=1, the condition in (11) is equivalent to using the Wald estimand as the upper bound for |θ|. 8The results are similar for larger values of Ln. Further results are available from the author upon request.
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1355 Table 3. 90% and 95% confidence intervals with using Ras in Theorem 9. Ln90% CI 95% CI 1[5741 25,829][4431 28,081] 2[5696 26,197][4417 28,588] 3[5652 26,487][4422 28,713] 4[5665 26,612][4418 28,846] Table 4. 90% and 95% confidence intervals without Tas in Theorem 11. Ln90% CI 95% CI 1[5781 Inf] [4485 Inf] 2[5741 81,973][4483 89,204] 3[5729 108,861][4455 118,218] 4[5725 84,999][4433 92,032] Table 3summarizes the confidence intervals with the IRA participation Ras an additional measurement discussed in Theorem 9.9It shows similar values to Table 2,and it can be because the IRA participation has only little identifying power on the local average treatment effect. Table 4summarizes the confidence intervals without using the measured treatment T,asinTheorem11.10 The lower bound of the confidence intervals does not change from those in Table 2, because the lower bound of the identified set does not depend on the measured treatment T. The upper bound is 3–4 times larger than those in Table 2, which is the cost of not using T.WhenLn=1,TVYbecomes zero, and then there is no finite upper bound on the local average treatment effect. In summary, the empirical results based on the proposed method are qualitatively similar to those results based on the Wald estimator. Given that this paper allows for a misclassification of the treatment variable, the empirical exercise has demonstrated the robustness of the existing results which ignore the measurement error (cf. Poterba, Venti, and Wise (1995)andAbadie (2003)). 6. Conclusion This paper studies the identifying power of an instrumental variable in the heterogeneous treatment effect framework when a binary treatment is mismeasured and endogenous. The assumptions in this framework are the monotonicity of the instrumen- 9In Table 3,IconstructIn1InKnas follows. Define Kn=4Lnand, for every l=1Ln, define Inl = {(ytr):y(l−1)/Ln<y≤yl/Lnt =0r =0},Inl+Ln={(ytr) :y(l−1)/Ln<y≤yl/Lnt =1r =0},Inl+2Ln= {(ytr):y(l−1)/Ln<y≤yl/Lnt=0r =1}, and Inl+3Ln={(ytr):y(l−1)/Ln<y≤yl/Lnt=1r =1}. 10In Table 4,IconstructIn1InKnas follows. Define Kn=Lnand, for every l=1Ln, define Inl ={(ytr):y(l−1)/Ln<y≤yl/Ln}.
1356 Takuya Ura Quantitative Economics 9 (2018) tal variable Zon the true treatment T∗and the exogeneity of Z. I use the total variation distance to characterize the identified set for the local average treatment effect E[Y1−Y0|T∗ 0<T ∗ 1]. I also provide an inference procedure for the local average treatment effect. Unlike the existing literature on measurement error, the identification strategy does not reply on a specific structure of the measurement error; the only assumption on the measurement error is its independence of the instrumental variable. There are several directions for future research. First, the choice of the partition In in Section 4, particularly the choice of Kn, is an interesting direction. To the best of my knowledge, the literature on many moment inequalities has not investigated how econometricians choose the numbers of the many moment inequalities. Second, it may be worthwhile to investigate the other parameter for the treatment effect. This paper has focused on the local average treatment effect, but the literature on heterogeneous treatment effect has emphasized the importance of choosing an adequate treatment effect parameter in order to answer relevant policy questions. Third, it is also interesting to investigate various assumptions on the measurement errors. In some empirical settings, for example, it may be reasonable to assume that the measurement error is one-directional (e.g., misclassification happens only when T∗=1). Fourth, it is not trivial how the analysis of this paper can be extended to an instrumental variable taking more than two values. For a general instrumental variable, it is always possible to focus on two values of the instrumental variable and apply the analysis of this paper to the subpopulation with the instrumental variable taking these two values. However, different pairs of the values can have different compliers, so that the parameter of interest is not common across different pairs, as discussed in Heckman and Vytlacil (2005). Appendix A: Multiplier bootstrap This section describes the construction of T(θπ) and c(αθπ) via the two-step multiplier bootstrap in Chernozhukov, Chetverikov, and Kato (2014). The two-step procedures involves inequality selection in the first step and testing in the second step. The first step uses β=βnas the size for selecting the moments. Chernozhukov, Chetverikov, and Kato (2014) imposed β<α/3and 1/βn≤C1log(n) where C1is the constant in Assumption 12. Define the moment functions based on (9)–(11). Define Hn={h1hpn−2}and g1(Wθπ)=− Z−π(V ) π(V )1−π(V )sgn(θ)Y g2(Wθπ)=Z−π(V ) π(V )1−π(V )sgn(θ)Y −|θ| g2+j(Wθπ)=Z−π(V ) π(V )1−π(V )|θ|hj(YTV)−sgn(θ)Y for 3≤j≤pn Then the approximated identified set is characterized by (θπ) ∈Θ×[01]:EPgj(Wθπ)for every j=1pn
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1357 The test statistic for the true parameter values being (θ π) is defined by T(θπ)=max 1≤j≤pn √nˆ mj(θπ) ˆσj(θπ) where ˆ mj(θπ) estimates mj(θπ) =EP[gj(Wθπ)],and ˆσ2 j(θπ) estimates the variance σ2 j(θπ) of √nˆ mj(θπ): ˆ mj(θπ) =n−1 n i=1 gj(Wiθπ) ˆσ2 j(θπ) =n−1 n i=1gj(Wiθπ)−ˆ mj(θπ)2 To conduct a multiplier bootstrap, generate nindependent standard normal random variables 1n. The centered bootstrap moments are ˆ mB j(θπ) =n−1 n i=1 igj(Wiθπ)−ˆ mj(θπ) The bootstrapped test statistic for inequality selection is defined by TB1(θπ) =max 1≤j≤pn √nˆ mB j(θπ) ˆσj(θπ) The critical value c1(βθπ)for inequality selection is defined as the conditional (1−β)- quantile of TB1(θπ) given {Wi}. Define ˆ JB=j=1pn:√nˆ mB j(θπ) ˆσj(θπ) >−2c1(βθπ) The bootstrapped test statistic is defined by TB(θπ) =max j∈ˆ JB √nˆ mB j(θπ) ˆσj(θπ) where TB(θπ) is defined as 0if ˆ JB=∅. The critical value c(αθπ) is defined as the conditional (1−α+2β)-quantile of TB(θπ) given {Wi}. Appendix B: Proofs Proof of Lemma 2 By equation (4), θ(P∗)EP∗[Y|Z]=θ(P∗)2P∗(T∗ 0<T ∗ 1)≥0,and|EP∗[Y|Z]| = |θ(P∗)P∗(T∗ 0<T∗ 1)|≤|θ(P∗)|.
1358 Takuya Ura Quantitative Economics 9 (2018) Proof of Lemma 3 I obtain f(YT)|Z=1−f(YT )|Z=0=P∗(T∗ 0<T∗ 1)(f(Y1T1)|T∗ 0<T∗ 1−f(Y0T0)|T∗ 0<T∗ 1)by the same logic as Theorem 1 in Imbens and Angrist (1994): f(YT)|Z=0=P∗T∗ 0=T∗ 1=1|Z=0f(YT)|Z=0T ∗ 0=T∗ 1=1 +P∗T∗ 0<T∗ 1|Z=0f(YT)|Z=0T ∗ 0<T∗ 1 +P∗T∗ 0=T∗ 1=0|Z=0f(YT)|Z=0T ∗ 0=T∗ 1=0 =P∗T∗ 0=T∗ 1=1f(Y1T1)|T∗ 0=T∗ 1=1 +P∗T∗ 0<T∗ 1f(Y0T0)|T∗ 0<T∗ 1 +P∗T∗ 0=T∗ 1=0f(Y0T0)|T∗ 0=T∗ 1=0 f(YT)|Z=1=P∗T∗ 0=T∗ 1=1f(Y1T1)|T∗ 0=T∗ 1=1 +P∗T∗ 0<T∗ 1f(Y1T1)|T∗ 0<T∗ 1 +P∗T∗ 0=T∗ 1=0f(Y0T0)|T∗ 0=T∗ 1=0 By the triangle inequality, TV(YT) =1 2 t=01f(YT )|Z=1(yt) −f(YT )|Z=0(yt)dμY(y) =1 2 t=01P∗T∗ 0<T∗ 1f(Y1T1)|T∗ 0<T∗ 1(yt) −f(Y0T0)|T∗ 0<T∗ 1(yt)dμY(y) =P∗T∗ 0<T∗ 11 2 t=01f(Y1T1)|T∗ 0<T∗ 1(yt) −f(Y0T0)|T∗ 0<T∗ 1(yt)dμY(y) ≤P∗T∗ 0<T∗ 11 2 t=01f(Y1T1)|T∗ 0<T∗ 1(yt) +f(Y0T0)|T∗ 0<T∗ 1(yt)dμY(y) =P∗T∗ 0<T∗ 1 Moreover, since T∗ 0≤T∗ 1almost surely, P∗T∗ 0<T∗ 1=1 2P∗T∗ 0=1−P∗T∗ 1=1+1 2P∗T∗ 0=0−P∗T∗ 1=0 =1 2 t∗=01fT∗|Z=1t∗−fT∗|Z=0t∗ =TVT∗
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1365 Abadie (2003) and Frölich (2007) showed that EPEP[X|ZV ]=EPZ−π(V ) π(V )1−π(V )X(15) for any random variable X.11 By Theorem 17 and equation (14), ΘI(P) is characterized by |θ|EPEPh(YTV ) |ZV ≤EPEP[Y|ZV ]for every h∈H θEPEP[Y|ZV ]≥0 |θ|≥EPEP[Y|ZV ] Since the second condition implies sgn(θ) =sgn(EP[EP[Y|ZV ]]), the above three conditions become |θ|EPEPh(YTV ) |ZV ≤sgn(θ)EPEP[Y|ZV ]for every h∈H sgn(θ)EPEP[Y|ZV ]≥0 |θ|≥sgn(θ)EPEP[Y|ZV ] By equation (15), ΘI(P) is characterized as in Lemma 12. Proof of Lemma 15 Condition (i) implies Assumption 14(i). Condition (ii) implies Assumption 14(iii). The rest of the proof is going to show Assumption 14(ii). Define h∗ P(ytv) =05× sgn(f(YT)|ZV =v(yt)) and h∗ Pn =argmaxh∈HnEP[EP[h(YT) |ZV ]].Thenh∗ P= 11The proof is as follows: EPZ−π(V ) π(V )1−π(V )X=EP1−π(V )Z+(Z −1)π(V ) π(V )1−π(V )X =EPZ π(V )X−EP1−Z 1−π(V )X =EPEP[ZX |V] π(V ) −EPEP(1−Z)X |V 1−π(V ) =EPπ(V )EP[X|Z=1V] π(V ) −EP1−π(V )EP[X|Z=0V] 1−π(V ) =EPEP[X|Z=1V]−EP[X|Z=0V] =EPEP[X|ZV ]
1366 Takuya Ura Quantitative Economics 9 (2018) argmaxh∈HEP[Z−π(V ) π(V )(1−π(V ))h(YTV )]. By the Hölder continuity of f(YT)|ZV , max (ytv)∈Ink f(YT )|ZV =v(yt) −min (ytv)∈Ink f(YT )|ZV =v(yt) ≤2D02D1 Knd Define Dn=2D0(2D1 Kn)d.ForInk with sup(ytv)∈Ink |f(YT)|ZV =v(yt)|>D n,theabove inequality implies that sgn(f(YT)|ZV =v(yt)) is constant on Ink. For those Ink,h∗=h∗ n on Ink. Then, on every Ink, either h∗=h∗ nor |f(YT )|ZV =v(yt)|≤Dn. Therefore, h∗ Pn(ytv)f(YT)|ZV =v(yt) ≥h∗ P(ytv)f(YT)|ZV =v(y t) +05×Dn Since EPZ−π(V ) π(V )1−π(V )h(YT V ) =EPEPh(YTV ) |ZV = h(ytv)f(YT)|ZV =v(y t)fV(v)μY(dy)μT(dt)μV(dv) = Kn k=1Ink h(ytv)f(YT)|ZV =v(y t)fV(v)μY(dy)μT(dt)μV(dv) it follows that EPZ−π(V ) π(V )1−π(V )h∗ Pn(YTV)≥EPZ−π(V ) π(V )1−π(V )h∗ P(YTV)+Dn Since Dnconverges to zero uniformly over P, Assumption 14(ii) holds. Proof of Theorem 16 The following theorem is taken from Corollary 5.1 and Theorem 6.1 in Chernozhukov, Chetverikov, and Kato (2014). Theorem 21. Given εn>0with εn→0and εnlogpn→∞,denote by H1n the set of (θπP)∈Θ×Π×P0for which π=P(Z |V=·)and max j=1pn mj(θ)/σj(θ) ≥(1+εn)2log(pn/α)/n (16) Under the assumptions in Theorem 16, (i)liminf n→∞ inf (θπP)∈Θ×Π×P0s.t.θ∈ΘI(P) and π=P(Z|V)PT(θπ)≤c(αθπ)≥1−α; (ii)lim n→∞ sup (θP)∈H1n PT(θπ)≤c(αθπ)=0
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1367 Theorem 16(i) follows from Pθ/∈Cθn(α +δ)≤Pθ/∈Cθn(α +δ)&π∈Cπn(δ)+Pπ/∈Cπn(δ) ≤PT(θπ)>c(αθπ)&π∈Cπn(δ)+Pπ/∈Cπn(δ) =PT(θπ)>c(αθπ)+Pπ/∈Cπn(δ)≤α+δ where the last inequality comes from Theorem 21(i) and Assumption 13(iv). Theorem 16(ii) is shown as follows. Denote by D2a constant for which σj(·)<D 2. Let (θP) be any element of Θ×P0with θ/∈ΘI(P). It suffices to show that (16) holds for sufficiently large n.Ifeither(6)or(7) is violated, then (16) holds for sufficiently large n. In the rest of the proof, I focus on the case where (8)isviolated.Thatis, sup h∈H EPZ−π(V ) π(V )1−π(V )|θ|h(YTV )−sgn(θ)Y >0 Since Hnconverges to Hin the sense of (12)and Z−π(V ) π(V )(1−π(V ))|θ|is bounded, it follows that, for sufficiently large n,thereish∈Hnsuch that EPZ−π(V ) π(V )1−π(V )|θ|h(YTV ) −sgn(θ)Y>0 Denoted by κthe value of the left-hand side in the above inequality. For sufficiently large n,thereisj=3pnsuch that mj(θ) > κ.Forsuchj,mj(θ)/σj(θ) > κ/D2. Therefore, (16) holds for sufficiently large n. References Abadie, A. (2003), “Semiparametric instrumental variable estimation of treatment response models.” Journal of Econometrics, 113, 231–263. [1338,1347,1353,1355,1365] Aigner, D. J. (1973), “Regression with a binary independent variable subject to errors of observation.” Journal of Econometrics, 1, 49–59. [1340] Amemiya, Y. (1985), “Instrumental variable estimator for the nonlinear errors-in- variables model.” Journal of Econometrics, 28, 273–289. [1341] Angrist, J. D., G. W. Imbens, and D. B. Rubin (1996), “Identification of causal effects using instrumental variables.” Journal of the American Statistical Association, 91, 444–455. [1336,1339,1341,1342] Angrist, J. D. and A. B. Krueger (2001), “Instrumental variables and the search for identification: From supply and demand to natural experiments.” Journal of Economic Perspectives, 15, 69–85. [1335] Balke, A. and J. Pearl (1997), “Bounds on treatment effects from studies with imperfect compliance.” Journal of the American Statistical Association, 92, 1171–1176. [1346] Black, D., S. Sanders, and L. Taylor (2003), “Measurement of higher education in the census and current population survey.” Journal of the American Statistical Association, 98, 545–554. [1335,1337,1339,1343]
1368 Takuya Ura Quantitative Economics 9 (2018) Black, D. A., M. C. Berger, and F. A. Scott (2000), “Bounding parameter estimates with nonclassical measurement error.” Journal of the American Statistical Association, 95, 739–748. [1340] Blundell, R., A. Gosling, H. Ichimura, and C. Meghir (2007), “Changes in the distribution of male and female wages accounting for employment composition using bounds.” Econometrica, 75, 323–363. [1341] Bollinger, C. R. (1996), “Bounding mean regressions when a binary regressor is mismeasured.” Journal of Econometrics, 73, 387–399. [1340] Bound, J., C. Brown, and N. Mathiowetz (2001), “Measurement error in survey data.” In Handbook of Econometrics, Vol. 5, Chapter 59 (J. Heckman and E. Leamer, eds.), 3705– 3843, Elsevier. [1337,1343] Calvi, R., A. Lewbel, and D. Tommasi (2017), “LATE with mismeasured or misspecified treatment: An application to women’s empowerment in India.” Working paper. [1341] Card, D. (2001), “Estimating the return to schooling: Progress on some persistent econometric problems.” Econometrica, 69, 1127–1160. [1335,1339] Carroll, R. J., D. Ruppert, L. A. Stefanski, and C. M. Crainiceanu (2012), Measurement Error in Nonlinear Models: A Modern Perspective, second edition. Chapman & Hall/CRC, Boca Raton. [1341] Chalak, K. (2017), “Instrumental variables methods with heterogeneity and mismeasured instruments.” Econometric Theory, 33, 69–104. [1341] Chernozhukov, V., D. Chetverikov, and K. Kato (2014), “Testing many moment inequalities.” Working paper. [1338,1351,1356,1366] de Chaisemartin, C. (2016), “Tolerating defiance? Local average treatment effects without monotonicity.” Quantitative Economics. (forthcoming). [1343] DiTraglia, F. J. and C. García-Jimeno (2015), “On mis-measured binary regressors: New results and some comments on the literature.” Working paper. [1340] Frazis, H. and M. A. Loewenstein (2003), “Estimating linear regressions with mismeasured, possibly endogenous, binary explanatory variables.” Journal of Econometrics, 117, 151–178. [1338,1340] Frölich, M. (2007), “Nonparametric IV estimation of local average treatment effects with covariates.” Journal of Econometrics, 139, 35–75. [1338,1347,1365] Gustman, A. L., T. Steinmeier, and N. Tabatabai (2008), “Do workers know about their pension plan type? Comparing workers’ and employers’ pension information.” In Overcoming the Saving Slump; how to Increase the Effectiveness of Financial Education and Saving Programs (A. Lusardi, ed.), 47–81, University of Chicago Press, Chicago. [1338, 1353] Hausman, J. A., W. K. Newey, H. Ichimura, and J. L. Powell (1991), “Identification and estimation of polynomial errors-in-variables models.” Journal of Econometrics, 50, 273– 295. [1338,1348]
Quantitative Economics 9 (2018) Heterogeneous treatment effects 1369 Hausman, J. A., W. K. Newey, and J. L. Powell (1995), “Nonlinear errors in variables estimation of some engel curves.” Journal of Econometrics, 65, 205–233. [1339] Heckman, J. J. and E. Vytlacil (2005), “Structural equations, treatment effects, and econometric policy evaluation.” Econometrica, 73, 669–738. [1346,1356] Henry, M., K. Jochmans, and B. Salanié (2015), “Inference on two-component mixtures under tail restrictions.” Econometric Theory. (forthcoming). [1340] Henry, M., Y. Kitamura, and B. Salanié (2014), “Partial identification of finite mixtures in econometric models.” Quantitative Economics, 5, 123–144. [1340] Hernandez, M. and S. Pudney (2007), “Measurement error in models of welfare participation.” Journal of Public Economics, 91, 327–341. [1339] Hsiao, C. (1989), “Consistent estimation for some nonlinear errors-in-variables models.” Journal of Econometrics, 41, 159–185. [1341] Hu, Y. (2008), “Identification and estimation of nonlinear models with misclassification error using instrumental variables: A general solution.” Journal of Econometrics, 144, 27– 61. [1339,1340] Hu, Y., J.-L. Shiu, and T. Woutersen (2015), “Identification and estimation of single index models with measurement error and endogeneity.” The Econometrics Journal, 18, 347– 362. [1341] Huber, M. and G. Mellace (2015), “Testing instrument validity for LATE identification based on inequality moment constraints.” Review of Economics and Statistics, 97, 398– 411. [1343,1346] Imai, K. and T. Yamamoto (2010), “Causal inference with differential measurement error: Nonparametric identification and sensitivity analysis.” American Journal of Political Science, 54, 543–560. [1340] Imbens, G. W. (2010), “Better LATE than nothing: Some comments on Deaton (2009) and Heckman and Urzua (2009).” Journal of Economic Literature, 48, 399–423. [1336] Imbens, G. W. and J. D. Angrist (1994), “Identification and estimation of local average treatment effects.” Econometrica, 62, 467–475. [1336,1339,1341,1342,1343,1344,1358] Kaestner, R., T. Joyce, and H. Wehbeh (1996), “The effect of maternal drug use on birth weight: Measurement error in binary variables.” Economic Inquiry, 34, 617–629. [1340] Kane, T. J., C. E. Rouse, and D. Staiger (1999), “Estimating returns to schooling when schooling is misreported.” Working paper. [1340] Kitagawa, T. (2010), “Testing for instrument independence in the selection model.” Working paper. [1341] Kitagawa, T. (2015), “A test for instrument validity.” Econometrica, 83, 2043–2063. [1346] Kreider, B. and J. V. Pepper (2007), “Disability and employment: Reevaluating the evidence in light of reporting errors.” Journal of the American Statistical Association, 102, 432–441. [1340]
1370 Takuya Ura Quantitative Economics 9 (2018) Kreider, B., J. V. Pepper, C. Gundersen, and D. Jolliffe (2012), “Identifying the effects of SNAP (food stamps) on child health outcomes when participation is endogenous and misreported.” Journal of the American Statistical Association, 107, 958–975. [1340] Lewbel, A. (1998), “Semiparametric latent variable model estimation with endogenous or mismeasured regressors.” Econometrica, 66, 105–121. [1341] Lewbel, A. (2007), “Estimation of average treatment effects with misclassification.” Econometrica, 75, 537–551. [1336,1339,1340] Mahajan, A. (2006), “Identification and estimation of regression models with misclassification.” Econometrica, 74, 631–665. [1336,1339,1340] Manski, C. F. (2003), Partial Identification of Probability Distributions. Springer-Verlag, New York. [1341] Molinari, F. (2008), “Partial identification of probability distributions with misclassified data.” Journal of Econometrics, 144, 81–117. [1340] Mourifié, I. Y. and Y. Wan (2016), “Testing local average treatment effect assumptions.” Review of Economics and Statistics. (forthcoming). [1346] Poterba, J. M., S. F. Venti, and D. A. Wise (1995), “Do 401(k) contributions crowd out other personal saving?” Journal of Public Economics, 58, 1–32. [1353,1355] Schennach, S. M. (2013), “Measurement error in nonlinear models—a review.” In Advances in Economics and Econometrics: Economic Theory,Vol.3(D.Acemoglu,M.Arellano, and E. Dekel, eds.), 296–337, Cambridge University Press. [1340] Song, S. (2015), “Semiparametric estimation of models with conditional moment restrictions in the presence of nonclassical measurement errors.” Journal of Econometrics, 185, 95–109. [1341] Song, S., S. M. Schennach, and H. White (2015), “Estimating nonseparable models with mismeasured endogenous variables on separable models with mismeasured endogenous variables.” Quantitative Economics, 6, 749–794. [1341] Ura, T. (2018), “Supplement to ‘Heterogeneous treatment effects with mismeasured endogenous treatment’.” Quantitative Economics Supplemental Material, 86, https://doi. org/10.3982/QE886.[1337,1343] Yanagi, T. (2017), “Inference on local average treatment effects for misclassified treatment.” Working paper. [1341] Co-editor Rosa L. Matzkin handled this manuscript. Manuscript received 19 May, 2017; final version accepted 19 February, 2018; available online 14 March, 2018.