Identification of average effects under magnitude and sign restrictions on confounding
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Chalak, Karim Article Identification of average effects under magnitude and sign restrictions on confounding Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Chalak, Karim (2019) : Identification of average effects under magnitude and sign restrictions on confounding, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 10, Iss. 4, pp. 1619-1657, https://doi.org/10.3982/QE689 This Version is available at: https://hdl.handle.net/10419/217176 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 10 (2019), 1619–1657 1759-7331/20191619 Identification of average effects under magnitude and sign restrictions on confounding Karim Chalak Department of Economics, University of Virginia This paper studies measuring various average effects of Xon Yin general structural systems with unobserved confounders U, a potential instrument Z,anda proxy Wfor U.WedonotrequireXor Zto be exogenous given the covariates or Wto be a perfect one-to-one mapping of U. We study the identification of coefficients in linear structures as well as covariate-conditioned average nonparametric discrete and marginal effects (e.g., average treatment effect on the treated), and local and marginal treatment effects. First, we characterize the bias, due to the omitted variables U, of (nonparametric) regression and instrumental variables estimands, thereby generalizing the classic linear regression omitted variable bias formula. We then study the identification of the average effects of Xon Ywhen Umay statistically depend on Xand Z. These average effects are point identified if the average direct effect of Uon Yis zero, in which case exogeneity holds, or if Wis a perfect proxy, in which case the ratio (contrast) of the average direct effect of Uon Yto the average effect of Uon Wis also identified. More generally, restricting how the average direct effect of Uon Ycompares in magnitude and/or sign to the average effect of Uon Wcan partially identify the average effects of Xon Y. These restrictions on confounding are weaker than requiring benchmark assumptions, such as exogeneity or a perfect proxy, and enable a sensitivity analysis. After discussing estimation and inference, we apply this framework to study earnings equations. Keywords. Causality, confounding, endogeneity, omitted variable bias, partial identification, proxy, sensitivity analysis. JEL classification. C31, C35, C36. Karim Chalak: [email protected] I thank the participants in the Northwestern Junior Festival on New Developments in Microeconometrics, Harvard Causal Inference Seminar, 2012 CA Econometrics Conference, 2013 BU-BC Green Line Econometrics Conference, 2013 North American Winter Meeting of the Econometric Society, NY Camp Econometrics VIII, 23rd meeting of the Midwest Econometrics Group, 9th Greater NY Metropolitan Area Econometrics Colloquium, Cowles Foundation Conference on Econometrics, and the seminars at BC, Cleveland Fed, UCSD, UCLA, USC, Pitt, IUPUI, UW-Milwaukee, Oxford, Royal Holloway, University of Leicester, UVA, University of Montreal, Georgetown, Virginia Tech, Vanderbilt, and UNC as well as Kate Antonovics, Andrew Beauchamp, Stéphane Bonhomme, Federico Ciliberto, Donald Cox, Julie Cullen, Stefan Hoderlein, Arthur Lewbel, Matthew Masten, Elie Tamer, and especially John Pepper for helpful comments. I acknowledge support from the Boston College Research Incentive and Expense Grants. I thank Rossella Calvi, Daniel Kim, and Tao Yang for excellent research assistance. Any errors are the author’s responsibility. ©2019 The Author. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE689
1620 Karim Chalak Quantitative Economics 10 (2019) 1. Introduction When measuring causal effects in observational studies, researchers often consider the unobserved variables that may jointly drive the cause and response of interest. For example, when estimating the financial return to education, researchers consider the unobserved individual “ability” that may jointly affect educational attainment and wage. Similarly, when estimating the elasticity of output with respect to the labor input, researchers consider the unobserved firm productivity that may jointly affect the input demands (e.g., capital and labor) and the output. A standard assumption that is useful to point identify average effects is the exogeneity (unconfoundedness) of the treatment or the instrument given the covariates. For example, to estimate the return to education researchers sometime assume that educational attainment, or an instrumental variable that is related to educational attainment such as the distance to a college, does not depend on ability given the covariates (see, e.g., Card (1995)). Similarly, to estimate production functions, researchers may consider using the prices of inputs as instrumental variables that are related to the inputs and unrelated to productivity (see, e.g., the discussion in Griliches and Mairesse (1998)). A second common assumption that is useful to point identify average effects requires that there is a perfect one-to-one proxy for the unobserved confounders. For example, a researcher may rely on a test score as a measure of ability (see, e.g., the discussion in Neal and Johnson (1996)). Similarly, a researcher may assume that, conditional on the capital input, investment or an intermediate input is a perfect proxy for the firm’s productivity (see, e.g., Olley and Pakes (1996)andLevinsohn and Petrin (2003)). These standard assumptions are not directly testable and researchers often ponder their validity. In particular, researchers sometimes question whether selection on unobservables leads conditional exogeneity to fail. For example, Carneiro and Heckman (2002) provided evidence suggesting that several commonly employed instruments in the ability literature may be endogenous and Griliches and Mairesse (1998) discussed how input prices may fail to be valid instruments when estimating production functions. Also, researchers often question whether a proxy is a perfect coding of the unobserved confounders. For example, a test score may be an error-laden measure of ability (see, e.g., Bollinger (2003)). Similarly, the firm’s investment or intermediate input may fail to be a strictly monotonic function of its productivity if there is an optimization or measurement error in how this proxy variable is determined or coded.1 Given the important role that the assumptions of conditional exogeneity and perfect proxy play in estimating average effects in a variety of empirical settings, it is useful to study the consequences of a possible departure from these benchmark assumptions. In order to do so, this paper characterizes the bias that standard estimands of various average effects would incur in the presence of omitted variables (unobserved confounders). It then demonstrates how restrictions on confounding that are weaker than requiring conditional exogeneity or a perfect proxy can partially identify these average effects. This enables a sensitivity analysis through which a researcher may gain 1Also, Ackerberg, Caves, and Frazer (2015) discussed how the perfect proxy assumption can sometimes render the input variables functionally dependent, complicating the identification of their effects on the output.
Quantitative Economics 10 (2019) Identification of average effects 1621 confidence in a causal effect estimate that is not highly sensitive to deviations from a maintained assumption. In particular, the paper studies identifying and estimating various conditional average effects of the treatment Xon the response Yin structural systems with unobserved confounders U, a potential (possibly invalid) instrument Z,and aproxyWfor U.WedonotrequireXor Zto be conditionally exogenous, and thus Umay statistically depend on Xand Z.Further,wedonotrequireWto be a perfect proxy, that is, a one-to-one mapping of U. The framework encompasses general specifications; we study the identification of coefficients in a linear structure as well as of covariate-conditioned average nonparametric discrete and marginal effects (e.g., average treatment effect on the treated), local average treatment effect, and marginal treatment effect. The analysis proceeds in two steps. The first step studies the consequences of omitted variables on the identification of average effects via standard estimands in the general specifications that this paper considers. In the case of linear homogenous effects, the linear regression omitted variable bias (OVB) representation is a classic result in econometrics (see, e.g., Stock and Watson (2010, Chapter 6); Wooldridge (2012,Chapter 3)). For instance, Angrist and Pischke (2009, p. 62) stated that the linear regression OVB formula “is one of the most important things to know about regression.” What is the analogue of the OVB formula in the cases of standard estimands, such as nonparametric regression and instrumental variables (IV) estimands (e.g., Wald (1940) or local IV estimands), for the various average effects described above? The first contribution of this paper is to characterize the OVB formula in these cases thereby generalizing the classic linear regression OVB representation. This enables studying the direction of the OVB, including in nonparametric nonseparable structures with heterogenous effects. The second step of the analysis demonstrates how an imperfect proxy Wfor Ucan be used to either point or partially identify various average effects of Xon Y.Inparticular, these average effects are point identified in two special cases. The first occurs if Xor Zis exogenous. It suffices for exogeneity that Uis unassociated (in a precise statistical sense) with Xor Z. This condition is testable under our assumptions since it implies that Wis unassociated with Xor Z. Alternatively, when Umay statistically depend on Xand Zas we allow, exogeneity holds if one assumes that the average direct (i.e., holding Xfixed) effect of Uon Yis zero (e.g., the average effect of ability on wage is zero). The second special case in which these average effects of Xon Yare point identified occurs if the proxy Wis a perfect one-to-one mapping of U(e.g., a test score is a perfect proxy for ability). In this case, the ratio (contrast) of the average direct effect of Uon Yto the average effect of Uon Wis also point identified. More generally, a researcher may impose restrictions on how the average direct effect of Uon Ycompares in magnitude and/or sign to the average effect of Uon Wthat are weaker than the restrictions obtained when assuming exogeneity or a perfect proxy. Moreover, this comparison may be informed by economic theory and evidence, as illustrated in the paper’s empirical application. The second contribution of this paper is to demonstrate how these magnitude and/or sign restrictions on confounding can point or partially identify the various average effects of Xon Y. This enables a researcher to analyze the sensitivity of the causal effects estimates to deviations from the benchmark assumptions of exogeneity
1622 Karim Chalak Quantitative Economics 10 (2019) and perfect proxy and can help clarify the extent to which the empirical estimates hinge on these identifying assumptions. The paper is organized as follows. Section 2describes the paper’s basic framework. Section 3states the data generation assumption. We derive the OVB formulas and characterize the sharp identification regions under restrictions on confounding for constant coefficients in a linear structure in Section 4, nonparametric average discrete and marginal effects in Section 5, and local and marginal treatment effects in Section 6. Section 7discusses estimation and inference. Section 8applies this paper’s framework to study the return to education and the black–white wage gap. Section 9concludes. Appendix2A in the Online Supplemental Material (Chalak (2019)) contains extensions. Mathematical proofs are gathered in the Online Supplemental Material in Appendix B. 2. Basic framework and overview 2.1 Linear equations with homogenous effects To illustrate the paper’s main ideas, consider an earnings structural equation (see, e.g., Mincer (1974)andCard (1999)), frequently employed in empirical work, given by Y=X¯ β+U¯ δY+U Y¯αY(1) The researcher observes realizations of the logarithm of hourly wage, Y, and of determinants Xof wage, such as years of education (and the level and square of years of experience). U, commonly referred to as “ability” in the literature, denotes unobserved skill, and UYcollects additional unobservables (disturbances). To introduce the main ideas in their simplest form, we let Ube a scalar and consider homogenous (constant) linear effects ¯ β,¯ δY,and ¯αY. Also, we leave any additional covariates implicit. However, as discussed below, we emphasize that this paper’s approach does not require homogenous effects or a parametric or separable specification. Our object of interest is the (average) effect ¯ βof Xon Y, for example, the (average) financial return to education. Although the return to education is homogenous in this example, and thus does not depend on U,abilityUis freely associated with Xand may cause Y(e.g., the educational attainment and wage may depend on ability). Thus, Uis an unobserved “confounder” or “omitted variable” and Xis potentially “endogenous.” The researcher observes realizations of a vector Zof potential instrumental variables. By definition, UYcollects the unobservables that may drive Yandareassumedtobe uncorrelated with Z(given the covariates) whereas U(which may be a vector more generally in Section 4) is freely correlated with Z. This allows a potential instrument for education, for example, the proximity to a college, to be invalid if it is correlated with 2The Appendices are located in the Online Supplemental Material and are available in the Replication File (Chalak (2019)). Section A.1 extends the constant coefficients analysis to discuss exogenous random coefficients and conditioning on covariates, additional empirical estimates, a panel structure, and proxies included in the Yequation. Section A.2 studies the special case where Uenters the Yand Wequations additively separably. Section A.3 studies the nonparametric nonseparable case with discrete U. Appendix B in the Online Supplemental Material collects the mathematical proofs.
Quantitative Economics 10 (2019) Identification of average effects 1623 ability U, for example, due to unobserved parental characteristics or choices. We let Z and Xhave the same dimension and Cov(ZX) be nonsingular. In particular, the most basic case arises when Zequals X. The linear IV regression OVB (or inconsistency) Bin recovering ¯ βis given by B≡Cov(ZX)−1Cov(ZY) −¯ β=Cov(ZX)−1Cov(ZU)¯ δY and this expression may help clarify the direction of the OVB. Exogeneity requires Cov(U ¯ δY+U Y¯αYZ)=0in which case B=0. By definition of Uand UY,Cov(UYZ)=0 and we allow Cov(UZ) =0. Thus, exogeneity is guaranteed to hold only if ¯ δY=0,that is, the average direct (holding Xfixed) effect of Uon Yis zero. The expression for Bis the IV analogue of the classic regression OVB formula and reduces to it in the special case when Z=X. The researcher may observe realizations of an error-laden proxy Wfor Ugiven by W=U¯ δW+U W¯αW(2) where the unobservables UWmay be correlated with UY,U,andX. For now, we consider a linear equation for Wwith constant ¯ δWand ¯αW. For example, Wmay denote the logarithm of a test score commonly used as a proxy for ability, such as Intelligence Quotient (IQ) or Knowledge of the World of Work (KWW). This parsimonious specification facilitates comparing the coefficients on Uin the Yand Wequations while maintaining the commonly used log-level specification for the wage equation. In particular, ¯ δY and ¯ δWare the semi-elasticities of the wage and test score with respect to ability,3that is, 100 ¯ δY%and 100 ¯ δW%are the (average) approximate percentage changes in the wage (with Xfixed) and test score due to a unit increase in U. Sometimes, researchers consider conditioning on the proxy to control for endogeneity. Provided ¯ δW=0, substituting for Uin equation (1)gives Y=X¯ β+W¯ δY ¯ δW+U Y¯αY−U W¯αW¯ δY ¯ δW (3) If UWis degenerate then Wis a perfect one-to-one proxy for Uand, provided Cov[UY (ZU)]=0, a linear IV regression of Yon (1XW)using instruments (1ZW) may point identify the (average) effect ¯ βof Xon Yas well as ¯ δY ¯ δW, the ratio of the (average) direct effect of Uon Yto the (average) effect of Uon W. This result fails to hold when UWis nondegenerate because a nonzero4Cov(UWZ|W)can lead to an IV regression bias5in recovering ¯ β.Moreover, ¯ βis “underidentified” in equation (3) since 3One could also consider standardizing the variables in equations (1) and (2), in which case the slope coefficients on the standardized ability denote standard deviation shifts in wage (holding Xfixed) and the test score respectively due to a standard deviation shift in ability. 4From W=U¯ δW+U W¯αW,Cov(UWU|W)is generally nonzero. Since Z(or X) and Uare freely correlated, it follows that UWis generally correlated with Z(or X)givenW. 5In general models, conditioning on Wmay, but need not, attenuate the regression bias (see, e.g., Wickens (1972), Battistin and Chesher (2014), and Ogburna and VanderWeele (2012)).
1624 Karim Chalak Quantitative Economics 10 (2019) Zand Xhave the same dimension (recall Zmay equal X) and there are fewer exogenous instruments for (XW)than is needed for an IV regression to point identify (¯ β¯ δY ¯ δW). Instead of assuming that Wis a perfect proxy with UWdegenerate and Cov[UY (ZU)]=0, we consider the weaker restriction Cov[(U YU W)Z]=0. Then the IV OVB is given by B=Cov(ZX)−1Cov(ZU)¯ δY=Cov(ZX)−1Cov(ZW ) ¯ δY ¯ δW Note that the condition Cov(UZ) =0, which ensures exogeneity (B=0), is testable under these assumptions since it implies Cov(W Z) =0. Importantly, if Cov(U Z) =0then the IV OVB is known up to the ratio ¯ δY ¯ δWof the (average) direct effect of Uon Yto the (average) effect of Uon W.Inparticular, ¯ βis characterized by ¯ β=Cov(ZX)−1Cov(ZY) −Cov(ZX)−1Cov(ZW ) ¯ δY ¯ δW This expression for ¯ βinvolves two linear IV regression estimands. It also involves the unknown ratio ¯ δY ¯ δW. As we show, analogous expressions obtain for various nonparametric average effects. 2.2 Magnitude and sign restrictions on confounding In the linear homogenous case above as well as the nonparametric heterogenous cases discussed below, we ask the following question: How does the average direct effect of U on Ycompare in magnitude or sign (or both) to the average effect of Uon W? The paper demonstrates how the answer to this question imposes restrictions on the magnitude and sign of confounding (e.g., on ¯ δY ¯ δWin equations (1)and(2)) that can point or partially identify the average effect of Xon Y(e.g., ¯ βin (1)). We do not require a particular answer to this question. Instead, we characterize the mapping6from every possible answer to the corresponding identification region for the average effect of Xon Y. To keep the scope of the paper manageable, we focus on restricting the support of, for example, ¯ δY ¯ δW rather than imposing more general prior distributions. In particular, when Umay statistically depend on Xand Z, the average effect of Xon Y(e.g., ¯ β) is point identified in the following three special cases. The first case is exogeneity which obtains when one assumes that the average direct effect of Uon Yis zero (e.g., ¯ δY ¯ δW=0). The second special case assumes that Wis a perfect proxy, in which case the ratio (e.g., ¯ δY ¯ δW) of the average direct effect of Uon Yto the average effect of Uon Wis also point identified. The third special case is proportional confounding which assumes that the average direct effect of Uon Yis equal to a known proportion of the average effect of Uon W(e.g., ¯ δY ¯ δW=d). The 6Leamer (1983) suggested the slogan “the mapping is the message.”
Quantitative Economics 10 (2019) Identification of average effects 1625 paper demonstrates that weaker restrictions on how the average direct effect of Uon Y compares in magnitude and/or sign to the average effect of Uon W(e.g., |¯ δY|≤|¯ δW|, 0≤¯ δY ¯ δW,0≤¯ δY ¯ δW≤1,or ¯ δY ¯ δW∈[dLdH]where [dLdH]contains the perfect proxy estimate of ¯ δY ¯ δWor its 95% confidence interval) can partially identify the average effect of Xon Y and it characterizes the resulting sharp identification region. In this sense, restrictions on confounding may be used to weaken standard assumptions, such as exogeneity or a perfect proxy. Sometimes, economic theory and/or evidence can provide guidance on sign and magnitude restrictions on confounding. For example, in the earnings and proxy equations (1)and(2), it may be reasonable to assume that, given the observables, a change in ability may cause an average direct percentage change (elasticity) in wage that is smaller in magnitude than the resulting average percentage change in the test score, that is, |¯ δY|≤|¯ δW|. Moreover, these average effects may be in the same direction, that is, 0≤¯ δY ¯ δW. These assumptions are in accord with several theoretical and empirical findings. For instance, Cawley, Heckman, and Vytlacil (2001) find that the fraction of wage variance explained by measures of cognitive ability is modest and that personality traits are correlated with earnings primarily through schooling attainment. Provided that ability measures, such as IQ or KWW, are sufficiently associated with unobserved ability U,this suggests that the average direct effect of Uon Ymay be modest. Second, when ability is not revealed to employers, they may statistically discriminate based on observables such as education (see, e.g., Altonji and Pierret (2001)andArcidiacono, Bayer, and Hizmo (2010)). This also suggests a modest average direct effect of Uon Y. Last, recall that if one assumes ¯ δY ¯ δW∈[dLdH]then an estimate for ¯ βcorresponds to each d∈[dLdH].Thus,in determining restrictions on confounding, a researcher may ask: what restrictions on ¯ δY ¯ δW are in accord with plausible features of ¯ β? For example, the empirical findings in this paper corroborate the assumption ¯ δY ¯ δW≤1in the earnings equation since values of ¯ δY ¯ δW that exceed 1lead to negative estimates of the average return to education and to an estimated black–white wage gap in favor of blacks, which is unlikely and inconsistent with the general findings in the literature. In a nutshell, magnitude and sign restrictions on confounding serve as a means for identification when stronger assumptions, such as exogeneity or a perfect proxy, may fail to hold7and they enable examining the sensitivity of a study’s estimates to deviations from these standard assumptions. Of course, a particular restriction on confounding may in turn fail to hold. For example, an imposed sign and/or magnitude restriction on ¯ δY ¯ δWmay be invalid or there may be additional omitted variables unaccounted for in the analysis. Thus, researchers may want to carefully consider a range of restrictions on confounding in a sensitivity analysis. Nevertheless, our goal here is to provide a framework in which restrictions on confounding can be used to weaken the benchmark assumptions of exogeneity and a perfect proxy which are employed in a vast literature, leading to more robust and credible causal estimates. 7Even if stronger or alternative assumptions hold, restrictions on confounding may yield tighter confidence intervals.
1626 Karim Chalak Quantitative Economics 10 (2019) 2.3 Nonparametric nonseparable equations An advantage of this paper’s approach is that it does not require a parametric or separable specification. Sections 5and 6give key nonparametric results. Section 5focuses on the case in which there are no excluded instruments (Z=X) and then generalizes the linear specification of Section 4to let the outcome Yand proxy Wbe generated by Y=r(XSUUY)and W=q(SUUW) (4) Here, the vector of observed covariates S, the scalar8confounder U, and the vector of unobservables UYinteract nonseparably with Xto drive Yaccording to the unknown nonparametric structural function r. For example, this generalizes the linear specification in the “correlated random coefficient” model.9Similarly, Uand the unobserved vector UW interact nonseparably with Sto drive Waccording to the unknown nonparametric function q. Suppose that10 UY⊥(U X)|Sso that the component Uof the “exogenous treatment” (U X) is unobserved and, unlike UY, may statistically depend on Xgiven the covariates S.Inparticular,ifU|Sis degenerate then there are no omitted variables and we obtain the standard assumption of conditional exogeneity (or unconfoundedness) UY⊥X|S. First, we characterize the OVB of nonparametric regression methods in recovering the covariate-conditioned average discrete or marginal effects of Xon Y,thereby generalizing the classic linear regression OVB formula to the nonparametric nonseparablecase.Second,weshowthatifWis a perfect proxy for U(with UWdegenerate and q strictly monotonic in Ugiven S) then the conditional average effect of Xon Yis point identified, as is the ratio of the conditional average effect of Uon Yto that of Uon W. More generally, if the proxy Wis imperfect and UW⊥(UX)|Sthen restrictions on how the magnitude or sign (or both) of the conditional average effect of Uon Ycontrasts with that of Uon Wcan partially identify the conditional average effect of Xon Y. Section 6focuses on the case where Xis binary and then generalizes the specification in Section 5by allowing Zto differ from X. Specifically, it augments equations (4) with a treatment selection equation where UXis an unobserved variable and the function νis unknown: X=1UX≤ν(ZS)(5) Suppose (UXUY)⊥(ZU)|S.IfU|Sis degenerate then there are no omitted variables and we obtain the standard assumptions of “monotonicity” and exogeneity, (UXUY)⊥Z|S,ofZ.Otherwise,Umay statistically depend on the potential instrument Zgiven S. First, we characterize the OVB of the Wald and local IV estimands for the conditional local and marginal treatment effects (LATE and MTE). Then we use restrictions on confounding, that are weaker than exogeneity or a perfect proxy, to partially identify the conditional LATE and MTE. 8The Online Supplemental Material in Appendix A.2 considers a vector Uthat enters radditively separably. 9The correlated random coefficient model restricts rsuch that Y=r(XUUY)=β(U UY)X + αY(UUY), with a random intercept αY(U UY)and slope β(UUY). For example, in this special case ability Uaffects the wage Ythrough both αY(·)and the linear return β(·)to education X(see, e.g., Card (2001)). 10A⊥B|S=sdenotes conditional independence given S=sas in Dawid (1979). ⊥denotes dependence.
Quantitative Economics 10 (2019) Identification of average effects 1633 5. Identification of average nonparametric effects Section 5focuses on the case in which there are no excluded instruments, that is, Z=X, and then extends Section 4’s analysis by removing the linearity assumption15 S.2. Here, we study the identification of the conditional average effect of Xon Ywhen changing x to x∗given X=x∗: ¯ βxx∗|x∗≡Erx∗UUY−r(xUUY)|X=x∗ For instance, for binary X,¯ β(01|1)is the average treatment effect on the treated. If ris differentiable in a scalar cause of interest, we set k=1to denote this variable by Xand we subsume, without loss of generality, the remaining causes into the implicit covariates. We then study the identification of the conditional average marginal effect of Xon Yat xgiven X=x: ¯ β(x|x) ≡E∂ ∂xr(xUUY)X=x In studying the identification of ¯ β(xx∗|x∗)and ¯ β(x|x), we use a shorthand notation for the difference and derivative of a nonparametric regression. Specifically, for random vectors Aand Bwith E(A) finite and band b∗in the support16 of B,welet RN ABbb∗≡EA|B=b∗−EA|B=b Further, when Bis a scalar and the derivative exists, we write RN AB(b) ≡∂ ∂bEA|B=b Theorem 5.1 characterizes the nonparametric bias B(xx∗|x∗)or B(x|x) of the nonparametric regression estimand RN YX(xx∗)or RN YX(x) in recovering the average effect ¯ β(xx∗|x∗)or ¯ β(x|x) in the presence of an omitted variable U. While both Uand UYcan generate heterogeneity in the response of Yto X,Uis the only source of endogeneity of X(or of “essential heterogeneity” in the nomenclature of Heckman, Urzua, and Vytlacil (2006)). Specifically, Theorem 5.1 imposes the local (at xand x∗) mean independence condition,17 Erx†uUY|U=uX =¨ x=Erx†uUY for x†¨ x∈xx∗and all u∈U¨ x(6a) 15Appendix A.2 in the Online Supplemental Material studies an intermediate case in which Uenters r and qadditively separably. 16Throughout, for random vectors Aand B, we denote the cumulative distribution function (cdf) of A by FA(·)and that of Aconditional on B=bby FA|B(·|b). We let the corresponding probability density or mass functions be fA(·)and fA|B(·|b), respectively. We denote the support of Aby Aand that of A|B=bby Ab. 17Mean independence conditions are often employed to identify average causal effects. See, for example, Manski (1990) and Heckman, Ichimura, and Todd (1998).
1634 Karim Chalak Quantitative Economics 10 (2019) In the case of the marginal effect ¯ β(x|x),Theorem5.1 further imposes the local (at x) condition, E∂ ∂xr(xuUY)U=uX =x=E∂ ∂xr(xuUY)for all u∈Ux(6b) Note that UY⊥(UX) implies18 (6a)and(6b). If Uis degenerate, then there are no omitted variables and UY⊥(UX) reduces to the standard exogeneity condition UY⊥X(or more generally UY⊥X|Sas in, e.g., Altonji and Matzkin (2005), Hoderlein and Mammen (2007), and Imbens and Newey (2009)19). In this special case, the weaker local mean independence conditions20 (6a)and(6b)sufficeforRN YX(xx∗)or RN YX(x) to point identify ¯ β(xx∗|x∗)or ¯ β(x|x). Similarly, Theorem 5.1 imposes the analogous condition to (6a)fortheproxyequation: Eq(uUW)|U=uX =¨ x=Eq(uUW)for ¨ x∈xx∗and u∈U¨ x(7) so that Wis an informative proxy with the mean dependence of Won Xat xand x∗ arising solely due to U.Here,UW⊥(UX) implies the local mean independence condition (7). Using Theorem 5.1’s characterization, Corollary 5.2 partially identifies ¯ β(xx∗|x∗)or ¯ β(x|x) by imposing magnitude and sign restrictions on the average marginal effects of the omitted variable Uon the response Yat (xu) and on the proxy variable Wat u, denoted by ¯ δY(u;x) ≡E∂ ∂ur(xuUY)and ¯ δW(u) ≡E∂ ∂uq(uUW) For brevity, Theorem 5.1 states the results in the case where rand qare differentiable in uand the distribution of Ugiven Xis continuous or, as a limiting case, degenerate.21 Theorem A.7 in Appendix A.3 in the Online Supplemental Material gives the results for discrete U, with sums replacing integrals. To proceed, we collect in Assumption B.1 regularity conditions that ensure that the moments and derivatives exist and justify interchanging the order of the derivative and integral in expressions such as RN YX(x) =∂ ∂x UxE[r(xuUY)]fU|X(u|x)du,whereweuse(6a). For this, B.1 also lets Ux be constant in a neighborhood of x(or Ux=Ux∗in the case of ¯ β(xx∗|x∗))toremovethe complication introduced by the boundary terms. To give a stronger simpler condition, 18UY⊥(UX) is not necessary, for instance, Theorem 5.1 permits Var(UY|X) to depend on X. 19Similar to Imbens and Newey (2009), one can consider covariates S2and a scalar unobserved S1recov- erable from a choice equation X=˜ p(Z S2S1)with ˜ pmonotonic in S1,suchthat(UYS1)⊥Z|S2, yielding UY⊥X|Swith S=(S1S 2). We allow but do not require this possibility. 20For example, if Uis degenerate at uand Xis binary then condition (6a) states that the potential outcomes r(0uUY)and r(1uUY)are mean independent of the treatment X. 21It may be convenient to view the case in which U|X=xis degenerate at u(x) as a limiting case for a sequence of absolutely continuous Fτ U|X(u|x) as τ→0. In particular, one can set FU|X(u|x) =H(u−u(x)) where H(·)is the Heaviside step function. Then, when u(x) is a differentiable function, ∂ ∂x FU|X(u|x) = −∂ ∂x u(x)δ(u −u(x)) where δ(·)is the Dirac delta function, with an impulse concentrated at u(x) (see, e.g., Bracewell (1986)).
Quantitative Economics 10 (2019) Identification of average effects 1635 given (6a), (6b), and (7), it suffices for B.1 that (i) Uis compact and Ux=Ufor all xin Xand that, for all values of the fixed argument(s) in (ii)–(iv), (ii) r(x·uy)and q(·uw) (resp., fU|X(·|x) and, for k=1,∂ ∂x fU|X(·|x) and ∂ ∂x r(x·uy)) are continuously differentiable (resp., continuous) on U, (iii) E[r(xuUY)q(uUW)]<∞,and(iv) ∂ ∂u r(xu·), ∂ ∂x r(xu·)for k=1,and ∂ ∂u q(u·)are each bounded in absolute value by integrable functions of uyand uw, respectively. Assumption B.1 in Appendix B in the Online Supplemental Material gives weaker local regularity conditions Theorem 5.1. Assume S.1 with m=l=1,x x∗∈X,and that FU|X(·|x) and FU|X(·|x∗) are absolutely continuous or,in the limit,degenerate. (i.a) If conditions B.1.i(a, b, c, d) and (6a)hold,then Bxx∗|x∗≡RN YXxx∗−¯ βxx∗|x∗=−Ux¯ δY(u;x)FU|Xu|x∗−FU|X(u|x)du (i.b) If conditions B.1.i(b, e, f, g) and (7)hold,then RN WXxx∗=−Ux¯ δW(u)FU|Xu|x∗−FU|X(u|x)du (ii) Set k=1. (ii.a) If conditions B.1.i(c, d), B.1.ii(a, b, c, d), (6a), and (6b)hold,then B(x|x) ≡RN YX(x) −¯ β(x|x) =−Ux¯ δY(u;x) ∂ ∂xFU|X(u|x)du (ii.b) If conditions B.1.i(f, g), B.1.ii(a, d, e), and (7)hold,then RN WX(x) =−Ux¯ δW(u) ∂ ∂xFU|X(u|x)du B(xx∗|x∗)and B(x|x) generalize the classic linear regression OVB formula to the nonparametric nonseparable case. These biases depend on the average marginal effect ¯ δY(u;x) of Uon Yand on the conditional distribution of U|X. This provides insight into the sign of the OVB. For instance, if ¯ δY(u;x) is nonnegative for a.e. u∈Ux(e.g., the average marginal effect of ability on wage is nonnegative) and the stochastic dominance relation FU|X(u|x∗)≤FU|X(u|x) for a.e. u∈Uxholds (e.g., the probability of low ability Uis small when education is high (x<x ∗)) then B(x x∗|x∗)is nonnegative. Under exogeneity, B(xx∗|x∗)=0and RN YX(xx∗)point identifies ¯ β(xx∗|x∗).This occurs if U⊥X, in which case RN WX(xx∗)=0,orif ¯ δY(u;x) =0for a.e. u∈Ux.Alternatively, suppose that W=q(UUW)≡˜ q(U) is a perfect proxy, with UWdegenerate and ˜ qstrictly monotonic in u(see, e.g., Olley and Pakes (1996)andGriliches and Mairesse (1998)). Then, substituting for U=˜ q−1(W ) in rand using condition (6a), we have that ¯ β(xx∗|x∗)is point identified (see, e.g., White and Chalak (2013, Theorem 4.2)): EEY|X=x∗W−E(Y|X=x W )|X=x∗=¯ βxx∗|x∗
1636 Karim Chalak Quantitative Economics 10 (2019) In this case, under UY⊥(UX),theratio ¯ δY(u;x) ¯ δW(u) is also point identified by ∂ ∂wE(Y|X=xW =w) =E∂ ∂ur(xuUY)∂ ∂w ˜ q−1(w)=¯ δY(u;x) ¯ δW(u) Last, if Wis an imperfect proxy as in Theorem 5.1 then ¯ β(xx∗|x∗)is point identified by RN YX(xx∗)−d(x)RN WX(xx∗)under proportional confounding, when ¯ δY(u;x) = d(x)¯ δW(u) for a.e. u∈Uxand d(x) is known. In this case, the average effects of Uon Y(at x)andWare assumed to be of a known proportion d(x) for a.e. u∈Ux. Analogous results hold for ¯ β(x). Corollary 5.2 characterizes the sharp identification regions for ¯ β(xx∗|x∗)and ¯ β(x|x) that obtain under weaker magnitude and/or sign restrictions on confounding. Corollary 5.2. Suppose that,for a.e.u∈Ux,¯ δY(u;x) =d(ux)¯ δW(u) with d(ux) ∈ D(x) ≡[dL(x)dH(x)]. (i) Under the conditions of Theorem 5.1(i), if ¯ δW(u)[FU|X(u|x∗)−FU|X(u|x)]is either nonpositive for a.e.u∈Uxor nonnegative for a.e.u∈Uxthen ¯ βxx∗|x∗∈BD(x)≡RN YXxx∗−RN WXxx∗d:d∈D(x) and this identification region is sharp. (ii) Under the conditions of Theorem 5.1(ii), if ¯ δW(u) ∂ ∂x FU|X(u|x) is either nonpositive for a.e.u∈Uxor nonnegative for a.e.u∈Uxthen ¯ β(x|x) ∈BD(x)≡RN YX(x) −RN WX(x)d :d∈D(x) and this identification region is sharp. If Uis degenerate, then exogeneity holds and Corollary 5.2’s bounds collapse to the nonparametric regression estimand. More generally, the conditions of Corollary 5.2 obtain if E[q(uUW)]is monotonic in u(e.g., on average, the test score is monotonic in ability) and FU|X(u|x∗)≤FU|X(u|x) for a.e. u∈Ux(e.g., the probability of low ability U is small when education is high). In particular, Corollary 5.2 analyzes the consequences of deviating from the exogeneity and perfect proxy assumptions by letting the set D(x) contain zero (recall that ¯ δY(·;x) =0ensures exogeneity) and/or estimates of ¯ δY(u;x) ¯ δW(u) for u∈Uxthat obtain under the perfect proxy assumption. The identification region B(D(x)) is sharp under the conditions22 in Corollary 5.2.Weleavestudyingtheconsequences of imposing stronger assumptions, such as UY⊥(UX) and UW⊥(UX), on the identification of ¯ β(xx∗|x∗)and ¯ β(x|x) to other work. To conclude, Section 5’s analysis removes the requirement that U|Sis degenerate in the conditional exogeneity condition UY⊥(XU)|Swhere Uis an omitted component 22As can be seen from the proof, B(D(x)) remains sharp if one strengthens the local conditions (6a), (6b) and (7) to require the stronger global mean independence conditions E[r(xuUY)|UX]=E[r(xuUY)] for all (xu) ∈X×Uand E[q(uUW)|UX]=E[q(uUW)]for all u∈U.
Quantitative Economics 10 (2019) Identification of average effects 1637 of the “exogenous treatment” (XU) given the covariates S. This analysis complements the results in Imbens (2003) who proposes a sensitivity analysis under an alternative weakening of exogeneity that views Uas an omitted covariate, with UY⊥X|(US) and U⊥S. Also, the results in Section 5relate to the literature that uses an error-laden measure23 Wof Uto point identify the effect of Uon Yunder auxiliary assumptions. Recent examples include Hu, Shiu, and Woutersen (2015,2016) who point identify the coefficient on a mismeasured endogenous variable in a single index model with exogenous instruments and either a separable equation for the latent variable or structural functions that are monotonic in certain unobservables. Last, multiple proxies for Uthat are mutually independent given U(see, e.g., Cunha, Heckman, and Schennach (2010)) may help identify the average nonparametric effect of Xon Y. Corollary 5.2’s bounds may be useful when multiple proxies are unavailable or mutually dependent given U. 6. Identification of local and marginal treatment effects Section 6focuses on the case where Xis binary and then extends Section 5’s analysis to allow Zto differ from X.Here,bothZand Xmay be endogenous. Specifically, we let Yand Wbe as in24 S.1 and consider a treatment Xgenerated via a threshold crossing selection equation, as in, for example, Heckman and Vytlacil (2005). As shown in Vytlacil (2002), under exogeneity of Z, S.3 is equivalent to the monotonicity assumption in, for example, Imbens and Angrist (1994). Assumption 3 (S.3). Assume S.1 and suppose further that Xis generated by25 X=1UX≤ν(Z) where νis an unknown real-valued function and UXis an unobserved random variable with FUX(·)absolutely continuous.We augment L≡(U XU WU YU)with UX. Under S.3, selection into treatment (X=1) holds if and only if ν(Z) exceeds UX. When interest attaches to a scalar potential instrument, we set =1to denote it by Zand we subsume, without loss of generality, the remaining potential instruments into the implicit covariates. We let FUX(·)be absolutely continuous to simplify the exposition. It is convenient to rewrite the equation for Yin its random coefficients form Y=r(1UUY)−r(0UUY)X+r(0UUY)≡β(UUY)X +αY(UUY) (8) Following the literature (e.g., Imbens and Angrist (1994)andHeckman and Vytlacil (2005)), we study the identification of the conditional local average treatment effect (LATE) ¯ βν(z) < UX≤νz∗z∗≡Eβ(UUY)|ν(z) < UX≤νz∗Z=z∗ 23This paper’s analysis does not require that the “measurement error” UWobeys UW⊥UY. 24Appendix A.2 in the Online Supplemental Material studies the case in which Uenters rand qadditively separably. 251{A}=1if Ais true and equals 0otherwise.
1638 Karim Chalak Quantitative Economics 10 (2019) This is the average treatment effect for the subpopulation with instrument Z=z∗and for whom X=0if Z=zwhereas X=1if Z=z∗.UnderUX⊥Z, averaging this local effect over the distribution of Zyields the LATE ¯ β(ν(z) < UX≤ν(z∗)).WhenZis binary, this latter effect is the average treatment effect for the “compliers” who receive the treatment (X=1) if and only if Z=1(see, e.g., Angrist, Imbens, and Rubin (1996)). Similarly, we study the identification of the conditional marginal treatment effect (MTE) ¯ βν(z)z≡Eβ(UUY)|UX=ν(z)Z =z Under UX⊥Z, averaging ¯ β(ν(z) ·)over the distribution of Zyields the MTE ¯ β(ν(z)),the average treatment effect for those who are indifferent toward receiving the treatment if Z=z. In studying the identification of LATE and MTE, we use the following succinct notation for the Wald and local instrumental variable (LIV) estimands. In particular, for random variable Band vectors Aand C, provided the means exist and the denominator is nonzero, define RWald AB|Ccc∗≡RN ACcc∗ RN BCcc∗≡EA|C=c∗−EA|C=c EB|C=c∗−E(B|C=c) Further, when Cis a scalar and the derivatives exist with nonzero denominator, let RLIV AB|C(c) ≡RN AC(c) RN BC(c) ≡ ∂ ∂c EA|C=c ∂ ∂c E(B|C=c) Theorem 6.1 characterizes the OVB of RWald YX|Z(zz∗)or RLIV YX|Z(z) in recovering ¯ β(ν(z) < UX≤ν(z∗)z∗)or ¯ β(ν(z) z) in the presence of an omitted variable U. It maintains that UX⊥(UZ) (9) Further, analogously to Theorem 5.1,Theorem6.1 restricts the local (at zand z∗) mean dependence of the random coefficients on (UZ) so that: Eα(uUY)|U=uZ =¨ z=Eα(uUY)and Eβ(uUY)|UXU =uZ =¨ z=Eβ(uUY)|UXfor ¨ z=zz∗and all u∈U¨ z(10) Note that (UXUY)⊥(UZ) implies conditions (9)and(10). Thus, Theorem 6.1 characterizes the bias B(ν(z) < UX≤ν(z∗)z∗)or B(ν(z)z) that arises when the exogenous component Uof the treatment (UX) is unobserved and possibly stochastically dependent on the instrument Zfor the endogenous treatment component X.IfUis degenerate, then there are no omitted variables and (UXUY)⊥(UZ) reduces to the standard instrument exogeneity condition (UXUY)⊥Zwhich ensures the assumptions in, for example, Heckman and Vytlacil (2005)andImbens and Angrist (1994). In this special case, condition (9) and the local mean independence condition (10)sufficefor
Quantitative Economics 10 (2019) Identification of average effects 1639 RWald YX|Z(zz∗)or RLIV YX|Z(z) to point identify ¯ β(ν(z) < UX≤ν(z∗)z∗)or ¯ β(ν(z) z).Analogously to α(uUY), we let the local mean dependence of Won Zarise solely due to U: Eq(uUW)|U=uZ =¨ z=Eq(uUW)for ¨ z=zz∗and all u∈U¨ z(11) Here, UW⊥(UZ) implies the local mean independence condition (11). Theorem 6.1 considers the case where rand qare differentiable in uand the distribution of Ugiven Zis continuous or, as a limiting case, degenerate. Here, too, Assumption B.2 collects regularity conditions that justify the operations involving derivatives and integrals and lets Uzbe constant in a neighborhood of z(or Uz=Uz∗inthecaseof LATE). For example, given (9), (10), and (11) (and provided the denominator RN XZ(z) = fUX(ν(z)) ∂ ∂z ν(z) is nonzero), it suffices for B.2 that (i) Uis compact and Uz=Ufor all zin Zand that, for all values of the fixed argument(s) in (ii)–(v), (ii) r(1{ux≤ν(z)}·uy)and q(·uw)(resp., fU|Z(·|z),E[β(·UY)|UX=ux],and ∂ ∂z fU|Z(·|z) for =1) are continuously differentiable (resp., continuous) on U, (iii) E[r(1{UX≤ν(z)}uUY)q(uUW)]<∞, (iv) E[β(uUY)|UX=·]and fUX(·)are continuous on UXand ν(·)is continuously differentiable on Z,and(v) ∂ ∂u r(1{·≤ν(z)}u·)and ∂ ∂u q(u·)are bounded in absolute value by an integrable function of (uxuy)and uw, respectively. Assumption B.2 in Appendix B in the Online Supplemental Material gives weaker local regularity conditions. We slightly abuse the previous ¯ δY(u;x) notation and denote the average marginal effects of Uon Y at (zu) and Wat uby ¯ δY(u;z) ≡E∂ ∂ur1UX≤ν(z)uUYand ¯ δW(u) ≡E∂ ∂uq(uUW) Theorem 6.1. Assume S.1 and S.3 with m=l=1,zz∗∈Z,Pr[ν(z) < UX≤ν(z∗)]>0, and that FU|Z(·|z) and FU|Z(·|z∗)are absolutely continuous or,in the limit,degenerate. (i.a) If conditions B.2.i(a, b, c, d), (9), and (10)hold,then Bν(z) < UX≤νz∗z∗ ≡RWald YX|Zzz∗−¯ βν(z) < UX≤νz∗z∗ =− 1 RN XZzz∗Uz¯ δY(u;z)FU|Zu|z∗−FU|Z(u|z)du (i.b) If conditions B.2.i(b, e, f, g) and (11)hold,then RWald WX|Zzz∗=− 1 RN XZzz∗Uz¯ δW(u)FU|Zu|z∗−FU|Z(u|z)du (ii) Set =1. (ii.a) If conditions B.2.i(c, d), B.2.ii(a, b, c, d, e), (9), and (10)hold,then Bν(z)z≡RLIV YX|Z(z) −¯ βν(z)z=− 1 RN XZ(z) Uz¯ δY(u;z) ∂ ∂z FU|Z(u|z)du (ii.b) If conditions B.2.i(f, g), B.2.ii(a, c, e, f), and (11)hold,then RLIV WX|Z(z) =− 1 RN XZ(z) Uz¯ δW(u) ∂ ∂z FU|Z(u|z)du
1640 Karim Chalak Quantitative Economics 10 (2019) Theorem 6.1 shows how the OVB of the Wald or LIV estimand for the conditional LATE or MTE depends on the average marginal effect of Uon Yand on the distribution of U|Z. For example, if ¯ δY(u;z) is nonnegative for a.e. u∈Uz(e.g., the average marginal effect of ability on wage is nonnegative) and FU|Z(u|z∗)≤FU|Z(u|z) for a.e. u∈Uz(e.g., the probability of low ability is small when in proximity to a college) then B(ν(z) < UX≤ ν(z∗)z∗)is nonnegative. The OVB B(ν(z)z) vanishes under exogeneity (e.g., when U⊥Z(and thus RLIV WX|Z(z) =0)or ¯ δY(u;z) =0for a.e. u∈Uz). Alternatively, if W=q(UUW)=˜ q(U) is aperfectproxy,withUWdegenerate and ˜ qstrictly monotonic, then using U=˜ q−1(W ), (9), and (10)gives E∂ ∂z E(Y|Z=zW )Z=z ∂ ∂z E(X|Z=z) =¯ βν(z)z In this case, under (UXUY)⊥(UZ),theratio ¯ δY(u;z) ¯ δW(u) is also point identified by ∂ ∂wE(Y|Z=zW =w) =E∂ ∂wr1UX≤ν(z)˜ q−1(w) UY=¯ δY(u;z) ¯ δW(u) Last, when Wis an imperfect proxy and proportional confounding holds (i.e., ¯ δY(u;z) = d(z)¯ δW(u) for a.e. u∈Uzwith d(z) known), then RLIV YX|Z(z) −RLIV WX|Z(z)d(z) point identifies ¯ β(ν(z) z). Analogous results hold for the LATE ¯ β(ν(z) < UX≤ν(z∗)z∗). Restrictions on confounding that are weaker than setting ¯ δY(u;z) ¯ δW(u) to 0(exogeneity) or to the perfect proxy estimate can partially identify ¯ β(ν(z) < UX≤ν(z∗)z∗)or ¯ β(ν(z) z). Corollary 6.2. Suppose that,for a.e.u∈Uz,¯ δY(u;z) =d(uz)¯ δW(u) with d(uz) ∈ D(z) ≡[dL(z) dH(z)].(i)Under the conditions of Theorem 6.1(i), if ¯ δW(u)[FU|Z(u|z∗)− FU|Z(u|z)]is either nonpositive for a.e.u∈Uzor nonnegative for a.e.u∈Uzthen ¯ βν(z) < UX≤νz∗z∗∈BD(z)≡RWald YX|Zzz∗−RWald WX|Zzz∗d:d∈D(z) and this identification region is sharp. (ii) Under the conditions of Theorem 6.1(ii), if ¯ δW(u) ∂ ∂z FU|Z(u|z) is either nonpositive for a.e.u∈Uzor nonnegative for a.e.u∈Uzthen ¯ βν(z)z∈BD(z)≡RLIV YX|Z(z) −RLIV WX|Z(z)d :d∈D(z) and this identification region is sharp. B(D(z)) is sharp under the conditions in Corollary 6.2.Weleavestudyingtheconsequences of imposing stronger assumptions, such as (UXUY)⊥(UZ) and UW⊥ (UZ), on the identification of ¯ β(ν(z) < UX≤ν(z∗)z∗)and ¯ β(ν(z)z) to other work.26 26See the comments following the proof of Corollary 6.2 on the sharpness of B(D(z)) if one strengthens the local conditions (10,11) to the global mean independence conditions E[α(u UY)|UZ]=E[α(u UY)], E[β(uUY)|UXUZ]=E[β(uUY)|UX], and E[q(uUY)|UZ]=E[q(uUY)]for all u∈U.
Quantitative Economics 10 (2019) Identification of average effects 1641 Last, Appendix A.2.2 in the Online Supplemental Material discusses how one may use the bounds on MTE to partially identify various average effects. In closing, the analysis in Sections 5and 6contributes to the literature on partial identification of nonparametric average effects when Xor Zare endogenous. In particular, Manski and Pepper (2000) assumed known bounds on the range of Yand that E[r(xUUY)|Z=z]is monotonic in z. They also consider having rbe monotonic in x. Okumura and Usui (2014) combined these assumptions for Z=Xalong with having rbe concave in x. Further, the conditions in Corollaries 5.2 and 6.2 resemble those in Manski and Pepper (2009, Lemma 3.1) who show that if ris monotonic in uand FU|W(u|w∗)≤FU|W(u|w) for all w≤w∗and uthen Wis a monotone IV. Sections 5and 6do not impose any of the above assumptions. Instead they use restrictions on confounding to partially identify various average effects. Last, one can build on the results in Section 6to study the identification of various average effects under restrictions on confounding in systems with discrete (nonbinary) or continuous Xand possibly mismeasured potential instruments (see, e.g., Schennach, White, and Chalak (2012)and Chalak (2017)). 7. Estimation and inference The identification regions in Corollaries 4.3,5.2,and6.2 are of the form B(D)={L(R;d) : d∈D}where the function L(R;d) is known up to a nuisance parameter d,whichispartially identified in a known set D,andRcollects (IV) regression estimands of Yand W on X(using instruments Z). For example, if D=[dL,dH]then the identification region for ¯ β(xx∗|x∗)is B[dLdH]≡RN YXxx∗−RN WXx x∗d:d∈[dLdH] and each element of B([dLdH])is a linear transformation of E(Y|X) and E(W |X)evaluated at x∗and x. We can estimate the (IV) regression estimands R, underlying each element ¯ b(d) =L(R;d) of B(D), using consistent and asymptotically normal parametric, semiparametric, or nonparametric (e.g., kernel) standard estimators ˆ R. We can then estimate the identification region B(D)consistently using ˆ B(D)={L( ˆ R;d) :d∈D}.Further, for each d∈D, we can derive the asymptotic distribution of L( ˆ R;d) as a linear transformation of ˆ Rand construct a 1−α(e.g., 95%) confidence interval C1−α(d) for L(R;d). Using Proposition 2 of Chernozhukov, Rigobon, and Stoker (2010), a 1−αconfidence region CI ¯ β1−αfor a partially identified parameter ¯ β∈B(D)then obtains by forming the union:27 CI ¯ β1−α= d∈D C1−α(d) We illustrate the above discussion in the context of the earnings equation specification used in Section 8.1.Inthiscase,X,Z, and the covariates Sare binary or discrete 27Alternatively, one can consider adapting the procedures in, for example, Imbens and Manski (2004) and Stoye (2009).
1642 Karim Chalak Quantitative Economics 10 (2019) variables, and Yand W(here Uand Ware scalar) are generated by Y=gX(X)¯γ+U¯ δY+U Y¯αYand W=U¯ δW+U W¯αW(12) We collect into the vectors GX≡gX(X),HZ≡hZ(Z),andGS≡gS(S) known flexible (e.g., power and threshold crossing) functions of X,Z,andS,respectively.Here,the average effect ¯ β(xx∗)is encoded by the linear transformation [gX(x∗)−gX(x)]¯γof ¯γ. As discussed in Section A.1.1 in Appendix A in the Online Supplemental Material, when E(HZ|S) and/or E[(G XWY)|S]is affine in GS, applying Theorem A.1 (the conditional on Sversion of Theorem 4.1), with GXand HZreplacing Xand Z,yields ¯γj=RYG|Hj − RWG|Hj ¯ δwhere we put G≡(G XG S)and H≡(H ZG S). The same characterization for ¯γjobtains under the following specification, with Cov[H(U YU W)]=0, Y=G X¯γ+G S¯ ψY+U¯ δY+U Y¯αYand W=G S¯ ψW+U¯ δW+U W¯αW(13) In either representation, each element in the identification region for ¯ β(xx∗), obtained under restrictions on confounding, is a linear transformation of (R YG|HR WG|H). To proceed, we first derive the asymptotic distribution of the plug-in estimator (ˆ R YG|Hˆ R WG|H)for (R YG|HR WG|H). This allows for H=G. For observations {AiBiCi}n i=1corresponding to generic random vector Aand random vectors Band Cof equal dimension, let ˜ Ai≡Ai−1 nn i=1Aiand denote the linear IV regression estimator and sample residuals by ˆ RAB|C≡1 n n i=1 ˜ Ci˜ B i−11 n n i=1 ˜ Ci˜ A iand ˆ AB|Ci ≡˜ A i−˜ B iˆ RAB|C The asymptotic distribution of √n( ˆ R YG|Hˆ R WG|H)obtains using standard arguments.28 For this, we put Q≡diag(E( ˜ H˜ G)E( ˜ H˜ G)). Theorem 7.1. Assume S.1(i) with m=1and that E[˜ H( ˜ G˜ Y ˜ W)]is finite and E( ˜ H˜ G) is nonsingular.Suppose further that: (i) 1 nn i=1˜ Hi˜ G i p →E( ˜ H˜ G),and (ii) n−1/2n i=1(˜ H iYG|Hi˜ H iWG|Hi)d →N(0Ξ),where Ξ≡E˜ H2 YG|H˜ HE˜ HYG|HWG|H˜ H E˜ HWG|HYG|H˜ HE˜ H2 WG|H˜ H is finite and positive definite. Then Λ≡Q−1ΞQ−1is finite and positive definite and √nˆ R YG|Hˆ R WG|H−R YG|HR WG|Hd →N(0Λ) 28See, for example, White (2001) for primitive (sampling and moment) conditions that ensure the law of large numbers and central limit theorem in conditions (i) and (ii) of Theorem 7.1.
Quantitative Economics 10 (2019) Identification of average effects 1649 parental education as an instrument instead. Last, in both IV specifications, conditioning on the subset GS=S1of the covariates yields generally similar bounds.36 In sum, the IV-based bounds are generally wider than, or comparable to, the above regression-based ones and yield especially wider confidence intervals. 8.4 Nonlinear return to education Returning to the regression-based estimates with HZ=GX, we allow for nonlinear yearspecific incremental return to education. Specifically, we let GXcontain binary indicators for having at least tyears of education, where t=218 as in the sample, instead of the total years of education. Thus, γtencodes the incremental return β(tt +1)to year t+1of education. Table 4reports the results. Column 1 reports the results of the regression estimator which is consistent under exogeneity. Column 2 reports the perfect proxy results, yielding the estimate 021 for ¯ δY ¯ δWwith s.e. 003. Under the weaker restriction 0≤¯ δY ¯ δW≤1, we find evidence37 for nonlinearity in the return to education, with the 12th, 16th, and 18th year, corresponding to obtaining a high school, college, and possibly a graduate degree, yielding a high average return. For example, the estimated bounds for the average return to the 12th year are [16%146%]with CI ¯γ11095 [−54%21%] and those for the 16th year are [133%195%]with CI ¯γ15095 [63%262%]. Similarly, the estimated bounds for the return to the 18th year are [139%149%]with CI ¯γ17095 [5%237%]and we cannot reject at comfortable significance levels that the width of this region is zero or, under the maintained assumptions, that a regression consistently estimates this return by 149% with robust s.e. 45%. In contrast, the estimated bounds for the return to the 13th year are [07%78%]with CI ¯γ12095 [−42%124%]. Figure 1 illustrates the nonlinearity in the return to education. In addition to the regression and perfect proxy estimates, it plots the estimated bounds and CI ¯γj095 for the incremental average returns to the 9th up to the 18th year of education under the restriction 0≤¯ δY ¯ δW≤1. Last, using this specification, the estimate of the identification region for the black–white wage gap under 0≤¯ δY ¯ δW≤1is similar to that in Table 1and given by [−178%19%]with CI ¯γ20095 [−216%61%]. 8.5 Discussion and summary This empirical analysis employs a parametric specification in which Uenters additively separably. Further, it assumes that there is one confounder Udenoting “ability,” which 36Setting GS=S1sometimes leads to tighter identification regions albeit with possibly wider confidence intervals (e.g., [−101%24%]with CI ¯γ4095 [−248%174%]for the average black–white wage gap in the first IV specification and [−06%81%]with CI ¯γ1095 [−30%103%]for the average return to education in the second IV specification). 37Although we do not conduct a formal test for linearity, we note that, under the restriction 0≤¯ δY ¯ δW≤1, the 95% CI for the partially identified return to the 16th year of education does not overlap with the 95% CI for the partially identified return to, for example, the 15th year and overlaps with that of the 17th year slightly.
1650 Karim Chalak Quantitative Economics 10 (2019) Table 4. Regression-based estimates of the log wage equation with year-specific education indicators conditional on covariates under restrictions on confounding. jˆ RYGj ˆ RY(GW )j ˆ Gj([01])ˆ Gj([−11]) 10 Educ ≥11 years 0118 0102 [00390118][00390198] (s.e.) and [p-value] (0042)(0042)[0011]– CI095 and CI ¯γj095 [00360201][00200184][−00570201][−00570308] 11 Educ ≥12 years 0146 0119 [00160146][00160276] (s.e.) and [p-value] (0033)(0032)[0000]– CI095 and CI ¯γj095 [00820210][00560182][−00540210][−00540359] 12 Educ ≥13 years 0078 0063 [00070078][00070148] (s.e.) and [p-value] (0024)(0023)[0000]– CI095 and CI ¯γj095 [00320124][00170109][−00420124][−00420205] 13 Educ ≥14 years 0034 0020 [−00350034][−00350103] (s.e.) and [p-value] (0032)(0032)[0000]– CI095 and CI ¯γj095 [−00290096][−00420081][−01010096][−01010177] 14 Educ ≥15 years −0020 −0028 [−0056−0020][−00560015] (s.e.) and [p-value] (0038)(0038)[0048]– CI095 and CI ¯γj095 [−00950054][−01010046][−01340054][−01340102] 15 Educ ≥16 years 0195 0183 [01330195][01330258] (s.e.) and [p-value] (0034)(0034)[0000]– CI095 and CI ¯γj095 [01290262][01170248][00630262][00630335] 16 Educ ≥17 years 0012 −0005 [−00700012][−00700093] (s.e.) and [p-value] (0039)(0039)[0000]– CI095 and CI ¯γj095 [−00650088][−00810071][−01470088][−01470178] 17 Educ ≥18 years 0149 0147 [01390149][01390159] (s.e.) and [p-value] (0045)(0045)[0512]– CI095 and CI ¯γj095 [00610237][00600234][00500237][00500257] 20 Black indicator −0178 −0137 [−01780019][−03740019] (s.e.) and [p-value] (0020)(0021)[0000]– CI095 and CI ¯γj095 [−0216−0139][−0178−0096][−02160061][−04230061] 46 log(KWW )0207 (s.e.) – (0032)–– CI095 [01450269] Note: The results extend the specification in Table 1to include in GXindicators for having at least tyears of education, where t=218 corresponding to the sample, instead of total years of education. For brevity, Table 4does not report the estimated bounds for the average return to education for t<11; these are often relatively imprecise with wide CI ¯γj095.Also, Table 4omits the estimates associated with experience; these are similar to those reported in Table 1. The remaining notes in Table 1apply analogously here. we proxy using log(KWW ),andthat0≤¯ δY ¯ δW≤1or |¯ δY|≤|¯ δW|. Of course, one should interpret the results carefully if these assumptions are suspected to fail. For example, the analysis relaxes the assumption of exogeneity by allowing ability to act as a confounder. But if other confounders are present and strong valid instruments or proxies for these are not available then additional assumptions are needed to (partially) iden-
Quantitative Economics 10 (2019) Identification of average effects 1651 Figure 1. Year-specific incremental return to education conditioning on covariates under restrictions on confounding.
1652 Karim Chalak Quantitative Economics 10 (2019) tify the average effects of X. Similarly, the analysis allows Wto be an imperfect proxy, with UWnondegenerate and conditionally uncorrelated with GX(or HZ)in(13). However, this can in turn fail, for example, if UWdenotes test taking skill (resp., access to counseling) and is conditionally correlated with education (resp., distance to school). Section A.1.2 in Appendix A in the Online Supplemental Material reports complementary results that obtain under some alternative assumptions, such as assuming that the measurement error in the proxy is classical or using the Wequation to substitute for U in the Yequation and then assuming that certain excluded covariates (e.g., parental education) from GSare valid instruments for W. Nevertheless, an advantage of the above empirical analysis is that it does not require several commonly employed assumptions thereby enabling a sensitivity analysis. Specifically, (1) it does not require regressor or instrument exogeneity or restrict the dependence of Uon Xor Z(given S), (2) it does not require a linear return to education, and (3) it permits a test score to be an error-laden proxy for unobserved ability, with possibly nonclassical measurement error. In sum, the estimated bounds for the black–white wage gap are relatively wide, suggesting that, under the imposed assumptions that are weaker than requiring exogeneity or a perfect proxy, this data set is inconclusive about the extent of discrimination in the labor market. In contrast, the average return to education for the black subpopulation may differ slightly from the nonblack subpopulation, if at all. Last, we find evidence suggesting a nonlinearity in the return to education, with graduation years yielding a high average return. 9. Conclusion This paper studies measuring average causal effects in general structural systems with unobserved confounders (omitted variables). We study the identification of coefficients in a linear structure, covariate-conditioned average nonparametric discrete and marginal effects (e.g., average treatment effect on the treated), and local and marginal treatment effects. The first contribution of this paper is to characterize the OVB of common (nonparametric) regression and IV (e.g., Wald and LIV) estimands for these various average effects, thereby generalizing the classic linear regression OVB formula. Using an imperfect proxy for the unobserved confounders, this paper then introduces magnitude and sign restrictions on confounding that are weaker than standard assumptions such as the conditional exogeneity of the treatment or the instrument or requiring a perfect proxy. The paper’s second contribution is to demonstrate how these restrictions on confounding can be used to partially identity average effects and to conduct a sensitivity analysis to deviations from the stronger benchmark assumptions. The paper discusses estimation and inference and applies its framework to study the return to education and the black–white wage gap. Extensions for future work include imposing distributional restrictions on confounding (e.g., a prior distribution on ¯ δ) and using restrictions on confounding to identify the distribution of a causal effect or features of it other than the mean. It is also of interest to apply this paper’s framework to estimate production functions.
Quantitative Economics 10 (2019) Identification of average effects 1653 References Ackerberg, D., K. Caves, and G. Frazer (2015), “Identification properties of recent production function estimators.” Econometrica, 83, 2411–2451. [1620] Altonji, J., T. Conley, T. Elder, and C. Taber (2011), “Methods for using selection on observed variables to address selection on unobserved variables.” Yale University Department of Economics Working Paper. [1632] Altonji, J. and R. Matzkin (2005), “Cross section and panel data estimators for nonseparable models with endogenous regressors.” Econometrica, 73, 1053–1102. [1634] Altonji, J. and C. Pierret (2001), “Employer learning and statistical discrimination.” Quarterly Journal of Economics, 116, 313–350. [1625] Angrist, J., G. Imbens, and D. Rubin (1996), “Identification of causal effects using instrumental variables.” (With discussion) Journal of the American Statistical Association, 91, 444–455. [1638] Angrist, J. and J. Pischke (2009), Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press. [1621] Arcidiacono, P., P. Bayer, and A. Hizmo (2010), “Beyond signaling and human capital: Education and the revelation of ability.” American Economic Journal: Applied Economics, 2, 76–104. [1625] Battistin, E. and A. Chesher (2014), “Treatment effect estimation with covariate measurement error.” Journal of Econometrics, 178, 707–715. [1623] Blackburn, M. and D. Neumark (1992), “Unobserved ability, efficiency wages, and interindustry wage differentials.” Quarterly Journal of Economics, 107, 1421–1436. [1632] Bollinger, C. (2003), “Measurement error in human capital and the black–white wage gap.” Review of Economics and Statistics, 85, 578–585. [1620,1632,1643] Bontemps, C., T. Magnac, and E. Maurin (2012), “Set identified linear models.” Econometrica, 80, 1129–1155. [1632] Bracewell, R. (1986), The Fourier Transform and Its Applications. McGraw-Hill, Inc. [1634] Card, D. (1995), “Using geographic variation in college proximity to estimate the return to schooling.” In Aspects of Labor Market Behaviour: Essays in Honour of John Vanderkamp (L. N. Christofides, E. K. Grant, and R. Swidinsky, eds.). University of Toronto Press, Toronto. [1620,1644,1645,1646,1647,1648] Card, D. (1999), “The causal effect of education on earnings.” In Handbook of Labor Economics, Vol. 3, Part A (O. Ashenfelter and D. Card, eds.). Elsevier. [1622,1643] Card, D. (2001), “Estimating the return to schooling: Progress on some persistent econometric problems.” Econometrica, 69, 1127–1160. [1626]
1654 Karim Chalak Quantitative Economics 10 (2019) Carneiro, P. and J. Heckman (2002), “The evidence on credit constraints in post secondary schooling.” The Economic Journal, 112, 705–734. [1620,1647,1648] Carneiro, P., J. Heckman, and D. Masterov (2005), “Understanding the sources of ethnic and racial wage gaps and their implications for policy.” In Handbook of Employment Discrimination Research: Rights and Realities (R. Nelson and L. Nielsen, eds.), 99–136, Springer, Amsterdam. [1644] Cawley, J., J. Heckman, and E. Vytlacil (2001), “Three observations on wages and measured cognitive ability.” Labour Economics, 8, 419–442. [1625] Chalak, K. (2012), “Identification without exogeneity under equiconfounding in linear recursive structural systems.” In Causality, Prediction, and Specification Analysis: Recent Advances and Future Directions—Essays in Honor of Halbert L. White, Jr. (X. Chen and N. Swanson, eds.), 27–55, Springer. [1629] Chalak, K. (2017), “Instrumental variables methods with heterogeneity and mismeasured instruments.” Econometric Theory, 33, 69–104. [1641] Chalak, K. (2019), “Supplement to ‘Identification of average effects under magnitude and sign restrictions on confounding’.” Quantitative Economics Supplemental Material, 10, https://doi.org/10.3982/QE689.[1622] Chernozhukov, V., R. Rigobon, and T. Stoker (2010), “Set identification and sensitivity analysis with Tobin regressors.” Quantitative Economics, 1, 255–277. [1641] Conley, T., C. Hansen, and P. Rossi (2012), “Plausibly exogenous.” Review of Economics and Statistics, 94, 260–272. [1631,1632] Cunha, F., J. Heckman, and S. Schennach (2010), “Estimating the technology of cognitive and noncognitive skill formation.” Econometrica, 78, 883–931. [1637] Dawid, A. P. (1979), “Conditional independence in statistical theory.” (With discussion) Journal of the Royal Statistical Society, Series B, 41, 1–31. [1626] Fryer, R. (2011), “Racial inequality in the 21st century: The declining significance of discrimination.” In Handbook of Labor Economics,Vol.4B(O.AshenfelterandD.Card, eds.), 855–971, Elsevier. [1644] Griliches, Z. and J. Mairesse (1998), “Production functions: The search for identification.” In Econometrics and Economic Theory in the 20th Century: The Ragnar Frisch Centennial Symposium (S. Strøm, ed.), 169–203, Cambridge University Press. [1620,1635] Halvorsen, R. and R. Palmquist (1980), “The interpretation of dummy variables in semilogarithmic equations.” American Economic Review, 70, 474–475. [1645] Heckman, J., H. Ichimura, and P. Todd (1998), “Matching as an econometric evaluation estimator.” Review of Economic Studies, 65, 261–294. [1633] Heckman, J., S. Urzua, and E. Vytlacil (2006), “Understanding instrumental variables in models with essential heterogeneity.” Review of Economics and Statistics, 88, 389–432. [1633]
Quantitative Economics 10 (2019) Identification of average effects 1655 Heckman, J. and E. Vytlacil (2005), “Structural equations, treatment effects, and econometric policy evaluation.” Econometrica, 73, 669–738. [1637,1638] Hoderlein, S. and E. Mammen (2007), “Identification of marginal effects in nonseparable models without monotonicity.” Econometrica, 75, 1513–1518. [1634] Hu, Y., J. Shiu, and T. Woutersen (2015), “Identification and estimation of single index models with measurement error and endogeneity.” Econometrics Journal, 18, 347–362. [1637] Hu, Y., J. Shiu, and T. Woutersen (2016), “Identification in nonseparable models with measurement error and endogeneity.” Economic Letters, 144, 33–36. [1637] Imbens, G. (2003), “Sensitivity to exogeneity assumptions in program evaluation.” The American Economic Review, 93, 126–132. [1637] Imbens, G. and J. Angrist (1994), “Identification and estimation of local average treatment effects.” Econometrica, 62, 467–476. [1637,1638] Imbens, G. and C. Manski (2004), “Confidence intervals for partially identified parameters.” Econometrica, 72, 1845–1857. [1641] Imbens, G. and W. Newey (2009), “Identification and estimation of triangular simultaneous equations models without additivity.” Econometrica, 77, 1481–1512. [1634] Klein, R. and F. Vella (2009), “A semiparametric model for binary response and continuous outcomes under index heteroscedasticity.” Journal of Applied Econometrics, 24, 735– 762. [1632] Klein, R. and F. Vella (2010), “Estimating a class of triangular simultaneous equations models without exclusion restrictions.” Journal of Econometrics, 154, 154–164. [1632] Klepper, S. and E. Leamer (1984), “Consistent sets of estimates for regressions with errors in all variables.” Econometrica, 52, 163–184. [1632] Lang, K. and M. Manove (2011), “Education and labor market discrimination.” American Economic Review, 101, 1467–1496. [1643] Leamer, E. (1983), “Let’s take the con out of econometrics.” American Economic Review, 73, 31–43. [1624] Levinsohn, J. and A. Petrin (2003), “Estimating production functions using inputs to control for unobservables.” Review of Economic Studies, 70, 317–341. [1620] Lewbel, A. (2012), “Using heteroscedasticity to identify and estimate mismeasured and endogenous regressor models.” Journal of Business and Economic Statistics, 30, 67–80. [1632] Manski, C. (1990), “Nonparametric bounds on treatment effects.” American Economic Review Papers and Proceedings, 80, 319–323. [1633] Manski, C. and J. Pepper (2000), “Monotone instrumental variables: With an application to the returns to schooling.” Econometrica, 68, 997–1010. [1641]
1656 Karim Chalak Quantitative Economics 10 (2019) Manski, C. and J. Pepper (2009), “More on monotone instrumental variables.” Econometrics Journal, 12, S200–S216. [1641] Mincer, J. (1974), Schooling, Experience, and Earning. National Bureau of Economic Research, New York, NY. [1622] Neal, D. and W. Johnson (1996), “The role of premarket factors in black–white wage differences.” Journal of Political Economy, 104, 869–895. [1620,1643] Nevo, A. and A. Rosen (2012), “Identification with imperfect instruments.” Review of Economics and Statistics, 94, 659–671. [1632] Ogburna, E. and T. VanderWeele (2012), “On the nondifferential misclassification of a binary confounder.” Epidemiology, 23, 433–439. [1623] Okumura, T. and E. Usui (2014), “Concave-monotone treatment response and monotone treatment selection: With an application to the returns to schooling.” Quantitative Economics, 5, 175–194. [1641] Olley, G. and A. Pakes (1996), “The dynamics of productivity in the telecommunications equipment industry.” Econometrica, 64, 1263–1297. [1620,1635] Reinhold, S. and T. Woutersen (2009), “Endogeneity and imperfect instruments: Estimating bounds for the effect of early childbearing on high school completion.” Working Paper, University of Arizona Department of Economics. [1632] Schennach, S., H. White, and K. Chalak (2012), “Local indirect least squares and average marginal effects in nonseparable structural systems.” Journal of Econometrics, 166, 282– 302. [1641] Stock, J. and M. Watson (2010), Introduction to Econometrics, third edition. Addison- Wesley. [1621] Stoye, J. (2009), “More on confidence intervals for partially identified parameters.” Econometrica, 77, 1299–1315. [1641] Vytlacil, E. (2002), “Independence, monotonicity, and latent index models: An equivalence result.” Econometrica, 70, 331–341. [1637] Wald, A. (1940), “The fitting of straight lines if both variables are subject to error.” Annals of Mathematical Statistics, 11, 284–300. [1621] White, H. (1980), “A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity.” Econometrica, 48, 817–838. [1643] White, H. (2001), Asymptotic Theory for Econometricians. Academic Press, New York, NY. [1642,1643] White, H. and K. Chalak (2013), “Identification and identification failure for treatment effects using structural systems.” Econometric Reviews, 32, 273–317. [1635] Wickens, M. (1972), “A note on the use of proxy variables.” Econometrica, 40, 759–761. [1623]
Quantitative Economics 10 (2019) Identification of average effects 1657 Wooldridge, J. (2012), Introductory Econometrics: A Modern Approach, fith edition. South-Western College Publishing. [1621,1644] Co-editor Rosa L. Matzkin handled this manuscript. Manuscript received 17 March, 2016; final version accepted 18 March, 2019; available online 29 March, 2019.