Semiparametric estimation of structural functions in nonseparable triangular models
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Chernozhukov, Victor; Fernández-Val, Iván; Newey, Whitney K.; Stouli, Sami; Vella, Francis Article Semiparametric estimation of structural functions in nonseparable triangular models Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Chernozhukov, Victor; Fernández-Val, Iván; Newey, Whitney K.; Stouli, Sami; Vella, Francis (2020) : Semiparametric estimation of structural functions in nonseparable triangular models, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 11, Iss. 2, pp. 503-533, https://doi.org/10.3982/QE1239 This Version is available at: https://hdl.handle.net/10419/217194 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 11 (2020), 503–533 1759-7331/20200503 Semiparametric estimation of structural functions in nonseparable triangular models Victor Chernozhukov Department of Economics, MIT Iván Fernández-Val Department of Economics, Boston University Whitney Newey Department of Economics, MIT Sami Stouli Department of Economics, University of Bristol Francis Vella Department of Economics, Georgetown University Triangular systems with nonadditively separable unobserved heterogeneity provide a theoretically appealing framework for the modeling of complex structural relationships. However, they are not commonly used in practice due to the need for exogenous variables with large support for identification, the curse of dimensionality in estimation, and the lack of inferential tools. This paper introduces two classes of semiparametric nonseparable triangular models that address these limitations. They are based on distribution and quantile regression modeling of the reduced form conditional distributions of the endogenous variables. We show that average, distribution, and quantile structural functions are identified in these systems through a control function approach that does not require a large support condition. We propose a computationally attractive three-stage procedure to estimate the structural functions where the first two stages consist of quantile or distribution regressions. We provide asymptotic theory and uniform inference methods for each stage. In particular, we derive functional central limit theorems and Victor Chernozhukov: [email protected] Iván Fernández-Val: [email protected] Whitney Newey: [email protected] Sami Stouli: [email protected] Francis Vella: [email protected] We are grateful to seminar participants at CeMMAP and the 2018 European Summer Meeting of the Econometric Society (Cologne) for useful comments, and to four anonymous referees who have also helped us greatly improve the paper. We thank Richard Blundell, Xiaohong Chen, and Dennis Kristensen for kindly sharing the data used in the empirical application. Research support from the National Science Foundation, the U.K. Economic and Social Research Council, and the Royal Economic Society is gratefully acknowledged. ©2020 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1239
504 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) bootstrap functional central limit theorems for the distribution regression estimators of the structural functions. These results establish the validity of the bootstrap for three-stage estimators of structural functions, and lead to simple inference algorithms. We illustrate the implementation and applicability of all our methods with numerical simulations and an empirical application to demand analysis. Keywords. Structural functions, nonseparable models, control function, quantile and distribution regression, semiparametric estimation, uniform inference. JEL classification. C14, C31, C35, C51. 1. Introduction Models with nonadditively separable disturbances provide an important vehicle for incorporating heterogenous effects. However, accounting for endogenous treatments in such a setting can be challenging. One methodology which has been successfully employed in a wide range of models with endogeneity is the use of control functions (see, for surveys, Imbens and Wooldridge (2009), Wooldridge (2015), and Blundell, Newey, and Vella (2019)). The underlying logic of this approach is to account for the endogeneity by including an appropriate control function in the conditioning variables. This paper proposes some relatively simple control function procedures to estimate objects of interest in a triangular model with nonseparable disturbances. Our approach to circumventing the inherent difficulties in nonparametric estimation associated with the curse of dimensionality is to build our models upon a semiparametric specification. This also alleviates the full support requirement on the control function conditional on the treatment variable needed for nonparametric identification. Our goal is thus to provide models and methods that are essentially parametric but still allow for nonseparable disturbances in order to address strong data requirements that come with nonparametric formulations. These models can be interpreted as “baseline” models on which series approximations can be built by adding additional terms. We consider two kinds of baseline models: quantile regression and distribution regression. These models allow the use of convenient and widely available methods to estimate objects of interest including average, distribution, and quantile structural/treatment effects. A main feature of the baseline models is that interaction terms included would not usually be present as leading terms in estimation. These included terms are products of a transformation of the control function with the endogenous treatment. Their presence is meant to allow for heterogeneity in the coefficient of the endogenous variable. Such heterogenous coefficient linear models are of interest in many settings, including demand analysis and estimation of returns to education, and provide a natural starting point for more general models that allow for nonlinear effects of the endogenous treatments. We use these baseline models to construct estimators of the average, distribution, and quantile structural functions based on parametric quantile and distribution regressions. These objects fully characterize the structural relationship between the endogenous treatment and the outcome of interest, and describe the average, distribution, and quantiles of the outcome across treatment values, had the treatment been exogenous. We also show how these baseline models can be expanded to include higher order
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 505 terms, leading to more flexible structural function specifications. The estimation procedure consists of three stages. First, we estimate the control function via quantile regression (QR) or distribution regression (DR) of the endogenous treatment on the exogenous covariates and variables that satisfy an exclusion restriction. Second, we estimate the reduced form distribution of the outcome conditional on the treatment, covariates, and estimated control function using DR or QR. Third, we construct estimators of the structural functions applying suitable functionals to the reduced form estimator from the second stage. We derive asymptotic theory for the estimators based on DR in all the stages using a trimming device that avoids tail estimation in the construction of the control function. We also establish the validity of the bootstrap for our inference on structural functions, which enables the formulation of convenient inference algorithms which we describe in detail. The modeling framework we propose thus allows us to address three key difficulties that have restricted the use of such models in empirical work—the curse of dimensionality, the full support condition for identification, and the lack of easily implementable inference methods—while simultaneously retaining important features of the original nonparametric formulation. We give an empirical application based on the estimation of Engel curves which illustrates how our approach leads to flexible estimates of all structural functions and their confidence regions. Our results for the average structural function in the linear random coefficients model are similar to Garen (1984). Florens, Heckman, Meghir, and Vytlacil (2008)gave identification results for a random coefficients model where the structural function is a polynomial in the endogenous treatment. Blundell and Powell (2003,2004) introduced the average structural function, and Imbens and Newey (2009) gave general models and results for a variety of objects of interest and control functions, including quantile structural functions, under a large support condition on a variable that satisfies an exclusion restriction. We formulate semiparametric specifications of their models that do not require this excluded variable to have continuous support. Our work is also related to the literature on identification and estimation in nonseparable triangular systems with as many unobservables as equations. Chesher (2003), Ma and Koenker (2006), Jun (2009), and Chernozhukov, Fernández-Val, and Kowalski (2015) considered identification and estimation of the structural function at quantiles of the unobservable in the outcome equation conditional on values of the control function; Stouli (2015) gave conditions for identification and estimation of the structural function at both marginal and conditional quantiles of the unobservable in the outcome equation given the control function, under a normalization on the distribution of the unobservable conditional on a specified value of the control function. These approaches do not apply to triangular systems with more unobservables than equations. In contrast, we consider semiparametric formulations which, if correctly specified, provide valid models for the determination of quantile structural effects irrespective of the dimensionality of unobserved heterogeneity in the outcome equation. Our work also complements the analysis and methods developed for the single-equation instrumental variable quantile regression model of Chernozhukov and Hansen (2005,2006), which rely on monotonicity of the structural function in a scalar disturbance and, therefore, do not apply to the class of models we consider. In contrast, triangular systems rely on monotonicity of the first stage reduced
506 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) form function in a scalar disturbance, but include more than one source of unobserved heterogeneity. In addition, control function methods can be used for any type of outcome variable, continuous, discrete, or mixed continuous-discrete, but require the endogenous treatment to be continuous. In contrast, the instrumental variable quantile regression approach requires a continuous outcome but applies to any type of treatment. The two approaches therefore provide complementary modeling frameworks and estimation tools for the empirical analysis of nonseparable models with endogenous treatment.1 This paper makes four main contributions to the existing literature. First, we establish identification of structural functions in both classes of baseline models, providing conditions that do not impose large support requirements on the excluded variable. D’Haultfoeuille and Février (2015) and Torgovitsky (2015) gave identification results in the presence of an instrument with small support, but require monotonicity of the structural function in a scalar disturbance. Instead we restrict the functional form of our models. Second, we derive a functional central limit theorem and a bootstrap functional central limit theorem for the two-stage DR estimators in the second stage. These results are uniform over compact regions of values of the outcome. To the best of our knowledge, this result is new. For example, Chernozhukov, Fernández-Val, and Kowalski (2015) derived similar results for two-stage quantile regression estimators but their results are pointwise over quantile indexes and are not applicable to the problem considered here. Our analysis builds on Chernozhukov, Fernández-Val, and Galichon (2010) and Chernozhukov, Fernández-Val, and Melly (2013), which established the properties of the DR estimators that we use in the first stage. The theory of the two-stage estimator, however, does not follow from these results using standard techniques due to the dimensionality and entropy properties of the first stage DR estimators. We follow the proof strategy proposed by Chernozhukov, Fernández-Val, and Kowalski (2015) to deal with these issues. Third, we derive functional central limit theorems and bootstrap functional central limit theorems for plug-in estimators of functionals of the distribution of the outcome conditional on the treatment, covariates and control function via functional delta method. These functionals include all the structural functions of interest. We also use a linear functional for the average structural function which had not been previously considered. Fourth, we show that this linear operator that relates the average of a random variable with its distribution is Hadamard differentiable. Our modeling framework and theoretical results are also of interest for the study of nonseparable triangular models in various alternative settings,2and will allow establishing the validity of bootstrap inference for the corresponding estimators. The rest of the paper is organized as follows. Section 2describes the baseline models and objects of interest. Section 3presents the estimation and inference methods. Section 4gives asymptotic theory. Section 5reports the results of an extensive empirical application to Engel curves, and provide implementation algorithms for all our methods. 1See Section 4.2 in Chernozhukov and Hansen (2013) for a related discussion. 2See Fernández-Val, van Vuuren, and Vella (2018) for an application to the analysis of nonseparable sample selection models with censored selection rules.
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 507 The identification proof is given in the Appendix. The Appendix in the Online Supplemental Material (Chernozhukov, Fernández-Val, Newey, Stouli, and Vella (2020)) contains supplemental material, including the proofs for asymptotic theory and results of numerical simulations calibrated to the application. 2. Modeling framework We begin with a brief review of the triangular nonseparable model and some inherent objects of interest. Let Ydenote an outcome variable of interest that can be continuous, discrete, or mixed continuous-discrete, Xa continuous endogenous treatment, Zavector of exogenous variables, εa structural disturbance vector of unknown dimension, and ηa scalar reduced form disturbance.3A general nonseparable triangular model takes the form Y=g(Xε) X=h(Zη) (ε η) indep of Z where η→ h(z η) is a one-to-one function for each z. This model implies that εand Xare independent conditional on ηand that ηis a one-to-one function of V=FX(X | Z), the cumulative distribution function (CDF) of Xconditional on Zevaluated at the observed variables. Thus, Vis a control function. Objects of interest in this model include the ASF, μ(x), quantile structural function (QSF), Q(τ x), and distribution structural function (DSF), G(yx),where μ(x) =g(xε)Fε(dε) and Q(τ x) =τth quantile of g(xε) G(yx) =Prg(xε) ≤y Here, μ( ˜ x) −μ(¯ x) is like an average treatment effect, Q(τ ˜ x) −Q(τ ¯ x) is like a quantile treatment effect, and G(y ˜ x) −G(y ¯ x) is like a distribution treatment effect from the treatment effects literature. If the support of Vconditional on X=xis the same as the marginal support of V, then these objects are nonparametrically identified4by μ(x) =E[Y|X=xV =v]FV(dv) and Q(τ x) =G←(τ x) G(yx) =FY(y |X=xV =v)FV(dv) 3In our empirical application, we use household level data to study the structural relationship between the share of expenditure on either food or leisure, Y, and the log of total expenditure, X, with gross earnings of the head of household as the excluded variable Z. Additional examples and a general economic motivation of nonseparable triangular models are given in Chesher (2003) and Imbens and Newey (2009), for instance. 4Full support of Vconditional on X=xneeded for nonparametric identification requires the excluded variable Zto have large support conditional on X=x; see Imbens and Newey (2009).
508 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) where G←(τ x) denotes the left inverse of y→ G(yx) for each x,thatis,G←(τx) := inf{y∈R:G(yx) ≥τ}. It is straightforward to extend this approach to allow for covariates in the model by further conditioning on or integrating over them. Suppose that Z1⊂Zis included in the structural equation, which is now g(XZ1ε). Under the assumption that εand V are jointly independent of Z,thenεwill be independent of Xand Z1conditional on V. Conditional on covariates and unconditional average structural functions are identified by μ(xz1)=E[Y|X=xZ1=z1V =v]FV(dv) and μ(x) =E[Y|X=xZ1=z1V =v]FZ1(dz1)FV(dv) Similarly, conditional on covariates and unconditional quantile and distribution structural functions are identified by Q(τ xz1)=G←(τxz1) G(y x z1)=FY(y |X=x Z1=z1V =v)FV(dv) and Q(τ x) =G←(τ x) G(yx) =FY(y |X=xZ1=z1V =v)FZ1(dz1)FV(dv) respectively. The structural functions can all be expressed as functionals of the control function V=FX(X |Z), the conditional mean function, E[Y|XZ1V], and the conditional CDF FY(Y |XZ1V). These reduced form functions thus constitute natural modeling targets in the context of triangular models. Without functional form restrictions, the curse of dimensionality makes them difficult to estimate, and the full support condition makes it difficult to achieve point identification of the structural functions. These difficulties motivate our specification of baseline parametric models in what follows. These baseline models provide good starting points for nonparametric estimation and may be of interest in their own right. 2.1 Quantile regression baseline We start with a simplified specification with one endogenous treatment X, one excluded variable Z, and a continuous outcome Y. We show below how additional excluded variables and covariates can be included. The baseline first stage is the QR model X=QX(V |Z) =π1(V ) +π2(V )Z V |Z∼U(01) (2.1)
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 509 Note that v→π1(v) and v→π2(v) are infinite dimensional parameters (functions), and the mapping v→ π1(v) +π2(v)z is strictly increasing for each z.5We can recover the control function Vfrom V=FX(X |Z) =Q−1 X(X |Z) or equivalently from V=FX(X |Z) =1 0 1π1(v) +π2(v)Z ≤Xdv This generalized inverse representation of the CDF is convenient for estimation because it does not require the conditional quantile function to be strictly increasing to be welldefined.6 The baseline second stage has a reduced form: Y=QY(U |XV ) U |XV ∼U(01) (2.2) QY(U |XV ) =β1(U) +β2(U)X +β3(U)−1(V ) +β4(U)X−1(V ) (2.3) where −1is the standard normal inverse CDF. This transformation is included to expand the support of Vand to encompass the normal system of equations as a special case (cf. Section E.2 of the Online Supplementary Material for a detailed derivation), but it can be replaced by any other strictly monotonic quantile function. We can recover the conditional CDF FYfrom FY(Y |XV ) =Q−1 Y(Y |XV ) or equivalently from FY(Y |XV ) =1 0 1β1(u) +β2(u)X +β3(u)−1(V ) +β4(u)X−1(V ) ≤Ydu An example of a structural model with reduced form (2.2)–(2.3) is the random coefficient model Y=g(Xε) =ε1+ε2X (2.4) with the restrictions εj=Qεj(U |XV ) =θj(U) +γj(U)−1(V ) U |XV ∼U(01)j ∈{12}(2.5) These restrictions include the control function assumption εj⊥⊥X|Vand a joint functional form restriction, where the unobservable Uisthesameforε1and ε2. Substituting in the second stage equation, Y=θ1(U) +θ2(U)X +γ1(U)−1(V ) +γ2(U)−1(V )X U |XV ∼U(01) 5The specification (2.1) restricts the support of Z,as∂π1(v)/∂v +z∂π2(v)/∂v > 0must hold for all (zv) for QXto be well specified. In particular, if Zhas support the real line, then (2.1) restricts ∂π2(v)/∂v =0 for all v.WhenZhas positive support as in our empirical application, ∂πj(v)/∂v > 0,j=12,issufficient for ∂π1(v)/∂v +z∂π2(v)/∂v > 0. All results in the paper allow for more general specifications such as QX(v | z) =r(z)π(v),wherer(z) is a vector of transformations of zincluding a 1as the first component, which alleviate the support restrictions on Zimplied by (2.1). 6There are other valid approaches for the specification and estimation of the control function. For instance, dual regression (Spady and Stouli (2018)) provides a parametric alternative for the modeling of FX(X |Z), and nonparametric estimation based on locally linear and series estimators was considered by Imbens and Newey (2009). Here, we focus on semiparametric approaches, QR and DR, because of their flexibility, and their well-established computational and theoretical properties.
510 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) which has the form of (2.2)–(2.3). A special case of the QR baseline is a heteroskedastic normal system of equations. We use this specification in the numerical simulations given in the Online Supplementary Material. The specification (2.2)–(2.3) is a baseline, or starting point, for a more general series approximation to the quantiles of Yconditional on Xand Vbased on including additional functions of Xand −1(V ). The baseline is unusual as it includes the interaction term −1(V )X; it is more usual to take the starting point to be (1−1(V ) X),which is linear in the regressors Xand −1(V ). The inclusion of the interaction term is motivated by allowing the coefficient of Xto vary with individuals, so that −1(V ) then interacts Xin the conditional distribution of ε2given the control functions. Augmenting the baseline with splines or power transformations of Xand −1(V ) and their interactions gives rise to more flexible semiparametric specifications. In this more general case, the inclusion of interaction terms is motivated by increasing the flexibility with which the coefficients of Xand its transformations can vary across individuals—as would arise, for instance, when including power transformations of Xand −1(V ) in the random coefficient models (2.4)and(2.5), respectively. All the parameters of both the baseline and augmented specifications can be identified when the support of Zis discrete, and estimated using the QR estimator (Koenker and Bassett (1978)). Thus the baseline (2.2)– (2.3) incorporates a flexible heterogeneous coefficients structure that can easily be augmented for the purpose of modeling more complex relationships, while only relying on weak identification conditions and preserving ease of estimation.7 The ASF of the baseline specification is μ(x) =1 0 E[Y|X=xV =v]dv =β1+β2x where the second equality follows by 1 0−1(v) dv =0and E[Y|XV ]=1 0 QY(u |XV )du =β1+β2X+β3−1(V ) +β4X−1(V ) (2.6) with βj:= 1 0βj(u)du,j∈{14}. The QSF does not appear to have a closed-form expression. It is the solution to Q(τ x) =G←(τ x) G(yx) =1 01 0 1β1(u) +β2(u)x +β3(u)−1(v) +β4(u)−1(v)x ≤ydudv Remark 1. Another way to arrive at that conditional mean specification (2.6)istostart with the random coefficients model Y=ε1+ε2Xand assume that the conditional mean of ε1and ε2given Vare linear in (1−1(V )).Then E[Y|XV ]=E[ε1|V]+E[ε2|V]X=¯ θ1+¯γ1−1(V ) +¯ θ2X+¯γ2−1(V )X 7Our formal definition of the model in Section 2.3 explicitly allows for augmented semiparametric baseline specifications, and Section 5and the Online Supplementary Material illustrate our methods when incorporating spline and power transformations of X,−1(V ) and their interactions.
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 517 yS=sup[y∈Y]} with mesh width δsuch that δ√n→0. For example, (3.12)and(3.14) are approximated by Qe S(τ x) =δ S s=11(ys≥0)−1 Ge(ysx)≥τ(3.15) and μe S(x) =δ S s=11(ys≥0)− Ge(ysx)(3.16) 3.4 Weighted bootstrap inference on structural functions We consider inference uniform over regions of values of (yxτ). We denote the region of interest as IGfor the DSF, IQfor the QSF, and Iμfor the ASF. Examples include: (1) The DSF, y→ Ge(yx),forfixedxand over y∈ Y⊂Y, by setting IG= Y×{x}. (2) The QSF, τ→ Qe(τ x) for fixed xand over τ∈ T⊂(01), by setting IQ= T×{x}, (3) The ASF, μe(x),overx∈ X⊂X, by setting Iμ= X. When the region of interest is not a finite set, we approximate it by a finite grid. All the details of the procedure we implement are summarized in Section 5.1. The weighted bootstrap versions of the DSF, QSF, and ASF estimators are obtained by rerunning the estimation procedure introduced in Section 3.3 with sampling weights drawn from a distribution that satisfies Assumption 3in Section 4;seeAlgorithm2in Section 5.1 for details. They can then be used to perform uniform inference over the region of interest. For instance, a (1−α)-confidence band for the DSF over the region IGcan be constructed as G(yx) ± kG(1−α)σG(y x) (y x) ∈IG(3.17) where σG(y x) is an estimator of σG(yx), the asymptotic standard deviation of G(yx), such as the rescaled weighted bootstrap interquartile range11 σG(y x) =IQR Ge(yx)/1349(3.18) and kG(1−α) denote a consistent estimator of the (1−α)-quantile of the maximal tstatistic tG(yx) IG=sup (yx)∈IG G(yx) −G(yx) σG(yx) 11An alternative is to use the bootstrap standard deviation, but its validity requires convergence of bootstrap moments in addition to convergence of the bootstrap distribution; cf. Remark 3.2 in Chernozhukov, Fernández-Val, and Melly (2013).
518 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) such as the (1−α)-quantile of the bootstrap draw of the maximal t-statistic te G(yx) IG=sup (yx)∈IG Ge(yx) − G(yx) σG(y x) (3.19) Confidence bands for the ASF can be constructed by a similar procedure, using the bootstrap draws of the ASF estimator. For the QSF, we can either use the same procedure based on the bootstrap draws of the QSF, or invert the confidence bands for the DSF following the generic method of Chernozhukov, Fernández-Val, Melly, and Wuthrich (2019). The first possibility works only when Yis continuous, whereas the second method is more generally applicable. We provide algorithms for the construction of the bands in Section 5.1. 4. Asymptotic theory We derive asymptotic theory for the estimators of the ASF, DSF, and QSF where both the first and second stages are based on DR. The theory for the estimators based on QR can be derived using similar arguments. In what follows, we shall use the following notation. We let the random vector A=(YXZWV) live on some probability space (Ω0F0P). Thus, the probability measure Pdetermines the law of Aor any of its elements. We also let A1An, i.i.d. copies of A, live on the complete probability space (Ω FP), which contains the infinite product of (Ω0F0P). Moreover, this probability space can be suitably enriched to carry also the random weights that appear in the weighted bootstrap. The distinction between the two laws Pand Pis helpful to simplify the notation in the proofs and in the analysis. Unless explicitly mentioned, all functions appearing in the statements are assumed to be measurable. We now state formally the assumptions. The first assumption is about sampling and the bootstrap weights. Assumption 3 (Sampling and bootstrap weights). (a) Sampling:the data {YiXiZi}n i=1 are a sample of size nof independent and identically distributed observations from the random vector (YXZ).(b)Bootstrap weights:(e1en)are i.i.d.draws from a random variable e≥0,with EP[e]=1,Var P[e]=1,and EP|e|2+δ<∞for some δ>0;live on the probability space (Ω FP);and are independent of the data {YiXiZi}n i=1for all n. The second assumption is about the first stage where we estimate the control function (xz) →ϑ0(x z) defined as ϑ0(xz) :=FX(x |z) with trimmed support V={ϑ0(xz) :(x z) ∈XZ}. We assume a logistic DR model for the conditional distribution of Xin the trimmed support X.
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 519 Assumption 4 (First stage). (a) Trimming:we consider a trimming rule defined by the tail indicator T=1(X ∈X) where X=[xx]for some −∞ <x<x<∞,such that P(T =1)>0.(b)Model:the distribution of Xconditional on Zfollows Assumption 1(b) with =Λ,where Λis the logit link function;the coefficients x→ π0(x) are three times continuously differentiable with uniformly bounded derivatives;Ris compact;and the minimum eigenvalue of EP[Λ(Rπ0(x))[1−Λ(Rπ0(x))]RR]is bounded away from zero uniformly over x∈X. For x∈X,let πe(x) ∈arg min π∈Rdim(R) −1 n n i=1 ei1(Xi≤x)logΛR iπ+1(Xi>x)log1−ΛR iπ and set ϑ0(xr) =Λrπ0(x); ϑe(xr) =Λrπe(x) if (xr) ∈XR,andϑ0(x r) = ϑe(xr) =0otherwise. Theorem 4 of Chernozhukov, Fernández-Val, and Kowalski (2015) established the asymptotic properties of the DR estimator of the control function. We repeat the result here as a lemma for completeness and to introduce notation that will be used in the results below. Let T(x):=1(x ∈X),fT∞:=supa∈A|T(x)f(a)|for any function f:A→ R,λ:=Λ(1−Λ), the density of the logistic distribution. Lemma 2(Firststage). Suppose that Assumptions 3and 4hold.Then (1) √n ϑe(xr) −ϑ0(x r)=1 √n n i=1 ei(Aixr)+oP(1)e(x r) in ∞(XR) (A xr) :=λrπ0(x)1{X≤x}−ΛRπ0(x) ×rEPΛRπ0(x)1−ΛRπ0(x)RR−1R EP(Axr)=0EPT(AXR)2<∞ where (x r) → e(xr) is a Gaussian process with uniformly continuous sample paths and covariance function given by EP[(A xr)(A ˜ x ˜ r)].(2)There exists ϑe:XR → [01]that obeys the same first-order representation uniformly over XR,is close to ϑein the sense that ϑe− ϑeT∞=oP(1/√n) and,with probability approaching one,belongs to a bounded function class such that the covering entropy satisfies12 logN Υ ·T∞−1/20<<1 12See Section C of the Online Supplementary Material for a definition of the covering entropy.
520 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) The next assumptions are about the second stage. We assume a logistic DR model for the conditional distribution of Ygiven (X Z1V), impose compactness and smoothness conditions, and provide sufficient conditions for identification of the parameters. Compactness is imposed over the trimmed supports and can be relaxed at the cost of more complicated and cumbersome proofs. The smoothness conditions are fairly tight. The assumptions on Ycover continuous, discrete, and mixed outcomes in the second stage. We denote partial derivatives as ∂xf(xy):=∂f(xy)/∂x. Assumption 5 (Second stage). (a) Model:the distribution of Yconditional on (XZ1V) follows Assumption 1(b) with =Λ.(b)Compactness and smoothness:the set XZW is compact;the set Yis either a compact interval in Ror a finite subset of R;Xhas a continuous conditional density function x→ fX(x |z) that is bounded above by a constant uniformly in z∈Z;if Yis an interval,then Yhas a conditional density function y→ fY(y |xz) that is uniformly continuous in y∈Yuniformly in (x z) ∈XZ,and bounded above by a constant uniformly in (xz) ∈XZ;the derivative vector ∂vw(xz1v) exists and its components are uniformly continuous in v∈Vuniformly in (xz1)∈XZ 1, and are bounded in absolute value by a constant,uniformly in (xwv)∈XZ 1V;and for all y∈Y,β0(y) ∈B,where Bis a compact subset of Rdim(W ).(c)Nondegeneracy:the matrix C(yv) :=CovP[fy(A) +gy(A)fv(A) +gv(A)]is finite and is of full rank for all yv ∈Y, where fy(A) :=ΛWβ0(y)−1(Y ≤y)WT and,for ˙ W=∂vw(XZ1v)|v=V, gy(A) :=EPΛWβ0(y)−1(Y ≤y)˙ W+λWβ0(y)˙ Wβ0(y)W T(aXR)|a=A For y∈Y,let β(y) =arg min β∈Rdim(W ) 1 n n i=1 TiρyYiβ Wi Wi=w(XiZ1i Vi) Vi= ϑ(XiRi) where ρy(YB) :=−1(Y ≤y)log Λ(B) +1(Y > y) log1−Λ(B) and ϑis the estimator of the control function in the unweighted sample; and βe(y) =arg min β∈Rdim(W ) 1 n n i=1 eiTiρyYiβ We i We i=wXiZ1i Ve i Ve i= ϑe(XiRi) where ϑeis the estimator of the control function in the weighted sample. The following lemma establishes a functional central limit theorem and a functional central limit theorem for the bootstrap for the estimator of the DR coefficients in the second stage. Let dw:= dim(W ),and∞(Y)be the set of all uniformly bounded real functions on Y, and define the matrix J(y) := EP[λ(W β0(y))W W T]for y∈Y.Weuse Pto denote bootstrap consistency, that is, weak convergence conditional on the data
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 521 in probability, which is formally defined in Section C.1 of the Online Supplementary Material. Lemma 3 (FCLT and bootstrap FCLT for β(y)). Under Assumptions 1–5,in ∞(Y)dw, √n β(y) −β0(y)J(y)−1G(y) and √n βe(y) − β(y)PJ(y)−1G(y) where y→G(y) is a dw-dimensional zero-mean Gaussian process with uniformly continuous sample paths and covariance function EPG(y)G(v)=C(yv) yv ∈Y We consider now the estimators of the main quantities of interest—the structural functions. Let Wx:= w(xZ1V), Wx:= w(xZ1 V),and We x:= w(xZ1 Ve).The DR estimator and bootstrap draw of the DSF in the trimmed support, GT(y x) = EP{Λ[β0(y)Wx]|T=1},are G(yx) =n i=1Λ[ β(y) Wxi]Ti/nT,and Ge(yx) =n i=1ei× Λ[ βe(y) We xi]Ti/ne T.LetpT:= P(T =1). The next result gives large sample theory for these estimators. Theorem 2 (FCLT and bootstrap FCLT for DSF). Under Assumptions 1–5,in ∞(YX ), √npT G(yx) −GT(yx)Z(yx) and √npT Ge(yx) − G(yx)PZ(yx) where (y x) →Z(yx) is a zero-mean Gaussian process with covariance function CovPΛW xβ0(y)+hyx(A) ΛW uβ0(v)+hvu(A) |T=1 with hyx(A) =EPλW xβ0(y)WxT−1fy(A) +gy(A) +EPλW xβ0(y)˙ W xβ0(y)T(aX R)|a=A When Yis continuous and y→GT(y x) is strictly increasing, we can also characterize the asymptotic distribution of Q(τ x), the estimator of the QSF in the trimmed support. Let gT(y x) be the density of y→GT(yx),T:={τ∈(01):Q(τ x) ∈YgT(Q(τx) x) > x ∈X}for fixed >0,andQT(τ x) the QSF in the trimmed support TX defined as QT(τ x) =Y+ 1GT(yx) ≤τdy −Y− 1GT(yx) ≥τdy The estimator and its bootstrap draw given in (3.11)–(3.12) follow the functional central limit theorem.
522 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) Theorem 3 (FCLT and Bootstrap FCLT for QSF). Assume that y→ GT(yx) is strictly increasing in Yand (y x) → GT(y x) is continuously differentiable in YX.Under Assumptions 1–5,in ∞(TX), √npT Q(τx) −QT(τ x)−ZQ(τx)x gTQ(τ x)xand √npT Qe(τ x) − Q(τx)P−ZQ(τx)x gTQ(τ x)x where (y x) →Z(yx) is the same Gaussian process as in Theorem 2. Finally, we consider the ASF in the trimmed support μT(x) =Y+1−GT(yx)ν(dy) −Y− GT(yx)ν(dy) The estimator and its bootstrap draw given in (3.13)–(3.14) follow the functional central limit theorem. Theorem 4 (FCLT and bootstrap FCLT for ASF). Under Assumptions 1–5,in ∞(X), √npTμ(x) −μT(x)−Y Z(yx)ν(dy) and √npTμe(x) −μ(x)P−Y Z(yx)ν(dy) where (y x) →Z(yx) is the same Gaussian process as in Theorem 2. 5. Implementation and application to estimation of Engel curves In this section, we provide algorithms for the implementation of our methods, and apply them to the estimation of a semiparametric nonseparable triangular model for Engel curves. We focus on the structural relationship between household’s total expenditure and household’s demand for two goods: food and leisure. We take the outcome Y to be the expenditure share on either food or leisure, and Xthe logarithm of total expenditure. Endogeneity in the estimation of Engel curves arises because the decision to consume a particular good may occur simultaneously with the allocation of income between consumption and savings. Following Blundell, Chen, and Kristensen (2007), we use the logarithm of gross earnings of the head of household as the variable that satisfies an exclusion restriction. We also include an additional binary covariate Z1accounting for the presence of children in the household. There is an extensive literature on Engel curve estimation (e.g., see Lewbel (2006) for a review), and the use of nonseparable triangular models for the identification and estimation of Engel curves has been considered in the recent literature. Blundell, Chen, and Kristensen (2007) estimate semi-nonparametrically Engel curves for several categories of expenditure, Imbens and Newey (2009) estimated the QSF nonparametrically
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 523 for food and leisure, and Chernozhukov, Fernández-Val, and Kowalski (2015) estimated Engel curves for alcohol accounting for censoring. For comparison purposes, we use the same dataset as these papers, the 1995 U.K. Family Expenditure Survey. We restrict the sample to 1,655 married or cohabiting couples with two or fewer children, in which the head of the household is employed and between the ages of 20 and 55 years. For this sample, we estimate the DSF, QSF, and ASF for both goods. Unlike Imbens and Newey (2009), we also account for the presence of children in the household and we impose semiparametric restrictions through our baseline models. In contrast to Chernozhukov, Fernández-Val, and Kowalski (2015), we do not impose separability between the control function and other regressors, and we estimate the structural functions. 5.1 Implementation of estimation and inference methods In order to guide implementation of our methods, we provide step-by-step implementation algorithms for the three-stage estimation procedure, weighted bootstrap, and the construction of uniform bands for the structural functions. All structural functions are estimated by both QR and DR methods, following exactly the description of the implementation presented in Section 3with the specifications r(Z) =(1Z),r1(Z1)=(1Z1), p(X) =(1X),andq(V ) =(1−1(V )). We implement our methods in the software R (R Development Core Team (2019)), using the open source quantreg Rpackage (Koenker (2018)) for QR, and the glm function for DR. In the empirical application, the regions of interest for the structural functions are X=[ QX(01) QX(09)]and Y=[ QY(01) QY(09)],where QX(u) and QY(u) are the sample u-quantiles of Xand Y, respectively. We approximate Xby a grid XKwith K=35,and Yby a grid Y15. We estimate the structural functions and perform uniform inference on the structural functions over the following regions: (1) For the QSF, Q(τ x),wetake T={02505075}, and then set: IQ= T X5. (2) For the DSF, G(yx),weset:IG= Y15 X3. (3) For the ASF, μ(x),weset:Iμ= X5. 5.1.1 Estimation Algorithm 1is implemented for estimation of structural functions. In Algorithm 1, the choice of the link function for DR, of for QR, and the size of the grids Mcan differ across stages and methods. In the empirical application, we implement the DR estimator using the logit link function, and we set =001 for QR, and M=599 throughout. For the third stage, we approximate the integrals (3.12)and(3.14) using S=599 points. For DR, since the estimated DSF may be non-monotonic in y,we apply rearrangement to y→ G(yx) at each value of xin the region of interest, using the Rearrangement R package (Graybill, Chen, Chernozhukov, Fernández-Val, and Galichon (2016)). Overall, for the empirical application we have found that the estimates are not very sensitive to Mand the choice of link function, and are also robust to varying values of and S. None of the methods uses trimming, that is, we set T=1a.s. Remark 2. For DR, the estimation of π(x) at each x=Xican be computationally expensive. Substantial gains in computational speed is achieved by first estimating π(x) in a grid XM, and then obtaining π(x) at each x=Xiby interpolation.
524 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) Algorithm 1 Three-stage estimation procedure. For i=1n,setei=1. First stage. [Control function estimation] (1) (QR) For in (005)(e.g., =001) and a fine mesh of Mvalues {=v1<···< vM=1−},estimate{πe(vm)}M m=1by solving (3.2). Then set Ve i= Fe X(Xi|Zi),i=1n, as in (3.1). (2) (DR) Estimate {π(Xi)}n i=1by solving (3.4). Then set Ve i= Fe X(Xi|Zi),i=1n, as in (3.3). Second stage. [Reduced-form CDF estimation] (1) (QR) (a) For in (005)(e.g., =001) and a fine mesh of Mvalues {= u1uM=1−}, estimate { βe(um)}M m=1by solving (3.6). (b) Obtain Fe Y(y |xZ1i Ve i)as in (3.5). (2) (DR) (a) For each ym∈YM, estimate { β(ym)}M m=1by solving (3.8). (b) Obtain Fe Y(y |xZ1i Ve i)as in (3.7). Third stage. [Structural functions estimation] For ne T=n i=1eiTiand a fine mesh of S values {inf[y∈Y]=y1<···<y S=sup[y∈Y]},compute Ge(yx) =1 ne T n i=1 ei Fe Yy|xZ1i Ve iTi Qe S(τ x) =δ S s=11(ys≥0)−1 Ge(ysx)≥τμe S(x) =δ S s=11(ys≥0)− Ge(ysx) Remark 3. All the estimation steps can also be implemented keeping Z1,orsome component of Z1, fixed as a conditioning variable. The estimated structural functions are then evaluated at values of the conditioning variable(s) of interest. Denoting the DSF estimator and bootstrap draw by G(yxz1)=n i=1 FY(y |xz1 Vi)Ti/nTand Ge(yxz1)=n i=1ei Fe Y(y |xz1 Ve i)Ti/ne T, the corresponding QSF and ASF estimators and bootstrap draws obtain upon substituting G(yxz1)and Ge(yxz1)for G(yx) and Ge(yx) in (3.9)–(3.10). Remark 4. For the QR specification, the estimator of the ASF in the second and third stages can be replaced by μ(x) =w(x ¯ Z10) β,where ¯ Z1=n i=1Z1i/n and βthe least squares estimator of the linear regression of Yon We i. Our numerical implementation in the Online Supplementary Material shows that estimates thus obtained are very similar to those formed according to (3.16). 5.1.2 Inference Algorithm 2is implemented in order to obtain Bweighted bootstrap versions of our estimators. The Bweighted bootstrap estimates are used for the con-
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 525 Algorithm 2 Weighted bootstrap. For b=1B, repeat the following steps: Step 0. Draw eb:={eib}n i=1i.i.d. from a random variable that satisfies Assumption 3(e.g., the standard exponential distribution). Step 1. Reestimate the control function Ve ib = Fe Xb(Xi|Zi)in the weighted sample, according to (3.1)–(3.2)or(3.3)–(3.4). Step 2. Reestimate the reduced form CDF Fe Yb in the weighted sample according to (3.5)– (3.6)or(3.7)–(3.8). Step 3. For ne Tb =n i=1eibTiand the fine mesh of Svalues specified in the third stage of Algorithm 1,compute Ge b(yx) =1 ne Tb n i=1 eib Fe Yby|x Z1i Ve ibTi Qe b(τ x) =δ S s=11(ys≥0)−1 Ge b(ysx)≥τμe b(x) =δ S s=11(ys≥0)− Ge b(ysx) struction of confidence bands for the structural functions, according to Algorithm 3.The resulting confidence bands are then valid uniformly over specified regions of interest. In the empirical application, we run B=199 bootstrap replications in Algorithm 2for both methods. We then implement Algorithm 3in order to perform uniform inference on the structural functions over the specified regions IQ,IG,andIμ. While we found confidence bands to be robust to the choice of Bin the empirical application, the choice of the regions of interest over which to construct the uniform bands has a noticeable effect on the width and shape of the bands. This is especially the case for DR, while QR inference appears to be less sensitive to the choice of regions of interest. In the Online Supplementary Material, we illustrate the effect of varying the definition of regions of interest on weighted bootstrap confidence bands. 5.2 Empirical results Figures 1–3show the QSF, ASF, and DSF for both goods.13 For each structural function, we report weighted bootstrap 90%-confidence bands that are uniform over the corresponding region specified above. Our empirical results illustrate that QR and DR specifications are able to capture different features of structural functions, and are therefore complementary. For food, both estimation methods deliver very similar QSF estimates, close to being linear, although linearity is not imposed in the estimation procedure. For leisure, the QSF and ASF estimated by DR are able to capture some nonlinearity which is absent from those obtained by QR. For QR, this reflects the specified linear structure of the ASF which also constrains the shape of the QSF. In addition, some degree of heteroskedasticity appears to be a feature of the structural model for both goods, although 13For graphical representation the QSF and ASF are interpolated by splines over Xand the DSF over Y.
526 Chernozhukov, Fernández-Val, Newey, Stouli, and Vella Quantitative Economics 11 (2020) Algorithm 3 Uniform inference for structural functions. Step 1. Given Bbootstrap draws {( Ge b(yx) Qe b(τ x) μe b(x))}B b=1, compute the standard errors of G(yx), Q(τx),andμ(x) as σG(y x) =IQR Ge b(yx)B b=1/1349 σQ(τ x) =IQR Qe b(τ x)B b=1/1349σμ(x) =IQRμe b(x)B b=1/1349 Step 2. For b=1B, compute the bootstrap draws of the maximal t-statistics for the D S F, Q S F, a n d A S F a s te Gb(yx) IG=sup (yx)∈IG Ge b(yx) − G(yx) σG(y x) te Qb(τx) IQ=sup (τx)∈IQ Qe b(τ x) − Q(τx) σQ(τx) te μb(x) Iμ=sup x∈Iμμe b(x) −μ(x) σμ(x) Step 3. (i) (DSF and ASF) Form (1−α)-confidence bands for the DSF and ASF as G(yx) ± kG(1−α)σG(y x) :(y x) ∈IGμ(x) ± kμ(1−α)σμ(x) :x∈Iμ where kG(1−α) is the sample (1−α)-quantile of {te Gb(yx)IG:1≤b≤B},and kμ(1− α) is the sample (1−α)-quantile of {te μb(x)Iμ:1≤b≤B}. (ii) (QSF) If Yis continuous, form a (1−α)-confidence band for the QSF as Q(τ x) ± kQ(1−α)σQ(τx) :(τ x) ∈IQ where kQ(1−α) is the sample (1−α)-quantile of {te Qb(τx)IQ:1≤b≤B}. Otherwise, form a (1−α)-confidence band for the QSF as G← U(τ x) G← L(τ x):(τx) ∈I← G where I← G=(τ x) : GL(yx) =τ (y x) ∈IG∩(τ x) : GU(yx) =τ (y x) ∈IG with GL(yx) = G(yx) − kG(1−α)σG(y x) GU(yx) = G(yx) + kG(1−α)σG(y x) much more markedly for leisure, so our methods are well suited for this problem. Increased dispersion across quantile levels in Figure 1is reflected by the increasing spread across probability levels between the two extreme DSF estimates in Figure 3,animportant feature of the data highlighted in Imbens and Newey (2009). Our baseline models naturally allow for the inclusion of transformations of covariates—for instance spline transformations—in order to account for potential nonlinear-
Quantitative Economics 11 (2020) Semiparametric estimation of structural functions 533 Newey, W. and S. Stouli (2018), “Control variables, discrete instruments, and identification of structural functions.” Available at arXiv:1809.05706.[514] R Development Core Team (2019), R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. [523] Spady, R. and S. Stouli (2018), “Dual regression.” Biometrika, 105, 1–18. [509] Stouli, S. (2015), Construction of Structural Functions in Conditional Independence Models. Ph.D. thesis, University College, London. [505] Torgovitsky, A. (2015), “Identification of nonseparable models using instruments with small support.” Econometrica, 83 (3), 1185–1197. [506] Wooldridge, J. (2015), “Control function methods in applied econometrics.” Journal of Human Resources, 50, 420–445. [504] Co-editor Christopher Taber handled this manuscript. Manuscript received 5 November, 2018; final version accepted 24 April, 2019; available online 1 July, 2019.