Linear regression with many controls of limited explanatory power
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Li, Chenchuan Mark; Müller, Ulrich K. Article Linear regression with many controls of limited explanatory power Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Li, Chenchuan Mark; Müller, Ulrich K. (2021) : Linear regression with many controls of limited explanatory power, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 12, Iss. 2, pp. 405-442, https://doi.org/10.3982/QE1577 This Version is available at: https://hdl.handle.net/10419/253605 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 12 (2021), 405–442 1759-7331/20210405 Linear regression with many controls of limited explanatory power Chenchuan (Mark)Li Department of Economics, Princeton University Ulrich K. Müller Department of Economics, Princeton University We consider inference about a scalar coefficient in a linear regression model. One previously considered approach to dealing with many controls imposes sparsity, that is, it is assumed known that nearly all control coefficients are (very nearly) zero. We instead impose a bound on the quadratic mean of the controls’ effect on the dependent variable, which also has an interpretation as an R2-type bound on the explanatory power of the controls. We develop a simple inference procedure that exploits this additional information in general heteroskedastic models. We study its asymptotic efficiency properties and compare it to a sparsity-based approach in a Monte Carlo study. The method is illustrated in three empirical applications. Keywords. High dimensional linear regression, L2 bound, invariance to linear reparameterizations. JEL classification. C12, C21. 1. Introduction A classic issue that arises frequently in applied econometrics is how to deal with a potentially large number of control variables in a linear regression. In observational studies, the plausibility of an unconfoundedness assumption often hinges on having correctly controlled for the value of predetermined variables, which might require including higher order interactions, leading to many control variables. As is well understood, excluding controls that have nonzero coefficients in general yields estimators with omitted variable bias, and corresponding confidence intervals with less than nominal coverage. In empirical practice, this issue is often addressed by reporting results from several specifications that vary in the number and identity of included control variables. A seemingly more systematic approach is to use a pretest to identify which controls have nonzero coefficients, such as testing down procedures, or information criteria, and then proceed with standard inference using only the selected controls. As stressed by Chenchuan (Mark) Li: [email protected] Ulrich K. Müller: [email protected] We thank Michal Kolesar, two anonymous referees, and participants at various workshops for useful comments and advice, and Guido Imbens for providing us with the data of Section 6.2. Müller gratefully acknowledges financial support from the National Science Foundation through grant SES-1627660. ©2021 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1577
406 Li and Müller Quantitative Economics 12 (2021) Leeb and Pötscher (2005) (also see Leeb and Pötscher (2008a,2008b) and the references therein), however, this does not yield uniformly valid inference: If a control coefficient is of order O(n−1/2)in a sample of size n, then it is not selected with probability one, yet it induces an omitted variable bias that is still large enough to yield oversized confidence intervals. This speaks to a broader theoretical result that in the regression model with Gaussian errors, a hypothesis test either overrejects for some value of the control coefficients, or its power is uniformly dominated by the “long regression” that simply includes all potential controls. Hence, an assumption on the control coefficients is necessary to make progress. In that context, the empirical practice of reporting several specifications amounts to two extremes: A specification that does not include a set of potential control variables is justified under the assumption that all coefficients are zero, while the specification with the control variables leaves them entirely unconstrained. A potentially more attractive middle ground is an assumption that the control coefficients are, in some sense, of limited magnitude. One formalization of this idea that has spawned a burgeoning literature is the assumption of sparsity (Tibshirani (1996), Fan and Li (2001), etc.). Most of the control coefficients are known to be zero (or very close to zero), but it is not known which ones. A standard Lasso implementation does not lead to valid inference about the coefficient of interest. But by combining a sparsity assumption on the control coefficients with a sparsity assumption on the correlations between the regressor of interest and the control variables, recent work by Belloni, Chernozhukov, and Hansen (2014) shows how a novel Lasso based “double selection procedure” does yield uniformly valid large-sample inference (also see Zhang and Zhang (2014)andvan de Geer, Bühlmann, Ritov, and Dezeure (2014) for related approaches). While this work is important progress, a sparsity assumption might not always be a compelling starting point: In social science applications, it is usually not obvious why the large majority of control coefficients should be very nearly zero. In addition, the sparsity restriction does not remain invariant to linear reparameterizations of the controls. For instance, in the context of technical controls that are functions of an underlying continuous variable, sparsity drives a distinction between specifying the controls as powers or Chebyshev polynomials, and when including a set of fixed effects, in general it matters which one is dropped to avoid perfect multicollinearity. Finally, in a Lasso implementation, the imposed degree of sparsity is implicitly controlled by a penalty parameter, which makes the small sample interpretation of the resulting inference less than straightforward. This paper develops an alternative approach that considers a priori upper bounds on the weighted average of squared control coefficients, rather than on the number of nonzero control coefficients. To be precise, consider constructing a confidence interval for the scalar parameter βfrom the observations {yixiqizi}n i=1,where yi=βxi+q iδ+z iγ+εi(1) the m×1control variables qiare the baseline specification, the p×1additional variables ziare potential additional control variables, and εiis a conditionally mean zero
Quantitative Economics 12 (2021) Linear regression with many controls 407 error term. To ease notation, assume that zihas been projected off qi,sothatz iγis the contribution of zito the conditional mean of yiafter having controlled for the baseline controls qi.Weimposethebound κ2=n−1 n i=1z iγ2≤¯κ2(2) The parameter κ2is the average of the squared mean effects z iγon yiinduced by zi, that is, κis the quadratic mean of the mean effects of zion yi, after controlling for the baseline controls qi. Small values of ¯κthus embed the a priori assumption that the explanatory power of the controls is small. Let ˆ βshort and ˆ βlong be the coefficients on xifrom a linear regression of yion (xiqi), and from a linear regression of yion (xiqizi), respectively. We combine the information in (ˆ βshortˆ βlong)andthebound(2) to develop a likelihood ratio (LR) procedure that is more informative than the usual confidence interval centered at ˆ βlong. Since (ˆ βshortˆ βlong)andthebound(2) are invariant to linear transformations of the additional controls, so is the new confidence interval. The new interval essentially reduces to the usual intervals centered at ˆ βshort and ˆ βlong for ¯κ=0and ¯κ→∞, respectively.1The intervals thus provide a continuous bridge between omitting the additional controls and including them with unconstrained coefficients. Choosing ¯κin practice is difficult. At the same time, it is arguably no more difficult than choosing the a priori degree of sparsity of γ, say. Typical implementations of sparsity-based inference use penalty-based implicit choices for the level of sparsity, which makes it even harder to relate to the effective constraint that is imposed.2And, as noted above, it is impossible to sharpen long regression based inference without additional constraints on γ, so the implicit assumption embedded in the choice of penalty terms is the substantive constraint that drives the validity of sparsity-based inference. In contrast, the interpretation of ¯κas the quadratic mean of the effect of zion yi makes it more explicitly interpretable. It might also be useful to consider the ratio of n¯κ2and the sum of squared residuals of a regression of yion qi;thisR2-type ratio is the upper bound on the fraction of the variability of yithat is explained by the effect of zi on yiunder the null hypothesis of β=0, after controlling for the baseline controls qi. Thus, beliefs about plausible upper bounds on the explanatory power of ziin terms of R2values directly translate into plausible values of ¯κ,andviceversa. Still, we expect that empirical researchers will typically not argue for a particular ¯κ, but report results for a range of values. In this way, readers learn about the sensitivity of the results to the additional controls in a more comprehensive manner compared to the ¯κ=0short regression and ¯κ→∞long regression extremes. It turns out that when the short regression rejects, and the long regression does not, then there is a unique ¯κ∗ LR such that for all ¯κ<¯κ∗ LR, the LR based test rejects, and for ¯κ>¯κ∗ LR, it does not. The resulting threshold value ¯κ∗ LR, and the associated R2-type ratio, thus form interpretable 1See Section 3for details. 2Indeed, recent work by Wüthrich and Zhu (2020) documents severe small sample size distortions of the LASSO-based post-double-selection method even in some very sparse designs.
408 Li and Müller Quantitative Economics 12 (2021) summaries about the robustness of the statistical significance of βto allowing for additional controls. Whether the value ¯κ∗ LR is substantively large or small depends on the particular situation at hand; see Section 6below for three empirical examples and discussion. Our suggested test and confidence interval is based on the Likelihood Ratio (LR) test statistic obtained from the large sample normality of (ˆ βshortˆ βlong)andtheboundon the omitted variable bias of ˆ βshort implied by (2). From an econometric theory perspective, it is interesting to investigate whether this simple “bivariate” approach comes close to efficiently exploiting the information contained in (2). To this end, we consider the Gaussian homoskedastic version of the regression model (1) and consider asymptotics where the number of additional controls pis of the same order of magnitude as the sample size n. Our main theoretical finding is that in this model, tests that depend on the data only through (ˆ βshortˆ βlong)are asymptotically efficient in a well-defined sense as long as κ=o(n−1/4).Thisratecorrespondstoaratioofnκ2to the sum of squared residuals of a regression of yion qiof order o(n−1/2). While converging to zero, this rate allows for finitely many nonzero coefficients of order o(n−1/4), which would lead to corresponding individual t-statistics that diverge at the rate o(n1/4). It also allows for a fraction o(n1/4)of control coefficients of the already problematic order O(n−1/2). Since we expect that our procedure is most valuable in cases where the additional control coefficients are not obviously relevant a priori, this limited efficiency result is thus still useful. The validity of the suggested inference does not depend on any assumptions about κ or ¯κ. L2penalties of the form (2) play a key role in ridge regression (Hoerl and Kennard (1970)), but our set-up uses (2) as a constraint on the nuisance parameter γonly. Furthermore, our focus is on hypothesis testing and confidence intervals, and ridge regression estimators do not automatically lead to shorter confidence intervals (see, for instance, Obenchain (1977)). Armstrong and Kolesár (2018) derive small sample minimax optimal confidence intervals in a class of Gaussian regression models with the regression function an element of a known convex set. As they point out in Section 4.1.2 of the corresponding working paper Armstrong and Kolesár (2016), their generic results could be applied to (1) under the bound (2), and we provide some comparison with the LR confidence interval in our Section 2.2 below. Our approach of exploiting an a priori bound on the value of a nuisance parameter is also related in spirit to the analysis of Conley, Hansen, and Rossi (2012), who consider instrumental variable estimation with an imperfect instrument that has a direct effect on the outcome of bounded magnitude. The rest of the paper is organized as follows. Section 2contains the analysis of the Gaussian linear regression model (1). In this model, bivariate LR inference is exact, and we analyze and compare its properties. Section 2.3 derives the asymptotic efficiency result for bivariate inference. Section 3discusses the implementation of feasible inference for non-normal, possibly heteroskedastic and clustered linear regressions. Section 4 contains two extensions: First, we discuss instrumental variable regression with a scalar instrument and a scalar endogenous variable, and second, how to further sharpen inference under an additional bound on the explanatory power in the population regression
Quantitative Economics 12 (2021) Linear regression with many controls 409 of xion the potential controls zi.Section5provides a small sample Monte Carlo analysis of our procedure and compares it to the double selection Lasso procedure proposed by Belloni, Chernozhukov, and Hansen (2014). Section 6provides a self-contained description of the suggested methodology, and applies it in three empirical illustrations. Section 7concludes. All proofs are collected in the Appendix. 2. Gaussian linear model 2.1 Set-up Write model (1)invectorformas y=xβ+Qδ+Zγ+ε(3) in obvious notation. To ease notation, assume that xand the additional controls Zhave been projected off the baseline controls Q(so that Qx=0and ZQ=0), and that xand Zare normalized to satisfy xx=nand ZZ=nIp. Our efficiency results focus on the simplest model where the regressors (xQZ)are nonstochastic and ε∼N(0In). We assume throughout that (xQZ)is of full column rank. The (1+m+p) vector of OLS estimators ⎛ ⎜ ⎝ˆ βlong ˆ δ ˆ γ⎞ ⎟ ⎠=⎛ ⎜ ⎝ n0x Z 0Q Q0 Zx0nIp ⎞ ⎟ ⎠ −1⎛ ⎜ ⎝ xy Qy Zy⎞ ⎟ ⎠∼N⎛ ⎜ ⎜ ⎝⎛ ⎜ ⎝ β δ γ⎞ ⎟ ⎠⎛ ⎜ ⎝ 10x Z 0Q Q0 Zx0nIp ⎞ ⎟ ⎠ −1⎞ ⎟ ⎟ ⎠(4) form a sufficient statistic. Inference about βthus becomes inference about one element of the mean of a p+m+1dimensional multivariate normal with known covariance matrix. Let Y=(yxQZ)∈R(2+m+p)n be the observed data, let ϑ=(βδγ)∈R1+m+p, and let ϕβ0(Y)∈{01}be nonrandomized level αtests of the null hypothesis H0:β=β0, where ϕβ0(Y)=1indicates rejection. A confidence set of level 1−αis obtained by “inverting” the family of tests ϕβ0, that is, by collecting the values of β0for which the test does not reject, CS(Y)={β0:ϕβ0(y)=0}. By Proposition 15.2 of van der Vaart (1998), for one-sided hypothesis tests about β, the uniformly most powerful test is simply based on the statistic ˆ βlong, and the uniformly most powerful unbiased test is based on the statistic |ˆ βlong|.ByPratt (1961), the inversion of these uniformly most powerful tests yield confidence sets of minimal expected length: Let (−∞U(Y)) be a confidence interval obtained from inverting one-sided tests of the form H0:β≥β0against Ha:β<β 0. For a given realization Yand true value β, the excess length of this interval is max(U(Y)−β 0)=∞ β(1−ϕβ0(Y))dβ0. By Tonelli’s theorem, Eϑ[∞ β(1−ϕβ0(Y))dβ0]= ∞ βEϑ[1−ϕβ0(Y)]dβ0, and the integrand on the right-hand side is minimized by a family of uniformly most powerful tests, indexed by β0. Similarly, for a two-sided test, the length of the resulting confidence set can be written as (1−ϕβ0(Y))dβ0,soweobtain Eϑ[(1−ϕβ0(Y))dβ0]=(1−Eϑ[ϕβ0(Y)])dβ0and the inversion of uniformly most powerful unbiased tests thus yield the confidence interval of shortest expected length
410 Li and Müller Quantitative Economics 12 (2021) among all unbiased confidence intervals. In the Gaussian model, no procedure whatsoever can therefore do better than simply running the “long regression” that includes all controls in a well-defined sense. 2.2 Bivariate inference problem In order to exploit the bound (2) for more informative inference, consider the coefficient estimator ˆ βshort from the regression of yon (xQ)that excludes the additional controls Z. Since Qx=0,ˆ βshort =xy/n.Letρ2=xZZx/n2,theobservedR2of a regression of xon Z. To avoid trivial complications in notation, assume 0<ρin the following. Straightforward algebra yields ˆ βlong ˆ βshort∼N β β+Δn−1Σ(ρ)Σ(ρ) =⎛ ⎝ 1 1−ρ21 11 ⎞ ⎠(5) where Δ=xZγ/n is the unknown omitted variable bias. Equation (5) is intuitive: the long regression provides an unbiased signal ˆ βlong about β, but with a variance that is larger than the (typically biased) signal ˆ βshort from the short regression.3If ρ→0,then Zis orthogonal to x, there is no bias from the short regression, and the two signals are identical, ˆ βlong =ˆ βshort. Notice that κ2=γγin (2) may be rewritten as κ2=γZxxZZx−1xZγ+γMργ =ρ−2Δ2+γMργ(6) where Mρ=In−Zx(xZZx)−1xZ.Theboundκ2≤¯κ2in (2) thus implies an upper bound on the omitted variable bias, |Δ|≤ρ¯κ(7) and this bound is sharp. This limit on the magnitude of the omitted variable bias in (5) makes ˆ βshort potentially valuable for inference about β, especially if ρis close to one (so that ˆ βshort is much less variable than ˆ βlong). We focus in the following on tests of H0:β=0, since the general case H0:β=β0 may be reduced to this case by subtracting β0from ˆ βlong and ˆ βshort. In terms of the localized parameters b=√nβ,d=√nρ−1Δ,and ¯ k=√n¯κ, the inference problem then becomes testing H0:b=0from observing the bivariate normal vector ˆ b=(ˆ blongˆ bshort)∼ (√nˆ βlong√nˆ βshort), ˆ blong ˆ bshort∼N b b+ρdΣ(ρ)|d|≤¯ k (8) 3This strict ranking of the variance of the short and long regression estimators holds because we consider fixed regressors; see Section 4.2 for discussion.
Quantitative Economics 12 (2021) Linear regression with many controls 411 Figure 1. Five percent critical value of LR(¯ k) as a function of ¯ k. The inference problem (8) is a fairly transparent small sample problem indexed by two known parameters (ρ ¯ k) ∈(01)×[0∞), and involves a one-dimensional unknown nuisance parameter d∈R. The second observation ˆ bshort augments the usual Gaussian shift experiment, and there are a variety of potential approaches to exploiting this additional information. We found that a simple but effective test of H0:b=0is generated by the generalized likelihood ratio statistic LR(¯ k) =min |˜ d|≤¯ kˆ blong ˆ bshort −ρ˜ dΣ(ρ)−1ˆ blong ˆ bshort −ρ˜ d −min ˜ b|˜ d|≤¯ kˆ blong −˜ b ˆ bshort −˜ b−ρ˜ dΣ(ρ)−1ˆ blong −˜ b ˆ bshort −˜ b−ρ˜ d(9) The level αcritical value cvρ(¯ k) is the largest 1−αquantile under (8)withb=0, maximized over |d|≤¯ k. Figure 1plots cvρ(¯ k) for ρ∈{05095099}, and Figure 2plots the rejection region of the resulting 5% level test for ρ=095 and ¯ k∈{01310}.For Figure 2. Acceptance regions of LR(¯ k) for ρ=095. Notes: The lines are the boundaries of the acceptance region. For all values of ¯ k,(00)is in the acceptance region.
412 Li and Müller Quantitative Economics 12 (2021) ¯ k=0, the LR test reduces to rejecting for large values of (ˆ bshort)2>cvρ(0)=1962,that is, it reduces to the usual t-test based on the short regression. More generally, whenever |ˆ bshort|ρ¯ k, that is the short regression coefficient is much larger than ρ¯ kin absolute value, then the LR test rejects. On the other hand, for |ˆ bshort|ρ¯ kand ¯ klarge, the LR test rejects when (1−ρ2)( ˆ blong)2>cvρ(¯ k), that is, whenever the long regression coefficient is too large in absolute value, with a critical value that is slightly larger than 1962.Once ¯ kis moderately large (say, larger than 8), the critical value cvρ(¯ k) stabilizes at cvρ(∞), and further increases of ¯ ksimply amount to an additional elongation of the acceptance region along the ˆ bshort-axis. To formally characterize the limit of the acceptance region for values of |ˆ bshort|≈ρ¯ k under larger and larger bounds ¯ k→∞, consider the observation ˆ bo=(ˆ blongˆ bo short)∼ (ˆ blongˆ bshort −ρs)with ˆ bshort,thatis, ˆ bo short is shifted by ρs relative to ˆ bshort.Undera corresponding reparameterization a=d−s,weobtain ˆ bo∼N b b+ρaΣ(ρ)a∈A(10) and the constraint a∈Aresulting from |d|≤¯ kdepends on the relationship between s and ¯ k. In particular, with s=¯ k,a∈A1=(−∞0], and this corresponds to the case where the bound ¯ kis very large and ˆ bshort is close to the bound ρ¯ k. Similarly, with s=−¯ k,ˆ bshort is close to −ρ¯ k, and the corresponding constraint in (10) becomes a∈A−1=[0∞). Finally, if ¯ k→∞and ¯ k−|s|→∞, so that the bound ¯ kis much larger than |ˆ bshort|, then a∈A0=Ris unrestricted in (10). For each of these three cases i∈{1−10},theLR(¯ k) statistic converges to LRo i=min ˜ a∈Aiˆ blong ˆ bo short −ρ˜ aΣ(ρ)−1ˆ blong ˆ bo short −ρ˜ a −min ˜ b˜ a∈Aiˆ blong −˜ b ˆ bo short −ρ˜ a−˜ bΣ(ρ)−1ˆ blong −˜ b ˆ bo short −ρ˜ a−˜ b We consider this LR approach attractive for a number of reasons. First, it is easy to implement (we discuss implementation issues in more detail in Section 3below). Second, inversion of the LR statistic for general null hypotheses H0:b=b0yields a confidence interval for bthat is translation equivariant, that is, the interval obtained from the observation (ˆ blong +c ˆ bshort +c) simply shifts the interval from (ˆ blongˆ bshort)by c, for any c∈R. Third, it yields confidence intervals that are close to minimal weighted expected length under a weighting function where dis uniform between [−¯ k ¯ k],amongall translation equivariant confidence intervals. This is shown in panel A of Table 1,which reports a lower bound on this weighted expected length for selected values of (ρ ¯ k), along with the weighted expected length of the LR interval. Given the tight link between the power of tests and their expected length discussed in Section 2.1 above, this implies that the LR tests are also close to maximizing the corresponding weighted average power. Fourth, as shown in panel B, it is reasonably close to being maximin in terms of
Quantitative Economics 12 (2021) Linear regression with many controls 419 ˆ βshort, and potentially an independent randomization device, as long as one restricts attention to ¯κ∗ n(Yn)whose distribution does not depend on the nuisance parameters (δω). The local asymptotic properties of the threshold estimator (14) implied by the LR test (13) corresponds to the small sample properties of ¯ k∗ LR(ˆ b)discussed at the end of Section 2.2, so similar to Theorem 2, its attractive features again extend to this larger class. And by setting ψn(Yn)in Corollary 1equal to a potentially recentered and rescaled estimator of βthat exploits the bound (2), the same holds for our suggested midpoint estimator ˆ βLR(¯κ). In summary, under pn→∞asymptotics, as long as τn=o(n−1/4), the quality of asymptotic inference in the Gaussian homoskedastic model is limited from above by the performance of bivariate procedures. The attractive small sample features of the LR approach discussed in Section 2.2 thus translate into attractive large sample inference. 3. Implementation in non-Gaussian and potentially heteroskedastic models In the Gaussian linear regression model, the bivariate tests introduced in Section 2.2 have exact small sample properties. But for applied use, it is important to have a valid implementation in non-Gaussian and potentially heteroskedastic models. With the regressors nonstochastic (or after conditioning on the regressors with a conditionally mean zero error term), the general model is still of the form (1), where now εi∼(0σ2 i) independent across i. Under weak technical conditions on the tails of the distribution of εi, on the sequence {σ2 i}n i=1and on the regressors {xiqizi}n i=1, a central limit theorem yields Ω−1/2 nˆ βlongn −βn ˆ βshortn −βn−Δn⇒N(0I2)(15) for some suitably defined Ωn, since (ˆ βlongn −βnˆ βshort −βn−Δn)are linear combinations of the heterogeneous but mean zero and independent random variables {εi}n i=1. We provide a corresponding result in the supplemental Appendix that allows for dependence among the εidue to clustering. Suppose ˆ Ωnis a consistent estimator of Ωnin the sense that Ω−1 nˆ Ωn p →I2.Thenatural LR statistic of H0:βn=0under the bound (2) then becomes LRn(¯κn)=min |˜ Δ|≤ρn¯κn/√x nxn/nˆ βlongn ˆ βshortn −˜ Δˆ Ω−1 nˆ βlongn ˆ βshortn −˜ Δ −min ˜ β|˜ Δ|≤ρn¯κn/√x nxn/nˆ βlongn −˜ β ˆ βshortn −˜ β−˜ Δˆ Ω−1 nˆ βlongn −˜ β ˆ βshortn −˜ β−˜ Δ where ρ2 nis the R2of a regression of xion zi. Exploiting the invariance of the LR statistic to reparameterizations, the distribution of LRn(¯κn)under the approximations (15)and
420 Li and Müller Quantitative Economics 12 (2021) ˆ Ωn=Ωnis effectively indexed by χ=(χ1χ2)with χ1=|Ω11 −Ω12| Ω11Ω22 −Ω2 12 (16) χ2=Ω11 Ω11Ω22 −Ω2 12 ρn¯κn x nxn/n(17) where Ωij is the i,jth element of Ωn, and under the null hypothesis of βn=0and Ω11Δn/Ω11Ω22 −Ω2 12 →g, the asymptotic distribution of LRn(¯κn)is equal to min |˜ g|≤χ2Z1 Z2+g−˜ gZ1 Z2+g−˜ g −min ˜ h|˜ g|≤χ2Z1−˜ h Z2+g−χ1˜ h−˜ gZ1−˜ h Z2+g−χ1˜ h−˜ g(18) where (Z1Z2)∼N(0I2). This limit distribution depends on the nuisance parameter |g|≤χ2, but a numerical calculation shows that its 1−αquantile is maximized at g=χ2 for α∈{00100501}and all χ1. It is hence straightforward to obtain the appropriated critical value cv(χ)via simulation, and we provide a corresponding look-up table in the replication files. Alternatively, a linear interpolation of the values in Table 2that only depend on χ1generate (slightly conservative) critical values or all χ2(cf. Figure 1). Either way, a subsequence argument then yields asymptotic validity of this feasible LR test ˆϕLRn(¯κnYn)=1[ LRn(¯κn)>cv(ˆ χn)],where ˆ χn=(ˆχn1ˆχn2)are as in (16)and(17), with the elements of Ωnreplaced by those of ˆ Ωn. Lemma 3. (a) If Ω−1 nˆ Ωn p →I2and (15)holds,then limsupn→∞ Eθn[ˆϕLRn(¯κnYn)]≤αfor all sequences θnwith βn=0and |Δn|≤ρn¯κn/x nxn/n. (b) Under the assumptions of Theorem 2,Eθn[ˆϕLRn(¯κnYn)]−Eθn[ϕLRn(¯κnYn)]→0. Note that the asymptotic validity in part (a) holds without any assumptions about the sequences pnor ¯κn. In particular, it is not required that pn/n →c∈(01)or ¯κn= o(n−1/4). In the Gaussian homoskedastic model, Ωnis equal to (x nxn)−1Σ(ρn), and in large samples, ˆϕLRn reduces to the bivariate LR test introduced in Section 2.2. Formally, Table 2. Interpolation table for upper bound on cv(χ). α\χ102581225∞ 001 6663 6931 7170 7218 7251 7287 7306 005 3845 3959 4081 4142 4174 4203 4219 010 2711 2750 2810 2870 2898 2926 2941 Note: Linear interpolation within each row yields slightly conservative asymptotic critical values for LRn(¯κn).
Quantitative Economics 12 (2021) Linear regression with many controls 421 part (b) of the Lemma shows that the large sample power properties of ˆϕLRn in the Gaussian homoskedastic model are equal to the small sample power properties of the LR test as introduced in Section 2.2.Thus, ˆϕLRn(¯κnYn)has the same asymptotic efficiency properties as ϕLRn(¯κnYn)discussed below Theorem 2, even among tests that depend on the data beyond the short and long regression coefficient. Given Lemma 3, the only obstacle to a straightforward implementation of the LR test in a more general model is the estimation of the asymptotic variance Ωn.Ifthenumber of controls pis fixed, or only slowly increasing with n, the usual heteroskedasticity robust White (1980)estimatorforΩnis consistent under reasonably weak assumptions. However, under asymptotics where pn/n →c∈(01), as employed for the asymptotic efficiency argument in Theorem 2,Cattaneo, Jansson, and Newey (2018a) show that the White (1980) estimator is no longer consistent, and Cattaneo, Jansson, and Newey (2018b) provide an alternative estimator that remains consistent. Alternatively, if the explanatory power of the additional controls is limited in the sense that κn=o(1),onemay also consistently estimate ˆ Ωnfrom the usual White formula based on the residuals from the short regression that only includes the baseline controls. This has the advantage of being readily implementable also with clustering. We provide a corresponding result in the supplemental Appendix. Given any value of ¯κ≥0, a confidence set for βis obtained by collecting the values for β0such that the test H0:β=β0based on the LR statistic does not reject (in the following, we drop nsubscripts again to ease notation). For ¯κ=0,andunderhomoskedasticity, this yields the same interval as obtained from standard short regression inference using the 22element of ˆ Ωas the variance estimator. In small samples, when ˆ Ωdoes not impose homoskedasticity, the confidence interval for ¯κ=0is centered at a slightly different value, since under heteroskedasticity, it is in general more efficient to estimate βby a linear combination of ˆ βlong and ˆ βshort that puts nonzero weight on ˆ βlong. For ¯κ→∞, the interval is exactly centered at ˆ βlong, but the LR test uses a slightly larger critical value, as discussed in Section 2.2 above. 4. Extensions 4.1 Instrumental variable regression Suppose the scalar regressor xiof interest in the linear regression (1)isendogenous,but we have access to a scalar instrument wi(wicould be a linear combination of a vector of instruments, such as in two stage least squares). As in the baseline model, we treat {wiqizi}n i=1as nonstochastic, or equivalently, we condition on their realization in the following. To simplify notation, let wibe orthogonal to the baseline controls qi. Assume that the data is generated via xi=ηwi+q iδx+z iγx+εxi(19) yi=βxi+q iδ+z iγ+εi(20) where (εxiεi)is mean-zero independent across i, but potentially heteroskedastic, and if εxi is correlated with εi, the regressor xiin (20) is endogenous.
422 Li and Müller Quantitative Economics 12 (2021) Let (ˆ βIV longˆ βIV short)be the IV estimators of βthat include or exclude the additional controls zi. These estimators involve the term n i=1wixi,whichunder(19) is stochastic and depends on the realization of εxi, complicating the description of their bias. In order to avoid these difficulties, we focus on their moment condition instead, as in Anderson and Rubin (1949). Let ˆ wz ibe the residuals of a regression of wion zi. The estimators ˆ βIV long and ˆ βIV short are identified from the two moment conditions E[n−1n i=1ˆ wz iεi]=0 and E[n−1n i=1wiεi]=0, respectively. Consider testing H0:β=0(nonzero values can be reduced to this case by subtracting β0xifrom yi). Similar to (15), under H0,theempirical moment conditions then satisfy ΩIV−1/2⎛ ⎜ ⎜ ⎜ ⎜ ⎝ n−1 n i=1ˆ wz iyi n−1 n i=1 wiyi−ΔIV ⎞ ⎟ ⎟ ⎟ ⎟ ⎠⇒N(0I2)(21) for some suitably defined ΩIV, which can be consistently estimated by ˆ ΩIV.The“bias” ΔIV in (21)isgivenbyΔIV =n−1n i=1wiz iγ, and a straightforward calculation shows that under (2), we have the sharp bound ΔIV≤¯κwZZZ−1Zw/n with w=(w1wn). Thus, under the null hypothesis of H0:β=0, the observations (21) have the identical structure as (15) of the previous section, and one can apply the LR test in entirely analogous fashion to obtain a valid large sample test that exploits the bound (2) to sharpen inference in instrumental variable regression. Our focus on the moments (21), rather than the estimators (ˆ βIV longˆ βIV short), has the additional appeal that no assumptions about the strength of the instrument are required. 4.2 Double bounds In Sections 2and 3, we have treated the regressors {xiqizi}n i=1as either nonstochastic, or the analysis conditioned on their value. In the simple Gaussian model of Section 2 with random regressors, the Gram matrix forms an ancillary statistic. It is textbook advice to condition inference on ancillary statistics in general and on the Gram matrix in particular (see, for instance, Chapter 2.2 in Cox and Hinkley (1974)), providing a rationale for our analysis. Furthermore, our approach does not require or depend on a model for the potentially stochastic properties of the regressors. This is attractive in so far as it relieves applied researchers from having to defend a particular data generating mechanism, and avoids a source of potential misspecification. We now discuss how one could exploit additional assumptions on the generation of the regressor of interest xito potentially further sharpen inference about β.Inparticular, assume that ˜ xiis generated by the linear model ˜ xi=q iδx+z iγx+εxi(22)
Quantitative Economics 12 (2021) Linear regression with many controls 423 where εxi is conditionally mean zero given {qizi}n i=1. The regressor xiis simply defined as the residuals of a least squares regression of ˜ xion qi,sothatconsistentwithournotation above Qx=0. We maintain, as in Sections 2and 3,thatεiin (1)isconditionally mean zero (so ˜ xiis not endogenous, and no instrument is required). Assume further that we are willing to assume that in addition to (2), also κ2 x=n−1 n i=1z iγx2≤¯κ2 x(23) so that ¯κxhas the interpretation of an upper bound on the quadratic mean of the effect of zion xi, after controlling for qi. This “double bounds” structure of limiting the population coefficients in both the regression of interest (1), and the auxiliary regression (22), parallels the assumptions validating the double Lasso procedure by Belloni, Chernozhukov, and Hansen (2014). As in the previous subsection, it is convenient to focus on the moment conditions defining the OLS estimators (ˆ βlongˆ βshort): Under weak regularity conditions, (2), (22), and (23) imply that under H0:β=0 ΩDbl−1/2⎛ ⎜ ⎜ ⎜ ⎜ ⎝ n−1 n i=1ˆ xz iyi n−1 n i=1 xiyi−ΔDbl ⎞ ⎟ ⎟ ⎟ ⎟ ⎠⇒N(0I2) (24) where ˆ xz iare the residuals of a regression of xion zi,andΔDbl satisfies the sharp bound |ΔDbl|≤¯κ·¯κx(see the supplemental Appendix for details). With an appropriate estimator ˆ ΩDbl, this again has the same structure as the problem discussed in Section 3,sothe LR test defined there can be used to exploit the additional information contained in (22) and (23). 5. Small sample simulations In this section, we use Monte Carlo simulations to evaluate the finite-sample properties of confidence intervals based on ˆϕLR,andwecompareittotheperformanceofthe Lasso-based post-double-selection technique of Belloni, Chernozhukov, and Hansen (2014) (abbreviated BCH in the following two sections). As in BCH’s Monte Carlo, we set the total number of observations to n=500,let p=200, and generate data from a model where the baseline control is simply a constant, yi=˜ δ1+˜ xiβ+˜ z iγ+εii=1n (25) with εi∼iid N(01)independent of {˜ xi˜ zi},and˜ ziis generated by the linear model ˜ xi=˜ z iμ+εx i(26) with εx i∼iid N(01)independent of {˜ zi}. To be consistent with our previous notation, we orthogonalize the regressors in (25) off the baseline control, that is, xi=
424 Li and Müller Quantitative Economics 12 (2021) ˜ xi−n−1n l=1˜ xland zi=(zi1zip)with zij =˜ zij −n−1n l=1˜ zlj ,sothat(25)implies the linear model (1) with an appropriate definition of δ1.Wesetβ=0throughout. Our designs vary according to the value of four parameters: the previously introduced ρ2∈{06095}and κ∈{0205}; the scalar η∈{0103}determines the degree of sparsity of γ=(γ1γp)and μ=(μ1μp);andν∈{0051}determines the overlap between the nonzero indices of γand μ. Specifically, γj=cγ1[j≤ηp],wherethe scalar cγis chosen such that the implied value of κ2is equal to the specified value, and μj=cμ1[η(1−ν)p+1≤j≤η(2−ν)p],j=1p,wherecμ∈Ris chosen such that the sample R2of a regression of xion ziis equal to ρ2. The parameter ηplays a crucial role for the BCH method, since the method requires that the number of nonzero values in γand μis not too large. In contrast, the test ˆϕLR remains numerically invariant to any linear reparameterizations of the regressors. Finally, the parameter νdetermines the omitted variable bias in the short regression coefficient ˆ βshort (which is the coefficient on xiin the regression of yion (1xi)). Under ν=0,there is no overlap, and the variables zij with nonzero coefficient γjare uncorrelated with the regressor of interest xi, so there is no omitted variable bias, at least over repeated samples with random regressors. In the other extreme, with ν=1, every variable zij with nonzero coefficient γjis correlated with xi, leading to a large omitted variable bias. We consider four types of confidence intervals for β. First, the usual confidence interval based on ˆ βshort. Second, the usual confidence interval based on ˆ βlong. Third, the confidence interval obtained by inverting the feasible test ˆϕLR introduced in Section 3, whereweset¯κequal to the actual value of κ. Fourth, the Lasso-based post-double- selection method “LPDS” from BCH, as specified in their Monte Carlo Section 4.2.For the first three types of methods, we estimate standard errors of (ˆ βlongˆ βshort)with the heteroskedasticity-robust estimator of Cattaneo, Jansson, and Newey (2018b). In addition, we report quartiles of the threshold value ¯κ∗ LR ∈[0∞)∪{+∞}computed from the family of tests ˆϕLR for each draw, defined to be zero if ¯κ=0does not lead to rejection, and +∞if none of the ¯κvalues lead to rejection. Table 3contains the results. The confidence interval based on ˆ βshort has coverage substantially below the nominal level whenever the overlap parameter νis positive. This shows that the considered values of κare large enough to severely distort inference that simply sets the control coefficients to zero. In contrast, the interval associated with ˆ βlong has size very close to the nominal level throughout, but at the cost of being fairly long. The LPDS method sometimes substantially undercovers even in the relatively sparse design with η=01. Apparently, the relatively small values of γjmake it difficult for the method to correctly pick up the η·p=20 nonzero coefficients, leading to a remaining omitted variable bias that is large enough to induce nonnegligible overrejections. When κ=05and ρ2=095,sothatbothγjand μjare relatively larger, the LPDS method reliably controls size in the η=01sparse design, but yields somewhat longer intervals compared to ˆ βlong.Thenewtests ˆϕLR(κY)control size throughout as they should, given that ¯κ=κtrivially implies κ≤¯κ, and yield intervals that are shorter than the ˆ βlonginterval when ν>0, with larger gains for larger values of ρand ν. Of course, with κunknown in practice, one cannot apply ˆϕLR(¯κY)with κ=¯κ.We expect that in practice, researchers will consider a range of values of ¯κto gauge the sensitivity of the results about β. Then, by construction, researchers will also consider ¯κ=κ,
Quantitative Economics 12 (2021) Linear regression with many controls 425 Table 3. Small sample properties for n=500 and p=200. ηρ 2νˆ βshort ˆ βlong LPDS ˆϕLR(κY)¯κ∗ LR Cov Lgth Cov Lgth Cov Lgth Cov Lgth Q1 Q2 Q3 κ=020 010 060 000 094 014 094 022 095 023 094 022 000 000 000 010 060 050 074 014 095 022 092 023 095 022 000 000 000 010 060 100 027 014 094 022 082 023 095 020 000 004 008 010 095 000 094 005 094 022 095 026 098 014 000 000 000 010 095 050 042 005 095 022 095 026 098 014 000 002 005 010 095 100 001 005 095 022 095 025 095 013 008 011 015 030 060 000 094 014 095 022 095 021 095 022 000 000 000 030 060 050 074 014 094 022 087 021 094 022 000 000 000 030 060 100 027 014 094 022 062 021 095 020 000 004 008 030 095 000 094 005 095 022 095 018 098 014 000 000 000 030 095 050 043 005 095 022 091 017 098 014 000 002 005 030 095 100 001 005 095 022 081 017 095 013 008 011 015 κ=050 010 060 000 092 014 094 022 095 024 094 022 000 000 000 010 060 050 013 014 094 022 077 024 094 022 003 008 012 010 060 100 000 014 094 022 038 024 095 022 022 026 031 010 095 000 091 005 095 022 095 027 095 022 000 000 000 010 095 050 000 005 094 022 095 026 096 019 013 016 020 010 095 100 000 005 094 022 095 025 095 015 037 041 044 030 060 000 091 014 094 022 095 022 094 022 000 000 000 030 060 050 012 014 094 022 050 022 094 022 004 008 012 030 060 100 000 014 094 022 003 022 095 022 022 027 031 030 095 000 092 005 094 022 095 018 095 022 000 000 000 030 095 050 000 005 094 022 074 018 096 019 013 016 020 030 095 100 000 005 094 022 042 017 095 015 037 041 044 Note: Entries are coverage and average length of 95% confidence intervals for β, and the quartiles of the distribution of ¯κ∗ LR. Rows correspond to different DGPs, with ηmeasuring the sparsity of the design, νthe overlap between the nonzero indices on ziin the regressions of yion ziand of xion zi,andρ2is R2of a regression of xion zi. The columns are different confidence intervals, with ˆ βshort and ˆ βlong the confidence interval based on short and long regression coefficients, ˆϕLR(¯κY) the LR based confidence interval developed in this paper that imposes the bound κ≤¯κ, and LPDS is BCH’s Lasso-based postdouble-selection procedure. Based on 20,000 Monte Carlo simulations. and the length of ˆϕLR(κY)in Table 3indicates that at that point, the interval exploiting the bound (2) is often considerably more informative than the interval based on ˆ βlong. In addition, researchers might compute the threshold value ¯κ∗ LR, defined such that ˆϕLR(¯κY)rejects for all ¯κ≤¯κ∗ LR. The last three columns in Table 3report the quartiles of the distribution of ¯κ∗ LR.Forν=1,aswellasfor(κ ν) =(0505), its median is always positive. Thus, in the majority of draws, researchers would have been able to conclude that small upper bounds for ¯κare empirically incompatible with β=0. Unreported results show that if the true value of βis nonzero, these medians become larger. Thus, the LR approach helps sharpen inference about βin a meaningful way.
426 Li and Müller Quantitative Economics 12 (2021) Still, looking over the table, it is tempting to conclude that one should use ˆ βshort whenever there is no overlap, ν=0, as this leads to the shortest intervals by far, and only slight size distortions. Similarly, if ν<1the quartiles of ¯κ∗ LR are much smaller than κ. However, it is not possible to consistently determine the value of νfrom the observations. This is the result of the asymptotic efficiency derivations in Section 2.3:Forsmall values of ¯κ, it is impossible to do better than to construct inference based on the bivariate statistics (ˆ βlongˆ βshort)at least in large samples, and as demonstrated there, the LR approach comes close to exploiting the information contained in this pair of statistics. 6. Empirical applications 6.1 Overview We illustrate the suggested method in three empirical examples based on studies by Macchiavello and Morjaria (2015), Donohue and Levitt (2001), and Imbens, Rubin, and Sacerdote (2001). Our examples serve to highlight the mechanics of the bivariate LR test and illustrate its empirical content. In particular, for each of the studies, we calculate the LR statistic LR(¯κ) over a grid of values for ¯κ, resulting in a family of confidence intervals in βindexed by ¯κ≥0. This family provides an explicit correspondence between assumptions on the control coefficients γand empirical conclusions about the parameter of interest β. As a by-product, we obtain the threshold value ¯κ∗ LR, so that the intervals exclude the zero-effect value β=0for ¯κ<¯κ∗ LR, and contain it otherwise. The intervals are computed in the following steps: 1. Let ybe the n×1vector of outcome variables, ˜ xthe scalar regressor of interest, Q the matrix of baseline controls, and ˜ Zthe matrix of additional controls of questionable relevance. Run the long regression of yon (˜ x˜ ZQ)to find the long coefficient ˆ βlong, and run the short regression of yon (˜ xQ)to find the short coefficient ˆ βshort. 2. Let xand Zdenote the residuals of ˜ x,˜ Zin a regression on Q. Define vi= ((xx)−1xi(ˇ xˇ x)−1ˇ xi),i=1n with ˇ xthe vector of residuals of a linear regression of xon Z. 3. Obtain an estimate ˆ Ωnof the 2×2covariance matrix of (ˆ βlongˆ βshort): (a) If the total number of controls (the number of elements in δand γ)issmall compared to the sample size, a standard heteroskedasticity robust estimator may be used ˆ Ωn= j i∈Gj viei i∈Gj viei(27) where the sets Cjpartition the indices i=1n into clusters (so that for independent samples, Cj={j}), and eiare the residuals of the long regression. (b) If the total number of controls is of the same order as the sample size (say, 5% or more), then without clustering, apply the Cattaneo, Jansson, and Newey
Quantitative Economics 12 (2021) Linear regression with many controls 427 (2018b)estimator ˆ Ωn= n i=1 n j=1 κij e2 iviv i where κij are the elements of the n×nmatrix (MM)−1with M=In− W(WW)−1Wand W=(QZ),denotes the element-by-element product, and eiare the residuals of the long regression. (c) If the number of baseline controls is small, and the number of additional controls is of the same order as the sample size (say, 5% or more), then under clustering, use (27)witheiequal to residuals of the short regression. 4. Compute ρ2=xZ(ZZ)−1Zx/(xx),theR2of a regression of xon Z, and for given ¯κ,χ1and χ2from equations (16)and(17)withΩij the elements of ˆ Ωn. Define the function h:R2→Rvia h(Y1Y2)=h0(Y1Y2)−h1(Y1Y2),where h0(Y1Y2)=Y2 1+1|Y2|>χ 2|Y2|−χ22 h1(Y1Y2)= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ (χ2+χ1Y1−Y2)2 1+χ2 1 if χ2+χ1Y1<Y 2 (χ2−χ1Y1+Y2)2 1+χ2 1 if χ2−χ1Y1<−Y2 0otherwise 5. Compute the level αcritical value cv as the 1−αquantile of h(Z1Z2+χ2),where (Z1Z2)are independent standard normals via simulation (or rely on the look-up table in the replication files, or on the slightly conservative interpolation from Table 2). 6. The level 1−αconfidence interval for βis formed by the values of β0that satisfy hsign(Ω11 −Ω12)ˆ βlong −β0 Ω11 Ω11(ˆ βshort −β0)−Ω12(ˆ βlong −β0) Ω11Ω11Ω22 −Ω2 12 ≤cv and the estimator ˆ βLR(¯κ) is the midpoint of this interval. 6.2 Macchiavello and Morjaria (2015) Macchiavello and Morjaria (2015) use data on African rose exports to identify reputational effects in markets without contract enforcement. The authors construct a model with the feature that binding incentive constraints yield observable proxies for the buyer–seller relationship value during periods of maximum temptation for sellers and buyers to undercut each other. They find, empirically, that this value proxy is correlated with relationship age but not outside prices, evidence that reputation constrains trade in the absence of enforcement. We apply our approach to determine the extent to which
428 Li and Müller Quantitative Economics 12 (2021) the correlation between relationship value and age is sensitive to Macchiavello and Morjaria’s (2015) choice of control variables. We treat the panel regression from Table 5, Column 8 of Macchiavello and Morjaria (2015) as the short regression, where relationships are the unit of observation and the time dimension corresponds to four growing seasons (years), for a total of n=372 observations. This regression has the log of the relationship value as the dependent variable, the regressor of interest is the log of relationship age, and the baseline controls are the maximum of the previous observed log auction value as well as relationship and season fixed-effects. This is a difference-in-differences model in which the main effect is identified by variation in sales across seasons for buyer–seller relationships of different ages. Macchiavello and Morjaria (2015) find that β, the coefficient on relationship age, is statistically significant at standard confidence levels. We investigate the sensitivity of these results to p=123 additional buyer ×season fixed effects. This specification is an extension of the baseline season controls that allows flexibility over buyers. One might imagine that because sellers are located in Kenya, but buyers are located globally, time trends might be more plausibly heterogeneous for buyers. Hence, to the extent that purchase patterns over seasons differed between buyers with relationships of various lengths for reasons unrelated to learning about seller quality, omitting these additional fixed effects could lead to bias in β. However, absent constraints on the coefficient of these additional controls, only variation from seller differences across seasons can identify the main effect β, so including these additional fixed effects in an unconstrained fashion leads to much less informative inference. Figure 3plots 90%,95%,and99% confidence intervals for βfrom our new procedure as a function of ¯κ, along with the point estimates ˆ βLR(¯κ). The standard errors are clustered by seller and are computed as described in Step 3c of Section 6.1. We see that the short regression strongly rejects, but the long regression does not. The largest value of ¯κthat still leads to rejection, ¯κ∗ LR,ofthe5% level test is indicated by a vertical line and equals ¯κ∗ LR =0203. Thus, since the outcome is measured in logs, as long as one believes that season-specific idiosyncratic buyer preferences not already captured by Macchiavello and Morjaria’s (2015) baseline controls induce on average changes in the Figure 3. LR Confidence intervals for βin Macchiavello and Morjaria (2015).
Quantitative Economics 12 (2021) Linear regression with many controls 435 Now ∞ j=1 1 4s2j j! (m +1) (m +j+1)≤∞ j=1 s2j j! (m +1) (m +j+1) ≤∞ j=1s2/mj j!=exps2/m−1 where the second inequality uses the elementary inequality (m+1)mj/(m+j+1)≤1 obtained from repeatedly applying (m+i+1)=(m+i)(m+i) ≤m(m+i) for all i≥0 and m>0. The result now follows from s2 m/m →0under sm=o(m1/2). For ease of notation, we omit the dependence on n(and p=pn), except for tn.From (11), it follows that nˆτ2=nˆ φˆ φwith ˆ φ∼N(τωn−1Ip−1)is distributed noncentral χ2 with p−1degrees of freedom and noncentrality parameter nτ2. Without loss of generality, assume ω=ι=(100).Then,with ˆω=ˆ φ/ˆ φ, from the density of ˆ φand using the notation of the proof of Lemma 1, Ln(tn)=Cexp−1 2nˆτOˆω−ιtn2dHp−1(O) for some constant Cthat does not depend on tn(and note that Ln(tn)does not depend on the realization of ˆω). Thus Ln(tn)/Ln(0)=expntnˆτˆωOι−1 2nt2 ndHp−1(O) We initially show the convergence under τ=0. It then suffices to show that E[(Ln(tn)/ Ln(0)−1)2]→1under ˆ φ∼N(0n−1Ip−1)and an arbitrary sequence tn=o(n−1/4).Observe that ELn(tn)/Ln(0)−12 =Eexptnˆ φOι−1 2nt2 ndHp−1(O)−12 =Eexptnˆ φOι−1 2nt2 ndHp−1(O)−1 ×exptnˆ φ˜ Oι−1 2nt2 ndHp−1(˜ O)−1 =Eexptnˆ φOι−1 2nt2 ndHp−1(O)exptnˆ φ˜ Oι−1 2nt2 ndHp−1(˜ O) −2·Eexptnˆ φOι−1 2nt2 ndHp−1(O)+1 =hn(tn)−2˜ hn(tn)+1
436 Li and Müller Quantitative Economics 12 (2021) In what follows, we show that hn(tn)→1. The convergence ˜ hn(tn)→1follows from the same arguments and is omitted for brevity. Tonelli’s theorem and ˆ φ∼N(0n−1Ip−1)imply hn(tn)=Eexptnˆ φ(Oι+˜ Oι)−nt2 ndHp−1(O)dHp−1(˜ O) =Eexptnˆ φ(Oι+˜ Oι)−nt2 ndHp−1(O)dHp−1(˜ O) =exp1 2nt2 nOι+˜ Oι2−nt2 ndHp−1(O)dHp−1(˜ O) =expnt2 n(Oι)˜ OιdHp−1(O)dHp−1(˜ O) =expnt2 nιOιdHp−1(O) Using the notation of Lemma 4, the formula for the normalizing constant of the von Mises–Fisher distribution (see, for instance, equation (9.3.4) of Mardia and Jupp (2000)) implies hn(tn)=2˜ p/2−1·I˜ p/2−1(nt2 n)·( ˜ p/2)/(nt2 n)˜ p/2−1where ˜ p=p−1. Application of Lemma 4with sn=nt2 nnow yields hn(tn)→1, since under tn=o(n−1/4)and p/n →c∈(01),s2 n=n2t4 n=o(p/2−1). This concludes the proof under τn=0. Now apply this very result to another sequence tn,tn=t n.ThenLn(t n)/Ln(0)p →1implies via LeCam’s first lemma (see, for instance, Lemma 6.4 in van der Vaart (1998)) that in the experiment of observing ˆτ2 n, the sequence τn=t nis contiguous to τn=0.Thus,Ln(tn)/Ln(0)p →1also holds under τn=t n=o(n−1/4)by definition of contiguity, which was to be shown. A.3 Proof of Theorem 2 Let n(Tn)be the log-likelihood ratio statistic based on Tnof testing H0:(baτn)= (b0a0τn0)against H1:(baτn)=(b1a1τn1).Lethjn =(bjbj+ρnaj)and hj= (bjbj+ρaj),j=01.From(11), n(Tn)=x nxnˆ βlongn ˆ βshortn −snΣ(ρn)−1(h1n −h0n)−1 2h 1nΣ(ρn)−1h1n +1 2h 0nΣ(ρn)−1h0n +logLn(τn1) Ln(τn0) and with 0(ˆ bo)the log-likelihood ratio statistic based on ˆ boof of testing H0:(b a) = (a0b0)against H1:(ba) =(b1a1), 0ˆ bo=ˆ blong bo shortΣ(ρ)−1(h1−h0)−1 2h 1Σ(ρ)−1h1+1 2h 0Σ(ρ)−1h0
Quantitative Economics 12 (2021) Linear regression with many controls 437 By Lemma 2, Ln(τn1) Ln(τn0)=Ln(τn1) Ln(0) Ln(0) Ln(τn0) p →1 and from ρn→ρ,x nxn(ˆ βlongnˆ βshortn)⇒ˆ boand hjn →hjfor j=01.Thus,under H0,n(Tn)⇒0(ˆ bo). This straightforwardly extends more generally to {n(Tn)}(ba)∈H⇒ {0(ˆ bo)}(ba)∈Hfor any finite H⊂R2. Thus, by Definition 9.1 in van der Vaart (1991), under the assumptions of the lemma, the sequence of experiments of observing Tnwith local parameter space (ba) ∈R2converges to the experiment of observing ˆ bo.Thefirst claim now follows from Theorem 15.1 in van der Vaart (1991). For the second claim, for given (ba), suppose Eθn[ϕn(¯κnYn)]→Eba[φ(ˆ bo)]along θn=θn1with τn=τn1=o(n−1/4).Letτn2=o(n−1/4)be another sequence, and denote θn2the corresponding sequence of θ. Suppose Eθn2[ϕn(¯κnYn)]does not converge to Eba[φ(ˆ bo)]. By Prohorov’s theorem (see, for instance, Theorem 2.4 in van der Vaart (1998)) and 0≤ϕn(¯κnYn)≤1, there exists a subsequence of nsuch that ϕn(¯κnYn)converges in distribution along that subsequence. Furthermore, by Lemma 2, the likelihood ratio statistic between the corresponding sequences θn1and θn2with identical values of (ba) converges in probability to one, and this convergence automatically holds jointly with ϕn(¯κnYn)along the subsequence. Thus, a trivial application of LeCam’s third lemma (see, for instance, Theorem 6.6 in van der Vaart (1998)) yields that under θn2,ϕn(¯κnYn)converges to the same weak limit as under θn1under the subsequence. But convergence in distribution implies convergence of expectations given that 0≤ϕn≤1, and the desired contradiction follows. A.4 Proof of Corollary 1 We use the same notation as the proof of Lemma 1, and momentarily drop the index nto ease notation. By assumption, the distribution of ψ(Y)only depends on ξthrough (βΔτ),soψ(Y)has the same distribution under ξand ξ0. By sufficiency, the conditional distribution of ψ(Y)given ˆ ζunder ξ0does not depend on ξ0, so by inverting the probability integral transform conditional on ˆ ζ,wecanwriteψ(Y)∼ψS(ˆ ζUS)for some function ψS:Rp+2→Rwith US∼[01]independent of ˆ ζ. Let ˆ Obe a random rotation matrix drawn from the Haar measure Hp−1, independent of (ˆ ζUS). Since by assumption, the distribution of ψ(Y)does not depend on ω,wehave ψS(ˆ ζUS)∼ψSˆ βlongˆ βshort(τˆ Oι+e)US ∼ψSˆ βlongˆ βshort(τι+e)ˆ OUS ∼ψSˆ βlongˆ βshortˆτιˆ OUS where the second equality follows from the spherical symmetry of the distribution of e. Since (ˆ βlongˆ βshortˆτιˆ O)is a one-to-one function of T,sowecanhencewriteψ(Y)∼ ˜ ψ(TUS)for some function ˜ ψ:R4→R. Now reintroducing nsubscripts, consider the sequence of experiments of observing (T nUS). Recalling that USis independent of Tn, these experiments converge to the
438 Li and Müller Quantitative Economics 12 (2021) limit experiment of observing (ˆ boUS)∈R3under the assumptions of the corollary, by the same arguments employed in the proof of Theorem 2. Thus, by the asymptotic representation theorem (Theorem 9.3 in van der Vaart (1998)), there exists a function ˜ ψo:R4→R∪{+∞}and uniform random variable Uindependent of (ˆ boUS)such that the limit distribution of ˜ ψn(TnUS)can be written as ˜ ψo(ˆ boUSU), for all (a b). Since the distribution of ˜ ψo(ˆ boUSU)conditional on ˆ bodoes not depend (a b),thereexists a function ψo:R3→R∪{+∞}so that ˜ ψo(ˆ boUSU)∼ψo(ˆ boU)for all (a b),aswasto be shown. A.5 Proof of Lemma 3 We will make use of the following lemma. Lemma 5. Let Hnand ˆ Hnbe the Choleski decompositions of Ωn=HnH nand ˆ Ωn= ˆ Hnˆ H n,respectively,and let wn=(ˆ βlongn −βnˆ βshortn −βn−Δn).Then under (15)and Ω−1 nˆ Ωn p →I2: (a) ˆ H−1 nHn p →I2 (b) ˆ H−1 nwn⇒N(0I2). Proof. (a) Note that ˆ H−1 nHnH nˆ H−1 n, by similarity, has the same eigenvalues as ˆ H−1 nˆ H−1 nHnH n=ˆ Ω−1 nΩn p →I2, so they both converge to one in probability. But ˆ H−1 nHnH nˆ H−1 nis symmetric, and all symmetric matrices with eigenvalues converging to one converge to the identity matrix. Thus ˆ H−1 nHnH nˆ H−1 n p →I2, and since ˆ H−1 nHnis lower triangular, this further implies ˆ H−1 nHn p →I2. (b) Note that Hnis related to Ω1/2 nvia Hn=Ω1/2 nOnfor some rotation matrix On. Thus, also H−1 nwn=O nΩ−1/2wn⇒N(0I2). (Suppose otherwise. Then, by the Cramér– Wold device, for some 2×1vector υand c∈R,liminfn→∞ |P(υO nΩ−1/2wn>c)− P(N(0υυ)>c)|>0. Pick a subsequence along which the liminf is attained, and On converges. Then we have a contradiction, because the continuous mapping theorem implies the convergence P(υO nΩ−1/2wn>c)−P(N(0υυ)>c)→0along that subsequence.) Invoking Lemma 5(a), also ˆ H−1 nwn=(ˆ H−1 nHn)H−1 nwn⇒N(0I2)by the continuous mapping theorem. (a) Write L(Z1Z2+gχ1χ2)for the expression in equation (18). Reparametrize χand gin (18)intermsof(rφu)∈[0∞)×[0π/2)×[01]via χ1=rcos(φ),χ2= rsin(φ) and u=g/χ2(with u=0if χ2=0). By a direct calculation, the limit of L(Z1Z2+ ur sin(φ)r cos(φ) r sin(φ)) as r→∞exists for almost all Z1,Z2and all (u φ) ∈[01]× [0π/2)and is equal to L∞(Z1Z2uφ)=⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩Z1−(1+u)tan φ2if Z1>(1+u)tanφ Z1+(1−u)tan φ2if Z1<−(1−u)tanφ 0otherwise.
Quantitative Economics 12 (2021) Linear regression with many controls 439 Correspondingly, limr→∞cv((r cos(φ)r sin(φ)) =cv∞(φ) exists, too, and satisfies sup0≤u≤1P(L∞(Z1Z2uφ) ≥cv∞(φ)) ≤α. (In general, this inequality is not sharp, since the definition of cv∞(φ) also requires P(L(Z1Z2+ur sin(φ)r cos(φ)r sin(φ)) ≥ cv∞(φ)) ≤αfor all finite r.) If r→∞and φ→π/2, then the limit still exists and is equal to L∞(Z1Z2uπ/2)=0. Suppose the assertion of the lemma is false. Then there exists a subsequence of n such that along that subsequence, lim n→∞EθnˆϕLRn(¯κnYn)=lim sup n→∞ EθnˆϕLRn(¯κnYn)>α Pick a sub-subsequence, such that with (rnφnun)the parameters computed from Ω=Ωnand gn=Ωn11Δn/Ωn11Ωn22 −Ω2 n12,(rnφnun)converge along that subsubsequencetosomevalue(r0φ0u0)∈[0∞]×[0π/2]×[01]. Correspondingly, let cv0be the limit of cv((rncos(φn) rnsin(φn)) along that sub-subsequence (which exists by the above observations also when rn→∞,evenwhenφ0=π/2). By Lemma 5(a), (ˆ Zn1ˆ Zn2)=ˆ H−1 nwn⇒(Z1Z2)∼N(0I2)by the continuous mapping theorem. Since ˆ H−1 n=⎛ ⎜ ⎜ ⎜ ⎜ ⎝ 1/ˆ Ωn11 0 −ˆ Ωn12 ˆ Ωn11ˆ Ωn11 ˆ Ωn22 −ˆ Ω2 n12 ˆ Ωn11 ˆ Ωn11 ˆ Ωn22 −ˆ Ω2 n12 ⎞ ⎟ ⎟ ⎟ ⎟ ⎠ the definitions of ˆχ1n =ˆχ1and ˆχ2n =ˆχ2yield LRn(¯κn)=min |˜ g|≤ˆχ2n ˆ Zn1 ˆ Zn2+ˆ gn−˜ g 2 −min ˜ h|˜ g|≤ˆχ2n ˆ Zn1−˜ h ˆ Zn2+ˆ gn−˜ g−ˆχ1n ˜ h 2 where ˆ gn=ˆ Ωn11Δn/ˆ Ωn11 ˆ Ωn22 −ˆ Ω2 n12.Let(ˆ rnˆ φnˆ un)∈[0∞)×[0π/2)×[01] be such that ˆχ1n =ˆ rncos(ˆ φn),ˆχ2n =ˆ rnsin(ˆ φn)and ˆ un=ˆ gn/ˆχ2n (with ˆ un=0if ˆχ2n =0). Write ˆ H−1 nHn p →I2from Lemma 5(b) element-by-element to conclude that ˆ Ω11n/Ω11n p →1,(ˆ Ω11n ˆ Ω22n −ˆ Ω2 12n)/(Ω11nΩ22n −Ω2 12n)p →1and (ˆ Ω12n −Ω12n)/ (Ω11nΩ22n −Ω2 12n)p →0. Therefore, also (ˆ rn−rn)/max(rn1)p →0,ˆ un−un p →0and ˆ φn−φn p →0. Thus, along the sub-subsequence defined above, by the continuous mapping theorem, LRn(¯κn)⇒L0=!LZ1u0r0cos(φ0)+Z2r0cos(φ)r0sin(φ0)if r0<∞ L∞(Z1Z2u0φ0)otherwise and Eθn[ˆϕLRn(¯κnYn)]→P(L0>cv0). But by the definition of cv0,P(L0>cv0)≤α, yielding the desired contradiction.
440 Li and Müller Quantitative Economics 12 (2021) References Anderson, T. W. and H. Rubin (1949), “Estimators of the parameters of a single equation in a complete set of stochastic equations.” The Annals of Mathematical Statistics, 21, 570–582. [422] Armstrong, T. B. and M. Kolesár (2016), “Optimal inference in a class of regression models.” arXiv:1511.06028v2.[408,415] Armstrong, T. B. and M. Kolesár (2018), “Optimal inference in a class of regression models.” Econometrica, 86. [408] Belloni, A., V. Chernozhukov, and C. Hansen (2014), “Inference on treatment effects after selection among high-dimensional controls.” The Review of Economic Studies,81(2), 608–650. [406,409,423,429] Cattaneo, M. D., M. Jansson, and W. K. Newey (2018a), “Alternative asymptotics and the partially linear model with many regressors.” Econometric Theory, 34 (2), 277–301. [421] Cattaneo, M. D., M. Jansson, and W. K. Newey (2018b), “Inference in linear regression models with many covariates and heteroscedasticity.” Journal of the American Statistical Association, 113 (523), 1350–1361. [421,424,426,427,432] Conley, T. G., C. B. Hansen, and P. E. Rossi (2012), “Plausibly exogenous.” Review of Economics and Statistics, 94 (1), 260–272. [408] Cox, D. R. and D. V. Hinkley (1974), Theoretical Statistics. Chapman&Hall/CRC, New York, NY. [422] Donoho, D. L. (1994), “Statistical estimation and optimal recovery.” The Annals of Statistics, 22 (1), 238–270. [416] Donohue, J. J. and S. D. Levitt (2001), “The impact of legalized abortion on crime.” Quarterly Journal of Economics, CXVI, 379–420. [426,429,430,433] Elliott, G., U. K. Müller, and M. W. Watson (2015), “Nearly optimal tests when a nuisance parameter is present under the null hypothesis.” Econometrica, 83, 771–811. [413,414] Fan, J. and R. Li (2001), “Variable selection via nonconcave penalized likelihood and its oracle properties.” Journal of the American Statistical Association, 96 (456), 1348–1360. [406] Foote, C. L. and C. F. Goetz (2008), “The impact of legalized abortion on crime: Comment.” The Quarterly Journal of Economics, 123, 407–423. [429] Hoerl, A. E. and R. W. Kennard (1970), “Ridge regression: Biased estimation for nonorthogonal problems.” Technometrics, 12 (1), 55–67. [408] Imbens, G. W., D. B. Rubin, and B. I. Sacerdote (2001), “Estimating the effect of unearned income on labor earnings, savings, and consumption: Evidence from a survey of lottery players.” American Economic Review, 91, 778–794. [426,431,432]
Quantitative Economics 12 (2021) Linear regression with many controls 441 Imbens, G. W. and D. B. Rubin (2015), Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press, New York, NY. [431] Imbens, G. W. and J. M. Wooldridge (2009), “Recent developments in the econometrics of program evaluation.” Journal of Economic Literature, 47, 5–86. [430,432,433] Joyce, T. (2004), “Did legalized abortion lower crime?” Journal of Human Resources,39 (1), 1–28. [430] Joyce, T. (2009), “A simple test of abortion and crime.” The Review of Economics and Statistics, 91 (1), 112–123. [430] Leeb, H. and B. M. Pötscher (2005), “Model selection and inference: Facts and fiction.” Econometric Theory, 21, 21–59. [406] Leeb, H. and B. M. Pötscher (2008a), “Can one estimate the unconditional distribution of post-model-selection estimators?” Econometric Theory, 24 (2), 338–376. [406] Leeb, H. and B. M. Pötscher (2008b), “Recent developments in model selection and related areas.” Econometric Theory, 24 (2), 319–322. [406] Li, C. (M.) and U. K. Müller (2021), “Supplement to ‘Linear regression with many controls of limited explanatory power’.” Quantitative Economics Supplemental Material, 12, https://doi.org/10.3982/QE1577.[415] Macchiavello, R. and A. Morjaria (2015), “The value of relationships: Evidence from a supply shock to Kenyan rose exports.” American Economic Review, 105, 2911–2945. [426, 427,428,429,433] Mardia, K. V. and P. E. Jupp (2000), Directional Statistics. Wiley Series in Probability and Statistics. John Wiley & Sons, Chichester. [436] Müller, U. K. and A. Norets (2016), “Credibility of confidence sets in nonstandard econometric problems.” Econometrica, 84, 2183–2213. [414] Müller, U. K. and Y. Wang (2019), “Nearly weighted risk minimal unbiased estimation.” Journal of Econometrics, 209, 18–34. [414] Obenchain, R. L. (1977), “Classical F-tests and confidence regions for ridge regression.” Technometrics, 19, 429–439. [408] Pratt, J. W. (1961), “Length of confidence intervals.” Journal of the American Statistical Association, 56, 549–567. [409] Tibshirani, R. (1996), “Regression shrinkage and selection via the lasso.” Journal of the Royal Statistical Society. Series B (Methodological), 58, 267–288. [406] van de Geer, S., P. Bühlmann, Y. Ritov, and R. Dezeure (2014), “On asymptotically optimal confidence regions and tests for high-dimensional models.” Annals of Statistics, 42, 1166–1202. [406] van der Vaart, A. W. (1991), “An asymptotic representation theorem.” International Statistical Review, 259, 97–121. [437]
442 Li and Müller Quantitative Economics 12 (2021) van der Vaart, A. W. (1998), Asymptotic Statistics. Cambridge University Press, Cambridge, UK. [409,436,437,438] White, H. (1980), “A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity.” Econometrica, 48, 817–830. [421] Wüthrich, K. and Y. Zhu (2020), “Omitted variable bias of lasso-based inference methods: A finite sample analysis.” arXiv:1903.08704.[407] Zhang, C.-H. and S. S. Zhang (2014), “Confidence intervals for low dimensional parameters in high dimensional linear models.” JournaloftheRoyalStatisticalSociety:SeriesB (Statistical Methodology), 76, 217–242. [406] Co-editor Andres Santos handled this manuscript. Manuscript received 15 March, 2020; final version accepted 19 November, 2020; available online 21 December, 2020.