scieee AI-readable full text Open interactive document viewer

Optimal HAR inference

Dou, Liyu

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Dou, Liyu Article Optimal HAR inference Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Dou, Liyu (2024) : Optimal HAR inference, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 15, Iss. 4, pp. 1107-1149, https://doi.org/10.3982/QE1762 This Version is available at: https://hdl.handle.net/10419/320317 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Quantitative Economics 15 (2024), 1107–1149 1759-7331/20241107 Optimal HAR inference Liyu Dou Singapore Management University and The Chinese University of Hong Kong, Shenzhen This paper considers the problem of deriving heteroskedasticity and autocorrelation robust (HAR) inference about a scalar parameter of interest. The main assumption is that there is a known upper bound on the degree of persistence in data. I derive finite-sample optimal tests in the Gaussian location model and show that the robustness-efficiency tradeoffs embedded in the optimal tests are essentially determined by the maximal persistence. I find that with an appropriate adjustment to the critical value, it is nearly optimal to use the so-called equalweighted cosine (EWC) test, where the long-run variance is estimated by projections onto qtype II cosines. The practical implications are an explicit link between the choice of qand assumptions on the underlying persistence, as well as a corresponding adjustment to the usual Student-tcritical value. I illustrate the results in two empirical examples. Keywords. Heteroskedasticity and autocorrelation robust inference, long-run variance. JEL classification. C12, C18, C22. 1. Introduction This paper considers the problem of deriving appropriate corrections to standard errors when conducting inference with autocorrelated data. The resulting heteroskedasticity and autocorrelation robust (HAR) inference has applications in OLS and GMM settings.1Computing HAR standard errors involves estimating the “long-run variance” (LRV) in econometric jargon. Classical references on HAR inference in econometrics include Newey and West (1987)andAndrews (1991), among many others. The Newey– West/Andrews approach is to use t-andF-tests based on consistent LRV estimators and to employ the critical values derived from the normal and chi-squared distributions. Liyu Dou: [email protected] I am deeply indebted to Ulrich Müller for posing the question and for his continuous help, support, and encouragement. I would like to thank two anonymous referees for constructive comments and suggestions, which have substantially improved the paper. I also thank Xu Cheng, Paul Ho, Bo Honoré, Michal Kolesár, Jia Li, Mikkel Plagborg-Møller, Mikkel Sølvsten, James Stock, Mark Watson, Ke-Li Xu, Jun Yu, and numerous participants in seminars for helpful comments and suggestions. I gratefully acknowledge financial support from the National Natural Science Foundation of China through Grant 72103176. 1For instance, OLS/GMM with HAR inference has been used in many econometric applications, such as testing long-horizon return predictability in finance (see, e.g., Koijen and Van Nieuwerburgh (2011) and Rapach and Zhou (2013)) and estimating impulse response functions by local projections in macroeconomics (see, e.g., Jordà (2005)). ©2024 The Author. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1762 1108 Liyu Dou Quantitative Economics 15 (2024) The resulting HAR standard errors are asymptotically justified in a large variety of circumstances. Small sample simulations,2however, show that the Newey–West/Andrews approach can lead to false rejections of the null far too often. A large subsequent literature (surveyed in Müller (2014)) employs alternative asymptotics that is often more accurate in finite samples and thus demonstrates better performance for controlling the null rejection rate. To implement these procedures in practice, however, the user must choose a tuning parameter. One example is the choice of bin the fixed-bscheme,3in which a fixed-bfraction of the sample size is used as the bandwidth in kernel LRV estimators. Another example is the choice of qin orthonormal series HAR tests,4in which the LRV is estimated by projections onto qmean-zero low-frequency orthonormal functions. The choice of the tuning parameter embeds a tradeoff between bias and variability of the LRV estimator. It subsequently leads to a size-power tradeoff in the resulting HAR inference. Previous studies address this tradeoff by restricting attention to HAR tests that are based on kernel and orthonormal series LRV estimators. They derive the optimal tuning parameter based on second-order asymptotics and under criteria that average functions of type I and type II errors with different weights.5It is not clear, however, whether the resulting HAR tests would remain optimal in finite samples if those restrictions were not imposed. Moreover, as demonstrated later, the choice of LRV estimator or HAR test is empirically relevant. Therefore, it would be useful to have guidelines for practitioners to implement HAR inference with certain senses of optimality. The purpose of this paper is to provide formal finite-sample efficiency results of HAR inference about a scalar parameter of interest, without restricting the class of tests and with commonly used notions of optimality in hypothesis testing. Specifically, I derive optimal (weighted average power maximizing scale invariant) HAR tests in the Gaussian location model, under nonparametric assumptions on the underlying spectral density. In addition, I find that with an appropriate adjustment to the critical value, it is nearly optimal to use a type of t-test with the LRV estimated by equal-weighted projections onto qtype II cosines, which is known as the equal-weighted cosine (EWC) test in the literature (cf. Müller (2004,2007), Lazarus, Lewis, Stock, and Watson (2018)). The main assumption in this paper is that there exists an upper bound on the degree of persistence in data. In time-series terminology and from a spectral perspective, this amounts to specifying a worst-case steepest, or “uniformly maximal” spectral density function fin class F, which is the collection of all plausible spectra and is of a nonparametric nature, as opposed to possibly strong parametric classes.6In theory, such prim2See, e.g., den Haan and Levin (1997,1998) for early Monte Carlo evidence of the large size distortions of HAR tests computed using the Newey–West/Andrews approach. 3See pioneering papers by Kiefer, Vogelsang, and Bunzel (2000) and Kiefer and Vogelsang (2002,2005). Also, see Jansson (2004), Müller (2004,2007), Phillips (2005), Phillips, Sun, and Jin (2006,2007), Sun, Phillips, and Jin (2008), Atchadé and Cattaneo (2011), Gonçalves and Vogelsang (2011), Sun and Kaplan (2012), Sun (2014a); and Sun (2014b), among many others. 4See, e.g., Müller (2004,2007), Phillips (2005), Ibragimov and Müller (2010); and Sun (2013), among many others. 5See, for example, Sun, Phillips, and Jin (2008) and Lazarus, Lewis, and Stock (2021). 6For parametric examples, Robinson (2005) assumes that the underlying persistence is of the “fractional” type and derives consistent LRV estimators under that class; Müller (2014) assumes that the underlying Quantitative Economics 15 (2024) Optimal HAR inference 1109 itives may contain smoothness restrictions (e.g., bounds on derivatives) and/or shape restrictions (e.g., monotonicity). But, for convenience in actual implementations, I suggest the practitioners follow the convention of measuring persistence by fitting a simple AR(1) model to data to determine a reasonable f. As it turns out, the finite-sample efficiency bounds, the best choice of q, and the critical value adjustment in the nearly optimal EWC test are essentially governed by the maximal persistence. The practical implication is that once the maximal persistence in data is appropriately determined, the EWC test with adjusted critical value can be used without much loss of efficiency. The resulting implementation is straightforward and only involves estimating an AR(1) model and making a simple adjustment to the Studenttcritical value for the EWC test. Furthermore, this procedure can be easily adapted to regression models. I discuss these practical matters in detail below in Section 5. In addition, I illustrate the implementation in two empirical examples in Section 6concerning confidence interval construction and hypothesis testing with autocorrelated data. This paper makes three main contributions. First, I establish a finite-sample theory of optimal HAR inference in the Gaussian location model under a simplifying approximation. To do so, I follow Müller (2014) and recast HAR inference as a problem of inference about the covariance matrix of a Gaussian vector. The spectrum, as an infinite-dimensional nuisance parameter, complicates the solution of the problem. To make progress, I use insights from the so-called least favorable approach and identify the “least favorable distribution” over the class F. The resulting optimal test embeds robustness-efficiency tradeoffs in hypothesis testing. This optimal tradeoff is a function of the underlying primitive F, namely the serial correlations one is willing to correct for under the null, and the alternative dependency that one desires the test to orient power toward. Second, I find that nearly optimal inference can be obtained by using the EWC test, but only after an adjustment to the Student-tcritical value. The practical implications are an explicit link between the choice of qand assumptions on the underlying spectrum, as well as a corresponding adjustment to the Student-tcritical value. In detail, consider a second-order stationary scalar time series yt. The spectral density of yt scaled by 2πis given by the function f:[−π,π]→[0, ∞).TotestH0:E[yt]=0 against H1:E[yt]=0, the EWC test uses a t-statistic tq EWC =Y0     q  j=1 Y2 j/q ,(1) where Y0is the sample mean of ytand Yj,j=1, 2, ,qare qweighted averages of ytas Yj=T−1√2T t=1cos(πj(t−1/2)/T)yt. These weighted averages can be approximately thought of as independently normally distributed, each with variance T−1f(πj/T ).As mentioned earlier, the choice of qembeds a bias and variance tradeoff of the LRV estimator q j=1Y2 j/q. The conventional wisdom is to choose qsufficiently small such that long-run property can be approximated by a stationary Gaussian AR(1) model, with coefficient arbitrarily close to one and derives uniformly valid inference methods that maximize weighted average power. 1110 Liyu Dou Quantitative Economics 15 (2024) Figure 1. Power function plot of a weighted average power (WAP) bound induced test, optimal EWC test, and size-adjusted EWC test using q=3. Notes: Under the alternative, the mean of ytis δT−1/2and ytfollows a Gaussian white noise. Under the null, the “uniformly maximal” function of Fcorresponds to an AR(1) with coefficient 0.8. Sample size T=100. {Yj}q j=1can be treated as i.i.d. normals. By doing so, one avoid possibly large bias in estimating the LRV, and the resulting EWC test has less size distortions when the Student-t critical value is employed. In contrast, the new EWC test suggests using a larger qand an appropriately enlarged critical value for more powerful inference. Both the choice of q and the critical value adjustment depend on the class F. Figure 1illustrates this second contribution in testing E[yt]=0, f∈Fagainst the local alternative E[yt]=δT−1/2for T=100, where ytfollows a Gaussian white noise and the “uniformly maximal” function of Fcorresponds to an AR(1) model with coefficient 0.8. In this context, to avoid size distortions larger than 0.01, one needs to choose q=3 when the Student-tcritical value is employed. The new EWC test, however, has q=6and inflates the Student-tcritical value by a factor of 1.13. Moreover, it is nearly as powerful as a weighted average power bound induced test. It has a 28.9% efficiency gain over the size-adjusted EWC test using q=3, in order to achieve the same power of 0.5.7 Third, I propose a simple adjustment to the critical value of the EWC test. The adjusted critical value is computed easily, by inverting a one-dimensional numerical integral. For practical convenience, I offer a rule of thumb to adjust the Student-tcritical value of the EWC test in Table 2, as follows. Under a series of classes Fwhere the largest persistence is parameterized as an AR(1) with coefficient ρ=1−c/T,Table1lists the 7By efficiency gain, I mean the increase of δ2in percent for the size-adjusted EWC test using q=3in order to achieve the same power of the new EWC test. I note that one cannot directly appeal to Pitman efficiency measure (the increase of the number of observations required to achieve the same power) in the context of Figure 1, since the sample size Tis fixed at 100. A different calculation, however, shows that for T=77 the size-adjusted EWC test using q=6 has power of around 0.5, under the same δsuch that the EWC test using q=3 yields power of 0.5 for T=100. Quantitative Economics 15 (2024) Optimal HAR inference 1111 Table 1. Optimal qand adjustment factor of the Student-tcritical value of level αEWC. c5 10203040 50 77 ρ0.95 0.9 0.8 0.7 0.6 0.5 0.23 α=0.05 (3, 1.55)( 4, 1.26)( 6, 1.13)( 7, 1.07)( 9, 1.05)( 12, 1.04)( 19, 1.02) Note: Based on a series of classes F, in which the “uniformly maximal” function corresponds an AR(1) with coefficient ρ=1−c/T. Sample size Tis 100, but the resulting optimal choice of q(almost) remains unchanged for fixed c≤40 and for T=200, 500, 1000. optimal choice of qand the adjustment factor of the Student-tcritical value for selected c(and ρfor fixed T). It turns out that the resulting optimal choice of q(almost) remains unchanged as Tvaries, for fixed c≤40. More interestingly, for fixed qand T,theadjustment factor does not change substantially under other types of F.Table2collects the adjustment factors in (augmented) Table 1for selected q. In the event that the same qis optimally chosen under different c, the largest adjustment factor (corresponding to the largest c) is suggested in the rule of thumb. In case researchers pick a value of q by some other means, I suggest adjusting the corresponding Student-tcritical value directly according to Table 2.Otherwise,Section5.1 provides guidance on determining a reasonable f, the subsequent critical value adjustment, and the selection of q. This paper relates to a large literature. First, unlike the majority of the HAR literature, I consider optimal HAR inferences without restricting the class of tests. Second, the majority of the literature addresses the sampling variability of LRV estimators via the so-called fixed-basymptotics, and further accounts for bias by higher-order adjustment to the fixed-bcritical value.8In contrast, I concurrently tackle bias and variance in estimating the LRV by a first-order adjustment in the spirit of employing fixed smoothing asymptotics under strong persistence in Sun (2014a). Even so, the resulting adjusted critical value is easily computed without simulations under a simplifying structure. Third, this paper contributes to the uniform size control literature developed by Müller (2014), Preinerstorfer and Pötscher (2016), Pötscher and Preinerstorfer (2018,2019), and Müller and Watson (2022). I analytically derive powerful tests that uniformly control size over arguably large classes of models, while Müller (2014) numerically determines powerful tests under a possibly restricted parametric class of models. Preinerstorfer and Pötscher (2016)andPötscher and Preinerstorfer (2018,2019) focus on size distortions and power Table 2. Rule of thumb for adjustment factor of the Student-tcritical value of level αEWC. q346891011121620 α=0.05 1.55 1.37 1.17 1.15 1.09 1.09 1.07 1.05 1.03 1.03 Note:Eachqis justified as the optimal choice of level αEWC test, under some class Fand for sample size T. An example of the corresponding class Fis the one in which the “uniformly maximal” function corresponds to an AR(1) model with coefficient ρ=1−c/T as in Table 1. Only the largest adjustment factor is displayed should the same qemerge as the optimal choice under different c. 8See, for example, Velasco and Robinson (2001), Sun, Phillips, and Jin (2008), Sun (2011,2013,2014c); and Lazarus, Lewis, and Stock (2021). 1112 Liyu Dou Quantitative Economics 15 (2024) deficiencies of given HAR tests, allowing for general classes of models. Müller and Watson (2022) derive finite-sample size control results in general spatial settings without the simplifying structure considered in this paper but offer limited analytical results on the efficiency side. The suggestion of using a larger qand enlarged critical values for the EWC test mirrors recent recommendations for nonparametric inference, such as those of Armstrong and Kolesár (2018,2020). In different contexts, Armstrong and Kolesár and I both stress the advantage of accepting bias in estimating a nonparametric function and then using a suitably adjusted critical value to account for the maximum bias. Our frameworks are, however, different. I consider a Gaussian experiment in which the heteroskedasticity is governed by an unknown nonparametric function and is thus more in the spirit of Lehmann and Stein (1948), while the main focus in Armstrong and Kolesár (2018)isan unknown regression function in the mean of a homoskedastic Gaussian experiment. The remainder of the paper is organized as follows. Section 2sets up the model and discusses preliminaries. Section 3derives efficiency results under an essential simplification, which are theoretically investigated in more general settings in Section 4.Section 5converts the theoretical insights into practical guidance and discusses the implementation in regression models. Section 6provides empirical illustrations with a selfcontained guide to implementation. Interested practitioners can skip the theoretical discussions and read Section 6directly. Proofs and computational details are provided in the Appendices. 2. Model and preliminaries The paper concerns inference about μin the location model, yt=μ+ut,t=1, 2, ,T,(2) where μis the population mean of ytand utis a mean-zero stationary Gaussian process with absolutely summable autocovariances γ(j)=E[utuy−j]. The spectrum of yt scaled by 2πis given by the even function f:[−π,π]→ [0, ∞)defined via f(λ)= ∞ j=−∞cos(jλ)γ(j).Withy=(y1,y2,,yT)and e=(1, 1, ,1), y∼Nμe,(f),(3) where (f)has elements (f)j,k=(2π)−1π −πf(λ)e−i(j−k)λdλwith i =√−1. The location model (2) is often considered a stylized setting to provide theoretical insights into HAR inference.9As simple as it is, this model is empirically relevant in a number of situations. For example, the statistical study of unconditional equal predictive ability (UEPA) concerning competing forecasts in financial and macroeconomic contexts amounts to testing an unconditional mean-zero condition in (2)withytbeing the produced loss differential series. 9For HAR studies based on the Gaussian location model see, for example, Velasco and Robinson (2001), Jansson (2004), Sun, Phillips, and Jin (2008), Sun (2011,2013,2014c), Müller (2014), Lazarus, Lewis, and Stock (2021), among many others. Quantitative Economics 15 (2024) Optimal HAR inference 1113 Throughout the paper, I mainly focus on presenting and analyzing new efficiency results for HAR inference in the univariate Gaussian location model (3), and discuss the implications for conducting inference about a nonconstant regressor in regression settings. The ideas and methods explored in the simplest model (3) can often be used as a foundation for studying HAR inference in multivariate location models and general GMM settings, potentially involving additional complications.10 Formal generalizations of this paper’s efficiency results along those lines are, however, not straightforward and beyond the scope of the paper. The HAR inference problem in (3) concerns testing H0:μ=0(otherwise,subtract the hypothesized mean from yt) against H1:μ= 0 based on the observation y.The derivation of powerful tests in this problem is complicated by the fact that the alternative is composite (μis not specified under H1) and the presence of the infinite-dimensional nuisance parameter f. I follow standard approaches to deal with μand mainly focus on tackling the nuisance parameter fin this paper. It is useful to take a spectral transformation of the model (3). In particular, as introduced in the Introduction, consider the one-to-one transformation from {yt}T t=1into the sample mean Y0=T−1T t=1ytand the T−1 weighted averages: Yj=T−1√2 T  t=1 cosπj(t−1/2)/Tyt,j=1, 2, ,T−1. (4) Define as the T×Tmatrix with first column equal to T−1e,and(j+1)th column with elements T−1√2cos(πj(t−1/2)/T),t=1, ,T,andι1as the first column of IT.Then Y=(Y0,Y1,,YT−1)=y∼Nμι1,0(f),(5) where 0(f)=(f). The HAR testing problem becomes H0:μ=0 against H1:μ=0 based on the observation Y. A common device for dealing with the composite alternative in the nature of μis to search for tests that maximize weighted average power over μ. For analytical tractability, I follow Müller (2014) to consider a Gaussian weighting function for μwith mean zero and variance η2. The scalar η2governs whether closer or distant alternatives are emphasized by the weighting function. For a given f, and thus known 0(f)1,1, the choice η2=(κ−1)0(f)1,1 effectively changes the testing problem to H 0:Y∼N(0, 0(f)) against H 1:Y∼N(0, 1(f)),where1(f)=0(f)+(κ−1)ι1ι 10(f)1,1.Thistransforms the problem into one of inference about covariance matrices. The hyperparameter κspecifies a weighted average power criterion. As argued by King (1987), it makes sense to choose κin a way such that good tests have approximately 50% weighted average power. The choice of κ=11 would induce the resulting best 5% level (infeasible) test (reject if Y2 0>3.840(f)1,1) to have power of approximately P(χ2 1>3.84/11)≈56%. Ithususeκ=11 throughout the implementations. 10For formal discussions of HAR inference in general GMM settings see, for example, Sun (2014b), Hwang and Sun (2017,2018). 1114 Liyu Dou Quantitative Economics 15 (2024) In most applications, it is reasonable to impose that if the null hypothesis is rejected for some observation Y, then it should also be rejected for the observation aY ,forany a>0. By standard testing theory,11 any test satisfying this scale invariance property can be written as a function of Ys=Y/√YY. The density of Ysunder H i,i=0, 1 is equal to (see Kariya (1980)andKing (1980)) hi,fys=Ci(f)−1/2ysi(f)−1ys−T/2(6) for some constant C. By restricting to scale invariant tests, the HAR testing problem has been further transformed into H 0:“Yshas density h0,f” against H 1:“Yshas density h1,f.” The problem remains nonstandard due to the presence of nuisance parameter f.Tomake progress, I first consider directing power at a flat spectrum f1=1(whitenoise)when demonstrating how to derive efficiency bounds in Section 3. In this case, the alternative H 1then becomes a single hypothesis H 1,f1:“Yshas density h1,f1,” where 1(f1)= κT−1diag(1, κ−1,,κ−1). Moreover, under the null, I assume fbelongs to an explicit function class Fand seek scale invariant tests that uniformly control size over F.InSection 4.1, I robustify insights from the white noise case to ones where power is directed at a nonflat spectrum ˜ f1, or when minimax bounds are concerned if fbelongs to a class G⊂Funder H1. The testing problem is now reduced to distinguishing the composite null H 0 from a single alternative. A well-known general solution to this type of problem proceedsasfollows(cf.Lehmann and Romano (2005)). Suppose is some probability distribution over F, and the composite null H 0is replaced by the single hypothesis H 0,:“Yshas density h0,fd(f).” Any ad hoc test ϕah that is known to be of level αunder H 0also controls size under H 0,, because ϕah(ys)h0,f(ys)d(f)dys= ϕah(ys)h0,f(ys)dysd(f)≤α. By Neyman–Pearson lemma, the likelihood ratio test of H 0,against H 1,f1, denoted by ϕ,f1, yields a bound on the power of ϕah.Furthermore, if ϕ,f1also controls size under H 0, then it must be the best test of H 0against H 1,f1and the resulting power bound is the lowest possible one. In the jargon of statistical testing, the distribution that yields the best test (should it exist) is called the “least favorable distribution,” and I denote it by ∗throughout the paper. Unfortunately, there is no systematic way of deriving such a distribution. I make progress along this line in the following sections. 3. Finite-sample efficiency results In this section, I impose a simplifying Whittle-type diagonal structure on the implied covariance matrices 0of the effective observation Y, specify a priori that Fpossesses a most persistent spectrum, and analytically derive the resulting least favorable distribution, and thus obtain the optimal test. More specifically, I make the following assumptions. 11See, for example, Chapter 6 in Lehmann and Romano (2005). Quantitative Economics 15 (2024) Optimal HAR inference 1121 If one desires to direct power at say ˜ f1instead of f1, then the critical value adjustments will not change for any given q, and the optimal selection of qnow makes (12) reach its largest value at {˜ ζj=(cva q)2κ−1˜ f1(πj/T )/˜ f1(0)wj}q j=1. In this sense, robustness and efficiency constraints are sequentially addressed. In contrast, the optimal test ϕ∗,˜ f1(see details in Corollary 4.4) and the resulting optimal weights {w∗ j}in the t-statistic form (9) vary simultaneously with ˜ f1, rendering it less operationally straightforward as compared to the weighted cosine tests. The above discussion, of course, applies to the class of EWC tests, for which wj=q−1 for any given q.Inthatcase,iffis further parameterized using the limiting local-to-zero spectra under local-to-unity asymptotics, that is, 1/(π2j2+c2)for some c>0, (cva q)2 mirrors the critical value of an Fstatistic (with p=1) from fixed-smoothing asymptotics under strong (local-to-unity) persistence in Sun (2014a). Moreover, (cva q)2is of a hump shape as a function of qfor any given and finite c(see Figure 3 in Sun (2014a)). The optimal choice of qthen amounts to exploiting that hump shape and picking the qsuch that the vector that stacks qreplicates of (cva q)2κ−1q−1and T−1−qzeros majorizes all other candidate vectors of the same form. In face of c=∞,cva qmonotonically decreases in q. As a result, one shall utilize all information in the data and optimally choose q= T−1. Intuitively, this qshall also balance the worst absolute bias (|f(0)−q−1q j=1f(πj/ T)|) and variability of the LRV estimator q−1q j=1Y2 jin a testing-optimal sense. Accordingly, the adjusted critical value cva qaccounts for this maximum bias in the spirit of Armstrong and Kolesár. However, since cva qis determined in a rather more complicated way than in their settings, I choose not to discuss further these bias-variance tradeoffs. Furthermore, I note that because the implied weights {w∗ j}q∗ j=1in the optimal test do not depend on the associated critical value cvq∗in a straightforward way, and thus may bring another layer of complications, neither will I formally compare q∗in the optimal test with the MSE or testing optimal qin the EWC class in this paper. Comment 5. One may wonder whether Theorem 3.3 is limited by only orienting power toward f1. As it turns out, the insights can be generalized to accommodate possibly more complex alternatives. I relegate the formal analysis along this line to Section 4.1. 3.2 The optimal EWC test By using higher-order expansions, Lazarus, Lewis, and Stock (2021) derive a size-power frontier for kernel and orthonormal series HAR tests under an asymptotic framework. The EWC test is shown to achieve that frontier in their context. It is, however, not clear how the EWC test performs in the current context. The efficiency bounds derived in the last section provide a natural benchmark to gauge the performance of an ad hoc test. In this section, I take up the EWC test as the ad hoc test and discuss its properties. I have two related goals. The first is to study the (weighted average) power properties of the EWC test relative to the optimal test in Theorem 3.3. As it turns out, the EWC test is close to optimal, under an appropriate choice of qand with the adjusted critical value, which is just discussed in Comment 4 above. I refer to this new EWC test as the optimal 1122 Liyu Dou Quantitative Economics 15 (2024) Table 4. Weighted average power (WAP) of the optimal test and the optimal EWC test. ρ0.50 0.60 0.70 0.80 0.90 0.95 0.98 qin optimal EWC 12 9 7 6431 critical value adj. factor 1.04 1.05 1.07 1.13 1.26 1.55 1.87 WAP of optimal EWC 0.507 0.492 0.472 0.438 0.351 0.233 0.088 WAP of optimal test 0.507 0.493 0.475 0.442 0.358 0.239 0.091 Note: The “uniformly maximal” function corresponds to an AR(1) with coefficient ρ.Nominallevelis5%. Sample size Tis 100. EWC hereafter. My second goal is to draw the following practical implications of the optimal EWC test via its comparison with the conventional EWC test: One should use the EWC test with a larger qand appropriately enlarged critical values for more powerful HAR inference. 3.2.1 PoweroftheoptimalEWCtest Consider the type of Fin Table 3, that is, the “uniformly maximal” function fof the class Fcorresponds to an AR(1) with coefficient ρ. As ρvaries, Tables 4displays the crucial ingredients (qand adjustment factor relative to Student-tcritical value) in the optimal EWC test, its resulting weighted average power and the efficiency bound induced by the optimal test. Two observations are immediate. First, the optimal EWC test is nearly as powerful as the optimal test. This observation remains when a different type of class Fis considered in Table 5. Second, the selections of qin the optimal EWC test and q∗in the optimal test are not necessarily equal. After all, as explained in Comments 3 and 4 above, they are determined in arguably different robustness-efficiency tradeoff mechanisms. It is worth noting that qin the optimal EWC test can be somewhat sensitive to numerical errors in evaluating (12)whenfis relatively flat. This, however, does not have any substantial consequence in theory. 3.2.2 Practical implications Recall that the conventional wisdom in implementing the EWC test is to use a sufficiently small qand to employ the Student-tcritical value. I find, however, that it is better to use a larger qand to employ an enlarged critical value. Take the example from Figure 2as an illustration: The “uniformly maximal” function fcorresponds to an AR(1) with coefficient 0.8 and the sample size is fixed to be 100. According to conventional wisdom, one needs to use q=3 in the usual EWC test to obtain size distortions less than 0.01. The optimal EWC test, however, selects a larger q=6 (highlighted in Table 4), and the corresponding Student-tcritical value must be inflated by a factor of 1.13 for exact size control. An apple-to-apple comparison then Table 5. Weighted average power (WAP) of the optimal test and the optimal EWC test. C10.0 5.6 3.2 1.8 1.0 0.6 0.2 0.1 WAP of optimal EWC 0.286 0.361 0.418 0.453 0.482 0.500 0.526 0.533 WAP of optimal test 0.288 0.364 0.419 0.457 0.484 0.501 0.527 0.536 Note: The “uniformly maximal” function of Fis f(φ)=exp(−Cφ).Nominallevelis5%.Tis 100. Quantitative Economics 15 (2024) Optimal HAR inference 1123 Figure 5. Power function plot of the test ϕ∗, the optimal and conventional EWC tests. Notes: Under the alternative, the mean of ytis δT−1/2(1−ρ1)−1and ytfollows a Gaussian AR(1) with coefficient ρ1. Under the null, the fof Fcorresponds to an AR(1) with coefficient 0.8. Sample size Tis 100. reveals that the size-adjusted weighted average power of the usual EWC test (0.39) has about 12% loss as compared to that of the optimal EWC test (0.438, as highlighted in Table 4). The superior power property of the optimal EWC test is further evident when local alternatives are considered. In particular, in the context of the above example, I consider μ=δT−1/2(1−ρ1)−1under the alternative. Panels (a) and (b) of Figure 5plot the power of the optimal test ϕ∗, the optimal EWC test, and the size-adjusted EWC test using q=3 for various δunder ρ1=0andρ1=0.8, respectively. As can be seen in panel (a), even though the optimal EWC test underrejects under the null, it is more powerful than the EWC test using q=3 in detecting local deviations from the null. Specifically, by using the optimal EWC test, a 28.9% efficiency improvement is obtained in order to achieve the same power of 0.5. In the case in which ρ1=0.8, the efficiency gain is larger (47.6%), since the optimal EWC test then exactly controls size by construction. Furthermore, given that the optimal EWC test is nearly as powerful as the overall optimal test ϕ∗ in terms of weighted average power under the white noise alternative, it is not surprising to see that the power functions of these two tests are almost identical. 4. Theoretical generalizations The finite-sample efficiency results in Section 3are derived under seemingly restrictive assumptions. In particular, when power is directed at the white noise alternative a prior, the optimal test ϕ∗possesses a precise sense of optimality, and the optimal EWC is numerically found to be nearly as powerful as ϕ∗. More importantly, the existing efficiency results are entirely based on the Whittle-type diagonal structure. It is natural to ask how limited these simplifying assumptions are in eliciting theoretical insights in HAR inference. Said differently, can the insights on efficiency in Section 3be generalized to more 1124 Liyu Dou Quantitative Economics 15 (2024) general settings? In this section, I take up these questions and discuss the theoretical generalizations. 4.1 Power directions and minimax efficiency results First of all, I devise optimal tests that direct power at a nonflat ˜ f1, and, more generally, a nonsingleton class G⊂Funder H1. For analytical tractability, I maintain Assumption 3.1 and also impose a Whittle-type structure on 1(f), which automatically holds under f1. Assumption 4.1. For all f∈G,1(f)=T−1diag(κf (0),f(π/T ),,f(π(T−1)/T). Adopting the conventional scale invariance and weighted average power maximizing criteria, I now seek powerful tests as functions of Ys=Y/√YYin the problem of Hd 0:Y∼N0, T−1diagf(0),f(π/T ),,fπ(T−1)/T,f∈F(13) against Hd 1,G:Y∼N0, T−1diagκf (0),f(π/T ),,fπ(T−1)/T,f∈G, under the following assumption that nests Assumption 3.2 as a special case (G={f1}). Assumption 4.2. (a) There exists a f∈Fsuch that f(0) f(φ)≤f(0) f(φ),for all φ∈[−π,π]and f∈F. (b) There exists a f∈Gsuch that f(0) f(φ)≥f(0) f(φ),for all φ∈[−π,π]and f∈G. (c) f(πj/T )/f(πj/T )≥f(π(j+1)/T)/f(π(j+1)/T),j=0, 1, ,T−2. (d) Fcontains all kinked functions defined by fθ(φ)=rθ(φ)f(φ),in which,for θ∈ [0, π],rθ(φ)=1if |φ|≤θand rθ(φ)∈[1, ∞)if |φ|>θ. (e) Gcontains all kinked functions defined by gθ(φ)=r θ(φ)f(φ),in which,for θ∈ [0, π],r θ(φ)=1if |φ|≤θand r θ(φ)∈[0, 1]if |φ|>θ. Intuitively, relative persistence between the null and alternative hypotheses matters in distinguishing them. In Section 3, because the alternative is fixed at a constant singleton, the “uniformly maximal” spectral density plays a vital role in deriving “least favorable” results. Now, under Assumption 4.2,f/f posits the largest possible relative persistence and is thus expected to be an important primitive in eliciting similar efficiency results. It turns out that this intuition is correct and I formalize it below as a minimax result. I follow Lehmann and Romano (2005) to define the necessary notation in the HAR context. Let 0and 1denote distributions of fover Fand G, respectively. Let ϕ0,1 be the most powerful level αweighted average power maximizing scale invariant test for testing H0,0against H1,1,inwhichH0,0and H1,1are simple hypotheses defined analogously to H 0,in Section 2,andletβ0,1be its weighted average power for a given κ. Suppose 0and 1are such that supf∈FE[ϕ0,1(y)] ≤αand inff∈GE[ϕ0,1(y)] = Quantitative Economics 15 (2024) Optimal HAR inference 1125 β0,1,thenϕ0,1maximizes inff∈GE[ϕ(y)] among all valid level αtests ϕof H0.The following theorem makes further optimal statements about the minimax weighted average power bound β0,1among all possible candidates of 0and 1. Theorem 4.3. Under Assumptions 3.1,4.1,and 4.2, (i) If κf(0)/f(0)≤f(π/T )/f(π/T ),then the smallest β0,1among all possible pairs of 0and 1is αand is attained by the trivial randomized test. (ii) If κf(0)/f(0)>f(f)(π/T )/f(π/T ),then the following test ϕ∗identifies the least favorable pair of distributions (∗ 0,∗ 1)in the sense that: β∗ 0,∗ 1≤β0,1 for all possible pairs of 0and 1,in which ϕ∗rejects for large values of Y2 0+f(0) q∗  j=1 Y2 j/f(πj/T ) Y2 0+κf(0) q∗  j=1 Y2 j/f(πj/T ) (14) for a unique 1≤q∗≤T−1, and with the critical value cvq∗such that the test is of level αunder f=fand attains its minimax power β∗ 0,∗ 1at f=f. The proof strategy of Theorem 4.3 is similar to that of Theorem 3.3.Iprovebothof them and the two immediate corollaries within a coherent framework in Appendix A. Corollary 4.4. Let G={˜ f1}for some ˜ f1that is not necessarily equal to the flat f1so that f=˜ f1in the sense of Assumption 4.2(b). Under Assumptions 3.1,4.1,4.2: (i) If κ˜ f1(0)/f(0)≤˜ f1(π/T )/f(π/T ),then the best weighted average power maximizing scale invariant test of H0:μ=0against H1:μ=0is the trivial randomized test. (ii) If κ˜ f1(0)/f(0)>˜ f1(π/T )/f(π/T ),the best level αweighted average power maximizing scale invariant test ϕ∗of H0:μ=0against H1:μ= 0rejects for large values of Y2 0+f(0) q∗  j=1 Y2 j/f(πj/T ) Y2 0+κ˜ f1(0) q∗  j=1 Y2 j/˜ f1(πj/T ) (15) for a unique 1≤q∗≤T−1, and with the critical value cvq∗such that the test is of level αunder f=f. 1126 Liyu Dou Quantitative Economics 15 (2024) Corollary 4.5. Let Fand Gbe sets of fsatisfying Assumption 4.2 and with f=f1.Under Assumptions 3.1 and 4.1,the weighted average power of the optimal test in Theorem 3.3 gives the minimax power bound in the sense of Theorem 4.3. 4.2 Near optimality of EWC tests It is found in Section 3.2 that the so-called optimal EWC test nearly archives the efficiency bound when power is directed at the white noise (f1). It is tempting to ask whether such findings remain in more realistic situations, especially given that efficiency bounds are derived in Section 4.1 when power is oriented toward nonflat alternatives. I first maintain the Whittle diagonal structure with AR(1) fas before (coefficient ρ0) but consider power directions at possibly nonmonotonic ˜ f1of AR(2) processes with roots ρ1and ρ2. In panel (a) of Figure 6, I do not endogenize the test statistics for each (ρ0,ρ1,ρ2)combination. More precisely, the boxplots are for the difference in weighted average powers of the ϕ∗and the new EWC tests that are initially designed to be (nearly) optimal at f1(as in Section 3) but are now at ˜ f1for various (ρ1,ρ2)’s. In contrast, panel (b) leverages Corollary 4.4 to endogenize ˜ f1in obtaining the efficiency bound and in selecting qfor the optimal EWC test. I note that for the existence of the optimal test in those cases I restrict the ranges of ρ1and ρ2such that Assumption 4.2(c) is satisfied, but the corresponding spectral density ˜ f1may still be nonmonotone. Displayed results in panel (a) suggest that the new EWC test is nearly as powerful as the test ϕ∗at AR(2) alternatives, even if both are derived under different rationales and neither possesses a well-defined Figure 6. Boxplots of weighted average power differences between ϕ∗and EWC tests. Notes: The “uniformly maximal” function of Fcorresponds to an AR(1) with coefficient ρ0∈[0.8, 0.9]. Powers are directed at AR(2) alternatives with roots ρ1and ρ2(−0.8 ≤ρ1(ρ2)≤0.8 in panel (a); 0≤ρ1≤0.8 and −0.8 ≤ρ2≤0 in panel (b)). Sample size Tis 100, and the level of significance is 5%. Quantitative Economics 15 (2024) Optimal HAR inference 1127 Figure 7. Boxplots of weighted average power for weighted cosine tests. Notes: The “uniformly maximal” function fcorresponds to an AR(1) with root ρ0. Powers are directed at AR(2) alternatives with roots ρ1and ρ2(−0.7 ≤ρ1(ρ2)≤0.7). 5% tests using qweighted cosines are considered, in which 1000 sets of random positive weights are used for each (ρ0,ρ1,ρ2).SamplesizeT is 100. sense of optimality by construction. Panel (b) corroborates the near optimality feature of the optimal EWC test when ϕ∗is by design optimal according to Corollary 4.4 at a restricted yet considerably large set of AR(2) alternatives. I note that the appearance of negative values in Figure 6might be attributed to numerical errors in calculating the critical values for both tests.13 It is noted in Comment 3 of Section 3.1 that both the optimal and EWC tests belong to the so-called weighted cosine tests. I investigate whether the near optimality is a unique feature of EWC tests in that class. Specifically, I consider such tests using qweighted cosines but with 1000 sets of random positive weights that are normalized to sum to one. The critical values are obtained such that these tests exactly control size at AR(1) f with coefficient ρ0. Figure 7displays boxplots of the resulting weighted average power at AR(2) alternatives with roots ρ1and ρ2.Notethatq∗is by construction 7 and 5 when power is directed at f1in panels (a) and (b), respectively. It is found that the dispersion of weighted average power is relatively small at q∗across all tests and all AR(2) alternatives, suggesting that there may exist a set of weighted cosine tests, including the EWC test, that nearly achieve the efficiency bounds. In fact, the EWC test included is not even the 13For numerical stability, I set the tolerance level to be 0.003 around 0.05 in order to obtain the critical values, so it is not surprising to have numerical errors in approximating rejection probabilities to be of the order 10−3. With a smaller tolerance level, the numerical integration of (12) using 2000-point Gaussian quadrature may still produce complex-valued numbers. Moreover, in obtaining cvq∗for ϕ∗,becausethe positiveness conditions, and thus the integral expression of (12) may not hold a priori for a given ˜ q,Ifirst simulate 100,000 ϕ∗under fto determine a preliminary q∗and then use a bisection method to obtain amoreprecisecvq∗while checking the positiveness conditions. In the above senses, the plots in Figure 6 are bound by small numerical and simulation errors, and negative values are not evidence to dismiss the theories. 1128 Liyu Dou Quantitative Economics 15 (2024) best one, as its qis not optimized yet. Furthermore, the comparisons of boxplots across different qhint that one might need to choose weights (test statistic) more judiciously when a relatively large qis used, and the theoretical insights so far recommend the EWC test as a good candidate so long as the adjusted critical value is adopted. 4.3 Relaxation of the Whittle-type approximation The theoretical discussions so far are entirely based on the Whittle-type diagonal structure. For both theoretical interest and practical relevance, it is natural to ask whether the above insights on optimal HAR inference continue to hold without that structure. To that end, I maintain the criteria of weighted average power maximizing and scale invariance and still direct power at the flat spectrum f1. The goal is to seek powerful tests as functions of Ys=Y/√YYin the problem of He 0:Y∼N0, 0(f),f∈F(16) against He 1,f1:Y∼N0, κT−1diag1, κ−1,,κ−1, where the superscript edenotes the exact model. Note that He 1,f1is identical to Hd 1,f1, because 1exactly becomes diagonal under f1. First of all, I note that it is, in general, difficult to derive the optimal test of (16). This is mainly due to the complicated manner by which fenters 0(f). In this case, even if it is true that the least favorable distribution puts a point mass on some function f∗∈F, its determination seems rather difficult. Despite so, one still can obtain bounds on the power of any size-controlling test by using the bounding approach of Elliott, Müller, and Watson (2015). Recall from Section 2that for any probability distribution over F,the likelihood ratio test of H 0,against H 1,f1yields such a power bound. If the power of a valid ad hoc test ϕah is close to the power bound for some ,thenϕah is known to be close to optimal, as no substantially more powerful test exists. It turns out that the insights from the diagonal model are useful in guessing a good and in suggesting the near optimality of the EWC test in the exact model. In particular, for a given ain [0, 1], let abe a point mass distribution on the kinked function fa(φ), as was defined in Assumption 3.2(c). For every a, the likelihood ratio test of H 0,aagainst H 1,f1yields a power bound. I numerically search for asuch that the resulting power bound is minimized. Denote this aby a†and the resulting by †. The power bound I employ to gauge the efficiency of ad hoc tests is then the power of ϕ†,f1=1Y0(fa†)Y−1Y1(f1)Y>cv, (17) for some cv such that E[ϕ†,f1]=αunder H 0,a. It turns out that the EWC test essentially achieves this bound, after appropriate critical value adjustment and optimally choosing q. Quantitative Economics 15 (2024) Optimal HAR inference 1129 For an EWC test with a given q, the null rejection probability at a given fand critical value cv now becomes P Y0     q  j=1 Y2 j/q ≥cv=PZ2 0 q  j=1 λj(f)Z2 j ≥1, (18) in which {λj(f)}q j=1are positive eigenvalues of M(cv,q)0,q(f)(normalized by the magnitude of the only negative eigenvalue),14 where 0,q(f)is the upper left (q+1)×(q+1) block matrix of 0(f)and M(cv,q)=diag(−1, cv2/q,cv2/q,,cv2/q).Bythesamearguments in Section 3,(18) is maximized at fsuch that all λj(f)’s are jointly minimized. The opaque mapping from λj(f)back to f, however, prevents us from explicitly identifying the null rejection probability maximizer(s) like funder Assumption 3.2, rendering the critical value adjustment generally infeasible. In spatial settings, Müller and Watson (2022) consider a parametric Fwith a well-defined bound on the parameter as a benchmark model to obtain feasible adjustment and then robustify the size-controlling property of the resulting EWC-type test in more general model classes using the above eigenvalue insights. I, however, take a numerical approach: I approximate fas a linear combination ˆ f of basis functions, numerically search the weights such that resulting ˆ fmaximizes (18) under an additional assumption that fis nonincreasing over [0, π],15 and obtain the critical value cva,e qaccordingly in the same way as in Section 3. See Appendix Bfor the computational details. I then proceed as in Section 3.2 to select qoptimally. In the context of Tables 3and 4(AR(1) f), it is found that the difference between cva qand cva,e qis considerably small and the largest size distortion of 5% level EWC test using cva qis of the order 10−3in the exact model. These numerical findings are considerably robust when different F’s are considered.16 Tables 6and 7summarize the weighted average power of the optimal EWC test and the weighted average power bound induced by (17), paralleling the exercises in Tables 4 and 5, respectively. As can be seen, for most Fconsidered, the optimal EWC test essentially achieves the corresponding weighted average power bound, and thus possesses a notion of near optimality in the spirit of Lemma 1 in Elliott, Müller, and Watson (2015). I note that the relatively larger difference between the weighted average power of the optimal EWC test and the corresponding bound (e.g., under large ρin Table 6and under large Cin Table 7) is not informative about the efficiency of the optimal EWC test, since it can arise either because the bound is far from the least upper bound, or because the ad hoc test is inefficient. The practical implications of using the EWC test from Section 3.2.2 remain. In the exact model and in the context of Figure 5, the conventional wisdom and the optimal 14See Lemma 1(i) in Müller and Watson (2022) for the proof that there is only one negative eigenvalue. Note the sign difference between M(cv,q)here and D(cv)in their context. 15The benchmark models considered in Müller and Watson (2022) all satisfy this shape restriction. 16I refer interested readers to Tables 7, 12, 13, 14, and 15 in an earlier version of this paper (cf. Dou (2020)) for numerical details. I choose not to include them in this version for the sake of space. 1130 Liyu Dou Quantitative Economics 15 (2024) Table 6. Weighted average power (WAP) bound and the WAP of the optimal EWC test. ρ0.50 0.60 0.70 0.80 0.90 0.95 0.98 0.99 WAP of optimal EWC 0.501 0.486 0.465 0.432 0.345 0.228 0.095 0.067 WAP bound 0.505 0.493 0.475 0.440 0.359 0.254 0.133 0.087 Note:Theffunction of Fcorresponds to an AR(1) with coefficient ρ.Allfin Fare nonincreasing over [0, π].Nominal level is 5%. Sample size Tis 100. EWC test continue to suggest using 3 and 6 for q, respectively, but the corresponding Student-tcritical value has to be enlarged by a slightly higher factor of 1.14 for exact size control. In terms of weighted average power, there is a 14% gain by using the optimal EWC test. This efficiency advantage is further evident when the local alternative μ= δT−1/2(1−ρ1)−1is considered with power directed at ρ1=0 and even with cva q,asin Figure 1. I reiterate the general takeaway here: One should use the EWC test with a larger qand appropriately enlarged critical values for more powerful HAR inference. 5. Practical implementation In this section, I discuss the practical implementation of the optimal EWC test in the location model and extend it to inference about a scalar parameter in regression models. 5.1 Location model Recall that the above theoretical discussions suggest the advantage of using a larger q and adjusted critical value when implementing the EWC test, and this new test possesses a notion of near optimality under pre-specified efficiency criteria and a smoothness class F. As a practical matter, one might like to estimate the smoothness class F,in particular the associated f, from data. Unfortunately, the attempt is not useful in theory. This is because the (nearly) optimal tests depend on F,anda“larger”Fincorporating sampling uncertainties potentially leads to a lower power. Put differently, one cannot estimate Fandstillcontrolsize(cf.Pötscher (2002)). But how to determine a reasonable fin actual implementations? Given the analogy between the optimal EWC test and the test considered in Sun (2014a)whenfis parameterized in the local-to-unity form (see Comment 4 in Section 3), I follow Sun (2014a)and suggest the practitioners calibrate fin the following way when testing about the population mean of an observed scalar time series {yt}T t=1. One first computes the OLS estimator Table 7. Weighted average power (WAP) bound and the WAP of the optimal EWC test. C10.0 5.6 3.2 1.8 1.0 0.6 0.2 0.1 WAP of optimal EWC 0.307 0.368 0.425 0.460 0.486 0.501 0.524 0.530 WAP bound 0.321 0.381 0.431 0.466 0.488 0.505 0.527 0.534 Note:Theffunction of Fis f(φ)=exp(−Cφ).Nominallevelis5%. Sample size T=100. Quantitative Economics 15 (2024) Optimal HAR inference 1137 of variable u=√1−s. The expression (22)atn=1 becomes J1(ζ1)=2 π1 0 du 1−u2+ζ1=2 πarcsin 1 1+ζ1 , which follows from a change of variable v=u/√1+ζ1and the fact that the antiderivative of (1−v2)−1/2is arcsinv. Lemma A.5. For 0<α<1, (a) cv1exists if and only if m(π/T )=κ−1.m(π/T )≶κ−1if and only if 1≶cv1. (b) cond1holds if m(π/T )=κ−1. Proof.(a)In(20)at˜ q=1, if m(π/T )=κ−1,theredoesnotexistacv1and αsuch that (20) holds. On the other hand, a rearrangement of the event in (20)at˜ q=1gives PZ2 0+Z2 1>Z2 0+κm(π/T )Z2 1cv1=α. It follows that m(π/T )≶κ−1if and only if 1 ≶cv1.Moreover,ifm(π/T )>κ −1,thesecond part of Lemma A.4 in conjunction with (20)at˜ q=1gives cv1=1 κm(π/T )sin2(απ/2)+cos2(απ/2), which always exists for every 0 <α<1. In a similar vein, if m(π/T )<κ −1,wehave cv1=1 κm(π/T )cos2(απ/2)+sin2(απ/2), which always exists. Thus, cv1exists if and only if m(π/T )= κ−1. (b) follows from the above that cond1holds if m(π/T )=κ−1. In what follows, I fold most of the parts corresponding to cvq>1 in the proofs. Partly, this is because they can be worked out by exactly symmetric arguments. Also, the most relevant results from these auxiliary lemmas in establishing part (2) of Theorems 3.3, 4.3, and Corollary 4.4 are when κ>m (π/T )−1, which is equivalent to cvq<1forallqby Lemmas A.5 and A.10. Lemma A.6. For m(π/T )=κ−1and 0<α<1, if cond˜ qis violated for some 1<˜ q≤T−2, then condqis also violated for any ˜ q+1≤q≤T−1. Proof. Suppose cond˜ qis violated while cond˜ q+1holds. We have max j=1,2,,˜ qcv˜ qκm(πj/T )−1≥0and min j=1,2,,˜ qcv˜ qκm(πj/T )−1≤0. (23) Consider minj=1,2,,˜ q+1{cv˜ q+1κm(πj/T )−1}>0. We must have cv˜ q+1<1; otherwise, (20) does not hold at ˜ q+1. On the other hand, 0 <minj=1,2,,˜ q+1{cv˜ q+1κm(πj/T )− 1138 Liyu Dou Quantitative Economics 15 (2024) 1}≤minj=1,2,,˜ q{cv˜ q+1κm(πj/T )−1}. This, in conjunction with the second part of (23), implies that cv˜ q<cv˜ q+1<1, which we next show is impossible. Suppose cv˜ q<cv˜ q+1<1 is true. Denote A+ ˜ q={j|1≤j≤˜ q,cv˜ qκm(πj/T )−1>0}(A+ ˜ q=∅;otherwise,(20)isviolated for ˜ q). Now (20)at˜ qgives α=P(1−cv˜ q)Z2 0>˜ q  j=1cv˜ qκm(πj/T )−1Z2 j =PZ2 0>1 1−cv˜ q j∈A+ ˜ qcv˜ qκm(πj/T )−1Z2 j+1 1−cv˜ q j/∈A+ ˜ qcv˜ qκm(πj/T )−1Z2 j ≥PZ2 0>1 1−cv˜ q j∈A+ ˜ qcv˜ qκm(πj/T )−1Z2 j(24) >PZ2 0>1 1−cv˜ q+1 j∈A+ ˜ qcv˜ q+1κm(πj/T )−1Z2 j(25) >PZ2 0>1 1−cv˜ q+1 j∈A+ ˜ qcv˜ q+1κm(πj/T )−1Z2 j +1 1−cv˜ q+1 j/∈A+ ˜ qcv˜ q+1κm(πj/T )−1Z2 j(26) >PZ2 0>1 1−cv˜ q+1 ˜ q+1  j=1cv˜ q+1κm(πj/T )−1Z2 j=α, where (24) is due to the fact that P(A≥C+B)≥P(A≥C)when A,B,Care independent random variables and B≤0 almost surely. The inequality (25) is due to Lemma A.1 and the fact that for any j∈A+ ˜ q,1 1−cv˜ q[cv˜ qκm(πj/T )−1]<1 1−cv˜ q+1[cv˜ q+1κm(πj/T )−1] under cv˜ q<cv˜ q+1<1. The inequality (26) is due to Lemma A.1. The other half of (23)correspondstocv˜ q+1>1 and the proof is exactly symmetric to the above. Overall, we have for m(π/T )=κand 0 <α<1, if cond˜ qis violated for some 1<˜ q≤T−2, then condqis also violated for any ˜ q+1≤q≤T−1 by inductions. Corollary A.7. For m(π/T )=κ−1and 0<α<1, if cond˜ qholds for some 3≤˜ q≤T−1, then condqalso holds for any 2≤q≤˜ q−1. Proof. This is the contrapositive statement of Lemma A.6. Corollary A.8. For m(π/T )=κ−1and 0<α<1, either one of the following will hold: (a) there exists a unique 1≤q∗≤T−2such that condqis satisfied for all 1≤q≤q∗ and violated for all q∗+1≤q≤T−1; (b) condqis satisfied for all 1≤q≤T−1. In this case,define q∗=T−1. Quantitative Economics 15 (2024) Optimal HAR inference 1139 Proof.Ifcond T−1holds, by Corollary A.7, (b) is true. Otherwise, if condT−2holds, then (a) is true with q∗=T−2. Otherwise, given that cond1always holds by Lemma A.5, backward inductions lead (a) to be true for a unique 1 ≤q∗≤T−3. Corollary A.9. If mequals 1and 0<α<1, then condqholds for all 1≤q≤T−1. Proof.Inthiscase, F=G={m}={f1}and condT−1is trivially satisfied. It follows that condqholds for all 1 ≤q≤T−1 by Corollary A.7. Lemma A.10. For m(π/T )=κ−1,0<α<1, and q∗as defined in Corollary A.8,either one of the following will hold: (a) cvq>1for all 1≤q≤q∗,and if q∗≥2, cvq+1>cvq,q=1, 2, ,q∗−1; (b) cvq<1for all 1≤q≤q∗,and if q∗≥2, cvq+1<cvq,q=1, 2, ,q∗−1. Proof. Lemma A.5 leads to the conclusions for q∗=1. We now focus on q∗≥2. Suppose κm(π/T )>1, then Lemma A.5 implies that cv1<1. Suppose there exists a ˜ q= min{q|2≤q≤q∗,cvq>1}.(Inotethatcvqcannot be 1 for any q≤q∗;otherwise,(20) cannot hold at the corresponding q.) Then we must have max j=1,2,,˜ q−1cv˜ q−1κm(πj/T )−1<max j=1,2,,˜ q−1cv˜ qκm(πj/T )−1 ≤max j=1,2,,˜ qcv˜ qκm(πj/T )−1<0. This is a contradiction, because minj=1,2,,˜ q−1{cv˜ q−1κm(πj/T )−1}>0. It subsequently implies that cvq<1forall1≤q≤q∗. Moreover, for each j=1, ,q∗,Rj(x)= [κm(πj/T )x−1]/(1−x)is monotonically increasing in (0, 1). (To see this, note that κm(πj/T )−1>κm (πj/T )cvq∗−1>0 for every 1 ≤j≤q∗.) For (20) to hold sequentially, we necessarily need cvq+1<cvq,q=1, 2, ,q∗−1. (Otherwise, the LHS of (20)would always be below αby Lemma A.1.) Part (b) is proved, and part (a) holds by symmetric arguments. Lemma A.11. For m(π/T )= κ−1,0<α<1, and q∗as defined in Corollary A.8,if additionally m(πj/T )≥m(π(j+1)/T),j=0, 1, ,T−2, and q∗<T−1, then κ−1(m(πj/T ))−1≥cvq∗for j>q ∗. Proof. Define Q(x,q)=P((1−x)Z2 0>q j=1[xκm(πj/T )−1]Z2 j).Givenm(πj/T )≥ m(π(j+1)/T),j=q∗+1, ,T−2, it suffices to show κ−1(m(π(q∗+1)/T))−1≥cvq∗. Suppose not; then we must have κ−1(m(π(q∗+1)/T))−1<cvq∗. Suppose κm(π/T )>1, then cvq∗<1 by Lemma A.10: Qcvq∗,q∗+1=P[1−cvq∗]Z2 0> q∗+1  j=1cvq∗κm(πj/T )−1Z2 j <P[1−cvq∗]Z2 0> q∗  j=1cvq∗κm(πj/T )−1Z2 j=Qcvq∗,q∗=α. 1140 Liyu Dou Quantitative Economics 15 (2024) On the other hand, 0 <κ −1(m(π(q∗+1)/T))−1<cvq∗<1. Then Qκ−1mπq∗+1/T−1,q∗+1 =P1−κ−1mπq∗+1/T−1Z2 0> q∗+1  j=1mπq∗+1/T−1m(πj/T )−1Z2 j =P1−κ−1mπq∗+1/T−1Z2 0> q∗  j=1mπq∗+1/T−1m(πj/T )−1Z2 j =Qκ−1mπq∗+1/T−1,q∗>Q cvq∗,q∗=α, where the last but one inequality follows from the fact that Q(·,q∗)is monotonically decreasing in (0, cvq∗)under κm(π/T )>1. By the continuity of Q(·,q∗+1) and the intermediate value theorem, there must exist a number, denoted by cvq∗+1, such that Q(cvq∗+1,q∗+1)=α. There is a contradiction, because condq∗+1now holds, violating Corollary A.8. Suppose κm(π/T )<1 instead, the proof follows by almost symmetric arguments as above. Overall, we have κ−1(m(π(q∗+1)/T))−1= minq∗+1≤j≤T−1κ−1(m(πj/T ))−1≥cvq∗. A.2 Proof of Theorem 4.3 Part 1 holds by the definition of β0,1and by simply recognizing that the alternative Hd 1,f(defined analogous to Hd 1,f1) is included in the null Hd 0. Any nontrivial sizecontrolling test thus cannot be more powerful than the trivial randomized test, which does not depend on any 0nor 1. In part 2, m(π/T )>κ −1. Under Assumption 4.2(a,b,c) and for 0 <α<1, Corollary A.8 shows that there exists a unique q∗such that either (i) condqholds for 1≤q≤q∗and is violated for q∗+1≤q≤T−1, or (ii) condqholds for all 1 ≤ q≤T−1,wherewedefineq∗=T−1. I conjecture that the pair of least favorable distributions (∗ 0,∗ 1) put probability masses on functions {f∗}⊂Fand {g∗}⊂ G,inwhichf∗(φ)=f(φ)1[|φ|≤πq∗/T]+a∗(φ)f(φ)1[|φ|>πq ∗/T]and g∗(φ)= f(φ)1[|φ|≤πq∗/T]+b∗(φ)f(φ)1[|φ|>πq ∗/T],fora∗(φ)≥1andb∗(φ)≤1suchthat a∗(φ)f(φ)/(b∗(φ)f(φ)) =(κcvq∗)−1for |φ|>πq ∗/T. Assumption 4.2(d,e) ensures that such two sets of functions are nonempty as long as the pair of functions (a∗,b∗)exists. This true by Lemma A.11 (m(φ)≤(κcvq∗)−1for |φ|>πq ∗/T). Now first let b∗(φ)=1forallφand a∗(φ)=(κcvq∗m(φ))−1. The best level αtest of Hd 0,f∗against Hd 1,g∗is ϕf∗,g∗=1Y2 0+ T−1  j=1 Y2 j/f ∗(πj/T ) Y2 0+κ T−1  j=1 Y2 j/g∗(πj/T ) >cv, Quantitative Economics 15 (2024) Optimal HAR inference 1141 for some cv ≥0suchthatEPY,f∗[ϕf∗,g∗(Ys)] =α,wherePY,˜ fdenotes the joint distribution of Yat f=˜ funder Hd 0. It follows that α=PY,f∗Y2 0+ T−1  j=1 Y2 j/f ∗(πj/T ) Y2 0+κ T−1  j=1 Y2 j/g∗(πj/T ) >cv =PY,f∗Y2 0+ T−1  j=1 Y2 j/f ∗(πj/T )>cvY2 0+κ T−1  j=1 Y2 j/g∗(πj/T ) =P(1−cv)Z2 0> T−1  j=1cvκf ∗(πj/T )/g∗(πj/T )−1Z2 j =P(1−cv)Z2 0> q∗  j=1cvκm(πj/T )−1Z2 j+ T−1  j=q∗+1 [cv/cvq∗−1]Z2 j, (27) where the last equality follows from the definition of f∗and g∗. Because Yis a continuous random vector, the critical value cv is unique. By matching (27)with (20)at˜ q=q∗,wehavecv =cvq∗.Also,theevents{Y2 0+T−1 j=1Y2 j/f ∗(πj/T ) Y2 0+κT−1 j=1Y2 j/g∗(πj/T )>cvq∗}and {Y2 0+q∗ j=1Y2 j/f(πj/T ) Y2 0+κq∗ j=1Y2 j/f(πj/T )>cvq∗}are equivalent PY,˜ f-almost surely, uniformly in ˜ f∈F.The rejection regions defined by ϕf∗,g∗and the optimal test statistic in (14) are thus identical. It remains to check the following conditions: (1) ϕf∗,g∗is also the best level αtest of Hd 0,∗ 0against Hd 1,∗ 1;(2)ϕf∗,g∗uniformly controls size Hd 0and has its largest size distortion at f;(3)ϕf∗,g∗attains β∗ 0,∗ 1at f; and (4) for any other (0,1),β∗ 0,∗ 1≤ β0,1. For (1), note that ϕf∗,g∗is of exact size αunder H0,∗ 0: ϕf∗,g∗(ys)h0,f(ys)d∗ 0(f)dys=ϕf∗,g∗(ys)h0,f∗(ys)dysh0,fd∗ 0(f) =EPY,f∗[ϕf∗,g∗(Ys)] =α, where the second equality holds because the set of distributions of ϕf∗,g∗is degenerate under ∗ 0. By the same logic, the rejection probabilities of ϕf∗,g∗under H1,g∗and 1142 Liyu Dou Quantitative Economics 15 (2024) H1,∗ 1are identical. Because the best level αtest of Hd 0,∗ 0against Hd 1,∗ 1is unique, (1) holds. For(2),consideragiven ˜ f∈F, EPY,˜ fϕf∗,g∗(Y)=PY,˜ fY2 0+ q∗  j=1 Y2 j/f(πj/T ) Y2 0+κ q∗  j=1 Y2 j/f(πj/T ) >cvq∗ =P[1−cvq∗]Z2 0> q∗  j=1cvq∗κm(πj/T )−1˜ f(πj/T ) f(πj/T )Z2 j =PZ2 0>1 1−cvq∗ q∗  j=1cvq∗κm(πj/T )−1˜ f(πj/T ) f(πj/T )Z2 j(28) ≤PZ2 0>1 1−cvq∗ q∗  j=1cvq∗κm(πj/T )−1Z2 j=α, where (28)followsfrom(b)inLemmaA.10 under the condition m(π/T )>κ −1, and the inequality follows from the definition of q∗and Lemma A.1 under Assumption 4.2(a). For(3),consideragiven ˜ g∈Gand let P1 Y,˜ gdenote the joint distribution of Yunder Hd 1, ˜ g, EP1 Y,˜ gϕf∗,g∗(Y)=P1 Y,˜ gY2 0+ q∗  j=1 Y2 j/f(πj/T ) Y2 0+κ q∗  j=1 Y2 j/f(πj/T ) >cvq∗ =P[1−cvq∗]Z2 0> q∗  j=1cvq∗−κ−1m(πj/T )−1˜ g(πj/T ) f(πj/T )Z2 j =PZ2 0>1 1−cvq∗ q∗  j=1cvq∗−κ−1m(πj/T )−1˜ g(πj/T ) f(πj/T )Z2 j ≥PZ2 0>1 1−cvq∗ q∗  j=1cvq∗−κ−1m(πj/T )−1Z2 j, where the inequality follows from Corollary A.2 under Assumption 4.2(b). In this sense, β∗ 0,∗ 1=EP1 Y,f [ϕf∗,g∗(Y)]. Quantitative Economics 15 (2024) Optimal HAR inference 1143 For (4), because ϕf∗,g∗uniformly controls size under Hd 0, it also controls size under Hd 0,0. Then by the definition of β0,1, β0,1≥ϕf∗,g∗ysh1, ˜ gysd1(˜ g)dys≥inf ˜ g∈GEP1 Y,˜ gϕf∗,g∗(Y)=β∗ 0,∗ 1. A.3 Proofs of Theorem 3.3 and Corollary 4.5 Proof of Corollary 4.5 follows immediately from those in Section A.2 with f=f1.Theorem 3.3 is a special case of Corollary 4.5 with G={f1}, and its results follow by realizing that β∗ 0,∗ 1, in that case, is simply the weighted average power of test (8)atf1. A.4 Proof of Corollary 4.4 Corollary 4.4 is a special case of those considered in Theorem 4.3 with G={˜ f1}. The proof thus follows directly from those in Section A.2. Appendix B: Computational details in Section 4.3 In this section, I explain in detail how to numerically identify the null rejection probability maximizer of the EWC test in testing (16). Let the n+1 node points {xi}n i=0define a partition of the interval I=[0, π]into nsubintervals Ii=[xi−1,xi],i=1, 2, ,n, each of length hi=xi−xi−1,andx0= 0, xn=π.LetC0(I)denote the space of continuous functions on I,andP1(Ii)denote the space of linear functions on Ii.Let{ςi}n i=0be a set of basis functions for the space Fhof continuous piecewise linear functions defined by Fh={f:f∈C0(I),f|Ii∈ P1(Ii)}. The basis functions {ςi}n i=0are normalized such that ςj(xi)=1[i=j],i,j= 0, 1, ,n. By approximating fvia ˆ f=n i=0f(xi)ςiand by (12), I approximate (18) by PZ2 0 q  j=1 λj(ˆ f)Z2 j ≥1=2 π1 01−u2(q−1)/2du     q  j=11−u2+λj(ˆ f) , (29) which is a function of the n-dimensional vector (f(x1),f(x2),,f(xn)).(Bynormalization, f(x0)=1.) With pre-computed {0(ςi)}n i=0,thecomputationof(29) takes very little computing time for each ˆ f, and it is feasible to obtain a global maximizer of (29) subject to implied constraints on (f(x1),f(x2),,f(xn))from a given F. I additionally assume that the underlying spectrum is nonincreasing over [0, π]. In actual implementations, I choose n=50, and {xi}50 i=0are log-spaced nodes in [0, π]. The basis functions {ςi}n i=0are chosen to be the hat functions ςi(x)=⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ (x−xi−1)/hiif x∈Ii, (xi+1−x)/hi+1if x∈Ii+1, 0otherwise. (30) 1144 Liyu Dou Quantitative Economics 15 (2024) I pre-compute {0(ςi)}n i=0with a 5000-point Gaussian quadrature for each nonzero element. Since each ςiis compactly supported, these numerical integrations are nearly precise. For every (f(x1),f(x2),,f(xn)),0(ˆ f)is simply a linear combination of these pre-computed covariance matrices. Note, however, that the ultimate objective function I will optimize is (29), which involves 0(f)implicitly through λj(ˆ f).Inanunreported exercise, for a given EWC test and a parametric AR(1) class Fwith coefficient varying over a fine grid, I compare the rejection probabilities following the described approximate procedure and an “exact” procedure in which each entry of 0(f)is evaluated by numerical integrations via Mathematica. The differences in the rejection probabilities are at most of the order 0.0001. I thus hold on to the above choice of nand {xi}50 i=0. I proceed in three steps to identify the null rejection probability maximizer for a fixed EWC test of (16), in terms of (f(x1),f(x2),,f(xn)): (i) Program up the null rejection probability at a given (f(x1),f(x2),,f(xn))∈ Rn +, and with a given qand reasonable cv (e.g., cva q)from(29), where 0(ˆ f)= n i=0f(xi)0(ςi)with pre-computed {0(ςi)}n i=0. (ii) Randomly draw 100 n-dimensional vectors (f(x1),f(x2),,f(xn))such that each vector corresponds to some f∈F. This is, in general, a challenging task since the number of numerical constraints to be checked increases exponentially with nfor higher-order smoothness constraints. For feasibility, I focus on two types of smoothness classes: the class Fin which fcorresponds is AR(1) with coefficient ρand f∈Fis nonincreasing over [0, π];andtheclassFin which f∈Fis Lipschitz continuous in logs with Lipschitz constant C.Itis not hard to see that for the first type, it suffices to check the monotonicity constraint consecutively and the lower boundedness condition. For the second type, by the result of Beliakov (2006), the complexity of checking the global Lipschitz condition is reduced to consecutive checking of local Lipschitz conditions. (iii) Use every n-dimensional vector drawn in Step (ii) as the initial condition to optimize the null rejection probability function programmed in Step (i), subject to linear constraints induced by smoothness class F(as described in Step (ii)). Under the above specifications, it takes about 1 to 2 minutes to complete the optimization using fmincon in MATLAB via parallel computing in 12 cores. References Andrews, Donald W. K. (1991), “Heteroskedasticity and autocorrelation consistent covariance matrix estimation.” Econometrica, 59 (3), 817–858. [1107,1132] Armstrong, Timothy B. and Michal Kolesár (2018), “Optimal inference in a class of regression models.” Econometrica, 86 (2), 655–683. [1112,1121] Quantitative Economics 15 (2024) Optimal HAR inference 1145 Armstrong, Timothy B. and Michal Kolesár (2020), “Simple and honest confidence intervals in nonparametric regression.” Quantitative Economics, 11 (1), 1–39. [1112,1131] Atchadé, Yves F. and Matias D. Cattaneo (2011), “Limit theorems for quadratic forms of Markov chains.” arXiv:1108.2743 [math]. [1108] Bakirov, Nail K. (1996), “Comparison theorems for distribution functions of quadratic forms of Gaussian vectors.” Theory of Probability & Its Applications, 40 (2), 340–348. [1136] Bakirov, Nail K. and Gabor J. Székely (2006), “Student’s t-test for Gaussian scale mixtures.” Journal of Mathematical Sciences, 139 (3), 6497–6505. [1120,1136] Beliakov, Gleb (2006), “Interpolation of Lipschitz functions.” Journal of Computational andAppliedMathematics, 196 (1), 20–44. [1144] Bureau of Labor Statistics (2024), Series LNS14000000, 1948:1 to 2019:12, https://beta. bls.gov/dataViewer/view/timeseries/LNS14000000. U.S. Bureau of Labor Statistics. Accessed January 9, 2024. [1133] Choudhuri, Nidhan, Subhashis Ghosal, and Anindya Roy (2004), “Contiguity of the Whittle measure for a Gaussian time series.” Biometrika, 91 (1), 211–218. [1115] den Haan, Wouter J. and Andrew T. Levin (1997), “A practitioner’s guide to robust covariance matrix estimation.” In Handbook of Statistics,Vol.15(G.S.MaddalaandC.R.Rao, eds.), 299–342, Elsevier. [1108] den Haan, Wouter J. and Andrew T. Levin (1998), “Vector autoregressive covariance matrix estimation.” Manuscript, Board of Governors of the Federal Reserve. [1108] Diebold, Francis X. and Robert S. Mariano (1995), “Comparing predictive accuracy.” Journal of Business & Economic Statistics, 13 (3), 134–144. [1134] Dou, Liyu (2020), “Optimal HAR inference.” Working paper, CUHK-Shenzhen. [1129] Elliott, Graham, Ulrich K. Müller, and Mark W. Watson (2015), “Nearly optimal tests when a nuisance parameter is present under the null hypothesis.” Econometrica,83(2), 771–811. [1128,1129] Giacomini, Raffaella and Halbert White (2006), “Tests of conditional predictive ability.” Econometrica, 74 (6), 1545–1578. [1135] Golubev, Georgi K., Michael Nussbaum, and Harrison H. Zhou (2010), “Asymptotic equivalence of spectral density estimation and Gaussian white noise.” The Annals of Statistics, 38 (1), 181–214. [1115] Gonçalves, Sílvia and Timothy J. Vogelsang (2011), “Block bootstrap HAC robust tests: The sophiscation of the naive bootstrap.” Econometric Theory, 27 (4), 745–791. [1108] 1146 Liyu Dou Quantitative Economics 15 (2024) Hwang, Jungbin and Yixiao Sun (2017), “Asymptotic F and t tests in an efficient GMM setting.” Journal of Econometrics, 198 (2), 277–295. [1113] Hwang, Jungbin and Yixiao Sun (2018), “Should we go one step further? An accurate comparison of one-step and two-step procedures in a generalized method of moments framework.” Journal of Econometrics, 207 (2), 381–405. [1113] Ibragimov, Rustam and Ulrich K. Müller (2010), “t -statistic based correlation and heterogeneity robust inference.” Journal of Business & Economic Statistics, 28 (4), 453–468. [1108,1132] Jansson, Michael (2004), “The error in rejection probability of simple autocorrelation robust tests.” Econometrica, 72 (3), 937–946. [1108,1112] Jordà, Òscar (2005), “Estimation and inference of impulse responses by local projections.” American Economic Review, 95 (1), 161–182. [1107] Kariya, Takeaki (1980), “Locally robust tests for serial correlation in least squares regression.” The Annals of Statistics, 8 (5), 1065–1070. [1114] Kiefer, Nicholas M., Timothy J. Vogelsang, and Helle Bunzel (2000), “Simple robust testing of regression hypotheses.” Econometrica, 68 (3), 695–714. [1108,1132] Kiefer, Nicholas M. and Timothy J. Vogelsang (2002), “Heteroskedasticityautocorrelation robust standard errors using the Bartlett kernel without truncation.” Econometrica, 70 (5), 2093–2095. [1108] Kiefer, Nicholas M. and Timothy J. Vogelsang (2005), “A new asymptotic theory for heteroskedasticity-autocorrelation robust tests.” Econometric Theory, 21 (6), 1130–1164. [1108] King, Maxwell L. (1980), “Robust tests for spherical symmetry and their application to least squares regression.” The Annals of Statistics, 8 (6), 1265–1271. [1114] King, Maxwell L. (1987), “Towards a theory of point optimal testing.” Econometric Reviews, 6 (2), 169–218. [1113] Koijen, Ralph S. J. and Stijn Van Nieuwerburgh (2011), “Predictability of returns and cash flows.” Annual Review of Financial Economics, 3 (1), 467–491. [1107] Lazarus, Eben, Daniel J. Lewis, James H. Stock, and Mark W. Watson (2018), “HAR inference: Recommendations for practice.” Journal of Business & Economic Statistics,36(4), 541–559. [1108,1132,1133,1134,1135] Lazarus, Eben, Daniel J. Lewis, and James H. Stock (2021), “The size-power tradeoff in HAR inference.” Econometrica, 89 (5), 2497–2516. [1108,1111,1112,1121,1134] Lehmann, Erich L. and Joseph P. Romano (2005), Testing Statistical Hypotheses,thirdedition edition. Springer, New York. ISBN 978-0-387-98864-1. [1114,1124] Lehmann, Erich L. and Charles Stein (1948), “Mostpowerful tests of composite hypotheses. I. Normal distributions.” The Annals of Mathematical Statistics, 19 (4), 495–516. [1112]