scieee AI-readable full text Open interactive document viewer

Maximum likelihood inference in weakly identified dynamic stochastic general equilibrium models

Andrews, Isaiah,Mikusheva, Anna

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Andrews, Isaiah; Mikusheva, Anna Article Maximum likelihood inference in weakly identified dynamic stochastic general equilibrium models Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Andrews, Isaiah; Mikusheva, Anna (2015) : Maximum likelihood inference in weakly identified dynamic stochastic general equilibrium models, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 6, Iss. 1, pp. 123-152, https://doi.org/10.3982/QE331 This Version is available at: https://hdl.handle.net/10419/150382 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/3.0/ Quantitative Economics 6 (2015), 123–152 1759-7331/20150123 Maximum likelihood inference in weakly identified dynamic stochastic general equilibrium models Isaiah Andrews Department of Economics, MIT Anna Mikusheva Department of Economics, MIT This paper examines the issue of weak identification in maximum likelihood, motivated by problems with estimation and inference in a multidimensional dynamic stochastic general equilibrium model. We show that two forms of the classical score (Lagrange multiplier) test for a simple hypothesis concerning the full parameter vector are robust to weak identification. We also suggest a test for a composite hypothesis regarding a subvector of parameters. The suggested subset test is shown to be asymptotically exact when the nuisance parameter is strongly identified. We pay particular attention to the question of how to estimate Fisher information and we make extensive use of martingale theory. Keywords. Maximum likelihood, C(α) test, score test, weak identification. JEL classification. C32. 1. Introduction In recent years, we have witnessed the rapid growth of the empirical literature on the highly parameterized, microfounded macro models known as dynamic stochastic general equilibrium (DSGE) models. A number of papers in this literature have considered estimating these models by maximum likelihood (see, for example, Altug (1989), Ingram, Kocherlakota, and Savin (1994), Ireland (2004), Lindé (2005), and McGrattan, Rogerson, and Wright (1997)). More recently, Bayesian estimation has become increasingly popular, due in large part to the difficulty of maximum likelihood estimation in many DSGE models. As Fernández-Villaverde (2010) points out in his survey of DSGE estimation, “likelihoods of DSGE models are full of local maxima and minima and of Isaiah Andrews: [email protected] Anna Mikusheva: [email protected] We would like to thank Patrik Guggenberger, Ulrich Muller, Whitney Newey, Serena Ng, Zhongjun Qu, Frank Schorfheide, Jim Stock, the anonymous referees, and seminar participants at the Winter Econometric Society Meeting in Chicago, Boston College, Canadian Econometric Study Group, Columbia, Harvard–MIT, NBER summer institute, Rice, Texas A&M, UC Berkeley, UC San Diego, UPenn, U Virginia, and Yale for helpful comments. We are grateful to Lynda Khalaf for guidance on implementing the procedure of Dufour, Khalaf, and Kichian (2013). Andrews gratefully acknowledges support from the Ford Foundation and the NSF Graduate Research Fellowship under Grant 1122374. Mikusheva gratefully acknowledges financial support from the Castle–Krob Career Development Chair and the Sloan Research Fellowship. Copyright ©2015 Isaiah Andrews and Anna Mikusheva. Licensed under the Creative Commons Attribution- NonCommercial License 3.0. Available at http://www.qeconomics.org. DOI: 10.3982/QE331 124 Andrews and Mikusheva Quantitative Economics 6 (2015) nearly flat surfacesthe standard errors of the estimates are notoriously difficult to compute and their asymptotic distribution a poor approximation to the small sample one.” The poor performance of maximum likelihood estimation has fueled growing concerns about weak identification in many DSGE models (see Canova and Sala (2009), Guerron-Quintana, Inoue, and Kilian (2013), Iskrev (2010), and Mavroeidis (2005)). In this paper, we consider the problem of weak identification in models estimated by maximum likelihood, focusing in particular on weakly identified DSGE models. Weak identification arises when the amount of information in the data about some parameter or group of parameters is small and is generally modeled in such a way that information about parameters accumulates slowly along some dimensions. This leads to the breakdown of the usual asymptotics for maximum likelihood, but is distinct from loss of point identification. We assume throughout that the models we consider are point-identified, and thus that changing the value of any parameter changes the distribution of the data, though the effect will be small for some parameters. We provide several examples illustrating ways in which weak identification may arise in a DSGE context.1 We focus on the problem of testing and confidence set construction in this context. We consider two different tasks. First, we examine the problem of testing a simple hypothesis on the full parameter vector. We suggest using particular forms of the classical Lagrange multiplier (LM) test, which we show are robust to weak identification. The assumptions needed for this result are extremely weak and cover a large number of cases, including all of our examples. An advantage of our approach is that we can remain agnostic about the source and nature of weak identification, and need not rely on any particular asymptotic embedding. The proof for these tests makes extensive use of martingale theory, particularly the fact that the score (i.e., the gradient of the log likelihood) is a martingale when evaluated at the true parameter value. Second, we turn to the problem of testing a subset of parameters without restricting the remaining parameters. The tests we suggest for a subset of parameters are particular forms of Rao’s score test and are asymptotically equivalent to Neyman’s C(α) test when identification is strong. Consequently, our tests are efficient when all parameters are strongly identified. We show that the suggested tests have a χ2asymptotic distribution as long as the nuisance parameter (i.e., the part of the parameter vector that we are not testing) is strongly identified, even when the tested parameter is weakly identified. By combining our procedure for concentrating out nuisance parameters that are strongly identified with projection over the remaining nuisance parameters, one obtains weak identification-robust tests more powerful than those based on projection alone. The paper also reveals a previously unnoticed fact concerning estimation of the Fisher information. White (1982) noted that in strongly identified models, the Fisher information can be estimated using either the Hessian of the likelihood or the quadratic variation of the score, and argued that a large discrepancy between these two estimates indicates model misspecification. We show in examples that weak identification leads to a distinct but related phenomenon. In particular, under weak identification, the appropriately normalized quadratic variation of the score converges to fixed positive-definite 1Due to space limitations, most of the examples are placed in a Supplement, available as a supplementary file on the journal website, http://qeconomics.org/supp/331/supplement.pdf. Quantitative Economics 6 (2015) Maximum likelihood inference 125 matrix while the Hessian converges in distribution to a random matrix. Thus, large disparities between different estimators of information may arise even in correctly specified models if identification is weak. The issue of weak identification in DSGE models was first highlighted by Mavroeidis (2005) and Canova and Sala (2009), who pointed out that the objective functions implied by many DSGE models are nearly flat in some directions. Weak identificationrobust inference procedures for log-linearized DSGE models were introduced by Dufour, Khalaf, and Kichian (2013; henceforth DKK), Guerron-Quintana, Inoue, and Killian (2013; henceforth GQIK), and Qu (forthcoming). With the exception of GQIK, these papers focus on tests for the full parameter vector and make extensive use of the projection method to construct confidence sets for subsets of the structural parameters that, given the high dimension of the parameter space in many DSGE models, has the potential to introduce a substantial amount of conservativeness in many applications. The LM tests we suggest in this paper can be applied whenever the correct likelihood is specified and, in particular, can accommodate nonlinear DSGE models, which are increasingly popular and cannot be treated by existing weak identification-robust methods. We compare our LM tests with the existing weak identification-robust methods from a theoretical perspective, and report an extensive simulation study in a smallscale DSGE model, demonstrating the advantages and disadvantages of different robust methods. In simulation, we find that our LM statistics have much higher power than the limited information tests suggested by DKK. The test statistic proposed by Qu (forthcoming) is almost indistinguishable from our LMestatistic, but is defined for a much more limited set of models. The test of GQIK has power comparable to the LM tests in our simulation example, but is highly computationally intensive and relies on the questionable assumption of strong identification of the reduced-form parameters. Furthermore, this test will typically be asymptotically inefficient under strong identification of the structural parameters. Structure of the paper. In Section 2, we discuss how weak identification can arise in DSGE models. Section 3introduces our notation as well as some results from martingale theory; it also discusses the difference between two alternative measures of information. Section 4suggests a test for the full parameter vector. Section 5suggests a test for a hypothesis about a subset of parameters under the assumption that the nuisance parameter is strongly identified. Section 6contains suggestions for applied researchers. Simulations supporting our theoretical results and comparing our procedures to existing alternatives are reported in Section 7. Section 8concludes. Proofs of secondary importance, additional derivations, and further examples can be found in the Supplement. Replication files are also available on the journal website, http://qeconomics.org/supp/ 331/code_and_data.zip. Throughout the rest of the paper, Idkis the k×kidentity matrix, I{·} is the indicator function, [·] stands for the quadratic variation of a martingale, and [··] stands for the joint quadratic variation of two martingales; ⇒denotes weak convergence (convergence in distribution), while p →stands for convergence in probability. 126 Andrews and Mikusheva Quantitative Economics 6 (2015) 2. Weak identification in DSGE models We begin by considering a highly stylized DSGE model that is much simpler than contemporary models designed to fit the data. Unlike most DSGE models used in empirical practice, this model can be solved analytically and allows us to demonstrate how weak identification can arise in a DSGE context. Assume we observe data on inflation πtand a measure of real activity xtfor periods t=1T. Assume that the dynamics of the data are described by the simple DSGE model bEtπt+1+κxt−πt+εt=0 −[rt−Etπt+1−ρat]+Etxt+1−xt=0(1) λrt−1+(1−λ)φππt+(1−λ)φxxt+ut=rt The first equation is a Phillips curve, the second is a linearized Euler equation, and the third is the monetary policy rule. For this section, we assume that the interest rate rtis not observed. The unobserved exogenous shocks atand utare generated by the law at=ρat−1+εat;ut=δut−1+εut (2) (εtεatεut)∼iidN(0Σ);Σ=diagσ2σ2 aσ2 u To solve the model analytically, in this section we make several simplifying assumptions. In particular, we assume that λ=0,φx=0,φπ=1 b,andσ2=0. The model then has six unknown scalar parameters: θ=(bκρδσ2 uσ2 a). In the Supplement, we solve the model (1) under these restrictions to obtain xt πt=⎛ ⎜ ⎝−b b+κ−δb bρ b+κ−ρb −bκ (b +κ−δb)(1−δb) bκρ (b +κ−ρb)(1−bρ) ⎞ ⎟ ⎠ut at =C(θ)ut at As we can see, the observed series xtand πtare weighted sums of two unobserved autoregressive processes with AR coefficients ρand δ, where the weights depend on b and κ. It is relatively easy to see that if 0<b<1,κ>0,σ2 u>0,σ2 a>0,and0<δ<ρ<1, then the six-dimensional parameter θis point-identified. Identification of the model fails when ρ=δ. Indeed, there are two peculiarities in this case: first, if ρ=δ,thenutand atshare the same autoregressive coefficient, and the dynamics of the observed series become insufficiently rich to disentangle the weight functions and separately identify band κ. Second, the 2×2matrix C(θ)becomes degenerate (of rank 1) at ρ=δ. We show in the Supplement that at ρ=δ, the parameter θloses 2degrees of identification. In this case, we can identify only a four-dimensional quantity: the two parameters ρand δ, and the two functions b b+κ−ρb ρ2σ2 a+σ2 uand κ 1−ρb ,but not the parameters b,κ,σ2 a,andσ2 useparately. Quantitative Economics 6 (2015) Maximum likelihood inference 127 If ρ=δ, underidentification precludes us from estimating the parameter θconsistently, and the usual asymptotic theory of maximum likelihood estimation does not apply. Even if ρ=δ, when the difference ρ−δis close to zero, we may have difficulty making reliable statistical inferences. In particular, the finite-sample size of many statistical tests may be quite far from the declared level and many conventional confidence sets may be misleading. To give a concrete example, consider the Wald statistic Wfor testing true hypothesis H0:θ=θ0. According to the usual asymptotic theory of maximum likelihood, if ρ=δ, then as the sample size Tincreases to infinity, the statistic Wconverges in distribution to a χ2 6under H0. If, on the other hand, ρ=δ, this convergence breaks down as the maximum likelihood estimator (MLE) is not consistent. Hence, the limit distribution of Wexperiences a discontinuity at ρ=δ. Since the finite-sample distribution of Wis continuous in the true parameter value, this implies that the convergence of Wto a χ2distribution is not uniform in the parameter ρ−δin a neighborhood of zero. Specifically, the closer ρ−δis to zero, the larger a sample is required to achieve a given accuracy of approximation of the distribution of Wby its asymptotic (χ2 6) limit. This phenomenon is called weak identification. To model the problems arising from weak identification, we can use a weak asymptotic embedding, considering a sequence of models such that ρ=δ+C √T,whereCis a constant and Tis the sample size. It is important to emphasize the conceptual essence of such an embedding: the researcher does not think that the parameters ρand δare changing with the sample size, but rather uses this embedding to obtain asymptotic approximations that reflect the trade-off between the proximity of the parameters ρand δ and the quality of the classical asymptotic approximations. When examining asymptotic behavior along sequences of models with ρ=δ+C √Tas T→∞, we often find that some statistics, like W, have limiting distributions that differ from the χ2limits obtained under classical asymptotic theory. This reflects the sensitivity of those statistics to finiteness of information along some dimensions. If, however, we find a statistic that converges to the same χ2limit even under weak asymptotics, we call such a statistic robust to weak identification. Later, we will show that certain score statistics are robust in this sense. Allowing the true parameter value to drift toward a point of nonidentification as the sample grows is one common way to model weak identification (see Andrews and Cheng (2012) on this), but there are other approaches. Under the approach of Stock and Wright (2000) for weakly identified generalized method of moments (GMM) models, for example, the objective function is modeled as indexed by the sample size and is taken to be asymptotically flat along some directions in the parameter space, thus not providing identification in the limit. This approach is not explicit about what parameter, if any, measures the proximity to identification failure; neither need it assume that there is any point of identification failure in finite samples. To cast the DSGE model discussed above into this framework, suppose for a moment that we know (or calibrate) the true values of ρ=δso that ρand δare excluded from the parameter space. This does not solve the weak identification problem since the sample still contains limited information about the two weak directions if the calibrated values of ρand δare close. At the same time, the model is now point-identified over the whole parameter space. 128 Andrews and Mikusheva Quantitative Economics 6 (2015) The Supplement gives several stylized examples illustrating different types of weak identification that may arise in a DSGE context. In particular, we show how weak identification can arise from insufficiently rich dynamics of the observed process, for example, when autoregressive coefficients for several processes are close to each other or when moving average coefficients nearly cancel with autoregressive roots in an autoregressive moving average (ARMA) process. We also give a examples of a weakly identified vector autoregression (VAR) model and a nonlinear model with a weakly identified regimeswitching mechanism. 3. Martingale methods in maximum likelihood Let XTbe the data available at time T. To allow for the possibility of a weak identification embedding, we consider a so-called scheme of series. In a scheme of series, we assume that we have a series of experiments indexed by the sample size: the data XTof sample size Tare generated by distribution fT(XT;θ0), which may change as Tgrows. In general, we assume that XT=(xT1xTT).LetFTt be a sigma algebra generated by the first tobservations XTt =(xT1xTt). We assume that the log likelihood of the model, T(XT;θ) =logfT(XT;θ) = T  t=1 logfT(xTt|FTt−1;θ) is known up to the k-dimensional parameter θ, which has true value θ0.Wefurtherassume that T(XT;θ) is twice continuously differentiable with respect to θ, and that the class of likelihood gradients {∂ ∂θT(XT;θ):θ∈Θ}and the class of second derivatives {∂2 ∂θ∂θT(XT;θ):θ∈Θ}are both locally dominated integrable. Our main object of study will be the score function ST(θ) =STT(θ) =∂ ∂θT(XTθ)= T  t=1 ∂ ∂θlog fT(xTt|FTt−1;θ) and we take sTt(θ) =STt(θ) −STt−1(θ) =∂ ∂θlogfT(xTt|FTt−1;θ) to denote the increment of the score. Under the assumption that we have correctly specified the model, we have that E(sTt(θ0)|FTt−1)=0almost surely. This in turn implies that for each T,the score taken at the true parameter value, STt(θ0), is a martingale with respect to filtration FTt. This is a generalization of the first informational equality due to Silvey (1961). Similarly, the second informational equality also generalizes to the dependent case. This equality states that we can calculate the (theoretical) Fisher information, IT(θ0), either as the expectation of the negative Hessian of the log likelihood or as the expectation of the outer product of the score. Fisher information plays a key role in the classical asymptotics for maximum likelihood, as it is directly related to the asymptotic variance of the MLE, and the second informational equality suggests two different ways of estimating it that are asymptotically equivalent in the classical context. To generalize the second informational equality to the dynamic context, following Barndorff-Nielsen and Quantitative Economics 6 (2015) Maximum likelihood inference 129 Sorensen (1991), we introduce two measures of information based on observed quantities. The first is the observed information and is equal to the negative Hessian of the log likelihood, IT(θ) =− ∂2 ∂θ∂θT(XT;θ) = T  t=1 iTt(θ) where iTt(θ) =− ∂2 ∂θ∂θlog fT(xTt|FTt−1;θ). The second is the incremental observed information and is equal to the quadratic variation of the score, JT(θ) =ST(θ)= T  t=1 sTt(θ)s Tt(θ) where as before sTt(θ) is the increment of ST(θ). Both observed measures IT(θ) and JT(θ) are unbiased estimates of the (theoretical) Fisher information for the whole sample: IT(θ0)=E(IT(θ0)) =E(JT(θ0)). Using these definitions, let AT(θ) =JT(θ) −IT(θ) be the difference between the two measures of observed information. The second informational equality implies that ATt(θ0)is a martingale with respect to FTt. Specifically, the increment of ATt(θ0)is aTt(θ0)=ATt(θ0)−ATt−1(θ0)=sTt(θ0)s Tt(θ0)−iTt(θ0) and a simple argument gives us that E(aTt|FTt−1)=0almost surely (a.s.). In the classical context, IT(θ0)and JT(θ0)are asymptotically equivalent, which plays a key role in the asymptotics of maximum likelihood. In the independent and identically distributed (i.i.d.) case, for example, the law of large lumbers implies that 1 TIT(θ0)p → −E( ∂2 ∂θ∂θlog f(xtθ0)) =I1(θ0)and 1 TJT(θ0)p →E( ∂ ∂θlogf(xtθ0)∂ ∂θ logf(xtθ0)) = I1(θ0). As a result of this asymptotic equivalence, the classical literature in the i.i.d. context uses these two measures of information more or less interchangeably. The classical literature in the dependent context makes use of a similar set of conditions to derive the asymptotic properties of the MLE, focusing in particular on the asymptotic negligibility of AT(θ0)relative to JT(θ0). For example, Hall and Heyde (1980) show that for θscalar, if the higher order derivatives of the log likelihood are asymptotically unimportant, JT(θ0)→∞a.s. and limsupT→∞JT(θ0)−1|AT(θ0)|<1a.s., then the MLE for θis strongly consistent. If, moreover, JT(θ0)−1IT(θ0)→1a.s., then the MLE is asymptotically normal and JT(θ0)1/2(ˆ θ−θ0)⇒N(01). We depart from this classical approach in that we consider weak identification. We find that in weakly identified models, the difference between our two measures of information is important and AT(θ0)is no longer negligible asymptotically compared to observed incremental information JT(θ0). Example 1. To illustrate this nonequivalence in a simple example, suppose we observe data Ytfor t∈{1T}, generated by the model Yt=(π +β)Yt−1+et−πet−1e t∼iidN(01) (3) 130 Andrews and Mikusheva Quantitative Economics 6 (2015) The true value of the parameter θ0=(β0π0)satisfies the restrictions |π0|<1, β0= 0,and|π0+β0|<1, which guarantee that the process is stationary and invertible. For simplicity we assume that Y0=0and e0=0, though the initial condition will not matter asymptotically. One can rewrite the model as (1−(π +β)L)Yt=(1−πL)et. It is easy to see that if β0=0, then the parameter πis not identified. Andrews and Cheng (2012) modeled weak identification using the drifting parameter value β0=C √T, leading to the parameter πbeing weakly identified. Consider the normalization matrix KT=diag(1/√T1).Then KTJT(θ0)K T p →Σand KTIT(θ0)K T⇒Σ+0ξ ξCη  where Σis a positive-definite matrix, while ξand ηare two Gaussian random variables.2 As we can see, the difference between the two information matrices is asymptotically nonnegligible compared with the information measure JT(θ0). The Supplement contains several examples of weakly identified models. For all of them, we observe the same phenomenon: the appropriately normalized quadratic variation of the score JTconverges in probability to a positive-definite matrix, while the Hessian normalized in the same way converges weakly to a random matrix. White (1982) shows that the two measures of information may differ if the likelihood is misspecified. As our examples show, even if the model is correctly specified these two measures may differ substantially if identification is weak. This result is quite different from that of White (1982). In particular, correct specification implies that EAT(θ0)=0, and it is this restriction that is tested by White’s information matrix test. In contrast, weak identification in correctly specified models is related to AT(θ0)being substantially volatile relative to JT(θ0)while maintaining the assumption that EAT(θ0)=0. Correct specification can still be tested by comparing the realized value of AT(θ0)to the metric implied by a consistent estimator of its variance. One may potentially create a test for weak identification based on a comparison of ATwith JT, though this is beyond the scope of the present paper. We will, however, treat nonpositive-definiteness of the Hessian as an informal sign of weak identification. 4. Test for full parameter vector In this section, we suggest tests for a simple hypothesis on the full parameter vector, H0:θ=θ0, which are robust to weak identification. We introduce our first assumption. Assumption 1. Assume that there exists a sequence of constant matrices KTsuch that (a) for all δ>0,T t=1E(KTsTt(θ0)I{KTsTt(θ0)>δ}|FTt−1)→0, (b) T t=1KTsTt(θ0)sTt(θ0)K T=KTJT(θ0)K T p →Σ,where Σis a constant positivedefinite matrix. 2Details can be found in the Supplement. Quantitative Economics 6 (2015) Maximum likelihood inference 137 Table 1. True parameter values for simulations. φxφπλρ δκσ aσuσ Calibrated value 228 202 0898 085 0103 010325 0265 0556 Parameter space lower bound 000−099 −099 0 0 0 0 Parameter space upper bound 10 10 099 099 099 1 1 1 1 econometrician is concerned with inference on the remaining nine parameters θ= (φxφπλρδκσaσuσ). Note that unlike in Section 2, here we take the interest rate rtto be observable and do not restrict the parameters other than b. For our simulation exercise, we draw samples from the model with parameters calibrated to ML estimates obtained using demeaned U.S. macro data from Smets and Wouters (2007). The ML estimate of the parameter ρis very close to 1, so since robustness to unit roots lies beyond the scope of the present paper, for our simulations we will instead use the smaller value ρ=085. Likewise, the ML estimate for κlies quite close to 0, which is the boundary of the parameter space for this parameter. To ensure that parameter-on-the-boundary issues do not greatly affect the distribution of classical test statistics, we increase the value of this parameter, taking κ=01. The baseline values of parameters used in the simulations are reported in Table 1.Thestructuralparameters are point-identified at this parameter value. We generate samples of size 300 from this model and then discard the first 100 observations, using only the last 200 for the remainder of the analysis. 7.1 Properties of classical ML testing We begin by examining the behavior of the classical maximum-likelihood-based statistics. Histograms for the ML estimator6show that the marginal distributions of the estimates for several parameters depart substantially from a normal distribution. We consider four variations on the Wald statistic for testing the simple hypothesis H0:θ=θ0, where θ0is the true value, corresponding to different estimators of the asymptotic variance, ˆ V, used in the quadratic form (ˆ θ−θ0)ˆ V−1(ˆ θ−θ0). In particular, Wald (IT(ˆ θ)) uses the inverse of the observed information, evaluated at ˆ θ, to estimate the asymptotic variance. Wald (IT(θ0)), on the other hand, evaluates the observed information at the true parameter value. Likewise, Wald (JT(ˆ θ))andWald(JT(θ0))useJ−1 Tas the estimator of the asymptotic variance, calculated at ˆ θand θ0, respectively. Under the usual strong identification assumptions for ML, all of these statistics should have a χ2 9distribution asymptotically. In simulation, however, the distribution of these statistics appears quite far from a χ2 9. Table 2lists sizes for nominal 5% and 10% tests (based on 2500 simulations), and shows that all versions of the Wald test we consider severely overreject. Taken together, these results strongly suggest that the usual approaches to ML estimation and inference are poorly behaved when applied to this DSGE model. 6Available from the authors by request. 138 Andrews and Mikusheva Quantitative Economics 6 (2015) Table 2. Simulated size of Wald tests for the nine-dimensional hypothesis H0:θ=θ0;basedon 2500 simulations. Wald (IT(θ0))Wald(IT(ˆ θ))Wald(JT(θ0))Wald(JT(ˆ θ)) Size of 5% test 3916% 4236% 4024% 4044% Size of 10% test 432% 4744% 4572% 4588% 7.2 Behavior of the information matrix In Section 3, we associated weak identification with the difference between two information measures AT(θ0)being large compared to JT(θ0). Note that observed incremental information JT(θ0)is an almost surely positive-definite matrix by construction, while AT(θ0)is a mean-zero random matrix. If AT(θ0)is negligible compared to JT(θ0), then the observed information IT(θ0)=JT(θ0)−AT(θ0)will be positive-definite for almost all realizations of the data. We can check positive-definiteness of IT(θ0)directly in simulations. Considering the observed information evaluated at the true value, we find that it has at least one negative eigenvalue in over 47% of simulation draws (based on 2500 simulations). While this falls far short of a formal test for weak identification, it is consistent with the idea that weak identification is the source of the poor behavior of ML estimation in this model. In line with the conjecture discussed above that the persistence and variance parameters may be well identified if we know the structural parameters (φxφπλκ), we find that the observed information for the five parameters (ρδσaσuσ)alone is positive-definite in all simulation draws. 7.3 Size of the LM tests We now turn to the weak identification-robust statistics discussed earlier in this paper. Under appropriate assumptions, we have that LMo(θ0)⇒χ2 9and LMe(θ0)⇒χ2 9, where LMo(θ0)is the LM statistic using the observed incremental information JT(θ0) and LMe(θ0)is calculated with the theoretical Fisher information IT(θ0). In Figure 1, we plot the cumulative distribution functions (CDFs) of the simulated distributions of LMo(θ0)and LMe(θ0)together with a χ2 9. Table 3reports the size of the LM tests. Two points are clear from these results: first, though our tests based on the LM statistics are not exact, the χ2approximation is very good for LMeand reasonable for LMo.Second, the LMestatistic has somewhat better finite-sample properties. We next consider the size of the two LM statistics for testing subsets of parameters. Specifically, as before, we consider a partition of the parameter vector, θ=(αβ),and consider the problem of testing H0:β=β0, treating αas a nuisance parameter. As discussed in Section 5, an important issue is whether the nuisance parameter αis weakly or strongly identified. While we are unaware of any formal tests for identification strength in DSGE models that ensure size control when used as pretests, there is a common perception that, fixing structural parameters like φx,φπ,λ,andκ, the parameters controlling the persistence and the variance of shocks will be well identified. Since this is consistent with our results from comparing different information measures, we treat these parameters as strongly identified. Quantitative Economics 6 (2015) Maximum likelihood inference 139 Figure 1. CDF of simulated LM statistics introduced in Theorem 1compared to χ2 9. Table 3. Simulated size (based on 1000 simulations) of a test for the full parameter vector and for five tests of composite hypotheses H0:β=β0, treating all other parameters as nuisance parameters. LMoLMe Tested Parameters 5% 10% 5% 10% All parameters 9% 153% 45% 91% (∗1) β=(φxφπλκ) 59% 111% 44% 81% (∗2) β=(φxφπλκρ) 63% 115% 51% 91% (∗3) β=(φxφπλκδ) 59% 116% 41% 86% (∗4) β=(φxφπλκσa)59% 108% 44% 82% (∗5) β=(φxφπλκσu)59% 112% 40% 81% (∗6) β=(φxφπλκσ) 72% 13% 49% 96% Note: Statistic LMorefers to the LM test using observed incremental information and statistic LMeuses theoretical Fisher information, and in both cases we plug in the restricted MLE for nuisance parameters. We consider testing six different composite hypotheses (corresponding to cases (∗1)–(∗6) in Table 3): a hypothesis on the four structural parameters (φxφπλκ) and five hypotheses on these four parameters plus each of the other five parameters taken one at a time, (φxφπλκρ),(φxφπλκδ), and so forth. In each case we follow the approach discussed in Section 5and plug in the restricted MLE for the parameters not under test, reducing the critical value appropriately. Our simulation results, reported in Table 3, are consistent with the assumption that the parameters (ρδσaσuσ) are strongly identified. In particular, we see that all the tests we consider for composite hy- 140 Andrews and Mikusheva Quantitative Economics 6 (2015) Table 4. 95% LMoconfidence intervals for parameters based on single draw of simulated data, where we treat the parameters (ρδσaσuσ)as well identified and project over the other parameters. Level φxφπλρδκσ aσuσ Lower 104 058 076 074 004 0 028 024 046 Upper 997 888 097 091 046 018 050 034 056 potheses control size fairly well, though the size control of the LMetestsisagainsomewhat better. 7.4 Calculation of confidence sets Despite weak identification, we can produce informative confidence sets. To illustrate this point, we take one random draw from the model, treat it as a sample, and report LMoconfidence intervals for each of our nine parameters separately in Table 4. To calculate these one-dimensional confidence intervals, we follow the approach discussed in Section 6, and first form four- and five-dimensional confidence sets by inverting the LMotests for the six composite hypotheses corresponding to (∗1)–(∗6) in Table 3, that is, for each group of parameters, we collect all values of β0such that the corresponding hypotheses H0:β=β0are not rejected. For example, in case (∗1), we construct a joint four-dimensional confidence set for parameters (φxφπλκ). For each group of tested parameters, we take 5·104draws uniformly at random over the parameter space for βformed by the Cartesian product of the one-dimensional parameter spaces given in Table 1, and keep those draws that are not rejected by the LMotest that plugs in the restricted MLE for the nuisance parameters (all parameters other than β). By projecting the (four-dimensional) convex set obtained for the case (∗1) on the subspace corresponding to each parameter separately, we obtain one-dimensional confidence sets for each of the parameters φx,φπ,λ,andκ. To obtain one-dimensional confidence sets for the remaining five parameters ρ,δ,σa,σu,andσ, we project the corresponding fivedimensional confidence sets obtained for cases (∗2)–(∗6) on the subspace corresponding to the parameter of interest. We can see that while the confidence intervals for many parameters are wide, in all instances they exclude some values and in most cases they cover only a small portion of the parameter space. 7.5 Alternative weak identification-robust methods Issues of weak identification in DSGE models have recently attracted the attention of econometricians, and several weak identification-robust methods for DSGE models have been suggested independently by Dufour, Khalaf, and Kichian (2013) (DKK), Guerron-Quintana, Inoue, and Kilian (2013)(GQIK),andQu(forthcoming). It is important to note that DKK and Qu focus primarily on testing the full parameter vector, while GQIK allow one to concentrate out strongly identified nuisance parameters. None of the competing papers offers procedures to determine which specific parameters are Quantitative Economics 6 (2015) Maximum likelihood inference 141 strongly identified. They all use projection for testing with weak nuisance parameters or parameters whose identification strength is unknown. Our method differs from the three approaches mentioned above in that it is valid in a general ML framework with potentially weak identification and is not restricted to log-linearized DSGE models. The LM statistics we propose can be used whenever we can evaluate the likelihood function. In contrast, the three approaches above are specially designed for log-linearized DSGE models that can be written as linear expectation equations. In general, these methods cannot be applied to the nonlinear DSGE models that are increasingly popular; see, for example, Fernández-Villaverde and Rubio- Ramírez (2011). Though the range of nonlinear DSGE models for which one can differentiate the likelihood function is quite limited at present, the number of such models is growing; see, for example, Amisano and Tristani (2011). The method closest to ours is the LM test suggested by Qu (forthcoming)forloglinearized DSGE models with normal errors. Qu (forthcoming) notices that in large samples, the Fourier transforms of the observed data at different frequencies are approximately independent Gaussian random variables with variance equal to the spectrum of the observed series; this allows him to write an approximate likelihood for the data in a very elegant way and to discuss the properties of the likelihood analytically. His statistic is almost the same as our statistic LMe(θ0)for testing the full parameter vector, the main difference being that Qu (forthcoming) uses an approximate likelihood, while we use the exact likelihood. Hence, we expect that the two statistics applied to a log-linearized DSGE model with normal errors should be very close provided Qu’s approximate likelihood is well behaved. GQIK consider models with a linear state-space representation and assume that the coefficients of the state-space representation, Υ=Υ(θ), are either strongly identified or not identified at all, while no assumption is made on the identification of the structural parameters θ. For testing a hypothesis H0:θ=θ0about the structural parameter vector, GQIK suggest testing the hypothesis  H0:Υ=Υ(θ0)about the reduced-form parameter, using the classical LR statistic and the usual χ2critical values with degrees of freedom equal to the dimensionality of the identified reduced-form parameter. The assumption of GQIK that the reduced-form parameters are strongly identified seems quite problematic in some DSGE applications and no test is available to check it. Schorfheide (2010) provides an example in which weak identification of the structural parameters leads to weakly identified reduced-form parameters. Unlike the tests suggested in this paper, the LR test proposed by GQIK is typically asymptotically inefficient under strong identification, since the dimension of the reduced-form parameter is usually higher than that of the structural parameter. GQIK also suggest a test based on Bayes factors, which we do not discuss here as it is less directly comparable to our approach. DKK propose a limited information approach based on a set of exclusion restrictions implied by a system of linear expectation equations, which they then test using a seemingly unrelated regression-based (SUR-based) F-statistic in the spirit of Stock and Wright (2000). Advantages of this approach are that a researcher has the freedom to choose which restrictions he or she wishes to use for inference, and that it does not require distributional assumptions on the error term and hence is robust to misspecifica- 142 Andrews and Mikusheva Quantitative Economics 6 (2015) tion. A disadvantage of the method is its limited ability to accommodate latent state variables. Furthermore, this limited information test may be expected to have lower power than full-information methods if the model is correctly specified. DKK also suggested a full-information ML method based on a VAR approximation to the DSGE solution, but the authors seem to prefer and advocate their limited information approach, so we focus on this method. 7.6 Power comparisons with alternative methods Here we compare the power of the alternative approaches to that of the proposed LM tests. As the alternative approaches deal primarily with testing the full parameter vector, we will focus on this case. Table 5reports actual size, while Figure 2shows (non-size-corrected) power curves for 5% tests based on the statistics LMo(θ0)and LMe(θ0),aversionofQu’s(forthcoming) LM test, the LR test introduced in GQIK, and the limited information (LI) test of DKK. Implementation details are discussed below. Power is calculated for alternatives that entail a change in one element of the parameter vector while the other elements remain at their null values. The label on each subplot denotes the parameter whose value changes under the alternative. First, we consider Qu’s (forthcoming) frequency-domain LM test.7Initial simulations showed that this test tended to overreject at some parameter values and that the degree of overrejection seemed to be related to how close ρwas to 1. At our baseline parameter value, a nominal 5% test based on Qu’s approach had size of approximately 8%, but if we increased ρto 09or 095, we obtained size of approximately 15% and 33%, respectively. While the tests proposed in this paper are not robust to unit roots, they did not show similar sensitivity to the choice of ρand had roughly the same size for a wide range of values for ρ. Qu suggested that the size distortions of the frequency-domain LM test were due to bias in the periodogram, and proposed a prewhitened version of his test that resolves these size issues in our context.8Our power simulations focus on this prewhitened (PW) test, which we call Qu’s PW LM test. Table 5. Simulated test size for the full parameter vector (number of simulations is 1000). Level LMo(θ0)LMe(θ0)Qu LM Qu PW LM GQIK DKK 5% 9% 45% 84% 48% 67% 64% 10% 153% 91% 136% 86% 118% 115% 7Qu’s test allows one to test hypotheses using only a subset of frequencies, if desired. For comparability with the other tests studied, we focus on results obtained using the whole spectrum. 8In private correspondence with the authors. The prewhitening procedure consists of simulating a long sample under the null and fitting a VAR(1)model to this simulated data. Letting Abe the matrix of VAR coefficients and XTbe the T×3matrix of data, one then applies Qu’s approach using the transformed data YT=XT(Id3−A·L),whereLdenotes the lag operator. Correspondingly, in all later expressions, the spectral density fθ(ω) is replaced by gθ(ω) =(Id3−A·exp(−iω))fθ(ω)(Id3−A·exp(−iω))∗,whereM∗ denotes the conjugate transpose of M. Quantitative Economics 6 (2015) Maximum likelihood inference 143 Figure 2. Power functions for 5% tests of the null hypothesis H0:θ=θ0for the following statistics: LMe(θ0),LM o(θ0), prewhitened version of Qu’s test, GQIK LR, and DKK test with Newey–West covariance matrix. Power is calculated based on 500 simulations. We find that the power function for Qu’s PW LM test is nearly indistinguishable from the power function for the LM statistic LMe(θ0)based on theoretical Fisher information. On the one hand, this may seem surprising, since the non-prewhitened version of Qu’s test had behavior (in particular, size) that differed substantially from that of the LMetest. On the other hand, Qu’s statistic has the same form as LMe(θ0)but is calculated with an 144 Andrews and Mikusheva Quantitative Economics 6 (2015) approximate likelihood while LMe(θ0)is calculated with the exact likelihood. The discrepancy between Qu’s original LM test and the LMetest is thus due to the difference between the approximate likelihood and the true likelihood. Insofar as the quasi-likelihood based on the prewhitened data offers a better approximation to the true likelihood, one would expect the behavior of the prewhitened LM test to be closer to that of LMe.Consistent with this interpretation, the correlation between the prewhitened version of Qu’s statistic and LMe(θ0)under the null is 09. In the GQIK LR approach, rather than testing a hypothesis about the ninedimensional structural parameter H0:θ=θ0, one instead tests a hypothesis about the reduced-form parameter (i.e., the coefficients of the state-space representation)  H0:Υ=Υ(θ0)using the LR statistic. While simulating GQIK’s method, we encountered several difficulties. First, it is not obvious how many degrees of freedom to use. Examining the solution of the model, we noticed that matrices of the state-space representation have numerous zeros. We imposed these zeros, which left us with 28 nonzero reducedform parameters. However, the effective dimensionality of the reduced-form parameter space is lower since some values of the reduced parameters are observationally equivalent. Hence, we used degrees of freedom equal to the rank of the Fisher information with respect to the state-space coefficients evaluated under the null, which leads us to think that the (local) dimensionality of the reduced-form parameter space is 18. The second difficulty is that computing the GQIK LR statistic is numerically very involved and time consuming, as noted by GQIK in their paper. To test a hypothesis on the full parameter vector, one must solve a high-dimensional nonlinear optimization problem, while no optimization is required for the other methods discussed here. From Figure 2, one can see that the GQIK test gives us power comparable to the LM tests for all considered alternatives. For the test of DKK, we consider the transformation of the data ξπt =bπt+1+κxt−πt ξxt =˜ ξxt −ρ˜ ξxt−1 ξrt =˜ ξrt −δ˜ ξrt−1 where ˜ ξxt =−[rt−πt+1]+xt+1−xt; ˜ ξrt =λrt−1+(1−λ)φππt+(1−λ)φxxt−rt The transformed data (ξπt ξxtξrt )comprise a linear combination of the uncorrelated structural error terms (εtεatεut)and the expectation errors Etπt+1−πt+1, Et−1πt−πt,Etxt+1−xt+1,andEt−1xt−xt. We base the test on the exclusion restriction that (ξπtξxtξrt )are not predictable by the instruments Yt−1=(πt−1xt−1rt−1).Itis easy to see that (ξπtξxtξrt )follows a (moving average) MA(1)process and hence that the heteroskedasticity and autocorrelation robust (HAC) formulation of DKK should be used. We calculate the DKK test using the Newey–West HAC estimator for the long-run covariance matrix (using three lags). DKK formulate the null in such a way that variances Quantitative Economics 6 (2015) Maximum likelihood inference 145 of the shocks do not enter, and the test is not supposed to have power against alternatives that differ only in these parameters. Hence, we do not depict the corresponding power functions. Based on Figure 2, the DKK test in our context is significantly less powerful than the other tests considered and has nearly flat power curves in the neighborhoods where the LM tests achieve almost 100% power. Power simulations on larger neighborhoods show that the DKK test has nontrivial power against some alternatives, but confirm that for the null and alternatives considered, it has substantially less power than the other tests we study. This lower power is to be expected given the limited-information nature of the test, and may be a reasonable price to pay for robustness to misspecification. 8. Conclusion This paper studies the problem of weak identification in DSGE models and explores how weak identification can arise in several examples. We show that two forms of the LM statistic may be used to construct robust tests for hypotheses about the full parameter vector, as well as hypotheses about subvectors of parameters for which the nuisance parameter is strongly identified. How to determine whether the nuisance parameter is strongly identified is an open question. We give suggestive evidence that the discrepancy between two measures of information may serve as an indication of weak identification, but further exploration of this issue is an important topic for future research. Appendix:Proofs We denote by superscript 0quantities evaluated at θ0=(α 0β 0).IntheTaylorexpansions used in the proofs, the expansion is assumed to be for each entry of the expanded matrix. Proof of Lemma 1. The proof follows closely the argument of Bhat (1974), starting with the Taylor expansion 0=Sα(ˆαβ0)=S0 α−I0 αα(ˆα−α0)−Iααα∗β0−I0 αα(ˆα−α0) where α∗is a convex combination of ˆαand α0. We may consider different α∗for different rows of Iαα. Assumption 2(b) helps to control the last term of this expansion, while Assumption 2(a) allows us to substitute JααT for IααT in the second term. Assumption 1 gives the central limit theorem (CLT) for KαT SαT . Lemma 2. Let MT=T t=1mtbe a multidimensional martingale with respect to sigma field Ftand let [X]tbe its quadratic variation.Assume that there is a sequence of diagonal matrices KTsuch that MTsatisfies the conditions of Assumption 3.Let mit be the ith component of mtand let KiT be the ith diagonal element of KT.For any i,j,l, KiT KjT KlT T  t=1 mit mjtmlt p →0 146 Andrews and Mikusheva Quantitative Economics 6 (2015) Proof of Theorem 2. For simplicity of notation, we assume in this proof that Cij =C for all i,j. The generalization of the proof to the case with different Cij ’s is obvious but tedious. According to the martingale CLT, Assumption 3implies that KαT S0 αKβT S0 βKαβT vecA0 αβ⇒(ξαξβξαβ ) (8) where the ξ’s are jointly normal with variance matrix ΣM. We Taylor expand Sβj(ˆαβ0),thejth component of vector Sβ(ˆαβ0), keeping in mind that I0 βjα=− ∂2 ∂βj∂α (α0β0), and receive KβjT Sβj(ˆαβ0)=KβjT S0 βj−KβjT I0 βjα(ˆα−α0) +1 2KβjT (ˆα−α0)I0 ααβj(ˆα−α0)+˜ Rj with residual ˜ Rj=KβjT 1 2(ˆα−α0)I∗ ααβj−I0 ααβj(ˆα−α0) where I0 ααβj=∂3 ∂α∂α∂βj(α0β0),I∗ ααβj=∂3 ∂α∂α∂βj(α∗β0),andα∗is a point between ˆα and α0. From Assumption 2(c), we have that K−1 αT |ˆα−α0|=Op(1). As a result, Assumption 4(c) makes the Taylor residual negligible: KβjT Sβj(ˆαβ0)=KβjT S0 βj−KβjT I0 βjα(ˆα−α0) +1 2KβjT (ˆα−α0)I0 ααβj(ˆα−α0)+op(1) We plug asymptotic statement (7) into this equation and get KβjT Sβj(ˆαβ0)=KβjT S0 βj−KβjT I0 βjαI0 αα−1S0 α +1 2KβjT S0 αI0 αα−1I0 ααβjI0 αα−1S0 α+op(1) Recall that by definition I0 βα =J0 βα −A0 βα. We use this substitution in the equation above and receive KβjT Sβj(ˆαβ0)=KβjT S0 βj−KβjT J0 βjαI0 αα−1S0 α+KβjT A0 βjαI0 αα−1S0 α(9) +1 2KβjT S0 αI0 αα−1I0 ααβjI0 αα−1S0 α+op(1) One can notice that we have the informational equality I0 ααβj=−A0 ααS0 βj−A0 αβjS0 α−S0 αA0 αβj+2 T  t=1 sαt s αt sβjt +Λααβj(10)