A machine learning projection method for macro-finance models
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Valaitis, Vytautas; Villa, Alessandro T. Article A machine learning projection method for macro-finance models Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Valaitis, Vytautas; Villa, Alessandro T. (2024) : A machine learning projection method for macro-finance models, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 15, Iss. 1, pp. 145-173, https://doi.org/10.3982/QE1403 This Version is available at: https://hdl.handle.net/10419/296356 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 15 (2024), 145–173 1759-7331/20240145 A machine learning projection method for macro-finance models Vytautas Valaitis School of Economics, University of Surrey Alessandro T. Villa Economic Research Department, Federal Reserve Bank of Chicago We use supervised machine learning to approximate the expectations typically contained in the optimality conditions of an economic model in the spirit of the parameterized expectations algorithm (PEA) with stochastic simulation. When the set of state variables is generated by a stochastic simulation, it is likely to suffer from multicollinearity. We show that a neural network-based expectations algorithm can deal efficiently with multicollinearity by extending the optimal debt management problem studied by Faraglia, Marcet, Oikonomou, and Scott (2019) to four maturities. We find that the optimal policy prescribes an active role for the newly added medium-term maturities, enabling the planner to raise financial income without increasing its total borrowing in response to expenditure shocks. Through this mechanism, the government effectively subsidizes the private sector during recessions. Keywords. Machine learning, incomplete markets, projection methods, optimal fiscal policy, maturity management. JEL classification. C63, D52, E32, E37, E62, G12. 1. Introduction In this paper, we exploit the computational gains that derive from the robustness to multicollinearity of neural networks to extend the optimal debt management problem studied by Faraglia et al. (2019) to four maturities. The hedging benefits provided by the additional maturities allow the government to respond to expenditure shocks by raising financial income without increasing the total outstanding debt. Through this mechanism, the government effectively subsidizes the private sector in recessions. Vytautas Valaitis: [email protected] Alessandro T. Villa: [email protected] We are thankful to Andrea Lanteri, Lukas Schmid, and Matthias Kehrig for their encouragement. The paper received the Student Award of the Society of Computational Economics and benefited from comments by Albert Marcet, Serguei Maliar, Lilia Maliar, Swapnil Singh, and seminar participants at the Duke Macro Breakfast, the 2018 Baltic Economic Conference, the Society for Computational Economics 24th International Conference, the 2018 Econometric Society Summer European Meeting, and the 2019 Econometric Society African meeting. Disclaimer: The views expressed in this paper do not represent the views of the Federal Reserve Bank of Chicago or the Federal Reserve System. Declaration of conflicts of interest: none. ©2024 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1403
146 Valaitis and Villa Quantitative Economics 15 (2024) We use a neural network (NN) in a supervised machine learning fashion to approximate the expectation terms typically contained in the optimality conditions of an economic model, in the spirit of the Parameterized Expectations Algorithm (PEA) with stochastic simulation, introduced by den Haan and Marcet (1990) and in a similar fashion to Duffy and McNelis (2001). On the one hand, stochastic simulation methods allow us to tackle problems with a high number of state variables, since they calculate solutions only in the states that are visited in equilibrium (i.e., the ergodic set). On the other hand, when the set of state variables is generated by a stochastic simulation, it is likely to suffer from multicollinearity. In this context, this paper makes two contributions. First, we show that an NN-based expectations algorithm can deal efficiently with multicollinearity by extending the optimal debt management problem studied by Faraglia et al. (2019) to four maturities. Second, we show that the optimal debt management policy prescribes an active role for the medium-term maturities, enabling the planner to raise financial income without increasing its total borrowing in response to expenditure shocks. We consider this problem a particularly interesting economic application that also poses significant computational challenges for four reasons. First, the number of state variables increases in the number and length of maturities available. Second, this class of problems includes forward-looking constraints, and the problem can be made recursive at the cost of adding even more state variables. Following Marcet and Marimon (2019), we formulate the recursive Lagrangian to solve for the time-inconsistent optimal policy under full commitment with multiple maturities. When markets are incomplete, the Ramsey planner needs to keep track of all promises made in the previous periods. Because of these reasons, optimal maturity management problems suffer from the curse of dimensionality (see Bellman (1961)). For example, the optimal debt management problem with four maturities considered in Section 4features 46 state variables. Third, because of the maturities, many of these state variables are multicollinear when the model is solved by using a stochastic simulation approach.1 Fourth, this class of problems does not have a stochastic steady state, as documented in Aiyagari, Marcet, Sargent, and Sappala (2002), and tends to frequently hit the borrowing and lending constraints. Such properties render the model particularly hard to solve using perturbation methods around a particular point.2 In Section 4, we use the methodology to study the optimal government debt management policy when the Ramsey planner can issue an increasing number of debt in1The state space includes lagged values of the same variables (e.g., lagged values of outstanding bonds and Lagrange multipliers). Multicollinearity in the state space might prevent standard regression-based algorithms from converging because the estimated regression coefficients may never stabilize due to high estimation variance and because misspecification of the true policy function under multicollinearity may lead to severe prediction bias, as we show in Section 2.5. Alternatively, people have used the stochastic simulation based on regularization (see Judd, Maliar, and Maliar (2011)) or have extended the PEA algorithm to condensed PEA; see Faraglia et al. (2019). In Section 5, we discuss how the NN-based expectations algorithm improves upon these methods. 2Bhandari, Evans, Golosov, and Sargent (2017b) propose a method that allows one to approximate a system around a current level of government debt, and Lustig, Sleet, and Yeltekin (2008) on the other hand, solve the optimal fiscal policy problem in incomplete markets with seven maturities up to 7 periods using value function iteration on a sparse grid.
Quantitative Economics 15 (2024) A machine learning projection method 147 struments with different maturities. Intuitively, the prices of longer maturities are typically more responsive to shocks than prices of shorter maturities. This differential response creates opportunities for hedging by borrowing in long-term and saving in shortterm bonds. In this case, the value of liabilities falls by more than the value of assets in response to negative shocks (see Angeletos (2002), Buera and Nicolini (2004)and Faraglia et al. (2019)). Additionally, the fact that short bond prices are not as responsive to shocks allows the planner to smooth the price of new debt issuance by rebalancing the portfolio toward the longer maturities in economic booms and toward the shorter maturities in recessions. We find that the planner actively uses the additional medium-term maturities to exploit both the hedging and the price smoothing benefits. The government holds leveraged positions in all bonds and rebalances the portfolio with more emphasis on the shorter maturities in recessions. We find that, when the number of available maturities increases from two to three (and four), the total amount of outstanding debt becomes procyclical. The additional maturities allow the government to respond to expenditure shocks by raising financial income without increasing the total outstanding debt. Through this mechanism, the government effectively subsidizes the private sector in recessions, resulting in higher leisure and less volatile labor taxes. 1.0.0.1 Literature review This paper contributes to two strands of literature: (i) numerical methods in economics and (ii) optimal fiscal policy. In terms of methods, this paper builds on the seminal work of den Haan and Marcet (1990), who introduced PEA. The idea of using neural networks to parameterize decision rules in a similar fashion to PEA goes back to Duffy and McNelis (2001). Our paper contributes to this literature by showing that an NN-based expectations algorithm can deal efficiently with multicollinearity by extending the optimal debt management problem studied by Faraglia et al. (2019) to more than two maturities. In particular, we exploit the computational gains to study the optimal government debt management problem of Faraglia et al. (2019) with three and four maturities, which yields new economic insights. Note that PEA has been extended more recently (see Faraglia, Marcet, Oikonomou, and Scott (2014)andFaraglia et al. (2019)) to deal with multicollinearity (condensed PEA) and overidentification (Forward-States PEA). Our methodology builds on condensed PEA and Forward-States PEA, in the context of optimal fiscal policy, allowing for machine learning to reduce the state space endogenously and handling multicollinearity effectively when a stochastic simulation approach is adopted. In contrast, condensed PEA achieves this result by introducing an external loop that tests a subset of the state space as a candidate to solve the model. In contemporaneous work, Maliar, Maliar, and Winant (2021)andMaliar and Maliar (2022) discuss how neural networks can handle multicollinearity. In particular, they do so in the context of the Krusell and Smith (1998) model. We complement their work by demonstrating the robustness to multicollinearity in the context of optimal fiscal policy. Additionally, we show that the interaction between the capability of a neural network to deal with multicollinearity and its flexibility in approximating generic policy functions plays an important role in generating unbiased predictions. In Section 2.5, we show that if a researcher precommits to approximate the policy functions with polynomials that
148 Valaitis and Villa Quantitative Economics 15 (2024) are misspecified then, under multicollinearity among the state variables, the predictions will be biased. Thanks to the flexible nonparametric nature of a neural network, which does not require making ex ante assumptions about the functional form of the policy functions, this problem disappears. PEA can potentially be used in combination with other standard econometric techniques that tackle the problem of multicollinearity, as in Judd, Maliar, and Maliar (2011). Similar to our paper, Judd, Maliar, and Maliar (2011) adopt a stochastic simulation approach and show how already established methods in econometrics can be used to alleviate the multicollinearity problem using a multicountry neoclassical growth model. We discuss the relation between our method and the methods of Faraglia et al. (2019)and Judd, Maliar, and Maliar (2011) in greater detail in Section 5. Other papers that use machine learning to solve economic models include Scheidegger and Bilionis (2019), Azinovic, Gaegauf, and Scheidegger (2022), FernándezVillaverde, Hurtado, and Nuño (2023), and Duarte, Duarte, and Silva (2023). FernándezVillaverde, Hurtado, and Nuño (2023) use deep neural networks to approximate the aggregate laws of motion in a heterogeneous agents model featuring strong nonlinearities and aggregate shocks. Duarte, Duarte, and Silva (2023) casts the economic model in continuous time and uses neural networks to approximate the Bellman equation. Maliar, Maliar, and Winant (2021)andAzinovic, Gaegauf, and Scheidegger (2022)approximate all the model equilibrium conditions using neural networks and use the simulated data to train them. Azinovic, Gaegauf, and Scheidegger (2022) solve the life-cycle model with borrowing constraints, aggregate shocks, and financial frictions using unsupervised machine learning. The main difference of our paper is to leverage on supervised machine learning to deal effectively with the problem of multicollinearity typical of stochastic simulation approaches. In this context, we show how our algorithm can alleviate the curse of dimensionality, allowing us to explore the problem of the optimal maturity structure of government debt in a more realistic environment. Our application also contributes to the strand of literature on optimal fiscal policy. In particular, it is relevant to the literature on the optimal maturity structure of government debt.3 Lustig, Sleet, and Yeltekin (2008) find that the optimal policy prescribes an almost exclusive role to the longest maturity in a model with no-lending constraints and a New Keynesian model where bonds are nominal. In our setting, we allow for government lending and study the hedging benefits of a choice between multiple maturities of real bonds. Bhandari et al. (2017b) study the optimal maturity structure in an open economy with two maturities, and Bigio, Nuño, and Passadore (2023) allow for a finite number of maturities in an economy with liquidity costs of issuing debt, where liquidity costs differ by maturity. Faraglia et al. (2019) is the closest paper to ours and studies the role of frictions in a closed economy with two types of bonds. Solving the Ramsey problem considered in this paper is particularly challenging, as the dimension of its state space increases significantly in function of the length of the maturities and the number of bonds. Moreover, this class of problem includes forward-looking constraints, 3Aiyagari et al. (2002), Angeletos (2002), Buera and Nicolini (2004), Lustig, Sleet, and Yeltekin (2008), Faraglia et al. (2019), Bhandari et al. (2017b), and Bigio, Nuño, and Passadore (2023).
Quantitative Economics 15 (2024) A machine learning projection method 149 so the commonly used recursive representation can not be adopted. Marcet and Marimon (2019) provide an alternative formulation to solve for the time-inconsistent optimal contract under full commitment: a recursive Lagrangian or saddle-point functional equation. The solution involves adding even more state variables to the original problem. These additional state variables, necessary to recursify the problem, create history dependence. In this context, we use our methodology to extend the literature to study optimal debt management with three and four maturities in a closed economy. We find that the optimal policy prescribes an active role for the medium-term bonds. The additional maturities enable the planner to raise financial revenue without increasing the total outstanding debt, in response to a positive expenditure shock. We show that, through this mechanism, the government uses the additional maturities to effectively subsidize the private sector in recessions, resulting in more leisure and less volatile labor taxes. The paper is organized as follows. Section 2is a user guide that introduces the reader to PEA, machine learning, and how to combine them in a simple Neoclassical Investment Model example. Section 3introduces the reader to the problem of multicollinearity using a one-bond economy studied in Aiyagari et al. (2002) and describes the details of the NN-based expectations algorithm using a general model with Nmaturities. Section 4presents and discusses the calibration and the quantitative results for the extended model with three and four maturities. Section 5discusses and compares the NNbased expectations algorithm to other state-of-the-art methods. Section 6concludes. 2. User guide:Machine learning and PEA This section serves as an introduction to supervised machine learning. Specifically, it focuses on how to use it to solve a dynamic economic model in a similar fashion to PEA with stochastic simulation.4Hence, the purpose of this section is solely to introduce the methodology in a simple environment. The method allows us to investigate more realistic models of increased complexity. Its benefits are highlighted in the application presented in Section 3and arise from the ability of the algorithm to approximate nonlinear policy functions in the presence of a large and multicollinear state space. 2.1 Environment The typical dynamic model contains intertemporal Euler equations, intratemporal Euler equations, and laws of motion fInter(ct,Xt)=βEg(ct+1,Xt+1)|Xt, fIntra(ct,Xt)=0, Xt+1=h(Xt,ct,ξt+1), 4For a general introduction to machine learning, the reader can refer to Hastie, Tibshirani, and Friedman (2009). For a course tailored to economists, the reader can refer to the lecture notes by Jesús FernándezVillaverde available here: https://www.sas.upenn.edu/~jesusfv/teaching.html.
150 Valaitis and Villa Quantitative Economics 15 (2024) where ct∈RCis a vector of Ccontrols (with Edynamic choices and C−Estatic choices), Xt∈RSis a vector of endogenous and exogenous state variables, β∈(0, 1)is a timediscount factor, fInter :RC×RS→RE,fIntra :RC×RS→RC−E,g:RC×RS→RE,and ξt+1∈RIis a vector of innovation shocks. For example, in the stochastic neoclassical investment model, ctcorresponds to consumption, fInter corresponds to the marginal utility of consumption, fIntra does not apply if the model does not include intratemporal choices (e.g., labor), g≡f(ct+1)(zt+1αKα−1 t+1+1−δ),Xt≡{Kt,zt}is a vector that contains capital stock and TFP, and h(Xt,ct,ξt+1)is a function that describes the laws of motion for capital stock, given by the resource constraint Kt+1=(1−δ)Kt−ct+ztKα tand the TFP Markov process, that is, Xt+1=h(Xt,ct,ξt+1)=Kt+1 logzt+1=(1−δ)Kt−ct+ztKα t ρlogzt+ξt+1. The typical PEA approximates the conditional expectations in the intertemporal Euler equations as polynomial functions of the state space Xt, ∀e∈[1, E]:Ege(ct+1,Xt+1)|Xt≃Pn(Xt;ηe). The polynomial typically used in the PEA is Pn(Xt;ηe)=expηe,0 + P p=1 S s=1ηe,p,s·(lnXs,t)p, where ηe=[ηe,0,ηe,1,1,,ηe,1,S,]. For a given sequence of exogenous aggregate shocks {ξt}T t=1, an initial guess of the polynomials’ parameters η1, the standard stochastic PEA (described in Algorithm 1) aims to find parameters ηn={ηn 1,,ηn E}that solve all Euler equations and all laws of motion. When X≡{Xt}T−1 t=T0is generated by a stochastic simulation as in Algorithm 1,thematrix XTXis often ill-conditioned.5Hence, with a finite-precision computer, the inverse of XTXcannot be computed reliably and it is challenging to compute the linear regression in line 9 of Algorithm 1. This problem potentially leads to jumps in the regression coefficients and failure to converge. Moreover, in the simple illustrative case of the neoclassical investment model, a firstorder polynomial (P=1) is enough to approximate the expectation term in the Euler equation. Generically speaking, richer models that feature a larger state space and nonlinearities require the use of higher-order approximation (P1) and/or cross-state terms. These circumstances further aggravate the multicollinearity problem as the matrix ˆ XTˆ X,with ˆ X≡{Xt,X2 t,}T t=0, is even more ill-conditioned. 2.2 Supervised machine learning In this paper, we use machine learning as a tool to learn how to represent the function that maps from the set of simulated state variables {Xt}T t=0to the set of simulated 5Let λ=λ1,,λSbe the vector of eigenvalues of the matrix XTX,suchthatλ1≥λ2≥ ··· ≥ λS≥0. Ill-conditioning refers to the fact that the ratio λ1/λnis large, implying the matrix is close to being singular.
Quantitative Economics 15 (2024) A machine learning projection method 151 Algorithm 1 Stochastic (simulations) PEA. Precondition: initial state X0,sequence{ξt}T t=0, initial guess η1 n, and dampening 0 < w<1 1: while ηi nconverges do 2: for t←0toTdo Generate X≡{Xt}T t=0 3: ct←Solve fIntra(ct,Xt)=0andfInter(ct,Xt)=βPn(Xt;ηn) 4: Xt+1←h(Xt,ct,ξt+1) 5: end for 6: for t←0toT−1do Generate Y≡{yt}T−1 t=0 7: yt←g(ct+1,Xt+1) 8: end for 9: ˆηi n←(XTX)−1XTYRegress to find new weights 10: ηi+1 n←w·ˆηi n+(1−w)·ηi nUpdate with dampening 11: end while terms {yt}T−1 t=0. For example, in the neoclassical investment model with log-utility over consumption, this would serve the purpose of representing the function P(Kt,zt)=Ec−1 t+1zt+1αKα−1 t+1+1−δ|Kt,zt. Machine learning proposes a flexible structure for the function Pand infers a function from the generated data {Xt}T−1 t=T0(which we label training data) to the set of generated examples {yt}T−1 t=T0(which we label training examples). This particular task of using machine learning to learn a function that maps from inputs to outputs based on training data and examples is referred in the literature as supervised learning. And neural networks are a powerful class of universal approximators able to deal with strong nonlinearities. 2.3 Fitting neural networks In the NN-based expectations algorithm, the equivalent of the regression phase is called the training phase. As described in Supplemental Appendix C (Valaitis and Villa (2024)), a neural network is characterized by unknown weights {w,ψ}.6Similar to a regression, the objective is to seek weights such that the neural network fits the samples {Xt,yt}T−1 t=0. More precisely, the problem is to find {w0,m,wm;m=1, 2, ,M},{ψ0,e,ψe;e=1, 2, ,E}, such that the sum of squares R(w,β)= E e=1 T−1 t=0yt,e−Fe(Xt;w,ψ)2 6Appendix C contains details about the neural network structure used in this section.
152 Valaitis and Villa Quantitative Economics 15 (2024) is minimized. In a standard linear regression setting, typically (but not necessarily) this problem is solved analytically. This problem could also be solved using a gradient iterative procedure (e.g., gradient descent). This approach is typically more robust to multicollinearity since it does not require inverting the matrix ˆ XTˆ X. An iteration nof gradient descent updates the weights of the neural network according to w(n+1) m=w(n) m−γr K k=1 ∂Rk(w) ∂wm ,(1) ψ(n+1) e=ψ(n) e−γr K k=1 ∂Rk(w) ∂ψe ,(2) where the gradient can be derived using the chain rule for differentiation. More specifically, the partial derivatives ∂Rk(w) ∂wmand ∂Rk(w) ∂βein equations (1)and(2) can be efficiently computed through a two-pass algorithm called backpropagation (Rumelhart, Hinton, and Williams (1986)). Backpropagation applies the chain-rule sequentially, iterating from the output layer to the input layer. Each neuron in the hidden layer receives and dispatches information only from and to neurons that are directly connected. For this reason, this process can be efficiently parallelized. When the backpropagation algorithm is applied to a single-layer neural network, it is known as the delta rule (Widrow and Hoff (1960)). One cycle through the full training samples is called a training epoch. In other words, completing a training epoch means that all training samples have had a chance to update the model parameters. Batch (or offline) learning builds the model digesting the entire training set at once, whereas online training allows the network to update the weights as new observations come in. The former is typically implemented by batch gradient descent, when the latter can typically handle larger training sets and is implemented by stochastic gradient descent. When the neural network weights are updated, the speed at which the model changes can be updated through the parameters γrin equations (1)and(2). The parameter γris called the learning rate and it is similar in spirit to a dampening parameter. Intuitively, it represents how quickly the model “learns.” It can either be a constant (for batch learning) or optimized dynamically at each update by minimizing the error function. Other aspects that can affect the fitting of the neural network are: (i) the initial weights, (ii) the problem of overfitting, (iii) inputs normalization, and (iv) the number of neurons. Initial neural network weights are chosen as near zero random values. Figure C.2 suggests that when the weight αis close to zero, the sigmoid approaches a linear function. This choice of initial weights allows the model to adapt to nonlinearities starting from the linear case.7In practical terms, we solve the model by first initializing the neural network to a simplified version of the model. For example, before solving the neoclassical investment model as described in Algorithm 2, it is possible to solve the model analytically (in this particular case, by setting δ=1), simulate an equilibrium sequence 7Substantial research effort has been put into choosing the initial weights depending on the specific neural network architecture (e.g., see Glorot and Bengio (2010) for deep neural networks).
Quantitative Economics 15 (2024) A machine learning projection method 159 Figure 3. Autocorrelation function of the equilibrium bond sequence. Note: The figure shows the autocorrelation function of bN t. The numbers are obtained after simulating the model equilibrium dynamics for T=5000. where μtis the Lagrange multiplier on the time tmeasurability constraint, and ξU,tand ξL,tare the Lagrange multipliers on the upper and the lower bounds, respectively. By issuing debt at time t, the government commits to increasing taxes and/or to reissuing debt at time t+N. When the government sets taxes between time tand time t+N,it needs to take into account its past actions in the form of all lags of the state variables up to N. More formally, the Ramsey planner’s state space Xtis Xt=gt,{μt−i}N i=1,bN t−iN−1 i=0. The state space contains 2N+1 variables, with many lags of the same state variable (e.g., μ), which tend to be highly correlated with each other. Moreover, equation (6) reveals that the Lagrange multiplier on the implementability constraint μtfollows a random walk, creating an additional source of multicollinearity between the state variables. We solve the model with maturity N=10, and we report in Figure 3the autocorrelation function of the simulated equilibrium bond’s sequence {bN t}. It is clear that the previous 10 lags of the same variable, which are all part of the state space, are highly correlated with each other in the simulated sequence. For this reason, the model is hardly solvable using PEA (Algorithm 1). In the literature, this problem has been tackled by an algorithm called condensed PEA. Condensed PEA proposes to approximate the expected values in equations ((5), (6), and (7)) using functions of a subset XC tof the state space Xt(XC tis also called the core set). These approximations are Et(uc,t+N)≃P1(XC t;η1),Et(uc,t+N−1)≃P2(XC t;η2)and Et(uc,t+Nμt+1)≃P3(XC t;η3), where both the functions and the core set (including its cardinality) are ex ante unknown. The subset XCof the information set Xis selected through an iterative procedure called condensed PEA. In essence, this method adds an additional loop to PEA and keeps extracting orthogonal components from the state space, similar to the Principle Component Analysis (PCA), but the number of factors does not have to be chosen ex ante. A more detailed description of the procedure can be
160 Valaitis and Villa Quantitative Economics 15 (2024) found in Section 5, Algorithm 4, where we compare our methodology to existing ones in the literature. In the next section, we present our methodology in a model with Nmaturities. Due to the presence of multiple lagged bonds, the multicollinearity problem is further accentuated. 3.2 Optimal maturity management with Nbonds The economy is populated by a representative household with preferences over consumption cand leisure l. The representative household chooses sequences of consumption {ct}∞ t=0and leisure {lt}∞ t=0to maximize its time-0 expected lifetime utility: E0 ∞ t=0 βtu(ct)+v(lt), subject to the budget constraint: N i=1 pi tbi t+1+ct=(1−τt)(1−lt)+ N i=1 pi−1 tbi t, where bi tindicates an i-periods maturity bond and pi tis its corresponding price. The only source of aggregate risk in the economy is an exogenous stream of government expenditures {gt}∞ t=0. In each period, the government can finance gtby: (i) levying a proportional labor tax τtand (ii) by issuing nonstate contingent bonds with maturity 1, ,N.The government’s budget constraint reads N i=1 pi−1 tbi t=τtht−gt+ N i=1 pi tbi t+1. 3.2.0.1 Sequential formulation of the Ramsey problem Combining the technology constraint, ct+gt=ht, with the household’s labor optimality condition, 1 −τt=vl,t/uc,t, yields an expression for surplus st≡τtht−gt=ct−(1−τt)ht=ct−vl,t uc,t(ct+gt). Substitute bonds prices pi,t, pinned down by the household’s Euler equations, to get N i=1 bi tEtβi−1uc,t+i−1 uc,t=st+ N i=1 bi t+1Etβiuc,t+i uc,t, with borrowing and lending limits17 ∀i:¯ M≥bi t+1,M≤bi t+1,¯ Mtotal ≥ N i=1 bi t+1,Mtotal ≤ N i=1 bi t+1. 17 ¯ MN≥bN t+1is the government saving constraint, which is equivalent to a household’s borrowing constraint.
Quantitative Economics 15 (2024) A machine learning projection method 161 The optimality conditions are ct:uc,t−vl,t+μtuc,t−vl,t+ucc,tc+vll,t(ct+gt) + N i=1 (μt−i−μt−i+1)bi t−i+1ucc,t=0, ∀i,bi t+1:μt=[Etuc,t+i]−1Etμt+1uc,t+i+ξi U,t βi−ξi L,t βi+ξTotal U,t βi−ξTotal L,t βi, μt: N i=1 bi tEtβi−1uc,t+i−1 uc,t=st+ N i=1 bi t+1Etβiuc,t+i uc,t, where ξU,tand ξL,tare the Lagrange multipliers on the upper and the lower bounds, respectively, and ξTotal U,tand ξTotal L,tare the Lagrange multipliers on the upper and the lower bounds on the total bond portfolio. In the following section, we describe our computational strategy in detail. Details on the implementation and results using Epstein–Zin preferences can be found in Appendix A. 3.3 NN-based expectations algorithm In this section, we describe the main algorithm, which is an extension of the basic idea illustrated in Section 2.4, applied to an optimal fiscal policy model with incomplete markets and multiple maturities. Here, we present the key steps, while implementation details can be found in Appendix B.1. There are Nbonds available with maturities from 1 to Nperiods. The state space at time tis It={gt,{{bi t−k}N−1 k=0}N i=1,{μt−k}N k=1}.Theneural network needs to approximate Et[uc,t+i],Et[μt+iuc,t+i],andEt[uc,t+i−1]in function of It. We model these relationships using one single-layer neural network AN N (It).In particular, if the long maturity is N>1, then the terms to approximate are AN N i 1(It)=E[uc,t+i|It]for i=[1, ,N], AN N i 2(It)=E[μt+1uc,t+i|It]for i=[1, ,N], AN N i 3(It)=E[uc,t+i−1|It]for i=[1, ,N]. For example, in the two-bond case there are six terms to approximate and, if the short bond has 1 period maturity, they reduce to five.18 Given starting values for μtand {bi t}N i=1 and initial weights for AN N , simulate a sequence of {ct}T t=1,{μt}T t=1and {{bi t+1}N i=1}T t=1as follows:19 18Use Sand Nto denote shortand long-bond maturities, respectively. The six terms are Et(uc,t+N), Et(uc,t+N−1),Et(uc,t+N−1μt+1),Et(uc,t+S),Et(uc,t+S−1), and Et(uc,t+Sμt+1). The term that does not require approximation in the latter case is Et(uc,t+S−1), which becomes just uc,twhen S=1. 19The network can be initially trained using an educated guess for {bi t+1}N i=1,ct,μt. It is important that the initial training sequence is not constant. More details can be found in Appendix B.1.
162 Valaitis and Villa Quantitative Economics 15 (2024) 1. As suggested by Maliar and Maliar (2003), we initially restrict the solution artificially within tight bounds on all debt instruments, and refine the solution gradually while we open the bounds slowly. These bounds are particularly important and initially need to be tight and open slowly, since the neural network at the beginning can only make accurate predictions around zero debt, that is, our initialization point. Additionally, we use penalty functions instead of the ξ-terms to avoid out of bound solutions.20 Since μtis identified by the first-order condition for bi t,it is overidentified if the number of available maturities is greater than one: ∀i:μt=AN N i 1(It)−1AN N i 2(It)+ξi U,t βi−ξi L,t βi+ξTotal U,t βi−ξTotal L,t βi. We tackle this problem by using the forward-states approach described in Faraglia et al. (2019). This involves approximating the expected value terms at time t+i with functions of the state variables that are relevant at t+1 instead of tand invoking the law of iterated expectations, such that we calculate EtAN N i(It+1)instead of AN N i(It).Thisisdoneintwosteps.First,wereplacetheAN N i(It)terms in the optimality conditions with EtAN N i(It+1)and, instead of approximating Et(uc,t+i),Et(uc,t+i−1),andEt(uc,t+iμt+1), we use the information set It+1to approximate Et+1(uc,t+i),Et+1(uc,t+i−1),andEt+1(uc,t+iμt+1). Then we use Gaussian quadrature to calculate the conditional expectations of the neural network evaluated at It+1. 2. To perform the stochastic simulation, choose Tbig enough and find {ct}T t=1,{μt}T t=1 and {{bi t+1}N i=1}T t=1that solve the following system of (N+2)Tequations: ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ μt=EtAN N i 1(It+1)−1 ×EtAN N i 2(It+1)+ξi U,t βi−ξi L,t βi+ξTotal U,t βi−ξTotal L,t βi,∀i, uc,t−vl,t+μtuc,t−vl,t+ucc,tc+vll,t(ct+gt) + N i=1 (μt−i−μt−i+1)bi t−i+1ucc,t=0, N i=1 bi tβi−1EtAN N i 3(It+1)=uc,tst+ N i=1 bi t+1βiEtAN N i 1(It+1). (8) The system of equations (8) contains multiple Lagrange multipliers (arising from the inequality constraints). This poses a significant computational challenge. Ideally, one would numerically solve the unconstrained model and then verify that the constraints do not bind and if, for example, MNbinds, set bN t+1=¯ MNand find the associated values for consumption and leisure. In a multiple-bond model, this is challenging because after setting bN t+1=¯ MN, one needs to check if other constraints do not bind in the recomputed solution, and if they do, enforce them and 20We also find that including ξterms explicitly in the training set improves prediction accuracy. More details can be found in Appendix B.1.
Quantitative Economics 15 (2024) A machine learning projection method 163 recalculate the solution again, and so on. To overcome this challenge, we augment the objective function with the following differentiable penalty function: ∀i:bi t+1= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ φ 2·bi t+1−¯ Mi2 +log1+φ·bi t+1−¯ Midbi t+1if bi t+1>¯ Mi, 0ifMi≤bi t+1≤¯ Mi, φ 2·Mi−bi t+12 +log1+φ·¯ Mi−bi t+1dbi t+1if bi t+1<Mi, where φcontrols the severity of the penalty. More details can be found in Appendix B. The system of equations (8) can be rederived after including the aforementioned penalty function. We solve the system of equations (8) using the Levenberg–Marquardt algorithm. Since this is a local solver, there is no guarantee that the system is solved globally given a particular initial guess. In our implementation, we attempt to solve the system for at most maxrep number of different starting points. If the solution errors are below our specified threshold, the algorithm proceeds with the solution and moves to the next period t. If the solution errors are not below our specified threshold, we pick the solution with the lowest error. 3. If the solution error in the stochastic simulation is large, or a reliable solution could not be found, the algorithm automatically restores the previous period neural network and performs the stochastic simulation with a reduced bound. More specifically, if an unreliable solution has been detected in iteration i, the algorithm restores the iteration i−1’s environment and performs the stochastic simulation with Boundi−1=α·Boundi−1+(1−α)·Boundi−2. 4. If the solution calculated shrinking the bound at iteration i−1 is still not satisfactory, the algorithm does not go back another iteration but uses the same neural network and tries to lower the Boundi−1again toward Boundi−2. Once a reliable solution is found, the algorithm proceeds to calculate the solution for iteration i again, but with Boundi=Boundi−1+(Boundi−1−Boundi−2). In this way, if an error is detected multiple times we guarantee that both Boundi and Boundi−1keep shrinking toward Boundi−2, and there should exist a point close enough to Boundi−2such that the system can be reliably solved with both Boundi−1and Boundi. 5. If the solution found at iteration iis satisfactory, the neural network enters the learning phase supervised by the implied model dynamics, the bounds are increased, and a new iteration starts.
164 Valaitis and Villa Quantitative Economics 15 (2024) We repeat this procedure until the neural network predictions converge and the simulated sequences of {bi t}N i=1and ctdo not change.21 Algorithm 3describes the algorithm in greater detail and Appendix B.1 contains more details. 4. Numerical results In this section, we exploit the computational gains that derive from the robustness to multicollinearity of the NN-based expectations algorithm to study the optimal maturity management problem of Section 3with four maturities of 1, 5, 10, and 15 periods. Specifically, we are interested in the effects on policy and allocations arising from the additional hedging opportunities with respect to a portfolio with only a short and a long maturity. We first present the calibration and then our numerical results. 4.1 Calibration We calibrate the model following the strategy of Faraglia et al. (2019). Specifically, we use additively separable utility in consumption and leisure u(c)=c1−γ 1−γ,v(l)=χl1−ηl 1−ηl with γ=1.5 and ηl=1.8, respectively. We calibrate χsuch that households spend on average 2/3 of their time endowment on leisure in the steady state, which gives a value of 2.87. We set βto 0.96 and for the sake of comparison, we follow the calibration strategy for gtfrom Faraglia et al. (2019). We assume that gtfollows an AR(1) process gt=μg+ ρggt−1+t,t∼N(0, σ2 g)with ρgequal to 0.95. Then we look for the value of μgsuch that government expenditure is on average equal to 25% of GDP. This gives a value of 0.0042. Lastly, we set the value for σgsuch that gtis always at least 15% and at most 35% of GDP in a simulated sample of ten thousand periods, which gives a value of 0.0031. Note that such parameterization is also broadly aligned with the estimates from the data.22 The government has four debt instruments at its disposal. We set maturities to 1, 5, 10, and 15 years and denote b1,b5,b10,b15 as short, medium, and long and very long bonds, respectively. In addition to debt limits on individual bonds, we introduce a total debt limit of ±100% of GDP both in our benchmark model with only short and long bonds and in our calibration with four bonds. A fixed limit on total debt allows us to make a fair comparison and isolate the effects of the hedging benefits of the additional bond on the household’s welfare. Table 1summarizes the parameter values. Before proceeding, it is worth noting that we tested our methodology with the twobond case. Our results in a two-bond model confirms the findings of Angeletos (2002) and Faraglia et al. (2019), where the optimal debt portfolio includes a negative short bond position and a positive long bond position, as shown in Table 2. Moreover, as also 21There is no need to check μt, which can be backed out analytically from the first-order condition for ct. 22We obtain very similar estimates using the sum of government consumption and gross investment from the NIPA tables.
Quantitative Economics 15 (2024) A machine learning projection method 165 Algorithm 3 NN-based expectations algorithm applied to optimal maturity management. Precondition: parameters from Table 1; utility functions u(c)=c1−γ 1−γ,uc(c)=c−γ,i=0, Bound(0)=0. 1: Simulate AR(1) process 2: gt+1←μg+ρg·gt+t+1 3: Create and train the NN using initial conditions 4: Net ←feedforwardnet(Num. Neurons) 5: Solve the model 6: while Bound(i)<B max OR OutofBoundIter <NumOutofBound do 7: Generate {ct}T t=1,{μt}T t=1,and{{bi t+1}N i=1}T t=1 8: for t←1toTdo 9: for r←1 to maxrep do 10: xg←{c(r)guess,b(r)1 guess,,b(r)N guess} 11: {ct(r),μt(r),{bi t+1(r)}N i=1,residuals (r)}←Solve (8)|{AN N (It+1),Bound (i),xg} 12: end for 13: r∗←minrresiduals(r) 14: {ct,μt,{bi t+1}N i=1}←{ct(r∗),μt(r∗),{bi t+1(r∗)}N i=1} 15: end for 16: if residuals(r∗)>threshold then Restart from line 7 with a smaller bound 17: end if 18: Train the NN using the new simulated sequences 19: It←{gt,{{bi t−k}N−1 k=0}N i=1,{μt−k}N k=1} 20: RHSi 1,t←uc,t+ifor i=[1...N] 21: RHSi 2,t←μt+iuc,t+ifor i=[1...N] 22: RHSi 3,t←uc,t+i−1for i=[1...N] 23: Net ←train(Net, It+1,RHS t) 24: Checking convergence and updating {bi old,t}N i=1and cold,t 25: errorb←max(|{bi old,t}N i=1−{bi t}N i=1|) 26: errorc←max(|cold,t−ct|) 27: if max(errorb,error c)<then Break 28: end if 29: {bi old,t}N i=1←{bi t}N i=1 30: cold,t←ct 31: Bound(i)←Bound(i)+BoundStep 32: if Bound(i)>¯ Mthen 33: Bound(i)←¯ M 34: OutofBoundIter ←OutofBoundIter +1 35: end if 36: i←i+1 37: end while
166 Valaitis and Villa Quantitative Economics 15 (2024) Table 1. Calibrated parameters. Parameter Value Preferences Discount factor β0.96 Risk aversion γ1.5 Labor disutility χ2.87 Leisure curvature ηl1.8 Government Average gtμg0.0042 Volatility of gtσg0.0031 Autocorr. of gtρg0.95 Debt limits ¯ M,M,¯ Mtotal,Mtotal ±100% of GDP shown in Table 2, the bond portfolio positions are large and volatile as in Buera and Nicolini (2004). 4.2 Optimal debt management with three and four bonds Tables 2and 3summarize the equilibrium outstanding debt-to-GDP ratio for each maturity and for each model with an increasing number of bonds. Moments are calculated given a sequence of government expenditure shocks with persistence and volatility specified in Table 1. As shown in Tables 2and 3, the optimal policy includes an active use of all available maturities. Table 2shows that the average position of each maturity is significantly different from zero and that bond positions are volatile, suggesting their active use responding to expenditure shocks. Table 3shows the correlations of all the maturities with government expenditure and among themselves. First, it shows the position of short maturity is positively correlated with expenditure shocks while the other maturities are negatively correlated. Second, short maturity is negatively correlated with all other maturities. These two together suggest that, in addition to holding a leveraged portfolio on average, it is optimal to rebalance the portfolio toward shorter maturities. As shown in Table 2. Selected bond moments: means and variances. Model E(b1/GDP)E(b5/GDP)E(b10/GDP)E(b15/GDP) 1 Bond 0.017 - - - 2Bonds −0.03 - 0.343 - 3Bonds −0.555 0.704 0.632 - 4Bonds −0.63 0.884 0.908 −0.173 σ(b1/GDP)σ(b5/GDP)σ(b10/GDP)σ(b15/GDP) 1 Bond 0.243 - - - 2 Bonds 0.1 - 0.122 - 3 Bonds 0.591 0.34 0.533 - 4 Bonds 0.218 0.266 0.27 0.374 Note: The table shows the average outstanding debt for each maturity. Moreover, the table also reports the standard deviations of each outstanding position.
Quantitative Economics 15 (2024) A machine learning projection method 167 Table 3. Selected bond moments: correlations. Model ρ(gt,b1 t)ρ(gt,b5 t)ρ(gt,b10 t)ρ(gt,b15 t) 1 Bond 0.549 - - - 2 Bonds 0.707 - −0.482 - 3Bonds 0.35 −0.181 −0.302 - 4 Bonds 0.762 −0.094 −0.212 −0.22 ρ(b1 t,b5 t)ρ(b1 t,b10 t)ρ(b10 t,b5 t)ρ(b15 t,b1 t)ρ(b15 t,b5 t)ρ(b15 t,b10 t) 1Bond----- - 2Bonds - −0.796 - - - - 3Bonds −0.944 −0.985 0.931 - - - 4Bonds −0.458 −0.565 0.918 0.047 −0.877 −0.82 Note: The table shows the correlations between each maturity of outstanding debt and government expenditure. Moreover, the table also reports the cross-correlations among the bonds. Table 4, the additional hedging benefits of the additional maturities are reflected in a higher average leisure and a lower consumption volatility, while the economy sustains a lower average consumption. Labor tax volatility and autocorrelation also decrease significantly, while the average level rises. Next, we inspect the economic mechanism of how hedging benefits provided by the additional maturities affect household allocations and taxes. As known since Angeletos (2002), differences in long and short bond prices provide a tool to hedge against shocks by borrowing in long bonds and accumulating assets in the short term. Since long prices are more volatile than short prices, when a negative shock hits, the value of government liabilities falls more than the value of government assets, thus providing insurance against negative shocks. In addition to decreasing the government’s liabilities, the differential response of long and short prices also affects the terms of issuing new debt. Since long prices fall more than shorter ones, it becomes cheaper for the planner to obtain funds by issuing shorter debt. This is why we observe portfolio rebalancing and a negative correlation between the long and short bonds. Table 5shows how optimal debt management affects government finances as we increase the number of debt instruments. To inspect how this rebalancing matters for the government’s budget, we decompose government income into labor tax income and net financial income, which is the inflow from issuing new bonds minus the outflow due to outstanding debt. Most imTable 4. Allocations and policies. Model E(ct)σ(ln(ct)) E(lt)σ(ln(lt)) E(τt)σ(ln(τt)) ρ(ln(τt),ln(τt−1)) 1 Bond 0.252 0.029 0.666 0.006 0.247 0.121 0.971 2 Bonds 0.250 0.029 0.668 0.004 0.255 0.106 0.929 3 Bonds 0.248 0.028 0.670 0.005 0.27 0.10 0.914 4 Bonds 0.247 0.027 0.671 0.006 0.274 0.091 0.841 Note: The table shows the effects of the optimal policy on consumption and leisure as the number of bonds increases.
168 Valaitis and Villa Quantitative Economics 15 (2024) Table 5. Government income and borrowing. Description Moment 1 Bond 2 Bonds 3 Bonds 4 Bonds Corr. Debt/GDP and gtρ(ibi t yt,gt)0.547 0.136 −0.079 −0.131 Corr. Net Financial Income and gt ρ(i(pi tbi t+1−pi−1 tbi t),gt)0.186 0.405 0.416 0.511 Corr. Net Financial Income (constant price) and gt ρ(i(E(pi t)bi t+1−E(pi−1 t)bi t),gt)0.078 0.11 −0.103 0.019 Av. Net Financial Income (%) E(ipi tbi t+1−pi−1 tbi t yt)−0.142 −0.845 −2.213 −2.569 Av. Labor Tax Income (%) E(τt(1−lt) yt)24.7 25.5 27.0 27.4 Note: The table shows selected moments from the models with one, two, three, and four maturities. The first row shows the correlation between the outstanding debt/GDP ratio and expenditure shocks. Rows two and three show the correlation between government financial income and expenditure shocks. The last two rows show the average net financial income and the average labor tax income. Net Financial Income is defined as the inflow from issuing new debt at the net of the cost of buying back the outstanding debt. Net Financial Income (constant price) is the counterfactual and corresponds to Net Financial Income holding bond prices fixed at their average values. portantly, as the number of maturities increases, the correlation between total debt and government expenditures changes sign, as shown in the first row of Table 5. In the oneand two-bond economy, the government borrows from the private sector to finance expenditure shocks. In the threeand four-bond economy, the government reduces its total debt to subsidize the private sector and smooth its consumption. At the same time, net financial income becomes even more positively correlated with gt and allows for smoother labor taxes, despite a falling total debt in bad times. The reduction of total debt together with rising financial income is achieved precisely because the planner holds leveraged positions and responds to expenditure shocks by substituting to short bonds. As further evidence of this mechanism, we construct a counterfactual measure of net financial income assuming that bond prices were fixed at their mean values. The counterfactual correlation is reported in the third row of Table 5. The low correlation here suggests that the comovement between net financial income and government expenditures is achieved by exploiting the differential response of short, medium, and long prices. This indicates that if prices were constant, portfolio rebalancing would have little effect on the cyclicality of financial income and the government’s budget. Looking at the averages in rows four and five, we see that as the number of maturities increases, the government becomes a net payer to the private sector and collects a larger share of its income in labor taxes. This happens because the increase in labor taxes outweights the decrease in average labor supply. Although average household labor income falls, the household is compensated for holding government debt. 5. Comparison with alternative methods There are other simulation-based numerical methods designed to address the issue of multicollinearity among state variables. In this section, we discuss and compare our method to the two most prominent ones: the Condensed PEA (C. PEA) used in Faraglia