Essays on Asset Pricing
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Jensen, Christian Skov Doctoral Thesis Essays on Asset Pricing PhD Series, No. 32.2018 Provided in Cooperation with: Copenhagen Business School (CBS) Suggested Citation: Jensen, Christian Skov (2018) : Essays on Asset Pricing, PhD Series, No. 32.2018, ISBN 9788793744110, Copenhagen Business School (CBS), Frederiksberg, https://hdl.handle.net/10398/9653 This Version is available at: https://hdl.handle.net/10419/209079 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/3.0/
ESSAYS ON ASSET PRICING Christian Skov Jensen Ph.D. School in Economics and Management PhD Series 32.2018 PhD Series 32-2018ESSAYS ON ASSET PRICING COPENHAGEN BUSINESS SCHOOL SOLBJERG PLADS 3 DK-2000 FREDERIKSBERG DANMARK WWW.CBS.DK ISSN 0906-6934 Print ISBN: 978-87-93744-10-3 Online ISBN: 978-87-93744-11-0
Essays on Asset Pricing Christian Skov Jensen A thesis presented for the degree of Doctor of Philosophy Supervisor: Lasse Heje Pedersen Ph.D. School in Economics and Management Copenhagen Business School
ii Christian Skov Jensen Essays on Asset Pricing 1st edition 2018 PhD Series 32.2018 Print ISBN: 978-87-93 744-10-3 Online ISBN: 978-87-93744-11-0 © Christian Skov Jensen ISSN 0906-6934 The PhD School in Economics and Management is an active national and international research environment at CBS for research degree students who deal with economics and management at business, industry and country level in a theoretical and empirical manner. All rights reserved. No parts of this book may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopying, recording, or by any information storage or retrieval system, without permission in writing from the publisher.
Preface The thesis consists of three chapters, which can be read independently. The common theme across all three chapters is the relation between asset prices and investor preferences and beliefs. The first chapter addresses the question of whether it is possible to recover physical probabilities, marginal utilities, and the discount rate from observed asset prices. We show when such a recovery is possible - and when it isn’t - using a simple but general “counting argument”. Recovery is possible when the number of states of the economy is no greater than the number of time periods. Our counting argument shows why recovery is impossible in most standard financial models where the state space grows as a multinomial tree. Nevertheless, we provide conditions under which recovery is possible in such an economy. While leaving probabilities fully free, we show that recovery is possible in an economy that evolves as a multinomial tree, if the number of parameters governing the stochastic discount factor is no greater then the number of time periods. The second chapter addresses the question of how financial market tail risks vary over time and how we can infer such tail risks from asset prices. We show how the market’s higher order moments can be estimated ex ante using options written on the market. We find that, the market’s higher order moments move together in the sense that skewness becomes more negative when kurtosis becomes more positive. In other words, there are times when higher-moment risk is high, in the sense that the return distribution is both substantially left skewed (due to the large negative skewness) and fat tailed (due to the large positive kurtosis). Interestingly, higher-moment risk tends to be high at times when volatility is low, suggesting that when volatility is low, risk hides in the tails of the market return distribution. We show that this systematic variation iii
in higher-moment risk has large implications for investors; for example, the tail loss probability of a volatility-targeting investor varies from 3.6% to 9.7%, entirely driven by changes in higher-moment risk. Lastly, we show that times when higher-moment risk is high are characterized by high market and funding liquidity, high turnover, and low expected future returns. The third chapter addresses the question of how investor risk aversion varies over time and how we can infer this risk aversion from asset prices. Using options written on the market and historical market returns, I present a new method for estimating the market’s time-varying risk aversion. Interestingly, I find that the market’s risk aversion is negatively related to variance, suggesting that the market became more risk tolerant during the recent 2008-2009 financial crisis. This finding is difficult to reconcile with the leading asset pricing models. Therefore, I discuss two possible explanations for this systematic variation in risk aversion. First, I find that my results are consistent with investors salience. At times of high volatility the expected return on the market is usually high relative to the risk-free return. The relatively high expected return on the market becomes salient for investors which induces heightened risk tolerance among investors. Second, I show how the systematic variation in risk aversion can arise if the stock market is not a perfect proxy for aggregate consumption. iv
Acknowledgements This thesis has benefited from discussions with more people than I can mention here. I am grateful to everyone who has helped me, inspired me, and made my time as a doctoral student so much fun. However, a few people deserve special recognitions. Above all, I am deeply indebted to my advisor Lasse Heje Pedersen who taught me the trade of being an academic within the area of financial economics. Working with Lasse was the most inspiring professional experience of my doctoral studies. I am also indebted to my secondary advisor David Lando who took me under his wings as my master thesis advisor. David encouraged me to do research and shoot for the stars. In the spring of 2015, I visited Peter Christoffersen at the University of Toronto. Peter quickly became a mentor and a friend and he has my deepest gratitude for his encouragement in my research. I am also deeply grateful to Ian Martin. Visiting Ian at the London School of Economics in the fall of 2017 was a great experience and a perfect preparation for the job market. Colleagues and friends made it easy for me to enjoy the past five years, for that I am grateful. My co-author Niels challenged everything I said and did. Discussions with Niels taught me a lot. My partner Pia deserves special thanks. Pia has supported me tremendously throughout my doctoral studies. I have enjoyed every aspect of it: as a partner, a travel companion, and as a professional sparring partner. Finally, I am grateful to my parents for their guidance, support, and encouragement in the choices I have made over the years. v
vi
Introduction and Summaries The central formula in asset pricing relates the price of an Arrow-Debreu security to an investor’s preferences and beliefs: Price of an Arrow-Debreu security = Preferences ×Beliefs We observe the prices of Arrow-Debreu securities in the option markets. But we do not directly observe the extent to which these prices are driven by preferences or beliefs. Decomposing and investigating preferences and beliefs is essential for understanding asset prices, and it is therefore the focus of this thesis. In chapter one, my co-authors and I develop a model in which we can disentangle the contribution in asset prices which is driven by preferences and beliefs. In chapter two, my co-author and I estimate investor beliefs and study how these beliefs vary over time. In chapter three, I estimate investor preferences and study how they co-vary with investor beliefs. 1 Summaries in English Generalized Recovery Decoding risk preferences and beliefs from asset prices has been viewed as impossible until Ross (2015) provided sufficient conditions for such a recovery. Ross’ recovery relies on two critical assumptions: (1) The economy evolves as a time-homogeneous Markov chain. (2) Preferences are time-separable. In this paper, we generalize Ross’ recovery theorem to handle a general probability distribution which makes no assumptions of time-homogeneity or Markovian behavior. We show when recovery is possible – and when it isn’t – using a simple “counting” vii
medfører øget risikotolerance blandt investorer. For at teste deres hypotese regresserer jeg forholdet mellem markedets forventede afkast of det risiko-frie afkast p˚a risikoaversion. Konsistent med resultaterne i Lian, Ma, and Wang (2018) finder jeg at, investorer bliver mere risikotolerante n˚ar forholdet mellem det forventede markedsafkast og det risikofrie afkast stiger. Dernæst viser jeg hvordan den systematiske variation i risikoaversion kan opst˚a, hvis aktiemarkedet ikke er en perfekt proxy for det samlet forbrug i økonomien. xiv
Contents Preface iii Acknowledgements v Introduction and Summaries vii 1 Summaries in English ............................ vii 2 Summaries in Danish ............................. xi 1 Generalized Recovery 1 1 Introduction .................................. 2 2 Ross’s Recovery Theorem .......................... 6 3 A Generalized Recovery Theorem ...................... 9 3.1 A Noah’s Arc Example: Two States and Two Dates ....... 9 3.2 General Case: Notation ....................... 11 3.3 Generalized Recovery ........................ 11 4 Generalized Recovery vs. Other Forms of Recovery ............ 14 4.1 Generalized Recovery in a Ross Economy ............. 14 4.2 Ross Recovery in our Generalized Economy ............ 17 4.3 Recovery in Infinite Horizon ..................... 17 4.4 Flat Term Structure and Risk Neutrality ............. 19 5 Closed-Form Recovery ............................ 20 6 Recovery in a Large State Space ...................... 22 6.1 A Large Discrete State Space .................... 23 6.2 Continuous State Space ....................... 24 7 Recovery in Specific Models: Examples .................. 25 7.1 Recovery in the Mehra and Prescott (1985) model ........ 26 7.2 Cox-Ross-Rubinstein and iid. consumption growth ........ 27 7.3 A non-stationary model without Markov structure ........ 28 8 Empirical Analysis .............................. 28 8.1 Data and Sample Selection ..................... 28 xv
8.2 Recovery Methodology ........................ 29 8.3 Computing Statistics of the Recovered Distribution ........ 29 8.4 Empirical Results ........................... 30 9 Conclusion .................................. 33 A Proofs ..................................... 43 B Details on Recovery in Mehra-Prescott ................... 50 C Computing State Prices Empirically .................... 51 C.1 The Bates (2000) Stochastic Volatility Model with Jumps . . . . 51 C.2 The Jackwerth (2004) “Fast and Stable” Method ......... 52 D Pricing Kernels used in Empirical Analysis ................ 53 2 Higher-Moment Risk 57 1 Introduction .................................. 58 2 Inferring Ex Ante Moments from Asset Prices ............... 62 2.1 Inferring Ex Ante Market Tail Probabilities ............ 65 3 Data and Empirical Implementation .................... 66 3.1 Estimating Market Moments .................... 66 3.2 Estimating Market Tail Probabilities ................ 68 4 Estimated Moments Predict Realized Moments .............. 69 5 Commonalities in Higher-Moment Risks .................. 73 6 Systematic Variation in Higher-Moment Risks ............... 74 7 Implications for Investors .......................... 76 8 What Explains Higher-Moment Risk? ................... 81 8.1 The Volatility Paradox and Intermediary Leverage ........ 81 8.2 Market Liquidity and Funding Liquidity. .............. 83 8.3 “Bubble” Characteristics ...................... 85 9 Conclusion: when volatility is low, risk hides in the tails ......... 86 A Ex Ante Physical Moments, Risk-Neutral Pricing, and Realized Moments106 B Appendix Tables ...............................108 3 The Market’s Time-Varying Risk Aversion 111 1 Introduction ..................................112 2 Inferring Financial Market Risk Aversion .................115 3 Data and Risk Neutral Distributions ....................119 4 Testing the Estimated Distributions ....................120 4.1 Empirical Risk Aversion Estimates .................123 4.2 Empirical Time Preference Estimates ...............123 5 Systematic Variation in Market Risk Aversion ...............125 xvi
6 Understanding the Market’s Time-Varying Risk Aversion ........127 6.1 Evidence of a Procyclical Market Risk Aversion ..........128 6.2 Salience Theory and Market Risk Aversion ............128 6.3 Consumption and Stock Market Correlation ............129 7 Conclusion ..................................131 Bibliography 143 xvii
xviii
Chapter 1 Generalized Recovery Co-authored with David Lando and Lasse Heje Pedersen Accepted for publication in the Journal of Financial Economics Abstract: We characterize when physical probabilities, marginal utilities, and the discount rate can be recovered from observed state prices for several future time periods. We make no assumptions of the probability distribution, thus generalizing the time-homogeneous stationary model of Ross (2015). Recovery is feasible when the number of maturities with observable prices is higher than the number of states of the economy (or the number of parameters characterizing the pricing kernel). When recovery is feasible, our model allows a closed-form linearized solution. We implement our model empirically, testing the predictive power of the recovered expected return and other recovered statistics. We are grateful for helpful comments from Jaroslav Borovicka, Peter Christoffersen, Horatio Cuesdeanu, Darrell Duffie, Lars Peter Hansen, Jens Jackwerth, Pia Mølgaard, Stephen Ross, Jose Scheinkman, Paul Schneider, Paul Whelan as well as from seminar participants at London Business School, University of Zurich, University of Konstanz, University of Toronto, ATP, HEC-McGill winter finance workshop 2015, SED summer conference 2016, EFA 2016, AFA 2017, and especially to Dan Petersen for guiding us through Sard’s theorem. All three authors gratefully acknowledge support from the FRIC Center for Financial Frictions (grant no. DNRF102) and Jensen and Pedersen from the European Research Council (ERC grant no. 312417). 1
1 Introduction The holy grail in financial economics is to decode probabilities and risk preferences from asset prices. This decoding has been viewed as impossible until Ross (2015) provided sufficient conditions for such a recovery in a time-homogeneous Markov economy (using the Perron-Frobenius Theorem). However, his recovery method has been criticized by Borovicka, Hansen, and Scheinkman (2016) (who also rely on Perron-Frobenius and results of Hansen and Scheinkman (2009)), arguing that Ross’s assumptions rule out realistic models. This paper sheds new light on this debate, both theoretically and empirically. Theoretically, we generalize the recovery theorem to handle a general probability distribution which makes no assumptions of time-homogeneity or Markovian behavior. We show when recovery is possible – and when it isn’t – using a simple “counting” argument (formalized based on Sard’s Theorem), which focuses the attention on the economics of the problem. When recovery is possible, we show that our recovery inversion from prices to probabilities and preferences can be implemented in closed form. We implement our method empirically using option data from 1996-2015 and study how the recovered expected returns predict future actual returns. To understand our method, note first that Ross (2015) assumes that state prices are known not just in each final state, but also starting from each possible current state as illustrated in Figure 1.1, Panel A. Simply put, he assumes that we know all prices today and all prices in all “parallel universes” with different starting points. Since we clearly cannot observe such parallel universes, Ross (2015) proposes to implement his model based on prices for several future time periods, relying on the assumption that all time periods have identical structures for prices and probabilities (time-homogeneity), illustrated in Figure 1.1, Panel B. In other words, Ross assumes that, if S&P 500 is at the level 2000, then one-period option prices do not depend on the calendar time at which this level is observed. We show that the recovery problem can be simplified by starting directly with the state prices for all future times given only the current state (Figure 1.1, Panel C). We impose no dynamic structure on the probabilities, allowing the probability distribution to be fully general at each future time, thus relaxing Ross’s time-homogeneity 2
assumption which is unlikely to be met empirically. We first show that when the number of states Sis no greater than the number of time periods T, then recovery is possible. To see the intuition, consider simply the number of equations and the number of unknowns: First, we have Sequations at each time period, one for each Arrow-Debreu price, for a total of ST equations. Second, we have 1 unknown discount rate, S−1 unknown marginal utilities, and S−1 unknown probabilities for each future time period. In conclusion, we have ST equations with 1 + (S−1) + (S−1)T=ST +S−Tunknowns. These equations are not linear, but we provide a precise sense in which we can essentially just count equations. Hence, recovery is possible when S≤T. To understand the intuition behind this result, note that, for each time period, we have Sequations and only S−1 probabilities. Hence, for each additional time period we have one extra equation that can help us recover the marginal utilities and discount rate — and the number of marginal utilities does not grow with the number of time periods. By focusing on square matrices, Ross’s model falls into the category S=Tso our counting argument explains why he finds recovery. However, our method applies under much more general conditions. We show that, when Ross’s time-homogeneity conditions are met, then our solution is the same as his and, generically, it is unique.1 On the other hand, when Ross’s conditions are not met, then our model can be solved while Ross’s cannot. Further, we illustrate that our solution is far simpler and allows a closed-form solution that is accurate when the discount rate is close to 1. To understand the economics of the condition S≤T, consider what happens if the economy evolves in a standard multinomial tree with no upper or lower bound on the state space: For each extra time period, we get at least two new states since we can go up from highest state and down from the lowest state. Therefore, in this case S > T, so we see that recovery is impossible because of the number of states is higher than the number of time periods. Hence, achieving recovery without further assumptions is typically impossible in most standard models of finance where the state space grows 1Generically means that the result holds for all parameters except on a “small” set of parameters of zero measure. For the measure-zero set of parameters where a certain matrix of prices has less than full rank such that there is a continuum of solutions to our generalized recovery problem, we show that the multi-period version of Ross’s problem also has a continuum of solutions. 3
in this way. In other words, our model provides a fundamentally different way – via our simple counting argument – to understand the critique of Borovicka, Hansen, and Scheinkman (2016) that recovery is impossible in standard models. Nevertheless, we show that recovery is possible even when S > T under certain conditions. While maintaining that probabilities can be fully general (and, indeed, allow growth), we assume that the utility function is given via a limited number of parameters. Again, we simply need to make our counting argument work. To do this, we show that, if the marginal utilities can be written as functions of Nparameters, then recovery is possible as long as N+ 1 < T. This large state-space framework is what we use empirically as discussed further below. We illustrate how our method works in the context of three specific models, namely Mehra and Prescott (1985a), Cox, Ross, and Rubinstein (1979), and a simple non- Markovian economy. For each economy, we generate model-implied prices and seek to recover natural probabilities and preferences using our method. This provides an illustration of how our method works, its robustness, and its shortcomings. For Mehra and Prescott (1985a), we show that S > T so general recovery is impossible, but, when we restrict the class of utility functions, then we achieve recovery. For the binomial Cox-Ross-Rubinstein model (the discrete-time version of Black and Scholes (1973)), we show that recovery is impossible even under restrictive utility specifications because consumption growth is iid., which leads to a flat term structure, a pricing matrix of a lower rank, and a continuum of solutions for probabilities and preferences. While the former two models fall in the setting of Borovicka, Hansen, and Scheinkman (2016) (with a non-zero martingale component), we also show how recovery is possible in the non-Markovian setting, which falls outside the framework of Borovicka, Hansen, and Scheinkman (2016) and Ross (2015), illustrating the generality of our framework in terms of the allowed probabilities. Finally, we implement our methodology empirically using a large data set of call and put options written on the S&P 500 stock market index over the time period 1996-2015. We estimate state price densities for multiple future horizons and recover probabilities and preferences each month. Based on the recovered probabilities, we derive the risk and expected return over the future month from the physical distribution of returns using four different methods. The recovered expected returns vary substan- 4
tially across specifications, challenging the empirical robustness of the results. The recovered expected returns have weak predictive power for the future realized returns, but the predictability is stronger when we exclude the global financial crisis. We can also recover ex ante volatilities, which have much stronger predictive power for future realized volatility. We note that a rejection of the recovered distribution is a rejection of the joint hypothesis of the general recovery methodology and the specific empirical choices including the state space and the available options. The literature on recovery theorems is quickly expanding.2Bakshi, Chabi-Yo, and Bakshi (2017) and Audrino, Huitema, and Ludwig (2014) empirically test the restrictions of Ross’s Recovery Theorem. Martin and Ross (2013) apply the recovery theorem in a term structure model in which the driving state variable is a stationary Markov chain, illustrating the role played by the (infinitely) long end of the yield curve, a role already recognized in Kazemi (1992). Several papers focus on generalizing the underlying Markov process to a continuous-time process with a continuum of values and an infinite horizon (Carr and Yu (2012), Linetsky and Qin (2016)) and Walden (2017) in particular derive intuitive results on the importance of recurrence. All these papers impose time-homogeneity of the underlying Markov process.3Qin and Linetsky (2017) go beyond the Markov assumption, discussing factorization of stochastic discount factors and recovery in a general semimartingale setting. These approaches require an infinite time horizon while our approach only requires the observed finite-maturity data. Indeed, the martingale decomposition used by Borovicka, Hansen, and Scheinkman (2016) is only defined over an infinite horizon, as is the recurrence condition used by Walden (2017), and the factorization of Qin and Linetsky (2017).4 2Prior to Ross (2015), the dynamics of the risk-neutral density and the physical density along with the pricing kernel has been extensively researched using historical option or equity market data (e.g., Jackwerth (2000), Jackwerth and Rubinstein (1996), Bollerslev and Todorov (2011), Ait-Sahalia and Lo (2000), Rosenberg and Engle (2002), Bliss and Panigirtzoglou (2004) and Christoffersen, Heston, and Jacobs (2013)). 3See also Schneider and Trojani (2017a) who focus on recovering moments of the physical distribution and Malamud (2016) who shows that knowledge of investor preferences is not necessarily enough to recover physical probabilities when option supply is noisy, but shows how recovery can may be feasible when the volatility of option supply shocks is also known. 4Said differently, if we observe a data from a finite number of time periods from an economy satisfying the conditions on Borovicka, Hansen, and Scheinkman (2016), then there is no unique Markov decomposition. Recurrence means that each state is being visited infinitely often so it can only be defined over an infinite horizon. The factorization of Qin and Linetsky (2017) relies on limits 5
transition probabilities from the observed state prices in Π: P=D−1ΠH−1(1.12) Since probabilities add up to 1, we can write P e =e, where e= (1,...,1)0is a vector of ones. Using this identity, we can simplify (1.12) such that it only depends on δand h: ΠH−1e=DPe =De = (δ, δ2, . . . , δT)0(1.13) To further manipulate this equation it will be convenient to work with a division of Π into block matrices: Π = hΠ1Π2i= Π11 Π12 Π21 Π22 (1.14) Here, Π1is a column vector of dimension T, where the first S−1 elements are denoted by Π11 and the rest of the vector is denoted Π21. Similarly, Π2is a T×(S−1) matrix, where the first S−1 rows are called Π12 and the last rows are called Π22. With this notation and the fact that H(1,1) = h(1) = 1, we can write (1.13) as Π1+ Π2 h−1 2 . . . h−1 S = δ . . . δT (1.15) where of course h−1 s=1 hs.Given that these equations are linear in the inverse marginal utilities h−1 s, it is tempting to solve for these. To solve for these S−1 marginal utilities, we consider the first S−1 equations Π11 + Π12 h−1 2 . . . h−1 S = δ . . . δS−1 (1.16) 12
with solution6 h−1 2 . . . h−1 S = Π−1 12 δ . . . δS−1 − π11 . . . πS−1,1 (1.17) Hence, if δwere known, we would be done. Since δis a discount rate, it is reasonable to assume that it is close to one over short time periods. We later use this insight to derive a closed-form approximation which is accurate as long as we have a reasonable sense of the size of δ. For now, we proceed for general unknown δ. We thus have the utility ratios given as a linear function of powers of δ. The remaining T−S+ 1 equations give us Π21 + Π22 h−1 2 . . . h−1 S = δS . . . δT (1.18) and from this we see that if we plug in the expression for the utility ratios found above, we end up with T−S+ 1 equations, each of which involves a polynomium in δof degree a most T. If T=S, then δis a root to a single polynomium so at most a finite number of solutions exist. If T > S, then generically no solution exists for general Arrow-Debreu prices Π since δmust simultaneously solve several polynomial equations (where “generically” means almost surely as defined just below Proposition 1). However, if the prices are generated by the model, then a solution exists and it will almost surely be unique. To be precise, we say that Π has been “generated by the model” if there exist δ,P, and Hsuch that Π can be found from the right-hand side of (1.11). The following theorem formalizes these insights (using Sard’s Theorem): Proposition 1 (Generalized Recovery).Consider an economy satisfying Assumption 1 with Arrow-Debreu prices for each of the Ttime periods and Sstates. The recovery problem has 1. a continuum of solutions if S > T; 2. at most Ssolutions if the submatrix Π2has full rank and S=T; 6Of course, to invert Π12 it must have full rank. As long as Π2has full rank, we can re-order the rows to ensure that Π12 also has full rank. 13
3. no solution generically in terms of an arbitrary positive matrix Πand S < T; 4. a unique solution generically if Πhas been generated by the model and S < T . The proof of this and all following propositions are in the appendix. The proposition states our results using the notion “generically,” which means that they fail to hold at most for a set of measure zero. Said differently, if someone picks parameters “at random,” then our results hold almost surely.7 Further, since Sard’s theorem is not a standard tool in asset pricing theory, some words here on the basic intuition behind our use the theorem are in order. To get started, consider a linear function f(x) = Ax from Rmto Rngiven by the n×m matrix A. We know that if n=mand Ahas full rank, then the image of Ais all of Rn, i.e., every point of Rnis being “hit” by A. If, however, n>m, then the image of A is a linear subspace of Rn, which is vanishingly small (has Lebesgue measure 0 in Rn). By Sard’s theorem, we can extend this result to a non-linear smooth function fand still conclude that, when n > m, the image of fis vanishingly small. Said differently, there exists no solution xto f(x) = ygenerically (i.e., if you pick a random ythen almost surely no solution exists).8 4 Generalized Recovery vs. Other Forms of Recovery Proposition 1provides a simple way to understand when recovery is possible, namely, essentially when the number of time periods Tis at least as large as the number of states S. We now show how our method relates to Ross’s method and other recovery results. 4.1 Generalized Recovery in a Ross Economy We first show that our method generalizes Ross’s recovery method in the sense that, if we are in a Ross economy, then any solution to Ross’s problem has a corresponding solution to our problem. 7We note that the fact that our results hold only generically is not a consequence of our solution method – indeed, there exist counter-examples for special sets of parameters as discussed in our examples. 8On a more technical note, Sard’s theorem in fact states, that if Mis the set of critical points of f(i.e., the set of points for which the Jacobian matrix of fhas rank strictly smaller than n), then f(M) has Lebesgue measure zero in Rn. When n > m all points are critical points, and therefore in this case f(M) is the same as the image of f, which is what we need for our proof. 14
It is important to be clear about the terminology here. In Ross’s recovery problem, physical transition probabilities are specified in terms of a one-period transition probability matrix ¯ Pwhich includes transition probabilities from states that we are currently not in (“parallel universes”). Our problem focuses on recovering the matrix Pof multi-period transition probabilities as seen from the state we are in at time 0, which we take to be state 1. We say that Pis generated from ¯ Pif the k’th row of P is equal to the first row of ¯ Pk. The same terminology can be applied to state prices, of course.9 Proposition 2 (Generalized Recovery Works in a Ross Economy).If observed prices Πover S=Ttime periods are generated by a Ross economy (i.e., an irreducible matrix ¯ Πof one-period state prices and probabilities Psatisfying Assumptions 1 and 2) then 1. The matrix Pgenerated from ¯ Pis a solution to our generalized recovery problem. 2. Pis a unique solution to our generalized recovery problem generically in the space of Ross price matrices ¯ Π. 3. If Π12 has full rank, then Ross’s parallel-universe prices ¯ Πcan be derived uniquely from multi-period prices Πobserved from the current state. Otherwise, there may exist a continuum of Ross prices ¯ Πconsistent with the observed prices. The rank condition is satisfied generically in the space of Ross price matrices. Part 1 of the proposition confirms that any solution to Ross’s recovery problem corresponds to a solution to our generalized problem. Part 2 of the proposition considers the deeper question of uniqueness. Ross establishes a unique solution while our generalized recovery solution in our earlier Proposition 1 only narrows the solution set down to at most S=Tsolutions. Interestingly, Proposition 2 shows that our method too yields a unique solution when prices come from a Ross economy, generically. Thus, in this sense, nothing is “lost” by using generalized recovery even when we are in a Ross economy. One way to understand this result is to note that Ross’s problem comes down to solving a characteristic polynomial, and, similarly, our generalized recovery problem 9The notion of generating Pfrom ¯ Pis based on the fact that, in a Ross economy, the matrix of probabilities of going from state ito state jin ktime periods is given by ¯ Pk. Likewise, the k-period state prices are given by ¯ Πk. 15
can be solved via the polynomial given by (1.18). Even though these polynomials come from different sets of equations, it turns out that they have the same roots when Ross’s assumptions are satisfied. Finally, part 3 of the proposition deals with the issue that some of our results only hold “generically,” that is, for almost all parameters. One might ask whether Ross also has a similar problem for the (small set of) remaining parameters. The answer turns out to be “yes,” and for a reason that has not yet been discussed in the context of Ross’s method. The issue is that Ross finds a unique solution given his parallel universe price matrix ¯ Π, but where does this matrix come from? In any real-world application, we start with observed prices Π over time as in our generalized recovery setting. When Ross implements his model empirically, he must first find his ¯ Π from the observed Π and then use his recovery method (but he does not consider the mathematics of the first step, getting ¯ Π from Π). Part 3 of the proposition shows that Ross has the same problem as we do for the small set of parameters where Π12 has less than full rank. In other words, his lack of uniqueness arises from the difficulty in finding the price matrix ¯ Π. Interestingly, this may have been unnoticed since Ross takes ¯ Π as given in his theoretical analysis (and shows that his recovery is unique for each ¯ Π). This last point is most clearly seen through an example: Consider two different one-period transition probability matrices, that are both irreducible: ¯ P= 1 3 1 3 1 3 1 3 1 3 1 3 1 3 1 3 1 3 and ¯ P0= 1 3 1 3 1 3 1 3 2 30 1 302 3 If we assume that the current state is state 1, then since all powers of the matrices ¯ P and ¯ P0have the same first row, namely (1 3,1 3,1 3), it follows that the matrices Pand P0 (i.e., the physical transition probabilities as seen from state 1) generated by ¯ Pand ¯ P0 become the same matrix P=P0= 1 3 1 3 1 3 1 3 1 3 1 3 1 3 1 3 1 3 For given discount factors Dand marginal utilities H, Π = DPH and Π0=DP0H 16
are then the same, and hence observing the 3 ×3 matrix of state prices Π would not allow us to distinguish between the physical transition matrices ¯ Pand ¯ P0. The problem is not mitigated by observing more periods. It is simply impossible in a world where we cannot observe parallel universe prices to distinguish between the two irreducible matrices. In our approach, we do not seek to recover the one period transition probabilities. Rather, we recover the matrix P, and our ability to do so depends on the rank of a submatrix the Π matrix. For example, if we let δ= 0.98, and let h1= 1, h2= 0.9, h3= 0.8, then the sub-matrix of state prices Π12 has rank 1, and this means that we would not have unique recovery either. 4.2 Ross Recovery in our Generalized Economy We now establish that our formulation is strictly more general, by showing that for many “typical” price matrices (e.g., those observed in the data), no solution exists for Ross’s recovery problem even though a solution exists for the generalized recovery problem. Proposition 3 (Generalized Recovery is More General).With S=T, there exists set of parameters with positive Lebesgue measure for the generalized recovery problem where no solution exists for Ross’s recovery problem. With S < T, generically among price matrices for the the generalized recovery problem, there exists no solution to Ross’s recovery problem. This proposition shows that generalized recovery may be useful because it can match a broader class of market prices, in addition to the basic advantage that it starts with the observed multi-period prices (rather than parallel universe prices). 4.3 Recovery in Infinite Horizon In addition to generalizing Ross’s method, our result also provides a simple and intuitive way of understanding why, for example, growth may present a challenge for recovery, cf. the critique of Borovicka, Hansen, and Scheinkman (2016) that recovery is infeasible in standard models. Indeed, we provide a simple counting argument: Suppose that the economy has growth such that, for each extra time period, the economy 17
can increase from the previously highest state and go down from the previously lowest state. Then we get two new states for each new time period, which implies that S > T such that recovery is impossible. Nevertheless, we can still achieve recovery in such a large state space if we consider a class of pricing kernels that is sufficiently low-dimensional as we discuss below in Section 6. Our argument is very different from that of Borovicka, Hansen, and Scheinkman (2016) who rely on a martingale decomposition, which requires infinite time horizon. Our counting argument is simple and is based on a finite horizon, consistent with the data observed in practice. Our finite-horizon recovery theorem is therefore also markedly distinct from the existing approaches that exist in continuous-time models in that we make no reference to, and have no need for, recurrence or stationarity conditions. In a diffusion setting, Walden (2017) shows the fundamental role of recurrence as a necessary condition for recovery in these models. Recurrence essentially means that each state is being visited infinitely often so it can only be defined over an infinite horizon. Recurrence bears some resemblance to Ross’ condition of irreducibility in that an infinite time extension of an irreducible chain would be recurrent. The result of Walden (2017) is intuitive since, when states are visited infinitely often, we have a chance to recover probabilities. Our approach can naturally be used to consider whether recovery is possible in a finite-time version of infinite-horizon process (i.e., even if a process is defined over an infinite horizon, we can ask what happens if we only see it over a couple of years). Further, we can show via some examples that recovery may even be possible for nonrecurrent processes or processes with growth. To give a simple example of this, consider a two period non-homogeneous Markov process with two states defined from the probability transition matrices for each time ¯ P(0,1) = 0.4 0.6 0.5 0.5 and, for t≥1,¯ P(t, t + 1) = 1 0 0 1 In the first period, the process either stays in its current state or jumps to the other state, but, after that, the process is absorbed in its current state. If we only observe prices for two time periods, then this is clearly the restriction of a non-recurrent process. 18
Given that S=T= 2, our counting argument shows that generalized recovery is feasible. We could also imagine a process with growth, starting in the “lowest state” 1 and evolving according to a transition matrix specified as an upward drifting process. To give a simple illustration, imagine Assumptions 1 and 2 hold and that the one-period transition matrix of physical probabilities across five states is given as ¯ P= 0.5 0.5 0 0 0 0.1 0.5 0.4 0 0 0 0.1 0.5 0.4 0 0 0 0.1 0.5 0.4 0 0 0 0.5 0.5 If we observe prices over five time periods, then our counting argument is satisfied S=T= 5, and we see that it is not growth per se which makes recovery impossible — it is the expanding state space necessary to accommodate models with growth that may cause problems. In summary, our results complement those in the literature in two ways. First, generalized recovery may work when other methods don’t and vice versa. Second, generalized recovery provides an economic intuition in finite economies while other methods do so in infinite-horizon economies. 4.4 Flat Term Structure and Risk Neutrality We finally note that the very special case of an observed flat term structure of interest rates has some special properties. In particular, with a flat term structure there exists a solution to the problem in which the representative agent is risk neutral, echoing an analogous result by Ross. To see this result, note that the price of a zero-coupon bond with maturity τis equal to the sum of the τ’th row of Π, which we write as (Πe)τ. Having a flat term structure means that the yield on the zero-coupon bonds does not depend on maturity, 19
i.e., that there exists a constant rsuch that 1 (1 + r)τ= (Πe)τ(1.19) Let the T×Smatrix Qcontain the risk-neutral transition probabilities seen from the starting state, i.e., the k’th row of Qgives us the risk-neutral probabilities of ending in the different states at date k. Proposition 4 (Flat Term Structure).Suppose that the term structure of interest rates is flat, i.e., there exists r > 0such that 1 (1+r)τ= (Πe)τfor all τ= 1, . . . , T. Then the recovery problem is solved with equal physical and risk-neutral probabilities, P=Q. This means that either the representative agent is risk neutral or the recovery problem has multiple solutions. We note that this result should be interpreted with caution. The knife-edge (i.e., measure zero) case of a flat term structure may well be generated by the knife-edge case of a price matrix Π with low rank, which implies that a continuum of solutions may exists and the representative agent may well be risk averse (as one would expect). Intuitively, a flat term structure may be generated by a Π with so much symmetry that it has a low rank. 5 Closed-Form Recovery The recovery problem is almost linear, except for the powers of the discount rate δ which enter into the problem as a polynomial. In practical implementations over the time horizons where options are liquid, a linear approximation provides an accurate approximation given that δis close to one. For instance, we know from the literature that δis close to 0.97 at an annual horizon. The linear approximation is straightforward. To linearize the discounting of δτ around a point δ0(say, δ0= 0.97), we write δτ≈aτ+bτδfor known constants aτand bτ. Based on the Taylor expansion δτ≈δτ 0+τδτ−1 0(δ−δ0),we have aτ=−(τ−1)δτ 0 and bτ=τδτ−1 0. As seen in Figure 1.2, the approximation is accurate for δ∈[0.94,1] for time horizons less than 2 years. 20
With the linearization of the polynomials in δ, the equations for the recovery problem (1.13) become the following: π11 . . . πT1 + π12 . . . π1S . . .. . . πT2. . . πTS h−1 2 . . . h−1 S = a1+b1δ . . . aT+bTδ (1.20) which we can rewrite as a system of Tequations in Sunknowns as −b1π12 . . . π1S . . .. . .. . . −bTπT2. . . πTS δ h−1 2 . . . h−1 S = a1−π11 . . . aT−πT1 (1.21) Rewriting this equation in matrix form as Bhδ=a−π1(1.22) we immediately see the closed-form solution hδ= B−1(a−π1) for S=T (B0B)−1B0(a−π1) for S < T (1.23) We see that, when S=T, we simply need to solve Slinear equations with Sunknowns. When S < T, we could simply just consider Sequations and ignore the remaining T−S equations. More broadly, if S < T and we start with prices Π that are not exactly generated by the model (e.g., because of noise in the data), then (1.23) provides the values of δ and the vector hthat best approximate a solution in the sense of least squares. The following theorem shows that the closed-form solution is accurate as long as the value of δ0is close to the true discount rate: Proposition 5 (Closed-Form Solution).If prices are generated by the model and B has full rank S≤Tthen the closed-form solution (1.23) approximates the true solution in the following sense: The distance between the true solution (¯ δ, ¯ h, ¯ P) and the 21
Proposition 4, where recovery is impossible.14 Concretely, the problem is that the price matrix Π from (1.34) is not full rank. Hence, as seen in Figure 1.3 Panel B, the objective of minimizing pricing errors has a continuum of solutions. In other words, recovery is not feasible. 7.3 A non-stationary model without Markov structure Lastly, we consider a model where the consumption growth Xtis not Markov. Specifically, we still consider the binomial tree described above in Sections 7.1–7.2, but now we let the probability of transitioning up/down from any state sat any time tdepend on the path taken from time 0 to time t. At each node at each path, we draw a random uniformly distributed probability for an “up” move, and, of course, assign one minus this probability to the next “down” node. We now seek to recover δand γ. As seen in Figure 1.3 Panel C, the objective function has a unique solution which again equals the true parameters ˆ δ= 0.98 and ˆγ= 4. Hence, recovery can be possible even when the driving process is non-stationary and non-Markovian, again under parametric assumptions about the utility function (i.e., a model outside the scope of Ross (2015) and Borovicka, Hansen, and Scheinkman (2016)). 8 Empirical Analysis This section describes our data, empirical methodology, and empirical findings. 8.1 Data and Sample Selection We use the Ivy DB database from OptionMetrics to extract information on standard call and put options written on the S&P 500 index for every last trading day of the month from January 1996 to December 2015. We obtain implied volatilities, strikes, and maturities, allowing us to back out market prices. As a proxy for the risk-free rate, we use the zero-coupon yield curve of the Ivy DB database, which is derived 14Iid. consumption growth and standard utility functions generally lead to a flat term structure because the price of a bond with τperiods to maturity can be written as Et(δτut+τ ut) = Et(Qs=1,...,τ δut+s ut+s−1) =: ( 1 1+r)τ, where the expected utility increments are the same for all sbecause they depend on consumption growth ct+s ct+s−1, which has constant expected value when it is iid. 28
from LIBOR rates and settlement prices of CME Eurodollar futures. We also obtain expected dividend payments, calculated under the assumption of a constant dividend yield over the life time of the option. We consider options with time to maturity between 10 and 360 days and apply standard filters, excluding contracts with zero open interest, zero trading volume, and quotes with best bid below $0.50, and options with implied volatility higher than 100%. 8.2 Recovery Methodology The Generalized Recovery Theorem relies on the knowledge of Arrow-Debreu state prices from the current initial state to all possible future states for several future time periods. Unfortunately, there is currently no market trading pure Arrow-Debreu securities. Therefore, we use options to back out Arrow-Debreu prices. Further, given the large number of states, we use the parametric kernel method from Section 6. To study the robustness of recovery, we consider two different methods for backing out Arrow-Debreu prices and two different specifications of the pricing kernel, for a total of four different recovered distributions and preferences. More specifically, we apply the following two methods of extracting Arrow-Debreu prices from options: (i) the parametric model of Bates (2000) and (ii) the nonparametric method of Jackwerth (2004). Each of the methods yields Arrow-Debreu prices across multiple time horizons and mutliple index levels for each day tas described in detail in Appendix C. Given these observed Arrow-Debreu prices, we recover preferences and probabilities based on the two different specifications of the pricing kernel that we denote “piecewise linear” and “polynomial” pricing kernels, respectively, as described in detail in Appendix D. 8.3 Computing Statistics of the Recovered Distribution Once we have recovered the probabilities of each state for each future time period, it is straightforward to compute any statistic under the physical probability distribution. If the level of the index at time tis St, then the state space consists of all integer values of the index between the minimum value (1 −2.5VIXt)Stand (1 + 4VIXt)St. Let Nt 29
denote the number of states as seen from time tand think of state 1 as the lowest state and Ntas the highest state. We compute the recovered expected excess return µtat time tby summing over the Ntpossible states: µt=EP t[rt,t+1]−rf t,t+1 = Nt X ν=1 pt+1,νrt+1,ν −rf t,t+1 (1.35) where rf t,t+1 is the risk-free rate, pt+1,ν is the recovered time-tconditional physical probability for the transition to state νat time t+ 1, rt+1,ν =St+1(ν) St−1 is the return in state ν, and St+1(ν) is the value of the index at time t+ 1 if state νis realized. We compute the contemporaneous unpredictable innovation in the conditional expected return as ∆µt+1 =µt+1 −Et[µt+1] (1.36) where we impose an AR(1)-process on the innovation to the risk premium Et[µt+1] = α0+α1µtbased on the regression µt+1 =α0+α1µt+εt+1 (1.37) The estimated persistence parameter α1is 0.3 at the monthly horizon. We compute the recovered conditional variance, VARP t(rt,t+1), analogously to how we computed the expected return and we denote the recovered volatility by σt= qVARP t(rt,t+1). 8.4 Empirical Results We next investigate the properties of the recovered probabilities based on each of our four methods. We first consider the recovered expected return. Table 1.1 shows the correlation matrix for the recovered expected returns based on each of our four methodologies as well as the VIX volatility index and the SVIX variable of Martin (2017). The good news is that all variables are positively correlated, as we would expect. The ææless good news is that the correlations between the different recovered expected returns are modest in magnitude, with an average pairwise correlation of only 0æ.5. This modest correlation is concerning because all these recovered expected 30
returns should be measures of the same thing, namely the market’s expected return at any given time. Figure 1.4 shows the time series variation of the recovered expected return based on one of the methodologies (we plot just one time series since it is difficult to look at all four together). These recovered expected returns do not look unreasonable, but we next try to test their predictability of actual realized returns. Specifically, we regress the ex post realized excess return on the ex ante recovered expected excess return, µt, and the ex post innovation in expected return, ∆µt+1: rt,t+1 =β0+β1µt+β2∆µt+1 +t,t+1 (1.38) where t+1 is a noise term. To understand this regression, note that we are interested in testing whether the recovered probabilities give rise to reasonable expected returns, that is, time-varying risk premia. For this, we want to test whether a higher ex ante expected return is associated with a higher ex post realized return (β1>0), whether an increase in the risk premium is associated with a contemporaneous drop in the price (β2<0), and whether the intercept as zero (β0= 0). Table 1.2 reports the results of this regression for each of our four recovery methodologies as well as using VIX and SVIX as the expected return over the full sample from 1997 to 2015. First, the intercept β0is insignificantly different from zero in most specifications, but significantly different from zero using method 2 and using VIX, providing evidence against these models. Second, β1is positive and marginally significant from 0 in model 1, but otherwise insignificantly different from zero, providing neither evidence in favor or against the models. The coefficient β2is highly significant and has the desired negative sign in all models. Further, as expected the absolute value of β2 is greater than one since a shock to the discount rate leads to a larger shock to the price (cf. Gordon’s growth model for the extreme example of a permanent shock). Table 1.3 reports the result of regression (1.38) over the sub-sample that excludes the global financial crisis (9/2008–7/2009), a sub-sample that has been considered in the literature (e.g., Martin (2017)). The results here are stronger and more consistent with theory. All the key parameters have the expected sign, the estimated coefficient β0is small and insignificant in all models, the estimated coefficient β1is positive and 31
marginally significant or insignificant, and β2is negative and significant. The reason that the models work better when we exclude the crisis is intuitive: During the crisis, there were several months in which the ex ante recovered expected return was high, but, nevertheless, the ex post realized return was negative and large in magnitude. It seems plausible that investors were scared at that time, which means that it is plausible that the true required return was indeed high, which in turn implies that the negative realized return was a negative surprise. Hence, one could argue that the model gets this period wrong for the “right” reason, but we don’t want to push this argument too far as the most compelling evidence is almost always that of using the full sample. Finally, we consider the recovered physical volatility as plotted in Figure 1.5. This recovered volatility looks reasonable. Further, the recovered volatilities are similar across the different methodologies with an average pairwise correlation of 0.95 and an average correlation to VIX of 0.92. It is not that surprising that volatilities can be recovered, but studying volatility provides a simple and powerful reality check of our method since the true future volatility is known with much less error than the expected return. Hence, we regress the ex post realized volatility on the ex ante recovered conditional volatility, σt: qVAR(rt,t+1) = β0+β1σt+t,t+1 (1.39) where the realized volatility pVAR(rt,t+1) is computed using close-to-close daily data over the 4 weeks from tto t+ 1 by OptionMetrics. We also run the same regression where we replace the recovered volatilities by the VIX volatility index. The theory predicts that β0= 0 and β1= 1. Table 1.4 reports the results. As seen in Table 1.4, the estimated intercept coefficient β0is insignificant for models 1 and 2, but significant for models 3 and 4. However, for all models, the intercept is smaller than that of VIX, suggesting that the recovered volatilities are less biased than VIX. The estimated slope coefficient β1is positive and highly significant for all models. Further, the estimated slope is close to the predicted value of 1, in particular closer than the estimated value for VIX. Lastly, we see that VIX has a slightly higher R2, 32
which may reflect that the recovery method introduces some noise in the volatility measures. In summary, we find substantial differences across the recovered probabilities based on different methodologies, and the predictive power for future returns appears weak in the full sample, but slightly stronger in the sample that excludes the gælobal financial crisis. The recovered volatilities predict well the future volatility in a way that is less biased than VIX, but slightly lower R2. We are able to reject that the recovered probabilities provide a perfect description of the future evolution of the market based onaBerkowitz (2001) test.15 This rejection could be due to the details of our implementation. For instance, while the true pricing kernel may depend on multiple factors, we assume that the state space is given by the level of S&P500 since we do not observe option prices depending simultaneously on multiple factors. 9 Conclusion We characterize when preferences and natural probabilities can be recovered from observed prices using a simple counting argument. We make no assumptions on the physical probability distribution, thus generalizing Ross (2015) who relies on strong time-homogeneity assumptions. In economies with growth, our counting argument immediately shows that recovery is generally not feasible. While this finding parallels results by Borovicka, Hansen, and Scheinkman (2016), our intuitive counting argument is fundamentally different and does not rely on the assumptions of an infinite-period time-homogeneous Markov setting, but, rather, is based on the general methods pioneered by Debreu (1970) for general equilibrium. To pursue recovery even in economies with growth, e.g., classical multinomial models, we show how our method can be used when the pricing kernel can be parameterized by a sufficiently low-dimensional parameter vector. When recovery is feasible, our model allows a closed-form linearized solution. We implement our model empiri- 15The details of this test are not reported for brievity. The idea is that, given the estimated distribution ˆ Ftof the excess return rt+1 at time t, the distribution of the transformed variable ut+1 = ˆ Ft(rt+1) should be uniform and the distribution of the further transformed variable xt+1 = Φ−1(ut+1) should be standard normal, which is tested by estimated the coefficients in the model xt+1 =c+βxt+t and perform a likelihood ratio test of the joint hypothesis that c=β= 0 and V ar(t) = 1. 33
cally using several different specifications, testing the predictive power of the recovered statistics. 34
Panel A. Ross’s Recovery Theorem: one period, two “parallel universes” t=1t=0 Current state Other state Panel B. Ross’s Recovery Theorem: time-homogeneous dynamic setting t=2t=1t=0 Current state Other state Panel C. Our Generalized Recovery: No assumptions about probabilities t=2t=1t=0 Current state Other state Figure 1.1: Generalized Recovery Framework. Panel A illustrates the idea behind Ross’s Recovery Theorem, namely that we start with information about all Arrow- Debreu prices in all initial states (not just the state we are currently in, but also prices in “parallel universes” where today’s state is different). Panel B shows how Ross moves to a dynamic setting by assuming time-homogeneity, that is, assuming that the prices and probabilities are the same for the two dotted lines, and so on for each of the other pairs of lines. Panel C illustrates our Generalized Recovery method, where we make no assumptions about the probabilities. 35
0.90 0.92 0.94 0.96 0.98 1.00 0.80 0.85 0.90 0.95 1.00 δ a+bδ δt Panel A: t= 2 years 0.90 0.92 0.94 0.96 0.98 1.00 0.95 0.96 0.97 0.98 0.99 1.00 δ a+bδ δt Panel B: t= 0.5 years Figure 1.2: Closed-Form Solution: Approximation Error. The figure shows that the generalized recovery problem is very close to being linear. We show that the only non-linearity comes from the discount rate δdue to the powers of time, δt. However, the function δ→δtis very close to being linear for the relevant range of annual discount rates, say δ∈[0.94,1], and the relevant time periods that we study. Panel A plots the discount function and the linear approximation around δ0= 0.97 given a horizon of t= 2 years. Panel B plots the same for a horizon of a half year. 36
Table 1.1: Correlation Matrix. This tables shows the pairwise correlations between the recovered conditional expected excess return for different specifications of marginal utilities and method for estimating risk-neutral prices; (i) µt,1: Bates and polynomial, (ii) µt,2: Bates and piecewise linear, (iii) µt,3: Jackwerth and polynomial, (iv) µt,4: Jackwerth and piecewise linear. We augment the table with pairwise correlations with the VIXtindex and the lower boundary on the equity premium, SVIXt, due to Martin (2017). µt,1µt,2µt,3µt,4VIXtSVIXt µt,11 0.359 0.393 0.392 0.534 0.485 µt,21 0.642 0.523 0.716 0.794 µt,31 0.642 0.784 0.830 µt,41 0.634 0.689 VIXt1 0.928 SVIXt1 37
spect to each other) form a subspace of dimension ST −2(T−S+ 1) since each point imposes T−S+ 1 equations (and saying that the points are in general position means that all these equations are independent). Therefore, Mis a manifold of dimension ST −2T+ 2Ssince the pair (δa, δb) depends on two parameters, and, for a given pair, there is a (ST −2T+ 2S−2)-dimensional subspace of possible Π (any two distinct points are always in general position). Hence, we see that dim(M) = ST −2T+ 2S < ST −T+S= dim(X) since S < T , which implies that G(M) has measure zero in X. Further, the prices where recovery is impossible, F(G(M)), have measure zero in the space of all prices generated by the model F(X) where we use the Lebesgue measure on Xto define a measure16 on F(X). Proof of Proposition 2. Let ¯ Π be an S×Stransition matrix corresponding to an irreducible matrix (as in Ross). Without loss of generality we assume that the current state is the first state. Since prices are generated by a Ross economy, the observed matrix Π of multiperiod prices is given as Π := (¯ Π)1 (¯ Π2)1 . . . (¯ ΠS)1 where (¯ Π)1denotes the first row of ¯ Π,(¯ Π2)1is the first row of ¯ Π2, etc. We want to show that all solutions to the eigenvalue problem for ¯ Π give rise to solutions to our system (both the “correct solution” and the ones that, by the Perron-Frobenius theorem, do not generate viable solutions). Observe that if z= (z1, . . . , zS)0is a (right) eigenvector of ¯ Π with corresponding eigenvalue δ, then Πz= (δz1, δ2z2, . . . , δSzS)0. If zis the eigenvector corresponding to the maximal eigenvalue of ¯ Π, then we know that it is strictly positive. Generically, in the space of matrices ¯ Π, the matrix is diagonaliz- 16We can define a measure on F(X) by µ∗(A) := µ(F−1(A)) for any set A, where µis the Lebesgue measure on X. 44
able with eigenvectors that contain no zeros and with distinct non-zero eigenvalues – in particular, it has full rank. Therefore, generically, even for the other eigenvectors, we have that the coordinates of zare non-zero, so we can normalize zto have first coordinate 1. Now let the Ross probability matrix be defined (as in Ross) ¯ P=1 δDiag−1(z)¯ ΠDiag(z) (1.40) with corresponding multi-period probabilities given by P:= (¯ P)1 (¯ P2)1 . . . (¯ PS)1 . Note that since the rows of ¯ Psum to 1, so do rows of P. Further, using (1.40), P= (1 δ1Diag(z)−1¯ Π1Diag(z))1 . . . (1 δSDiag(z)−1¯ ΠSDiag(z))1 = (1 δ1¯ Π1Diag(z))1 . . . (1 δS¯ ΠSDiag(z))1 =D−1ΠDiag(z), where the second equality uses that z1= 1 and that we only consider the first rows, and the last equation uses our maintained notation D=Diag(δ,...,δS). We note that this equation is the same as our equation (1.12), which means that all solutions to Ross’s eigenvalue problem for the matrix ¯ Π also appear as solutions to our equations. The fact that Pgenerated from the Ross solution ¯ Pis a solution to the generalized problem required no assumptions other than irreducibility, and this proves part 1 of the theorem. To obtain uniqueness also of our solution, note that, generically, there are Seigenvectors for Ross’s matrix from which a matrix Pcan be generated using (1.40). Each of these solutions can be used to generate a solution Pto our problem, as shown above. The S−1 solutions are “fake” in the sense that they imply that some marginal utilities (elements in the eigenvector zabove) are negative. Hence, these solutions are also fake in the context of the generalized recovery framework. Given that Ross’s equations 45
yield a total of Spossible solutions to our problem, of which S−1 are fake, we have a unique viable solution (by Proposition 1) if we can ensure that Π12 has full rank. This follows from the generic property of ¯ Π as being diagnonalizable with distinct, non-zero eigenvalues. In fact, we can show the stronger statement that Π has full rank: Consider the diagonalization of Ross’s price matrix as ¯ Π = V ZV 0, where Z= diag(z1, ..., zS) is the matrix of eigenvalues and Vis the matrix of eigenvectors. The k’th row in the generalized-recovery pricing matrix is the first row (still assuming that the starting state is 1) of ¯ Πk=V ZkV0. Letting vdenote the first row in V, we see that the k’th row of Π is vZkV0= (v1zk 1, ..., vSzk S)V0so Π = 1... 1 . . .. . . zT−1 1... zT−1 S v1z10 ... 0vSzS V0(1.41) Therefore, Π is full rank generically because it is the product of three full-rank matrices. Indeed, the first matrix is a Vandermonde matrix, which is full rank when the z’s are non-zero and different, which is true generically. The second matrix is clearly also full-rank since the v’s are also non-zero generically, and the third matrix is full rank by construction. Hence our set of equations can have no more than Ssolutions, and since S−1 of these are “fake”, we have unique recovery of the solution corresponding to Ross’s solution also, generically. To see how to derive ¯ Π in an economy where Π arises from a time-homogeneous Ross economy, note that the following equation set must hold: (Π)2 . . . (Π)S | {z } (S−1)×S = (Π)1 . . . (Π)S−1 | {z } (S−1)×S ¯ Π (1.42) where (Π)iis the i’th row of Π. Further, using the notation from (1.14) for blocks of Π and denoting the first row of ¯ Π by ¯ Π1and remaining rows by ¯ Π2, we can rewrite 46
this equation as (Π)2 . . . (Π)S =hΠ11 Π12i ¯ Π1 ¯ Π2 (1.43) Given that ¯ Π1is known (because the one-period state prices from state 1 are observed), it is useful to further rewrite this system as (Π)2 . . . (Π)S −Π11 ¯ Π1= Π12 ¯ Π2(1.44) Hence, when Π12 is full rank, the Ross price matrix ¯ Π2can be derived uniquely and explicitly by pre-multiplying by (Π12)−1. We have already shown in Part 2, that Π12 has full rank generically. If Π12 does not have full rank, there exists a non-zero vector v∈RS−1for which Π12v= 0. In this case, if we start from a solution for which ¯ Π2has strictly positive elements, we can pick > 0 small enough that adding v to a row of ¯ Π2yields a perturbed matrix ¯ Π 2whose elements are also strictly positive. Clearly, ¯ Π 2 also satisfies (1.44), and hence the Ross price matrix is not unique, showing part 3. Proof of Proposition 3. Consider first the case where S < T. The dimension of the parameter set (transition probabilities + utility parameters) generating the generalized-recovery price matrix Π is ST −T+S, which is strictly greater than the dimension S2of the parameter space generating price matrices in Ross’s homogeneous case. Hence, generically no time-homogeneous solution can generate a generalized recovery price Π. Our framework is also more general in the the case S=T. Recalling that pτi denotes the probability of going from the current state 1 to state iin τperiods, it is clear that in a time-homogeneous setting we must have p22 ≥p11p12, i.e., the probability of going from state 1 to state 2 in two periods is (conservatively) bounded below by the probability obtained by considering the particular path that stays in state 1 in the first time period and then jumps to state 2 in the second. However, such a bound need not apply for the true probabilities if the transition probabilities are not 47
time-homogeneous. The set of parameters that can generate Π matrices that are not attainable from homogeneous transition probabilities is clearly of Lebesgue measure greater than zero in the S2−dimensional parameter space. Proof of Proposition 4. Let Rdenote the diagonal matrix whose k’th diagonal element is 1 (1+r)k. Having a flat term structure means that the matrix Π of state prices as seen from a particular starting state can be written as Π = RQ which defines Qas a stochastic matrix (i.e., with rows that sum to 1). Clearly, by letting δ= 1/(1 + r) and having risk-neutrality, i.e. H=IS(the identity matrix of dimension S), we obtain a solution to our recovery problem Π = RQ =DPH =RPIS=RP by setting P=Q. Proof of Proposition 5. The result follows from the following lemma. Lemma 1. Suppose that x∗∈Rnis defined by f(x∗)=0for a differentiable function f:Rn→Rnwith full rank of the Jacobian df in the neighborhood of x∗, and xis defined as the solution to the equation, f(¯x) + df(¯x)(x−¯x) = 0, where fhas been linearized around ¯x=x∗+ ∆x ε for ∆x∈Rnand ε∈R. Then x=x∗+o(ε)for ε→0. Proof of Lemma 1. Since we have x= ¯x−df−1f(¯x) we see that, as ε→0, x−x∗ ε=¯x−x∗ ε−df−1f(¯x)−f(x∗) ε→∆x−df−1df∆x= 0 (1.45) Proof of Proposition 6. Following the same logic as the proof of Proposition 1, we note that the set Xof all (δ, θ, P) is a manifold-with-boundary of dimension S· T−T+N+ 1. The discount rate, marginal utility parameters, and probabilities map into prices, which we denote by F(δ, θ, P) = DPH = Π, where, as before, D= 48
diag(δ, ..., δT) and H= diag(h1(θ), h2(θ), ..., hS(θ))), and Fis C∞. Since N+ 1 < T, the image F(X) has Lebesgue measure zero in RT×Sby Sard’s theorem, proving part 1. Turning to part 2, we first note that Pcan be uniquely recovered from (¯ θ, Π) using equation (1.12), where ¯ θ= (δ, θ). Therefore, we can focus on (¯ θ, Π), studying the solutions to Π(h−1 1(θ), ..., h−1 S(θ))0= (δ, ..., δT)0. For two different choices of the parameters (¯ θa,¯ θb) and a single set of prices Π, we consider the triplet (¯ θa,¯ θb,Π). We are interested in showing that the different parameters cannot both be consistent with the same prices, generically. To show this, we consider the space Mwhere the reverse is true, hoping to show that Mis “small.” Specifically, Mis the set of triplets where Π is of full rank and both discount rates are consistent with the prices, that is, there exists (unique) Pi(i=a, b) such that DaPaHa=DbPbHb= Π. Given that probabilities can be uniquely recovered from prices and parameters, we have a smooth map Gfrom Mto Xby mapping any triplet (¯ θa,¯ θb,Π) to (δa, θa, Pa). The image of this map consists exactly of those elements of Xfor which Fis not injective. The proof is complete if we can show that this image has Lebesgue measure zero, which follows again by Sard’s theorem if we can show that the dimension of M is strictly smaller than S·T−T+N+ 1. To study the dimension of M, consider first VΠ:= {Π(h−1 1(θ), ..., h−1 S(θ))0|θ∈Θ}, which is an N-dimensional submanifold of RTfor Π of full rank and given that his a one-to-one embedding. We note that we can think of Mas the space of triplets such that VΠcontains both the points (δa, δ2 a, ..., δT a)0and (δb, δ2 b, ..., δT b)0, where the corresponding θ’s are given uniquely from the definition of VΠsince Π is full rank and his one-to-one. The set of all those Π ∈RT×Ssuch that VΠpasses through two given points of RTform a subspace of dimension ST −2(T−N) since each point imposes T−Nequations. Therefore, Mis a manifold of dimension ST −2T+ 2N+ 2.Hence, we see that G(X) has measure zero in Xand F(G(X)) has measure zero in F(X). 49
B Details on Recovery in Mehra-Prescott Let Π = π0,d 0,1π1,u 0,10 0 0 0 0 . . . 0 0 0 0 . . . 0 0 0 π0,d 0,2π1,d 0,2π1,u 0,2π2,u 0,20. . . 0 0 0 0 . . . 0 . . .. . .. . .. . .. . .. . .. . .. . . . . .. . .. . .. . .. . . . . . 0 0 0 0 0 0 0 . . . 0π0,d 0,T π1,d 0,T π1,u 0,T . . . πT,u 0,T (1.46) where πk,u 0,t is the state price of making a total of k“up” moves in tperiods where the last move was “up,” that is, the Arrow-Debreu price for the state st= (yt, xt)=(ukdt−k, u). Similarly, πk,d 0,t is the state price of making a total of k“up” moves in tperiods where the last move was “down”. Π has dimension T×(PT t=1 2t). This implies that the h−1(γ) vector of inverse marginal utility ratios must be (PT t=1 2t)-dimensional. We fix this in the following way. We let h−1(γ) = h(y0 1)γ(y1 1)γ(y0 2)γ(y1 2)γ(y1 2)γ(y2 2)γ. . . (yT T)γi0 (1.47) where yk t=ukdt−kis the level of aggregate consumption when making a total of k “up” moves in tperiods and γis the risk-aversion parameter that we wish to recover. There is no closed-form solution to the non-linear case of CRRA preferences. In order to obtain model estimates we sort to a numerical exercise, that is to minimize the objective function g: min γ,δ g(γ, δ) := norm Πh−1(γ)− δ δ2 . . . δT (1.48) s.t. γ ∈R+ δ∈(0,1] Based on the recovered (γ, δ) that solve this minimizition problem, we can recover the 50
natural probabilities from (1.33). C Computing State Prices Empirically Before we can recover probabilities, we need to know the Arrow-Debreu prices or, said differently, characterize the risk-neutral distribution. There exist many ways to do this in practice based on observed option prices, including various interpolation methods. We implement two methods; (i) the parametric stochastic volatility model of Bates (2000) and (ii) the non-parametric “Fast and Stable” method of Jackwerth (2004). C.1 The Bates (2000) Stochastic Volatility Model with Jumps To ensure that we start with an arbitrage-free collection of Arrow-Debreu prices by strike and maturity, we use the model of Bates (2000) to derive state prices from observed option prices. This parametric approach puts structure on the tails of the risk-neutral density, which also allows us to extrapolate outside the range of observable option quotes. While the Bates (2000) model may not be the “true” specification of the economy, we simply use this framework as a standard method in the literature to compute state prices, and, consistent with this pragmatic view, we allow parameters to change over time (which also avoids look-ahead bias). In this model, the risk-neutral process for the price of the underlying asset, St, and the instantaneous variance, Vt, are assumed to be of the form dSt/St= (rf−d−λ¯ k)dt +pVtdZt+kdqt(1.49) dVt= (α−βVt)dt +σvpVtdZvt (1.50) where Ztand Zvt are Brownian motions with correlation ρ, and qtis a Poisson counting process that captures the risk of jumps in the price. The jumps occur with intensity λ and each jump causes the price to be multiplied by the factor 1+k, which is lognormally distributed, i.e., ln(1 + k)∼N(ln(1 + ¯ k)1 2δ2, δ2). Further, rfis the risk-free rate and dis the dividend yield. We calibrate these model parameters every fourth Wednesday as follows:17 On each 17We use data for every fourth Wednesday as a compromise between (i) the tradition in the asset pricing literature on return predictability of focusing on monthly returns, and (ii) the tradition in the 51
day, given the current level of the market Stand the risk-free term structure rf t,t+τ, we find the model parameters (α, β, λ, ¯ k, σv, δ) and state variable Vtthat minimize the vega-weighted squared pricing errors for fifty call and put options, following the methodology of Trolle and Schwartz (2009). The fifty chosen call/put options are those with the highest volumes. We allow the model parameters to vary over time since we simply use the model to smooth observed option prices (that may be noisy) such that they are arbitrage-free. Once we have obtained model estimates, we compute the risk-neutral density f(τ, Sτ) for any time τperiods into the future and state Sτgiven the current time state Stas: f(τ;Sτ) = 1 πZ∞ 0Sτ St−iu ψ(τ, u)du (1.51) that is, by integrating the characteristic function ψnumerically using the Gauss- Laguerre quadrature method. Knowing the risk-neutral density, the corresponding state price density π(τ;ST) is the density discounted by the τ-period risk-free rate rf t,t+τ: π(τ;Sτ) = e−rf t,t+τf(T;Sτ) (1.52) This completes the computation of state prices. Indeed, we think of π(τ;Sτ) as the Arrow-Debreu prices we need as starting point for our recovery for each index level. For example π(1,2000) is the Arrow-Debreu price of receiving $1 in one year of the S&P500 is between 2000 and 2001. We consider the grid of maturities and index levels described in Section 8.2. C.2 The Jackwerth (2004) “Fast and Stable” Method We are interested in converting a (noisy) sparse set of implied volatilities into a full risk-neutral distribution. In section C.1 we imposed a parametric form on the implied volatility surface through a stochastic volatility model with jumps. In this section we refrain from imposing any structure on implied volatilities, that is, we fit a nonparametric method to implied volatilities. The method we have chosen is the “Fast and Stable” method of Jackwerth (2004). This method has a single tuning parameter, λ, which simultaneously controls the smoothness of the function and the fit to observed option literature of focusing on Wednesdays, where among other reasons option liquidity is high. 52
implied volatilities. Clearly, there is a trade-off in choosing the value of the tuning parameter, which is: the smoother the function the worse the fit to observations. We therefore control the smoothness of the fit by imposing two conditions; (i) the estimated implied volatilities gives rise to a non-negative risk-neutral distribution, (ii) the risk-neutral distribution is unimodal in the range from 0.8 to 1.2 in moneyness (defined as St/S0, the index level at time trelative to the current index level). Under these conditions we minimize the objective function: min σs 1 2(S+ 1) S X s=1 σ00 s2+λ 2I I X i=1 (σi−¯σi)2(1.53) Where Sis the number of states. σsis the implied volatility associated with state s.σ00 sis the second derivative of the implied volatility function with respect to strike prices. i= 1, ..., I is the index for the observed implied volatilities and ¯σiis the i’th observed implied volatility. As seen from (1.53), if λis high then the fit to observations will be good compared to when λis low. We therefore choose the highest value of λ which satisfies our two conditions described above. See Jackwerth (2004) for further comments on the method. Once a smooth function for the implied volatilities is obtained we can back out a risk-neutral distribution by evaluating the Black and Scholes (1973) formula in the estimated implied volatilities and then differentiate the resulting call function twice with respect to strike prices as explained in Breeden and Litzenberger (1978). The Fast and Stable method estimates a single option maturity at a time. In the period from January 1996 until December 2015 we have at least 7 maturities on any given last trading day of the month. In the framework of Proposition 6this allows us to parameterize the pricing kernel with up to 6 parameters and still obtain generalized recovery. D Pricing Kernels used in Empirical Analysis Piecewise linear. The inverse marginal utilities are piecewise linear over states. Given the initial state 1 at time 0 the τ-period inverse marginal utility ratio in state 53
are usually “calm” times as measured by variance. We find that the HRI is positively related to the past two year return. The relation is statistically significant at a 99% level, showing that the return distribution is more left skewed and fat tailed subsequent to a “good” period where prices have increased significantly. Fourth, we show that higher-moment risk has large economic implications for investors. To understand the importance of higher-moment risk, we study the portfolio risk of a volatility-targeting investor who holds a portfolio of cash and the market. The investor adjusts the portfolio weights to achieve a constant volatility of σvol target. Despite having constant variance, the riskiness of the portfolio varies substantially over time as higher moment risk varies. Because higher moment risk is high when variance is low, the portfolio is the riskiest when market variance is low. To understand the economic magnitude of the systematic variation in highermoment risks, we estimate the probability that the return on the volatility-targeting investor’s portfolio is less than −2σvol target. The monthly probability peaked on June 30th 2014 with a probability of 9.7%, almost three times the size of its low, on February 27th 2008, where the probability was 3.6%. Furthermore, the average probability of a −2σvol target event is 6.6%, which is large compared to the 2.5% that is implied by a normal distribution. Similarly, the probability of a portfolio return that is less than −3σvol target peaked on November 30th 2006 with a probability of 3.6%, which is four times the size of its low on February 27th 2008, when the probability was 0.76%. These probabilities are also far above what is implied by a normal distribution, which is 0.13%. Furthermore, we find that the probability of a portfolio return that is less than −2σvol target for the volatility-targeting investor is negatively correlated with variance with a correlation coefficient of −0.70 and 95% bootstrapped confidence bounds of [−0.78,−0.65]. This strong negative correlation further emphasizes the importance of considering higher-moment risks in portfolio choice problems. For example, this finding can help explain why Moreira and Muir (2017b) find that investors can earn high Sharpe-ratios by moving wealth into the market at times when variance is low and moving wealth out of the market when variance increases. The relative (to variance) high expected return in calm times may be compensation for elevated higher-moment risks. 60
Our fifth main result shows how higher-moment risk is associated with several economic drivers. First, our results are closely related to the volatility paradox (Bunnermeier and Sannikov, 2014), which is the notion that systematic risk is high when variance is low. In their model, risk increases when variance is low because specialized investors are more levered. We therefore investigate how the level of financial intermediary leverage is associated with higher-moment risk. In particular, we test if financial intermediaries are more levered when variance is low, and if such variation in financial intermediary leverage can explain our observed variation in higher moment risk. Using the measure of financial intermediary leverage from He, Kelly, and Manela (2016), we find no relation between higher-moment risks and aggregate financial intermediary leverage. We next investigate how higher-moment risk is related to market illiquidity and funding illiquidity. We find that higher-moment risks are positively associated with both market and funding liquidity. Specifically, using the average value-weighted bidask spread of S&P 500 constituents as a proxy for market illiquidity, we find that times when the average bid-ask spread is low are times when higher-moment risks are high. Similarly, using the TED spread as a proxy for funding illiquidity, we find that a low TED spread is associated with high higher-moment risks. Lastly, we investigate how higher-moment risks are related to previously suggested measures of “bubble” characteristics and market valuation. We consider the “bubble” characteristics: acceleration (Greenwood, Shleifer, and You (2017)), turnover (Chen, Hong, and Stein (2001)), issuance percentage (Pontiff and Woodgate (2008)), and the market valuation measures: CAPE, the dividend-price ratio, and cay (Lettau and Ludvigson (2001)). We find that higher-moment risk is positively related to price acceleration: there is more higher-moment risk when the recent price path is more convex. Also, higher turnover after market run-ups is associated with more highermoment risk. Furthermore, there is more higher-moment risk when cay (Lettau and Ludvigson, 2001) is high. We find no conclusive relation between higher-moment risks and CAPE, the dividend-price ratio, or equity issuance. Our paper relates to and extends the existing literature on estimating time-varying market tail risk by integrating two different approaches. Previous research on tail risk is based on either (1) physical moments based on backward looking information 61
or (2) risk-neutral moments based on forward looking option prices. We show that physical higher-moment risks can be estimated in a forward looking manner, and in real time, which complements the existing literature that uses historical (backward looking) returns to estimate tail risks; e.g., using realized returns, Bollerslev and Todorov (2011) suggest using high frequency intraday returns and fit an extreme value distribution to the tails of returns. Also, Kelly and Jiang (2014) estimate market wide tail risks from the cross-section of firm-level returns. Our paper also relates to the literature that studies tail risk using option prices. However, while the existing literature studies tail risk using risk-neutral moments (e.g. Siriwardane (2015), Gao, Gao, and Song (2017), Gao, Lu, and Song (2017), Bates (2000), and Schneider and Trojani (2017b)), we study tail risk using physical moments. Thereby, we can investigate physical tail probabilities and study which economic drivers can explain the time-varying patterns in higher-moment risks. In summary, higher-moment risks can be measured in real time, and a single factor explains 90% of the joint variation in higher order moments. Furthermore, times when higher-moment risks are high are characterized by: (1) low variance, (2) large (and accelerating) recent price run-ups, (3) low market and funding frictions, (4) high turnover, and (5) low future expected returns. The paper proceeds as follows: Section 2covers the theory behind how we estimate higher order moments and tail probabilities. Section 3covers the data and the empirical implementation. Section 4investigates the relation between our ex ante moments and ex post realized moments. Section 5studies the commonalities in higher order moments. Section 6investigates the systematic patterns in higher-moment risks. Section 7studies the implications of time-varying higher-moment risks for investors. Section 8studies the economic drivers of higher-moment risks. Section 9concludes the paper. 2 Inferring Ex Ante Moments from Asset Prices We consider an economy where agents can trade two assets, a risk-free asset and a risky asset. The risk-free asset earns a gross risk-free rate of return Rf t,T between time tand time T. The risky asset has a price of Sand earns a random gross return Rt,T . The risky asset pays dividends, Dt,T , between time tand time Tsuch that its gross 62
return is Rt,T = (ST+Dt,T )/St. Starting from the standard asset pricing formula, we can relate risk-neutral and physical expected values of the time Trandom payoff, XT, as Et[XTmt,T ] = E∗ t[XT]/Rf t,T (2.1) where the asterisk denotes risk-neutral expectation and mt,T is a stochastic discount factor. If we define the time Trandom payoff, Xt,T (n), in the following way Xt,T (n) = Rn t,T m−1 t,T (2.2) then equation (2.1) implies that the n’th moment of the risky asset’s physical return distribution can be expressed in terms of the risk-neutral expectation of Xt,T (n): Et[Rn t,T ] = Et[Rn t,T m−1 t,T | {z } Xt,T (n) mt,T ] = E∗ t[Rn t,T m−1 t,T | {z } Xt,T (n) ]/Rf t,T (2.3) So if we know the pricing kernel m, then we can derive all moments of Rt,T directly from risk-neutral pricing of the claim to Xt,T (n). Following Martin (2017), we compute the physical expected value of Rn t,T from the point of view of an unconstrained rational power-utility investor who chooses to be fully invested in the market. This investor has initial wealth W0and terminal wealth WT=W0Rt,T . Given the investor’s utility function, U(x) = x1−γ/(1 −γ), with relative risk-aversion, γ, we can determine the investor’s stochastic discount factor. Specifically, combining the first order condition from the investor’s portfolio choice problem with the fact that the investor holds the market, the stochastic discount factor becomes proportional to R−γ t,T : mt,T =kR−γ t,T (2.4) for some constant kwhich is unobservable to us. However, we do not need to learn k to estimate physical moments; we can correct for kby rewriting (2.3) in the following way. First, setting n= 0 in (2.2) we get Xt,T (0) = m−1 t,T and the standard asset pricing 63
formula (2.1) then implies the relation: E∗ t[m−1 t,T ] = Rf t,T (2.5) Then, inserting (2.5) and (2.4) into (2.3), we obtain an expression of the n’th physical moment perceived by an unconstrained rational power utility investor who chooses to be fully invested in the market: Et[Rn t,T ] = E∗ t[Rn t,T m−1 t,T z }| { Rγ t,T /k] E∗ t[Rγ t,T /k | {z } m−1 t,T ]=E∗ t[Rn+γ t,T ] E∗ t[Rγ t,T ](2.6) since kis a constant. The relation between physical and risk-neutral moments shown in (2.6) is central to our empirical analysis. The key insight is that we can estimate the n’th physical moment directly from risk-neutral pricing of Rγ t,T and Rn+γ t,T . Furthermore, by pricing claims to the payoffs Rm+γ t,T for m∈ {1, ..., n}, we can then estimate standardized moments. To understand how we estimate standardized moments from (2.6), recall the notion of the n’th standardized moment formula: n’th standardized moment of Rt,T =EtRt,T −Et[Rt,T ] Var[Rt,T ]1/2n(2.7) Expanding (2.7) and replacing physical moments with risk-neutral counterparts as presented in equation (2.6), we can arrive at expressions for all physical standardized moments as functions of risk-neutral moments. For example, the third standardized physical moment (skewness) can be expressed in terms of risk-neutral moments by first expanding (2.7) with n= 3: Skewnesst,T =Et[R3 t,T ]−3Et[Rt,T ]Et[R2 t,T ]+2Et[Rt,T ]3 (Et[R2 t,T ]−Et[Rt,T ]2)3/2(2.8) and then replacing the physical moments in (2.8) with the risk-neutral counterparts using equation (2.6). Similar expressions can be written up for other higher order 64
moments of interest, as seen in Appendix A. Importantly, the right-hand-side of (2.6) consists of asset prices which can be estimated directly from current and observable call and put options written on the risky asset. Hence, higher order moments can be estimated in real time, without using historical realized returns or accounting data. 2.1 Inferring Ex Ante Market Tail Probabilities Next, we show how we estimate ex ante tail probabilities from option prices written on the market. To understand our approach, note first that the probability at time tof a market return that is lower than αat time Tcan be written as the physical expectation of an indicator function in the following way Pt(Rt,T < α) = Et[1{Rt,T <α}] (2.9) Using the standard asset pricing formula in (2.1), we can rewrite the probability in terms of the risk-neutral measure by adjusting the right hand side of equation (2.9) for the inverse of the stochastic discount factor in (2.4) Pt(Rt,T < α) = E∗ t[Rγ t,T 1{Rt,T <α}] E∗ t[Rγ t,T ](2.10) The right hand side of (2.10) is an asset price that has the simple representation presented in Proposition 8, which generalizes Result 2 in Martin (2017) from log-utility to general power utility for any level of relative risk-aversion. Proposition 8. For the unconstrained rational power utility investor who wants to hold the market, the conditional physical probability that market return from time tto Tis lower than αis: Pt(Rt,T < α) = Rf t,T E∗ t[Rγ t,T ]"αγput0 t,T (αSt−Dt,T )−γ St αγ−1putt,T (αSt−Dt,T ) (2.11) +ZαSt−Dt,T 0 γ(γ−1) S2 tK+Dt,T Stγ−2 putt,T (K)dK#(2.12) where put0 t,T (αSt−Dt,T )is the first derivative of the put option price with strike αSt− Dt,T . 65
Proof. The results of Breeden and Litzenberger (1978) imply the equality E∗ t[Rγ t,T 1{Rt,T <α}] = Rf t,T Z∞ 0K+Dt,T Stγ 1{K<αSt−Dt,T }put00 t,T (K)dK (2.13) where put00 t,T (K) is the second derivative of the put option price written on the underlying process S. Splitting the integral at αSt−Dt,T we have E∗ t[Rγ t,T 1{Rt,T <α}] = Rf t,T ZαSt−Dt,T 0K+Dt,T Stγ put00 t,T (K)dK (2.14) Proposition 8then follows from using integration by parts twice. 3 Data and Empirical Implementation We use the Ivy DB database from OptionMetrics to collect information on call options and put options that are written on the S&P 500 index for the last trading day of every month. The data ranges from January 1996 to December 2015. We obtain implied volatilities, strikes, closing bid-prices, closing ask-prices, and maturities. We proxy the risk-free rate with the zero-coupon yield curve from the Ivy DB database, which is derived from the LIBOR rates and settlement prices of CME Eurodollar futures. We also obtain expected dividend payments. We consider options with times to maturity between 10 and 360 calender days, and apply common filters, excluding contracts with zero open interest, zero trading volume, quotes with best bid below $0.50, and options with implied volatility higher than 100%. We use daily realized returns to estimate realized daily moments. We also estimate monthly moments from monthly returns. In Appendix A, we discuss the estimation of realized moments in detail. 3.1 Estimating Market Moments There is a large body of literature devoted to pricing asset derivatives such as those in (2.6), using observable option prices written on the asset. Indeed, Breeden and Litzenberger (1978), Bakshi and Madan (2000), and Bakshi, Kapadia, and Madan (2003) show that the arbitrage free price of a claim on some future (twice differentiable) payoff can be expressed in terms of a continuum of put and call option prices. Specifically 66
for our purposes, using the results of Breeden and Litzenberger (1978), Martin (2017) shows that we can write the n’th physical moment of Rt,T as Et[Rn t,T ] = E∗ t[Rn+γ t,T ] E∗ t[Rγ t,T ]=(Rf t,T )n+γ+Rf t,T [p(n+γ) + c(n+γ)] (Rf t,T )γ+Rf t,T [p(γ) + c(γ)] (2.15) with p(θ) = ZFt,T 0 θ(θ−1) Sθ tStRf t,T −Ft,T +Kθ−2putt,T (K)dK (2.16) c(θ) = Z∞ Ft,T θ(θ−1) Sθ tStRf t,T −Ft,T +Kθ−2callt,T (K)dK (2.17) where Ft,T is the forward price and callt,T (K) and putt,T (K) are call and put option prices written on the risky asset at time twith horizon T−tand strike K. In practice, we do not observe a continuum of call and put options and therefore (2.15) must be numerically approximated. Let Ft,T be the forward price and, using the notation from Martin (2017), we can write the price, Ωt,T (K), at time tof an out-of-the money option with strike Kand maturity Tas Ωt,T (K) = callt,T (K) if K≥Ft,T putt,T (K) if K < Ft,T (2.18) We let K1, ..., KNbe the (increasing) sequence of observable strikes for the Nout-of-the money put and call options and define ∆Ki=Ki+1−Ki−1 2with ∆Ki= Ki+1 −Kiif i= 1 Ki−Ki−1if i=N. (2.19) We approximate the integrals in (2.16) by observable sums such that the n’th physical moment becomes: Et[Rn t,T ] = (Rf t,T )n+γ+Rf t,T hPN i=1 (n+γ)(n+γ−1) Sn+γ(StRf t,T −Ft,T +Ki)n+γ−2Ωt,T (Ki)∆Kii (Rf t,T )γ+Rf t,T hPN i=1 γ(γ−1) Sγ(StRf t,T −Ft,T +Ki)γ−2Ωt,T (Ki)∆Kii (2.20) 67
In summary, combining equation (2.20) with the standardized moment formula in equation (2.7), we can express standardized physical moments in terms of the derivatives prices written on the risky asset. When we estimate physical moments for a given horizon, say T, for which we do not observe put and call prices, we linearly interpolate the (standardized) moments between the two closest horizons available in the data. In a few cases, we need to extrapolate to obtain moments for the desired horizon. Our benchmark investor has power utility and a coefficient of relative risk-aversion of 3, that is, γ= 3. This level of risk-aversion as the benchmark is motivated by the results of Bliss and Panigirtzoglou (2004), i.e., using our sample we replicate their results and find that 3 is the optimal option-implied level of risk aversion when matching realized returns at the monthly horizon. We also estimate moments for the risk-neutral investor, the log-utility investor, and the power utility investor with a risk-aversion coefficient of 5. Figure 2.1 shows monthly higher order moments and Table 2.1 shows the moment summary statistics. The average ex ante estimated skewness is negative for both horizons and all levels of risk aversion, suggesting that the physical distributions are left skewed. Consistent with the results of Neuberger (2012), we find that average skewness is not diminishing in the horizon, in the sense that skewness is close to the same on a monthly and quarterly horizon. Similarly, average kurtosis is larger than 3 for both horizons and all levels of risk aversion, which means that the physical distributions are leptokurtic; that is, the tails of the physical return distributions are fatter than what is implied by a normal distribution. 3.2 Estimating Market Tail Probabilities The main challenge when implementing Proposition 8is that we are required to estimate the first derivative of the put option price written on the risky asset at strike αSt−Dt,T . To handle a sparse and discrete set of observed option prices, we smoothen observed option prices using a Gaussian kernel smoothening procedure. Specifically, we smoothen implied volatilities around the strike αSt−Dt,T and choose the kernel bandwidth to minimize the squared errors between the observed and estimated im- 68
plied volatilities under the constraint that the estimated option prices do not allow for arbitrage. Given a smooth set of option prices around the strike αSt−Dt,T , we compute the first derivative as the slope between the two adjacent prices: put0 t,T (αSt−Dt,T ) = putt,T (αSt−Dt,T +h)−putt,T (αSt−Dt,T −h) 2h(2.21) where his the chosen grid step size in the discretization. Let K1, ..., KMbe the (increasing) sequence of observable strikes for the Mout-of- the money put options where KMis the observed strike that is closest to αSt−Dt,T . We approximate the integral in Proposition 8by the observable sum: M X i=1 γ(γ−1) S2 tKi+Dt,T Stγ−2 putt,T (Ki)∆Ki(2.22) Inserting (2.21) and (2.22) into Proposition 8, we can estimate physical probabilities. 4 Estimated Moments Predict Realized Moments In this section, we show that the ex ante higher order moments estimated using the methods described in Sections 2and 3predict ex post realized higher order moments. We start with a simple sorting exercise. For each moment, we first sort ex post realized monthly returns into a “low” or “high” bucket depending on whether the ex ante moment is lower or higher than its median time series value. Next, we estimate the ex post moments for each bucket; for example, we estimate moments using the monthly ex post returns sorted into the “high” bucket. Figure 2.2 shows the monthly ex post realized moments of the two buckets for all moments. The ex post realized returns sorted into the “high” buckets exhibit in-sample higher moment values, suggesting that our ex ante moments predict ex post moments, for example, the “high” bucket for kurtosis has an in-sample kurtosis of 5.93, while the “low” bucket has a kurtosis of 2.85. Next, we test more formally the relation between ex ante and ex post moments. Specifically, we conduct two tests which differ in the way we estimate ex post realized higher order moments. First, we test if the bucket values following our sorting exercise 69
between market run-ups and subsequent higher-moment risks we run a set of regressions of ex ante moments onto the past two year return,2rt−24,t =Rt−24,t −1: Mt,T =β0+β1rt−24,t +t,T (2.28) where the moments, Mt,T , are variance, skewness, kurtosis, hyperkurtosis, hyperskewness, and the higher-moment risk index (HRI). Panel B of Table 2.6 shows the β1 coefficients of regression (2.28) and in Panel C of Table 2.6 we show β1coefficients of regression (2.28) when controlling for the lagged ex ante moment. We find a negative and significant relation between past returns and variance. This finding is consistent with the intuition that times after market run-ups are “calm” times where risk, as measured by variance, is low. Looking at skewness, we find a statistically significant and negative relation with past returns, implying that the return distribution tilts to the right and leaves more probability mass in the left tail of the return distribution subsequent to market run-ups. Similarly, kurtosis is statistically significant and positive in past returns, hyperskewness is negative in past returns, and hyperkurtosis is positive in past returns. The results are quantitatively similar for monthly and quarterly moments. Panel C of Table 2.6 shows that controlling for lagged risk does not change our results. We still find strong significant systematic variation in higher order moments. 7 Implications for Investors The results presented in Table 2.4, Table 2.5, and Table 2.6 show that times when variance is low are times when the market’s return distribution is highly left skewed (due to large negative skewness and hyperskewness) and fat tailed (due to large positive kurtosis and hyperkurtosis). That higher-moment risks are high at times when variance is low runs counter to the way we usually think about risk, i.e., we often equate risk with variance, saying that risk is high at times when variance is high. To better understand the importance of higher-moment risks, we next investigate portfolio risks for two investors who both hold a portfolio of cash and the market. The first investor holds a constant notional in the market. The probability that 2This is similar to the market run-up period of Greenwood, Shleifer, and You (2017). 76
the investor’s portfolio realizes an unexpected return (the shock to the portfolio), rshock t,T =Rt,T −Et[Rt,T ], less than αis: Pt(rshock t,T < α) (2.29) The constant notional investor is exposed to both time-varying variance risk and timevarying higher-moment risks; that is, the probability that the portfolio realizes an unexpected return less than αdepends on both conditional variance and conditional higher order moments. The second investor targets a constant level of portfolio volatility, i.e., the investor moves wealth in and out of the market such that the portfolio has constant volatility. Such volatility-targeting strategies are common practice and have been shown to generate high risk-adjusted returns (e.g. Moskowitz, Ooi, and Pedersen (2012), Asness, Frazzini, and Pedersen (2012), Moreira and Muir (2017a), and Moreira and Muir (2017b)). If σt,T is the market’s conditional volatility, rt,T =Rt,T −1 is the return on the market, and rf t,T is the risk-free rate of return, then rvol target t,T is the return on the volatility-targeting investor’s portfolio who targets a constant volatility of σvol target: rvol target t,T =σvol target σt,T | {z } ωt,T rt,T +1−σvol target σt,T rf t,T (2.30) where ωt,T is the fraction of wealth held in the market. If ωt,T >1, the investor levers up by borrowing cash to invest more than all the initial wealth in the market. We assume for simplicity that the investor is unconstrained. The unexpected return of the volatility-targeting investor’s portfolio is rvol target,shock t,T =rvol target t,T −Et[rvol target t,T ] which can be rewritten as: rvol target,shock t,T =ωt,T rt,T + (1 −ωt,T )rf t,T −ωt,T Et[rt,T ] + (1 −ωt,T )rf t,T (2.31) =ωt,T rshock t,T (2.32) The probability that the volatility-targeting investor’s portfolio realizes an unexpected 77
return less than αis: Pt(rvol target,shock t,T < α) = Ptσvol target σt,T rshock t,T < α(2.33) =Ptrshock t,T <α σvol target σt,T (2.34) For example, if σvol target = 5% and α=−10%, then the probability that the volatilitytargeting investor’s portfolio realizes a return that is 10% lower than expected is Ptrshock t,T <−2σt,T .3The volatility-targeting investor’s portfolio is only exposed to time-varying higher-moment risks, that is, given a level of σvol target, the probability that the investor’s portfolio realizes a return less than αdepends only on conditional higher order moments. Time-varying variance risk is eliminated by targeting a constant level of portfolio volatility. Recall that σt,t+his the ex ante volatility from time tto t+h, and we then define ¯σhas the time series average of σt,t+h. For example, the time-series average of monthly volatility for the S&P 500 index is ¯σh= 5.0%. Figure 2.5 shows time-series plots of monthly probabilities, as shown in (2.29) and (2.33), where α=−2¯σmonth =−10.1% and the volatility-target is σvol target = 5.0%. The top figure shows the probabilities of −2σt,t+1 drops in the market, which are the probabilities of the volatility-targeting investor’s portfolio return. The horizontal line shows the probability of a −2σt,t+1 drop in the market implied by a normal distribution, which is 2.5%. The shaded area between the two lines is higher-moment risk; that is, the excess probability of a tail event due to negative skewness, excess kurtosis, and all other higher order moments. Interestingly, the probabilities, in excess of what is implied by a normal distribution, range from 1.1% to 7.2%, showing that time-varying higher-moment risks have large economic implications for the risk of the volatilitytargeting investor’s portfolio. The probability of a −2σt,t+1 drop peaked on June 30th 2014 with a probability of 9.7%, almost three times the size of its low on February 27th 2008, where the probability was 3.6%. The systematic variation in the tail probabilities, from 3.6% at high variance times to 9.7% at low variance times, emphasizes that investors who manage risk by managing variance are implicitly imposing more risk 3Notice that this probability is not necessarily the same as the probability of a portfolio return of −10%. In the example, the probability of a portfolio return of −10% is Ptrshock t,T <−2σt,T −Et[rt,T ]. 78
into their portfolio when variance is low. The bottom figure shows the probabilities of −2¯σ=−10.1% drops in the market along with the probabilities implied by a normal distribution. The shaded area between the two lines is higher-moment risk for the constant notional investor. The probability of a −10.1% drop in the market is, as expected, high when variance is high. Importantly, higher-moment risk also contributes to the portfolio risk for the constant notional investor, and the economic magnitude is large. For example, the probabilities, in excess of what is implied by a normal distribution, range from 0.5% on October 31st 2006 to 4.8% on August 31st 2015. On August 31st 2015, the total probability of a −10.1% drop was 9.90%, which means that, on that day, 48% of the probability mass in the left tail of the return distribution beyond −10.1% was due to higher-moment risk. Figure 2.6 shows time-series plots of monthly probabilities, as shown in (2.29) and (2.33), where α=−3¯σmonth =−15.1%. The probability of a portfolio return that is less than −3σt,t+1 peaked on November 30th 2006 with a probability of 3.6%, which is four times the size of its low on February 27th 2008, where the probability was 0.8%. These probabilities are far from what is implied by a normal distribution, which is 0.13%. Specifically, the average probability of a −3σt,t+1 event is 1.8%, which is fourteen times higher than what is implied by the normal distribution. Figure 2.6 shows that higher-moment risk is even more important when evaluating the probability of events further out in the lower tail of the return distribution; that is, the relative amount of probability mass in the lower tail that is due to higher-moment risk increases the further we go out in the tail. The probabilities co-move in the sense that, when the probability of a −2σt,T event is high, then the probability of a −2¯σevent is low. To further investigate these patterns, Panel A of Table 2.7 reports correlations between variance and the probability of a portfolio return that is less than αfor the constant notional investor and the volatilitytargeting investor. The first column of Panel A shows the correlations between variance and the probability that the market realizes an unexpected return less than −2¯σh(the probability of a constant notional investor) Pt(rshock t,t+h<−2¯σh) (2.35) 79
The correlations range from 0.94 to 0.98 with tight bootstrapped confidence bounds, showing that the conditional probability that the market realizes a return less than −10.1% monthly or −15.1% quarterly is highly correlated with conditional variance, which would be expected just from looking at Figure 2.5. Panel A of Table 2.7 also reports correlations between variance and the probability of a portfolio return that is less than αfor the volatility-targeting investor. Specifically, we estimate the correlations between variance and the probabilities of a −2σt,t+h Pt(rshock t,t+h<−2σt,t+h) (2.36) This probability is equivalent to the probability in (2.33) with α σvol target =−2. Interestingly, the correlations in the last two column of Panel A are all negative and range from −0.70 to −0.44, with tight bootstrapped confidence bounds. These high negative correlations show that the portfolio of the volatility-targeting investor is most risky at times when variance is low, even though the investor has eliminated all dependencies on variance in the portfolio. This finding can help explain why Moreira and Muir (2017a) and Moreira and Muir (2017b) find that investors can earn high Sharpe ratios by moving wealth into the market at times of low variance and moving wealth out of the market when variance increases (in some sense mimicking a volatility targeting strategy). The relatively (to variance) high expected return in calm periods may be compensation for the elevated higher-moment risks. To better understand the systematic variation in higher-moment risks, we next investigate the relation between tail probabilities and past returns. Specifically, we regress tail probabilities onto past two year returns, e.g. the probability of a −2σt,t+1 drop as Pt(rshock t,t+h<−2σt,t+h) = β0+β1rt−24,t +t,T (2.37) Panel B of Table 2.7 reports β1coefficients from regressions such as in (2.37). We find that the probability of both a −2σt,t+1 and a −3σt,t+1 drop in the market is statistically significant and positively related to past returns. The economic magnitude is such that a 50% market run-up over the past two years implies a 1% higher probability of 80
a monthly −2σt,t+1 drop in the market. Furthermore, the monthly probability of a −10% drop in the market is negatively related to past returns, which is to be expected, because this probability is highly correlated to variance, as shown in Table 2.6, and periods after market run-ups are usually associated with low variance. Panel C of Table 2.7 reports β1coefficients from regressions such as in (2.37) when controlling for lagged probabilities. Controlling for lagged probabilities does not change our results: high past two year returns imply higher current tail probabilities for the volatilitytargeting investor. Our finding that market run-ups are related to contemporaneously higher highermoment risks supplements the existing literature that relates market run-ups to subsequent (realized) market “crashes”, e.g. Greenwood, Shleifer, and You (2017). Specifically, we find that the probability of an x% drop in the market decreases in past returns. High past returns means low current volatility and a low probability of a subsequent x% drop in the market price. However, conditional on variance, the probability of an x% drop in the market increases in past returns. 8 What Explains Higher-Moment Risk? In this section we investigate three possible explanations for the systematic variation in higher-moment risk. First, we investigate how higher-moment risk is associated with financial intermediary leverage. Second, we study how market and funding liquidity relates to higher-moment risk. Third, we investigate how higher-moment risk is associated with common “bubble” characteristics. Throughout this section, we will focus on monthly horizon ex ante higher-moment risk. 8.1 The Volatility Paradox and Intermediary Leverage The volatility paradox is the phenomenon that endogenous risk is high even though exogenous risk is low (Brunnermeier and Sannikov (2014)). Loosely speaking, exogenous risk can be seen as variance and endogenous risk can be seen as higher-moment risk. When variance is low, investors take on more risk in their positions, for instance through leverage, which creates endogenous risk. This negative relation between higher-moment risk and variance is closely related to our empirical findings, 81
we therefore test if our finding can be linked to the economic drivers suggested by Brunnermeier and Sannikov (2014). One way in which this endogenous risk may arise is through intermediary leverage.4 We test if financial intermediary leverage can help explain higher-moment risks by running the following regression: Mt,T =β0+β1Leveraget+t,T (2.38) where the risk, Mt,T , is variance, skewness, kurtosis, hyperkurtosis, hyperskewness, and the higher-moment risk index (HRI). Leverage is the financial intermediary leverage ratio of He, Kelly, and Manela (2016). Regression (2.38) relates aggregate financial intermediary leverage to contemporaneous higher-moment risks. Panel A of Table 2.8 shows the results of regression (2.38). We find that leverage is positively associated with contemporaneous ex ante variance, which is consistent with financial intermediary leverage being counter-cyclical, as noted in He, Kelly, and Manela (2016). The first column of Panel A shows that aggregate financial intermediary leverage is not related to the HRI: we find a regression coefficient pf −0.15 which is statistically insignificant. Decomposing higher-moment risks into individual moments, we do not find a significant relation between financial intermediary leverage and individual higher order moments. Overall, aggregate leverage does not help explain higher-moment risks. Next, we test if conditional (on variance) financial intermediary leverage is associated with higher-moment risks. We run the regression: Mt,T =β0+β1Leveraget+β2Variancet,T +t,T (2.39) Panel B of Table 2.8 reports the results of regression (2.39). Interestingly, we find that, conditioning on ex ante variance, financial intermediary leverage can help explain contemporaneous higher-moment risks. We find that skewness, kurtosis, hyperskewness, and hyperkurtosis all load statistically significantly on financial intermediary leverage, with negative signs for skewness and hyperskewness and positive signs for kurtosis 4Several papers have shown that financial intermediary leverage is associated with asset returns, e.g. He, Kelly, and Manela (2016), He and Krishnamurthy (2013), Adrian and Boyarchenko (2012), and Adrian, Etula, and Muir (2014). 82
and hyperkurtosis. Furthermore, the HRI is positively related to conditional financial intermediary leverage. Given a level of ex ante variance, higher leverage is associated with higher contemporaneous higher-moment risks. Panel C in Table 2.8 reports regression (2.39) when controlling for lagged risk. Controlling for lagged risk does not change our results. Aggregate financial intermediary leverage is in general not associated with higher-moment risks. Given a level of ex ante variance, and controlling for lagged risk, higher leverage is associated with higher contemporaneous higher-moment risks. 8.2 Market Liquidity and Funding Liquidity. Several previous papers link market liquidity and funding liquidity to aspects of the stock market’s return distribution. Christoffersen, Feunou, Jeon, and Ornthanalai (2016) suggest market illiquidity as an economic factor driving risk-neutral market variance and jump risks, or equivalently, higher order moments. They argue that market illiquidity is the common culprit of market price drops in cases when the price drop happened without news about fundamentals, and it is therefore a reasonable economic driver of market moments. Brunnermeier and Pedersen (2009) show that, from a theoretical point of view, stocks with low market (and funding) liquidity have high variance because they are associated with high margin requirements. Furthermore, Danilova and Julliard (2015) develop a model in which volatility and illiquidity are jointly determined by the same equilibrium forces. First, we test if high market illiquidity is associated with high contemporaneous ex ante variance. Thereafter, we investigate the relation between market illiquidity and higher-moment risks. When testing the relation between higher-moment risks (or variance) and market illiquidity, we run the regression: Mt,T =β0+β1Bid-ask spreadt+t,T (2.40) where the risk, Mt,T , is variance, skewness, kurtosis, hyperkurtosis, hyperskewness, and the higher-moment risk index (HRI). As a proxy for market illiquidity, we follow Christoffersen, Feunou, Jeon, and Ornthanalai (2016), and use the average value- 83
weighted bid-ask spread of constituents of the S&P 500 index. Panel A of Table 2.9 reports the results of regression (2.40). We find that higher market illiquidity is associated with higher contemporaneous ex ante variance, which is consistent with the model of Brunnermeier and Pedersen (2009). The effect is statistically significant at a 99% level and controlling for lagged variance does not change the result. The HRI is negatively related to market illiquidity with a regression coefficient of −1.22, which is statistically significant at a 99% level. When we control for lagged HRI, we still get a negative relation between market illiquidity and higher-moment risks, but the relation is insignificant. The negative relation between the HRI and market illiquidity shows that higher-moment risks tend to be high at times when the market is most liquid. Next, we test the relation between funding illiquidity and higher-moment risks. We run the regression: Mt,T =β0+β1TED spreadt+t,T (2.41) where the risk, Mt,T , is variance, skewness, kurtosis, hyperkurtosis, hyperskewness, and the higher-moment risk index (HRI). The TED spread is a common proxy for funding illiquidity, e.g. Frazzini and Pedersen (2014). The TED spread is the three month LIBOR intrabank interest rate minus the three month T-bill interest rate and it is available from the St. Louis FED. Panel B of Table 2.9 reports the results of regression (2.41). Contemporaneous ex ante variance is positively related to funding illiquidity, higher TED spread is associated with higher ex ante variance. We find that the HRI is negatively related to funding illiquidity which means that, higher-moment risks are high at times when there is low friction in the funding market. Controlling for the lagged HRI does not change our result. Figure 2.7 shows time-series plots of market illiquidity and funding illiquidity with the HRI. Consistent with the results presented in Table 2.9, we see that the HRI is negatively correlated with both the bid-ask spread and the TED spread. 84
8.3 “Bubble” Characteristics A range of macroeconomic variables have been proposed as possible indicators of increased market “crash” risks, or equivalently, increased higher-moment risks. A partial list of the variables include the suggestion of Chen, Hong, and Stein (2001), who suggest turnover, Pontiff and Woodgate (2008), who use issuance as a characteristic, and Greenwood, Shleifer, and You (2017), who propose price acceleration as a highermoment risk characteristic. In this section we investigate the relation between common market “crash” indicators and contemporaneous ex ante higher-moment risks. We therefore run regressions on the form: HRIt,T =β0+β1Characteristict+t,T (2.42) where the “bubble” characteristics are: 1) The Greenwood, Shleifer, and You (2017) variable acceleration, which is defined as the annualized past two year return minus the return of the first year of the two year return. Acceleration captures the convexity in the recent price path and a high value of acceleration is intended to be associated with high contemporaneous ex ante higher-moment risks. 2) Issuance as the percentage of firms in the S&P 500 index that issued equity in the past year. We follow Greenwood, Shleifer, and You (2017), and define an equity issuance as the event that a firm’s split-adjusted share count increased by five percent or more. 3) Market turnover. The market valuation measures are: 4) CAPE, the Shiller cyclically adjusted price-earnings ratio. 5) The dividend price ratio as the past two year dividends divided by the current market price. 6) Cay, the Lettau and Ludvigson (2001) log consumption - aggregate wealth ratio. Table 2.10 reports the results of regression (2.42). Marginally, we find that cay is negatively and significantly related to the HRI. Calm times when expected returns, as proxied by cay, are low are times when higher-moment risks are high. Interacting turnover with the past two year return, we find that turnover is positively related to the HRI. This finding is consistent with the findings in Chen, Hong, and Stein (2001), that is, subsequent to market run-ups, a higher turnover is associated with higher higher-moment risks. For issuance, we find that, subsequent to market run-ups, a higher level of equity 85
0.00 0.02 0.04 0.06 0.08 0.10 0.12 Month Probability of a −2σt month event 1/1996 1/1998 1/2000 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 0.00 0.02 0.04 0.06 0.08 0.10 0.12 Probability of a −2σt month event Probability of a −2σt month event − normal distribution 0.00 0.05 0.10 0.15 0.20 0.25 Month Probability of a −2σmonth event 1/1996 1/1998 1/2000 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 0.00 0.05 0.10 0.15 0.20 0.25 Probability of a −2σmonth event Probability of a −2σmonth event − normal distribution cor = −0.68 Figure 2.5: Market tail loss probabilities – two sigma. The top figure shows portfolio tail loss probabilities for the volatility-targeting investor; that is, the probability of an unexpected return lower than −2σmonth t. The dashed line is the tail loss probabilities implied by a normal distribution. The shaded area between the lines is higher-moment risk, that is, the part of the tail loss probability that is entirely driven by changes in higher order moments. The bottom figure shows portfolio tail loss probabilities for the constant notional investor. Here, σmonth tis the conditional monthly ex ante variance and ¯σmonth is the time-series average of σmonth t. In our sample, ¯σmonth = 5.0%. 92
0.00 0.01 0.02 0.03 0.04 Month Probability of a −3σt month event 1/1996 1/1998 1/2000 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 0.00 0.01 0.02 0.03 0.04 Probability of a −3σt month event Probability of a −3σt month event − normal distribution 0.00 0.05 0.10 0.15 Month Probability of a −3σmonth event 1/1996 1/1998 1/2000 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 0.00 0.05 0.10 0.15 Probability of a −3σmonth event Probability of a −3σmonth event − normal distribution cor = −0.68 Figure 2.6: Market tail loss probabilities – three sigma. The top figure shows portfolio tail loss probabilities for the volatility-targeting investor; that is, the probability of an unexpected return lower than −3σmonth t. The dashed line is the tail loss probabilities implied by a normal distribution. The shaded area between the lines is higher-moment risk, that is, the part of the tail loss probability that is entirely driven by changes in higher order moments. The bottom figure shows portfolio tail loss probabilities for the constant notional investor. Here, σmonth tis the conditional monthly ex ante variance and ¯σmonth is the time-series average of σmonth t. In our sample, ¯σmonth = 5.0%. 93
−5 0 5 10 Higher−moment risk index 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 −1 0 1 2 3 Bid−ask spread Higher−moment risk index Bid−ask spread cor = −0.59 −2 0 2 4 6 8 10 Higher−moment risk index 1/1996 1/1998 1/2000 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 0.0 0.5 1.0 1.5 2.0 2.5 3.0 TED spread Higher−moment risk index TED spread cor = −0.31 Figure 2.7: Higher-moment risk index and market and funding illiquidity. The top figure shows time series plots of the HRI and market illiquidity (proxied by the average value-weighted bid-ask spread of S&P 500 constituents). The bottom figure shows time series plots of the HRI and funding illiquidity (proxied by the TED spread). 94
−5 0 5 10 Higher−moment risk index 1/1996 1/1998 1/2000 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 −0.04 0.00 0.02 0.04 0.06 0.08 Cay Higher−moment risk index Cay cor = −0.53 Figure 2.8: Higher-moment risk index and cay. The figure shows time series plots of the HRI and cay. −5 0 5 10 Higher−moment risk index 1/1996 1/1998 1/2000 1/2002 1/2004 1/2006 1/2008 1/2010 1/2012 1/2014 1/2016 −400 −200 0 200 400 600 Turnover×r(t−24, t) Higher−moment risk index Turnover×r(t−24, t) cor = 0.44 Figure 2.9: Higher-moment risk index and turnover. The figure shows time series plots of the HRI and turnover times the past two year return. 95
Table 2.1: Moment Summary Statistics. In this table we report the average time-series values for ex ante estimated moments: excess return (ER−Rf), standard deviation (St. dev.), skewness (Skew), kurtosis (Kurt), hyperskewness (Hskew), and hyperkurtosis (Hkurt). We estimate ex ante moments from the point of view of a riskneutral investor (γ= 0), a log-utility investor (γ= 1), and two power-utility investors (γ= 3, γ= 5). Annualized (%) Horizon Risk-aversion ER−Rf St. dev. Skew Kurt Hskew Hkurt Month γ= 0 0 21.07 -1.45 8.90 -46.58 347.41 Month γ= 1 4.44 19.89 -1.31 8.25 -41.01 307.44 Month γ= 3 12.00 18.33 -1.08 7.20 -31.16 233.07 Month γ= 5 18.36 17.32 -0.89 6.43 -23.47 175.88 Quarter γ= 0 0 21.07 -1.17 5.78 -20.58 110.36 Quarter γ= 1 4.44 19.59 -1.09 5.57 -18.64 97.45 Quarter γ= 3 11.48 17.55 -0.95 5.23 -15.56 80.69 Quarter γ= 5 17.12 16.25 -0.81 4.94 -12.93 68.19 96
Table 2.2: Ex Ante Conditional Moments Predict Ex Post Realized Moments. Panel A reports ex post moments for monthly returns sorted into a low or high bucket based on the ex ante moment. Panel B reports correlations between our ex ante moments and ex post realized moments. Panel C reports correlations between our ex ante kurtosis and hyperkurtosis with ex post left kurtosis and left hyperkurtosis. We also report correlations between historical moments and ex post moments. We report bootstrapped standard errors in the appendix and significance as; * when p < 0.1, ** when p < 0.05, and *** when p < 0.01. Panel A: Sorting on ex ante monthly moments Our moments Historical moments Low ex ante High ex ante Low ex ante High ex ante Variance (%) 0.08∗∗∗ 0.31∗∗∗ 0.07∗∗∗ 0.28∗∗∗ Skewness −0.83 −0.48 −0.49 −0.75 Kurtosis 2.85 5.93∗∗ 3.35 4.46 Hyperskewness −19.03∗∗ −4.24 −5.40 −9.01 Hyperkurtosis 15.70 95.24∗∗∗ 15.34 40.10 Panel B: Correlation between ex ante moments and ex post realized moments Our moments Historical moments Month Quarter Month Quarter Variance 0.67∗∗∗ 0.49∗∗∗ 0.72∗∗∗ 0.46∗∗∗ Skewness 0.21∗∗∗ 0.25∗∗∗ 0.07 0.24∗∗∗ Kurtosis −0.01 0.00 0.06 0.04 Hyperskewness 0.17∗∗∗ 0.20∗∗∗ 0.07 0.15∗∗∗ Hyperkurtosis −0.03 0.06 0.03 0.01 Panel C: Left kurtosis and left hyperkurtosis Our moments Historical moments Month Quarter Month Quarter Left kurtosis 0.19∗∗∗ 0.26∗∗∗ 0.02 0.14∗∗ Left hyperkurtosis 0.17∗∗∗ 0.21∗∗∗ 0.05 0.13∗∗ 97
Table 2.3: Ex Ante Conditional Moments Predict Ex Post Realized Moments — Robustness. Panel A reports correlations between our ex ante moments and ex post realized moments when we remove observations that overlap with the period from August 1st 2008 to July 31st 2009. Panel B reports correlations between our ex ante moments (estimated with different levels of relative risk aversion) and ex post realized moments. Panel C reports correlations when controlling for risk-neutral moments or historical moments. We report bootstrapped standard errors in the appendix and significance as; * when p < 0.1, ** when p < 0.05, and *** when p < 0.01.. Panel A: Excluding August 1st 2008 to July 31st 2009 Our moments Historical moments Month Quarter Month Quarter Variance 0.52∗∗∗ 0.49∗∗∗ 0.51∗∗∗ 0.42∗∗∗ Skewness 0.23∗∗∗ 0.23∗∗∗ 0.11 0.25∗∗∗ Kurtosis −0.01 0.00 0.07 0.06 Hyperskewness 0.18∗∗∗ 0.18∗∗∗ 0.09 0.13∗∗∗ Hyperkurtosis −0.03 0.04 0.04 0.01 Panel B: Other levels of risk-aversion γ= 1 γ= 5 Month Quarter Month Quarter Variance 0.67∗∗∗ 0.48∗∗∗ 0.67∗∗∗ 0.48∗∗∗ Skewness 0.20∗∗∗ 0.25∗∗∗ 0.21∗∗∗ 0.23∗∗∗ Kurtosis −0.03 0.00 0.01 0.02 Hyperskewness 0.13∗∗∗ 0.17∗∗∗ 0.20∗∗∗ 0.22∗∗∗ Hyperkurtosis −0.06 0.04 −0.01 0.07 Panel C: Marginal correlations Controlling for Controlling for risk-neutral moments historical moments Month Quarter Month Quarter Variance 0.09 0.12 0.18∗∗ 0.20∗∗∗ Skewness 0.09 0.01 0.20∗∗∗ 0.17∗∗∗ Kurtosis 0.11∗∗ −0.02 −0.02 0.01 Hyperskewness 0.16∗∗∗ 0.14∗∗∗ 0.16∗∗∗ 0.17∗∗ Hyperkurtosis 0.09∗∗ 0.07 −0.03 0.05 98
Table 2.4: Correlations Between S&P 500 Moments. Panel A reports pairwise correlations between monthly S&P 500 moments. Expected return (Er), variance (Var), skewness (Skew), kurtosis (Kurt), hyperskewness (Hskew), and hyperkurtosis (Hkurt). Panel B shows the correlation between quarterly horizon moments. We report 95% bootstrapped confidence bounds in brackets. Panel A: Month Er Var −Skew Kurt −Hskew Hkurt Er 1 0.99 −0.46 −0.50 −0.48 −0.41 [0.99,1] [−0.56,−0.37] [−0.57,−0.46] [−0.55,−0.43] [−0.48,−0.37] Var 1 −0.52 −0.54 −0.51 −0.44 [−0.60,−0.43] [−0.60,−0.50] [−0.58,−0.47] [−0.51,−0.40] -Skew 1 0.80 0.78 0.66 [0.76,0.84] [0.74,0.82] [0.60,0.72] Kurt 1 0.97 0.93 [0.95,0.98] [0.90,0.95] -Hskew 1 0.98 [0.97,0.98] Hkurt 1 Panel A: Quarter Er Var −Skew Kurt −Hskew Hkurt Er 1 0.99 −0.46 −0.54 −0.58 −0.56 [0.99,0.99] [−0.55,−0.37] [−0.60,−0.49] [−0.64,−0.54] [−0.62,−0.52] Var 1 −0.54 −0.62 −0.65 −0.62 [−0.62,−0.47] [−0.67,−0.57] [−0.71,−0.61] [−0.68,−0.57] -Skew 1 0.83 0.86 0.74 [0.79,0.86] [0.82,0.90] [0.68,0.79] Kurt 1 0.96 0.94 [0.95,0.97] [0.92,0.96] -Hskew 1 0.97 [0.96,0.98] Hkurt 1 99
Table 2.5: Principal Components of Higher-Moment Risks. We estimate the four principal components (PC) spanning the space of monthly (Panel A) and quarterly (Panel B) skewness (Skew), kurtosis (Kurt), hyperskewness (Hskew), and hyperkurtosis (Hkurt). Panel A reports the loadings on each of the monthly moments. Panel B reports the loadings on each of the quarterly moments. The last column of Panel A shows that the first principal component (PC 1) explains 89% of the variation in monthly higher order moments. Similarly, the last column of Panel B shows that 91% of the variation in quarterly higher order moments is captured by the first principal component. Panel A: Month Skew Kurt Hskew Hkurt Variation explained PC 1 eigenvector −0.45 0.52 −0.52 0.50 89% PC 2 eigenvector 0.85 0.07 −0.20 0.48 10% PC 3 eigenvector −0.23 −0.83 −0.16 0.48 1% PC 4 eigenvector −0.13 0.19 0.81 0.54 0% PC 1 correlation −0.85 0.98 −0.99 0.95 Panel B: Quarter Skew Kurt Hskew Hkurt Variation explained PC 1 eigenvector −0.47 0.51 −0.52 0.50 91% PC 2 eigenvector −0.84 −0.16 0.11 −0.51 7% PC 3 eigenvector −0.13 −0.84 −0.32 0.40 2% PC 4 eigenvector −0.25 0.01 0.78 0.57 0% PC 1 correlation −0.89 0.98 −0.99 0.96 100
Table 2.6: Cyclicality in Higher-Moment Risks. Panel A reports correlations between ex ante variance and the higher-moment risk index (HRI). We report bootstrapped 95% confidence intervals in brackets. Panel B reports β1coefficients when regressing physical moments onto the past two year returns: Mt,T =β0+β1rt−24,t +t,T where the moment Mt,T is variance (Var), skewness (Skew), kurtosis (Kurt), hyperskewness (Hskew), hyperkurtosis (Hkurt), and the higher-moment risk index. Panel C reports the regression when controlling for lagged moments. We report t-statistics in parentheses and significance as; * when p < 0.1, ** when p < 0.05, and *** when p < 0.01. We correct standard errors for autocorrelation using Newey and West (1987). Panel A: Variance and the higher-moment risk index Horizon HRI Month −0.53 95% CI [−0.60,−0.48] Quarter −0.64 95% CI [−0.69,−0.59] Panel B: Past return and higher-moment risks Horizon HRI Var (%) Skew Kurt Hskew Hkurt Month 2.26∗∗∗ −0.30∗−0.61∗∗∗ 2.75∗∗∗ −24.57∗∗∗ 201.65∗∗ (sd) (0.79) (0.16) (0.14) (1.01) (9.41) (98.06) Quarter 2.13∗∗ −0.68∗−0.45∗∗∗ 1.02∗−9.62∗∗ 50.63∗∗ (sd) (0.88) (0.39) (0.12) (0.56) (3.83) (25.23) Panel C: Past return and higher-moment risks — controlling for lagged risk Horizon HRI Var (%) Skew Kurt Hskew Hkurt Month 0.94∗∗∗ −0.07∗∗ −0.23∗∗∗ 1.16∗∗∗ −11.11∗∗ 109.07∗ (sd) (0.34) (0.03) (0.06) (0.39) (4.53) (57.91) Quarter 0.47∗∗ −0.14∗∗ −0.15∗∗∗ 0.25∗−2.04∗∗ 11.98∗∗ (sd) (0.19) (0.06) (0.04) (0.14) (0.86) (5.33) 101
B Appendix Tables Table AI: Ex Ante Conditional Moments Predict Ex Post Realized Moments (Test statistics). This Table reports test statistics for the results reported in Table 2.2. Panel A reports p-values from the bootstrapped distribution. Panel B reports bootstrapped standard errors for the correlation coefficient between ex ante moments and ex post realized moments. Panel C reports bootstrapped standard errors for the correlation coefficient between ex ante kurtosis and ex post realized left kurtosis. Panel C reports also reports correlations for hyperkurtosis. Panel A: Sorting on ex ante monthly moments Our moments Historical moments Low ex ante High ex ante Low ex ante High ex ante Variance 0.00 0.00 0.00 0.01 Skewness 0.19 0.46 0.54 0.74 Kurtosis 0.21 0.03 0.46 0.20 Hyperskewness 0.02 0.45 0.47 0.75 Hyperkurtosis 0.39 0.01 0.38 0.20 Panel B: Correlation between ex ante moments and ex post realized moments Our moments Historical moments Month Quarter Month Quarter Variance 0.07 0.09 0.09 0.10 Skewness 0.06 0.06 0.07 0.06 Kurtosis 0.06 0.07 0.05 0.04 Hyperskewness 0.05 0.06 0.06 0.04 Hyperkurtosis 0.05 0.06 0.04 0.03 Panel C: Left kurtosis and left hyperkurtosis Our moments Historical moments Month Quarter Month Quarter Left kurtosis 0.06 0.07 0.07 0.06 Left hyperkurtosis 0.06 0.06 0.07 0.06 108
Table AII: Ex Ante Conditional Moments Predict Ex Post Realized Moments — Robustness (Test statistics). Panel A reports bootstrapped standard errors for the correlations between our ex ante moments and ex post realized moments when we remove observations that overlap with the period from August 1, 2008 to July 31, 2009. Panel B reports bootstrapped standard errors for the correlations between our ex ante moments (estimated with different levels of relative risk aversion) and ex post realized moments. Panel C reports bootstrapped standard errors for the correlations when controlling for risk-neutral moments or historical moments. We report bootstrapped standard errors in the appendix and significance as; * when p < 0.1, ** when p < 0.05, and *** when p < 0.01.. Panel A: Excluding August 1, 2008 to July 31, 2009 Our moments Historical moments Month Quarter Month Quarter Variance 0.06 0.05 0.07 0.08 Skewness 0.06 0.05 0.07 0.06 Kurtosis 0.06 0.08 0.06 0.07 Hyperskewness 0.05 0.06 0.07 0.07 Hyperkurtosis 0.05 0.07 0.06 0.05 Panel B: Other levels of risk-aversion γ= 1 γ= 5 Month Quarter Month Quarter Variance 0.07 0.10 0.06 0.09 Skewness 0.06 0.06 0.06 0.06 Kurtosis 0.06 0.06 0.05 0.06 Hyperskewness 0.05 0.05 0.05 0.07 Hyperkurtosis 0.05 0.05 0.05 0.06 Panel C: Marginal correlations Controlling for Controlling for risk-neutral moments historical moments Month Quarter Month Quarter Variance 0.17 0.10 0.12 0.06 Skewness 0.06 0.06 0.06 0.06 Kurtosis 0.05 0.06 0.06 0.07 Hyperskewness 0.05 0.06 0.05 0.06 Hyperkurtosis 0.04 0.06 0.05 0.06 109
110
Chapter 3 The Market’s Time-Varying Risk Aversion Abstract: I present a new method for estimating the market’s time-varying risk aversion using historical market returns and option prices written on the S&P 500 index. Market risk aversion varies in a systematic way; it tends to be low during times of financial market distress, e.g., during the 2008-2009 financial crisis, and it tends to be high at times when the market is considered to be calm as measured by variance. This systematic variation in market risk aversion is difficult to reconcile with the leading asset pricing models. I discuss several possible explanations for these time-varying patterns in risk aversion including investor salience and time-varying correlations between aggregate consumption and the market. I am grateful for helpful comments from Niels Joachim Gormsen, David Lando, and Lasse Heje Pedersen. I gratefully acknowledge support from the FRIC Center for Financial Frictions (grant no. DNRF102) and from the European Research Council (ERC grant no. 312417). 111
1 Introduction Risk averse market participants require compensation for taking on market risk. The amount of compensation required depends on the level of investor risk aversion and the expected distribution of future market returns. Understanding how risk aversion and the expected distribution of future market returns co-vary over time is therefore important in understanding market prices. In this paper, I provide a new method for estimating the market’s time-varying risk aversion. My methodology allows me to investigate the co-movements between market risk aversion and the physical distribution of market returns. I arrive at the following two main results: (i) Market risk aversion varies over time and tends to be low when volatility is high, e.g., during the recent 2008-2009 financial crisis. (ii) Market risk aversion tends to be high at times when market tail risk is high. This systematic variation in market risk aversion is difficult to reconcile with the leading asset pricing models. In the last part of the paper, I discuss possible explanations for the systematic variation in risk aversion. Specifically, I consider salience theory as a behavioral explanation and time-varying correlation between the market and aggregate consumption as a rational explanation. Before I go into the details on my results, it is instructive to understand how I estimate the market’s time-varying risk aversion. As in Martin (2017) and Gormsen and Jensen (2017b), I consider the preferences of a power utility investor who chooses to be fully invested in the market. I estimate option implied risk neutral distributions for multiple horizons corresponding to the last trading day of the month from January 1996 until December 2015. Using these risk neutral distributions, I estimate the time series of risk aversion coefficients which best match the historical realized monthly returns on the market while still being consistent with observable option prices. To do so, I rely on the generalized recovery methodology developed in Jensen, Lando, and Pedersen (2017) combined with the Berkowitz test as used in Bliss and Panigirtzoglou (2004). Specifically, I estimate the power utility investor’s time preference parameter by minimizing the Berkowitz test statistic under the constraint that, for a given value of the time preference parameter, the time-varying risk aversion coefficients solves the generalized recovery equation set. 112
As my first main result, I show that the market’s risk aversion varies systematically with variance. Specifically, risk aversion is negatively correlated with market variance with a correlation coefficient of −0.65 and tight bootstrapped 95% confidence bounds of [−0.70,−0.62].1The negative correlation suggests that market participants are more risk tolerant at times when the market is generally considered highly risky, i.e., during times of high market volatility. This finding is hard to reconcile with the leading asset pricing models. For example, Campbell and Cochrane (1993) explain asset prices with a risk aversion that is countercyclical. To understand why it is reasonable that market risk aversion is low at times when variance is high, consider the following heuristic example. Under power utility and log-normality of the market return distribution then the expected excess return on the market is equal to γσ2, where γis market risk aversion and σ2is market variance. Suppose risk aversion is equal to 3 and constant, which is a commonly chosen value in the financial economics literature, see e.g. Bliss and Panigirtzoglou (2004) and Gormsen and Jensen (2017b). Then, during the peak of the financial crisis in 2008-2009 where annualized realized market variance reached almost 60%, the implied annual expected return on the market was about 180%. It is hard to believe that the market participants believed that, in expectation, the market was going to bounce back with an increase of 180% from the peak of the crisis. Therefore, if annualized expected returns were truly lower than 180% during the financial crisis, then either risk aversion is lower than 3 and constant or it fluctuates and was low during the crisis. I find that risk aversion on average is 2.77 and that it falls to 1.49 during the peak of the financial crisis. Using the heuristic argument above, this level of risk aversion implies an expected excess return on the market during the peak of the crisis of about 90% annually and 7.5% on a monthly horizon. As my second main result, I show that the market’s risk aversion varies systematically with market tail risk. Specifically, I find that risk aversion is: (i) negatively correlated with market skewness with a correlation of −0.42 with bootstrapped 95% confidence bounds of [−0.51,−0.31] and (ii) positively correlated with market kurtosis with a correlation of 0.62 and bootstrapped confidence bounds of [0.54,0.70]. This 1Bliss and Panigirtzoglou (2004) also find evidence of a risk aversion which is negatively related to variance. 113
finding suggests that the investor becomes more risk averse as the market’s higher order moments become more risky, i.e., there is more probability mass in the tail of the market return distribution. Next, I discuss evidence of this systematic variation in risk aversion that is already in the finance literature. For example, in a recent paper by Moreira and Muir (2017b), the authors show that investors can earn large alphas by timing market volatility. They find that investors should exit the market when volatility increases and enter the market when it drops. Their result arises because expected excess returns are high relative to variance during low variance periods, i.e., the ratio ERt σ2 tvaries over time and becomes high when variance is low. Their findings imply a time variation in market risk aversion which is consistent with the results I present. To see why, recall that we can express the expected excess return on the market as ERt=γtσ2 twhich implies a market risk aversion of γt=ERt σ2 t. Their results are therefore reminiscent of a risk aversion that is low at times when volatility is high. Similarly, Gormsen and Jensen (2017a) also find evidence that the market price of risk is high during low volatility periods. In the final part of the paper, I discuss two possible explanations for the systematic variation in market risk aversion. First, I show that my results are consistent with aspects of salience theory by Bordalo, Gennaioli, and Shleifer (2012). Specifically, I follow Lian, Ma, and Wang (2018) who argue that, at times of low interest rates the relatively high expected returns on risky assets are salient, and this salience on the upside of a higher return on the risky asset induces heightened risk tolerance and “reaching for yield” tendencies among investors. I therefore regress the ratio of expected gross returns on the market to gross risk-free returns, Et(Rt,T )/Rf t,T , onto risk aversion. Consistently with the findings of Lian, Ma, and Wang (2018), I find that investors become more risk tolerant as the ratio of expected returns to risk-free returns increases. Lastly, I discuss how the systematic variation in risk aversion can arise if the stock market is not a perfect proxy for aggregate consumption. Specifically, for an investor with power utility who cares about aggregate consumption, then the stock market implied risk aversion can be expressed as αt=ρtγc tσc t/σtwhere ρtis the correlation between the stock market and aggregate consumption, σc tis consumption volatility, 114
γc tis the investor’s risk aversion, and σtis the volatility of the stock market. As an illustrative example, I consider the case where risk aversion and consumption volatility are constant, i.e., γc t=γcand σc t=σc. In this case, the stock market implied risk aversion is proportional to the ratio of the correlation between the stock market and aggregate consumption and the volatility of the stock market. During times of financial distress, stock market volatility increases, if the correlation between the stock market and aggregate consumption does not increase enough to offset the increase in stock market volatility then αtwill decrease. My paper relates to and extends the existing literature on estimating market risk aversion. I present a new method for estimating the market’s risk aversion by integrating the Berkowitz test with the generalized recovery method of Jensen, Lando, and Pedersen (2017). Previous research which use option prises to estimate time-varying risk aversion is based either (1) a full identification (e.g. a parameterization through a GARCH model) of the market’s physical return distribution from historical returns or (2) parameterization of investor preferences, e.g. Ross (2015) and Jensen, Lando, and Pedersen (2017). My method allows for the identification of time-varying risk aversion while leaving the physical probability distribution fully free, which compliments the existing literature that imposes structure on the physical return distribution, e.g. Barone-Adesi, Engle, and Mancini (2008), Jackwerth (2004), Jackwerth (2000), Ait- Sahalia and Lo (2000). My paper also relates to Bliss and Panigirtzoglou (2004) who use the Berkowitz test, but they use it to estimate a constant risk aversion. The paper proceeds as follows: Section 2describes how I estimate the market’s time-varying risk aversion. Section 3covers data and details on the empirical implementation. Section 4studies the empirical results of the time-varying market risk aversion. Section 5studies the relation between market risk aversion and the market’s physical moments. Section 6investigates what might explain the systematic variation in the market’s time-varying risk aversion. Section 7concludes the paper. 2 Inferring Financial Market Risk Aversion I consider the preferences and beliefs of a power utility investor who optimally chooses to be fully invested in the market. The economy consists of a risk-free asset with gross 115
risk-free return of Rf t,T and a risky asset, the stock market, with price S, dividends Dt,T , and gross returns Rt,T =ST+Dt,T St (3.1) The power utility investor’s utility function at time tis Ut(x) = x1−γ(t)/(1−γ(t)) where γ(t) is the investor’s (possibly time-varying) risk aversion coefficient. The investor has initial wealth W0and terminal wealth W0Rt,T . Given that the power utility investor chooses to be fully invested in the stock market, I can express the investor’s stochastic discount factor in the following way: mt,T (Rt,T ) = δ(t, T)R−γ(t) t,T (3.2) where δ(t, T) is a time-preference parameter and γ(t) is the market’s risk aversion2. Table 3.1 reports the five different specifications of the stochastic discount factor in (3.2) which I investigate in this paper. Next, I show how I infer the market’s time-varying risk aversion. I combine the results of two related papers, the generalized recovery methodology of Jensen, Lando, and Pedersen (2017) and the Berkowitz test as used in Bliss and Panigirtzoglou (2004). The main methodological objective of both these papers is to backward engineer preference parameters from observable asset prices,3in this paper I exploit these methods in a joint setting. I start by fixing some notation. I write the standard asset pricing formula in the following common way: πt,T (r) = pt,T (r)mt,T (r) (3.3) where πt,T is the time tand T−thorizon known state price density, pt,T is the corresponding unknown physical probability density, and mt,T is the unknown stochastic discount factor. In Section 3, I discuss how I estimate state price densities using option prices written on the market. Given the standard asset pricing formula, I can write 2I assume that the market’s risk aversion is independent of the horizon. There are several papers that address horizon dependent risk aversion, see e.g., Bliss and Panigirtzoglou (2004) and Lazarus (2018). The methodology I use in this paper can accommodate a horizon specific risk aversion, however, I choose to focus on the variation in the time series of risk aversion rather than the term structure of risk aversion. 3 116
the market’s probability distribution function, say F, as Ft,T (r) = Zr −∞ pt,T (x)dx =Zr −∞ πt,T (x) mt,T (x)dx (3.4) If I insert the power utility investor’s stochastic discount factor from (3.2) into (3.4), then I can rewrite the market’s distribution function as follows Ft,T (r) = Zr −∞ πt,T (x)xγ(t) δ(t, T)dx (3.5) Given values of the parameters δ(t, T) and γ(t), I can estimate the market’s probability distribution as perceived by the power utility investor who chooses to invest everything in the market. The objective now is figuring out what the true values of δ(t, T ) and γ(t) are. To estimate these true values of the preference parameters, I follow Bliss and Panigirtzoglou (2004) and use the so-called Berkowitz test, cf. Berkowitz (2001). The idea behind the Berkowitz test is that, for the true values ˆ δ(t, T) and ˆγ(t), the distribution of ut,T =ˆ Ft,T (Rt,T ) is uniform and the distribution yt,T = Φ−1(ut,T ) is standard normal. Here ˆ Ft,T (Rt,T ) denotes the distribution function in (3.5) with the true values, ˆ δ(t, T) and ˆγ(t), inserted. In the Berkowitz test, I estimate the coefficients in the regression model: yt,T = ˆa+ˆ βyt−1,T−1+t,T , t,T ∼N(0,ˆσ) (3.6) and perform a likelihood ratio test of the joint hypothesis that a=β= 0 and Var(t,T ) = 1. The hypothesis that b= 0 is natural when considering non-overlapping returns, for overlapping returns see e.g. Bliss and Panigirtzoglou (2004) for a thorough discussion of the test. It is also worth noticing that, even though there might be momentum effects in returns, then we will still want b= 0 because the true distribution should take these momentum effects into account. The Berkowitz likelihood ratio test for non-overlapping returns is then: LR =−2(LL(0,0,1) −LL(ˆa, ˆ β, ˆσ)) ∼χ2 3(3.7) 117
Therefore, in specification (2) from Table 3.1, I fix δ(t, T)=0.98T−t.8The second column of Table 3.2 shows the Berkowitz test for specification (2). The constant time preference parameter with an annualized value of 0.98 is strongly rejected by the data. This result means that, if δ(t, T) = 0.98T−t, there is no time series of γ(t) which can simultaneously match the historical returns on the market and the observable option prices. From the point of view of a power utility investor who wants to hold the market, if we want to match both option prices and returns, then we need to accept that time preferences are on average above one. It is worth noting that, even though δ(t, T) is above one, the setting I adopt will always (and by constriction) match interest rates. To understand how, recall that the law of one price gives us the following relation: Et[mt,T (Rt,T )] = 1 Rf t,T (3.16) which for the power utility investor is equivalent to: Ethδ(t, T)R−γ(t) t,T i=Z∞ −∞ pt,T (x)δ(t, T)x−γ(t)dx (3.17) From the standard asset pricing formula in (3.3), I am given the relation pt,T (r) = πt,T (r)/mt,T (r) which is equal to πt,T (r)rγ(t)/δ(t, T ) for the power utility investor. Inserting this expression of physical probabilities into (3.17), I get: Z∞ −∞ πt,T (x)xγ(t) δ(t, T)δ(t, T)x−γ(t)dx =Z∞ −∞ πt,T (x)dx =1 Rf t,T Z∞ −∞ qt,T (x)dx (3.18) where qt,T (r) is the risk-neutral distribution. The integral on the rhs. of (3.18) is equal to one since qt,T (r) is a probability distribution. Therefore, for any level of δ(t, T), my setting will by construction match interest rates. Next, I want to better understand the scenarios in which δ(t, T )>1. Therefore, I consider the following equation which again arises from the law of one price and a 8E.g. Campbell and Cochrane (1993) derive an annualized value of the time preference parameter close to 0.98. 124
power utility investor who choses to hold the market δ(t, T)Z∞ −∞ pt,T (x)x−γ(t)dx =1 Rf t,T (3.19) Now, if R∞ −∞ pt,T (x)x−γ(t)dx < 1 Rf t,T then δ(t, T) must, by the law of one price, be above one. This scenario can for example arise if the physical return distribution has most of its probability mass in the states where the market increases in value, that is, when R∞ 1pt,T (x)dx is large (close to one). In this case, the integral R∞ −∞ pt,T (x)x−γ(t)dx could potentially be lower than 1/Rf t,T because x−γ(t)<1 for values of x > 1 and γ(t)>0. Therefore, a value of δ(t, T ) which is larger than one can arise in a scenario where the probability of a positive return on the market is large. That is, the investor is more willing to postpone consumption into the future when the probability of a future good state (a state with high consumption) is high.9This result is highly unintuitive. I expected the exact opposite result, that the investor is willing to postpone (smooth) consumption if the probability of a bad state in the future is high. The result may arise for several reasons. For example, it might be a consequence of a misspecified functional form of the stochastic discount factor. Also, it can be because the state variable is wrong, that is, market returns might not be a perfect proxy for aggregate consumption. There are of course several other possible explanations. 5 Systematic Variation in Market Risk Aversion Figure 3.1 shows the market’s time-varying risk aversion. As seen from the figure, risk aversion was low during the 2008-2009 financial crisis and high both in the years leading up to the crisis and in the post crisis period of 2010-2015. This variation in risk aversion suggests that the market became less risk averse during the peak of the financial crisis and more risk averse after the crisis. In this section, I investigate these time-varying patterns of market risk aversion. First, using methods developed in Martin (2017) and Gormsen and Jensen (2017b), I estimate the market’s physical return distribution. Table 3.4 shows the summary 9Similarly, a value of δ(t, T ) which is lower than one can arise in a scenario where the probability of a negative return on the market is large. That is, the investor is more willing to consume today if the probability of a future bad state (a state with low consumption) is high. 125
statistics for the first four moments of the physical return distribution: expected excess return, variance, skewness, and kurtosis. Looking at the last row of Table 3.4 we see the market’s time-varying moments implied by specification (5) of Table 3.1. The physical return distribution implied by this specification of the stochastic discount factor is on average negatively skewed and exhibit excess kurtosis. Results are similar for specification (3) and (4). Figure 3.2 shows a time-series plot of the market’s time-varying risk aversion and its physical conditional variance. Clearly, risk aversion tend to be low at turbulent times when variance is high. Panel A of Table 3.5 shows the correlations between the market’s time-varying risk aversion and the market’s physical moments. The correlation between market risk aversion and variance is −0.65 with bootstrapped 95% confidence bounds of [−0.70,−0.62]. Figure 3.3 shows the market’s risk aversion and its physical conditional skewness. Risk aversion and skewness are negatively correlated with a correlation of −0.42 and 95% bootstrapped confidence bounds of [−0.51,−0.31] as shown in Table 3.5. The negative correlation between risk aversion and skewness suggests that the market is more risk averse at times when market tail risk is high. The expected return on the market is negatively related to risk aversion with a correlation coefficient of −0.56 and bootstrapped 95% confidence bounds of [−0.61,−0.52]. This negative relation suggests that the expected return on the market is high during times of financial market distress when risk aversion is low. Similarly, the market’s Sharpe ratio tends to be high during periods of high volatility. To better understand why the market’s Sharpe ratio is high at times when risk aversion is low, notice that, under power utility and log-normal market returns, the expected excess return on the market is ERt=γtσ2 t. Rewriting this equation leads to the following expression of risk aversion: γt=SRt/σt(3.20) where SRt=ERt/σtis the market’s Sharpe ratio. Even though the market’s Sharpe ratio might increase at times of financial market distress when volatility also increases, then the market’s risk aversion can decrease if the increase in the Sharpe ratio is too low to offset the higher market volatility. 126
6 Understanding the Market’s Time-Varying Risk Aversion As shown in the previous section, market risk aversion varies in a systematic way, that is, it tends to be high at times when volatility is low and vice versa. A natural question to ask is; does it make sense at all that market risk aversion is low during times of financial distress when volatility is high? It is definitely difficult to reconcile this systematic variation in market risk aversion with the leading asset pricing models. For example, Campbell and Cochrane (1993) use habit formation as a mechanism for time-varying risk aversion and find that a countercyclical risk aversion can help explain asset prices and the behavior of the stock market. As a start to understanding this systematic variation in market risk aversion, it is instructive to think about the following heuristic argument: standard textbook asset pricing tells us that, under power utility and a log-normal assumption on market returns, the expected excess return on the market is: ERt=γtσ2 t(3.21) During the financial crisis of 2008-2009, annualized monthly variance reached 60%. If I assume that risk aversion is constant and at a level of 3, which is a common choice in the financial literature10, then (3.21) implies an annualized expected excess return of 180% and a monthly expected excess return of 15%. It is hard to believe that the market expected the excess returns on the monthly horizon was as high as 15% during the peak of financial crisis, not to mention the extremely high annualized expected excess returns. Therefore, since we can estimate market variance rather accurately, then if we want expected excess returns to be at a realistic level during times of high volatility, then market risk aversion should go down when volatility spikes. As noted above, this systematic variation in market risk aversion is difficult to reconcile with the leading asset pricing models. However, when looking at the intersection between option prices and market returns, this variation seems to be a persistent pattern and it is therefore interesting to understand how it can arise in our data. 10See e.g. Bliss and Panigirtzoglou (2004) and Gormsen and Jensen (2017b) 127
6.1 Evidence of a Procyclical Market Risk Aversion In a recent paper, Moreira and Muir (2017b) show that investors can earn large alphas by timing market volatility. They find that investors should exit the market when volatility increases and enter the market when it drops. Their result arises because expected excess returns are high relative to variance during low variance periods, i.e., the ratio ERt σ2 tvaries over time and becomes high when variance is low. Similarly, Gormsen and Jensen (2017a) also find evidence that the market price of risk is high during low volatility periods. These findings have implications for how we should think about the time variation in market risk aversion. To see why, consider again equation (3.21) which tells us that the expected excess return to variance ratio is equivalent to market risk aversion and their results are therefore reminiscent of a risk aversion that is high at times when volatility is low. The results of Moreira and Muir (2017b) can also in part be explained by the fact that market tail risk is higher at times of low volatility, see e.g. Gormsen and Jensen (2017b). The higher expected excess return to variance during low volatility periods could simply be compensation for higher tail risk (or a combination of higher tail risk and higher risk aversion). Nevertheless, it is an artifact of the data that risk aversion tends to be high at times when the market is generally considered to be calm. In the following two subsections, I discuss two possible explanations for these time-varying patterns. First, I investigate a behavioral explanation through investor salience. Second, I consider a rational explanation through time-varying correlations between the market and aggregate consumption. 6.2 Salience Theory and Market Risk Aversion In this subsection, I discuss investor salience as a possible explanation for the systematic variation in market risk aversion which I document in this paper. Bordalo, Gennaioli, and Shleifer (2012) develop a model in which the focus of investors are drawn to unusual, different or odd events. Specifically, an investor is risk-seeking if his attention is drawn to the “upside” of a lottery and risk-averse if his attention is drawn to the “downside” of a lottery. They refer to the state which has drawn the investor’s 128
attention as the salient state or salient outcome. A practical example of salience theory can be found in Lian, Ma, and Wang (2018). The authors show that, investors have greater appetite for risk taking when interest rates are low. They argue that, at times of low interest rates the relatively high expected returns on risky assets become salient, and this salience on the upside of a higher return on the risky asset induces heightened risk tolerance and “reaching for yield” tendencies. Also, they refer to Weber’s law and argue that investors tend to evaluate assets in terms of proportional returns rather than by differences. In my setting both the risk-free rate and the expected return on the market fluctuate. Therefore, to test how their hypothesis fits into my setting, I regress my measure of risk aversion onto the ratio of gross expected returns to gross risk-free returns:11 γ(t) = β0+β1 Et(Rt,T ) Rf t,T +t,T (3.22) The idea is that, as expected returns on the risky asset increase relative to the safe return on the risk-free asset, investors will find the higher return on the risky asset salient and therefore increase their risk tolerance which results in a lower γ(t). Table 3.6 presents the results of regression (3.22). The coefficient β1is −6.97 with 95% bootstrapped confidence bounds of [−8.63,−5.92]. This results suggests that an increase in the ratio of gross expected return on the market to risk-free gross return of, say 0.1, results in a decrease in risk aversion of almost 0.7, that is, as expected returns on the market increase relative to the return on the risk-free asset, the investor becomes more risk tolerant. 6.3 Consumption and Stock Market Correlation A natural critique of the method I apply in this paper is that the power utility investor only cares about how the stock market develops and gains all utility from market movements. In this subsection, I discuss what can drive my results if I in fact have 11Lian, Ma, and Wang (2018) argue that, when looking at gross returns the effect of salience is lower than when considering ’normal’ returns because the ratio of, say 1.1/1.05, is much smaller than the ratio 0.1/0.05. Nevertheless, they find evidence that investors still exhibit salient preferences when considering gross returns. I do not consider ’normal’ returns because they, at the end of my sample, are very close to zero. 129
postulated the wrong state variable. Suppose for now that the true state variable is aggregate consumption and that the stock market is not a prefect proxy for aggregate consumption. Also, suppose that risk aversion is actually constant, say γc, the power utility investor’s ’true’ stochastic discount factor is then: mt,T =δT−tCT Ctγc (3.23) where Ctis aggregate consumption at time t. If I assume that consumption is lognormally distributed with constant variance, σ2 c, then in the standard Merton (1973) model, I can express expected return on the stock market in the following way: ERt=γcσtσcρt(3.24) where ρtis the time-varying correlation between the stock market and aggregate consumption. Keeping the stock market volatility constant, then as the correlation between the stock market and aggregate consumption increases, so does the required expected return on the stock market. Now, if I define the time-varying parameter αt=ρtγcσc σt 12 then I can rewrite (3.24) as ERt=αtσ2 t(3.25) which is equivalent to imposing that the stock market is the state variable, log-normally distributed, and that the power utility investor has time-varying risk aversion αt. Therefore, even though risk aversion is constant at γc, the stock market implied risk aversion can be time-varying because the correlation between consumption and the stock market is time-varying. As a consequence of the assumptions made in this subsection, the time-varying correlation between consumption and the stock market is proportional to the market’s Sharpe ratio, ρt∝ERt σt, and the stock market implied time-varying risk aversion is proportional to the consumption to stock market correlation over stock market volatility, 12In a (more realistic) setting where risk aversion (γc t), consumption volatility, and the correlation between the market and aggregate consumption are time-varying, then αt=ρtγc tσc t σtis the market’s implied risk aversion as seen from the point of view of a power utility investor who holds the market. Clearly, in this setting, I cannot distinguish between the contributions in the time variation of αtwhich comes from either of the time-varying parameters. 130
αt∝ρt σt. The correlation between consumption and the stock market is proportional to the Sharpe ratio of the market and is therefore high during high volatility periods. For instance, during the recent financial crisis when stock market volatility spiked, then so did the correlation between consumption and the stock market. However, the increase in correlation between consumption and the stock market was not enough to offset the increase in stock market volatility and as a consequence stock market risk aversion decreased. Figure 3.4 shows the time series of implied correlations between consumption and the stock market when consumption risk aversion is constant at 38 as implied by the equity premium puzzle, cf. Mehra and Prescott (1985b), and monthly consumption volatility is constant at 0.0075. Clearly, a setting in which both risk aversion and consumption volatility are constant is too simplified to fully capture the real world. Many of the conclusions which I draw in this subsection can be altered simply by assuming that, for example, risk aversion (γc) is time-varying or consumption volatility is time-varying. Nevertheless, the discussion in this subsection highlights that the time variation in market risk aversion, which I document in this paper, can be a consequence of a misinterpretation of the implied risk aversion. That is, the risk aversion implied by the market is time-varying and tends to be low at times when volatility is high because the mapping from the true state variable, aggregate consumption, onto the market is time-varying, e.g., through time-varying correlations. 7 Conclusion I present a new method for estimating the market’s time-varying risk aversion directly from option written on the market and historical market returns. My method combines the generalized recovery method of Jensen, Lando, and Pedersen (2017) and the Berkowitz test as used in Bliss and Panigirtzoglou (2004). The estimated market risk aversion coefficients are on average 2.77 and vary from 1.49 to 5.03 over the period from 1996 to 2016. Interestingly, market risk aversion is negatively related to variance. During the peak of the financial crisis in 2008-2009 when market volatility was high, risk aversion was at its all-time low. These co-movements 131
between risk aversion and variance are difficult to reconcile with the leading asset pricing models. I discuss two possible explanations for the negative relation between risk aversion and variance. First, I show that my results are consistent with salience theory. Specifically, during periods of high volatility when expected returns on the market are high relative to the risk-free returns, investors find the relatively high expected return on the market salient and therefore become more risk tolerant. Secondly, I show that my results can arise if the stock market is a poor proxy for aggregate consumption. 132
Table 3.1: Specifications of the stochastic discount factor. This table reports five different specifications of the CRRA stochastic discount factor that I use in this paper. The stochastic discount factors take the form: mt,T (Rt,T ) = δ(t, T)R−γ(t) t,T (3.26) Here δ(t, T) is a time preference parameter and γ(t) is the coefficient of relative risk aversion. Parameterization # Comments δ(t, T)γ(t) (1) Risk neutral distribution 1/Rf t,T 0 (2) Fixed time preferences 0.98T−tγt Time-varying risk aversion (3) Time-varying time preferences (Bliss and Panigirtzoglou (2004)) δT−t tγ Constant risk aversion (4) Constant time preferences δT−tγt Time-varying risk aversion (5) Time preferences affine in the inverse of the gross risk-free return 1/Rf t,T +δT−t−1γt Time-varying risk aversion 133