scieee AI-readable full text Open interactive document viewer

Principal and Independent Component Analysis in Financial Time Series

José Miguel Rodrigues Teixeira Salgado

Full text

Family comes in all shapes and sizes. — The Family Book, Todd Parr To Ana (and Matilde and João) with endless Love... ABSTRACT In this work we consider the application of a plethora of Econophysics techniques to multivariate financial time series, particularly the Correlation matrix, the Forecastable Component Analysis, the Mutual Information, the Kullback-Leibler Divergence, the Approximate Entropy, the Distance Correlation and the Hurst exponent. The key idea was not to compare their differences but more to find their “joint strength” by combining their different views of time series. We applied these techniques to two different scenarios: one, more local, to 12 stocks quoted in the Portuguese Stock Market (PSI-20); the other one, more global, to 23 world stock markets. Also, we have studied and used “sliding windows” of different sizes. The motivation and importance of this kind of analysis relies on the well known multi-fractal behaviour that financial data exhibits. We started by confirming some results found in literature, namely the ones from random matrix theory and the ones for the Hurst exponent. In this case, and based in previous results, we propose that the PSI-20 is becoming more mature. Distance correlation have shown to be a good complement to entropy measures like Mutual Information or Kullback-Leibler divergence. Approximate entropy, as a stand alone method, have shown potential complementarity with Distance correlation in the case of the stocks from PSI-20 index. To our knowledge, it is the first time that energy statistics is applied to the PSI-20 data. Is is interesting to note that this measure, and this is corroborated by Approximate entropy results, proposes two well defined behaviour for the PSI-20 stocks. One period, from 2000 to 2007, relatively calm, with low variation of Distance Correlation between stocks, and another period, from 2007 till now, much more agitated in what concerns this measure. Unfortunately, we cannot say the same for the Distance Correlation results applied to the World Markets set. Nevertheless, we can find strong regional correlation for most of the markets. Some, but only a few, can be considered more global markets, with influence in all the others. There is, in that sense, a strong connection between the NorthAmerican markets and most of the European ones. That correlation has become higher since 2007, complementing the idea that the markets are more connected. For Mutual Information or Kullback-Leibler Divergence the results are very sharp and we can clearly match high entropy values with real events. Some of them are only important for specific stocks or markets, but some others, more related to recession periods, are independent of a specific stock or market. In general, a trend common to most markets is the progressive growing correlation over time. One possible reason to this is the progressive globalisation of markets, where the arbitrage opportunities are reduced due to more efficient markets. Also, the information we got from Hurst exponent was vital to confirm that stocks and markets are getting more and more mature, that is, less autocorrelated. iii RESUMO Neste trabalho consideramos a aplicação de algumas técnicas da Econofísica às séries financeiras temporais multivariadas, nomeadamente consideramos as técnicas das matrizes aleatórias como a matriz de correlação, as técnicas da análise de componentes, da informação mútua, da divergência de Kullback-Leibler, da entropia aproximada, da distância de correlação e do expoente de Hurst. A ideia fundamental não foi comparar as suas diferenças mas sim encontrar as suas “forças conjuntas” ao combinar a forma como cada técnica “vê” as séries temporais. Estas técnicas foram aplicadas em dois cenários distintos: um, mais local, a 12 ações cotadas no PSI-20, o índice da Bolsa portuguesa; o outro, mais global, foi aplicado a 23 mercados de diferentes países. Ainda, usou-se aqui uma técnica de cálculo por “janelas” temporais dado o conhecido comportamento multifractal dos dados financeiros. Começamos por confirmar os resultados conhecidos da literatura para as matrizes aleatórias e para o expoente de Hurst. Neste último caso, e baseados nos resultados anteriores, propomos que o PSI-20 está a tornar-se um mercado mais maduro. A Distância de Correlação provou ser uma medida com boa complementaridade com medidas de entropia como a Informação Mútua ou a divergência de Kullback-Leibler. A Entropia Aproximada, por si só, mostrou uma boa complementaridade com a Distância de Correlação na aplicação às ações do PSI-20. Que tenhamos conhecimento, é a primeira vez que a Distância de Correlação é aplicada ao PSI-20. É interessante notar que esta medida, e isto é corroborado pelos resultados da Entropia Aproximada, propõe dois períodos comportamentais bem definidos: um, de 2000 a2007, com pequenas variações e valores também pequenos e outro, com grandes variações e com valores muito elevados de correlação entre as ações do PSI-20. Contudo, esta observação não permanece quando aplicamos a mesma medida aos mercados mundiais. Todavia, encontramos correlações regionais fortes para a maior parte dos mercados. Alguns mercados, embora poucos, podem ser vistos como globais já que influenciam todos os outros. Neste sentido, é de referir a forte ligação dos mercados norte-americanos com os mercados europeus. Esta correlação continua a crescer desde 2007, ajudando a complementar a ideia de que os mercados estão mais ligados. Para a Informação Mútua ou para a divergência de Kullback-Leibler os resultados são muito claros. Conseguimos ligar os valores mais elevados da entropia a acontecimentos reais. Uns, mais restritos, e portanto, influenciando apenas ações ou mercados pontuais; outros, mais globais, deixando a sua marca em todas as ações/mercados. Em geral, uma tendência comum a todos os mercados é o aumento gradual temporal da correlação. Uma possível razão pode ter a ver com a progressiva globalização dos mercados, onde as oportunidades de arbitragem estão reduzidas devido ao facto dos mercados serem cada vez mais eficientes. A informação que obtivemos a partir do expoente de Hurst foi vital para confirmar a informação de que os mercados estão cada vez mais maduros, isto é, menos autocorrelacionados. iv ACKNOWLEDGEMENTS I owe, firstly, many thanks to my advisor, José Abílio Oliveira Matos, for being so helpful, patience, dedicated and committed to this project. Most of the time that I was lost, he was there to keep us up, was not his motto “Be Prepared”! In second place I wish to thank my family, my teachers and some friends, not necessarily by this order of importance:: • To the scouts from my Group in Guimarães (an endless list started by Alexandre, Ernesto, Manel, Miguel and Samuel) for, most of the times without knowing, keeping me up; • To Ricardo Gama for his friendship, even at distance, from the times since the Master degree; • To some of my teachers, particularly Prof. Eduardo Laje and my master thesis advisor, Prof. Silvio Gama, from whom, without no pain, I got some of the most important lessons in my life; • To my colleagues from IPG, particularly A. Martins, C. Rosa, J.C. Miranda, P. Costa and P. Vieira, for helping me to keep up my scientific motivation, for, at some times, their hospitality or for, at other times, just sharing meals and/or coffees; • To my nephew and nieces, particularly my godsons Francisca and Dinis, but also Beatriz and Carolina, for their joy and life; • To my grandfather, António Augusto Cordeiro Rodrigues, for reminding me all the time to accomplish this purpose; • To my parents, Sr. Salgado and D. Conceição, and my mother-in-law, D. Isabel, for their continuous love, concern, support and understanding; • To my beloved Ana, Matilde and João, for being unique and precious, for their love, joy, patience and... for everything!, and without whom all this effort would seem totally senseless. v CONTENTS 1 introduction 1 1.1Motivation..................................... 1 1.2Econophysics ................................... 1 1.2.1Briefhistory................................ 2 1.2.2WhyEconophysics? ........................... 4 1.2.3Current Econophysics efforts . . . . . . . . . . . . . . . . . . . . . . 5 1.3Objectives ..................................... 6 1.4Contributions ................................... 6 1.5ThesisOutline................................... 7 2 definitions and background 9 2.1SettingtheStage.................................. 9 2.1.1Dataandmodels ............................. 9 2.1.2Financial time series analysis . . . . . . . . . . . . . . . . . . . . . . 10 2.1.3Random Walk Hypothesis and the Brownian Motion . . . . . . . . 11 2.1.4Stylizedempiricalfacts ......................... 12 2.1.5Market Crashes or “When things go terribly wrong” . . . . . . . . 14 2.2StochasticProcesses................................ 19 2.2.1Randomvariables ............................ 19 2.2.2Stochasticprocesses ........................... 20 2.3RandomMatrixTheory ............................. 21 2.3.1Returnsstatistics ............................. 22 2.3.2Thecorrelationmatrix.......................... 23 2.3.3Eigenvalues and eigenvectors . . . . . . . . . . . . . . . . . . . . . . 24 2.4ComponentAnalysis............................... 29 2.4.1Principal Component Analysis . . . . . . . . . . . . . . . . . . . . . 29 2.4.2Independent Component Analysis . . . . . . . . . . . . . . . . . . . 30 2.4.3Forecastable Component Analysis (ForeCA) . . . . . . . . . . . . . 32 2.5Entropy....................................... 33 2.5.1Definition ................................. 34 2.5.2Entropy different incantations . . . . . . . . . . . . . . . . . . . . . 35 2.5.3MutualInformation ........................... 37 2.5.4Kullback-Leibler Divergence . . . . . . . . . . . . . . . . . . . . . . 37 2.5.5ApproximateEntropy .......................... 38 2.6EnergyStatistics.................................. 39 2.6.1Definitions................................. 40 2.6.2Properties ................................. 42 2.6.3BrownianCovariance .......................... 43 2.7FractionalBrownianMotion........................... 44 2.8OtherMethods .................................. 46 2.9Methodologies................................... 47 2.9.1Data Analysis Methodology . . . . . . . . . . . . . . . . . . . . . . . 47 2.9.2Computational Methodology . . . . . . . . . . . . . . . . . . . . . . 48 vii viii contents 3 data 51 3.1DataConsiderations ............................... 51 3.2DataSets...................................... 52 3.2.1PSI-20 set ................................. 52 3.2.2WorldMarketsset ............................ 54 3.3Eventsofinterest ................................. 55 4 portuguese standard index (psi-20)analysis 57 4.1PSI-20 Index.................................... 57 4.1.1PSI-20 evolution.............................. 57 4.1.2A random PSI-20 ............................. 58 4.2Dynamic analysis of PSI-20 using sliding windows . . . . . . . . . . . . . 59 4.2.1Stepsizedecision............................. 59 4.2.2Windowsizedecision .......................... 60 4.3Results ....................................... 63 4.3.1RandomMatrix.............................. 63 4.3.2ComponentAnalysis........................... 66 4.3.3Entropy .................................. 69 4.3.4DistanceCorrelation........................... 71 4.3.5HurstExponent.............................. 73 4.4ConcludingRemarks............................... 75 5 world markets analysis 77 5.1Introduction.................................... 77 5.2Results ....................................... 77 5.2.1RandomMatrix.............................. 77 5.2.2ComponentAnalysis........................... 80 5.2.3Entropy .................................. 83 5.2.4DistanceCorrelation........................... 86 5.2.5HurstExponent.............................. 97 5.3ConcludingRemarks............................... 99 6 conclusions and future work 101 6.1Conclusions .................................... 101 6.2Futurework .................................... 103 a data 105 a.1PSI-20 Stocks.................................... 106 a.2Markets....................................... 118 b catalogue of results 141 b.1Markets Index versus Crisis Dates . . . . . . . . . . . . . . . . . . . . . . . 142 b.2Distance Correlation for PSI-20 ......................... 145 c package description 149 c.1Hash ........................................ 149 c.2PerformanceAnalytics .............................. 149 c.3Zoo ......................................... 150 c.4Pracma....................................... 150 c.5Energy ....................................... 151 c.6Lattice ....................................... 151 contents ix c.7Xts.......................................... 152 c.8xtsExtra....................................... 152 c.9entropy....................................... 152 c.10 ForeCA....................................... 153 d software 155 d.1MarketsMatrixcode ............................... 155 d.2Returnscode.................................... 156 d.3Eigenvaluescode ................................. 157 d.4ApproximateEntropycode ........................... 159 d.5DistanceCorrelationcode ............................ 160 d.6Plotscode ..................................... 161 d.7Kullback-Leibler Divergence code . . . . . . . . . . . . . . . . . . . . . . . 164 d.8MutualInformationcode ............................ 165 d.9ForeCacode .................................... 166 d.10 Marchenko-Pasturcode ............................. 166 bibliography 169 4 introduction Poincaré established the foundations of the chaotic behaviour. The study of chaos turned out to be a major branch of theoretical physics (see Mandelbrot [1977] and Mandelbrot [1982]). For a beautiful and colourful presentation see Peitgen et al. [1992]. More recently chaos theory turned to economy. It was not until the 1990s that physicists started seriously turning to this interdisciplinary subject. Nowadays studies of chaos, self-organized criticality, cellular automata and neural networks are seriously taken into account, as economical and financial tools. 1.2.2Why Econophysics? When addressing the need for a new discipline that merges Physics and Economy two main reasons prevail: 1. The limitations of the traditional approach of Economics/Finance; 2. The advantages of the empirical method used in Physics. In the limitations side we must include the Efficient Market Hypothesis (EMH), by Fama [1970], whose basis is the random walk hypothesis, with independent and identically distributed increments. Despite its popularity, this principle is strongly controversial and has been successively questioned, since it represents a idealization that can hardly be verified. It states, in simple words, that the price variation is random as a result of the activity of the traders who attempt to make profit (arbitrage opportunities); the application of their strategies induces a feedback dynamic in the market, randomising the stock-price. In fact, the idea that markets are rational, from which this theory departs, is a theoretical construction that can be easily violated. Another example stands from the no risk-less Capital Asset Pricing Model (CAPM), by Black and Scholes [1973], which cannot be applied if investors differ in their expectations and if they cannot borrow limitless amount of money at the same interest rate. Also, we could include in this side the so called rationality of economic agents. In the advantages side, we must refer that the appeal from Physics relies on the methodology frequently applied, mainly focused on an experimental basis, which makes the crucial difference between these disciplines. Physicists have learned to be suspicious about axioms and models. If empirical observation is incompatible with the model, the model must be reviewed or discarded, even if it is conceptually beautiful or mathematically convenient. In reality, markets are not efficient, humans tend to be over-focused in the short term and blind in the long term, and errors get amplified through social pressure and herding, ultimately leading to collective irrationality, panic and crashes. Free markets can be, in this sense, actually more like bad tempered or wild markets. It would seem to be foolish to believe that the market can impose its own self-discipline. To sum up, we may say, following Stanley [1999], that the interest of physicists in economic and financial fields, also coined as “statistical finance” is due to three main factors: 1. Economic fluctuations affect everybody, which means that their implications are ubiquitous; 1.2 econophysics 5 2. Methods and concepts developed in the study of fluctuation systems might yield new results; 3. Existence of large data sets in economic/financial domain, which in some cases contains hundreds of millions of events. 1.2.3Current Econophysics efforts It has been proven that reliance on models based on incorrect axioms has clear and tremendous effects. For example, the Black-Scholes model [Black and Scholes,1973] assumes that price changes have a Gaussian distribution, i.e. the probability of extreme events is deemed negligible. Unwarranted use of this model on stock markets led to the October 1987 crash. Ironically, it is the very use of this crash-free Black-Scholes model that “crashed” the market! In the recent sub-prime crisis of 2008 also, the problem lay in part in the development of structured financial products that packaged sub-prime risk into seemingly respectable high-yield investments. The models used to price them were fundamentally flawed: they underestimated the probability of the multiple borrowers would default on their loans simultaneously. In other words, these models again neglected the possibility of a global crisis, even as they contributed to triggering one. Surprisingly, there is no framework in classical economics to understand wild markets, even though their existence is so obvious to the layman. Physicists, on the other hand, have developed several models allowing one to understand how small perturbations can lead to wild effects. The theory of complexity, developed in the physics literature over the last thirty years, shows that although a system may have an optimum state (such as a state of lowest energy), this is sometimes so hard to identify that the system in fact never settles there. This three key ideas presents briefly some of the current efforts in Econophysics [Bentes,2010]: • Statistical characterization of the stochastic process of price changes of a financial asset: this is an active area, and attempts are ongoing to develop the most satisfactory stochastic model describing all the features encountered in empirical analyses. One important accomplishment in this area is an almost complete consensus concerning the finiteness of the second moment of price changes. This has been a long standing problem in finance, and its resolution has come about because of the renewed interest in the empirical study of financial systems. • Development of a theoretical model that is able to encompass all the essential features of real financial markets. Several models have been proposed, and some of the main properties of the stochastic dynamics of stock price are reproduced by these models as, for example, the leptokurtic ’fat-tailed’ non-Gaussian shape of the distribution of price differences. Parallel attempts in the modelling of financial markets have been developed by economists. • Time correlation of a financial series. The detection of the presence of a higherorder correlation in price changes has motivated a reconsideration of some beliefs of what is termed technical analysis. 6 introduction 1.3 objectives The main objective of this work is to apply Econophysics techniques derived from Information and Random Matrix Theories in the study of financial data. The Econophysics techniques applied in this work are twofold: measures of “disorder”/complexity and measures of coherence (for a discussion of coherence and persistence in the scope of financial time series see Ausloos [2001]). The measures of “disorder” and complexity are the different forms of entropy (as defined by Shannon [1948], Rényi [1961], Theil [1967], Tsallis [1988]orSchreiber [2000]). Measures of coherence can be obtained from Random Matrix Theory such as the covariance matrix (see financial applications by Plerou et al. [2000]orLaloux et al. [2000]). The main focus of this thesis is placed, then, on a plethora of measures for the following reasons: 1. They allow us to predict how the market indices will evolve; 2. They add to the portfolio of techniques used to study financial time series; 3. They allow us to characterise the specific features of each market index; 4. They are measures of how markets perceive risk. Each technique captures different nuances of the signal evolution. The use of different tools at the same times allow us to have more confidence in the obtained results, avoiding the several pitfalls of using a single technique. This work carries several types of analyses, from entropy to correlation matrix analysis between different stocks or markets indices. All analyses were performed on daily data from Portuguese PSI-20 stocks and on worldwide markets indices. The daily indices were used as benchmarks for the different stocks or markets studied. Only world markets indices and stock prices from Portuguese Stock Market were used but it should be noted that the same techniques are applicable to other type of financial assets data. We hope that the combination of both families of techniques gives a complementary view of the data in order to search for early warning information and for signs of information transfer by measuring in a quantitative way the transfer of information between stocks or markets. 1.4 contributions The main contributions of this thesis are: 1. All of the seven methods applied have shown interesting and complementary features so that we can not discard none of these methods. 2. Distance Correlation have shown to be a good complement to entropy measures like Mutual Information or Kullback-Leibler Divergence. 3. Approximate Entropy, as a stand alone method, have shown potential complementarity with Distance Correlation in the case of PSI-20 stocks. 4. Hurst Exponent results were vital to confirm that stocks and markets are getting more and more mature, that is, less autocorrelated. 1.5 thesis outline 7 1.5 thesis outline This thesis is organized as follows: • Chapter 2provides a background to some mathematical tools needed, particularly those concerned with Random Matrix Theory (RMT), their eigenvalue analysis and the calculation of the correlation coefficients as the elements of the correlation matrix; also, provides background for those tools related with component analysis like Principal Component Analysis (PCA), Independent Component Analysis (ICA) and Forecastable Component Analysis (ForeCA) and their definition and application to financial time series, namely the entropy and mutual information concepts; finally, some background is given in relatively new tools like the Approximate Entropy and the Energy Statistics and an more old tool like the Hurst Exponent; • Chapter 3considers the data used in this thesis; • Chapter 4characterizes the PSI-20, Portuguese stock market, and applies the methods defined in Chapter 2; also, some concluding remarks are exposed; • in Chapter 5are applied the methods defined in Chapter 2to a vast number of World markets indices; also, again, some concluding remarks are highlighted; • finally, Chapter 6draws the conclusions about the use of these methods in financial time series and propose some work to be done in future studies. In order to keep this text clear and readable, some subjects and results, although interesting, have been placed in Appendix. 2 DEFINITIONS AND BACKGROUND “A very small cause which escapes our notice determines a considerable effect that we cannot fail to see, and then we say that the effect is due to chance.” - Henri Poincaré In this chapter are presented and defined, with mathematical rigour, the tools used in this thesis. Since the main interest is the study of financial time series we start with stochastic processes, firstly developed in the scope of Statistical Physics. Following, are introduced the techniques derived from Random Matrix Theory, Component Analysis, Entropy and Information Theory and Energy Statistics. At the end of the chapter are presented the data and computational methodologies used with these techniques. 2.1 setting the stage Although we must take into account that human beings and particles may behave in a significantly different manner, there is an obvious temptation to create an analogy between economic phenomena (considered a result of the interaction among many heterogeneous agents) and Statistical Mechanics. So, when we talk about basic tools of Econophysics, we are talking about probabilistic and statistical methods often taken from Statistical Physics and/or from Applied Mathematics. 2.1.1Data and models There are, generally, two main routes to problem solving in science: • to use a model and, from there, study the real data to infer the consequences; • to look at the data and from there infer a model. The approach followed in Econophysics is typically the second one, that is, to look first at the data and then to get the best model that describes it. This empirical overview of the data tends to be a first approximation to study a subject. Despite this approach, one of the implicit goals of Econophysics, is to merge these two routes and make a bridge between Econophysics and Economics: data are only useful within an interpretative framework. As with other complex systems, economics, and especially finance has lots of data available. To analyse these data, we have to summarise and reduce them to manage their complexity. In this work we will consider equally spaced data but with one day time interval, which will be named a trading day. The frequency of data must be taken into account because of the granularity effect, that is, as we can see from the literature, measures for different scales yield different results. 9 10 definitions and background 2.1.2Financial time series analysis When studying financial time series the aim is to “understand” them with the ultimate goal to “predict” them (for a good reference on the subject follow Tsay [2005], or, more general, Chatfield [2003]). By this understanding we mean one of these two views: • to model in a mathematical way the time series, that is to say, to represent reality using appropriate mathematical formulae; • to find a set of plausible causes interesting enough to explain the time series behaviour. Also, our starting point includes the common idea that financial time series are intrinsically non-stationary. In Econophysics, it is not usual to study the original financial series. This approach has its drawbacks, although. The one that comes first to mind is that we cannot study stationarity, that is, the long term information. The focus, instead, goes to a transformed quantity (as in the financial literature) named one-day returns. Sometimes these are called log-returns to distinguish them from a similar quantity without the logarithm being applied, xi−xi−1 xi−1. In what follows in this work, returns means always the log-returns. The main reason to use the log-returns has to do with the additive process associated to the time series. For an asset, that is, any good to which we can give a price, with an associated time series xwe have the following definition: Definition 1.Let xibe the value of a time series xat time i.Returns are defined as: ηi=log xi xi−1, (1) where ηiis the return at time step i. Since xiare asset values, they are positive and thus the returns are always well defined. The use of the ratio between two consecutive values makes the quantity dimensionless and the use of logarithms gives a different sign to gains and losses. The distribution of returns was first modelled for bonds, Bachelier [1900], as a Normal distribution, P(r)=1 √2πσ2e−r2 2σ2(2) where σ2is the variance of the distribution. Returns can be used to compare different series, to search for patterns both exclusive to some series only or for the whole group of series. We can, also, use them to give us a new perception of the involved correlations. Also, of interest to a better understanding of the following sections, is the definition of financial volatility. Volatility, σ, corresponds to standard deviation and is a measure for the variation of a price of a financial instrument over time. Definition 2.The annualized volatility σis the standard deviation of the financial instrument’s yearly logarithmic returns. 2.1 setting the stage 11 Therefore, if the daily logarithmic returns of a stock have a standard deviation of σd and the time period of returns is P, the annualized volatility is σ=σd √P. (3) The Equation (3) converts returns or volatility measures from one time period to another assuming a particular underlying model or process because it is an extrapolation of a random walk, or Wiener process, whose steps have finite variance. More generally, though, for natural stochastic processes, the precise relationship between volatility measures for different time periods is more complicated. Some use the Lévy stability exponent αto extrapolate natural processes: σT=T1/ασ. (4) If α=2 we get a Wiener process scaling relation [Mandelbrot,1963]. 2.1.3Random Walk Hypothesis and the Brownian Motion “What if the time series were similar to a random walk?”, or, “It is possible to predict future price movements using the past price movements?” are long asked questions by experts and laymen. Another view of the complexity/disorder is the (fractional) Brownian motion, that appeared in Bachelier PhD thesis, in 1900, [Bachelier,1900], when studying the Paris Stock Exchange as a way to describe the evolution of the financial assets. Louis Bachelier, who firstly proposed a theory of stock market fluctuations, reached the conclusion that “the mathematical expectation of the speculator is zero” and described this condition as a “fair game”. He gave the distribution function the name for what is now known as the Wiener stochastic process (the stochastic process that underlies Brownian Motion) linking it mathematically with the diffusion equation. Feller [1968], called it the BachelierWiener process. This work states that the second order moments of the increments of a heat/diffusion process scale as E{(X(t2)−X(t1))2}∝|t2−t1|, (5) where Xis the stochastic process under study. Henri Poincaré, Bachelier´s advisor, observed that "M. Bachelier has evidenced an original and precise mind [but] the subject is somewhat remote from those our other candidates are in the habit of treating". Nevertheless, his thesis anticipated many of the mathematical discoveries made later by Wiener and Markov, and outlined the importance of such ideas in today’s financial markets, stating that "it is evident that the present theory solves the majority of problems in the study of speculation by the calculus of probability". Later, works from Hurst in the 50’s and Mandelbrot in the 60’s gave rise to the fractional Brownian motion, a generalization of the Brownian motion, firstly described by Bachelier. The Hurst exponent has become an important estimation sign of the financial data disorder or complexity. These two concepts, entropy and fractional Brownian motion, provide a measure of financial data disorder or complexity [Matos et al.,2006]. 12 definitions and background In the seventies, Black, Scholes and Robert Morton, [Black and Scholes,1973], following the ideas of Osborne [1959], Osborne [1977] and Samuelson [1973], modelled the share price as a stochastic process known as a geometric Brownian motion. They also established the isomorphism between the standard deviation of the fluctuations in price of a financial instrument and investment risk. Nowadays, a modern version of Bachelier’s theory is still routinely used in financial literature. This theory predicts a Gaussian probability distribution for stock-price fluctuations. The random walk hypothesis, with independent and identically distributed increments, is the basis of the Efficient Market Hypothesis Fama [1970], as we stated in Chapter 1. Present in Econophysics is the conviction about scaling arguments coming from the study of systems in critical states (see, for instance, Mantegna and Stanley [1995], Cont et al. [1997] or Di Matteo et al. [2005]). The empirical study of those distributions led also to the analysis of distributions of economic shocks, growth rate variations, firm and city sizes. In all these measures scaling laws were found, thus giving confidence that the same type of analysis could be applied to the study of the distributions used to characterise complex systems. 2.1.4Stylized empirical facts Physicists interest in analysing financial data has been to find common or universal regularities in the time series (a different approach from those of the economists doing traditional statistical analysis of financial data). The results of their empirical studies showed that the apparently random variations in time series share some statistical properties which are interesting, non-trivial and common for various values and time periods. These are called stylized empirical facts. The concept of “stylized facts” was introduced in macroeconomics around 1960 by Nicholas Kaldor, who advocated that a scientist studying a phenomenon “should be free to start off with a stylized view of the facts”. In his work, Kaldor [1957] isolated several statistical facts characterizing macroeconomic growth over long periods and in several countries, and took these robust patterns as a starting point for theoretical modelling. This expression has thus been adopted to describe empirical facts that arose in statistical studies of financial time series and that seem to be persistent across various time periods, places, markets or assets. Stylized facts are, then, obtained by taking a common denominator among the properties observed in different markets and financial instruments. By doing so, one gains in generality but tends to lose in precision of the statements one can make about asset returns. Indeed, stylized facts are usually formulated in terms of qualitative properties of asset returns and may not be precise enough to distinguish among different parametric models Cont [2001]. One can find many different lists of these facts in several reviews (see Bollerslev et al. [1994]orCont [2001]). 1. Absence of autocorrelations: linear autocorrelations of asset returns are often insignificant, except for very small intra-day time scales ( 20 minutes) for which microstructure effects come into play. The auto-correlation of log returns rapidly decays to zero for τ≥15 minutes, which supports the Efficient Market Hypothesis. When 2.1 setting the stage 13 τis increased, weekly and monthly returns exhibit some auto-correlation but the statistical evidence varies from sample to sample. 2. Heavy/Fat tails: the distribution of returns seems to display a power-law or Paretolike tail, with a tail index which is finite, between 2 −5 for most data sets studied [Gabaix et al.,2003]. This excludes stable laws with infinite variance and the normal distribution. However, the precise form of the tails is difficult to determine as Mandelbrot [1963] pointed out. The Gaussian/Normal distribution is a special case of the more general Lévy distributions, and is often used as an approximation to log-normal distributions. In contrast, these distributions display power-law decay in the tails and this is related to the fractal nature of financial data [Higushi, 1988], where uni-fractal processes, such as fractional Brownian motion [Mantegna and Stanley,2000,Bouchaud and Potters,2003] and simple multi-fractal processes (see [Lux,2004] and Calvet and Fisher [2002]) have been considered for financial data. The "fat tails" can only be obtained by "nonperturbative" methods, mainly by numerical ones, since they contain the deviations from the usual Gaussian approximations [Nolan,2006]. 3. Gain/loss asymmetry: one observes large draw downs in stock prices and stock index values but not equally large upward movements. 4. Aggregational Gaussianity: as one increases the time scale tover which returns are calculated, their distribution looks more and more like a normal distribution, meaning that the shape of the distribution is not the same at different time scales. The fact that the shape of the distribution changes with τmakes it clear that the random process underlying prices must have non-trivial temporal structure. 5. Intermittency: returns display, at any time scale, a high degree of variability. This is quantified by the presence of irregular bursts in time series of a wide variety of volatility estimators. 6. Volatility clustering: different measures of volatility display a positive autocorrelation over several days, which quantifies the fact that high-volatility events tend to cluster in time, and decays roughly as a power law with an exponent between 0.1 and 0.3. Price fluctuations are not identically distributed and the properties of the distribution, such as the absolute return or variance, change with time. To sum up, large changes tend to be followed by large changes, and analogously for small changes. 7. Existence of nonlinear correlation: Abhyankar et al. [1997] found nonlinear dependence in the four important stock-market indices. Also, Ammermann and Patterson [2003] have shown that nonlinear dependencies play a significant role in the returns for a broad range of financial time series (see http://finance.martinsewell. com/stylized-facts/nonlinearity/ for more details). 8. Conditional heavy tails: even after correcting returns for volatility clustering, the residual time series still exhibit heavy tails. However, the tails are less heavy than in the unconditional distribution of returns. 20 definitions and background If X=Ythen we get the variance of X: VarX=CX,X. (7) The standard deviation of the random variable Xis the square root of variance σX=pVarX. (8) The correlation coefficient of two random variables Xand Yis rX,Y=CX,Y σXσY, (9) where σXand σYare the standard deviations of two stock return series. It is a common measure of the dependence between the return series of the two stocks. The elements of the correlation matrix are restricted to the domain −1≤cij ≤+1: for 0 <cij ≤+1 the stocks are correlated (in a positive way), for −1≤cij <0 the stocks are anti-correlated (correlated in a negative way), and for cij =0 the stocks are uncorrelated. The crosscorrelation defined above calculates the dependence between the return series in the whole period of the sample data. 2.2.2Stochastic processes Definition 6.Let (Ω,F,P)be a probability space. A stochastic process is a collection {X(t)|t∈T}of random variables X(t)defined on (Ω,F,P), where Tis a set, called the index set of the process. Tis usually (but not always) a subset of R. One can also think of a stochastic process as a function X= (X(t,ω)) in two variables: t∈Tand ω∈Ω, such that for each t,Xt(ω):=X(t,ω)is a random variable on (Ω,F,P). Given any t, the possible values of X(t)are called the states of the process at t. The set of all states (for all t) of a stochastic process is called its state space. If Tis discrete, then the stochastic process is a discrete-time process. If Tis an interval of R, then {X(t)|t∈T} is a continuous-time process. If Tcan be linearly ordered, then tis also known as the time. Let X(t)and Y(t)be stochastic processes, with t∈Tand Tbeing the index set. Definition 7.The mean η(t)of X(t)is the expected value of the random variable X(t) ηX(t) = E{X(t)}. (10) The cross-correlation of two processes X(t)and Y(t)is RXY(t1,t2) = E{X(t1)Y(t2)}. (11) The autocorrelation R(t1,t2)of X(t)is the expect value of the product X(t1)X(t2) R(t1,t2) = E{X(t1)X(t2)}. (12) 2.3 random matrix theory 21 The cross-covariance of two processes X(t)and Y(t)is CXY(t1,t2) = E{X(t1)Y(t2)}−ηX(t1)ηY(t2). (13) The autocovariance C(t1,t2)of X(t)is the covariance of the random variables X(t1)and X(t2) C(t1,t2) = R(t1,t2)−η(t1)η(t2). (14) The ratio r(t1,t2) = C(t1,t2) pC(t1,t1)C(t2,t2)(15) is the correlation coefficient of the process X(t). 2.3 random matrix theory The R/S, DFA and Geometric Brownian Motion methods that will be considered in Section 2.7are suitable for analysing univariate data. But, as the stock-market data are essentially multivariate time-series data, it is worth to look for other instruments. Also, in the multivariate signal processing problem, one key issue might be when instabilities occur in signal patterns and how we might determine if the fluctuations are damped, remain at low level, or combine in some way as to cause a major event, e.g. a market crash. Crashes are also interesting since the market dynamics changes during the event (see Mendes et al. [2003], Araújo and Louçã [2006]). Random matrix theory (RMT) is concerned with the study of large-dimensional matrices, in particular with their eigenvalues, eigenvectors and singular values, whose entries are sampled according to known probability densities. The interest in random matrices appeared in the context of multivariate statistics with the works of Wishart and Hsu in the 30´s, but it was only in the 50´s, with Wigner (Wigner [1955] and Wigner [1958]), who introduced random matrix ensembles and derived the first asymptotic result although in the context of nuclear physics. It seems that the problem of interpreting the correlations among large amounts of spectroscopic data on the energy levels, whose exact nature is unknown, is similar of interpreting the correlations among different stocks returns. Therefore, with the minimal assumption of a random Hamiltonian, given by a real symmetric matrix with independent random elements, a series of predictions can be made. In 1967, a seminal paper by Marchenko and Pastur [Marchenko and Pastur,1967] on the spectrum of empirical correlation matrices gave birth to many interesting applications in very different contexts. However, its central objective, as a new statistical tool to analyse large dimensional data sets, only became fully relevant more recently, when the computational storage and handling of huge amounts of data became common to almost all human activity. In fact, the correlations among stock returns have also been addressed by means of the random matrix theory. The quest for the causes that explain the dynamics of Nquantities in a financial context, say for instance, the daily returns of the different stocks of the PSI-20, brought a great development to this subject. 22 definitions and background 2.3.1Returns statistics As stated before, in Econophysics the focus goes to returns. As already know, their distribution is not Gaussian and has fat tails, decaying as a power law. The empirical probability distribution function of the returns on short time scales (from high frequency data to a few days, where we still can assume that the returns have zero mean) can be satisfactory fit by a Student-t distribution [Bouchaud and Potters,2003]: P(r)=1 √π Γ1+µ 2 Γµ 2 aµ (r2+a2)1+µ 2 , (16) where ais related to the variance of the distribution, σ2=a2/(µ−2), and µmoves in the interval [3,5](Plerou et al. [1999], Gopikrishnan et al. [1999]). On longer time scales, from a few weeks to months, the returns distribution approaches a Gaussian [Bouchaud and Potters,2003]. However, we have to point out two restrictions: 1. The returns cannot be used as independently drawn Student random variables, that is to say, returns are far from being considered independent and identically distributed (i.i.d.) random variables: from empirical evidence, it is known that asset returns are clearly not independent as they exhibit certain patterns; 2. Because of their nature there is diminishing predictability of data that are further away from the present. In other words, the volatility of financial returns is itself a dynamical variable over time, having a broad distribution of characteristic frequencies. Formally, the returns at time tcan be represented by the product of a volatility component σtand a directional component ξt[Bouchaud and Potters,2003]: rt=σtξt, (17) where, for instance, the ξtare such that now are i.i.d. random variables with unit variance and σtis a positive random variable with both fast and slow components. Or vice-versa, because, in fact, a Student-t variable can be written as in Equation (17) where the ξis Gaussian and σis an inverse Gamma random variable. Indeed, σtand ξtcannot be considered independent. From the literature (see Bouchaud and Potters [2003] for a review) we know that when considering stock markets, negative past returns tend to increase future volatilities and vice-versa: this is the “leverage” effect, coined by Black in 1976, which tells us that the average of quantities such as ξtσt+τis negative when τ>0. But, going back to Equation (17) and considering the first assumption, the slow part of σtis actually a long memory process such that it correlation function decays as a slow power-law of the time lag τ: σtσt+τ−σ2∝τ−υ.υv0.1 (18) In the more general case of a multivariate distribution of returns there is a need to extend these previous results to a multivariate ambient, where there are Ncorrelated stocks and a joint distribution of simultaneous returns rt 1,rt 2,..., rt N. All marginals of 2.3 random matrix theory 23 this joint distribution must resemble the Student-t distribution, Equation (16), and it must be compatible with the true correlation matrix of the returns: Cij =Z∏ k [drk]rirjP(r1,r2,...,rN). (19) This previous result, Equation (19), leads us to the “copula specification problem” in quantitative finance, that is, a multivariate probability distribution of Nrandom variables uiall having a uniform marginal probability distribution in [0,1]. Further developments about this “copula specification problem” are out of the scope of this thesis. 2.3.2The correlation matrix “Correlation” is defined as “a relation existing between phenomena or things or between mathematical or statistical variables which tend to vary, be associated, or occur together in a way not expected on the basis of chance alone”1. When we discuss about correlations in stock prices, we are interested in the relations between variables such as close prices and transaction volumes, for instance, and more importantly how these relations affect the nature of the statistical distributions which govern the prices variation in the time series. We pay, now, our attention to the estimation of the correlations between the price movements of different assets (for a recent review, Fraham and Jaekel [2008]). Denoting by Tthe total number of observations of each of the Nquantities, say, thinking about stock returns, Tis the total number of trading days in the sampled data. The realization of the ith quantity (i=1,..., N) at “time” t(t=1,..., T) will be rt i. Now, the normalized T×Nmatrix of returns, denoted as X, will be: Xti =rt i √T. If we want to characterize the correlations between these quantities, the simplest form is to compute the Pearson estimator of the correlation matrix: Eij =1 T T ∑ t=1 rt irt j≡XTXij , (20) where Eis the empirical correlation matrix, most probably different from the “true” correlation matrix C: ρt ij = <rt irt j>< rt i>< rt j> rh<rt2 i>−<rt i>2ih<rt2 j>−<rt j>2i, (21) where the <... >gives a time average over the consecutive trading days included in the return vectors. These correlation coefficients fulfill the condition −1≤ρij ≤1 and form an N×Ncorrelation matrix Ct, which serves as the basis of further analyses. Apart for dimensionality, correlation and covariance are very similar concepts. 1In Merriam-Webster Online Dictionary. Retrieved July 31,2014, from http://www.merriamwebster.com/dictionary/correlations 24 definitions and background We also present, here, the covariance matrix with variable weights at time T, over an horizon M,σT(M), that is given by: σT ij (M) = ∑M s=0Wsri,T−srj,T−s ∑M s=0Ws , (22) where ri,tis the value of return riat time t, and Wsis the weight given for the covariance at delay s, (time T−s). The weight vector, W, can be used to have decreasing components since higher weights are attributed to moments closer to the time being analysed. One example traditionally used and the same that is used in this work is Wi=Ri, with 0 <R<1. Then we have ∑T s=0WT−s=RT 1−RT, and Wicorresponds to a geometric series. Typical values (see Litterman and Winkelmann [1998]) are R=0.9 and T=20. Some interesting studies using correlation matrix forecasts of financial asset returns have been done in financial risk management (Embrechts et al. [2002] and Bouchaud and Potters [2003]). In market maturities, Matos et al. [2006] and Sharkasi et al. [2006a], studied the behaviour of eigenvalues of the covariance matrices around crashes and also studied the ratio of the dominant (first eigenvalue) to the sub-dominant (second eigenvalue) for emerging and mature markets. Their results showed that mature markets react to crashes in a different way than emerging ones which, as suggested before, take longer to recover than mature markets. Their investigation also suggests that the second largest eigenvalue may thus be expected to provide additional information on market movements. In more recent years, there are increasing works concentrated on the variation of the cross correlations between market equities over time. Di Matteo et al. [2010] have investigated the evolution of the correlation structure among 395 stocks quoted on the U.S. equity market from 1996 to 2009, in which the connected links among stocks are built by a topologically constrained graph approach. They found that the stocks have increased correlations in the period of larger market instabilities. Fenn et al. [2011] have used the RMT method to analyse the time evolutions of the correlations between the market equity indices of 28 geographical regions from 1999 to 2010, and they also observe the increase of the correlations between several different markets after the credit crisis of 2007-2008. 2.3.3Eigenvalues and eigenvectors The empirical determination of a correlation matrix is a difficult task. If one considers Nassets, the correlation matrix contains N(N−1)/2 mathematically independent elements, which must be determined from Ntime-series of length T. If Tis not very large compared to N, then generally the determination of the covariances is noisy, and therefore the empirical correlation matrix is to a large extent random. The smallest eigenvalues of the matrix are the most sensitive to this ‘noise’. But the eigenvectors corresponding to these smallest eigenvalues determine the minimum risk portfolios in Markowitz theory [Laloux et al.,2000]. It is thus important to distinguish “signal” from “noise” or, in other words, to extract the eigenvectors and eigenvalues of the correlation matrix 2.3 random matrix theory 25 containing real information (those important for risk control), from those which do not contain any useful information and are unstable in time. It is, then, useful to compare the properties of an empirical correlation matrix to a “null hypothesis” - a random matrix which arises, for instance, from a finite timeseries of strictly uncorrelated assets. Deviations from the random matrix case might then suggest the presence of true information. The eigenvalues and eigenvectors of random matrices approach a well-defined functional form in the limit when Ntends to infinity. It is then possible to compare the distribution of empirically determined eigenvalues to the distribution that would be expected if the data were completely random. Obtaining the difference between Eand C was really the goal of the Marchenko and Pastur effort [Marchenko and Pastur,1967]. This difference may be found considering the ratio between Nand T: q=N T. (23) • If Nand Tare about the same order, that is, q∼O(1), then TrE−1=TrC−1/(1−q) [Bouchaud and Potters,2011]. • If Nis small compared to T, then we expect that the Pearson estimator Eis close to its “true” value and so a good estimator of TrC−1is TrE−1. This is the case when q→0, where we get the “true” density of the eigenvalues. • In the opposite, the asymptotic limit, the spectrum of the eigenvalues (their empirical density) is mostly distorted when compared to the “true” density. When T,N→∞the spectrum has some degree of universality with respect to the distribution of the rt i´s. The correlation matrix defined in Equation (20) is a N×Nsymmetric matrix and so we can diagonalize it. This is the beginning of the relationship between Random Matrix Theory and the Principal Component Analysis. Three Classical Results The asymptotic behaviour of random matrices attracted more attention and it was quickly realized that this behaviour is often independent of the distribution of the entries. Furthermore, the limiting distribution typically takes non-zero values only on a bounded interval, displaying sharp edges. Until recently, the majority of the results established were concerned with the spectra, or eigenvalue distributions, of such matrices. But now, the study of the eigenvectors of random matrices also starts to become relevant. Of interest are both the global regime, which refers to statistics on the entire set of eigenvalues, and the local regime, concerned with spacings between individual eigenvalues. In this thesis, we will briefly consider the three classical results and their behaviour in these regimes: 1. Wigner’s semicircle law for the eigenvalues of symmetric or Hermitian matrices; 2. the Marchenko-Pastur law for the eigenvalues of sample covariance matrices; 3. the Tracy-Widom distribution for the largest eigenvalue of Gaussian unitary matrices. 26 definitions and background Wigner’s semicircle law, for example, can be considered universal in the sense that the eigenvalue distribution of a Symmetric or Hermitian matrix with i.i.d. entries, properly normalized, converges to the same density regardless of the underlying distribution of the matrix entries. Also, in this asymptotic limit, the eigenvalues are almost surely supported on the interval [-2,2], illustrating the sharp edges behaviour mentioned before. Historically, results such as Wigner’s semicircle law, were initially discovered for specific matrix ensembles and later were extended to more general classes of matrices. As another example, the circular law for the eigenvalues of a non-symmetric matrix with i.i.d. entries was initially established for Gaussian entries in 1965, but only in 2008 was it fully expanded to arbitrary densities. From a practical standpoint, the benefits of universality are clear, given that the same result can be applied to a vast class of problems. Sharp edges are also important for practical applications. Here, the hope is to use the behaviour of random matrices to separate signals from noise. In such applications, the finite size of the matrices of interest poses a problem when adapting asymptotic results valid for matrices of infinite size. Nonetheless, an eigenvalue that appears significantly outside of the asymptotic range is a good indicator of non-random behaviour. The spectral properties of random matrices are one interesting application of the Central Limit Theorem. In fact, and just considering the simplest ensemble of random matrices, the one where all elements of the matrix Hare i.i.d. random variables and the only constraint being the matrix symmetry (Hij =Hji), in the limit of very large matrices, the distribution of its eigenvalues has universal properties, which can be considered independent of the distribution of the elements of the matrix. So, let us consider a square symmetric matrix H,N×N. The statistics of the eigenvalues λαof large random matrices, in particular the density of eigenvalues ρ(λ), is defined as: ρN(λ)=1 N∑N α=1δ(λ−λα), (24) where λαare the eigenvalues of the N×N symmetric matrix Hunder study and δis the Dirac function. We will need the “resolvent” G(λ)of the matrix H, defined as: Gij (λ)=1 λI−Hij , (25) where Iis the identity matrix. The trace of G(λ), using the eigenvalues of H, is: TrG (λ)= N ∑ α=1 1 λ−λα . (26) And the deduction goes through (see, for a full explanation, [Bouchaud and Potters, 2003]), until we get ρ(λ)=1 2πσ2p4σ2−λ2,|λ| ≤ 2σ(27) which is the “semi-circle” law derived by Wigner in the late fifties of the XX century. In finance we often see correlation matrices C, which are positive definite. Ccan be written as C=HHT, where HTdesignates the transpose. As His, generally, a rectangular matrix of size M×Nwhere Mis the assets number and Nthe observations days, 2.3 random matrix theory 27 then Cwill be M×M. If N=Mthen to get the eigenvalues from Cwe just need to obtain them from H:λC=λ2 H, that is, ρ(λC)dλC=2ρ(λH)dλH, and, by Equation (27), ρ(λC)=1 2πσ2s4σ2−λC λC, 0 ≤λC≤4σ2(28) However, usually N6=M, then we can obtain similar formula if we consider that in the limit N,M→∞, ρ(λC)=Q 2πσ2p(λmax −λC) (λC−λmin) λC(29) and λmax min =σ21+1/Q±2p1/Q(30) with a ratio Q=N M1, λe [λmin,λmax]and σ2being the variance of the elements of C. From Equation (29), and taking into attention that N→∞, we can predict the following: a. The lower “edge” of the spectrum is positive (except the case Q=1 where λmin =0 and therefore it diverges); for the other cases there is no eigenvalue between 0 and λmin. Near this edge the density of the eigenvalues exhibits a sharp maximum; b. The density of eigenvalues vanishes above a certain upper edge λmax. We can treat Equation (25) in a more general way. We will need to define the “resolvent” GH(z)of the matrix H, most well known by Stieltjes transform, as: GH(z)=1 NTr h(zI−H)−1i, (31) where zis a complex number and Iis the identity matrix. Then, the eigenvalues spectrum would be, ρN(λ)=lim e→0 1 π=(GH(λ−ie)) , (32) with =being the imaginary part of the complex number. When Ntends to infinity, in the limit, we almost surely have a unique and well defined density ρ∞(λ)[Bouchaud and Potters,2011]. This asymptotic result, under certain conditions, can be used to describe the eigenvalue density of a single instance. This is probably the cause to RMT great success. Eigenvalues in literature In the last fifteen years, several authors have been applying RMT in a tentative to understand the structure of financial correlation matrices in such a highly random setting. For a first lecture on the problematic Gallucio et al. [1998] will do. Plerou et al. [1999] shown that for the correlation matrix of 406 companies in the S&P index, on daily data, from 1991 to 1996, only seven out of the 406 eigenvalues were clearly significant with respect to a random null hypothesis, that is, the statistics of the most of the eigenvalues of the correlation matrix calculated from stock return series agree with the predictions 28 definitions and background of random matrix theory, but with deviations for a few of the largest eigenvalues, and their corresponding eigenvectors. This was also observed in other studies: Laloux et al. [1999], Laloux et al. [2000], Plerou et al. [2000], Plerou et al. [2001], Plerou et al. [2002], Sharifi et al. [2004] and Wilcox and Gebbie [2004]. Also, in these studies, the correlation (or covariance) matrices of financial time series appeared to contain such a large amount of noise that the eigenvalue structure could essentially be regarded as random. However, some previous studies, see as an example [Gopikrishnan et al.,1999], have focused only on the largest eigenvalue with no attention paid to the others. Extended work by [Plerou et al.,1999] was conducted to explain information contained in the deviating eigenvalues, which revealed that the largest eigenvalue corresponds to a market wide influence to all stocks and the remaining deviating eigenvalues correspond to conventionally identified business sectors. This also suggested that it is possible to improve estimates by setting the insignificant eigenvalues to zero, mimicking a common noise-reduction method used in signal processing. Wilcox and Gebbie [2004] examined the composition of all the eigenvalues of ten years of Johannesburg Stock Exchange. The authors concluded that the leading, that is, the first three, eigenvalues may be interpreted in terms of independent trading strategies with long range correlations indicating a role not just for one but also for a small number of the dominant eigenvalues. This means that only a few of the larger eigenvalues might carry collective information. All these results strongly suggest that eigenvalues of correlation matrix falling under the Marchenko-Pastur distribution contain no genuine information about the financial markets. Hence, one should systematically filter out such noise from the correlations for more accurate estimations of, for instance, future portfolio risk. Following Wilcox and Gebbie [2004], Sharkasi et al. [2006a] we will consider the three larger eigenvalues and its respective eigenvectors as carrying meaningful information. Further, Kwapien et al. [2005] investigated the distribution of eigenvalues of correlation matrices for equally-separated time windows with respect to the German DAX in order to study, quantitatively, the relation between stock price movements and properties of the distribution of the corresponding index motion. They reported that the importance of an eigenvalue is related to the correlation strength of different stocks, which means that the more aggregated the market behaviour, the larger the first eigenvalue (the maximum eigenvalue). In this context, another relevant study is the one done by Drozdz et al. [2007] with a comparison between empirical data and random matrix theory. Dynamics of the top eigenvector The Wigner and the Marchenko-Pastur ensembles are in some sense maximally random as no prior information about the matrices is assumed. But, for stock markets, it is intuitive that stocks are sensitive, for example, to global news about the economy. So, we must have some, at least one, common factor to all stocks. A reasonable null-hypothesis is that the true correlation matrix is: Cii =1, Cij =¯ ρ,∀i6=j. (33) 2.4 component analysis 29 This corresponds to add a rank one perturbation matrix to the empirical correlation matrix with one large eigenvalue N¯ ρand N−1 zero eigenvalues. When Nρ1, the empirical correlation matrix will also have a large eigenvalue close to Nρ. But, what happens when N¯ ρis not very large compared to unity? That case was solved in great detail in 2005 (Bouchaud and Potters [2011]). There it was considered a more general case where the true correlation matrix has kspecial eigenvalues, called “spikes”. So, in general, financial covariance matrices are such that a few large eigenvalues are well separated from the “bulk”, where all other eigenvalues reside. So, again, we expect to have a large eigenvalue λmax ≈N¯ ρwhen stocks are correlated on average. The associated eigenvector is the so-called “market mode”, that is to say, in a first view, all stocks move in the same direction. Plerou et al. [1999] and Plerou et al. [2002] found that the distribution of eigenvector components for the eigenvectors corresponding to the eigenvalues outside the RMT bound displayed systematic deviations from the RMT prediction and that these “deviating eigenvectors” were stable in time. They analyzed the components of the deviating eigenvectors and found that the largest eigenvalue corresponded to an influence common to all stocks. Their analysis of the remaining deviating eigenvectors showed distinct groups, whose identities corresponded to conventionally-identified business sectors. The important question, here, is then if and if yes how do these λmax and ~ Vmax behave in time. 2.4 component analysis Reducing the parameter space is a commonly used approach for successfully modelling multivariate time series, because the number of parameters involved increases quickly with the dimension of the series. Several methods are available to perform dimension reduction, including the canonical correlation analysis (CCA) of Box and Tiao [1977], the factor models of Peña and Box [1987], the independent components analysis (ICA) of Back and Weigend [1997], and the principal components analysis (PCA) of Stock and Watson [2002]. These methods seek linear combinations that have certain characteristics useful in model building: for instance, the CCA produces linear combinations that rank from the most predictable to the least predictable. 2.4.1Principal Component Analysis PCA invention is attributed to Karl Pearson (1901) who created this as an analogue of the principal axes theorem in mechanics; it was later independently developed and named by Harold Hotelling in the 1930s. The method is mostly used as a tool in exploratory data analysis and for making predictive models. In fact, PCA is closely related to RMT, since it is also done through eigenvalue decomposition of the correlation (or covariance) matrix of the return series. This method uses an orthogonal transformation to convert a set of possible correlated returns into several uncorrelated components, which are ranked by their explanatory power for the total variance of the system. 36 definitions and background Definition 11.Let Pebe a partition of disjoint boxes Pj, of size length ≤e, over the support of measure µ. If we consider µ(Pj) = pjthen ∼ Hq(Pe) = 1 1−qlog ∑ j pq j(46) is the q-order Rényi entropy for the partition Pe. Note for q=1 we have to apply the l’Hopital rule where we get ∼ H1(Pe) = −p∑ j pjlog pj. (47) ∼ H1(Pe)is thus the Shannon entropy as defined in Equation (43). In contrast to the other Rényi entropies is additive, i.e., if the probabilities can be factorised into independent factors, the entropy of the joint process is the sum of the entropies of the independent processes. 2.5.2.2Kolmogorov-Sinai entropy The Rényi entropies gain even more relevance when they are applied to transition probabilities, Equation (45). We apply the same reasoning as before: apply a partition Peon the dynamic range of the observable, and introduce the joint probability pi1,i2,...,imthat at an arbitrary time nthe observable falls into the interval Ii1, at time n+1 fall into interval Ii2, and so on. Definition 12.The block entropies of block size mis Hq(m,Pe) = 1 1−qlog ∑ i1,i2,...,im pq i1,i2,...,im. (48) The order-q entropies are then hq=sup P lim m→∞ 1 mHq(m,Pe)⇔hq=sup P lim m→∞hq(m,Pe), (49) where hq(m,Pe):=Hq(m+1, Pe)−Hq(m,Pe),hq(0, Pe) = Hq(0, Pe). (50) In the original sense only h1was called the Kolmogorov-Sinai entropy [Kolmogorov, 1958,Sinai,1959], but since the idea is the same, the name was extended to cover all the other Rényi entropies. Kolmogorov and Sinai were the first to consider correlations in time in information theory. The limit q→0 gives the topological entropy h0. As D0, the fractal dimension of the support of the measure, just counts the number of non-empty boxes in partition, h0gives just a measure of the different orbits, not of their relative importance as we get with h1. Another extension of entropy, related with Rényi entropies, is Tsallis non extensive entropy [Tsallis,1988], with applications to economics described in Tsallis et al. [2003]. 2.5 entropy 37 2.5.3Mutual Information Gaussian processes can be completely defined by second order statistics, namely the mean and the variance, but when talking about non-Gaussian processes higher order statistics are needed. We will make use of second order statistics Correlation Coefficient and the high order statistics known as Mutual Information (MI) to measure the dependency between two random variables. In fact, the Mutual Information, though hard to compute, is a natural measure of the independence between random variables. MI accounts for the whole dependency structure and not only the covariance. We can define the Mutual Information by the entropies H(X),H(Y)and H(X,Y)(see for example Papoulis [1985]): MI (X;Y)=H(X)−H(X|Y)(51) H(X|Y)=H(X,Y)−H(Y)(52) MI (X;X)=H(X). (53) Mutual Information is always non-negative and zero if and only if the variables are statistically independent. 2.5.4Kullback-Leibler Divergence Following the 1951 classical paper of S. Kullback and R.A. Leibler entitled “On information and sufficiency” [Kullback and Leibler,1951] it is presented the Kullback-Leibler divergence. Kullback and Leibler were concerned with the statistical problem of discrimination, by considering a measure of the “distance” or “divergence” between statistical populations in terms of their measure of information. For independent signals, the joint probability can be factorized into the product of the marginal probabilities. Therefore, the independent components can be found by minimizing the Kullback-Leibler divergence, or distance, between the joint probability and marginal probabilities of the output signals [Amari et al.,1996]. Hence, the goal of finding statistically independent components can be expressed in several ways: look for a set of directions that factorize the joint probabilities and, then, find a set of “interesting” directions with minimum mutual information. Where the mutual information between variables vanish, they are statistically independent. The goal of finding interesting directions is similar to projection pursuit (Friedman and Tukey [1974] and Huber [1985]). In the knowledge discovery and data mining community the term "interestingness" (Ripley [1996]) is also used to denote unexpectedness (Silberschatz and Tuzhilin [1996]). Assuming that Hi,i=1, 2, is the hypothesis that xwas selected from the population whose density function is fi,i=1,2, then we define log f1(x) f2(x)(54) 38 definitions and background as the information in xfor discriminating between H1and H2. In their seminal paper (Kullback and Leibler [1951]), they have denoted by I(1,2)the mean information for discrimination between H1and H2per observation from f1, i.e., I(1,2)=KLx(f1,f2)=Zf1(x)log f1(x) f2(x). (55) This quantity, in Equation (55) is called the Kullback-Leibler divergence and is denoted by KL (f1,f2), despite the fact that, originally, Kullback and Leibler denoted J(1,2)=KL (f1,f2)+KL (f2,f1)(56) as the divergence between f1and f2. Now, let us consider some properties of Kullback-Leibler divergence: •KL (f1,f2)≥0 with KL (f1,f2)=0 if and only if f1(x)=f2(x)almost everywhere; •KL (f1,f2)6=KL (f2,f1), that is, KL (f1,f2)is not symmetric; •KL (f1,f2)is additive for independent random events: KLxy (f1,f2)=KLx(f1,f2)+ KLy(f1,f2), being Xand Yindependent variables; For most densities f1and f2,KL (f1,f2)needs to be computed numerically. One exception is when f1and f2are both Gaussian distributions. In the univariate case, the Kullback-Leibler divergence between two Gaussian distributions p,qwith means µ1,µ2and variances σ2 1,σ2 2, is given by KL (p,q)=log σ1 σ2+σ2 1+(µ1−µ2)2 2σ2 2−1 2. (57) In the multivariate case, the Kullback-Leibler divergence between multivariate Gaussian distributions p,qis given by: KL (p,q)=0.5 hlog (det(Σ2)/det(Σ1))+tr Σ−1 2Σ1+(µ2−µ1)´Σ−1 2(µ2−µ1)−Ni, (58) with mean vectors µ1,µ2and covariance matrices Σ1,Σ2. 2.5.5Approximate Entropy The Approximate Entropy (ApEn) method is an information theory based estimate of the complexity of a time series introduced by Steve Pincus [Pincus,1991], formally based on the evaluation of joint probabilities, in a way similar to the entropy of Eckmann and Ruelle [Eckman and Ruelle,1985]. The original motivation and main feature, however, was not to characterize an underlying chaotic dynamics, rather to provide a robust model-independent measure of the randomness of a time series of real data, possibly - as it is usually in practical cases - from a limited data set affected by a superimposed noise. ApEn has been used by now to analyse data obtained from very different sources. See, for instance, Ho et al. [1997]. These authors point some weaknesses to ApEn, namely its 2.6 energy statistics 39 strong dependence on sequence length and its poor self-consistency (i.e., the observation that ApEn for one data set is larger than ApEn for another for a given choice of parameters should, but does not, hold true for other parameters choices). Given a sequence of Nnumbers {u(j)}={u(1),u(2),..., u(N)}, with equally spaced times tj+1−tj≡ 4t=const, one first extracts the sequences with embedding dimension m, that is, x(i)={u(i),u(i+1), ..., u(i+m−1)}, with 1 ≤i≤N−m+1. The ApEn is then computed as ApEn =Φm(r)−Φm+1(r), (59) where ris a real number representing a threshold distance between series, and the quantity Φm(r)is defined as Φm(r)=<ln [Cm i(r)] >= N−m+1 ∑ i=1 ln Cm i(r) N−m+1. (60) Here Cm i(r)is the probability that the series x(i)is closer to a generic series x(j)with (j≤N−m+1)than the threshold r, Cm i(r)=N[d(i,j)≤r] N−m+1, (61) with N[d(i,j)≤r]the number of sequences x(j)close to x(i)less than r. As definition of distance between two sequences, the maximum difference (in modulus) between the respective elements is used, d(i,j)=max k=1,2,...,m(|u(j+k−1)−u(i+k−1)|). (62) For a somewhat more mathematical presentation of this subject see Rukhin [2000]. Only more recently this method as been introduced to financial time series (Pincus and Kalman [2004] and Pincus [2008]). 2.6 energy statistics Energy statistics and energy distance are concepts developed by Székely et al. [2007] and were born in the more broad field of independence [Bakirov et al.,2006]. Energy statistics is based on the notion of potential energy as presented by Newton. Statistical observations are like heavenly bodies governed by a statistical potential energy which is zero only when an underlying statistical null hypothesis is present. In this way, energy statistics are functions of distances between statistical observations. Distance correlation is a recent multivariate dependence coefficients approach to the problem of measuring the dependence between random vectors, even if they are arbitrary and/or not of equal dimension. The pertinence of this measure to this work relies on the fact that an interesting approach to measure complicated dependence structures in multivariate data (see, for instance, Embrechts et al. [2002] or Feuerverger [1993]) is to study their vectors independence. 40 definitions and background 2.6.1Definitions Energy distance was introduced in 1985 and is a (statistical) distance between probability distributions. If Xand Yare independent random vectors in Rdwith cumulative distribution functions Fand Grespectively, then the energy distance between these distributions is: D(F,G)=2EkX−Yk−EkX−X´k−EkY−Y´k(63) where X,X´ and Y,Y´ are independent and identically distributed. D(F,G)=0 if and only if Xand Yare identically distributed. Later, Székely et al, based on this energy statistics, developed the concept of distance covariance (dCov) as the square root of ν2 n=1 n2 n ∑ k,l=1 AklBkl, (64) where Akl and Bkl are linear functions of the pairwise distance between sample elements. The distance correlation goes beyond the classical Pearson product-moment correlation, ρ, when in the multivariate environment because the diagonal covariance matrix generated implies independence but it is not a sufficient condition for independence. Over the years other methods have been proposed, and one of them, most notably proposed by Rényi called maximal correlation. For all distributions with finite first moments, the distance correlation Rgeneralizes the idea of correlation in, at least, two ways: 1.R(X,Y)is defined for Xand Yin arbitrary dimensions; 2.R(X,Y)=0 characterizes independence of Xand Y. This coefficient R(X,Y)satisfies 0 ≤R(X,Y)≤1 and R(X,Y)=0 only if Xand Yare independent. In this way distance covariance and distance correlation provide a natural extension of Pearson product-moment covariance σX,Yand correlation ρ. Let Xin Rpand Yin Rqbe random vectors, where pand qare positive integers. We will also denote fXas the characteristic function of X,fYas the characteristic function of Yand fX,Yas the joint characteristic function of Xand Y.Xand Yare independent if and only if fX,Y=fXfY, in what concerns characteristic functions. So, it is a natural idea to try to find a suitable norm to measure the distance between fX,Yand fXfY. Székely and Rizzo [2009] defined a measure of dependence ν2(X,Y;w)=kfX,Y(t,s)−fX(t)fY(s)k2 w, (65) that is, ν2(X,Y;w)=ZRp+q|fX,Y(t,s)−fX(t)fY(s)|2w(t,s)dt ds, (66) with a suitable choice of an arbitrary positive weight function w(t,s)so that this measure of dependence is analogous to classical covariance, but with the property that ν2(X,Y;w)=0 if and only if Xand Yare independent. 2.6 energy statistics 41 Definition 13.The distance covariance (dCov) between random vectors Xand Ywith finite first moments (that is EkXkp<∞and EkYkq<∞) is the non-negative number ν(X,Y)defined by ν2(X,Y)=kfX,Y(t,s)−fX(t)fY(s)k2, (67) where tand sare vectors. Similarly, Definition 14.Distance variance (dVar) is defined as the square root of ν2(X)=ν2(X,X)= kfX,X(t,s)−fX(t)fX(s)k2. By definition of the norm k.k, it is clear that ν(X,Y)≥0 and ν(X,Y)=0 if and only if Xand Yare independent. We can now define distance correlation. Definition 15.The distance correlation (dCor) between random vectors Xand Ywith finite first moments is the non-negative number R(X,Y)defined by R2(X,Y)=           ν2(X,Y) √ν2(X)ν2(Y),ν2(X)ν2(Y)>0; 0, ν2(X)ν2(Y)=0. (68) Remains the problem of the calculus of these quantities. To define the distance dependence statistics we consider a random sample (X,Y)={(XK,YK):k=1, ..., n}of n i.i.d random vectors (X,Y)from the joint distribution of the random vectors Xand Rp and Yand Rq. Then to compute the Euclidean distance matrices (akl)=|Xk−Xl|p and (bkl)=|Yk−Yl|pwe define Akl =akl −¯ ak.−¯ a.l+¯ a..,k,l=1, ..., n, where ¯ ak.=1 n n ∑ l=1 akl,¯ a.l=1 n n ∑ k=1 akl,¯ a.. =1 n2 n ∑ k,l=1 akl. (69) Similarly we define Bkl =bkl −¯ bk.−¯ b.l+¯ b..,k,l=1, ..., n. Definition 16.The non-negative sample distance covariance νn(X,Y)and sample distance correlation Rn(X,Y)are defined by ν2 n(X,Y)=1 n2 n ∑ k,l=1 AklBkl, (70) and R2 n(X,Y)=           ν2 n(X,Y) √ν2 n(X)ν2 n(Y),ν2 n(X)ν2 n(Y)>0; 0, ν2 n(X)ν2 n(Y)=0, (71) 42 definitions and background respectively, and where the sample distance variance is defined by ν2 n(X)=ν2 n(X,X)=1 n2 n ∑ k,l=1 A2 kl. (72) 2.6.2Properties Here, we will show some properties taken from the theorems in Székely and Rizzo [2009] and from previous results in Székely et al. [2007]. Theorem 17.If (X,Y)is a sample from the joint distribution of (X,Y), then ν2 n(X,Y)= kfn X,Y(t,s)−fn X(t)fn Y(s)k2. We must remark that this result is an alternative way of calculating Equation (70) but, as stated in the literature, a much harder and time consuming way. Theorem 18.If E |X|p<∞and E |Y|q<∞, then almost surely lim n→∞νn(X,Y)=ν(X,Y). Corollary 19.If E |X|p+|Y|q<∞, then almost surely lim n→∞R2 n(X,Y)=R2(X,Y). Theorem 20.For random vectors X ∈Rpand Y ∈Rqsuch that E |X|p+|Y|q<∞, the following properties hold: (i) 0≤R(X,Y)≤1, and R =0if and only if X and Y are independent. (ii) ν(X)=0implies that X =E[X], almost surely. (iii) If X and Y are independent, then if ν(X+Y)≤ν(X)+ν(Y). Equality holds if and only if one of the random vectors X or Y is constant. Proof of this last statement can be found in Székely and Rizzo [2009]. Theorem 21.(i) ν(X,Y)≥0. (ii) ν(X,Y)=0if and only if every sample observation is identical. (iii) 0≤Rn(X,Y)≤1. (iv) Rn(X,Y)=1implies that the dimensions of the linear subspaces spanned by Xand Y respectively are almost surely equal, and if we assume that these subspaces are equal, then in this subspace Y=a+bXC for some vector a, non-zero real number b and orthogonal matrix C. When considering that (X,Y)has a bivariate normal distribution, there is a deterministic relation between Rand |ρ|. Theorem 22.If X and Y are standard normal, with correlation ρ=ρ(X,Y), then: (i) R (X,Y)≤|ρ|, (ii) R2(X,Y)=ρarcsin ρ+√1−ρ2−ρarcsin(ρ/2)−√4−ρ2+1 1+π/3−√3, (iii) inf æ6=0 R(X,Y) |ρ|=lim ρ→0 R(X,Y) |ρ|=1 2(1+π/3−√3)1/2 ∼ =0.89066. 2.6 energy statistics 43 2.6.3Brownian Covariance To define Brownian covariance, let Wbe a two-sided one-dimensional Brownian motion/Wiener process with expectation zero and covariance function |s|+|t|−|s−t|=2 min (s,t),t,s≥0. (73) Comparing to the standard Wiener process, this is twice the covariance. Definition 23.The Brownian covariance or the Wiener covariance of two real-valued random variables Xand Ywith finite second moments is a non-negative number defined by its square ω2(X,Y)=Cov2 W(X,Y)=E[XWX´WYW´Y´W´], (74) where (W,W´)does not depend on (X,Y,X´,Y´). It is interesting to note that if in CovWwe replace Wby the identity function, id, then Covid (X,Y)=|Cov(X,Y)|=|σX,Y|, the absolute value of Pearson´s product-moment covariance. While the standardized product-moment covariance, Pearson correlation (ρ), measures the degree of linear relationship between two real-valued variables, we shall see that standardized Brownian covariance measures the degree of all kinds of possible relationships between two real-valued random variables. We will extend now the definition of CovW(X,Y)to random processes in higher dimensions. If Xis an Rp−valued random variable, and U(s)is a random process defined for all s∈Rpand independent of X, define the U−centered version of Xby XU=U(X)−E[U(X)|U], (75) whenever the conditional expectation exists. Definition 24.If Xis an Rp−valued random variable, Yis an Rq−valued random variable and U(s)and V(t)are arbitrary random processes defined for all s∈Rp,t∈Rq, then the (U,V)covariance of (X,Y)is defined as the non-negative number whose square is Cov2 U,V(X,Y)=E[XUX´UYV‘Y´V´], (76) whenever the right-hand side is non-negative and finite. In particular, if Wand W´ are independent Brownian motions with covariance function as Equation (73) on Rpand Rqrespectively, the Brownian covariance of Xand Yis defined by ω2(X,Y)=Cov2 W(X,Y)=Cov2 W,W´(X,Y). (77) Similarly, for random variables with finite variance the Brownian variance is ω(X)=VarW(X)=CovW(X,X). (78) Definition 25.The Brownian correlation is defined as 44 definitions and background CorW(X,Y)=ω(X,Y) pω(X)ω(Y)(79) whenever the denominator is not zero; otherwise CorW(X,Y)=0. We finish this part with the surprising result from the next theorem. Theorem 26.For arbitrary X ∈Rpand Y ∈Rqwith finite second moments ω(X,Y)=ν(X,Y). To summarise the results from Székely et al. [2007], distance covariance and distance correlation are natural extensions and generalizations of classical Pearson covariance and correlation in possibly three ways. 1. In one direction, the ability to measure linear association to all types of dependence relations was extended; 2. In another direction, the bivariate measure to a single scalar measure of dependence between random vectors in arbitrary dimension was also extended; 3. In addition to the obvious theoretical advantages, there are the practical advantages that dCov and dCor statistics are computationally simple and applicable in arbitrary dimension not constrained by sample size. Probably dCov is not the only possible or the only reasonable extension with the above mentioned properties, but this extension was received as a natural generalization of Pearson’s covariance in the sense that the covariance of random vectors was defined with respect to a pair of random processes, and if these random processes are i.i.d. Brownian motions, which is a very natural choice, then we arrive at the distance covariance; on the other hand, if we choose the simplest non-random functions, a pair of identity functions (degenerate random processes), then we arrive at Pearson’s covariance. To sum up, distance correlation extends the properties of classical correlation to multivariate analysis and the general hypothesis of independence. 2.7 fractional brownian motion Two of the most important and simple models of probability theory and financial econometrics are the random walk and the Martingale theory. They assume that the future price changes only depend on the past price changes. Their main characteristic is that the returns are uncorrelated. But are they truly uncorrelated or are there long-time correlations in the financial time series? This question has been studied especially since it may lead to deeper insights about the underlying processes that generate the time series (see, for instance, Lo [1991], Ding et al. [1993] and Harvey [1993] or, for a more recent review, Doukhan et al. [2003]). Depending on the scientific field there are, typically, more then ten measures to quantify the long-time correlations. In the financial literature we find two methods: the Rescaled Range analysis (R/S) and the detrended fluctuation analysis (DFA). For further details see Taqqu et al. [1995]. 2.7 fractional brownian motion 45 In the 50’s, Hurst, while analysing hydrological flows, proposed a single exponent to characterise time variation in time series [Hurst,1951]. This approach is a generalisation of Brownian motion later called fractional Brownian motion [Mandelbrot and Van Ness, 1968], and is characterised by a single exponent, called Hurst exponent. Another way of estimating the Hurst exponent was introduced via DFA by Peng et al. [1994] while studying DNA patterns and their characteristics. In order to measure the strength of trends or “persistence” in different processes, the rescaled range (R/S) analysis to calculate the Hurst exponent can be used. One studies the rate of change of the rescaled range with the change of the length of time over which measurements are made. We divide the time series ξtof length Tinto Nperiods of length τsuch that Nτ=T. For each period i=1,2,..., Ncontaining τobservations, the cumulative deviation is X(τ)= iτ ∑ t=(i−1)τ+1 (ξt−hξit), (80) where hξitis the mean within the time-period and is given by hξit=1 τ iτ ∑ t=(i−1)τ+1 ξt. (81) The range in the i−th time period is given by R(τ)=max X(τ)−min X(τ), and the standard deviation is given by S(τ)="1 τ iτ ∑ t=(i−1)τ+1 (ξt−hξit)2#1/2 . (82) Then R(τ)/S(τ)is asymptotically given by a power-law R(τ)/S(τ)=kτH(83) where kis a constant and Hthe Hurst exponent. In general, “persistent” behaviour with fractal properties is characterized by a Hurst exponent 0.5 <H≤1, random behaviour by H=0.5 and “anti-persistent” behaviour by 0 ≤H<0.5. Usually the Equation (83) is rewritten in terms of logarithms, log (R(τ)/S(τ)) = Hlog (τ)+log (k), and the Hurst exponent is determined from the slope. In the DFA−nmethod, the time-series ξtof length Tis first divided into Nnonoverlapping periods of length τsuch that Nτ=T. In each period i=1,2,..., Nthe time-series is first fitted through a polynomial function zn(t)=antn+an−1tn−1+a0, called the local trend. In this thesis we use a quadratic function n=2 as our fit function. Then it is detrended by subtracting the local trend, in order to compute the fluctuation function, F(τ)="1 τ iτ ∑ t=(i−1)τ+1 (ξt−hξit)2#1/2 . (84) 52 data 3.2 data sets 3.2.1PSI-20 set The PSI-20 set is formed by twelve stocks that were obtained from the PSI-20 Index, which is a price index calculation based on 20 stocks obtained from the universe of Portuguese companies listed to trade on the Main Market and was designed to became the underlying element of futures and options contracts. The choice criteria were two: • the availability of data in the period 2001-2014, to maximize the days where all the stocks were in the market; • the best PSI-20 representation, that is, stocks from almost all the sectors and from different importance. In Table 3are summarized the stocks used with their respective business sector. Data and summary statistics on the markets studied are recorded and are presented in Appendix A. Abrev. Stock Name Sector ^BES Banco Espírito Santo Financial Services ^BPI Banco Português de Investimento Financial Services ^EDP Energias de Portugal Electricity ^JMT Jerónimo Martins Distribution ^EGL Mota-Engil Construction ^NBA Novabase Technological Services ^PTI Portucel Paper ^PTC Portugal Telecom Telecommunications ^SEM Semapa Paper ^SONC Sonae Com Telecommunications ^SON Sonae SGPS Distribution ^ZON Zon Optimus Media Table 3: PSI-20 set business sectors The data used in this study are the close values and its log returns from these 12 stocks and cover the period common to all stocks from January 25,2001 to September 13,2013 for a total of 3362 observations. For a more close look to PSI-20 stocks degree of importance, based on their stock market capitalization, we can see in Table 4their “top ten” classification between 2000 and 2013. As we can see, from the 12 chosen stocks, only sensibly half are represented in this top ten. The idea, here, was to choose representative stocks. It is also possible to analyse particular stock “movements” in this classification but this is out of scope of this study. 3.2 data sets 53 Position 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 1st PTC PTC PTC PTC PTC EDP EDP EDP JMT 2nd PTC EDP EDP PTC EDP PTC JMT JMT 3rd EDP EDP BES EDP PTC PTC EDP EDP EDP EDP 4th BES BES EDP BES BES BES BES PTC JMT BES BES 5th BES BES PTC 6th ZON BPI BES JMT PTC 7th BPI ZON ZON BPI BES BES PTI PTC 8th BPI SON BPI BPI BPI JMT ZON 9th BPI ZON SON SON SON SON SON PTI SON PTI 10th ZON SON JMT JMT JMT ZON JMT BPI BPI PTI BPI SON Table 4: PSI-20 set top-ten classification 3.2.1.1Stock splits and other corrections In order to obtain correct data we needed to study the stocks history, namely the stock splits. Stock splits are conceptually a simple corporate event that consists in the division of each share into a higher number of shares of smaller par value. These operations have long been a part of financial markets. Abrev. Stock-Split Rights Issues Exceptions Date Last Price Next Price Date LP NP Goal ^BES 2000-Jul-11 25.70 17.35 2002-Feb-06 14.35 11.40 2006-Apr-27 15.00 11.59 2009-Mar-19 5.54 3.65 2012-Apr-16 1.05 0.65 ^BPI 2000-Oct-30 3.99 3.82 2006-Mar-13 4.24 5.33 take over threat (BCP) 2008-Jun-20 2.92 2.81 ^EDP 2000-Jul-17 17.95 3.64 ^JMT 2007-May-28 22.00 4.54 2004-Jun-08 9.78 8.64 ^EGL 2001-Jan-23 8.35 1.66 2000-Aug-07 11.30 11.40 ^NBA ^PTI 2001-Jan-22 7.35 1.44 2001-Sep-04 0.91 0.90 ^PTC ^SEM 2000-Sep-14 19.98 3.96 ^SONC ^SON 2000-Jun-21 50.61 9.65 2005-12-27 1.22 0.95 spin-off Sonae Industria ^ZON 2005-Jun-14 3.38 6.77 social capital reduction Table 5: PSI-20 stock splits Portugal, for instance, witnessed 26 of these operations from 1999 (the year the Euro was introduced) to June 2003 essentially due to a legislative change that took place when the corporate law was adapted for the change from Escudo to Euro [Pereira and Cutelo, 54 data 2010]. Stock splits are associated with positive abnormal returns in the short run (around the announcement dates and ex-dates). If a company has undergone stock splits over its lifetime, comparing historical stock prices to those of the present day would not accurately reflect performance. For this reason, we must compare split-adjusted share prices. For discerning and analysing the real performance of the stock, it is standard to adjust the old prices to reflect the splits. In other words, we have to find the present equivalent of the past prices. In Table 5are shown the main operations concerning the twelve PSI-20 stocks studied. This information is partially adapted from Pereira and Cutelo [2010]. 3.2.2World Markets set The choice of the markets used in this study was driven by the goal of studying major markets across the world in an effort to ensure that tests and conclusions could be as general as possible. In Table 6we summarise the markets used in this study. Data and summary statistics on the markets studied are recorded and are presented in Appendix A. Abrev. Index Name Country Region ^AEX Amsterdam Exchange Index Netherlands Europe ^ASX Australian Securities Exchange Australia Asia/Pacific ^ATX Austrian Traded Index Austria Europe ^BSESN Bombay Stock Exchange India Asia/Pacific ^BVSP Bovespa - Bolsa de Valores de S. Paulo Brazil America ^CAC Compagnie des Agents de Change France Europe ^DAX Deutscher Aktien Index Germany Europe ^DJI Dow Jones Industrial Average United States America ^FTSE Footsie United Kingdom Europe ^HSI Hang Seng Index Hong Kong Asia/Pacific ^IBEX Índice Bursátil Espanol Spain Europe ^IXIC Nasdaq Composite United States America ^JKSE Jakarta Stock Exchange - Composite Index Indonesia Asia/Pacific ^KOSPI Seoul Composite South Korea Asia/Pacific ^MERVAL Mercado de Valores de Buenos Aires Argentina America ^MIB Milano Italia Borsa Italy Europe ^MXX IPC - Mexican Stock Exchange Index Mexico America ^NIK Nikkei Tokyo Japan Asia/Pacific ^PSI20 Portuguese Stock Index Portugal Europe ^SPY S&P 500 United States America ^SSMI Swiss Market Switzerland Europe ^STOXX DJ Euro Stoxx 50 Europe ^STRAITS Straits Times Singapore Asia/Pacific Table 6: World Markets Set We have considered here the major and most active markets worldwide from America (North and South), Asia/Pacific, Africa and Europe. The data used in this work are the 3.3 events of interest 55 daily Close values for these 23 markets obtained from January 2,2001 to September 25, 2013. In the chapters that follow when we refer the values for markets and/or compare them we are actually comparing the (log-) return of the chosen index for that market. This decision was made in order to simplify the language. Subsequently, we obtained the “common data”, i.e., the subset of days where all the markets are open, excluding local holidays and periods where the transaction of any market was suspended. Regardless these strict criteria, the data used in this work make for a total of 2965 common daily Close values. 3.3 events of interest As noted in Chapter 2, Section 2.9.1, a sliding window approach will be used to analyse and calculate the values for the different measures for the data sets. This will help us to confine the search for “early warning signs” to a few windows before and after the events of interest. Also, some “neutral” events are going to be explored using the same methodology in order to perform a comparative analysis. The chosen events of interest are the recession dates proposed by NBER (see SubSection 2.1.5in Section 2.1in Chapter 2). So, we are going to look in more detail the following periods: • from 14-02-2001 until 09-11-2001, the first XXI recession and the respective before and after recession periods: from 04-01-2001 until 13-02-2001 and from 12-11-2001 until 17-01-2002; • from 16-11-2007 until 17-06-2009, the second XXI recession and the respective before and after recession periods: from 02-08-2007 until 14-11-2007 and from 18-062009 until 09-09-2009; These before and after periods were chosen to be, approximately, about 20% each of the total recession period. This criterion was due to the availability of the data (mainly for the before recession period). For the “neutral” periods we considered the following two: • from 19-02-2004 until 26-08-2004, the first neutral period and the respective before and after neutral periods: from 08-01-2004 until 18-02-2004 and from 27-08-2004 until 08-10-2004; • from 07-06-2011 until 13-03-2013, the second XXI neutral period and the respective before and after neutral periods: from 30-12-2010 until 26-05-2011 and from 14-032013 until 25-06-2013; In the next two chapters the techniques presented in Chapter 2will be applied to the data sets presented in this chapter. 4 PORTUGUESE STANDARD INDEX (PSI-2 0 ) ANALYSIS “One of the funny things about the stock market is that every time one person buys, another sells, and both think they are astute”. William Feather In this chapter we will apply the mathematical tools presented/described in Chapter 2 to the PSI-20 data set. Let us start by presenting some of the features of this index. 4.1 psi-20 index The Portuguese Stock Index PSI-20 is the national benchmark index, reflecting the price evolution of the 20 largest most liquid assets selected from the set of companies listed on the Portuguese Main Market. The rules for construction of PSI-20 are published PSI [2003], but can be summarised briefly as giving a different weight to each asset belonging to the index, such that no asset has more than 20% of the total weight. PSI-20 had its beginning in January 4th, 1993. Figure 4shows the PSI-20 index evolution from January 24,2000 to September 25, 2013. 2000 2005 2010 4000 8000 12000 time Close value Psi−20 Index Figure 4: PSI-20 from 2000 to 2014 4.1.1PSI-20 evolution After the 2000 peak (roughly corresponding to the dotcom bubble burst), we essentially assist to a decline in the index value until the end of 2002. Additionally, the sub-sample period January 2,2001 to November 23,2001 was characterized by a climate of economic and political instability in Europe and United States due to the high value of the Dollar against the Euro, the Israel-Palestinian conflict, and the terrorist attacks on September 11, 2001 and the subsequent climate of uncertainty, with negative impacts on the financial markets, including the Portuguese stock market. 57 58 portuguese standard index (psi-20)analysis In this period the PSI-20 index declined by 24,42 per cent. Between 2002 and 2007 we assisted to world markets recovery, but in 2008, with the mortgage and sub-prime crises, the world markets in general, and PSI-20 in particular, went down once again. Some ups and downs are found between 2009 and 2011, with the market/investors probably still “astonished” with what had happened before. In the first quarter of 2011 another fall, a period coincident with the international assistance program applied to Portugal. Finally, from the beginning of the second quarter of 2012 we are having some recovery signals in the PSI-20 index. 4.1.2A random PSI-20 Now, we generated a shuffled data by randomly reordering the full return time series for the PSI-20 index. This process destroys the temporal correlations between the return time series but preserves the distribution of returns for each series was we can see in Figure 5. 2000 2005 2010 −0.10 0.00 0.05 0.10 time Value Psi−20 Returns (a) PSI-20 returns 2000 2005 2010 −0.10 0.00 0.05 0.10 time Value Random psi−20 Returns (b) Random PSI-20 returns Figure 5: Real vs Random PSI-20 returns. To try to highlight interesting features in the correlations, we compare the real PSI-20 close values to a corresponding distribution for randomly shuffled returns (a random PSI-20 close values). For a visual comparing between these markets we present Figure 6. 2000 2005 2010 4000 8000 12000 time Close value Real psi−20 vs Random psi−20 Figure 6: Real versus Random PSI-20 close values 4.2 dynamic analysis of psi-20 using sliding windows 59 As we are going to work all the time with returns, now we show their values along time and their distribution (see Figure 7). According to Rege et al. [2013] the distribution of the returns of the PSI-20 exhibits much higher kurtosis and extreme values than the Normal distribution do. They also found that the best fit is provided by the Student t and the Generalized Hyperbolic distributions. 2000 2005 2010 −0.10 0.00 0.05 0.10 time Value Psi−20 Returns (a) PSI-20 returns −0.15 −0.10 −0.05 0.00 0.05 0.10 0.15 0 10 30 PSI−20 returns density N = 2024 Bandwidth = 0.001843 Density (b) PSI-20 returns density Figure 7: PSI-20 returns time series and their distribution. A broader and earlier study reaching the same conclusions but applied to a “World Market Index” was done by Fergusson and Platen [2006]. 4.2 dynamic analysis of psi-20 using sliding windows In Section 2.9.1a sliding/rolling windows approach was introduced. The nature of the approach (i.e. based on the interval characterisation) means that we can apply these techniques to different intervals of fixed size (20, 60 and 120 points, corresponding, approximately, to 1month, 3months and 6months of data). Each one of these sub-intervals is characterised by different results. The purpose of this analysis on different scales is to test the dependence of the results on the granularity of the data, since we expect different behaviours at different scales for financial time series. 4.2.1Step size decision The first analysis was done on the step size, that is, the number of data points used to “slide” the window. To illustrate this, we consider, for instance, Figure 8where are shown the Distance Correlation window values versus the window step size for the PSI-20 stocks BES and BPI. These results serve only, at this stage, for comparison terms. Each point represents the Distance Correlation value in the centre of a sliding window, moved along the series. We can see for all the calculated steps (5, 10 and 20), that the Distance Correlation values remain essentially the same. So, this is not a distinguishable criterion to have into account. 60 portuguese standard index (psi-20)analysis Eventually, the more readable value is for the 20 steps case. 2002 2004 2006 2008 2010 2012 2014 0.2 0.4 0.6 0.8 time dcor.BESBPI (a) Step_5 2002 2004 2006 2008 2010 2012 2014 0.2 0.4 0.6 0.8 time dcor.BESBPI (b) Step_10 2002 2004 2006 2008 2010 2012 2014 0.2 0.4 0.6 0.8 time dcor.BESBPI (c) Step_20 Figure 8: Distance Correlation values for different steps 4.2.2Window size decision The other studied criterion is the window size. Does the results, in general, remain the same despite the size of the window? Taking into account the recommendation by Fenn et al. [2011], the size should be Q∼O(1)that is to say T=12. On the other side, we are talking about companies, so, T=60 represent approximately 3 months of data, and this is a relevant period with almost all the companies presenting quarterly reports. Example 1 In Figure 9it is possible to compare the effect of having two different size sliding windows. The 20 days window gives higher Distance Correlation values but it is harder to read than the 60 day one. It is notable that the Distance Correlation value goes down as the window size goes bigger (see Figure 9). Are we loosing relevant information by choosing one or another size? A possible answer can be pointed later when we will try to identify the events corresponding to peaks or valleys. Example 2 For another example (see Figure 10), the same happens if we consider the World Markets set. It can be seen for the different sliding windows that the Distance Correlation values 4.2 dynamic analysis of psi-20 using sliding windows 61 2002 2004 2006 2008 2010 2012 2014 0.2 0.4 0.6 0.8 time dcor.BESBPI (a) Size 20 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 time dcor.BESEDP (b) Size 60 Figure 9: DCor values for different “sliding” windows size between AEX and ASX suffer significantly as the window size gets bigger. Eventually, the more readable values are for the 120 sliding window, but for this case the Distance Correlation is more smoother and weaker than the previous sizes. 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 time dcor.AEX_ASX (a) Size 20 2002 2004 2006 2008 2010 2012 0.2 0.3 0.4 0.5 0.6 time dcor.AEX_ASX (b) Size 60 2002 2004 2006 2008 2010 2012 0.1 0.2 0.3 0.4 0.5 time dcor.AEX_ASX (c) Size 120 Figure 10: Markets DCor values for different “sliding” windows size Despite that, for instance, it is easier to understand what happens to the correlation between these two markets. We can, roughly, define three typical behaviours for this relationship: the first, corresponding to periods of world crisis, between 2000 and mid 2001 and between nearly 2007 and 2008, where the correlation goes up; the second, corresponding to non-crisis periods, between mid 2001 and late 2006 and between 2008 and 68 portuguese standard index (psi-20)analysis −0.0015 0.0010 −0.0015 0.0010 ForeC1 ForeC2 1 2 3 45 6 7 8 910 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 3132 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428429 430 431 432 433 434 435 436 437 438439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 14571458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 18251826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 18611862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 18781879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 19141915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 19511952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 19661967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 19931994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 20082009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 20692070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 22422243 2244 2245 2246 2247 2248 2249 2250 2251 2252 2253 2254 2255 2256 2257 2258 2259 2260 2261 2262 2263 2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 2312 2313 2314 2315 2316 2317 2318 2319 2320 2321 2322 2323 2324 2325 2326 2327 2328 2329 2330 2331 2332 2333 2334 2335 2336 2337 2338 2339 2340 2341 2342 2343 2344 2345 2346 2347 2348 2349 2350 2351 2352 2353 2354 2355 2356 2357 2358 2359 2360 2361 2362 2363 2364 2365 2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 2376 2377 2378 2379 2380 2381 2382 2383 2384 2385 2386 2387 2388 2389 2390 2391 2392 2393 2394 2395 2396 2397 23982399 2400 2401 2402 2403 2404 2405 2406 2407 2408 2409 2410 2411 2412 2413 2414 2415 2416 2417 2418 2419 2420 2421 2422 2423 2424 2425 2426 2427 2428 2429 2430 2431 2432 2433 2434 2435 2436 2437 2438 2439 2440 2441 2442 2443 2444 2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 2465 2466 2467 2468 2469 2470 2471 2472 2473 2474 2475 2476 2477 2478 2479 2480 2481 2482 2483 2484 2485 2486 24872488 2489 2490 2491 2492 2493 2494 2495 2496 2497 2498 2499 2500 2501 2502 2503 2504 2505 2506 2507 2508 2509 2510 2511 2512 2513 2514 2515 2516 2517 2518 2519 2520 2521 2522 2523 2524 2525 2526 2527 2528 2529 2530 2531 2532 2533 2534 2535 2536 2537 2538 2539 2540 2541 2542 2543 2544 2545 2546 2547 2548 2549 2550 2551 2552 2553 2554 2555 2556 2557 2558 2559 2560 2561 2562 2563 2564 2565 2566 2567 2568 2569 2570 2571 2572 2573 2574 2575 2576 2577 2578 2579 2580 2581 2582 2583 2584 2585 2586 2587 2588 2589 2590 2591 2592 2593 2594 2595 2596 2597 2598 2599 2600 2601 2602 2603 2604 2605 2606 2607 2608 2609 2610 2611 2612 2613 2614 2615 2616 2617 2618 2619 2620 2621 2622 2623 2624 2625 2626 2627 2628 2629 2630 2631 2632 2633 2634 2635 2636 2637 2638 2639 2640 2641 2642 2643 2644 2645 2646 2647 2648 2649 2650 2651 2652 2653 2654 26552656 2657 2658 2659 2660 2661 2662 2663 2664 2665 2666 2667 2668 2669 2670 2671 2672 2673 2674 2675 2676 2677 2678 2679 2680 2681 2682 2683 2684 2685 2686 2687 2688 2689 2690 2691 2692 2693 2694 2695 2696 2697 2698 2699 2700 2701 2702 2703 2704 2705 2706 2707 2708 2709 2710 2711 2712 2713 2714 2715 2716 2717 2718 2719 2720 2721 2722 2723 2724 2725 2726 2727 2728 2729 2730 2731 2732 2733 2734 2735 2736 2737 2738 2739 2740 2741 2742 2743 2744 2745 2746 2747 2748 2749 2750 2751 2752 2753 2754 2755 2756 2757 2758 2759 2760 2761 2762 2763 2764 2765 2766 2767 2768 2769 2770 2771 2772 2773 2774 2775 2776 2777 2778 2779 2780 2781 2782 2783 2784 2785 2786 2787 2788 2789 2790 2791 2792 2793 27942795 2796 2797 2798 2799 2800 2801 2802 28032804 2805 2806 2807 2808 2809 2810 2811 2812 2813 2814 2815 2816 2817 2818 2819 2820 2821 2822 2823 2824 2825 2826 2827 2828 2829 2830 2831 2832 2833 2834 2835 2836 2837 2838 2839 2840 2841 2842 28432844 2845 2846 2847 2848 2849 2850 2851 2852 2853 2854 2855 2856 2857 2858 2859 2860 2861 2862 2863 2864 2865 2866 2867 2868 2869 2870 2871 2872 2873 2874 2875 2876 2877 2878 2879 2880 2881 2882 2883 2884 2885 2886 2887 2888 2889 2890 2891 2892 2893 2894 2895 2896 2897 2898 2899 2900 2901 2902 2903 2904 2905 2906 2907 2908 2909 2910 2911 2912 2913 2914 2915 2916 2917 2918 2919 2920 29212922 2923 2924 2925 2926 2927 2928 2929 2930 2931 2932 2933 2934 2935 2936 2937 2938 2939 2940 2941 2942 2943 2944 2945 2946 2947 2948 2949 2950 2951 2952 2953 2954 2955 2956 2957 2958 2959 2960 2961 2962 2963 2964 2965 29662967 29682969 2970 2971 2972 2973 2974 2975 2976 2977 2978 2979 2980 2981 2982 2983 2984 2985 2986 2987 2988 2989 2990 2991 2992 2993 2994 2995 2996 2997 2998 2999 3000 3001 3002 3003 3004 3005 3006 3007 3008 3009 3010 3011 3012 3013 3014 3015 3016 3017 3018 3019 3020 3021 3022 3023 3024 3025 3026 3027 3028 3029 30303031 3032 3033 3034 3035 3036 3037 3038 3039 3040 3041 3042 3043 3044 3045 3046 3047 3048 3049 3050 3051 3052 3053 3054 3055 3056 3057 3058 3059 3060 3061 3062 3063 3064 3065 3066 3067 3068 3069 3070 3071 3072 3073 3074 3075 3076 3077 3078 3079 3080 3081 3082 3083 3084 3085 3086 3087 3088 3089 30903091 3092 3093 3094 3095 3096 3097 30983099 3100 3101 3102 3103 3104 31053106 3107 3108 3109 3110 3111 3112 3113 3114 3115 3116 3117 3118 3119 3120 3121 3122 3123 3124 3125 3126 3127 3128 3129 3130 3131 3132 3133 3134 3135 3136 3137 3138 3139 3140 3141 3142 3143 3144 3145 3146 3147 3148 3149 3150 3151 3152 3153 3154 3155 3156 3157 3158 3159 3160 3161 3162 3163 3164 3165 3166 3167 3168 3169 3170 3171 3172 3173 3174 3175 3176 3177 3178 3179 3180 3181 3182 3183 3184 3185 3186 3187 3188 3189 3190 3191 3192 3193 3194 3195 3196 3197 3198 3199 3200 3201 3202 3203 3204 3205 3206 3207 3208 3209 3210 3211 3212 3213 3214 3215 3216 3217 3218 −40 0 40 −40 0 40 Series 1 Series 2 Series 3 Series 4 Series 5 Series 6 Series 7 Series 8 Series 9 Series 10 Series 11 Series 12 ForeC1 Forecastability Ω ^(xt) (in %) 0.0 1.0 2.0 Series 1 Series 10 Forecastability Ω ^(xt) (in %) 0.0 1.0 2.0 ForeC1 0.0 0.2 0.4 p−value (H0: white noise) 1 white noise Series 1 Series 10 0.00 0.15 0.30 p−value (H0: white noise) 2 white noise Figure 16: ForeCA stocks global results 4.3 results 69 4.3.3Entropy 4.3.3.1Mutual Information The Mutual Information between the stocks set was calculated using an R library called “entropy”. We got abnormal values, the peaks, during 2001 and during 2008-2009, which corresponds to the first and second recession periods although the first recession period is not so notorious in the BES-BPI case (see Figure 17). 2002 2004 2006 2008 2010 2012 2014 0.0000 0.0005 0.0010 0.0015 time MI.BESBPI BES_BPI Mutual Information (a) MI for BES_BPI 2002 2004 2006 2008 2010 2012 2014 0.0000 0.0010 time MI.EDPZON EDP_ZON Mutual Information (b) MI for EDP_ZON 2002 2004 2006 2008 2010 2012 2014 0.0000 0.0010 0.0020 time MI.JMTSON JMT_SON Mutual Information (c) MI for JMT_SON 2002 2004 2006 2008 2010 2012 2014 0.0000 0.0010 time MI.PTCZON PTC_ZON Mutual Information (d) MI for PTC_ZON Figure 17:MI for PSI-20 stock pairs Also, it is interesting to see that in the BES-BPI case we can find a peak in the first quarter of 2006, related to the aborted take-over attempt by Banco Comercial Português over BPI, and that from the second recession period until now there are some peaks due, probably, to the fact that this second recession became a financial system crisis bringing turbulence over financial institutions. In the EDP-ZON and PTC-ZON cases there is a common peak in the first quarter of 2003 that we attribute to the split of PT Multimedia (now known by ZON) from PT. For the comparative periods proposed in Chapter 3, namely 2004 and from 2011 until 2013, there are no interesting peaks, apart from the one reported before for the BES-BPI case. 4.3.3.2Kullback-Leibler divergence The Kullback-Leibler divergence for the stocks set was calculated using an R library called “entropy” and are shown in Figure 18. 70 portuguese standard index (psi-20)analysis 2002 2004 2006 2008 2010 2012 2014 0.000 0.002 0.004 0.006 time KL.BESBPI BES−BPI KL_Divergence (a) KLDiv for BES_BPI 2002 2004 2006 2008 2010 2012 2014 0.000 0.004 time KL.EDPZON EDP−ZON KL_Divergence (b) KLDiv for EDP_ZON 2002 2004 2006 2008 2010 2012 2014 0.000 0.004 0.008 time KL.JMTSON JMT−SON KL_Divergence (c) KLDiv for JMT_SON 2002 2004 2006 2008 2010 2012 2014 0.000 0.002 0.004 0.006 time KL.PTCZON PTC−ZON KL_Divergence (d) KLDiv for PTC_ZON Figure 18:KLDiv for PSI-20 stock pairs The results are almost the same as the ones obtained for the Mutual Information. This is probably due to the fact that these two measures are very similar. So, the conclusions extracted for the Mutual Information technique can be adopted to the Kulback-Leibler divergence technique conclusions. 4.3.3.3Approximate Entropy Approximate Entropy (ApEn) was proposed and is being used as a measure of systems complexity. In this way, ApEn is a “regularity statistic” that quantifies the unpredictability of fluctuations in a time series. Intuitively, then, the presence of repetitive patterns of fluctuation in a time series should render it more predictable than a time series in which such patterns are absent. ApEn value reflects the likelihood that “similar” patterns of observations will not be followed by additional “similar” observations. A time series containing many repetitive patterns has a relatively small ApEn; a less predictable time series has a higher entropy value. Our results suggests that the stock time series are highly unpredictable with significant ApEn values variations during time as we can see in Figure 19. The results are very irregular, nevertheless we can infer, by inspection, two distinct periods: one, from 2000 to 2008, with higher ApEn variations and another, more calm, from 2009 to present. Obviously, no rule dominates alone, so we can observe a very interesting exception with PTC, being the lower ApEn variations from 2000 to 2006. 4.3 results 71 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 0.8 time ApEn_semapa (a) ApEn for SEM 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 0.8 time ApEn_edp (b) ApEn for EDP 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 0.8 time ApEn_jeronimomartins (c) ApEn for JMT 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 0.8 time ApEn_portugaltelecom (d) ApEn for PTC Figure 19:ApEn for PSI-20 stocks A closer look, using the recession periods, tells us that the ApEn has an atypical behaviour tendency, diminishing as the period goes through. The exceptions are in the first recession period for EDP and PTC (Figure 19). 4.3.4Distance Correlation Here are presented the results obtained with Distance Correlation. In a general way, for most of the observed correlations the most striking fact seems so evident that we can propose a division between a relatively stable period from 2000 to 2007, with the maximum correlation values being well under the correlation values present in a quite unstable period from 2007 until present (see Figure 20). The exception is Novabase (NBA) as we can see from Figure 21. One possible reason to this behaviour may be the fact that NBA was not a full-time PSI-20 stock between 2000 and 2014. This division suggests by one hand that the magnitudes of the two recessions are quite distinct and that the time series are now much more correlated. This means that an important event will spread easily. In the recession periods we see the Distance Correlation values going down with time. showing the same tendency already observed in Approximate Entropy. For a complete “catalogue” of results on PSI-20 please refer to the Appendix B. 72 portuguese standard index (psi-20)analysis 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 0.8 time dcor.BESEGL (a) Distance Correlation pair BES-EGL 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 time dcor.BESSEM (b) Distance Correlation pair BES-SEM 2002 2004 2006 2008 2010 2012 0.1 0.3 0.5 0.7 time dcor.EGLSON (c) Distance Correlation pair EGL-SON 2002 2004 2006 2008 2010 2012 0.2 0.3 0.4 0.5 0.6 time dcor.PTIZON (d) Distance Correlation pair PTI-ZON Figure 20:DCov for PSI-20 stock pairs 2002 2004 2006 2008 2010 2012 0.20 0.30 0.40 time dcor.JMTNBA (a) Distance Correlation pair JMT-NBA 2002 2004 2006 2008 2010 2012 0.2 0.3 0.4 0.5 0.6 time dcor.NBAZON (b) Distance Correlation pair NBA-ZON 2002 2004 2006 2008 2010 2012 0.2 0.3 0.4 0.5 0.6 time dcor.NBAPTI (c) Distance Correlation pair NBA-PTI 2002 2004 2006 2008 2010 2012 0.2 0.3 0.4 0.5 time dcor.NBAPTC (d) Distance Correlation pair NBA-PTC Figure 21:DCov for PSI-20 stock pairs 4.3 results 73 4.3.5Hurst Exponent Here we present some results on PSI-20 data set for Hurst exponent calculated using detrended fluctuation analysis (DFA). But, first of all, for the robustness and liability of the results let us show the fluctuation function (Figure 22) obtained for the PSI-20 index. The linear fit over all windows from all scales (see explanation in Section 2.7) gives a Pearson correlation coefficient of 0.998 and a standard-deviation (assuming the errors normally distributed) of 0.004 taken for the log-log results. Hurst exponent is obtained by fitting a power law to the DFA function <F(t)> computed in the sliding window. Pearson Correlation coefficients are computed for the fit in each case. 0.01 0.1 1 10 100 1000 scale Fluctuation function Linear best fit Figure 22: PSI-20 fluctuation function Let us now consider in Figure 23 some Hurst exponent calculations for some PSI-20 stocks. Their values are, typically, around 0.5 and 0.7 meaning that there is a small long memory process present in these stocks. The correlation coefficient r(t)is also plotted for each point revealing the quality of the fit where the Hexponent is evaluated; in all graphics the correlation coefficient is near 1. All correlation coefficients, r(t), may be seen to fall in the range 0.95 −1, giving us confidence in the power law behaviour of <F(t)>. Of interest are the observed “abrupt valleys” in all four plots, namely the ones that are common for BES, BPI and PTC in the beginning of 2006. These, and all the other present “abrupt valleys” should have a event related meaning. For a global Hurst exponent for the stocks we can view Table 10. It is noticeable that half of the Hurst exponents, H, are under or above 0.5, meaning that there is some diversity in stocks maturity and in independence from past results. EDP is the best example of a stock that does not follow trends, that is, have “anti-persistence” behaviour. Others examples could be SEM or even PTC, PTI and SON, all corresponding to classical business sectors. On the other hand we see NBA and SONC having the most “persistent” behaviour. These stocks correspond to technological companies, that is, belonging to a more “turbulent” business sector. The same can be said about BES and BPI, from the financial sector, another “turbulent” business sector. 74 portuguese standard index (psi-20)analysis 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 2000 2002 2004 2006 2008 2010 2012 2014 time (years) BES Evolution - Hurst exponent (window size 120) H(t) r(t) (a) Hurst exponent for BES 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 2000 2002 2004 2006 2008 2010 2012 2014 time (years) BPI Evolution - Hurst exponent (window size 120) H(t) r(t) (b) Hurst exponent for BPI 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 2000 2002 2004 2006 2008 2010 2012 2014 time (years) PORTUGALTELECOM Evolution - Hurst exponent (window size 120) H(t) r(t) (c) Hurst exponent for PTC 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 2000 2002 2004 2006 2008 2010 2012 2014 time (years) SONAEC Evolution - Hurst exponent (window size 120) H(t) r(t) (d) Hurst exponent for SONC Figure 23: Hurst exponent for PSI-20 stocks Stock H R σH ^BES 0.525 0.998 0.00443 ^BPI 0.53 0.999 0.00302 ^EDP 0.392 0.975 0.0121 ^JMT 0.505 0.999 0.00309 ^EGL 0.495 0.999 0.00341 ^NBA 0.567 0.998 0.0053 ^PTI 0.472 0.991 0.00839 ^PTC 0.462 0.997 0.00454 ^SEM 0.437 0.992 0.00727 ^SONC 0.559 0.999 0.00307 ^SON 0.473 0.996 0.00581 ^ZON 0.501 0.998 0.00469 Table 10: Hurst exponent for PSI-20 stocks 4.4 concluding remarks 75 4.4 concluding remarks In this chapter some results found in literature were confirmed, namely the ones from random matrix theory and the ones for Hurst exponent. For Mutual Information or Kullback-Leibler Divergence the results are very sharp and a event related comparison was applied to find out the coincidences. This analysis has shown that we can match the more interesting values obtained with real events. To our knowledge, it is the first time that energy statistics is applied to the PSI-20 data. It is interesting to note that this measure proposes two well defined behaviour for the PSI-20 stocks. One period, from 2000 to 2007, relatively calm, with low variation of Distance Correlation between stocks, and another period, from 2007 till now, much more agitated in what concerns this measure. Nevertheless, besides the proposal that the stocks are much more correlated in this period, and that this happen because of the global recession, it is only possible to suggest that the Distance Correlation values tend to diminish after the most important event take place. Distance Correlation proposal is complemented by Approximate Entropy. Also, this measure, proposes these two well defined periods. When, in periods of crisis, ApEn becomes agitated with higher variations but also diminishing with time. 5 WORLD MARKETS ANALYSIS “I compare her (Fortune) to one of those raging rivers, which when in flood overflows the plains, sweeping away trees and buildings, bearing away the soil from place to place; everything flies before it, all yield to its violence, without being able in any way to withstand it; and yet, though its nature be such, it does not follow therefore that men, when the weather becomes fair, shall not make provision, both with defences and barriers, in such a manner that, rising again, the waters may pass away by canal, and their force be neither so unrestrained nor so dangerous. So it happens with fortune, who shows her power where valour has not prepared to resist her, and thither she turns her forces where she knows that barriers and defences have not been raised to constrain her.” Niccolò Machiavelli, The Prince , Chapter XXV 5.1 introduction In this chapter we will apply the mathematical tools presented in the Chapter 2to the World Markets set. The data used in this study was taken from a set of worldwide market indices, enumerated in Chapter 3, and are constituted by the daily close values for the respective indices. As it is usual in this kind of analysis, the results come from the analysis of the returns ηi=log xi xi−1. In Appendix Awe can observe the returns for all the 23 markets. Looking at the returns helps us to look only to relative variation and not to absolute values. In fact, these markets are quite different in absolute values, as it can be seen. 5.2 results Applying the techniques from Chapter 2we reach a set of results that we will show and interpret in this Section. 5.2.1Random Matrix For this set we consider 2965/5 =589 samples by sequentially sliding a window of T=20 days by 5 days (roughly one month calculated week by week). For each period, we look at the empirical correlation matrix of the N=23 markets during that period. The quality factor is therefore Q=T/N=20/23 =0.87. We started by comparing the real eigenvalues density with the theoretical one as proposed by Marchenko and Pastur [1967] (see Figure 24). 77 84 world markets analysis Kullback-Leibler divergence The Kullback-Leibler divergence for the World markets set was calculated using an R library called “entropy” and are shown in Figure 31. 2002 2004 2006 2008 2010 2012 2014 0.0000 0.0015 0.0030 time KL.AEXPSI AEX_PSI KL Divergence (a) KLDiv for AEX_PSI 2002 2004 2006 2008 2010 2012 2014 0.0000 0.0010 time KL.AEXPSI CAC_DAX KL_Divergence (b) KLDiv for CAC_DAX 2002 2004 2006 2008 2010 2012 2014 0.000 0.002 0.004 time KL.AEXPSI DJI_IXIC KL_Divergence (c) KLDiv for DJI_IXIC 2002 2004 2006 2008 2010 2012 2014 0.000 0.005 0.010 0.015 time KL.STOXXSTRAITS (d) KLDiv for STOXX_STRAITS Figure 31:KLDiv for World markets pairs The results are almost the same as the ones obtained for the Mutual Information. This is probably due to the fact that these two measures are very similar. So, the conclusions extracted for the Mutual Information technique can be adopted to the Kulback-Leibler divergence technique conclusions. Approximate Entropy Here are presented the results obtained with Approximate Entropy for World Markets set. To analyse possible regional patterns we dedicated some attention to European region dividing the results in European markets and non-European markets. Our results suggests that all the time series seem highly unpredictable with significant ApEn values variations during time as we can see in Figure 32 and Figure 33. Despite this unpredictability ApEn seems to peak at the beginning of recession periods and then goes down with time, although this is more notorious in the second one. 5.2 results 85 2002 2004 2006 2008 2010 2012 0.6 0.7 0.8 0.9 1.0 time ApEn_CAC (a) ApEn for CAC 2002 2004 2006 2008 2010 2012 0.65 0.75 0.85 0.95 time ApEn_IBEX (b) ApEn for IBEX 2002 2004 2006 2008 2010 2012 0.6 0.8 1.0 time ApEn_PSI (c) ApEn for PSI-20 2002 2004 2006 2008 2010 2012 0.7 0.8 0.9 1.0 time ApEn_SSMI (d) ApEn for SSMI Figure 32: Approximate Entropy for European markets 2002 2004 2006 2008 2010 2012 0.7 0.8 0.9 1.0 1.1 time ApEn_ASX (a) ApEn for ASX 2002 2004 2006 2008 2010 2012 0.70 0.80 0.90 1.00 time ApEn_BVSP (b) ApEn for BVSP 2002 2004 2006 2008 2010 2012 0.7 0.8 0.9 1.0 1.1 time ApEn_DJI (c) ApEn for DJI 2002 2004 2006 2008 2010 2012 0.6 0.7 0.8 0.9 1.0 time ApEn_IXIC (d) ApEn for IXIC Figure 33: Approximate Entropy for non-European markets 86 world markets analysis 5.2.4Distance Correlation Here are presented some of the results obtained for Distance Correlation. For a complete “catalogue” of results concerning PSI-20 please refer to the Appendix B. Asia-Pacific Markets ASX For the ASX market we can observe that there is no high correlation with any other market. Almost all the correlations goes between 0.3and 0.7. As an example (see Figure 34) it is shown the correlation between ASX and HSI. 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 time dcor.ASX_HSI Figure 34: Distance Correlation for the ASX_HSI pair BSESN For this market we can only find a little different correlation relationship with the HSI market (Figure 35). The correlation goes up until 2008 and goes down from 2008 on, but does not leave the interval 0.3to 0.7, apart from some peaks reaching 0.8in 2008. For all 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.BSESN_HSI Figure 35: Distance Correlation for the BSESN_HSI pair 5.2 results 87 the other market it is not easy to find a pattern. Almost all the correlations are between 0.3and 0.7for most of the time series. HSI, JKSE and NIK For this market we can find interesting correlation relationship with the BSESN market, as commented before. Also, there are some pertinent comments on the correlation with some of the Asian markets: with NIK the correlation remains between 0.4and 0.8until 2007 (see Figure 36), but going down, and then, jumps to 0.5to 0.8and starts going down until now. The same transition in 2007 happens with other markets like JKSE but then remaining more “constant” before and after that year. For all the other markets it 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.HSINIK Figure 36: Distance Correlation for the HSI_NIK pair is not easy to find a pattern. Almost all the correlations are between 0.3-0.7. KOSPI For the KOSPI market we can find a pertinent correlation with NIK in Figure 37. The correlation remains between 0.5and 0.8until 2007, and then, jumps to 0.6to 0.9between 2007 and 2011 and, after that, starts to oscillate in a no characteristic way. 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.KOSPINIK Figure 37: Distance Correlation for the KOSPI_NIK pair 88 world markets analysis European Markets AEX For the AEX market we can observe that there is a very high correlation with the other European markets, being the PSI-20 the exception, with correlation values typically 20% under. For the AEX_ATX pair it is possible to observe (see Figure 38) an interesting behaviour. 2002 2004 2006 2008 2010 2012 0.2 0.4 0.6 0.8 time dcor.AEX_ATX Figure 38: Distance Correlation for the AEX_ATX pair (60 days window width) From 2007, corresponding to the crisis beginning, the correlation between these two markets grew from about 0.6to 0.8, clearly showing more correlation. Apart from the European country markets there is only a very high correlation between AEX and STOXX, as we can see in Figure 39. 2002 2004 2006 2008 2010 2012 2014 0.6 0.7 0.8 0.9 1.0 time dcor.AEX_STOXX Figure 39: Distance Correlation for the AEX_STOXX pair ATX As AEX we can observe a very high correlation with the other European markets (for an example, see Figure 40), although only from 2008, jumping roughly from 0.5to 0.8. In the PSI or SSMI case this jump also appears but fades quickly (see Figure 41). 5.2 results 89 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.ATX_IBEX Figure 40: Distance Correlation for the ATX_IBEX pair 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.ATX_PSI Figure 41: Distance Correlation for the ATX_PSI pair Apart form the European country set, as with AEX, there is only a very high correlation between ATX and STOXX, but, again, only beginning in 2008 (Figure 42). CAC For the CAC market we can observe a very high correlation with the other European markets, from above 0.8, being the PSI-20 the only exception, with correlations varying between 0.5and 0.8. Another interesting relationship is with STOXX (Figure 43). We can also observe correlations between 0.5and 0.8for the relations with the North American subset (DJI, IXIC and SPY) and the Latin-American subset (BVSP, MERVAL and MXX). See, as an example, CAC versus DJI (Figure 44). For the other world markets we observe correlations between 0.4and 0.8. 90 world markets analysis 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.ATX_STOXX Figure 42: Distance Correlation for the ATX_STOXX pair 2002 2004 2006 2008 2010 2012 2014 0.75 0.85 0.95 time dcor.CACSTOXX Figure 43: Distance Correlation for the CAC_STOXX pair 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.CACDJI Figure 44: Distance Correlation for the CAC_DJI pair 5.2 results 91 DAX For the DAX market we can observe a very high correlation with the other European markets, from above 0.8, being the exceptions the PSI-20, with correlations varying between 0.4and 0.8and the SSMI, with correlations between 0.7and 0.8. Another interesting relationship is with IBEX with the correlation jumping to 0.8only from 2005 but going down more recently (Figure 45). 2002 2004 2006 2008 2010 2012 2014 0.4 0.6 0.8 1.0 time dcor.DAXIBEX Figure 45: Distance Correlation for the DAX_IBEX pair We can also observe correlations between 0.4to 0.8for the relations with the North American subset (DJI, IXIC and SPY) and the Latin-American subset (BVSP, MERVAL and MXX). See, as an example, DAX versus SPY (Figure 46). For the other world markets 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.DAXSPY Figure 46: Distance Correlation for the DAX_SPY pair we observe correlations between 0.3and 0.7. FTSE For the FTSE market we can observe a very high correlation with the other European markets, from above 0.8, being the exceptions the PSI-20 as can be noted in Figure 47, with correlations varying between 0.4and 0.8(but varying in time). 92 world markets analysis 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.FTSEPSI Figure 47: Distance Correlation for the FTSE_PSI pair About FTSE and MIB, the correlation remains around 0.8until 2011, and then, going down to 0.7(see Figure 48). We observe the same interesting relationship with IBEX, as 2002 2004 2006 2008 2010 2012 2014 0.4 0.6 0.8 1.0 time dcor.FTSEMIB Figure 48: Distance Correlation for the FTSE_MIB pair happened with DAX and IBEX, with the correlation jumping to 0.8only from 2005 but then going down from 2011. We can also observe correlations between 0.3and 0.7from the year 2000 until 2007 for the relations with the Latin-American subset (BVSP, MERVAL and MXX). More recently happens that the correlation goes up for correlations values around 0.7from 2007 until 2012 and finally starting going down from 2012. See, for example the correlation with MERVAL (Figure 49). We can also observe correlations between 0.4and 0.8for the relations with the North American subset (DJI, IXIC and SPY), getting higher from 2007. For the other world markets we observe correlations between 0.3and 0.7. IBEX For IBEX we can observe a very high correlation with the other European markets, from above 0.8, but only since 2005. The exceptions are the PSI and the SSMI. The first, because the 2005 jump is not so abrupt and because the correlation (apart from peaks) never goes higher then 0.8. The later because of the jump also being not so abrupt and because the 5.2 results 93 2002 2004 2006 2008 2010 2012 2014 0.3 0.5 0.7 0.9 time dcor.FTSEMERVAL Figure 49: Distance Correlation for the FTSE_MERVAL pair correlation stays around 0.8only until 2011. From that year on the correlation starts to go down. We can also observe correlations between 0.3and 0.8for the relations with the North American subset (DJI, IXIC and SPY) and with the Latin-American subset (BVSP, MERVAL and MXX), getting higher from 2007 and lower from 2011. For the other world markets we observe correlations between 0.3and 0.7. MIB and SSMI For MIB market we can observe a very high correlation with the other European markets and in a lower grade with the North American subset. Generally, we observe a diminishing correlation from 2011, for all the world markets. The correlations for these markets are, typically, between 0.3and 0.7. We can apply to SSMI almost the same observations as we did for MIB market. PSI-20 and STOXX Nothing more relevant to say. 5.2.4.1Latin-American Markets BVSP For the BVSP market we can observe that there is a high correlation, although variable, with the other five markets from North or Latin-America. As an example we show the correlation between BVSP and MERVAL (see Figure 50). For the other seventeen world markets nothing interestingly different from the correlation variation between 0.3and 0.7can be observed. MERVAL For this market we can observe, with the other five markets from North or LatinAmerica, that there is a time varying correlation: between 0.3and 0.7, from 2000 to 2006; going up, between 0.5and 0.8, from 2006 to 2009; going up, again, between 0.7 6 CONCLUSIONS AND FUTURE WORK "Prediction is very difficult, especially about the future" - Niels Bohr “It’s too early to tell”, Zhou Enlai, Chinese premiere in the 1960s, about the impact of the French revolution In this chapter all the results obtained in Chapter 4and in Chapter 5are merged and put into perspective in order to compose a coherent line of conclusions. 6.1 conclusions In this work we have addressed the analysis of financial time series from an econophysical point of view. Financial data presents complex behaviour which needs to be decomposed effectively, that is, the breakdown of financial signals into component elements, in order to determine the nature of the fluctuations observed. This was done using a number of techniques: • random matrix theory like the Correlation matrix; • component analysis like the Forecastable Component Analysis; • entropy measures like the Mutual Information, the Kullback-Leibler divergence and the Approximate entropy; • energy statistics like the Distance Correlation; • fractional Brownian motion like the Hurst exponent. These techniques are twofold: measures of “disorder”/complexity and measures of coherence. We found that these techniques are in a sense complementary, that is, each provides a different view over the financial data studied, but they can be placed under the umbrella of Econophysics measures. If entropy is disorder, implying lack of a common trading strategy, then coherence implies cooperative, or at least common tendencies in behaviour. We use the Correlation matrix as a measure of coherence among a closely related set of stocks or markets. Coherence can be either observed between each financial time series, like in Forecastable Component Analysis, Approximate entropy or Hurst exponent, or between different financial time series like in Mutual Information, Kullback-Leibler divergence, Distance Correlation or Correlation matrix. Also, there were studied and used “sliding windows” of different sizes. The motivation and importance of this kind of analysis is the well known multi-fractal behaviour that financial data exhibits (see Lux [2004]). This was reflected in the output for 20, 60 and 120 trading days windows, that is, sensibly 1, 3 and 6 trading days (in months). A natural extension of this analysis is to consider other window sizes. 101 102 conclusions and future work The first application of the techniques was to a set of 12 stocks from the PSI-20, the Portuguese index of the 20 most liquid assets of the Portuguese Stock market. PSI-20 index main characteristics are described in Appendix A. The Portuguese case is chosen both for: a) regional relevance; b) relatively little previous study and c) its relevance as a showcase both as an emerging young/mature market and its relevance to discuss features on the techniques presented. The global results are presented in Chapter 4and Chapter 5. We started by confirming some results found in literature, namely the ones from random matrix theory and the ones for the Hurst exponent. In this case, and based in previous results, we can go further and propose that the PSI-20 is becoming more mature. Indeed, it is noticeable when comparing the results for three and eight years ago (Matos et al. [2004], Matos et al. [2006] and Gomes [2012] ). It is safe to propose that an increasing number of markets achieving or mimicking mature behaviour relatively rapidly, irrespectively of their trading capability, which suggests that windows of opportunity are narrowing for investors since the arbitrage opportunities are reduced due to more efficient markets. To our knowledge, it is the first time that energy statistics is applied to the PSI-20 data. It is interesting to note that this measure, and this is corroborated by Approximate entropy results, proposes two well defined behaviour for the PSI-20 stocks. One period, from 2000 to 2007, relatively calm, with low variation of Distance Correlation between stocks, and another period, from 2007 till now, much more agitated in what concerns this measure. In Chapter 5we have applied the above Econophysics tools to the study of the World Markets set. In this Chapter, we confirm some results found in literature, namely the ones from random matrix theory and the ones for Hurst exponent. In this case, and based in previous results, we can go further and propose that all the world markets are becoming more mature, that is to say that they are becoming more transparent. Indeed, it is noticeable when comparing with the results obtained in a previous study [Matos, 2006]. For Mutual Information or Kullback-Leibler Divergence the conclusions are similar to the ones obtained from PSI-20 stocks analysis. Indeed, there are certain events that are clearly reflected in all markets, as expected since most events are due to external causes, and thus independent of the specific market. One event where this is clearly seen is the 9/11 (September 11th, 2001) attack against the World Trade Centre towers in Manhattan, NY, corresponding to the first XXI century recession. In all the markets this is clearly seen, both in markets present here and in Appendix B, where the same type of analysis reveals the same dominant stripe appearing around September 2001 and around 2008 when the second recession of XXI century happened. It is, also, interesting to note that the results from energy statistics are not so well defined as with PSI-20 stocks. Despite that, we can find strong regional correlation for most of the markets and some, but a few, more global influence markets. There is, also, a strong connection between the North-American markets and most of the European ones. That correlation became higher since 2007. 6.2 future work 103 Distance Correlation proposal is not complemented here with Approximate Entropy like it was for the PSI-20 stocks, which is somewhat disappointing because the pattern for stocks was very well defined. In general, a trend common to most markets is the progressive correlation over time for most of the studied markets. One possible reason to this is the progressive globalisation of markets, where the arbitrage opportunities are reduced due to more efficient markets. Also, the information we got from Hurst exponent was vital to confirm that stocks and markets are getting more and more mature, that is, less autocorrelated. Would Bachelier liked this? A good overall conclusion must include the understanding that we can not discard none of these methods. All of them show merits and the complementarity between them is an objective to pursue. Distance correlation have shown to be a good complement to entropy measures like Mutual Information or Kullback-Leibler divergence. Approximate entropy, as a stand alone method, have shown potential complementarity with Distance correlation. The recession periods and in a comparative view, the chosen non-recession periods, have shown that these Econophysics tools behave quite differently in recession and nonrecession times. This is a quite hopeful sign for the times to come. 6.2 future work This work opened some new “windows” in the horizon, namely, to other variants of the techniques presented in this work that were not fully explored but have shown potential for further studies. These new “windows” are discriminated next. 1. The scale dependency can be further extended into comparing the detail levels. Instead of the whole time series, we must use the time dependent covariance matrix. 2. When studying the covariance matrix and its most significant eigenvalues, we could study the evolution of eigenvectors. This type of analysis should be useful to pick sudden jumps when the main eigenvectors changes suddenly, instead of smooth time dependency. 3. New libraries are needed for Mutual Information or Kullback-Leibler divergence calculation. Two good starting points are the R libraries “infotheo” and “FNN”. 4. Forecastable Component Analysis deserves a more profound study, that was not possible in this work. 5. Approximate entropy peaks in periods of crisis, becoming agitated and with higher variations. For the World markets set a closer look is a work in progress. 6. Finally, we have studied and used “sliding windows” of different sizes. The motivation and importance of this kind of analysis is the well known multi-fractal behaviour that financial data exhibits [Calvet and Fisher,2002]. A natural extension to this question is to consider other window and step sizes. A DATA In this Appendix we visualise and present for each stock or market studied: • Country and name of the index • Historical index values. • Historical return values. • Statistical information: Observations, Minimum and Maximum, measures of central tendency like Arithmetic Mean, Geometric Mean, Median and Quartiles, Confidence Interval (95%), dispersion measures like variance and Standard Deviation, and Skewness and Kurtosis. As previously described, all analyses deal with returns, as e.g. prices can be problematical due to currency exchanges. For each stock or market, therefore, we illustrate the original time series and the returns. The same scale is used for all plots to place comparisons in a context where they can be understood. 105 106 data a.1 psi-20 stocks BES Banco Espírito Santo (BES) Year Returns value Stock value 0 5 10 15 Close Values −0.15 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(BES returns, ci=0.95, digits=8) NA Observations 3218.00000000 NAs 0.00000000 Minimum -0.55961579 Quartile 1 -0.00659163 Median 0.00000000 Arithmetic Mean -0.00093816 Geometric Mean -0.00129075 Quartile 3 0.00548269 Maximum 0.15290767 SE Mean 0.00043587 LCL Mean (0.95) -0.00179277 UCL Mean (0.95) -0.00008355 Variance 0.00061136 Stdev 0.02472571 Skewness -5.52336083 Kurtosis 115.05597353 A.1 psi-20 stocks 107 BPI Banco Português de Investimento (BPI) Year Returns value Stock value 246 Close Values −0.1 0.0 0.1 0.2 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(BPI returns, ci=0.95, digits=8) NA Observations 3218.00000000 NAs 0.00000000 Minimum -0.11705656 Quartile 1 -0.00972062 Median 0.00000000 Arithmetic Mean -0.00044468 Geometric Mean -0.00067470 Quartile 3 0.00840047 Maximum 0.23021660 SE Mean 0.00037934 LCL Mean (0.95) -0.00118844 UCL Mean (0.95) 0.00029908 Variance 0.00046306 Stdev 0.02151874 Skewness 0.63221621 Kurtosis 8.68241189 108 data EDP Energias de Portugal (EDP) Year Returns value Stock value 2 3 4 5 Close Values −0.15 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(EDP returns, ci=0.95, digits=8) NA Observations 3218.00000000 NAs 0.00000000 Minimum -0.17788696 Quartile 1 -0.00840047 Median 0.00000000 Arithmetic Mean -0.00007049 Geometric Mean -0.00020413 Quartile 3 0.00841225 Maximum 0.12568822 SE Mean 0.00028786 LCL Mean (0.95) -0.00063490 UCL Mean (0.95) 0.00049393 Variance 0.00026666 Stdev 0.01632977 Skewness -0.09063438 Kurtosis 8.95731757 A.1 psi-20 stocks 109 EGL Mota Engil (EGL) Year Returns value Stock value 123456 Close Values −0.100.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(EGL returns, ci=0.95, digits=8) NA Observations 3218.00000000 NAs 0.00000000 Minimum -0.10500331 Quartile 1 -0.00828173 Median 0.00000000 Arithmetic Mean 0.00016655 Geometric Mean -0.00002486 Quartile 3 0.00843887 Maximum 0.18392284 SE Mean 0.00034573 LCL Mean (0.95) -0.00051131 UCL Mean (0.95) 0.00084442 Variance 0.00038464 Stdev 0.01961214 Skewness 0.46309715 Kurtosis 7.34153549 116 data SONC Sonae Com (SONC) Year Returns value Stock value 1 2 3 4 5 6 7 Close Values −0.1 0.0 0.1 0.2 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(SONC returns, ci=0.95, digits=8) NA Observations 3218.00000000 NAs 0.00000000 Minimum -0.18015000 Quartile 1 -0.01010110 Median 0.00000000 Arithmetic Mean -0.00042678 Geometric Mean -0.00067183 Quartile 3 0.00816331 Maximum 0.18571715 SE Mean 0.00039073 LCL Mean (0.95) -0.00119289 UCL Mean (0.95) 0.00033933 Variance 0.00049130 Stdev 0.02216523 Skewness 0.34516349 Kurtosis 7.86558480 A.1 psi-20 stocks 117 ZON Zon Multimédia (ZON) Year Returns value Stock value 2 4 6 8 10 12 Close Values −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(ZON returns, ci=0.95, digits=8) NA Observations 3218.00000000 NAs 0.00000000 Minimum -0.11687436 Quartile 1 -0.00847463 Median 0.00000000 Arithmetic Mean -0.00031704 Geometric Mean -0.00051066 Quartile 3 0.00809721 Maximum 0.14673408 SE Mean 0.00034725 LCL Mean (0.95) -0.00099789 UCL Mean (0.95) 0.00036382 Variance 0.00038804 Stdev 0.01969870 Skewness 0.28515419 Kurtosis 6.49151035 118 data a.2 markets AEX Netherlands (AEX Index) Year Returns value Index value 200 400 600 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(AEX returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1127 Quartile 1 -0.0086 Median 0.0002 Arithmetic Mean -0.0003 Geometric Mean -0.0004 Quartile 3 0.0084 Maximum 0.1129 SE Mean 0.0004 LCL Mean (0.95) -0.0011 UCL Mean (0.95) 0.0006 Variance 0.0004 Stdev 0.0196 Skewness 0.1986 Kurtosis 5.5145 A.2 markets 119 ASX Australia (ASX Index) Year Returns value Index value 10 20 30 40 50 60 Index −0.100.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(ASX returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1275 Quartile 1 -0.0086 Median 0.0003 Arithmetic Mean 0.0005 Geometric Mean 0.0003 Quartile 3 0.0099 Maximum 0.1775 SE Mean 0.0004 LCL Mean (0.95) -0.0004 UCL Mean (0.95) 0.0013 Variance 0.0004 Stdev 0.0200 Skewness 0.0843 Kurtosis 9.3220 120 data ATX Austria (ATX Index) Year Returns value Index value 1000 3000 5000 Index −0.1 0.0 0.1 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(ATX returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1294 Quartile 1 -0.0072 Median 0.0011 Arithmetic Mean 0.0004 Geometric Mean 0.0002 Quartile 3 0.0092 Maximum 0.1789 SE Mean 0.0004 LCL Mean (0.95) -0.0004 UCL Mean (0.95) 0.0013 Variance 0.0004 Stdev 0.0198 Skewness 0.2031 Kurtosis 13.1087 A.2 markets 121 BSESN India (BSESN Index) Year Returns value Index value 5000 15000 Index −0.1 0.0 0.1 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(BSESN returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1718 Quartile 1 -0.0084 Median 0.0012 Arithmetic Mean 0.0008 Geometric Mean 0.0006 Quartile 3 0.0103 Maximum 0.1599 SE Mean 0.0005 LCL Mean (0.95) -0.0001 UCL Mean (0.95) 0.0017 Variance 0.0004 Stdev 0.0203 Skewness -0.2492 Kurtosis 8.0857 122 data BVSP Brazil (BVSP Index) Year Returns value Index value 20000 60000 Index −0.100.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(BVSP returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1321 Quartile 1 -0.0110 Median 0.0006 Arithmetic Mean 0.0006 Geometric Mean 0.0003 Quartile 3 0.0129 Maximum 0.1687 SE Mean 0.0005 LCL Mean (0.95) -0.0004 UCL Mean (0.95) 0.0016 Variance 0.0005 Stdev 0.0233 Skewness 0.1234 Kurtosis 5.4069 A.2 markets 123 CAC France (CAC Index) Year Returns value Index value 3000 5000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(CAC returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.0961 Quartile 1 -0.0087 Median 0.0003 Arithmetic Mean -0.0002 Geometric Mean -0.0003 Quartile 3 0.0090 Maximum 0.1330 SE Mean 0.0004 LCL Mean (0.95) -0.0010 UCL Mean (0.95) 0.0007 Variance 0.0004 Stdev 0.0193 Skewness 0.2561 Kurtosis 5.3707 124 data DAX Germany (DAX Index) Year Returns value Index value 2000 6000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(DAX returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1137 Quartile 1 -0.0091 Median 0.0009 Arithmetic Mean 0.0002 Geometric Mean 0.0000 Quartile 3 0.0094 Maximum 0.1346 SE Mean 0.0004 LCL Mean (0.95) -0.0007 UCL Mean (0.95) 0.0010 Variance 0.0004 Stdev 0.0200 Skewness 0.0526 Kurtosis 4.7335 A.2 markets 125 DJI United States (DJI Index) Year Returns value Index value 8000 12000 16000 Index −0.1 0.0 0.1 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(DJI returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1592 Quartile 1 -0.0065 Median 0.0005 Arithmetic Mean 0.0002 Geometric Mean 0.0000 Quartile 3 0.0066 Maximum 0.1604 SE Mean 0.0003 LCL Mean (0.95) -0.0005 UCL Mean (0.95) 0.0008 Variance 0.0002 Stdev 0.0157 Skewness -0.0279 Kurtosis 15.0527 132 data MERVAL Argentina (MERVAL Index) Year Returns value Index value 01000 3000 5000 Index −0.2−0.1 0.0 0.1 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(MERVAL returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1959 Quartile 1 -0.0110 Median 0.0010 Arithmetic Mean 0.0012 Geometric Mean 0.0008 Quartile 3 0.0133 Maximum 0.2310 SE Mean 0.0006 LCL Mean (0.95) -0.0001 UCL Mean (0.95) 0.0024 Variance 0.0008 Stdev 0.0278 Skewness 0.0518 Kurtosis 7.3188 A.2 markets 133 MIB Italia (MIB Index) Year Returns value Index value 20000 40000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(MIB returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1291 Quartile 1 -0.0088 Median 0.0006 Arithmetic Mean -0.0004 Geometric Mean -0.0006 Quartile 3 0.0085 Maximum 0.1447 SE Mean 0.0004 LCL Mean (0.95) -0.0013 UCL Mean (0.95) 0.0004 Variance 0.0004 Stdev 0.0197 Skewness -0.0899 Kurtosis 6.3869 134 data MXX Mexico (MXX Index) Year Returns value Index value 10000 30000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(MXX returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.0966 Quartile 1 -0.0067 Median 0.0014 Arithmetic Mean 0.0010 Geometric Mean 0.0008 Quartile 3 0.0087 Maximum 0.1259 SE Mean 0.0004 LCL Mean (0.95) 0.0002 UCL Mean (0.95) 0.0017 Variance 0.0003 Stdev 0.0167 Skewness 0.1871 Kurtosis 6.9050 A.2 markets 135 NIK Japan (NIK Index) Year Returns value Index value 800012000 18000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(NIK returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1211 Quartile 1 -0.0092 Median 0.0005 Arithmetic Mean 0.0000 Geometric Mean -0.0002 Quartile 3 0.0100 Maximum 0.1367 SE Mean 0.0004 LCL Mean (0.95) -0.0008 UCL Mean (0.95) 0.0009 Variance 0.0004 Stdev 0.0197 Skewness -0.4147 Kurtosis 7.0111 136 data PSI Portugal (PSI Index) Year Returns value Index value 4000 8000 12000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(PSI returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1378 Quartile 1 -0.0063 Median 0.0007 Arithmetic Mean -0.0003 Geometric Mean -0.0004 Quartile 3 0.0063 Maximum 0.1407 SE Mean 0.0003 LCL Mean (0.95) -0.0010 UCL Mean (0.95) 0.0004 Variance 0.0002 Stdev 0.0156 Skewness -0.3625 Kurtosis 15.4539 A.2 markets 137 SPY United States (SPY Index) Year Returns value Index value 80100 140 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(SPY returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1036 Quartile 1 -0.0064 Median 0.0006 Arithmetic Mean 0.0001 Geometric Mean 0.0000 Quartile 3 0.0072 Maximum 0.1207 SE Mean 0.0004 LCL Mean (0.95) -0.0006 UCL Mean (0.95) 0.0008 Variance 0.0003 Stdev 0.0160 Skewness -0.1062 Kurtosis 7.3934 138 data SSMI Switzerland (SSMI Index) Year Returns value Index value 4000 6000 8000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns > table.Stats(SSMI returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1274 Quartile 1 -0.0069 Median 0.0004 Arithmetic Mean 0.0000 Geometric Mean -0.0001 Quartile 3 0.0075 Maximum 0.1576 SE Mean 0.0004 LCL Mean (0.95) -0.0007 UCL Mean (0.95) 0.0007 Variance 0.0003 Stdev 0.0159 Skewness 0.2232 Kurtosis 10.4162 A.2 markets 139 STOXX Europe (STOXX Index) Year Returns value Index value 200030004000 Index −0.10 0.00 0.10 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(STOXX returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.1067 Quartile 1 -0.0088 Median 0.0000 Arithmetic Mean -0.0002 Geometric Mean -0.0004 Quartile 3 0.0089 Maximum 0.1295 SE Mean 0.0004 LCL Mean (0.95) -0.0011 UCL Mean (0.95) 0.0006 Variance 0.0004 Stdev 0.0194 Skewness 0.1935 Kurtosis 4.9081 140 data STRAITS Singapore (STRAITS Index) Year Returns value Index value 2 4 6 Index −0.2 0.00.10.2 2002 2004 2006 2008 2010 2012 2014 Returns table.Stats(STRAITS returns, ci=0.95, digits=4) NA Observations 2024.0000 NAs 0.0000 Minimum -0.2600 Quartile 1 -0.0058 Median 0.0000 Arithmetic Mean 0.0004 Geometric Mean 0.0001 Quartile 3 0.0060 Maximum 0.1948 SE Mean 0.0005 LCL Mean (0.95) -0.0006 UCL Mean (0.95) 0.0014 Variance 0.0005 Stdev 0.0229 Skewness -0.6769 Kurtosis 31.5261 141