scieee AI-readable full text Open interactive document viewer

Return and risk of pairs trading using a simulation-based Bayesian procedure for predicting stable ratios of stock prices

Ardia, David,Gatarek, Lukasz T.,Hoogerheide, Lennart,van Dijk, Herman K.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Ardia, David; Gatarek, Lukasz T.; Hoogerheide, Lennart; van Dijk, Herman K. Article Return and risk of pairs trading using a simulation-based Bayesian procedure for predicting stable ratios of stock prices Econometrics Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Ardia, David; Gatarek, Lukasz T.; Hoogerheide, Lennart; van Dijk, Herman K. (2016) : Return and risk of pairs trading using a simulation-based Bayesian procedure for predicting stable ratios of stock prices, Econometrics, ISSN 2225-1146, MDPI, Basel, Vol. 4, Iss. 1, pp. 1-19, https://doi.org/10.3390/econometrics4010014 This Version is available at: https://hdl.handle.net/10419/171867 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/ econometrics Article Return and Risk of Pairs Trading Using a Simulation-Based Bayesian Procedure for Predicting Stable Ratios of Stock Prices David Ardia 1,2, Lukasz T. Gatarek 3, Lennart Hoogerheide 4and Herman K. van Dijk 4,5,* 1Institute of Financial Analysis, University of Neuchatel, Neuchatel, 2000, Switzerland; [email protected] 2Department of Finance, Insurance and Real Estate, Laval University, Quebec City, G1V 0A6, Canada 3Institute of Econometrics and Statistics, Faculty of Economics and Sociology, University of Lodz, Lodz, 90-255, Poland; [email protected] 4Department of Econometrics and Tinbergen Institute, Vrije Universiteit Amsterdam, Amsterdam, 1081 HV, The Netherlands; [email protected] 5Econometric Institute, Erasmus University Rotterdam, Rotterdam, 3062 PA, The Netherlands *Correspondence: [email protected].nl; Tel.: +31-(0)10-4088900 Academic Editors: Nalan Ba¸stürk, Francesco Ravazzolo and Roberto Casarin Received date: 3 September 2015; Accepted: 28 January 2016; Published: 10 March 2016 Abstract: We investigate the direct connection between the uncertainty related to estimated stable ratios of stock prices and risk and return of two pairs trading strategies: a conditional statistical arbitrage method and an implicit arbitrage one. A simulation-based Bayesian procedure is introduced for predicting stable stock price ratios, defined in a cointegration model. Using this class of models and the proposed inferential technique, we are able to connect estimation and model uncertainty with risk and return of stock trading. In terms of methodology, we show the effect that using an encompassing prior, which is shown to be equivalent to a Jeffreys’ prior, has under an orthogonal normalization for the selection of pairs of cointegrated stock prices and further, its effect for the estimation and prediction of the spread between cointegrated stock prices. We distinguish between models with a normal and Student tdistribution since the latter typically provides a better description of daily changes of prices on financial markets. As an empirical application, stocks are used that are ingredients of the Dow Jones Composite Average index. The results show that normalization has little effect on the selection of pairs of cointegrated stocks on the basis of Bayes factors. However, the results stress the importance of the orthogonal normalization for the estimation and prediction of the spread—the deviation from the equilibrium relationship—which leads to better results in terms of profit per capital engagement and risk than using a standard linear normalization. Keywords: Bayesian analysis; cointegration; linear normalization; orthogonal normalization; pairs trading; statistical arbitrage JEL: C11; C15; C32; C58; G17 1. Introduction In this paper we consider statistical arbitrage strategies. Such strategies presume that the patterns observed in the historical data are expected to be repeated in the future. That is, a statistical arbitrage is a purely advanced descriptive approach designed to exploit market inefficiencies. Econometrics 2016,4, 14; doi:10.3390/econometrics4010014 www.mdpi.com/journal/econometrics Econometrics 2016,4, 14 2 of 19 Khandani and Lo [1] consider a specific strategy—first proposed by Lehmann [2] and Lo and MacKinlay [3]—that can be analyzed directly using individual equities returns. Given a collection of securities, they consider a long/short market-neutral equity strategy consisting of an equal dollar amount of long and short positions, where at each rebalancing interval, the long positions are made up of “losers” (underperforming stocks, relative to some market average) and the short positions are made up of “winners” (outperforming stocks, relative to the same market average). By buying yesterday’s losers and selling yesterday’s winners at each date, such a strategy actively bets on mean reversion across all stocks, profiting from reversals that occur within the rebalancing interval. For this reason, such strategies have been called “contrarian” trading strategies that benefit from market overreaction, i.e., when underperformance is followed by positive returns and vice-versa for outperformance. The same key idea is the basis of pairs trading strategies, which constitute another form of statistical arbitrage strategies. The idea of pairs trading relies on long-term equilibrium among a pair of stocks. If such an equilibrium exists, then it is presumed that a specific linear combination of prices reverts to zero. A trading rule can be set up to exploit the temporary deviations (spread) to generate profit. When the spread between two assets is positive it is sold; that is, the outperforming stock is shorted and the long position is opened in the underperforming stock. In the opposite case, when the spread is negative: one buys. Gatev et al. [4] investigate the performance of this arbitrage rule over a period of 40 years and they find huge empirical evidence in favor of it. It is fundamental for the pairs trading strategy to precisely estimate the current and expected spread among the stock prices. In this paper we interpret spread as the temporary deviation from the equilibrium in a cointegration model. Equilibrium in a cointegration model is interpreted as time series behavior that is characterized by stable, or otherwise stated stationary, long-run relations to which actual series return after temporary deviations. This approach differs from Gatev et al. [4], who implement a nonparametric framework. These authors choose a matching partner for each stock by finding the security that minimizes the sum of squared deviations between the two normalized price series; pairs are thus formed by exhaustive matching between normalized daily prices, where price includes reinvested dividends. However, as argued above, in the cointegration analysis that we perform, the spread between two assets is modeled as the temporary deviation from the long-run stable relations among the time series of asset prices. This deviation is computed as a linear combination of stock prices, where the weights in the linear combination are given by the cointegrating vector. Long-run stability also implies that there exists a finite uncertainty in the predictibility of stock prices that can be used in devising trading strategies. Therefore, pairs trading strategies are strongly dependent on the stability of ratios of pairs of stocks. The estimated and predicted spreads are both computed from the estimated cointegration model. We introduce a simulation-based Bayesian estimation procedure that allows us to combine estimation and model uncertainty in a natural way with decision uncertainty associated with a decision process like a trading strategy. For the Bayesian estimation of the cointegration model, we work with a Metropolis-Hastings (M-H) type of sampler derived under an encompassing prior where we show that the encompassing prior is equivalent under certain conditions to the well-known Jeffreys’ or Information matrix prior. This sampling algorithm is derived by Kleibergen and Van Dijk [5] for the Simultaneous Equations Model and extended by Kleibergen and Paap [6] for the cointegration model. The latter authors specify a linear normalization to identify the parameters in the model. However, Strachan and Van Dijk [7] point at possible distortions of prior beliefs associated with the linear normalization. Moreover, in our application we find out that the distribution of the spread is particularly sensitive to the choice of normalization. Therefore we make use of an alternative normalization, the orthogonal normalization, in order to identify the parameters in the cointegration model. Given that one is usually only interested in a linear combination of price series, this normalization is a natural one since it treats the variables in the series in a symmetric way. More details are given in Section 3. Hence, we implement the M-H Econometrics 2016,4, 14 3 of 19 sampler for the cointegration model under this normalization. We compare the performance of the pairs trading strategy under the orthogonal normalization with the performance of the counterpart under the linear normalization and find that, for our set of data, the orthogonal normalization is highly favored over the linear normalization with respect to the profitability and risk of the trading strategies. The results imply that within the statistical arbitrage approach of pairs trading based on the cointegration model, the normalization is not only a useful device easing the parameter identification but it primarily becomes an important part of the model. To take into account the non-normality of the conditional distribution of daily returns, we extend our approach of using the normal distribution to the case of the Student-tdistribution. The outline of the paper is as follows. In Section 2the conditional and implicit statistical arbitrage approaches are discussed. In Section 3our Bayesian analysis of the cointegration model under the encompassing prior is explained. In Section 4we consider an empirical application using stocks in the Dow Jones Composite Average index. Section 5concludes. The appendices contain technical derivations and additional tables with detailed results from our empirical application. 2. Pairs Trading: Implicit and Conditional Statistical Arbitrage Suppose that there exists a statistical fair price relationship [8] between the prices yt,1 and yt,2 of two stocks, where the spread st=β1yt,1 +β2yt,2 (1) is the deviation from this statistical fair price relationship, or “statistical mispricing”, at the end of day t. In this paper we consider two types of trading strategies that are based upon the existence of such a long-run equilibrium relationship: conditional statistical arbitrage (CSA) and implicit statistical arbitrage (ISA), where we use the classification of Burgess [8]. We will implement these strategies in such a way that at the end of each day the holding is updated, after which the holding is kept constant for a day. In the CSA strategy the desired holding at the end of day tis given by CSA(st,k) = sign(E(∆st+1|It)) |E(∆st+1|It)|k, (2) where Itis the information set at the end of day t, and where we consider k=0 and k=1. A positive value of CSA(st,k)means that we buy CSA(st,k)spreads and a negative value of CSA(st,k)means that we short CSA(st,k)spreads. That is, if β1>0 and β2<0, then a positive value of CSA(st,k)means that we buy β1×CSA(st,k)of stock 1 and short (−β2)×CSA(st,k)of stock 2. For k=1 the obvious intuition of the CSA strategy is that we want to invest more in periods with larger expected profits. In this way we consider the accuracy of the used method. In the case of k=0 we only look at the sign of the expected change in the next day. In this way we consider the directional accuracy of the used method. Note that the expectation in (2) is taken over the distribution of ∆st+1(given the information set Itand the ‘fixed’ values of β1and β2). In the sequel of this paper, we use the posterior median to obtain estimates of model parameters, where the expectation in (2) will still be taken given these “fixed” estimated values. We use the posterior median, since the posterior distribution has Cauchy type tails in one of the model specifications that we investigate and these Cauchy type tails imply that the coefficients have no posterior means. In the ISA strategy the desired holding at the end of day tis given by: ISA(st) = −st. (3) A positive value of ISA(st)means that we buy ISA(st)spreads and a negative value of ISA(st) means that we short ISA(st)spreads. Or equivalently, a negative value of stmeans that we buy −st spreads and a positive value of stmeans that we short stspreads. That is, if β1>0 and β2<0, then a positive value of ISA(st)means that we buy β1×ISA(st) = β1×(−st)of stock 1 and short Econometrics 2016,4, 14 4 of 19 (−β2)×(−ISA(st)) = (−β2)×stof stock 2. In the sequel of this paper, we will substitute the posterior medians of β1and β2to obtain an estimate of the spread in (1). The CSA and ISA strategies raise several questions. First, how do we define such long-run equilibrium relationships? How are the coefficients β1and β2estimated? Second, how do we find pairs of stocks that satisfy such a long-run equilibrium relationship? Third, how do we estimate how the stock prices adjust towards their long-run equilibrium relationship? In the next section, we consider how our Bayesian analysis of the cointegration model (under linear or orthogonal normalization) provides answers to all these questions. In order to answer the first and third questions we use the posterior distribution (more precisely, the posterior median) of the parameters in the cointegration model. In order to answer the second question we compute the Bayes factor of a model with a cointegration relationship versus a model without a cointegration relationship for a large number of pairs of stocks. At this point, we stress why we make use of the CSA and ISA strategies, rather than the approach of Gatev et al. [4]. In the strategy of Gatev et al. [4] a holding is taken as soon as it is found that a pair of prices has substantially diverged. After that, the holding remains constant until the prices have completely converged to the equilibrium relationship. A disadvantage of that trading strategy is that there is not much trading going on (i.e., in most periods there is no trading at all), which makes it more difficult to investigate the difference in quality between different models given a finite period, or equivalently a very long period may be required to be able to find substantially credible differences in trading results between models. 3. Bayesian Analysis of the Cointegration Model Under Linear and Orthogonal Normalization Consider a vector autoregressive model of order 1 (VAR(1)) for an n-dimensional vector of time series {Yt}T t=1 Yt=ΦYt−1+εt, (4) εtis an independent n-dimensional vector normal process with zero mean and n×npositive definite symmetric (PDS) covariance matrix Σ. We will consider two alternative distributions for εt: a multivariate normal distribution and a multivariate Student’s tdistribution. Φis an n×nmatrix with with autoregressive coefficients. The initial values in Y1are assumed fixed. The VAR model in (4) can be written in error correction form ∆Yt=Π0Yt−1+εt, (5) where Π0=Φ−In(with Inthe n×nidentity matrix) is the long-run multiplier matrix, see e.g., Johansen [9] and Kleibergen and Paap [6]. If Πis a zero matrix, the series Ytcontains nunit roots and there is no opportunity for long term predictibility with finite uncertainty. If the matrix Πhas full rank, the univariate series in Ytare stationary and long-run equilibrium relations are assumed to hold. Cointegration appears if the rank Πequals rwith 0 <r<n. The matrix Π0can be written as the outer product of two full rank n×r matrices α0and β: Π0=α0β0. The matrix βcontains the cointegration vectors, which reflect the stationary long-run (equilibrium) relations between the univariate series in Yt; that is, each element of β0Ytcan be interpreted as a temporary deviation from a long-run (equilibrium) relations. The matrix αcontains the adjustment parameters, which indicate the speed of adjustment to the long-run (equilibrium) relations. To save on notation, we write (5) in matrix notation ∆Y=Y−1Π+ε Econometrics 2016,4, 14 5 of 19 with (T−1)×nmatrices ∆Y= (∆Y2, . . . , ∆YT)0,Y−1= (Y1, . . . , YT−1)0and ε= (ε2, . . . , εT). Under the cointegration restriction Π=βα, this model is given by: 1 ∆Y=Y−1βα +ε. (6) The individual parameters in βα are non-identified as βα =βBB−1αfor any nonsingular r×rmatrix B. That is, postmultiplying βby an invertible matrix Band premultiplying αby its inverse leaves the matrix βα unchanged. Therefore, r×ridentification restrictions are required to identify the elements of βand α, so that these become estimable. In this paper we will consider two different normalization restrictions for identification purposes. The first normalization is the linear normalization, which is commonly used, where we have β= Ir β2!. (7) That is, the r×relements of the first rrows must form an identity matrix. The intuition behind this normalization is that for the case of two series it is assumed that the second series has an effect on the first series that is similar to the case of the linear regression model, where on measures the effect of a right-hand side explanatory variable on a left-hand side dependent variable. The second normalization is the orthogonal normalization, where we have β0β=Ir. (8) Here the interpretation is that the two series are treated as symmetrically effective and only the linear combination matters. This normalization and interpretation comes natural for a set of time series of different, symmetrically treated prices, where one is mainly interested in stable linear combinations. In this paper we consider the case of n=2 time series (of stock prices) in Yt= (yt,1,yt,2), where the rank of Πis equal to r=1: ∆yt,1 ∆yt,2 != α1 α2!(β1yt−1,1 +β2yt−1,2) + εt,1 εt,2 !(9) with spread st=β1yt,1 +β2yt,2 (10) and E(∆yt+1,1|It) E(∆yt+1,2|It)!= α1 α2!(β1yt,1 +β2yt,2) = α1 α2!st, (11) so that E(∆st+1|It) = β1E(∆yt+1,1) + β2E(∆yt+1,2) = (α1β1+α2β2)st. (12) From (10) and (12) it is clear that our ISA trading strategy depends on β1and β2, whereas our CSA trading strategy also depends on α1and α2. Under the linear normalization we have β1=1: β= 1 β2!, (13) 1We also considered models with a constant term inside the cointegration relationship and/or drift terms in the model equation. The inclusion of such terms did not change the conclusions of our paper. Econometrics 2016,4, 14 6 of 19 whereas under the orthogonal normalization we have β0β=β2 1+β2 2=1, (14) which is (under the further identification restriction β2≥0) equivalent with β2=q1−β2 1. (15) Since the adjustment coefficients α1and α2may be close to 0, there may be substantial uncertainty about the equilibrium relationship. The linear normalization allows β2to take values in (−∞,∞), whereas the orthogonal normalization allows β1to take values in [−1, 1]and β2in [0, 1]. One may argue that the spread under the linear normalization is just a re-scaled version of the spread under the orthogonal normalization (where the spread under the linear normalization would result by dividing the spread under the orthogonal normalization by β1). However, we will consider a moving window, where the parameters will be updated every day, so that the re-scaling factor is not constant over time. Therefore, the profit/loss of the ISA strategy under the linear normalization is not just a re-scaled version of the profit/loss of the ISA strategy under the orthogonal normalization. Further, we estimate the parameters using their posterior median, where the posterior median of β2 under the linear normalization will typically differ from the ratio of the posterior medians of β2and β1 under the orthogonal normalization. The profit/loss of the strategies under the linear normalization may be much affected by a small number of days at which the β2is estimated very large (in an absolute sense), whereas under the orthogonal normalization the profit/loss may be more evenly affected by the different days, as (the estimates of) β1and β2can not ‘escape’ to extreme values outside [−1, 1]×[0, 1]. 3.1. The Encompassing and Jeffreys’ Framework for Prior Specification and Posterior Simulation As mentioned above, we consider the case of n=2 time series (of stock prices), where the rank of Π=βα is equal to r=1. That is, the matrix Πneeds to satisfy a reduced rank restriction. A natural way to specify a prior for αand βis given by the encompassing framework, in which one first specifies a prior on Πwithout imposing a reduced rank restriction and then obtains the prior in our model as the conditional prior of Πgiven that the rank of Πis equal to 1. As singular values are generalized eigenvalues of non-symmetric matrices, they are a natural way to represent the rank of a matrix. Using singular values we can artificially construct the full rank specification of Πvia an auxiliary parameter given by the (n−r)×(n−r)matrix λ;i.e.,λis a scalar in our case with n=2 and r=1. The reduced rank matrix βα is extended into the full rank specification: Π=βα +β⊥λα⊥, (16) where β⊥and α0 ⊥are n×(n−r)matrices that are specified such that β0β⊥≡0, β0 ⊥β⊥≡In−r,α⊥α0≡ 0 and α⊥α0 ⊥≡In−r. The full rank specification encompasses the reduced rank case given by λ=0. In this framework the probability p(λ=0|Y)can be interpreted as a measure quantifying the likelihood of reduced rank. The specification in (26) is obtained using the singular value decomposition Π=USV0of Π, where the n×nmatrices Uand Vare orthogonal such that U0U=In and V0V=Inand the n×nmatrix Sis diagonal and has the singular values of Πon its diagonal in a decreasing order. To derive the elements of equation (26) in terms of parameters Πwe partition Πaccording to the specifics of the chosen normalization. Under the linear normalization, we partition the matrices U,S and Vas follows U= U11 U12 U21 U22!,S= S10 0S2!, and V0= V0 11 V0 21 V0 12 V0 22!. Econometrics 2016,4, 14 7 of 19 The matrices in decomposition (26) in terms of the blocks of U,Sand Vare given by α=U11S1V0 11 V0 21,α⊥= (V22V0 22)1/2V−1 22 0V0 12 V0 22, β2=U21U−1 11 ,β⊥= U12 U22!U−1 22 (U22U0 22)1/2, λ= (U22U0 22)−1/2U22S2V0 22(V22V0 22)−1/2. Under the orthogonal normalization, the matrices are partitioned as U=U1U2,S= S10 0S2!and V0= V0 1 V0 2!, and the following relations hold: α=S1V0 1,α⊥=V0 2 β=U1,β⊥=U2, λ=S2. Under the orthogonal normalization λis directly equal to S2, whereas under the linear normalization it is just a rotation of S2. In both cases restriction λ=0 is equivalent with restricting the n−rsmallest singular values of Πto 0. The prior on (α,β)is equal to the conditional prior of the parameters (α,β,λ)given that λ=0, which is proportional to the joint prior for (α,β,λ)evaluated at λ=0: p(α,β)∝p(α,β,λ)|λ=0∝p(Π(α,β,λ))|λ=0|J(Π,(α,β,λ))||λ=0, (17) where |λ=0stands for evaluated in λ=0, where J(Π,(α,β,λ)) denotes the Jacobian of the transformation from Πto (α,β,λ). Kleibergen and Paap [6] derive the closed form expression for the determinant of the Jacobian |J(Π,(α,β,λ))|for the general case of nvariables and reduced rank runder the linear normalization. In Appendix B the Jacobian is derived under the orthogonal normalization of β. Ba¸stürk et al. [10] prove that under certain conditions the encompassing prior is equivalent to Jeffreys’ prior in the cointegration model with normally distributed innovations, irrespective of the normalization applied. We emphasize this equivalence, since the use of the information matrix or Jeffreys’ prior is more well-known than the encompassing approach. Since the information matrix prior may yield certain desirable properties of the posterior, we conclude that an encompassing approach may also serve this purpose. In a similar fashion, the posterior of (α,β)is equal to the conditional posterior of the parameters (α,β,λ)given that λ=0, which is proportional to the joint posterior for (α,β,λ)evaluated at λ=0: p(α,β|Y) = p(α,β|λ=0, Y) ∝p(α,β,λ|Y)|λ=0=p(Π(α,β,λ|Y))|λ=0|J(Π,(α,β,λ))||λ=0, (18) where the detailed expression for p(Π(α,β,λ)|Y)is given by Kleibergen and Paap [6], and where p(α,β,λ|Y) = p(Π(α,β,λ)|Y)|Π=βα+β⊥λα⊥|J(Π,(α,β,λ))|. (19) For Bayesian estimation of the cointegration model we need an algorithm to sample from the posterior density in (18). However this posterior densities does not belong to any known Econometrics 2016,4, 14 8 of 19 class of distributions, see Kleibergen and Paap [6], and as such can not be sampled directly. The idea of the Metropolis-Hastings (M-H) algorithm is to generate draws from the target density by constructing a Markov chain of which the distribution converges to the target distribution, using draws from a candidate density and an acceptance-rejection scheme. Kleibergen and Paap [6] present the M-H algorithm to sample from (18) for the cointegration model with normally distributed disturbances under the linear normalization. In this algorithm (19) is used to form a candidate density. The general outline of this sampling algorithm is presented in Appendix A. Appendix B presents the approach to evaluate the acceptance-rejection weights under the orthogonal normalization. The posteriors of the coefficients under the linear normalization have Cauchy type tails, so that there exist no posterior means for the coefficients. Therefore, we estimate the coefficients using the posterior median (which we do under both normalizations to keep the comparison between the normalizations as fair as possible). Given that the time series considered have a non-normal shape, we also consider the model under a multivariate Student’s tdistribution for the innovations εt. Then the M-H algorithms are straightforwardly extended, see Geweke [11]. Since we make use of the independence-chain Metropolis-Hastings algorithm, the simulation of candidate draws and the evaluation of the importance weights (to be used in the probability of accepting the candidate draw) can be easily performed in a parallel fashion. This would enormously increase the speed of our computations. Only the final step of the method, the actual acceptance or rejection of candidate draws, can not be performed in a parallel fashion. But this step takes relatively very little computing time. As an alternative, one can make use of importance sampling, where the whole method can be performed in a parallel fashion. 3.2. Bayes Factors We evaluate the Bayes factor of rank 1 versus rank 2 and the Bayes factor of rank 0 versus rank 2. The Bayes factor of rank 1 versus rank 0 is obviously given by the ratio of these Bayes factors. For the evaluation of these Bayes factors we extend the method of Kleibergen and Paap [6] who evaluate the Bayes factor as the Savage-Dickey density ratio, see Dickey [12] and Verdinelli and Wasserman [13] to the case of orthogonal normalization. The Bayes factor for the restricted model with λ=0 (where Πhas rank 0 or 1) versus the unrestricted model with unrestricted λ(where Πhas rank 2) equals the ratio of the marginal posterior density of λ, and the marginal prior density of λ, both evaluated in λ=0. However, in the case of our diffuse prior specification this Bayes factor for rank reduction is not defined, as the marginal prior density of λis improper. Therefore, we follow Chao and Phillips [14] who use as prior height (2π)−(2nr−r2)/2 to construct their posterior information criterium (PIC). We assume equal prior probabilities 1 3for the rank 0, 1 or 2, so that the Bayes factor is equal to the posterior odds, the ratio of posterior model probabilities. For pairs of stock prices we will mostly observe that the estimated posterior model probability is highest for rank 0, the case of two random walk processes without cointegration. Only for a small fraction of pairs, we will observed that the estimated posterior model probability is highest for rank 1, the case of two cointegrated random walk processes. 4. Empirical Application The CSA and ISA strategies are applied to components of the Dow Jones Composite Average index. We work with daily closing prices recorded over the period of one year, from 1 January 2009 until 31 December 2009. We consider the 65 stocks with the highest liquidity. First, we identify cointegrated pairs based on the estimated posterior probability of cointegration (i.e.,Πhaving rank 1) computed for the first half year of the data. That is, among the 65×64 2=2080 pairs we select the 10 pairs with the highest Bayes factor of rank 1 versus rank 0 (where these Bayes factors are larger than 1) for both the linear and orthogonal normalization. The 10 pairs are identical for both normalizations; these pairs are given by Table 1. Second, those pairs are used in the CSA and ISA Econometrics 2016,4, 14 15 of 19 ∂vecΠ ∂(vecα)0= (Ip⊗β) + ∂vecΠ ∂(vecα⊥)0 ∂vecα⊥ ∂(vecα)0(38) and ∂vecΠ ∂(vecα⊥)0= (I⊗β⊥λ). (39) If we assume c=Ir00and c⊥=0Ip−r, we have α⊥=c0 ⊥Ir−α0(αc)0−1c0and ∂vecα⊥ ∂(vecα)0=c(αc)−1⊗(c(αc)−1αc⊥)0−c(αc)−1⊗c0 ⊥Kr,p, (40) so that ∂vecΠ ∂(vecα)0= (Ip⊗β) + c(αc)−1⊗β⊥λc0 ⊥(c(αc)−1α)0−IpKr,p. (41) Then for β2we obtain ∂vecΠ ∂(vecβ2)0=∂vecΠ ∂(vecβ)0 ∂vecβ ∂(vecβ1)0 ∂vecβ1 ∂(vecβ2)0+∂vecΠ ∂(vecβ⊥)0 vecβ⊥ ∂(vecβ)0 ∂vecβ ∂(vecβ1)0 ∂vecβ1 ∂(vecβ2)0= (α0⊗In)"Ir⊗ Ir 0(n−r)×r!# ∂vecβ1 ∂(vecβ2)0+ (α0 ⊥λ0⊗In)vecβ⊥ ∂(vecβ)0"Ir⊗ Ir 0(n−r)×r!# ∂vecβ1 ∂(vecβ2)0 The formula for ∂vecβ1 ∂(vecβ2)0is derived based on the orthogonal normalization condition. We have: Ir=β0 1β1+β0 2β2 0=d(β0 1β1) + d(β0 2β2) 0= (β1⊗Ir)dvecβ0 1+ (Ir⊗β0 1)dvecβ1 +(β2⊗Ir)dvecβ0 2+ (Ir⊗β0 2)dvecβ2 As Krvec(A) = vec(A0), see we have 0=Kr(Ir⊗β0 1)dvecβ1+ (Ir⊗β0 1)dvecβ1+Kr(Ir⊗β0 2)dvecβ2+ (Ir⊗β0 2)dvecβ2 =(Kr+Ir)(Ir⊗β0 1)dvecβ1+ (Kr+Ir)(Ir⊗β0 2)dvecβ2 =2Nr(Ir⊗β0 1)dvecβ1+2Nr(Ir⊗β0 2)dvecβ2, (42) where Nr=1 2(Ir+Kr). Thus we obtain 0=2Nr(Ir⊗β0 1)dvecβ1+2Nr(Ir⊗β0 2)dvecβ2(43) and ∂vecβ1 ∂(vecβ2)0=−Nr(Ir⊗β0 1)−1Nr(Ir⊗β0 2). (44) Further, because β=U1and β⊥=U2we can derive vecβ⊥ ∂(vecβ)0based on the orthomorphic transformation between U=U1U2and ˜ X, where ˜ X= (I+U)−1(I−U)and U= (I+˜ X)−1(I− ˜ X). We find that Econometrics 2016,4, 14 16 of 19 vecβ⊥ ∂(vecβ)0=∂vecβ⊥ ∂(vec ˜ X)0 ∂vec ˜ X ∂(vecβ)0 =− I⊗ 0 Ip−r!!(Ip+U)0⊗(Ip+˜ X)−1×"− I⊗ Ir 0!!(Ip+˜ X)0⊗(Ip+U)−1#(45) where U1=UIr00and U2=U0Ip−r0and d(I+U)−1(I−U)=−(I+U)−1dU(I+ U)−1(I−U)−(I+U)−1dU. For ∂vecΠ ∂(vecλ)0we obtain ∂vecΠ ∂(vecλ)0=α0 ⊥⊗β⊥. (46) Appendix C: Tables for Ten Pairs of Stocks Table 4. Performance evaluation measure Pro f itability in (22)–(23), which is the ratio of cumulative income to the average daily capital absorption (in %), under a normal distribution for the innovations. The meaning of the ten pairs of stocks is indicated in Table 1. Outperforming normalization in boldface. k=0 (Directional Accuracy) k=1 (Accuracy) CSA ISA CSA 0% 20% 30% 40% 50% 60% 0% 20% 30% 40% 50% 60% linear normalization of β AA-OSG 19.5 25.2 15.5 23.2 25.5 40.1 22.5 4.2 2.1 3.5 2.1 7 5.2 DUK-IBM 7.3 17 16.2 22.9 16.7 74.3 6.5 0.9 1.8 1.5 2.5 2.3 7.5 DUK-OSG −3 10.5 15.9 47.7 46.2 82.7 24.9 2.9 4.3 4.9 9.2 11.4 16.1 NI-NSC 24.9 35.9 40.4 33.5 33.9 38.8 7.1 13.2 19.1 25 25.8 33.8 37.4 NI-OSG 11.2 19.9 22.3 34.1 47.9 70.6 33.2 3.4 4.8 5.7 7.7 10 14.7 CNP-OSG 14.8 24.8 32 29.7 23 37.3 10.3 1.6 2.1 2.7 2.8 3.5 8.8 MO-UPS −5.8 −5.9 10.1 23.3 30.5 16.6 19.8 1.4 1.5 3 4.3 5.2 4.6 NI-R 39.2 32.9 22.9 14.4 20 13.4 24.1 17.8 15.2 8.3 4.7 3.8 3.3 NI-UNP 11.2 9.1 8.3 18.6 31.6 44 6.4 0.9 1.3 1.6 2.8 4.7 6.9 NI-UTX 12.9 9.5 9.4 18.3 32.1 48.6 6.5 1.1 1.5 1.9 3 4.7 7.7 orthogonal normalization of β AA-OSG 108.7 138.6 145.2 155.2 167.4 181.2 133.3 229 293.3 324.6 358.6 397.1 428.2 DUK-IBM 23.6 36.3 32.6 60.7 82.3 364 14.9 19.7 34.7 36.4 65.9 112.9 626.1 DUK-OSG 80.9 85.7 85.3 124.3 136.5 59.7 62.2 77.5 96.5 100.7 130.7 174.4 12.1 NI-NSC 46.4 54.8 59.8 81.2 81.5 80.9 17.5 32.3 44.4 56.4 82.1 95.1 132.8 NI-OSG 67.1 100.3 94.2 113.3 160.4 150.2 65.5 74 96.6 108.9 130.2 164.4 103.5 CNP-OSG 25.5 21.4 16.9 17.4 6.4 9.2 8.8 3.4 2.7 1.8 2.8 1.1 2.1 MO-UPS 0.2 4.4 2.8 15.6 32 66.7 17.9 1 1.1 1.3 2.8 4.6 7.5 NI-R 23.7 12.9 2.5 −10.2 −9.4 12.4 17.2 9.7 7.7 5.5 1 1.5 2.6 NI-UNP 14.2 7.5 1.5 16.8 30.7 21.6 5.8 21 1.1 2.7 4.7 3.8 NI-UTX 180.7 186.1 195.3 191.3 147.1 176.1 118 229.2 252.3 265.7 269.4 225.1 288.1 Econometrics 2016,4, 14 17 of 19 Table 5. Performance evaluation measure Pro f itability in (22)–(23), which is the ratio of cumulative income to the average daily capital absorption (in %), under a Student’s t-distribution for the innovations. The meaning of the ten pairs of stocks is indicated in Table 1. k=0 (Directional Accuracy) k=1 (Accuracy) CSA ISA CSA 0% 20% 30% 40% 50% 60% 0% 20% 30% 40% 50% 60% linear normalization of β AA-OSG 24.5 31.1 18.8 24.7 28.8 41.6 24.6 2.5 3.2 3 3.5 4 6.4 DUK-IBM 10.5 16.5 15.4 25.3 17.8 95.4 7 1.3 2.4 2.1 3.2 3.3 14.9 DUK-OSG 13.8 27 35.8 45.2 47.7 82.6 24.7 2.8 4.5 5.1 7.1 8 11.7 NI-NSC 41.5 53.1 73.2 71.3 67.6 79.8 17.8 34.8 48.2 70.4 87.7 93.8 106.3 NI-OSG 29 35.7 42.8 51 77 86.8 34.1 4.8 6.9 8 9.9 13.7 18 CNP-OSG 17.1 24.3 34.2 17.7 26.4 61.2 10.3 1.5 2.2 3 1.7 3.1 11.6 MO-UPS 9.9 7.9 20.8 39.8 44 48.7 20.5 2.3 2.2 3.7 5.6 6.1 5.6 NI-R 32.4 28.1 24.6 4.2 11.5 14.8 23.9 17.4 15.9 16.2 1.5 2.6 3.3 NI-UNP 4.9 8.3 7.4 18.6 31.5 50.2 6.5 0.9 1.7 1.7 2.4 4.3 8.2 NI-UTX 19.7 17.2 16.8 23 25.2 29.8 6.6 4.5 3.9 2.9 2.7 3 3.5 orthogonal normalization of β AA-OSG 118.3 150.9 170.1 182.5 192.7 213.2 141.1 245.1 314 351.8 391.9 427.4 466.3 DUK-IBM 25.9 29.7 26.6 63.1 98.4 350.3 14.9 19.8 30.4 35 65.3 97.7 608.5 DUK-OSG 78.4 84.2 96.7 115.3 148.5 101.9 60.5 73.7 90.5 102.5 121.3 133.5 16.4 NI-NSC 43.5 54.6 58.2 69.9 53 16.1 12.8 23.4 41 41.5 54.8 25.2 7.8 NI-OSG 55.5 77 74.3 101.5 152.1 119.6 54.8 53.6 69.9 80.8 114 145.1 100.2 CNP-OSG 29.7 26.9 18.9 20.8 15.4 6.6 9.7 5.8 7.3 2.5 3.3 2.9 2.9 MO-UPS 11.1 4.9 2.3 15.4 30.7 68.9 18.5 1.7 1.5 1.5 2.9 4.6 7.7 NI-R 33.4 22.1 13.8 2.6 -12.4 1.7 16.9 9.3 6.7 5 0.7 1.3 1.8 NI-UNP 16 15 9.2 35.4 40.4 48.8 6.5 2.9 3.4 3.4 6.9 7.8 11.3 NI-UTX 190.9 198.8 206.6 211.3 196.5 213.4 127.5 248.5 272.8 289.7 299.3 293.8 333.8 Table 6. Risk evaluation measure Risk in (24), fraction of trading days with decreases in cumulative income, under a normal distribution for the innovations. The meaning of the ten pairs of stocks is indicated in Table 1. Outperforming normalization in boldface. k=0 (Directional Accuracy) k=1 (Accuracy) CSA ISA CSA 0% 20% 30% 40% 50% 60% 0% 20% 30% 40% 50% 60% linear normalization of β AA-OSG 0.42 0.43 0.42 0.46 0.41 0.39 0.48 0.49 0.48 0.43 0.46 0.42 0.39 DUK-IBM 0.46 0.44 0.46 0.43 0.36 0 0.49 0.55 0.55 0.52 0.5 0.5 0 DUK-OSG 0.55 0.48 0.49 0.41 0.44 0.36 0.5 0.58 0.5 0.52 0.43 0.44 0.36 NI-NSC 0.4 0.38 0.37 0.43 0.53 0.54 0.48 0.5 0.42 0.4 0.48 0.53 0.54 NI-OSG 0.46 0.43 0.43 0.4 0.4 0.36 0.42 0.48 0.44 0.44 0.41 0.4 0.36 CNP-OSG 0.42 0.34 0.29 0.3 0.29 0.33 0.44 0.48 0.39 0.34 0.33 0.29 0.33 MO-UPS 0.52 0.52 0.47 0.42 0.39 0.42 0.5 0.58 0.58 0.53 0.46 0.43 0.46 NI-R 0.38 0.39 0.42 0.45 0.45 0.45 0.49 0.42 0.44 0.47 0.49 0.47 0.47 NI-UNP 0.41 0.45 0.43 0.45 0.42 0.33 0.44 0.46 0.49 0.43 0.45 0.42 0.33 NI-UTX 0.42 0.47 0.44 0.48 0.46 0.33 0.45 0.5 0.49 0.44 0.48 0.46 0.33 orthogonal distribution of β AA-OSG 0.45 0.42 0.42 0.4 0.39 0.37 0.4 0.51 0.48 0.49 0.45 0.45 0.4 DUK-IBM 0.46 0.42 0.46 0.4 0.36 0 0.49 0.54 0.48 0.52 0.43 0.36 0 DUK-OSG 0.4 0.44 0.44 0.38 0.42 0.36 0.4 0.49 0.52 0.51 0.45 0.48 0.36 NI-NSC 0.36 0.36 0.35 0.35 0.39 0.46 0.44 0.44 0.42 0.39 0.39 0.43 0.54 NI-OSG 0.39 0.32 0.36 0.36 0.31 0.27 0.37 0.44 0.39 0.41 0.39 0.34 0.23 CNP-OSG 0.4 0.39 0.4 0.42 0.42 0.4 0.45 0.44 0.41 0.42 0.45 0.42 0.4 MO-UPS 0.52 0.52 0.54 0.49 0.44 0.31 0.53 0.58 0.58 0.6 0.56 0.5 0.34 NI-R 0.45 0.46 0.5 0.54 0.57 0.51 0.5 0.49 0.52 0.56 0.61 0.63 0.56 NI-UNP 0.44 0.46 0.54 0.5 0.46 0.47 0.46 0.52 0.52 0.59 0.57 0.54 0.53 NI-UTX 0.22 0.25 0.22 0.23 0.26 0.29 0.41 0.35 0.34 0.32 0.32 0.35 0.38 Econometrics 2016,4, 14 18 of 19 Table 7. Risk evaluation measure Risk in (24), fraction of trading days with decreases in cumulative income, under a Student’s tdistribution for the innovations. The meaning of the ten pairs of stocks is indicated in Table 1. k=0 (Directional Accuracy) k=1 (Accuracy) CSA ISA CSA 0% 20% 30% 40% 50% 60% 0% 20% 30% 40% 50% 60% linear normalization of β AA-OSG 0.45 0.46 0.49 0.46 0.42 0.37 0.48 0.48 0.48 0.49 0.46 0.42 0.37 DUK-IBM 0.45 0.48 0.49 0.46 0.43 0 0.5 0.54 0.55 0.53 0.5 0.43 0 DUK-OSG 0.51 0.46 0.46 0.38 0.44 0.36 0.49 0.54 0.47 0.48 0.41 0.44 0.36 NI-NSC 0.37 0.32 0.27 0.39 0.41 0.4 0.48 0.45 0.38 0.36 0.43 0.47 0.47 NI-OSG 0.41 0.39 0.39 0.38 0.32 0.33 0.44 0.45 0.4 0.4 0.39 0.34 0.36 CNP-OSG 0.45 0.38 0.33 0.34 0.29 0.25 0.46 0.5 0.42 0.39 0.42 0.33 0.25 MO-UPS 0.49 0.49 0.47 0.41 0.42 0.42 0.49 0.53 0.55 0.53 0.48 0.51 0.46 NI-R 0.42 0.41 0.44 0.49 0.5 0.47 0.48 0.45 0.44 0.48 0.53 0.52 0.49 NI-UNP 0.45 0.45 0.4 0.4 0.38 0.25 0.45 0.54 0.51 0.46 0.47 0.46 0.25 NI-UTX 0.43 0.44 0.45 0.43 0.41 0.4 0.5 0.49 0.48 0.48 0.47 0.44 0.44 orthogonal normalization of β AA-OSG 0.46 0.41 0.39 0.37 0.35 0.34 0.42 0.49 0.43 0.42 0.38 0.37 0.36 DUK-IBM 0.42 0.43 0.44 0.34 0.23 0 0.45 0.51 0.52 0.52 0.38 0.23 0 DUK-OSG 0.39 0.41 0.4 0.37 0.27 0.18 0.4 0.51 0.52 0.51 0.46 0.37 0.27 NI-NSC 0.29 0.32 0.32 0.35 0.4 0.5 0.43 0.39 0.37 0.32 0.31 0.35 0.42 NI-OSG 0.44 0.4 0.38 0.39 0.32 0.38 0.42 0.5 0.45 0.44 0.43 0.35 0.43 CNP-OSG 0.37 0.38 0.35 0.38 0.36 0.4 0.44 0.43 0.42 0.38 0.38 0.36 0.4 MO-UPS 0.52 0.53 0.54 0.49 0.43 0.3 0.5 0.55 0.55 0.56 0.51 0.43 0.3 NI-R 0.42 0.45 0.49 0.53 0.57 0.55 0.5 0.5 0.5 0.51 0.56 0.59 0.55 NI-UNP 0.43 0.47 0.48 0.4 0.44 0.4 0.47 0.46 0.47 0.48 0.4 0.44 0.4 NI-UTX 0.24 0.25 0.24 0.25 0.25 0.24 0.49 0.38 0.4 0.37 0.37 0.38 0.36 References 1. Khandani, A.E.; Lo, A.W. What happened to the quants in August 2007? J. Invest. Manag. 2007,5, 29–78. 2. Lehmann, B. Fads, martingales and market efficiency. Q. J. Econ. 1990,105, 1–28. 3. Lo, A.W.; MacKinlay, A.C. When are contrarian profits due to stock market overreaction? Rev. Financ. Stud. 1990,3, 175–206. 4. Gatev, E.; Goetzmann, W.N.; Rouwenhorst, K.G. Pairs trading: Performance of a relative-value arbitrage rule. Rev. Financ. Stud. 2006,19, 797–827. 5. Kleibergen, F.R.; van Dijk, H.K. Bayesian simultaneous equation analysis using reduced rank structures. Econom. Theory 1998,14, 701–743. 6. Kleibergen, F.R.; Paap, R. Priors, posteriors and Bayes factors for a Bayesian analysis of cointegration. J. Econom. 2002,111, 223–249. 7. Strachan, R.W.; van Dijk, H.K. Valuing Structure, Model Uncertainty and Model Averaging in Vector Autoregressive Processes; Econometric Institute Report EI 2004-23; Erasmus University Rotterdam: Rotterdam, the Netherlands, 2004. 8. Burgess, A.N. A Computational Methodology for Modelling the Dynamics of Statistical Arbitrage. Ph.D. Thesis, University of London, London Business School, London, UK, 1999. 9. Johansen, S. Estimation and hypothesis testing of cointegration vectors in Gaussian vector autoregressive models. Econometrica 1991,59, 1551–1580. 10. Ba¸stürk, N.; Hoogerheide, L.F.; Kleijn, R.; van Dijk, H.K. Corresponding author: H.K. van Dijk, Department of Econometrics and Tinbergen Institute, Vrije Universiteit Amsterdam and Econometric Institute, Erasmus University Rotterdam. Prior ignorance, likelihood shape and posterior existence in a cointegration model. Unpublished working paper, 2015. 11. Geweke, J. Bayesian treatment of independent Student-tlinear model. J. Appl. Econom. 1993,8, 19–40. 12. Dickey, J. The weighted likelihood ratio, linear hypothesis on normal location parameters. Ann. Math. Stat. 1971,42, 204–223. 13. Verdinelli, I.; Wasserman, L. Computing Bayes factors using a generalization of the Savage-Dickey density ratio. J. Am. Stat. Assoc. 1995,90, 614–618. 14. Chao, J.C.; Phillips, P.C.B. Model selection in partially nonstationary vector autoregressive processes with reduced rank structure. J. Econom. 1999,91, 227–271. Econometrics 2016,4, 14 19 of 19 15. Hoogerheide, L.F.; Opschoor, A.; van Dijk, H.K. A class of adaptive importance sampling weighted EM algorithms for efficient and robust posterior and predictive simulation. J. Econom. 2012,171, 101–120. 16. West, K.D.; Edison, H.J.; Cho, D. A utility-based comparison of some models of exchange rate volatility. J. Int. Econ. 1993,35, 23–45. 17. Marquering, W.; Verbeek, M. The economic value of predicting stock index returns and volatility. J. Financ. Q. Anal. 2004,39, 407–429. 18. Furmston, T.; Hailes, S.; Morton, A.J. A Bayesian Residual-Based Test for Cointegration. 2013. Available online: http://arxiv.org/abs/1311.0524 (accessed on 21 February 2016). 19. Bracegirdle, C.; Barber, D. Bayesian Conditional Cointegration. 2012. Available online: http://arxiv.org/abs/1206.6459 (accessed on 21 February 2016). 20. Kleibergen, F.R.; van Dijk, H.K. On the shape of the likelihood/posterior in cointegration models. Econom. Theory 1994,10, 514–551. 21. Zellner, A. An Introduction to Bayesian Inference in Econometrics; Wiley: New York, NY, USA, 1971. 22. Chen, M.H. Importance-weighted marginal Bayesian posterior density estimation. J. Am. Stat. Assoc. 1994, 89, 818–824. c 2016 by the authors; licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons by Attribution (CC-BY) license (http://creativecommons.org/licenses/by/4.0/).