Comparing and selecting performance measures using rank correlations
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Caporin, Massimiliano; Lisi, Francesco Article Comparing and selecting performance measures using rank correlations Economics: The Open-Access, Open-Assessment E-Journal Provided in Cooperation with: Kiel Institute for the World Economy – Leibniz Center for Research on Global Economic Challenges Suggested Citation: Caporin, Massimiliano; Lisi, Francesco (2011) : Comparing and selecting performance measures using rank correlations, Economics: The Open-Access, Open-Assessment E-Journal, ISSN 1864-6042, Kiel Institute for the World Economy (IfW), Kiel, Vol. 5, Iss. 2011-10, pp. 1-34, https://doi.org/10.5018/economics-ejournal.ja.2011-10 This Version is available at: https://hdl.handle.net/10419/48833 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by-nc/2.0/de/deed.en
Vol. 5, 2011-10 | August 2, 2011 | http://dx.doi.org/10.5018/economics-ejournal.ja.2011-10 Vol. 5, 2011-10 | August 2, 2011 | http://dx.doi.org/10.5018/economics-ejournal.ja.2011-10 Comparing and Selecting Performance Measures Using Rank Correlations Massimiliano Caporin Department of Economics and Management "Marco Fanno", University of Padova Francesco Lisi Department of Statistical Sciences, University of Padova Abstract The financial economics literature proposes dozens of performance measures to be used, for instance, to compare, analyse, rank and select assets. There is thus a problem: which measures should be considered? We extend the current literature by comparing a large set of performance measures over more than one thousand of equities included in the Standard & Poor’s 1500 index. We evaluate performance measures by mean of rank correlations, exploiting the possible dynamic evolution of the rank correlations, and proposing a method for the identification of the subset of measures which are not equivalent. Our empirical study highlights that recent and more flexible measures provide different asset ranks compared to classical approaches, and that the set of equivalent performance measures is not stable over time. JEL C10, G11, C40 Keywords Performance measurement; rank correlations; comparing performance measures Correspondence Massimiliano Caporin, Dipartimento di Scienze Economiche "Marco Fanno", Via del Santo, 33, 35123, Padova, Italy; e-mail: [email protected] Citation Massimiliano Caporin and Francesco Lisi (2011). Comparing and Selecting Performance Measures Using Rank Correlations. Economics: The Open-Access, Open-Assessment E-Journal, Vol. 5, 2011-10. doi:10.5018/economics-ejournal.ja.2011-10. http://dx.doi.org/10.5018/economics-ejournal.ja.2011-10 © Author(s) 2011. Licensed under a Creative Commons License - Attribution-NonCommercial 2.0 Germany
conomics: The Open-Access, Open-Assessment E-Journal 1 Introduction Since the pioneering works of Sharpe (1964 and 1966) and Treynor (1965), the topic of performance measurement has attracted considerable interest in the financial economic literature. From a general viewpoint, we may identify, among others, two fundamental topics covered by performance measurement. The first considers the returns of financial assets, and aims to define and interpret ratios or indices, the performance measures or reward-to-risk ratios, for the purpose of determining the assets’ risk/return trade-off. The second analyses returns of managed portfolios and focuses on the introduction and use of models and approaches which make possible to infer the choices made by investment managers. For examples on the second topic see Knight and Satchell (2002) and the references therein, the literature on style analysis (see Sharpe, 1992, among others) and the contributions related to conditional CAPM approaches, including Ferson and Schadt (1996), Avramov and Chordia (2006). This study deals with the first issue. We focus on the comparison of performance measures based on the returns of specific assets. The approaches proposed by this strand of the literature may be considered as tools for portfolio managers and agents facing investment decisions. Performance measures are here used as tools for selecting a relatively small number of assets with given features (such as small drawdowns or high return...) for a subsequent allocation possibly using a generalization of the Markowitz approach. Alternatively, performance measures may be used to select a number of assets for the direct application of naïve portfolio allocation rules, such as the equally weighted one (see De Miguel et al., 2009). The financial economics literature proposes also to use performance measures as objective functions for determining the weights of an optimal portfolio. We will not pursue this objective, but for an example of such an approach see Farinelli et al. (2008, 2009). A relevant point is still open and has recently attracted some interest: which performance measure should be used? In fact, many reward-to-risk ratios have been proposed. Besides the well-known Sharpe, Sortino and Treynor indices, a number of alternative measures are available, such as the Omega index (Shadwick and Keating, 2002), the Rachev ratio (Rachev et al., 2003), and the FT ratios (Farinelli and Tibiletti, 2003), among others. Their number is increasing over time, www.economics-ejournal.org 1
conomics: The Open-Access, Open-Assessment E-Journal and new indices are designed to meet specific requirements, for example Pedersen and Rudholm-Alfvin (2003), or with the purpose of overcoming the limitations of the oldest measures. Some examples are given by the need of increasing the robustness of performance measures with respect to deviations from normality, or of introducing measures more appropriate for agents characterized by loss aversion (Gemmill et al., 2006) or by aggressiveness (Farinelli and Tibiletti, 2003). The comparison of alternative performances, generally using rank correlations, have already been considered. In particular, we refer to Gemmill et al. (2006), Eling and Schuhmacher (2007), Eling (2008), and Eling et al. (2011). These contributions use a simple and effective approach for deciding which measures to use: in order to compare alternative indices, they verify whether they rank assets differently. Performance measures providing equivalent rankings are redundant and may thus be discarded. Following this method, we may identify a restricted set of performance measures carrying different information on the risk/return trade-off. In this work we follow the empirical approach of Eling and Schuhmacher (2007) and provide three main contributions. The first one extends and completes the cited paper by broadening the set of performance measures to be compared. In particular, we include performance measures based on partial moments (Farinelli and Tibiletti (2003) and Rachev et al. (2003), as in Eling et al. (2011), and on loss aversions (Gemmill et al., 2006). In addition, we base our analysis on equities, rather than on managed portfolios as in Eling and Schuhmacher (2007), Eling (2008), and Eling et al. (2011). With respect to these issues, and differently from Eling and Schuhmacher (2007), we find cases of low rank correlation across performance measures, and then we argue that the equivalence relations may depend on the kind of assets considered and on the sample period. We also introduce four new performance measures: the expected return over range, where the risk measure is given by the maximum range; the VaR ratio, which is the ratio of the upper and lower quantiles of a given return distribution; and two performance measures derived from a utility function with loss aversion. The second contribution is associated with a different topic: the stability over time in the rankings induced by different performance measures. We will try to answer the question: "Are rank correlations time-varying?" To that end, we compare the rank correlations computed both over samples of different length, and over rolling windows. Our analysis extends the studies of Eling and Schuhmacher www.economics-ejournal.org 2
conomics: The Open-Access, Open-Assessment E-Journal (2007), and Eling (2008) that did not consider the rolling approach but evaluate the rank correlations on the full sample and on a two or five years sample. We show that, for our data, the rank correlations are not time invariant and are influenced by the sample size. Therefore, on the one side, appropriate tools for comparing and selecting performance measures are needed, while, on the other side, these dynamics could be exploited within an asset management framework. Building on this new evidence, for the third contribution of this work, we tackle the topic of the redundancy of the performance measures in a dynamic context. Given a set of N performances measures, we propose a way to reduce them in order to consider only those which really carry different information. In our empirical study we start, in the most general case, with 80 measures and, using a procedure based on the asymptotic distribution of the rank correlation coefficient, we conclude that 57 measures are redundant since they carry information similar to the 23 we select. In connection with the second outcome of this paper, we also infer that the set of performance measures carrying relevant information may be time-varying as well. This additional piece of information could be proven to be extremely relevant for periodic rotation or rebalancing of managed portfolios using asset screens. Given that the allocation choices of portfolio managers and agents are generally taken at a low frequency (monthly to quarterly) in this paper we work with monthly data, but analysis at different frequencies may be considered. Moreover, we assume that the series of interest are characterized by deviations from normality (which, for equities, is one of the well-known stylized facts, see Cont, 2001, among others), and that the risk and reward measures presented below are estimated with their sample counterparts without introducing a parametric model. The rest of the paper is organized as follows. Section 2 lists the performance measures that will be considered, describes the dataset, and discusses some problems connected to the selection bias. In Section 3 we report the results of the analysis concerning the correlations between different performance measures and we show how to obtain the set of the measures that are significantly different. Our final conclusions are presented in Section 4. www.economics-ejournal.org 3
conomics: The Open-Access, Open-Assessment E-Journal 2 Performance Measures List and Dataset Description From a general viewpoint, performance measures can be defined as ratios between a reward measure and a risk measure, and their value can be interpreted as the reward per unit of risk. Despite a general agreement on what a performance measure is, a number of choices are available for the reward and risk measures to be considered, as well as for the type of variables to be used for their evaluation. In order to provide a general setup, we start by introducing some notation: we denote by Ri,t the (nominal) log-return of asset i in period t ; Rf,t is the riskfree investment return (it is time-varying since we consider it as a pure risk-free investment within each period); RB,t identifies the return of a benchmark investment; XT t=1 is the sequence of observations of the variable Xt from time 1 to time T ; E[Xp] is the moment of order p of X ; E[g(X)p] is the moment of order p of the function g(X) ; σ[X] is the volatility of X ; and, E[Xp|Y] is the conditional moment of order pof X. The performance measures presented below will be defined over a variable Xi,t that takes one of the following values Xi,t= Ri,t Ri,t−Rf,t Ri,t−RB,t .(1) These cases represent three possible relevant dimensions for performance measurement, not necessarily mutually exclusive: nominal returns (relevant for agents focusing on purely risk investments), excess returns with respect to a risk-free returns (for investors considering also a risk-free investment), deviations from a benchmark (relevant within an active management framework). We now describe the performance measures we consider, grouping them from a statistical point of view (thus separating the use of general risk measures, from ratios based on partial moments and quantiles, and those derived from utility functions). Of course, different and more detailed classifications could have been used, Aftalion and Poncet (2003), Le Sourd (2007), and Cogneau and Hubner (2009a, 2009b). Nevertheless, we prefer to maintain a limited and simple structure, and our selection of performance measures includes quantities designed to capture deviations from normality as well as to take into account agents’ behaviour. www.economics-ejournal.org 4
conomics: The Open-Access, Open-Assessment E-Journal 2.1 Traditional Performance Measures and Other Unclassified Measures This first set of performance measures contains the most known and traditional indices: – the Sharpe ratio, introduced by Sharpe (1966, 1994): Sh(Xi,t) = E[Xi,t] σ[Xi,t[; (2) – the Treynor index , Treynor (1965), defined for nominal returns and excess returns only: Tr (Xi,t) = E[Xi,t] βi ,(3) where βiis estimated through a CAPM regression; – the Appraisal ratio, defined as: AR(Xi,t) = αi σ[εi,t],(4) where αi is the intercept of a CAPM regression and σ[εi,t] denotes the volatility of the CAPM residuals; – the expected return over Mean Absolute Deviation ratio of Konno and Yamazaki (1991): ERMAD(Xi,t) = E[Xi,t] E[|Xi,t−E[Xi,t]|].(5) We also include here some performance measures which are not consistent with the following groups and are defined as ratios between the first order moment of Xi,tand a risk measure: – the return over MiniMax ration of Young (1998): ERMM (Xi,t) = E[Xi,t] maxmaxXT t=1,−minXT t=1; (6) – the expected return over the range ratio, which, to our knowledge has never been considered in previous studies: ERR(Xi,t) = E[Xi,t] maxXT t=1−minXT t=1 .(7) www.economics-ejournal.org 5
conomics: The Open-Access, Open-Assessment E-Journal Finally, we include here also the Risk Adjusted Performance (RAP), or M2 index of Modigliani and Modigliani (1997): M2= (E[Ri,t]−E[RB,t]) σ[RB,t] σ[Ri,t]+E[Rf,t]−E[RB,t].(8) 2.2 Measures Based on Drawdown This set contains measures based on risk indices focusing on the drawdown, which is define as Dt(Xi,t) = min(Dt−1+Xi,t,0)D0=0.(9) Given the observations for Xi,tt=1, ...T , the drawdown Dt(Xi,t) or simply Dt represents, at time t , the maximum loss an investor may have suffered from 1 to t . Risk measures are defined ordering drawdowns and computing quantities such as the maximum drawdown, D1(Xi,t) = minDT t=1 , or the second largest drawdown D2(Xi,t) = minDT t=1−D1(Xi,t) , and so on. We also assume D1(Xi,t)<0 . We consider three indices based on drawdowns: – the Calmar ratio of Young (1991): CR(Xi,t) = E[Xi,t] −D1(Xi,t); (10) – the Sterling ratio, introduced by Kestner (1996): SR(Xi,t;w) = E[Xi,t] −1 w∑w j=1Dj(Xi,t); (11) where w is a parameter that identifies the number of values used in the computation of the risk index; – the Burke ratio, due to Burke (1994): BR(Xi,t;w) = E[Xi,t] 1 w∑w j=1[Dj(X,t)]21 2 .(12) In the Burke and Sterling ratios, Eling and Schuhmacher (2007) fix the value of w between 1 and 10 Differently, we prefer linking the number of drawdowns to the sample dimension as w=T 20 ,T 10 where [a]denotes the nearest integer of a. www.economics-ejournal.org 6
conomics: The Open-Access, Open-Assessment E-Journal 2.3 Measures Based on Partial Moments We also analyze performance measures based on partial moments: - the Sortino ratio, Sortino and Van der Meer (1991): Sr (Xi,t) = E[Xi,t] Eh(min(Xi,t,0))2i1 2 ; (13) – the Kappa 3 measure of Kaplan and Knowles (2004): K3(Xi,t) = E[Xi,t] Eh(min(Xi,t,0))3i1 3 .(14) – the Farinelli and Tibiletti (2003) ratio, or FT ratio: FT (Xi,t;b,p,q) = Eh(Xi,t−b)+pi1 p Eh(Xi,t−b)−qi1 q (15) where (Xi,t−b)+=max(Xi,t−b,0) , (Xi,t−b)−=max(b−Xi,t,0) . The threshold return level b , and the partial moment orders p and q are calibrated following Farinelli and Tibiletti (2003) in order to match them with possible investors’ styles or preferences: p=0.5 and q=2 for a defensive investor; p=1.5 and q=2 for a conservative investor; p=q=1 for a moderate investor (note that this combination makes the FT (Xi,t;b,1,1) equivalent to the Omega index of Shadwick and Keating (2002)); p=2 and q=1.5 for a growth investor; p=3 and q=0.5 for an aggressive investor; in addition, p=1 and q=2 defines the Upside Potential Ratio of Sortino et al. (1999). Finally, we consider the following cases for the threshold return, b={−0.02,0,0.02} , where the −2% and 2% values may represent the choices of a less risk averse and a more risk averse investor, respectively. 2.4 Measures Based on Quantiles A class of performance measures similar to the previous one replaces partial moments with reward and variability measures based on quantiles (see Rachev www.economics-ejournal.org 7
conomics: The Open-Access, Open-Assessment E-Journal Table 2: Rank correlations across selected performances measures - drawdowns and quantile based measures. The first column reports the pair of performance measures compared and the corresponding parameter value. The other columns report the rank correlations across three return types (asset returns, excess returns with respect to a risk free investment, and deviations between asset returns and a benchmark investment), and three sample dimensions (36, 60 and 120 months). Bold values identify rank correlations below the minimum threshold of 0.822 defined in Section 3, and denote relevant differences across the ranks induced by the two performance measures compared. Rank correlations Returns Excess Deviations from Returns benchmark Window length 36 60 120 36 60 120 36 60 120 Sterling (5%) and (10%) 1.000 0.997 0.998 1.000 1.000 0.998 1.000 0.999 0.992 Burke (5%) and (10%) 0.997 0.974 0.880 0.996 0.969 0.866 0.995 0.979 0.903 VR index (5%) and (10%) 0.993 0.991 0.985 0.992 0.992 0.990 0.994 0.994 0.987 VaR Ratio (5%) and (10%) 0.727 0.699 0.600 0.728 0.701 0.611 0.795 0.743 0.586 STARR (5%) and (10%) 0.997 0.998 0.997 0.998 0.998 0.998 0.998 0.999 0.998 www.economics-ejournal.org 14
conomics: The Open-Access, Open-Assessment E-Journal Note that in our case, the rank correlation is computed between rankings induced by performance measures within the set of the N considered assets. Thus ?? holds for N large, since, in general, a large number of assets is analysed within equity screening programs. This allows us to define the required threshold for RS as R∗ S(α) = expln1+RS 1−RS+2Z1−αq1 N−2−1 expln1+RS 1−RS+2Z1−αq1 N−2+1 ,(29) where Z1−α is the (1−α)− th quantile of a standard normal distribution. Such a quantity, corresponds thus to the critical value for the null hypothesis reported above. Such a choice allows a more direct interpretation of results, without resorting to the Fisher transformation of all quantities. In our analysis, with N=1236 in the static case, and α=1% , the threshold (or critical value) defining the low correlation is 0.822 . We further note that the sample size plays a relevant role in the definition of the critical value. For a very small number of assets, say below 50, the critical value would be quite large, easily leading to an acceptance of the null. However, the normal use of performance measures within an equity ranking program (an equity screening rule) involves evaluations over hundreds of assets, thus increasing the power of the test. 3.1 Within Group Analysis In this section we report, analyze and comment on the rank correlation between performance measures that differ only for the parameters included in their definition. The purpose of this section is to provide a first reduction of the number of performance measures included in Table 1. The first group we consider includes some measures based on drawdowns: the Sterling and Burke indices. These two quantities depend on the number of returns used for their computation. In the previous section we suggested the use of at least two values associated with 5% and 10% of the sample dimension. Given these two sets of performance measures, we evaluate whether the sample size used in the computation of the indices provides a different ranking across the assets. The results are reported in the first and second row of Table 2. The rank correlations www.economics-ejournal.org 15
conomics: The Open-Access, Open-Assessment E-Journal Table 3: Rank correlations across selected performances measures - Generalized Rachev Ratios. The first column reports the set of performance measures considered within each row. The first and second rows report the average rank correlation across the Generalized Rachev ratios for the parameter combinations associated to Moderate, Conservative, Growth and Defensive investors with a given quantile level. The third row reports the average rank correlation between the two groups associated to the first and second row. The other three rows report the average rank correlation between the Generalized Rachev ratios for Aggressive investors with respect to other indices. Columns from 2 to 10 contain the rank correlations values across three return types (asset returns, excess returns with respect to a risk free investment, and deviations between asset returns and a benchmark investment), and three sample dimensions (36, 60 and 120 months). Bold values identify rank correlations below the minimum threshold of 0.822 defined in Section 3, and denote relevant differences across the ranks induced by the performance measures compared within each row. Rank correlations Returns Excess Deviations from Returns benchmark Window length 36 60 120 36 60 120 36 60 120 Within GR (10%) excl. Aggressive 0.989 0.986 0.982 0.989 0.986 0.982 0.99 0.986 0.981 Within GR (5%) excl. Aggressive 0.997 0.995 0.992 0.997 0.995 0.992 0.998 0.995 0.991 Between GR (5%) and GR (10%) excl. Aggressive 0.957 0.955 0.955 0.956 0.955 0.956 0.962 0.956 0.955 GR Aggressive (10%) wrt other GR (10%) 0.941 0.898 0.848 0.955 0.912 0.871 0.772 0.699 0.788 GR Aggressive (5%) wrt other GR (5%) -0.474 0.233 0.946 -0.489 0.150 0.937 0.047 0.660 0.954 GR Aggressive (5%) - GR Aggressive (10%) -0.452 0.192 0.853 -0.459 0.128 0.868 -0.072 0.415 0.810 www.economics-ejournal.org 16
conomics: The Open-Access, Open-Assessment E-Journal show evidence of equivalent informative content of the performance measures with respect to the number of returns used for the evaluation of the Burke and Sterling indices. Results do not change with respect to the sample dimension or to the return used for the evaluation. We conclude that there is no need to consider the Sterling and Burke indices computed over different numbers of drawdowns. This result confirms the findings of Eling and Schuhmacher (2007). The second set of performance measures we analyze includes the quantile based measures, with the exclusion of the Generalized Rachev ratios. Table 2 reports the rank correlations between the VR index, the VaR ratios and the STARR ratio at the 5% and 10% quantile levels. Results show that the VR index and the STARR ratios should be considered with a single quantile level (rank correlation is always higher than 0.985) while the VaR ratio should be considered with both the 5% and 10% quantiles, given that the rank correlation is lower than 0.822 in all cases and also reaches a minimum close to 0.6 with a 10 years sample dimension (irrespective of the return used). Table 3 reports the rank correlations across the Generalized Rachev ratios. We recall that we computed 10 different GR ratios combining five parameter combinations (Aggressive, Growth, Moderate, Conservative and Defensive) with two quantile levels (5% and 10%). We distinguished two groups, separating the effect of the Aggressive indices. Our analysis points out that this last parameter combination is the most sensible to the sample dimension, providing results different from the other GR ratios when the sample used is medium to small (3 or 5 years). The difference tends to be canceled with the sample set to 10 years, with the exclusion of the case of the evaluation of deviations from the benchmark. Differently, the other GR ratios (Growth to Defensive) are almost equivalent (the smallest rank correlation is equal to 0.955). In addition, the effect of the quantile level is minor. Building on these results, we chose to include the Moderate GR ratio at the 10% level when the sample dimension is large (10 years). In contrast, when the sample is small or medium, the GR for Aggressive investors will also be considered (again at the 10% level). Following the performance measure groups previously introduced, we move then to measures based on partial moments that include the indices of Sortino, the Kappa 3 index and the FT ratios. Similarly to the Generalized Rachev ratios, we group the FT performance measures into two sets, separately considering the www.economics-ejournal.org 17
conomics: The Open-Access, Open-Assessment E-Journal Table 4: Rank correlations across selected performances measures - Farinelli-Tibiletti Ratios. The first column reports the set of performance measures considered within each row. We considered two groups: the first is composed by the parameter combinations associated to Defensive, Conservative, Moderate and Growth investors (in all possible pairs); the second group contains the pairs of performance measures involving at least one measure associated to Aggressive investors. "Within" stands for the average of rank correlations across the pairs of measures with the same minimum acceptable return. "Between" stands for the average rank correlations across the pairs of measures with different minimum acceptable returns. As an example, the line "Within -0.02" in the first group, contains the averages of rank correlations over the following pairs: UPR-Defensive, UPR-Conservative, UPR-Moderate, UPR-Growth, Defensive-Conservative, Defensive-Moderate, Defensive-Growth, Conservative-Moderate, Conservative-Growth, Moderate-Growth. Columns from 2 to 10 contain the rank correlations values across three return types (asset returns, excess returns with respect to a risk free investment, and deviations between asset returns and a benchmark investment), and three sample dimensions (36, 60 and 120 months). Bold values identify rank correlations below the minimum threshold of 0.822 defined in Section 3, and denote relevant differences across the ranks induced by the performance measures compared within each row. Rank correlations Returns Excess Deviations from Returns benchmark Window length 36 60 120 36 60 120 36 60 120 Within and Between UPR, Defensive, Conservative, Moderate and Growth Within -0.02 0.941 0.944 0.914 0.938 0.942 0.907 0.956 0.960 0.927 Within 0 0.915 0.919 0.871 0.912 0.917 0.872 0.928 0.926 0.882 Within 0.02 0.915 0.926 0.922 0.918 0.929 0.931 0.902 0.907 0.915 Between -0.02 and 0 0.865 0.843 0.716 0.848 0.825 0.682 0.881 0.843 0.725 Between -0.02 and 0.02 0.557 0.428 0.147 0.534 0.406 0.140 0.581 0.414 0.213 Between 0 and 0.02 0.790 0.736 0.649 0.794 0.744 0.682 0.798 0.735 0.701 Within Aggressive and Between Aggressive and Defensive, Conservative, Moderate and Growth Within -0.02 -0.268 0.045 0.498 -0.365 -0.100 0.356 -0.111 0.067 0.470 Within 0 0.383 0.05 -0.108 0.532 0.255 0.163 0.014 0.027 -0.011 Within 0.02 0.857 0.853 0.819 0.874 0.864 0.840 0.852 0.847 0.823 Between -0.02 and 0 0.058 0.025 0.103 0.086 0.057 0.152 -0.047 0.025 0.146 Between -0.02 and 0.02 0.125 0.104 -0.029 0.112 0.073 -0.041 0.171 0.09 0.018 Between 0 and 0.02 0.513 0.349 0.242 0.574 0.421 0.333 0.386 0.346 0.269 www.economics-ejournal.org 18
conomics: The Open-Access, Open-Assessment E-Journal Aggressive parameter combination. The results are reported in Table 4, where the first group includes the parameter combinations Growth, Moderate, Conservative, Defensive as well as the Upside Potential Ratio (which is a special case of the FT index as we previously argued). Our analysis shows that these parameter combinations do not provide additional information or relevant differences in the ranking of the underlying assets (first to third rows). The result is marginally influenced by the sample length and the kind of return used in the evaluation of the indices. On the other hand, the threshold used in the index construction matters, making the indices sensibly different in terms of assets ranking (fourth to sixth rows of Table 4). In fact, the rank correlations across indices computed over different thresholds are generally small and always lower than 0.822. When considering the Aggressive parameter combination, the rank correlations are always small, and sometimes negative (this is a consequence of limited relevance given to the risk by that parameter combination). In addition, they are affected by both the sample dimension and the return type. Summarizing, we suggest considering the FT Moderate index (or Omega index) together with the Aggressive parameter combination, under all three of the thresholds considered. For the Sortino and Kappa 3 indices, the rank correlation with respect to the Omega index is higher than 0.98 and therefore the two indices are not considered. Moving to the performance measures based on utility functions (Table 5), we first note that the MRAR indices with risk aversion set to 10 and 50 are almost equivalent. Therefore, we decide to focus on the measure with risk aversion set to 2 and 10. By contrast, in the LAP measures, the Hwang-Satchell, Moderate and Growth parameter combinations are almost equivalent while the Conservative case is very close to them. In order to provide a selection of measures which is limited, internally consistent, and that maximizes the difference across parameter combinations, we suggest focusing on the cases Defensive, Moderate and Aggressive. Within each group, we suggest considering all performance measures even if the Moderate case reports a high within-group average rank correlation. Finally, we consider a further group composed by most of the traditional performance measures. Table 6 includes the rank correlation of these indices with the ranking induced by the Sharpe ratio. As we may observe, all indices are almost identical to the Sharpe ratio in terms of ranking of the assets. Some minor exceptions are the Appraisal ratio and the M2 index for the 3 year sample. www.economics-ejournal.org 19
conomics: The Open-Access, Open-Assessment E-Journal Table 5: Rank correlations across selected performances measures - utility based performance measures. The first column reports the set of performance measures considered within each row. We separately consider the MRAR measures and the Loss Aversion Performance measures. The last one are grouped depending on their parameters in Hwang-Satchell (HS), Defensive, Conservative, Moderate, Growth and Aggressive. Apart the Hwang-Satchell case, the groups do not include the LAP S measure which is equivalent to the FT Moderate Index with threshold set at zero. For MRAR measures we report the rank correlation coefficients. For LAP groups we report the average rank correlation within each group and between each pair of groups. Columns from 2 to 10 contain the rank correlations values across three return types (asset returns, excess returns with respect to a risk free investment, and deviations between asset returns and a benchmark investment), and three sample dimensions (36, 60 and 120 months). Bold values identify rank correlations below the minimum threshold of 0.822 defined in Section 3, and denote relevant differences across the ranks induced by the performance measures compared within each row. Rank correlations Returns Excess Deviations from Returns benchmark Window length 36 60 120 36 60 120 36 60 120 MRAR 2 - MRAR 10 0.252 -0.076 -0.006 0.462 0.068 0.256 0.021 -0.124 0.024 MRAR 2 - MRAR 50 0.168 -0.132 0.003 0.404 0.047 0.260 -0.153 -0.160 0.052 MRAR 10 - MRAR 50 0.894 0.858 0.955 0.940 0.922 0.974 0.707 0.806 0.97 LAP - Within HS 0.951 0.939 0.685 0.950 0.938 0.685 0.932 0.932 0.715 LAP - Within Defensive -0.067 -0.054 0.081 -0.052 -0.046 0.074 -0.101 -0.069 0.052 LAP - Within Conservative 0.813 0.759 0.534 0.853 0.776 0.510 0.777 0.767 0.494 LAP - Within Moderate 0.958 0.927 0.670 0.957 0.929 0.665 0.944 0.924 0.666 LAP - Within Growth 0.897 0.864 0.432 0.898 0.861 0.446 0.850 0.837 0.469 LAP - Within Aggressive 0.523 0.591 0.508 0.477 0.580 0.520 0.408 0.514 0.517 LAP - Between HS-Defensive 0.081 0.027 -0.121 0.078 0.012 -0.125 0.061 -0.008 -0.145 LAP - Between HS-Conservative 0.741 0.698 0.154 0.751 0.703 0.123 0.684 0.634 0.064 LAP - Between HS-Moderate 0.948 0.928 0.671 0.947 0.929 0.669 0.935 0.925 0.686 LAP - Between HS-Growth 0.835 0.838 0.562 0.831 0.837 0.567 0.787 0.807 0.590 LAP - Between HS-Aggressive 0.339 0.434 0.477 0.322 0.439 0.506 0.240 0.367 0.531 LAP - Between Defensive-Conservative 0.156 0.104 0.173 0.141 0.097 0.161 0.100 0.071 0.156 LAP - Between Defensive-Moderate 0.108 0.059 -0.023 0.103 0.045 -0.031 0.076 0.013 -0.061 LAP - Between Defensive-Growth 0.052 0.008 -0.042 0.044 -0.003 -0.048 0.036 -0.021 -0.060 LAP - Between Defensive-Aggressive -0.051 -0.074 -0.145 -0.055 -0.077 -0.146 -0.037 -0.078 -0.139 LAP - Between Conservative-Moderate 0.817 0.783 0.417 0.832 0.791 0.388 0.769 0.738 0.326 LAP - Between Conservative-Growth 0.734 0.672 0.253 0.745 0.679 0.237 0.673 0.623 0.205 LAP - Between Conservative-Aggressive 0.128 0.134 -0.179 0.127 0.146 -0.177 0.028 0.043 -0.199 LAP - Between Moderate-Growth 0.831 0.819 0.536 0.824 0.817 0.541 0.787 0.791 0.568 LAP - Between Moderate-Aggressive 0.264 0.333 0.257 0.249 0.341 0.284 0.174 0.277 0.319 LAP - Between Growth-Aggressive 0.485 0.556 0.442 0.463 0.551 0.461 0.392 0.481 0.467 www.economics-ejournal.org 20
conomics: The Open-Access, Open-Assessment E-Journal Table 6: Rank correlation of traditional and similar performance measures with the Sharpe ratio over different sample length. The first column reports the performance measure which is compared to the Sharpe ratio. Columns from 2 to 10 contain the rank correlations values across three return types (asset returns, excess returns with respect to a risk free investment, and deviations between asset returns and a benchmark investment), and three sample dimensions (36, 60 and 120 months). Bold values identify rank correlations below the minimum threshold of 0.822 defined in Section 3, and denote relevant differences across the ranks induced by the performance measures reported in each row and the Sharpe ratio. Rank correlations Returns Excess Deviations from Returns benchmark Window length 36 60 120 36 60 120 36 60 120 Treynor 0.950 0.934 0.858 0.945 0.940 0.906 — — — Appraisal ratio 0.792 0.940 0.906 0.555 0.915 0.893 — — — ERMAD 0.999 0.999 0.997 0.999 0.999 0.998 0.999 0.999 0.998 ERR 0.995 0.992 0.966 0.994 0.994 0.985 0.994 0.989 0.978 ERMM 0.991 0.987 0.956 0.990 0.990 0.978 0.992 0.986 0.969 M2 0.719 0.928 0.964 — — — — — — www.economics-ejournal.org 21
conomics: The Open-Access, Open-Assessment E-Journal Overall, we may infer that the Treynor index, the Appraisal ratio and the indices replacing the standard deviation in the Sharpe with a proxy are all equivalent. We thus suggest introducing in the following analysis only the Sharpe ratio. Notably, this result is in line with the findings of Eling and Schuhmacher (2007). In our case, the rank correlations are not as high as shown by these authors. Furthermore, our results point out that the equivalence across the selected performance measures is not influenced by the return used for the evaluation and only scarcely affected by the sample dimension. After this within-group analysis, we select the following performance measures: the Sharpe ratio; the Calmar ratio; the Sterling Ratio and the Burke ratio computed over the 5% of the sample dimension; the VR index and the STARR at the 5% quantile; the VaR ratio at both the 5% and 10% quantiles; the Generalized Rachev ratio with Moderate parameter combination at the 10% quantile level (one single index - the Aggressive index is included only if the evaluation window is small); the FT Moderate and Aggressive indices under all three threshold levels (6 indices); the MRAR index with risk aversion set to 2 and 10; and the LAP measures for Defensive, Moderate and Aggressive parameter combinations (9 indices). On the whole, the total number of selected measures is 26. 3.2 Descriptive and Rolling Analysis of Selected Measures We run additional correlation analysis on the reduced set of performance measures identified in the previous section. As a first outcome, we highlight that some of the measures are still highly correlated. In particular, we report in Table 7 the correlation between the Sharpe ratio and some selected measures. As shown in the table, we may infer that the Calmar ratio, the Sterling ratio (5%), the VR Index (5%), and the STARR (5%) are all equivalent to the Sharpe ratio. These findings confirm the results of Eling and Schuhmacher (2007) and are in line with the findings of Ortobelli et al. (2005) showing that traditional risk measures induce indifference across performance measures where the reward index is the average return. However, we obtain rather different rank correlations for Omega, with values going down to 0.536 and high rank correlation for long samples (120 months) only. Note that these differences are pronounced if we compute the Omega www.economics-ejournal.org 22
conomics: The Open-Access, Open-Assessment E-Journal Table 7: Rank correlation of selected measures with the Sharpe ratio. The first column reports the performance measure which is compared to the Sharpe ratio. Columns from 2 to 10 contain the rank correlations values across three return types (asset returns, excess returns with respect to a risk free investment, and deviations between asset returns and a benchmark investment), and three sample dimensions (36, 60 and 120 months). Bold values identify rank correlations below the minimum threshold of 0.822 defined in Section 3, and denote relevant differences across the ranks induced by the performance measures reported in each row and the Sharpe ratio. Rank correlations Returns Excess Deviations from Returns benchmark Window length 36 60 120 36 60 120 36 60 120 Calmar ratio 0.983 0.977 0.930 0.982 0.978 0.965 0.987 0.977 0.946 Sterling ratio (5%) 0.982 0.978 0.932 0.785 0.903 0.912 0.833 0.951 0.911 VR Index (5%) 0.993 0.991 0.983 0.993 0.993 0.992 0.994 0.994 0.990 STARR (5%) 0.996 0.994 0.983 0.997 0.996 0.994 0.995 0.996 0.993 Omega 0.723 0.933 0.991 0.532 0.866 0.991 0.865 0.907 0.993 www.economics-ejournal.org 23
conomics: The Open-Access, Open-Assessment E-Journal References Aftalion, F., and Poncet, P. (2003).Les techniques de mesure de performance. Economica, Paris. Avramov, D., and Chordia, T. (2006). Asset pricing models and financial market anomalies. Review of Financial Studies, 19 (3): 1001–1040. http://ideas.repec.org/a/oup/rfinst/v19y2006i3p1001-1040.html Barberis, N., Hwang, M., and Santos, T. (2001). Prospect theory and asset prices, Quarterly Journal of Economics, 116 (1): 1– 53.http://ideas.repec.org/a/tpr/qjecon/v116y2001i1p1-53.html Biglova, A., Ortobelli, S., Rachev, S., and Stoyanov, S. (2004). Different approaches to risk estimation in portfolio theory, Journal of Portfolio Management, 31 (1): 103–112. http://www.iijournals.com/doi/abs/10.3905/jpm.2004.443328 Burke, G. (1994). A sharper Sharpe ratio, Futures, 23 (3): 56. Cogneau, P., and Hubner, G. (2009a). The (more than) 100 ways to measure portfolio performance. Part 1: standardized risk-adjusted measures, Journal of Performance Measurement, 13 (4). http://orbi.ulg.ac.be/handle/2268/2782 Cogneau, P., and Hubner, G. (2009b). The (more than) 100 ways to measure portfolio performance. Part 2: special measures and comparison, Journal of Performance Measurement, 14 (1). http://orbi.ulg.ac.be/handle/2268/30494 Cont, R. (2001). Empirical properties of asset returns: stylized facts and statistical issues, Quantitative Finance, 1 (2): 223–236. http://citeseer.ist.psu.edu/viewdoc/summary?doi=10.1.1.16.5992 De Miguel, V., Garlappi, L., and Uppal, R. (2009). Optimal versus naive diversification: how inefficient is the 1/N portfolio strategy? Review of Financial Studies, 22 (5): 1915–1953. http://ideas.repec.org/a/oup/rfinst/v22y2009i5p1915-1953.html www.economics-ejournal.org 30
conomics: The Open-Access, Open-Assessment E-Journal Eling, M. (2008). Does the measure matters in the mutual fund industry? Financial Analyst Journal, 64 (3): 54–66. http://papers.ssrn.com/sol3/papers.cfm4?abstractid=1266187 Eling, M., and F. Schuhmacher (2007). Does the choice of performance measure influence the evaluation of hedge funds? Journal of Banking and Finance, 31: 2632–2647. http://ideas.repec.org/a/eee/jbfina/v31y2007i9p2632-2647.html Eling, M., Farinelli, S., Rossello, D., and Tibiletti, L. (2011). One-size or tailor-made performance ratios for ranking hedge funds, Journal of Derivatives and Hedge Funds, 16: 267–277. http://www.palgravejournals.com/jdhf/journal/v16/n4/full/jdhf201020a.html Farinelli, S., and Tibiletti, L. (2003). Upside and downside risk with a benchmark, Atlantic Economic Journal, Anthology Section, 31 (4): 387. http://ideas.repec.org/a/kap/atlecj/v31y2003i4p387-387.html Farinelli, S., Ferreira, M., Rossello, D., Thoeny, M., and Tibiletti, L. (2008). Beyond Sharpe ratio: optimal asset allocation using different performance ratios, Journal of Banking and Finance, 32: 2057–2063. http://ideas.repec.org/a/eee/jbfina/v32y2008i10p2057-2063.html Farinelli, S., Ferreira, M., Rossello, D., Thoeny, M., and Tibiletti, L. (2009). Optimal asset allocation aid system: from "one-size" vs "taylor-made" performance ratio, European Journal of Operational Research, 192: 209–215. Ferson, W.E., and Schadt, R. (1996). Measuring fund strategy and performance in changing economic conditions, Journal of Finance, 51: 425–461. http://ideas.repec.org/a/bla/jfinan/v51y1996i2p425-61.html Fisher, R.A. (1915). Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population, Biometrika, 10: 507–521. http://www.jstor.org/stable/2331838 Gemmill, G., Hwang, S., and Salmon, M. (2006). Performance measurement with loss aversion, Journal of Asset Management, 7 (3): 190–207. http://www.palgrave-journals.com/jam/journal/v7/n3/abs/2240213a.html www.economics-ejournal.org 31
conomics: The Open-Access, Open-Assessment E-Journal Hwang, S., and Salmon, M. (2003). An analysis of performance measures using copulae, in: Knigth, J., and Satchell, S. (eds), Performance measurement in finance: firms, funds and managers, Butterworth-Heinemann Finance, Quantitative Finance Series. Jensen, M. (1968). The performance of mutual funds in the period 1945-1968, Journal of Finance, 23 (2): 389–416. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.139.6166 Kahnemann, D., and Tversky, A. (1979). Prospect theory: an analysis of decision under risk, Econometrica, 47: 263–291. http://ideas.repec.org/a/ecm/emetrp/v47y1979i2p263-91.html Kaplan, P.D., and Knowles, J.A. (2004). Kappa: A Generalized Downside RiskAdjusted Performance Measure, Morningstar Associates and York Hedge Fund Strategies, January. Kestner, L.N. (1996). Getting a handle on true performance, Futures, 25 (1): 44–46. http://www.allbusiness.com/personal-finance/investing-tradingfutures/538174-1.html Knigth, J., and Satchell, S. (eds) (2002). Performance measurement in finance: firms, funds and managers, Butterworth-Heinemann Finance, Quantitative Finance Series. Konno, H., and Yamazaki, H. (1991). Mean-absolute deviation portfolio optimization model and its application to Tokyo stock market, Management Science, 37: 519–531. http://www.jstor.org/stable/2632458 Le Sourd, V. (2007). Performance measurment for traditional investment. Financial Analysts Journal, 58 (4): 36–52. Lintner, J. (1965). The valuation of risky assets and the selection of risky investment in stock portfolios and capital budgets, Review of Economics and Statistics, 47: 13–37. http://www.jstor.org/stable/1924119 Markowitz, H. (1959). Portfolio selection: efficient diversification of investments, John Wiley. www.economics-ejournal.org 32
conomics: The Open-Access, Open-Assessment E-Journal Modigliani, F., and Modigliani, L. (1997). Risk-adjusted performance – how to measure it and why, Journal of Portfolio Management, 23 (2): 45–54. Mossin, J. (1969). Security pricing and investment criteria in competitive markets, American Economic Review, 59: 749–756. http://ideas.repec.org/a/aea/aecrev/v59y1969i5p749-56.html Ortobelli, S., Rachev, S., Stoyanov, S., Fabozzi, F.J., and Biglova, A. (2005). The proper use of risk measures in portfolio theory, International Journal of Theoretical and Applied Finance, 8 (8): 1107–1133. http://ideas.repec.org/a/wsi/ijtafx/v08y2005i08p1107-1133.html Pedersen, C.S., and Rudholm-Alfvin, T. (2003). Selecting a risk-adjusted shareholder performance measure, Journal of Asset Management, 4 (3): 152–172. http://www.palgrave-journals.com/jam/journal/v4/n3/abs/2240101a.html Rachev, S., Martin, D., and Siboulet, F. (2003). Phi-alpha optimal portfolios and Extreme Risk Management, Wilmott Magazine of Finance, November, 70–83. http://www.wilmott.com/pdfs/060530martin.pd f Shadwick, W.F., and Keating, C. (2002). A universal performance measure, Journal of Performance Measurement, 6 (3): 59–84. Sharma, M. (2004). A.I.R.A.P. - Alternative RAPMs for alternative investments, Journal of Investment Management, 3 (4). https://www.joim.com/abstract.asp?ArtID=116 Sharpe,W.F. (1964). Capital asset prices: A theory of market equilibrium under. Conditions of risk, Journal of Finance, 19: 425–442. http://www.jstor.org/stable/2977928 Sharpe, W.F. (1966). Mutual fund performance, Journal of Business, 39 (1): 119– 138. http://finance.martinsewell.com/fund-performance/Sharpe1966.pdf Sharpe, W.F. (1992). Asset allocation: management style and performance measurement, Journal of Portfolio Management, 18 (2): 7–19. http://www.iijournals.com/doi/abs/10.3905/jpm.1992.409394 www.economics-ejournal.org 33
conomics: The Open-Access, Open-Assessment E-Journal Sharpe, W.F. (1994). The Sharpe ratio, Journal of Portfolio Management, Fall: 45–58. http://www.stanford.edu/ wfsharpe/art/sr/sr.htm Sortino, F.A. (2001). Managing downside risk in financial markets, ButterworthHeinemann Finance, Oxford. Sortino, F.A., and van der Meer, R. (1991). Downside risk, Journal of Portfolio Management, 17 (Spring): 27–31. http://www.iijournals.com/doi/abs/10.3905/jpm.1991.409343 Sortino, F.A., van der Meer, R., and Plantinga, A. (1999). The Dutch triangle, Journal of Portfolio Management, 26 (Fall): 50–58. Treynor, J.L. (1965). How to rate management of investment funds, Harvard Business Review, 43 (1): 63–75. Young, T.W. (1991). Calmar ratio: A smoother tool, Futures, 20 (1): 40. http://www.allbusiness.com/business-finance/equity-fundingstock/261555-1.html Young, M.R. (1998). A MiniMax portfolio selection rule with linear programming solution, Management Science, 44: 673–683. http://www.jstor.org/stable/2634472 www.economics-ejournal.org 34
Please note: You are most sincerely encouraged to participate in the open assessment of this article. You can do so by either recommending the article or by posting your comments. Please go to: http://dx.doi.org/10.5018/economics-ejournal.ja.2011-10 The Editor © Author(s) 2011. Licensed under a Creative Commons License - Attribution-NonCommercial 2.0 Germany