scieee AI-readable full text Open interactive document viewer

Evaluation of strategy portfolios

Wang, Anlan

Abstract

People usually create a portfolio in order to diversify the risk coming from individual investments. To get a high yield with a good level of diversification, investors usually seek professional advice from portfolio managers. However, the true performance of an optimized portfolio usually depends on the correctness of the estimates of the distribution of future returns, which is often a matter of luck rather than skill. Thus, the optimization models may not be better than randomly selected portfolios. Our aim is to find how the so-called strategy portfolios, i.e., portfolios obtained by some decision optimized for a long-run horizon, perform compared to a benchmark, namely, a random investment, under specific market conditions. For this purpose, we evaluate several portfolio strategies over two periods of crisis: the subprime mortgage crisis and the Covid-19 pandemic, as well as run a moving window analysis over a longer horizon. In each case, the results are compared with the performance of random-weight portfolios. We find that if the strategy is minimization, the portfolios perform well; however, for the maximization of the objectives, the results are rather mixed.

Full text

Vol.:(0123456789) Computational Management Science (2024) 21:17 https://doi.org/10.1007/s10287-023-00497-5 1 3 ORIGINAL PAPER Evaluation ofstrategy portfolios AnlanWang1· AlešKresta1· TomášTichý1 Received: 31 October 2022 / Accepted: 13 December 2023 / Published online: 20 January 2024 © The Author(s) 2024 Abstract People usually create a portfolio in order to diversify the risk coming from individual investments. To get a high yield with a good level of diversification, investors usually seek professional advice from portfolio managers. However, the true performance of an optimized portfolio usually depends on the correctness of the estimates of the distribution of future returns, which is often a matter of luck rather than skill. Thus, the optimization models may not be better than randomly selected portfolios. Our aim is to find how the so-called strategy portfolios, i.e., portfolios obtained by some decision optimized for a long-run horizon, perform compared to a benchmark, namely, a random investment, under specific market conditions. For this purpose, we evaluate several portfolio strategies over two periods of crisis: the subprime mortgage crisis and the Covid-19 pandemic, as well as run a moving window analysis over a longer horizon. In each case, the results are compared with the performance of random-weight portfolios. We find that if the strategy is minimization, the portfolios perform well; however, for the maximization of the objectives, the results are rather mixed. Keywords Portfolio optimization· Financial crisis· Random weights· Performance measure· Risk measure Aleš Kresta and Tomáš Tichý have contributed equally to this work. * Tomáš Tichý tomas.tich[email protected] Anlan Wang anlan.w[email protected] Aleš Kresta [email protected] 1 Department ofFinance, VSB–Technical University ofOstrava, Sokolská třída 33, 70200Ostrava, CzechRepublic A.Wang et al. 1 3 17 Page 2 of 27 1 Introduction Since the introduction of the mean–variance portfolio theory by Markowitz (1952), according to which rational investors should always optimize the ratio of expected return and risk level of their investment portfolio, particular researchers as well as practitioners are still trying to beat the rest of the market. Despite their effort, the question how to estimate future returns and their riskiness and determine the optimal weights in the portfolio remains mostly unanswered. If not beating the market completely, one might at least assume that highly skilled portfolio managers should be able to earn higher returns than random investments do. However, many scholars are rather sceptical about that and assume that the effort to find an optimal investment strategy is useless, since even A blindfolded monkey throwing darts at a newspaper’s financial pages could select a portfolio that would do just as well as one carefully selected by experts, as pointed out by Malkiel (1973). A good reason for such a conclusion might be the lack of information about the future performance of firms, the presence of bubbles, or the arrival of various unanticipated shocks, such as the recent subprime mortgage crisis, the Covid-19 pandemic, or the war in the Ukraine. One typical case of evaluating the performance of a strategy portfolio was proposed by DeMiguel etal. (2009). In their empirical study, three performance measures, i.e., the Sharpe ratio, the certainty-equivalent return, and portfolio turnover are implemented to evaluate the performance of the portfolio, based on the rolling-window approach. Specifically, in order to test whether a given performance measure of two distinct strategies is statistically distinguishable, they use Z-tests, a method based on Jobson and Korkie (1981). However, a crucial assumption to apply Z-tests here is the normality of the distribution of the test statistic – unfortunately, such an assumption is commonly violated in practice, as discussed also by Demsar (2006). Since the independence condition is not truly satisfied in a rolling-window out-of-sample validation, DeMiguel etal. (2009) use the Monte Carlo approach to simulate a sufficient dataset of monthly returns, which have i.i.d. normal distribution over time. As a result, the test statistic in their case is asymptotically normally distributed. Similarly, Martínez-Nieto etal. (2021) also compare the performance of the portfolios of 11 diversification strategies under hypothesis testing. However, they point out that the evaluation of the performance measures of strategy portfolios can lead to the rejection of the hypothesis of the normality and equality of the variance. So, to prevent false results, they apply nonparametric Friedman tests instead of the parametric ones. A commonly used approach for statistical validation is to compare the expost returns of a given strategy with a benchmark. It is also common to use the bootstrapping method to make statistical inferences about the hypothesis that the applied strategy does (not) perform better than the benchmark. The simplest case can be the comparison of mean returns. As explained by White (2000), if more strategies are tested, one must also take into account data snooping bias. Moreover, various authors have proposed several tools for hypothesis testing of the 1 3 Evaluation ofstrategy portfolios Page 3 of 27 17 performance, such as the Sharpe ratio discussed in Ledoit and Wolf (2008) or the variance of returns by Ledoit and Wolf (2011). In both papers the bootstrapping method is applied. Besides that, Kim and Lee (2016) also derive a closed-form expression for the probability distribution of the Sharpe ratio of a uniformly distributed random portfolio. In the above-mentioned papers, the measure of the portfolio performance focuses on the Sharpe ratio, which measures the relationship between the excess return (over the risk-free return) and its standard deviation, see Sharpe (1966). However, as discussed by Eling and Schuhmacher (2007), the Sharpe ratio is not always the most appropriate performance measure. On the other hand, as one can see in the relevant literature on the portfolio selection problem, in order to verify the performance of the portfolio optimization models, it is important to select the proper benchmark. For the same dataset, the success of an optimization model can be tested through the comparison of the performance of the obtained strategy portfolios with that of the benchmark portfolio. For example, in DeMiguel etal. (2009), a naive portfolio is employed as the benchmark to measure the performance of several optimization strategies which are designed to reduce the estimation errors on returns. Another commonly applied benchmark in practice is a market index; Solares etal. (2019) point out that the main contraindication to the use of market indexes as benchmarks is that the profitability of portfolios is often compared to that of popular indexes. Hence, most investors require reaching or even outperforming the yields of such indexes over time. Since modern portfolio theory was proposed by Markowitz in the 1950’s, more and more additional constraints have been incorporated into the classical mean–variance model. In such cases, it can be useful to consider the classical mean–variance model as the benchmark for the enhanced models, such as in Fulga (2016), who proposes an approach which takes the loss aversion preferences into account. Among other commonly used benchmarks, we can name, e.g. the mean–CVaR model and the mean–VaR model, see Rankovic etal. (2016) and Lwin etal. (2017), respectively. In this context, our aim is to verify how the so-called strategy portfolios, i.e., portfolios obtained by some decision optimized for a long-run horizon, perform compared to a benchmark, namely, a random investment, under specific market conditions. In order to verify the performance of portfolio optimization models, we select a benchmark similarly to Kim and Lee (2016). That is, we employ uniformly distributed random portfolios, although, in our approach, we do not provide a closed-form expression for the distribution of the risk and performance measures, but rather rely on Monte Carlo simulation so that we can also consider other measure of the performance and the risk. The random-weight portfolio, as it literally says, is defined as the portfolio in which the component assets are invested at random weights. First, we consider two windows with financial market crises as the beginning of the empirical comparison of the strategy and random portfolios. In particular, we use historical daily adjusted closing prices of the DJIA index for (i) 2006–2009, i.e., before and during the subprime mortgage crisis in the US and subsequent global financial crisis, a long-term bear market; (ii) 2018–2021, i.e., the period before and during the A.Wang et al. 1 3 17 Page 4 of 27 Covid-19 Pandemic, which was rather a sudden shock with a relatively quick recovery. Next, we apply a rolling window for (iii) 2006–2021, i.e., the period covering both previously mentioned subperiods, as well as the time in between, to carry out a complex multi-period analysis. This paper is structured as follows. After introducing the topic, in Sect.2 the basis for the portfolio optimization and performance evaluation is reviewed. Next, in Sect.3, three empirical case studies are presented. Finally, we summarize the empirical results and draw some conclusions. 2 Portfolio optimization models andtheir evaluation In this section, we review all the strategies we apply in the empirical part, as well as all the performance measures used for their evaluation. Basically, we assume a finite number of assets, i=1, …,N , for which the time series of returns are available, R i,t= P i,t −P i,t− 1 P i,t−1 . As a benchmark, we use the equally weighted portfolio. Naive strategy (NAIVE) The simplest equally weighted portfolio strategy is the naive one, for which each component has an equal relative weight: w naive i = 1 N . 2.1 Portfolio optimization models When a given investor looks for an optimal portfolio, he or she can use various measures of the reward and the risk, or performance ratios in general, according to particular preferences. However, all standard portfolio optimization strategies share the same framework: they either maximize desired outcomes or minimize unwanted ones or even, optimize their ratio. While the preferred outcome is either the expected return or expected excess return over some benchmark, such as the risk-less rate, the unwanted outcome is usually an appropriate risk measure (or regret measure, when being more general), the definition of which is more complex, especially since it should capture the mutual behaviour of particular items in the portfolio and the diversification effect as well. In this subsection, we summarize the portfolio optimization models considered in our empirical studies. The mean–variance portfolio According to Markowitz (1952), there are only two parameters to be estimated: the portfolio’s expected return and its variance, though to calculate them, one needs to estimate parameters for each feasible investment first, including the Pearson measure of linear dependency, i.e., the correlation matrix (or directly the covariance matrix). Then, each investor should invest on the efficient set (i.e., one parameter cannot be improved, unless the other one is worsened), whereas the optimal trade-off of returns and risk should be detected due to a particular risk attitude. However, for simplicity, we mostly consider either the lower (left) edge of the efficient set, the minimum variance portfolio, (1) 𝜎2 p → min, 1 3 Evaluation ofstrategy portfolios Page 5 of 27 17 with 𝜎 2 p = ∑N i= 1 ∑N j= 1wi⋅𝜎i,j⋅w j , or the upper (right) edge of the efficient set, the maximum expected return portfolio, with 𝔼 (R p )= ∑N i=1 w i ⋅𝔼(R i) , or possibly with PR(x), a so called performance ratio, being a function of the portfolio’s return, its variance, and possibly other factors and obviously with asset weights wi subject to ∑N i=1 w i = 1 and wi≥0 , with i=1, …,N . Notwithstanding, the remaining question to be answered is how these parameters should be estimated. Maximum expected return/Minimum variance with historical sample estimation (XRHS, NVHS) The simplest approach to the estimation of the parameters is to directly use the historical observations in the sample, without any extra adjustments. And so, the expected return is 𝔼 (Ri)= 1 K∑K k=1 Ri, k , the standard deviation is 𝜎 i= � 1 K−1∑ K k=1 � Ri,k−𝔼(Ri) � 2 , and the covariace is 𝜎 i,j= 1 K−1∑K k=1 �Ri,k−𝔼(Ri) � ⋅( R j,k −𝔼(R j ) ) . Minimum variance with Bayes–Stein shrinkage estimation (NVBS) In order to reduce the estimation errors, we might make a subjective (a priori) assumption about the shape of the asset’s return distribution. Thus, the resulting (a posteriori) assumption of the shape of the probability distribution is then a combination of the a priori assumption and the probability distribution of the observed sample. Here, we apply the shrinkage suggested by Jorion (1986) in the Bayesian portfolio selection problem. Thus, the expected return of asset i under the Bayes–Stein estimation is and the shrinkage factor is Next, setting the precision of the shrinkage factor 𝜍 = K𝜉 1 −𝜉 , we can also reformulate the historical covariance matrix C into its Bayes–Stein estimation: Maximum expected return/Minimum variance with fuzzy estimation (XRFU, NVFU) In order to handle the uncertainty about the distribution of asset returns, Tanaka etal. (2000) propose using a fuzzy probability model employing fuzzy set theory. In particular, a grade of possibility (denoted as hk ) is introduced to reflect the degree of similarity between the future state of the asset return and the same asset’s kth historical return: h k=0.1 +0.3 k−1 K−1, where K is the length of the estimation period. As a result, we can estimate the fuzzy-weighted expected return of the ith asset: (2) 𝔼(Rp) → max, (3) PR (Rp,𝜎 2 p ,⋅)→ max, 𝔼( RBS i) =(1−𝜉)𝔼(  R i )+𝜉𝔼(  R ) 𝜉 = N+2 N+2+K(𝔼(R i )−𝔼(R))TC−1(𝔼(R i )−𝔼(R)) . C BS = C⋅(1+1 K+𝜍)+𝜍 K(K+1+𝜍) 1n1 T n 1T n � C−11 n . A.Wang et al. 1 3 17 Page 6 of 27 and the fuzzy weighted items of the covariance matrix C : where Ri,k is the kth historical observation of the ith asset’s returns. Minimization of alternative risk measures Besides minimizing the variance of the portfolio, one can minimize various alternative risk measures. Minimum absolute deviation (NMAD) Within this approach, following (Konno and Yamazaki 1991), the variance of the portfolio is replaced by the mean absolute deviation, which turns the original quadratic optimization model into a linear one: subject to standard conditions.1 Minimum conditional value at risk (NCVAR) Assuming that the agents are risk averse, they treat downside risk negatively, but upside risk positively. In other words, it can be useful to replace the variance by a measure which considers only unfavourable deviations. An approach penalizing the downside risk is due to Rockafellar and Uryasev (2000), who propose using, for portfolio optimization purposes, an function which is an approximation of a so called conditional value at risk measure: where w is the vector of weights, VaR𝛼(w) = min{𝛾∶ ℙ [f(w,x) ≤ 𝛾] ≥ 𝛼}. If we replace uk = [ −wTR k −VaR 𝛼 (w) ]+ , we can formulate the optimization problem as follows: subject to uk ≥ 0 and wTRk+VaR𝛼(w) + uk ≥ 0 . 𝔼 �RFU i�= ∑K k=1hkRi,k ∑ K k=1 hk 𝜎 FU i,j=∑K k=1hk�𝔼(RFU i)−Ri,k� � 𝔼(RFU j)−Rj,k � ∑ K k=1 h k 1 T T ∑ t=1 N ∑ i=1| | Ri,t−𝔼(Ri) | | wi→ min, CVaR 𝛼(w) = VaR𝛼(w) + 1 K(1−𝛼) K ∑ k=1 [ −wTRk−VaR𝛼(w) ] + , VaR 𝛼(w) + 1 K(1−𝛼) K ∑ k=1 uk→ min, 1 For this purpose we use a toolbox of Matlab, https:// www. mathw orks. com/ help/ finan ce/ meanabsol utedevia tionportf oliooptim izati on. html?s_ tid= CRUX_ topnav. 1 3 Evaluation ofstrategy portfolios Page 7 of 27 17 Maximization of performance measures When maximizing a performance measure in (3), one can consider various combinations of reward, risk, and even dependency measures, see, e.g. Ortobelli and Tichy (2015). Maximum Sharpe ratio (XSR) The Sharpe ratio (SR) is a reward-to-variability ratio that is used to measure the adjusted return of a portfolio, and is defined as the difference between the portfolio return over the risk-free rate divided by the standard deviation of the portfolio returns: SR (Rp)= 𝔼(R p −R B ) 𝜎(R p ) , where RB is the benchmark return for a given horizon, such as a risk-free rate. Maximum Rachev ratio (XRR) The Rachev ratio (RR), see Biglova etal. (2004), is a performance ratio which uses the ratio between the CVaR of the opposite of the excess return at a given confidence level and the CVaR of the excess return at another confidence level: RR (Rp,𝛼,𝛽)= CVaR 𝛽 (R B −R p ) CVaR 𝛼 (R p −R B) , where RB is the return of the benchmark. Maximum STARR ratio (XSTARR) The STARR ratio was introduced by Martin et al. (2003) as a particular case of the Rachev ratio with 𝛽=1 STARR (Rp,𝛼)= 𝔼(R p −R B ) CVaR 𝛼 (R p −R B) . Maximum Sortino–Satchell ratio (XSS) Another attempt to stress downside deviation was proposed by Sortino and Satchell (2001): SS (Rp,u)= 𝔼(R p −R B ) ( 𝔼(R B −R p )u +) (1∕u) , where u is the order of the lower partial moment. Maximum Farinelli–Tibiletti ratio (XFT) The Farinelli–Tibiletti ratio connects partial moments of different orders so that the portfolio reward is measured by an upper partial moment rather than the expected value of the portfolio returns: FT (Rp,𝛾,𝛿)= (𝔼(Rp−RB) 𝛾 +) 1∕𝛾 ( 𝔼(R B −R p )𝛿 +) (1∕𝛿) , where 𝛾 ≥ 1 and 𝛿≤1 are the orders of the partial moments which reflect an investor’s different attitudes toward outperformance and underperformance. Specifically, Farinelli and Tibitelli (2008) propose using two different values, 𝛾=2, 𝛿=0.5 and 𝛾=0.5, 𝛿=2 . 2.2 Additional performance measures Jensen’s alpha (JA) Jensen’s alpha was first used by Jensen (1968) as a measure for the evaluation of the management of mutual funds in the 60’s. In essence, it is an ex-post alpha that measures the excess return of a portfolio over the theoretical expected return predicted by the capital asset pricing model (CAPM): 𝛼J=Rp−[Rf+𝛽p(Rm−Rf)] . Treynor ratio (TR) The Treynor ratio is defined as the excess return over the riskfree rate per unit of risk. However, as opposed to other ratios, it also uses the assumptions of CAPM, in particular, the market risk measured by 𝛽p : TR (Rp)= 𝔼(R p −R f ) 𝛽 p . A.Wang et al. 1 3 17 Page 8 of 27 Calmar ratio (CR) Another alternative is to replace the standard deviation by the maximum drawdown (MAXDD). Such a ratio is called Calmar (California managed accounts reports): CR (Rp,T)= 𝔼(R p −R f ) MAXDD(T) . 2.3 Random‑weight portfolios andtheir usage forranking thestrategy portfolios In a random-weight portfolio, just as it says literally, the weights of the assets in the portfolio are generated randomly, most commonly by Monte Carlo simulation. For generating N random weights, we choose Algorithm2 presented in Tervonen and Lahdelma (2007) who actually refer to David (1970). In this approach, we generate y∈[ 0, 1 ]N−1 as N−1 real numbers uniformly distributed in the interval [0, 1] and sort these numbers so that 0 ≤ y1 ≤ … ≤ yN−1 . Then, after hypothetically adding the values zero and one to them, we can obtain the vector w of random weights as the differences of the successive pairs of numbers in y: w=(y1,y2−y1,y3−y2,…,yN−1−yN−2,1−yN−1) , where ∑N i=1 w i = 1 for each generated portfolio. Knowing historical data, we can calculate the performance of a portfolio in terms of the above-mentioned measures of risk and performance for each of the randomly generated portfolios. After obtaining a sufficiently large number of random-weight portfolios, and calculating the measures of their risk and performance, we can get a good estimate of their distributions. This is specifically useful in the out-of-sample period, in which we can evaluate the performance of a selected strategy by calculating its ranking among these randomly generated portfolios. The ranking is always calculated on sorted values from the best (i.e. the lowest values of risk measures and the highest values of the performance measures) to the worst (i.e. the highest values of risk measures and lowest values of performance measures). In particular, a ranking of 1 would mean that only 1% of randomly generated portfolios perform better, and 99% of randomly generated portfolios perform worse. The obtained rankings can be also linked to the hypothesis testing suggested by Kim and Lee (2016) for the Sharpe ratio. Simply speaking, the ranking among the 5% of the best-performing random portfolios would mean the possibility of rejecting the null hypothesis that the strategy selects the stock weights randomly. However, when doing multiple tests, one must also be careful about the possibility of data snooping bias, see, e.g. White (2000), and some correction adjustments should be used, such as, e.g. the Bonferroni or Šidák correction. 3 Empirical analysis ofselected strategy portfolios In this section, we empirically evaluate selected portfolio optimization models introduced above under specific settings, see Table1. First, we evaluate their historical performance during the subprime mortgage crisis (2007–2009). Next, we study the consequences of the Covid-19 pandemic period (2018–2021). These two case studies demonstrate the procedure we later apply on the basis of a rolling window. In each case study, we use historical daily adjusted closing prices of stocks included in 1 3 Evaluation ofstrategy portfolios Page 9 of 27 17 the DJIA index (we consider only the 27 stocks included in DJIA as of January 3, 2006, as the stocks of General Motors Corporation, Hewlett-Packard Company, and United Technologies Corporation were excluded due to missing data). The evolution of the DJIA index during the relevant period is depicted in Fig.1. The vertical line splits the data into in-sample and out-of-sample (Cases 1 & 2) or the date on which the out-of-sample rolling windows start (bottom). For the rolling window approach, we employ the longer in-sample period for the estimation and optimization. Table 1 Summary of analysed periods Case study In-sample Out-of-sample Case 1 3.1.2006–10.8.2007 13.8.2007–2.3.2009 Case 2 6.3.2018–10.10.2019 11.10.2019–30.4.2021 Case 3 3.1.2006–21.12.2009 22.12.2009–17.12.2010 rolling-window approach 4.1.2006–22.12.2009 23.12.2009–18.12.2010 2607 overlapping ⋮ ⋮ 1000/250-day periods 13.5.2016–4.5.2020 5.5.2020–30.4.2021 5000,00 10000,00 15000,00 20000,00 25000,00 30000,00 35000,00 Fig. 1 Evolution of the DJIA (Case 1–top left, Case 2 – top right, Case 3–bottom) [Source: own elaboration of data from finance.yahoo.com] A.Wang et al. 1 3 17 Page 16 of 27 ratio. Changing the relation of these parameters, see, e.g. strategies FT05 and FT2, in which the values of the parameters are reverted, made one perform better in risk measures and the other in performance measures. 3.4 Rolling window approach (case study 3) The previous two cases were based on the single-period dataset only. However, one might wonder whether the results are stable and what would happen if the period is changed. Therefore, here, we evaluate the strategy portfolios by applying a rolling window analysis. In this way, we can increase the robustness of the results. The chosen dataset covers all the data from January 3rd, 2006 to April 30th, 2021. We always take 1,000 days as the in-sample period and 250 days as the out-of-sample period. Then, we keep moving the beginning of the window by one day until the end of the dataset. In all, we obtain 2,607 overlapping periods, see Table1. Firstly, we analyse the obtained weights. In Table 6, we summarize the average number of stocks held within the strategy with a weight higher than 0.1%, the Table 4 Out-of-sample ranking in risk measures (case study 2) STD MAD CVAR1 CVAR5 CVAR10 ML LPM MAXDD Average NVHS 0 0 0 0 0 0 0 0 0 NVBS 0 0 0 0 0 0 0 0 0 NVFU 0 0 0 0 0 0 0 0 0 NMAD 0 0 0 0 0 0 0 0 0 NCVAR1 0 1 0 1 0 9 0 0 1 NCVAR5 0 0 0 0 0 0 0 0 0 NCAV10 0 0 0 0 0 0 0 0 0 XRHS 100 100 82 99 100 96 98 88 95 XRFU 86 39 82 74 48 100 67 15 64 XSR 0 0 0 0 0 3 0 0 0 XRR5 97 98 75 99 98 100 100 1 84 XRR50 0 0 0 0 0 25 0 0 3 XRR75 0 0 0 0 0 0 0 0 0 XSTARR5 0 0 0 0 0 5 0 0 1 XSS08 0 0 0 0 0 12 0 0 2 XSS2 1 1 1 0 0 28 0 0 4 XSS25 1 1 1 0 0 22 0 0 3 XFT05 82 62 75 90 78 99 85 18 74 XFT2 0 0 0 0 0 9 0 0 1 XOMG 0 0 0 0 0 17 0 0 2 XFT10 85 40 100 88 72 100 98 62 81 XFT001 6 4 0 12 5 0 1 0 4 NAIVE 44 41 50 45 44 48 49 48 46 Average 22 17 20 22 19 29 22 10 1 3 Evaluation ofstrategy portfolios Page 17 of 27 17 Table 5 Out-of-sample ranking in performance measures (case study 2) MR JA TR SR SS08 SOR SS25 FT05 FT2 OMG FT10 FT001 RR5 RR50 RR75 STARR5 CR Average NVHS 100 98 100 100 100 100 100 99 81 95 28 100 100 100 97 100 100 94 NVBS 100 98 100 100 100 100 100 99 81 95 28 100 100 100 97 100 100 94 NVFU 100 98 100 100 100 100 100 99 64 90 15 100 100 100 97 100 100 92 NMAD 100 97 100 100 100 100 100 88 42 67 5 99 100 100 93 100 100 88 NCVAR1 96 44 30 60 78 76 77 100 47 85 46 100 100 71 20 77 67 69 NCVAR5 100 97 100 100 100 100 100 100 82 99 42 100 100 100 99 100 100 95 NCAV10 100 90 96 100 100 100 100 100 83 100 39 100 100 100 91 100 100 94 XRHS 0 0 0 18 8 1 0 0 0 0 0 69 0 26 99 1 0 13 XRFU 96 90 91 99 95 98 99 0 83 27 100 0 15 96 94 98 90 75 XSR 88 29 16 34 29 31 35 42 8 9 31 78 56 38 43 30 35 37 XRR5 0 0 0 0 0 0 0 100 73 100 80 94 61 0 0 0 0 30 XRR50 81 24 12 19 17 24 30 64 58 49 91 69 83 16 9 18 21 40 XRR75 100 88 93 98 99 99 99 90 9 25 19 99 100 99 94 99 98 83 XSTARR5 91 29 15 33 36 34 40 79 34 40 87 83 50 36 31 26 29 45 XSS08 92 33 18 33 27 39 47 57 25 24 58 75 99 33 8 36 44 44 XSS2 56 11 4 9 4 7 10 14 18 9 65 8 13 7 12 6 7 15 XSS25 56 12 5 11 7 9 12 34 38 26 86 29 3 9 19 7 6 22 XFT05 2 3 4 3 1 3 4 20 78 64 78 4 38 1 1 4 0 18 XFT2 94 50 41 56 42 50 57 16 17 12 57 13 8 52 51 48 33 41 XOMG 58 13 5 8 6 8 11 51 44 40 71 69 21 6 8 6 5 25 XFT10 9 14 16 19 10 37 51 46 100 99 100 0 96 3 1 19 11 37 XFT001 100 71 67 99 99 99 98 30 52 73 33 0 15 99 91 99 41 69 NAIVE 49 50 50 47 47 49 50 52 76 70 78 36 52 44 45 48 49 52 Average 72 50 46 54 52 55 57 60 52 56 54 62 61 54 52 53 49 A.Wang et al. 1 3 17 Page 18 of 27 standard deviation of the number of stocks, and most importantly the average change of the weights in the portfolio between two successive overlapping periods. The higher the average number of stocks, the more diversified the portfolio suggested by the strategy, on average. From Table6 we can see that the most diversified portfolio is suggested by the equally weighted strategy, which invests in all stocks equally. A relatively high number of stocks, and thus a good diversification, is also suggested by the strategies minimizing the variance, MAD, and maximizing the FT ratio (for 𝛾=10 , 𝛿=0.01 and 𝛾=0.5 , 𝛿=2 ). These strategies hold on average 10 stocks out of 27 with a standard deviation of around 3. The smallest number of stocks held in the portfolio are obviously in strategies maximizing the expected return (XRHS, XRFU), which always selected just one stock. For these strategies, the average change in portfolio’s composition is around 5%, which means that these strategies select a different stock on average every 20th day. The historical sample estimation is however more stable than the fuzzy approach (4.8% versus 5.8%). Table 6 The descriptive statistics of the portfolio weights Portfolio strategy Average # stocks with weight>0.1% Standard deviation of # stocks with weight>0.1% Average change in portfolio weights (%) NVHS 10.997 3.6716 0.5 NVBS 10.997 3.6716 0.5 NVFU 11.263 3.9096 0.6 NMAD 12.107 4.1531 1.1 NCVAR1 6.729 2.4929 0.6 NCVAR5 7.821 2.7119 0.8 NCVAR10 8.916 3.2326 0.7 XRHS 1.000 0.0000 4.8 XRFU 1.000 0.0000 5.8 XSR 5.140 1.8058 2.6 XRR5 2.727 1.3154 21.1 XRR50 6.105 2.0727 3.2 XRR75 6.535 1.8887 3.5 XSTARR5 4.498 1.5389 3.5 XSS08 7.475 2.5840 9.0 XSS2 5.255 1.8627 5.3 XSS25 4.809 1.7453 4.4 XFT05 9.515 6.2713 39.8 XFT2 8.573 3.2309 45.5 XOMG 8.530 5.4771 23.5 XFT10 10.512 3.2805 63.4 XFT001 3.374 5.0637 69.3 NAIVE 27.000 0.0000 0.0 Average 7.864 2.6948 13.5 1 3 Evaluation ofstrategy portfolios Page 19 of 27 17 The portfolios which have the largest changes in their structure are obtained from the strategies maximizing the FT ratios, followed by the XRR5 strategy. For strategies maximizing the FT ratios, the change is enormous, as moving the in-sample period by a single day causes a difference of 40%–70% in the portfolio’s composition. The exception is the special case of the max-FT strategy with 𝛾=1 , 𝛿=1 , i.e. the maximum-Omega-index strategy. The average change of its composition is comparable to that of the maximum-RR strategy with 𝛼=𝛽=5% (around 20%). The other strategies have much smaller changes, typically smaller than 5%. The smallest changes, and thus also the greatest stability, are in the strategies minimizing the variance and CVaR (below 1%). We can also notice a positive relationship between the average number of stocks included in the portfolio and its standard deviation, while we cannot observe the same for the average change in portfolio weights. We consider the average change in portfolio weights as more informative, as the stocks with a low weight (but still larger than 0.1% ) can have a high impact on the standard deviation of the number of stocks held within the strategy. Next, we simulate 30,000 random-weight portfolios and for each of 2,607 overlapping 250-day investment periods we calculate the risk and performance measures for these strategies and random portfolios. In each overlapping period, we compare the performance of the obtained strategy portfolios with that of the random-weight portfolios. Specifically, we make the comparisons by calculating the rankings of the strategy portfolios. The ranking is always calculated on sorted values from the best (i.e. the lowest values of the risk measures and the highest values of the performance measures) to the worst (i.e. the highest values of the risk measures and the lowest values of the performance measures). Thus, e.g. a ranking of 1 means that only 1% (i.e. 300) of the randomly generated portfolios performed better, and 99% (i.e. 29,700) of them performed worse. It must be mentioned that the out-of-sample periods are overlapping, and thus the observed rankings are not independent and the whole period does not have the same weight in the analysis. Out of the 2,857 days in the out-of-sample periods, the first and the last 250 days have lower weights in the analysis as these participate in fewer rolling windows. Nevertheless, our idea was to improve the robustness of the results and answer the question of what would happen on average during the analysed period of December 2009 to April 2021. In Tables7 and 8 we summarize the values of the mean relative ranking of the obtained strategy portfolios considering the risk measures and the performance measures, respectively. Based on the results in Table7, one can see the following. Although we could assume the mean relative ranking of the equally-weighted (NAIVE) strategy to be around 50%, we can see that the strategy performed better than random-weight portfolios by ranking as 33th–45th out of 100, depending on the chosen risk measure. As for the portfolios obtained from strategies minimizing the variance (NVHS, NVBS and NVFU), MAD (NMAD) and CVaR (NCVAR1, NCVAR5 and NCVAR10), their risk is lowered in the out-of-sample period and we see that for these strategies, the values of their average relative rankings are very low, except for the minimum-CVaR strategies, which generally performed worse, but still within 10% of A.Wang et al. 1 3 17 Page 20 of 27 the best randomly generated portfolios. We can also see that the maximum one-day loss (ML) and maximum drawdown (MDD) are stricter measures, as the risk minimization strategies ranked worse in these criteria compared to other risk measures, however, still within the best 20% (ML) and 30% (MDD) of the random portfolios. Maximization strategies, on the other hand, rank either around the middle (ranking between 40 and 60) or very poorly (ranking above 80). The worst strategies with respect to the risk measures are the strategies maximizing the expected return (XRHS performing slightly better than XRFU), maximizing the Rachev ratio with 𝛼=5% , and maximizing some of FT ratios. Not surprisingly, the strategies maximizing the expected return perform poorly in the out-of-sample period in terms of the risk measures. The poor performance of XRR5 is probably due to the low values of the parameters, as when the values of parameters are increased we can see that the the XRR strategy portfolios perform better, see the performances of XRR50 and XRR75. A relatively good performance can be seen for maximizing the SS ratio with u=0.8 and the Rachev ratio with 𝛼=𝛽=50% and maximizing the Omega Table 7 Rolling-window out-of-sample mean ranking in risk measures STD MAD CVAR1 CVAR5 CVAR10 ML LPM MAXDD Average NVHS 0 1 4 1 1 15 1 28 6 NVBS 0 1 4 1 1 15 1 28 6 NVFU 0 1 5 2 2 16 2 27 7 NMAD 0 0 4 1 1 15 1 29 6 NCVAR1 5 6 9 3 4 18 6 28 10 NCVAR5 2 5 7 1 2 17 1 30 8 NCAV10 1 3 8 1 2 19 2 31 8 XRHS 89 89 87 87 88 83 90 70 85 XRFU 97 97 94 95 96 92 97 80 94 XSR 53 53 53 48 46 56 53 47 51 XRR5 91 91 86 88 87 81 90 78 86 XRR50 49 49 48 43 42 52 49 44 47 XRR75 52 50 49 47 45 51 50 51 49 XSTARR5 53 55 52 47 46 57 53 48 51 XSS08 45 44 46 41 40 51 46 42 44 XSS2 52 53 50 46 46 56 52 44 50 XSS25 54 55 51 48 48 57 53 44 51 XFT05 88 88 84 86 86 81 88 62 83 XFT2 60 61 57 57 55 55 60 56 58 XOMG 47 48 46 43 42 53 47 40 46 XFT10 88 89 82 85 85 76 86 81 84 XFT001 88 89 85 85 85 85 86 81 85 NAIVE 36 33 45 41 39 45 41 43 40 Average 46 46 46 43 43 50 46 48 1 3 Evaluation ofstrategy portfolios Page 21 of 27 17 Table 8 Rolling-window out-of-sample mean ranking in performance measures MR JA TR SR SS08 SOR SS25 FT05 FT2 OMG FT10 FT001 RR5 RR50 RR75 STARR5 CR Average NVHS 57 37 33 43 44 43 43 44 38 42 54 47 34 43 44 42 44 43 NVBS 57 37 33 43 44 43 43 44 38 42 54 47 34 43 44 42 44 43 NVFU 58 39 35 45 47 45 45 48 40 44 54 53 39 46 46 44 46 46 NMAD 57 38 33 43 44 43 43 45 45 48 53 45 36 43 44 42 45 44 NCVAR1 54 40 36 44 45 43 43 48 37 45 47 49 24 44 46 42 42 43 NCVAR5 54 36 30 41 42 41 41 47 41 46 56 56 31 42 42 40 42 43 NCAV10 55 36 31 42 44 43 42 51 44 51 59 54 33 43 43 41 44 44 XRHS 33 31 31 50 50 49 48 51 29 40 24 62 33 50 54 48 46 43 XRFU 44 44 44 60 59 59 58 54 33 43 27 66 36 60 63 59 57 51 XSR 40 28 26 40 42 40 40 54 48 53 47 57 39 40 39 39 42 42 XRR5 55 56 57 75 75 74 74 61 43 50 49 57 35 75 76 74 71 62 XRR50 40 28 25 38 40 39 39 54 49 54 47 55 37 39 37 38 41 41 XRR75 36 25 23 34 36 35 35 49 45 49 48 50 42 34 33 33 38 38 XSTARR5 42 28 25 42 44 42 42 56 49 54 47 61 41 42 40 41 44 44 XSS08 38 28 25 37 39 37 37 52 49 53 46 51 37 37 35 36 40 40 XSS2 39 27 25 39 41 39 39 53 49 53 47 57 38 39 38 37 41 41 XSS25 39 27 25 39 41 39 39 53 49 54 46 58 38 39 39 38 40 41 XFT05 25 27 30 47 48 47 46 44 37 42 43 46 29 47 48 46 43 41 XFT2 50 42 40 52 54 52 51 61 48 57 41 64 41 52 51 51 51 50 XOMG 38 27 25 37 39 37 37 50 46 50 43 52 37 38 36 36 39 39 XFT10 48 48 48 58 58 57 57 59 39 49 37 60 40 58 60 57 57 52 XFT001 44 47 48 68 68 66 66 47 37 43 56 34 30 68 71 66 65 54 NAIVE 50 51 51 46 44 46 47 45 61 53 64 37 60 45 44 48 47 49 Average 46 36 34 46 47 46 46 51 43 48 47 53 37 46 47 45 46 A.Wang et al. 1 3 17 Page 22 of 27 index with a ranking lower than 50 for all the measures except the maximum oneday loss. In Table8 we show the rankings in terms of the performance measures. As is obvious at first sight, there is no strategy with a superior ranking, as there was in the case of the risk measures. Still, there are some interesting results. First, the equally weighted (naive) strategy performed literally on average with an average ranking of around 50. Second, the risk minimization strategies performed well, usually ranking around 43–46, except for the mean return (MR), the FT10 ratio, and the FT001 ratio. Thus, these strategies do not provide a high return, but the decrease in their riskiness improves the values of the performance ratios. The maximization strategies performed generally better than risk minimization strategies, with rankings of around 38–43, except for XRR5, XRFU and some of the FT strategies. We can see that for maximizing the expected return, the historical sample works better than fuzzy estimation. As the best strategies with minimal average rankings, we can identify the strategies maximizing the Rachev ratio with 𝛼=𝛽=75% and the Omega index. From the results it is possible to see that the higher values of 𝛼 and 𝛽 work better for the Rachev ratio than the lower one (5%). It should also be noted that the Omega index is similar to the Rachev ratio with 𝛼=𝛽=50% , whose rankings in risk and performance measures were very similar. The latter is the ratio of the mean of above-median returns to the mean of belowmedian losses, while the first is the ratio of the mean of positive returns to the mean of positive losses, both adjusted for minimum acceptable return or benchmark ( RB ). However, comparing the portfolio’s composition and its stability, see Table6, the XRR50 strategy is on average less diversified but more stable over time. We also analysed the evolution of the rankings over time. Due to limitations of space, in Figs.4 and 5 we only illustrate the situation for some selected portfolio strategies and risk and performance measures. However, we comment on all the findings. In principle, the blue areas in the figures are constructed by the different percentiles of the values of the corresponding risk and performance measures for all the 30,000 generated random-weight portfolios. The darkest blue area represents 50 % of the performances of random portfolios around the median (i.e. the boundaries are the 25th and 75th percentiles), the lightest blue area represents 99 % of the performances of random portfolios around the median (i.e. the boundaries are the 0.5th and 99.5th percentiles). The medium dark blue area represents 95 % of the performances of random portfolios around the median (i.e. the boundaries are the 2.5th and 97.5th percentiles). When the value of the risk measure for strategy is below the blue area, or when the value of the performance measure for strategy is above the blue area, we can consider the strategy portfolio as lowering the risk or improving the performance with an equivalent ranking of 1 (lightest blue area), 3 (medium dark area) and 25 (darkest area). For risk measures, in Fig.4 we can see that the equally weighted strategy portfolio (NAIVE) seems to be usually in the middle of the dark-blue area. However, from Table7 we know it performed better. The values of the out-of-sample standard deviation of returns are similar for all minimum-variance strategies (NVHS, NVBS and NVFU). Similar results can also be obtained for the other risk measures, which are illustrated by the CVaR and maximum drawdown. Generally, all the obtained 1 3 Evaluation ofstrategy portfolios Page 23 of 27 17 0,05 0,1 0,15 0,2 0,25 0,3 0,35 0,4 0,45 0,5 Time 07.04.2010 19.07.2010 27.10.2010 08.02.2011 20.05.2011 31.08.2011 12.12.2011 26.03.2012 06.07.2012 16.10.2012 31.01.2013 14.05.2013 23.08.2013 04.12.2013 19.03.2014 30.06.2014 09.10.2014 22.01.2015 05.05.2015 14.08.2015 24.11.2015 09.03.2016 20.06.2016 29.09.2016 11.01.2017 25.04.2017 04.08.2017 14.11.2017 28.02.2018 11.06.2018 20.09.2018 03.01.2019 16.04.2019 29.07.2019 06.11.2019 20.02.2020 STD 50,0%95,0%99,0%NVHNVBS NVFNAIVE 0 0,01 0,02 0,03 0,04 0,05 0,06 0,07 0,08 Time 15.04.2010 04.08.2010 22.11.2010 15.03.2011 05.07.2011 21.10.2011 13.02.2012 04.06.2012 21.09.2012 15.01.2013 07.05.2013 26.08.2013 13.12.2013 07.04.2014 28.07.2014 13.11.2014 09.03.2015 26.06.2015 15.10.2015 05.02.2016 26.05.2016 15.09.2016 05.01.2017 27.04.2017 16.08.2017 05.12.2017 28.03.2018 18.07.2018 05.11.2018 28.02.2019 19.06.2019 08.10.2019 29.01.2020 CVAR5 50,0% 95,0% 99,0% NCVAR5 NAIVE 0 0,1 0,2 0,3 0,4 0,5 0,6 Time 13.04.2010 29.07.2010 12.11.2010 03.03.2011 20.06.2011 05.10.2011 24.01.2012 10.05.2012 27.08.2012 14.12.2012 05.04.2013 23.07.2013 06.11.2013 26.02.2014 13.06.2014 30.09.2014 16.01.2015 06.05.2015 21.08.2015 08.12.2015 29.03.2016 14.07.2016 28.10.2016 16.02.2017 06.06.2017 21.09.2017 09.01.2018 27.04.2018 14.08.2018 29.11.2018 21.03.2019 09.07.2019 23.10.2019 11.02.2020 MAXDD 50,0% 95,0% 99,0% NVHNMAD NCVAR5 NAIVE Fig. 4 Evolution of the ranking of the risk-minimization strategies A.Wang et al. 1 3 17 Page 24 of 27 -1 -0,5 0 0,5 1 1,5 2 2,5 3 3,5 4 Time 09.04.2010 23.07.2010 04.11.2010 18.02.2011 06.06.2011 19.09.2011 03.01.2012 18.04.2012 01.08.2012 15.11.2012 05.03.2013 18.06.2013 01.10.2013 15.01.2014 01.05.2014 14.08.2014 26.11.2014 16.03.2015 29.06.2015 12.10.2015 27.01.2016 11.05.2016 24.08.2016 07.12.2016 24.03.2017 10.07.2017 20.10.2017 06.02.2018 22.05.2018 05.09.2018 19.12.2018 05.04.2019 22.07.2019 01.11.2019 19.02.2020 SR 50,0%95,0%99,0%XSRNAIVE -0,02 0 0,02 0,04 0,06 0,08 0,1 0,12 Time 19.04.2010 10.08.2010 01.12.2010 25.03.2011 19.07.2011 08.11.2011 05.03.2012 26.06.2012 17.10.2012 13.02.2013 07.06.2013 30.09.2013 23.01.2014 16.05.2014 09.09.2014 31.12.2014 27.04.2015 18.08.2015 09.12.2015 05.04.2016 27.07.2016 16.11.2016 14.03.2017 06.07.2017 26.10.2017 21.02.2018 14.06.2018 05.10.2018 31.01.2019 24.05.2019 17.09.2019 09.01.2020 STARR5 50,0%95,0%99,0%XSTARR5NAIVE 0,7 0,8 0,9 1 1,1 1,2 1,3 1,4 1,5 1,6 Time 13.04.2010 29.07.2010 12.11.2010 03.03.2011 20.06.2011 05.10.2011 24.01.2012 10.05.2012 27.08.2012 14.12.2012 05.04.2013 23.07.2013 06.11.2013 26.02.2014 13.06.2014 30.09.2014 16.01.2015 06.05.2015 21.08.2015 08.12.2015 29.03.2016 14.07.2016 28.10.2016 16.02.2017 06.06.2017 21.09.2017 09.01.2018 27.04.2018 14.08.2018 29.11.2018 21.03.2019 09.07.2019 23.10.2019 11.02.2020 OMG 50,0%95,0% 99,0% XOMG NAIVE Fig. 5 Evolution of the ranking of the performance-maximization strategies 1 3 Evaluation ofstrategy portfolios Page 25 of 27 17 minimum-risk strategy portfolios lower the corresponding risk measures out-ofsample in almost the whole analysed period. This corresponds to the average rankings shown in Table7. The evolutions of the performance measures for strategy portfolios are usually blurry, as the values of the performance ratios change from day to day, as illustrated by the maximum Omega index strategy in Fig.5. In the figure, we have also chosen two relatively stable strategies with a low average change of portfolio weights: strategies maximizing the Sharpe and Rachev ratios, see Table6. As one can see, maximizing the performance ratio in the in-sample period does not always mean improving the same performance ratio in the out-of-sample period. For example, the strategy portfolio maximizing the Sharpe ratio in the in-sample period leads to a good performance from July 2010 to July 2011, then a poor performance from July, 2011 to October 2013, then a superior performance from October 2013 to July 2015, and so on. In fact, most of the strategies have a similar evolution, with overperformance and underperformance at the same periods and thus react to the changes in the markets in a similar way. 4 Conclusion Investors usually seek professional advice from portfolio managers in order to get a high yield with a good level of diversification. Ther key idea is that the optimization models should provide an advantage over just randomly selected portfolios. In this context, our aim has been to determine how the so-called strategy portfolios, i.e. portfolios obtained by specific optimization rules, perform and whether they can be regarded as performing better than the non-random ones. For this purpose, we have evaluated 23 portfolio strategies over two crisis periods: the subprime mortgage crisis and the Covid-19 pandemic, as well as runing a moving window analysis over a longer horizon. The results of the first case study show that the strategies minimizing the classical risk measures in-sample (NVHS, NVBS, NVFU, NMAD) performed well in the out-of-sample period in terms of the applied risk measures, but in terms of the performance measures, these strategy portfolios performed randomly. However, minimizing the CVaR in the in-sample period resulted in low risk as well as good performance in the out-of-sample period. Observing the results of the second case study, we found the following differences: The minimizing strategies work in about the same way or slightly better in terms of risk measures, while extremely poor in terms of performance measures. The maximizing strategies also show about the same risk, while much worse performance. In the third case we applied a rolling window analysis based on a larger dataset. Concerning the applied risk measures, we found that the NAIVE portfolio performed better than the median of random-weight portfolios, which exceeded our expectations, as we expected the performance to be around the median. Moreover, the minimizing strategies worked very well, while the maximizing strategies had either poor results or were close to the median. As concerns the performance measures, most strategies were very close to the median values.