Actionable Forecasting and Activity Monitoring: Applications to Financial Trading
Full text
Faculty of Science of University of Porto Actionable Forecasting and Activity Monitoring: applications to financial trading Luís Baía Masters in Engineering Mathematics Supervisor: Luís Torgo August 26, 2015
Actionable Forecasting and Activity Monitoring: applications to financial trading Luís Baía Masters in Engineering Mathematics August 26, 2015
Abstract The thesis addresses a particular class of decision making problems. We name these tasks Actionable Forecasting problems. The main distinguishing feature of this type of decision problems is the fact that decisions are to be taken based on predictions of a numeric variable. Examples of such tasks include for instance the medical diagnosis of a patient based on predictions of some numerical indicator, or deciding which trading action should be taken depending on the prediction of the future evolution of the market prices. We study and compare two different alternative ways of addressing these decision problems: (i) using standard regression models to forecast the numeric variable and on a second step transform these numeric predictions into a decision according to some pre-defined and deterministic decision rules; and (ii) use models that directly forecast the right decision using classification models thus ignoring the intermediate numeric forecasting task. The objective of this study is to determine if both strategies provide identical results or if there is any particular advantage worth being considered that may distinguish each alternative. We also consider some potential limitations of each alternative, where we consider solutions such as the usage of cost-benefit matrices as well as re-sampling algorithms. We carry out two major studies to compare both alternatives to solve actionable forecasting tasks: (i) one involving a large set of generic and non-temporal tasks; (ii) and a second involving financial trading problems that use price time series with the goal of deciding whether to invest or not in the market with short and long positions. Decision making in the context of financial trading is a very relevant problem with very high economic impact. We argue that these tasks can be solved as a special instance of actionable forecasting problems. We have gathered enough experimental evidence to support the conclusion that classification models may be preferable to model the generic tasks, mainly when used together with cost-benefit matrices. We have also observed some differences according to the number of possible decisions per task. The higher the number of possible decisions/actions to make, the higher will be the advantage of the classification models over the alternative of using regression models. With respect to the specific tasks of financial trading, both modelling alternatives revealed similar potential. The usage of re-sampling on such tasks brings too much risk to the models while using cost-benefit matrices on the classification models was beneficial once again. The last topic addressed in this thesis was the evaluation of financial trading systems. Based on the theoretical framework of activity monitoring, we have proposed a formalisation of financial trading as an instance of these data mining tasks. Using this formalisation we have described algorithms that allow to obtain the ideal timings for holding market positions given some trading preference criteria. These ideal timings can be used as an optimal benchmark against which real trading records can be compared to. The main advantage of this evaluation framework is its adaptability to the investor’s preference bii
ii ases and also the interpretability of the evaluation outcomes. We present the evaluation framework and test it using trading records of real investors.
Acknowledgements I would like to thank my parents and family for guiding and providing me with the wisdom to successfully choose, start and finalise this path. Whatever my successful accomplishments may be, they will also be my family’s. When it comes to my academic environment, I wish to express my sincere gratitude to my supervisor, Luís Torgo, not only for his great knowledge and expertise, but also for his constant support and guidance. I have significantly improved my scientific knowledge and taken the first steps towards the path of becoming an enthusiastic researcher. Certainly, a wonderful and enlightening experience. Last, but not least, a heartfelt thank you to the close friends that endured and grew with me during the bachelor’s and the master’s, as well as the friends that have kept in touch through the past few years. A final and kind word for my beloved partner who has been present through all my academic career and supported all my efforts. Thank you my dear. Luís Baía iii
iv
Contents 1 Introduction 1 1.1 LiteratureRevision................................ 2 1.1.1 Actionable Forecasting . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.1.2 Evaluation of Trading Performance . . . . . . . . . . . . . . . . . . . 4 1.2 ThesisOrganisation ............................... 6 2 Actionable Forecasting 7 2.1 ProblemFormalization.............................. 7 2.2 Proposedapproaches............................... 8 2.2.1 The Classification Approach: A limitation . . . . . . . . . . . . . . . 9 2.3 MaterialandMethods .............................. 10 2.3.1 TheTasks................................. 10 2.3.2 Evaluation Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.3.3 TheModels................................ 14 2.4 The Experimental Methodology . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.5 Analysis of the usage of cost-benefit matrices . . . . . . . . . . . . . . . . . 17 2.5.1 Wilcoxon test: Best per metric and per data set . . . . . . . . . . . . 18 2.5.2 Wilcoxon test: Best per metric, per data set and per type of model 23 2.5.3 Post-hoc Nemenyi test: Average Ranks . . . . . . . . . . . . . . . . . 27 2.5.4 Conclusions................................ 31 2.6 Comparison of Regression and Classification modelling approaches . . . . . 33 2.6.1 Wilcoxon test: Best per metric and per data set . . . . . . . . . . . . 34 2.6.2 Post-hoc Nemenyi test: Average Ranks . . . . . . . . . . . . . . . . . 35 2.7 Conclusions.................................... 43 3 An Application to Financial Trading 45 3.1 MaterialandMethods .............................. 46 3.1.1 TheTask ................................. 46 3.1.2 Evaluation Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 3.1.3 TheModels................................ 48 3.2 The Experimental Methodology . . . . . . . . . . . . . . . . . . . . . . . . . 49 3.3 Hypothesistesting ................................ 49 3.3.1 Hypothesis 1: Re-sampling the data sets - Classification . . . . . . . 50 3.3.2 Hypothesis 1: Re-sampling the data sets - Regression . . . . . . . . . 53 3.3.3 Hypothesis 2: Adding cost-benefits . . . . . . . . . . . . . . . . . . . 57 3.3.4 Conclusions................................ 62 3.4 Comparison of Classification and Regression modelling approaches . . . . . 64 3.4.1 Wilcoxon test: Best per metric and per data set . . . . . . . . . . . . 64 v
2Introduction or losses as a consequence. Any new study/discovery regarding trading systems may be crucial for any investor, motivating our analysis for this particular application. Therefore, the second main contribution of the thesis is to conduct another exhaustive experimental study to evaluate the advantages and disadvantages of both modelling approaches in the context of financial trading. Our last contribution in this thesis is also related with financial trading. Namely, we address the key issue of how to evaluate financial trading decisions. Given a historic record of trading signals produced by some trading system, it is not an easy task to evaluate the performance of that trading system. The most common approaches is to have an investor analysing several metrics at the same time and deciding whether the observed scores are good enough for his own preference criteria. This procedure is very subjective and most of the trading metrics can not be adapted to the investor’s trading policy (level of risk aversion as well as the target return per trade). In the thesis we propose a benchmark driven by the preference criteria of the investors against which we can compare any concrete trading system. This benchmark is derived based on a proposed formalisation of financial trading as a special instance of some data mining tasks known as activity monitoring. This formalisation and the derived benchmark provide new tools for looking at the performance of trading systems that we claim to be more interpretable for investors as they directly incorporate the investors’ preference bias in terms of trading results. 1.1 Literature Revision In this section we describe some of the work that has been carried out within the scientific areas that the thesis is related with. We divide the literature revision in two parts, one for the actionable forecasting tasks and a second one regarding the evaluation of a trading system. 1.1.1 Actionable Forecasting To the best of our knowledge there are no research works directly comparing the two modelling approaches we are proposing for actionable forecasting tasks. There is some research regarding the usage of a single or a set of models using the same approach in a specific task, but never considering models of both approaches altogether. Nevertheless, there are some concepts that are strongly related with the problem of Actionable Forecasting and thus we provide here a short review of some of the main works in these areas. In Actionable Forecasting tasks the ultimate goal is to predict an action/decision. However, there may be classes/decisions more important than others, or there may exist information on some costs and benefits associated with each decision. Another related problem is that of unbalanced distributions of the target variable in the context of prediction models. Frequently, in decision making, some of the decisions are less common, and moreover these tend to be more relevant. Branco et al. (2015) presents
1.1 Literature Revision 3 a survey of existing techniques for handling these situations of unbalanced target variables. Although most of the existing work considers classification tasks (nominal target variables), this paper also describes methods designed to handle similar problems within regression tasks (numeric target variables). Regarding the problem of rare classes in the target variable Weiss (2004) and Weiss (2005) are good references on this topic. We were not able to find generic studies on the problem of actionable forecasting. Still, there are a few works that address related problems on specific domains, such as medical, trading or other types of tasks. Nevertheless this type of research does not address the specific question were are aiming in this thesis: what is the best general methodology to address this type of decision making tasks. Regarding the specific area of financial trading, since we devote a whole chapter to this type of tasks, we will now list some work that is somewhat related with our work. In terms of the usage of regression models for trading systems, neural networks models have shown to be a promising tool for forecasting time series (namely the prices of some assets). However, it requires a large effort in terms of model tuning and can be computationally demanding. Gençay (1999) has observed that a simple feed-forward network fails to statistically outperform the random walk model when the input variables are just the past returns. However, with the simple addition of a moving average it can statistically outperform this baseline. Ghazali et al. (2011); Shin and Ghosh (1995) and Sermpinis et al. (2012) have proposed variants and generalisations of neural networks that present decent improvements over the standard versions of these models. Most research seems to indicate that these standard versions of the neural networks models may present poor results, yet some small variations in terms of the structure of the model may greatly improve their performance. Another popular model in this area are the k-nearest neighbours (KNN). Gençay (1999) has observed that the regression KNN model statistically outperforms the Random Walk model by merely using the past return as predictors, unlike the feed-forward neural network. Lee et al. (2012) has used the nearest-neighbour-based approach for churn predictions. Even though this problem is different from financial trading, there are some similarities. Detecting an uncommon event the soonest possible can be seen as an analogy to the trading problem, i.e. detecting a buy or sell signal the earliest possible to obtain the maximum possible profit. In this article, this model achieved very optimistic and positive results. Support Vector Machines (SVM) and Multivariate Adaptive Regression Splines (MARS) have been the subject of study by Kao et al. (2013). The author also tested these models incorporated with wavelets, where some interesting results were obtained. Furthermore, Lu et al. (2009) has studied the combination of using Independent Component Analysis to reduce the noise and randomness of the data before applying standard SVM models, and has observed a slight improvement of the results. Fernandes (2002) has compared the usage of several regression models in financial timeseries forecasting. This author has tested regression trees, linear models, neural networks,
4Introduction etc., in several financial data sets, and has also considered the bootstrap aggregation of those models. Evidence was found for arguing that these regression models can outperform the baseline benchmarks such as the random walk model or the average model, with statistical significance. In terms of using classification models in the context of financial trading we have found fewer research works. Chang et al. (2009) have studied the combination of piecewise linear models with back-propagation artificial classification neural networks (PLR-BPN). The experimental results were interesting in terms of the amount of profit obtained. However, Luo and Chen (2013) have seen that PLR-BPN is outperformed by the combination of PLR with the well known Support Vector Machine model. Ma et al. (2012) have implemented cost matrices in back-propagation neural networks. In several tasks, an increase of the utility score of a model came at the cost of a decrease in the accuracy level. However, this accuracy decrease may not be relevant for financial trading, where the main goal is to avoid serious errors like suggestion to buy some assets when we should sell it, as this type of errors may have very serious economic impact. Teixeira and de Oliveira (2010) have studied the classification KNN model and also the combination of this model with indicators such as the RSI filter, stop-gain and stop-loss criteria. All the tested models outperformed the used benchmark particularly the models combined with the above indicators, that obtained an overall better performance. Atsalakis and Valavanis (2009) contains a very detailed state of art regarding stock market forecasting techniques. The authors list the work carried out by several researchers. Namely, for each referred work the authors list the used data sets, the chosen input variables and, most importantly, a summary of all the used modelling techniques as well as which models were compared against each other. Information regarding the usage of preprocessing techniques and the training method was also given. In none of the referred works we have seen a comparison between regression models and classification ones in the context of financial trading. Several techniques are explored but the answer to the question proposed in the thesis seems to be absent from the research works we have encountered in our review of the state of the art in these areas. Moreover, it seems regression models are being more frequently used than the classification ones to address trading tasks. 1.1.2 Evaluation of Trading Performance Another major topic of this thesis is the problem of evaluating a trading system. Whether a trading system tries to forecast the future variation of the prices or directly the final trading decision, the ultimate goal of these systems is to make the correct (and more profitable) trading decisions. In this context, from an investor perspective it is not very interesting to evaluate a system through metrics like the Root Mean Squared Error (in case there was a regression model trying to forecast the future price) or the Accuracy (how many correct decisions did the trading system made), but rather what were the
1.1 Literature Revision 5 financial results of the decisions taken by the system. This is typically checked through the analyses of several specific trading metrics over a testing period. By doing so, the investor may for instance check the return and the risk associated with the trading system, which are measures that are absolutely crucial for investors. Pardo (2011) is a well know reference that includes some performance summaries for trading systems, using trading metrics such as the total net profit, the ratio between the average profit over the average loss, maximum draw down, profit factor and several others. The Sharpe ratio is a well known metric used to describe the risk associated with a trading system. There are more helpful measures to evaluate the risk associated with a trading system, but Ferruz et al. (2006) has seen that there is a strong correlation (higher than 0.9) between all possible combinations between the Sharpe, Treynor and Jensen risk measures when dealing with Spanish investment funds. This suggests that even though we may use several different risk measures, we will eventually obtain similar conclusions among them. Therefore, we will give special focus to the Sharpe ratio during the thesis. Unarguably, it has been one of the most used performance criteria to evaluate traders. It has been thoroughly studied by several researchers, namely by Bailey and Lopez de Prado (2012), Lo (2002) and Pav (2014). Essentially, this metric is able to check whether the returns of a portfolio are due to smart investment decisions or a result of excess risk. Its formula depends on a benchmark, where usually the null benchmark is chosen1. Despite the existence of several measures either characterising the profit or the risk of a trading system, only when considering a large set of metrics together one can properly infer about the quality of the system. Folger (2012) presents a possible guide on how to evaluate the performance of a trading system, where several metrics are considered. The author strongly advises to go over a thorough study regarding a vast number of metrics, involving the profit and the risk of the trading system. This established procedure to evaluate the performance of a trading system, and the fact that most trading metrics have no dependency on the investor preference criteria in terms of profitability and risk, motivates the following question: if an investor A is willing to accept higher risk for higher returns and, on the other hand, another investor B prefers a more conservative policy, how can they determine if the performance of a certain trading system is the best for their preferences just by looking at these standard metrics? Answering this question is not easy as the metrics do not tell the investor how far are they from the optimal trading record from the perspective of their criteria. Luo and Chen (2013) has scratched the surface of this question, by using the piecewise linear representation (PLR) to obtain, on fully known data, what apparently would be the nearperfect points to trade. Even though it was not mentioned in his work, this could be used as some optimal benchmark to evaluate a trading system. However, this method depends on a very complex and hard to interpret parameter that will increase the number of optimal 1Typically, the null benchmark is a trader that has no signals, i.e., holds the position all the time leading to zero profits and zero losses
6Introduction turning points the lower the parameter gets. Figure 1.1 illustrates the usage of the PLR to obtain the “optimal” benchmark for a specific time-series and for different values of δ. Figure 1.1: PLR used in Luo and Chen (2013) to obtain some “optimal” turning points With the exception of this work that is slightly related with this concept of optimal benchmark, we have not found any other research literature that addresses this important question from the point of view of investors. This issue is one of the main contributions of the thesis. 1.2 Thesis Organisation This thesis is organised as follows. Chapter 2presents the concept of Actionable Forecasting, one of the main contributions of the thesis. We describe this type of applications and propose two alternative methods to address these problems. We study their advantages and disadvantages and present the results of an extensive empirical comparison of these alternatives using a large set of generic tasks and an extensive set of modelling tools. In Chapter 3we focus on one particular instance of the actionable forecasting tasks - financial trading. This is an application with high impact in our society. The characteristics of this application requires some modifications to our proposals. We compare the two alternative methods to solve actionable forecasting tasks in this particular application, with the goal of checking if their particularities change somehow the conclusions we have reached for the general case. Chapter 4describes our final contribution that addresses the evaluation of trading systems. We propose a formalisation of financial trading as a special instance of activity monitoring data mining tasks. Based on this formalisation, we provide means of obtaining the optimal trading actions given the economic preferences of an investor. We describe this optimal benchmark and apply it to evaluate the trading record of a real investor, highlighting the advantages of our proposal. Finally, in Chapter 5we briefly describe the main conclusions and contributions of this thesis, and outline possible directions for future research developments related with our thesis.
Chapter 2 Actionable Forecasting Actionable Forecasting can be defined as the process of predicting the right action, when the action itself is a function of a predicted numerical variable. This problem is very common in several distinct areas, such as medicine, trading, industry market, etc. For instance, consider the process of deciding which type of treatment should be applied to a patient with cancer based on the future volume of the tumour. Naturally, there is no way of knowing exactly this value, however a decision must be made. The “simplest” procedure would be to somehow predict the growth of the volume and based on that forecast make a decision. As we will see, there are other approaches to deal with this problem and one of the main contributions in the thesis is to properly compare two of the most used approaches. Another example application, whose importance is so high that will be separately studied in a further section, is financial trading. As an investor, deciding whether to open a short/long position or to hold takes serious risks. In principle, if the investor could be sure if a certain asset would increase or drop its value, then the decision would be straightforward to make. However, a decision must be made without knowing the numerical value of the asset in the future. Once again, the “simplest” procedure of somehow predicting the numerical value and then make the decision would work, but other approaches are also valid. 2.1 Problem Formalization The problem of decision making based on forecasts of a numerical (continuous) value can be formalized as follows. We assume there is an unknown function that maps the values of ppredictor variables into the values of a certain numeric variable Y. Let fbe this unknown function that receives as input a vector xwith the values of the ppredictors and returns the value of the target numeric variable Y, whose values are supposed to depend on these predictors, 7
8Actionable Forecasting f:Rp→R x7→ f(x). We also assume that based on the values of this variable Ysome decisions need to be made. Let gbe another function that given the values of this target numeric variable transforms them into actions/decisions, g:R→ A ={a1, a2, a3, . . . } Y7→ g(Y). where Arepresents a set of possible actions. In our target applications, functions fand gare very different. Function gis known and deterministic, in the sense that it is part of the domain background knowledge. Function f is unknown and uncertain. The only information we have about function fis a historical record of mappings from xinto Y, i.e. a data set that can be used to learn an approximation of the function f. We call this class of problems actionable forecasting tasks. At this stage it is important to remark that other related problems exist. For instance, there may exist applications for which gis also uncertain, or it may even vary with time. These alternative decision making scenarios are not addressed in this thesis. Our goal is to address the specific decision tasks that was described previously in the above formalization. 2.2 Proposed approaches Given that the variable Yis numeric, a prediction of its value could be obtained using some existing multiple regression tool. This means that given a data set Dr={hxi, Yiin i=1}we can use some regression tool to obtain a model ˆr(x)that is an approximation of f. From an operational perspective this would mean that given a test case qfor which a decision needs to be made we would proceed by first using ˆrto obtain a prediction for Yand then apply gto this predicted value to get the predicted action/decision, i.e. q7→ ˆr(q)7→ g(ˆr(q)). Given the deterministic nature of g, we can use an alternative process for reaching decisions. More specifically, we can build an alternative data set Dc={hxi, g(Yi)in i=1}, where the target variable is the decision associated with each known Yvalue in the historical record of data. This means that we have a nominal target variable, i.e. we are facing a classification task. Once again, we can use some standard classification tool to obtain an approximation ˆcof the unknown function that maps the predictors into the correct actions/decisions. Once such model is obtained, we can use it given a query case
2.2 Proposed approaches 9 qto directly estimate the correct decision by applying the learned model to the case, i.e. q7→ ˆc(q). Independently of the approach followed, the final goal of the applications we are targeting is always to make correct decisions. This means that whatever process we use to reach a decision, it will be evaluated in terms of the "quality" of the decisions it generates. In this context, it seems that the classification approach, by having as target variable the decisions, would be easier to bias towards optimal actions. However, this approach completely ignores the intermediate numeric variable that is supposed to influence decisions, though one may argue that information on the relationship between Yand the decisions is "encoded" when building the training set Dcby using as target the values of g(Yi). On the other hand, while the regression approach is focused on obtaining accurate predictions of Y, it completely ignores questions like eventual different cost/benefits of the different possible decisions that could be easily encoded into the classification tasks. All these potential trade-offs motivate the current study. The main goal of this chapter is to compare these two approaches in the context of this type of decision tasks. 2.2.1 The Classification Approach: A limitation A problem that may hinder the performance of the classification modelling approach is presented and discussed. These models face a theoretical problem as these algorithms do not distinguish among the different types of errors. Consider a task with three possible classes/decisions C1, C2, C3, that result from a deterministic mapping of a numerical response Y, where C1corresponds to a value lower than k1, the class C3to a value higher than k2> k1and the class C2to the the values in between. Typically, confusing C1with C3is a more serious mistake than C1with C2. However, the way a classification model is built makes all the errors equally weighed, as they assume a nominal target variable with no ordering among its values. This problem led us to consider an alternative to our base modelling approaches . We have considered a frequently used approach to this issue. Namely, we have used a costbenefit matrix that allows us to distinguish between the different types of classification errors (e.g. Torgo and Gama (1997)). Using this matrix, and given a probabilistic classifier, we can predict for each test case the class that maximises the utility instead of the class that has the highest probability. We have used the following procedure to obtain the cost-benefit matrices for our generic tasks. Correctly predicted decisions are rewarded with 1. On the other hand, any wrong predictions are penalised as minus the number of classes away from the correct one, e.g. forecasting as C5aC1class is penalized with −4, while predicting C2on a true C3will receive a score of −1. Table 2.1 shows an example of such a cost-benefit matrix for a generic data set with 5possible actions
10 Actionable Forecasting C1 C2 C3 C4 C5 C1 1.00 -1.00 -2.00 -3.00 -4.00 C2 -1.00 1.00 -1.00 -2.00 -3.00 C3 -2.00 -1.00 1.00 -1.00 -2.00 C4 -3.00 -2.00 -1.00 1.00 -1.00 C5 -4.00 -3.00 -2.00 -1.00 1.00 Table 2.1: Cost-Benefit matrix for a 5class task. We have also thoroughly tested the hypothesis that using cost-benefit matrices to implement utility maximisation would improve the performance of the models. The results will be presented after the experimental methodology is described. 2.3 Material and Methods Having described the concept of the Regression and Classification approaches for the Actionable Forecasting problem, our next goal is to compare both of them in in different experimental settings. We have chosen an empirical comparison because it may be difficult, perhaps impossible, to compare them in a theoretical way since each task and each modelling tool present their own theoretical properties, with some being extremely complex and distinct. Hence, any chance of incorporating several distinct tasks and models on the same study would not be viable. The main issues involved in the experimental comparison will now be presented. Firstly, the tasks that will be used in our study are detailed, describing both the available data sets and the variables involved in the process. Secondly, we address the important decision of how the alternatives will be evaluated and compared by describing the evaluation metrics used in our experiments. Finally, the modelling tools that will compose each modelling approach will be listed. All the work of the thesis was performed using the R programming language R Core Team (2014). 2.3.1 The Tasks In this chapter the goal is to perform a generic study, minimising the eventual influence of problems such as the class imbalance or the existence of decisions much more important than others. These questions will be addressed in the next chapter, namely when dealing with trading tasks. In the experiments of this chapter several regression tasks will be considered in which the gfunction, that maps the numerical response into actions, will consist of defining thresholds that will make all the respective classes (i.e. decisions) somewhat balanced in terms of frequency. Eight data sets of completely different contexts were used, whose number of observations vary between 1503 and 10617 (some brief details about each data
2.3 Material and Methods 11 set are given in Table 2.2). Each data set will lead to four different tasks, resulting from the application of four different variants of the gfunction. Essentially, for each data set, a version with 2,3,5and 10 classes/decisions is formed. The point is to ensure a rich study by not only having distinct contexts, but also a diversified set of possible actions per context. Behaviours such as one approach being preferable to predict a smaller set of actions while the other approach being fitter to model data sets with more actions may thus be captured. Data set Description (Nr.Cases; Nr.Predictors) Abalone (9;4177) Predicting the age of abalone from physical measurements. The age of abalone is determined by cutting the shell through the cone, staining it, and counting the number of rings through a microscope. Airfoil (6;1503) The NASA data set comprises different size NACA 0012 airfoils at various wind tunnel speeds and angles of attack. The span of the airfoil and the observer position were the same in all of the experiments. The goal is to predict the pressure. Concrete (9;1030) Predict the actual concrete compressive strength for a given mixture under a specific age that was determined from laboratory. Bank8FM (9;4499) A simulation of how bank-customers choose their banks. Tasks are based on predicting the fraction of bank customers who leave the bank because of full queues. CCP (5;9568) The data set contains data points from a Combined Cycle Power Plant. Features consist of hourly average ambient variables Temperature, Ambient Pressure, Relative Humidity and Exhaust Vacuum to predict the net hourly electrical energy output of the plant. CpuSmall (13;8192) Relative CPU Performance Data, described in terms of its cycle time, memory size, etc. Delta Ailerons (6;7129) The task of controlling the ailerons of a F16 aircraft Naval (13,8000) A numerical simulator of a naval vessel produced a 16feature data set containing the Gas Turbine measures at steady state of the physical asset. The goal is to predict Gas Turbine Turbine decay state coefficient Table 2.2: Brief description of each task. Table 2.3 gives some information of the created data sets. All consisted of tasks whose response variable was originally numeric. However, a new classification variable was added by mapping the numerical response using the threshold function g. In this way, regression models can be used to predict the numerical response, where the prediction will be applied in the known function gto construct a decision, while classification models can be directly used to predict the “artificial” variable. Note that the relative frequency of each class is not the same across all the data sets with the same number of classes.
18 Actionable Forecasting As a side note, the differences between the standard models and the ones with costs should become more evident with the increase of the number of possible actions/decisions. If there are more classes to be predicted, then there will be more chances to confuse distant classes, allowing the cost-benefit matrices to have a higher influence. Figure 2.1: Analysis of the relative frequencies on a standard random forest classification model (default specifications of the used package) with and without the usage of costbenefit matrices. Before starting the main analysis, we illustrate the effect of the use of cost-benefit matrices with a simple example on a Actionable Forecasting tasks. We have considered a regression task and categorised the numeric response into 10 ordered classes. In Figure 2.1, we have the relative frequencies of the predictions of a RandomForest model (default specifications used), with and without the usage of cost-benefit matrices. Eighty percent of the data set was used as the training set, with the remaining part being used as the testing sample. The relative frequency of the model with costs is superior in the central classes and lower on the extreme ones. Even though it seems that the model without costs is more frequently closer to the true relative frequency. one should remember the embedded numeric variable on the response target, implying that we should not only look for the right decision but also try to minimise the severity of a mistake when it occurs. 2.5.1 Wilcoxon test: Best per metric and per data set The first part of this study consists of performing a Wilcoxon test between the best classification variant without costs against the best one with costs per data set and per metric.
2.5 Analysis of the usage of cost-benefit matrices 19 In other words, for each individual metric we are selecting the best model variant from the alternatives using cost-benefit matrices and comparing it against the best model variant not using these matrices. This study may provide some insight about the impact of the usage of this cost-sensitive approach on classification models, but one should keep in mind that we are only comparing the top performer of each approach (cost vs no-costs) and it may be dangerous to infer any global conclusion by merely looking at theese top variants In Figure 2.2, we can see the results of the comparison in terms of Accuracy. For each data set, we have used the Wilcoxon test to check the statistical significance of the observed differences. The presence of an asterisk implies that the null hypothesis of the statistical test was rejected, thus indicating that one model significantly outperformed the other. The red bars refer to the standard models while the blue ones are for the models using costs. Figure 2.2: Best classification variant without costs (red) against the best classification variant with costs (blue) for the Accuracy score, where the presence of an asterisk implies that the respective variant performed significantly better than the one it is competing against, according to a Wilcoxon test with α= 0.05. There is a total of three significant wins of the model variants with costs against zero of the standard models. This comes as a surprise, since the standard models are optimised for this metric. However, only the best representative of each modelling approach is being considered, meaning that this behaviour may not occur across all the models
20 Actionable Forecasting overall. Moreover, the three significant wins occur only in tasks with 10 classes. One possible explanation is that it is theoretically expected for the cost-benefit matrices to have a higher influence on tasks with more classes, where the implicit ordering among the class values is more important, providing more space for a positive impact for using costs. In Figures 2.3 and 2.4 we see analogous plots for the Precision and Recall metrics. In terms of Precision, the model variants with costs have obtained several significant wins while some balance has been observed in the Recall metric. These graphs tell us that the new models (with costs) are increasing, in average, the precision level of each class individually without compromising their ability to properly detect, in average, more observations of each class when compared to the standard ones (without costs). These metrics may be quite useful in tasks whose classes of the target variable are not equally important. This behaviour could be explained by the fact that the models with embedded costs will predict more cases as the middle classes1. On the contrary, these models forecast an observation on the extreme classes (such as C1and C10) only when there is a high level of confidence for that prediction. This happens because confusing those extreme classes would incur on a very large cost, thus leading to more “central” forecasts in order to reduce the expected cost. Therefore, the precision on the “non central” classes should be quite high while the precision for the most central ones should decrease. On the other hand, due to more accurate, safer and consequently, fewer predictions on the extreme classes of the models with costs, the Recall will fall, with some increase being verified on the central classes. When considering the macro version of these metrics (the average across all the classes), our results suggest that using costs improves the Precision level without any serious change being verified on the Recall metric. This may be considered as a decent improvement when using cost-benefit matrices, though one should know the reasoning behind this improvement. One should also note that most of the significant wins occur on tasks with 10 classes, the ones where the usage of costs may have a higher influence on the final predictions, thus increasing the differences between the models. To sum up, the usage of costs is biasing the forecasts to the more central classes, allowing the predictions on the remaining classes to occur solely when there is a high degree of confidence. Therefore, we are seeing higher values of Precision when applying costs to the standard models. That could have come at the cost of some decrease on the levels of Recall, but that was not confirmed in our analysis. Even though we are just looking at the best model variant, we should see this behaviour, mainly the increase on the Precision metric, for most of the remaining models (something to be analyzed in the next parts of this study), since it is theoretically expected to see the models with costs 1By middle class, we mean the classes that result from the mapping of the central values of the numeric intermediate variable. For instance, in a 10 class task, the middle ones would be the C4, C5and C6, since they would result from the most central values of the internal numeric target variable.
2.5 Analysis of the usage of cost-benefit matrices 21 Figure 2.3: Best classification variant without costs (red) against the best classification variant with costs (blue) for the Precision score, where the presence of asterisk implies that the respective variant performed significantly better than the one it is competing against, according to a Wilcoxon test with α= 0.05. Figure 2.4: Best classification variant without costs (red) against the best classification variant with costs (blue) for the Recall score, where the presence of asterisk implies that the respective variant performed significantly better than the one it is competing against, according to a Wilcoxon test with α= 0.05.
22 Actionable Forecasting making decisions more often in the central classes in order to avoid the heavy costs given to confusing classes at maximum distances (such as C1and C10). Regarding the F-Measure, which is a metric that equally weighs Macro-Recall and Macro-Precision into a single metric, we should be seeing some trend favouring the models with costs, since they achieved better values of Precision without severely compromising the Recall. In Figure 2.5, we can see the results for the F-measure metric with α= 1. Indeed, the results confirm our expectations, where once again, most of the significant wins occurred on the tasks with 10 classes. Figure 2.5: Best classification variant without costs (red) against the best classification variant with costs (blue) for the F-Measure score, where the presence of asterisk implies that the respective variant performed significantly better than the one it is competing against, according to a Wilcoxon test with α= 0.05. Finally, we analyse the Utility measure in Figure 2.6. There are 8against 1significant wins clearly demonstrating some trend favouring the models with costs. This should be expected, since the models using cost-benefit matrices are using information that should help in maximising the utility of their predictions. Hence, one may argue that it is not a fair comparison, but the point here is to empirically prove the expected advantage of the models with costs in terms of Utility, and that it may come without harming (or it could actually improve) some other metrics like the Accuracy and Precision. As in the previous cases, most of the significant wins using costs occur on the tasks with more classes.
2.5 Analysis of the usage of cost-benefit matrices 23 Figure 2.6: Best classification variant without costs (red) against the best classification variant with costs (blue) for the Utility score, where the presence of asterisk implies that the respective variant performed significantly better than the one it is competing against, according to a Wilcoxon test with α= 0.05. To sum up, in this first part of the analysis, we have seen that both approaches are somewhat balanced with some advantages for the usage of costs, at least when selecting the best representative of each modelling approach. In terms of Accuracy, there was some small advantage of the models using cost-benefit matrices whilst in terms of Precision the difference was more noticeable. Some equilibrium was observed in the Recall metric. Finally, when it comes to the Utility score, the new models are obtaining considerably better values, thus indicating that adding costs to the baseline models is worth considering and may add some value to the whole set of classification models. In the next part of the study, we will move from comparing the best overall variants (with and without costs) to the best variants per modelling technique, such as the Support Vector Machines, Random Forests, etc. 2.5.2 Wilcoxon test: Best per metric, per data set and per type of model In the second part of the analysis of the impact of the use of costs on the performance of classification models, we present a different type of study. Instead of selecting the best variant across all the models for the two alternatives (costs vs no-costs), we will select the
24 Actionable Forecasting best for each type of model. This will allow us to check if the conclusions of the previous section still hold if we focus on a single algorithm (e.g. SVM or Random Forest). Using the same type of graphs as in the previous section would entail a very large number of plots, so we have decided to use a different visual representation of the results of this experiment. Our solution is given in Figure 2.7. We divide the tasks according to the number of classes, meaning that the numbers 2,3,5,10 at grey indicate which type of task is under analysis. The columns on the right indicate the metric being used to evaluate the models. On the left, we see the type of models, repeated for each metric. Each subplot illustrates the number of significant and non significant wins across the 8tasks (there are 8data sets with 2classes, 8data sets with 3classes, etc) per type of model. Finally, the colours indicate the following: •Strong Blue: Significant Wins for Models without costs; •Weak Blue - Non Significant Wins for Models without costs; •Yellow - Draws (meaning both top model variants had the same exact score); •Weak Green - Non Significant Wins for Models with costs; •Strong Green - Significant Wins for Models with costs. Let us consider, for instance, the subplot for the Accuracy metric for data sets with 3 classes: i) We can see that the NNET (neural networks) type of model returns 1significant win for the models without costs, 5non significant wins for the models without costs and 2non significant wins for the models with costs; ii) Regarding the Adaboost type on the same subplot, we can see 2significant wins and 3non significant wins for the models without cost plus 2draws and 1non significant win for the model variants with costs. Note that each single bar corresponds to the result of 8Wilcoxon statistical tests, each test performed on a different data set. This plot is quite dense and some time should be devoted to it. A first obvious observation is the abundance of yellow (draws) on the tasks with two classes. This is expected because with two classes the use of cost should only make sense if the data set is not balanced and a higher reward is given to the rarer class. Only the Airfoil and CCP data sets (see Table 2.3) have have this characteristic, which may explain the existence of a significant win with the SVM models in the Precision metric. One may also question the reason for the different behaviour of SVMs in two-class tasks. If we look at the list of models used in our experimental study (Table 2.6), we see that the SVM type is by far the one with more variants, up to a total of 32, while the other types have nearly half that amount2. Naturally, when comparing 32 variants without costs with 32 with costs, the chance for the best of each group to be different from each other increases, even in tasks where cost-benefits would have little influence. 2The kernel of the SVM model has a great impact on the model itself. By choosing two types of kernel, the number of variants nearly doubles
2.5 Analysis of the usage of cost-benefit matrices 25 Figure 2.7: Segmented by type of model, by metric and by the number of classes of the response variable of each task, a Wilcoxon test is performed between the best model variant of each modelling tool (Classification without costs vs Classification with costs). Since there are 8 data sets in each segment, there are 8 results per segment. Each bar shows the results for each segment, where each one of the five colours is associated to a type of win (significant of the classification approach, non-significant, draw, non-significant for the regression approach and significant). The length of each colour describes the number of times that type of win occurred. Metric by metric, we will draw some conclusions and see if they are coherent with the analysis from the previous subsection. In terms of Accuracy, it is hard to pinpoint any special advantage of any approach. There are models such as the Random Forests that seem to benefit more from the usage of costs, but we see slightly the opposite in models like the SVM and NNET and in a more severe way for the AdaBoost type with a large number of significant wins for the models not using costs. The number of classes seems to have little influence on this comparison for the accuracy metric, though there seems to be a slight improvement for the standard models for a higher number of classes. These remarks are coherent with the conclusions of the previous section, where some balance was observed, with slighter higher differences for tasks with more classes, with the exception of
26 Actionable Forecasting the AdaBoost type, where the usage of costs decreases the Accuracy the higher the number of classes. Concerning Precision the models with costs are clearly better across all the type of models, with the exception of Adaboost. This could be explained by the structure of the Adaboost models, since at each learning iteration of these models, every wrong prediction will be given a higher weight, somehow decreasing the severity of their mistakes. In the work of Desai and M Jadav (2012), we see that the standard AdaBoost models have their levels of performance decreased when cost sensitive extensions are added. In summary, with the exception of Adaboost, we have confirmed the results of the previous section where we had observed that the usage of costs-benefits improved the results in terms of Precision. Regarding Recall we observe that the models without costs obtain better results, particularly for a higher number of classes. Still, this general trend is somewhat contradicted by KNN and Random Forests where we observe more balanced results between the two alternatives. As with Precision, this general trend of the results was already observed in the previous section. Finally, in terms of Utility, all models except Adaboost obtain much better results when using costs. Once again, this observation is more marked in data sets with more classes. To sum up, this second type of analysis corroborates the conclusions from the first analysis. With the exception of Adaboost we have observed the same type of effect of the usage of cost-benefit matrices on the selected metrics. In terms of Precision and Utility it is clear that the higher the number of classes, the better the models with costs will be. On the other hand, the opposite behaviour is observed for the Accuracy metric, where more classes per data set seem to favour the models without costs. This is theoretically expected since with more classes, the new models will tend to prefer forecasts of the central values/classes in order to decrease the severity of their mistakes. Naturally, this may penalise the accuracy levels, since the predicted class may not correspond anymore to the one with the highest predicted probability. In appendix A.1, we present the results of another similar experimental study. Compared to the current study, instead of comparing the best variant of each modelling technique using costs and without costs, we compare the median score of all variants of the technique. Our goal with this other experiment was to show the results of the comparison that should be expected by a user that is not willing to spend that much time looking for the best variant of each modelling algorithm. The results are similar to the ones observed using the top modelling variant per group. However, the abundance of the green colour is stronger. There are more cases in which the models with costs are now outperforming the ones without. The AdaBoost type is the most evident case, where the dominance of the blue colour is replaced by the green one. This study suggests that if a user that is not able to conduct an extensive search for the best parameters of a classification model, then by using costs they should achieve better performances.
2.5 Analysis of the usage of cost-benefit matrices 27 2.5.3 Post-hoc Nemenyi test: Average Ranks So far, we have evaluated and compared, task by task, the performance between the top modelling variants of either approach (with and without costs). This has the potential limitation of some model variant to be chosen as the best, due to some randomness (by fitting the task unusually well, where the same model variant would achieve poor performances on other tasks). This possibility raises the question if this model variant was properly representing its respective modelling approach. Therefore, we introduce the last component of this study. The objective is to compare the best model variants of each modelling approach in terms of their overall performance across all the tasks of the same type (same number of classes), instead of conducting a separate test task by task. Hence, we may not only reinforce the quality of the conclusions drawn in the previous analysis, but also evaluate both approaches in terms of the robustness and adaptability of their model variants (since the top modelling variants will now be the ones who can perform well in several tasks at once). In order to do that, we will use the Friedman test followed by a post-hoc Nemenyi, segmented by the number of classes of each data set. For the models without and with costs, for each set of tasks (2,3,5,10 classes) and for each metric, we will rank each model variant according to its score (the best variant without costs for the accuracy metric for the Airfoil2 task receives rank 1, the second receives rank two, etc. The same is done for the models with costs). Then the procedure is repeated for the other tasks with the same number of classes, where the rankings of each model will be averaged. The five top models of each approach will be selected thus creating a new set of 10 models. Finally, the Friedman test followed by a post-hoc Nemenyi will be applied to this new set. By doing so, we are picking the best variants of each modelling approach in terms of their overall performance across all the tasks of the same type (same number of classes) and we may check whether these 5+5models present similar average ranks among themselves or quite distinct ones, thus allowing to infer that one approach is performing better than the other across several tasks for a specific metric. This methodology should become clearer after analysing Figure 2.8. Since in the 2class tasks, the influence of the cost-benefit matrices on the predictions of the models is quite low, we start directly by analysing the tasks with three classes. As an illustrative example, let us fully analyse the Precision sub-image (the second sub-image of Figure 2.8): The labels on each side are the name of the models that are connected to an average rank value through a line. This is the average ranking that the corresponding model obtained across all the data sets with three classes. Therefore, we see that the model named as “CLASS_RF_2.v8” has achieved the best score for the Precision metric obtaining an average ranking of approximately 2. Regarding the interpretation of the name of the models, “class” stands for classification, the text after the first underscore stands for the type of model (which in this case stands for RandomForest), while the number after the
34 Actionable Forecasting their performance as observed in the previous section. 2.6.1 Wilcoxon test: Best per metric and per data set Similarly to the comparison of classification models with and without cost-benefits, we will compare and carry out a Wilcoxon statistical test between the top modelling variant of each modelling approach (classification and regression) per task. This type of comparison simulates a setup where the user is able to almost exhaustively search for the optimal parameters that will lead to the best performance possible per task. This comparison will be able to identify any characteristic that may make one modelling approach more desirable than the other. In Figure 2.10, the results are shown for the Accuracy metric. There are 13 significant wins for the classification modelling approach against 2of the regression one. This is an overwhelming victory, though one should keep in mind that most of the tasks that verified the Wilcoxon’s test are the ones with more classes in the response variable. This suggests that most of the classification wins are occurring due to the usage of the post-optimised models with cost-benefit matrices, since with two or three classes, the number of wins falls to three cases. The regression seems to be more competitive with fewer classes Figure 2.10: Best classification variant (red) against the best regression variant (blue) for the Accuracy score, where the presence of asterisk implies that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05.
2.6 Comparison of Regression and Classification modelling approaches 35 Regarding Precision (Figure 2.11) the results appear to be quite similar to Accuracy. Another overwhelming victory of the classification modelling approach was obtained with 11 against 2significant wins. Figure 2.12 shows the results in terms of Recall, where we can observe that the classification modelling approach has obtained 9significant wins against 2of the regression models. Therefore, it is expected that the F-Measure that equally weights Recall and Precision into a single metric will favour the Classification models, since the latter have outperformed in both cases. In Figure 2.13 we can confirm these expectations. Lastly, when analysing the Utility score in Figure 2.14, we observe the most significant victory of the classification approach with 13 significant wins against 2of the regression alternative. These results provide plenty evidence that, if the users are willing to make some efforts in terms of searching the best modelling variant available in terms of tuning the parameters and testing several types of models, then they should find their best models for an actionable forecasting problem in the Classification approach, namely when using cost-benefits. There is, however, a drawback of this type of general conclusions. By looking at each task separately, we may be disregarding some models that obtained good scores across most tasks without ever being the best model available. These models may also be of interest, since they present a stable performance across all tasks thus seeming to be able to adapt to most tasks. This motivates our second analysis that looks at the average rankings of each model. 2.6.2 Post-hoc Nemenyi test: Average Ranks Finally, we analyse the average rankings of the best five average models per modelling approach. The main goal is to draw some conclusions regarding the robustness and adaptability that both classification or regression models may offer and determine if one approach outperforms the other in this respect. Unlike in the analysis of cost vs no cost, we will analyse not only the tasks with 3and 10 classes, but also the ones with 2and 5classes. The outcome depending on the number of classes is not obvious nor predictable, thus motivating the analysis case by case. However, we do know that by adding the model variants with costs to our set of classification models, their results were clearly improved, mainly on tasks with a higher number of classes. Therefore, the regression models should have more difficulties outperforming the classification modelling approach on data sets whose response variable has a larger domain. In Figure 2.15, we can see the obtained results for all the tasks whose response variable is binary. The first observation is the lack of any significant results across all the metrics. However, some conclusions may still be drawn. In terms of Accuracy, Recall and Utility, the top five models (lowest average rank) belong to the classification modelling approach
36 Actionable Forecasting Figure 2.11: Best classification variant (red) against the best regression variant (blue) for the Precision score, where the presence of asterisk implies that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05. Figure 2.12: Best classification variant (red) against the best regression variant (blue) for the Recall score, where the presence of asterisk implies that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05.
2.6 Comparison of Regression and Classification modelling approaches 37 Figure 2.13: Best classification variant (red) against the best regression variant (blue) for the F-Measure score, where the presence of asterisk implies that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05. (except the fifth position on the last metric) while in terms of Precision, the results are more even. Although the differences are small, the classification models have some advantage. A note of interest is that the Random Forest type completely dominated the results for these tasks and most of the top classification model variants are the standard versions without the addition of cost-benefits matrices. The dominance of the Random Forest type was a constant when evaluating the classification models with and without costs, and now it seems to be the best regression modelling tool. This matter will be addressed again when analysing the remaining tasks. The last point was rather expected, since with a binary response variable, whose classes have quite similar relative frequencies, there is not much space for the cost-benefit matrices to have a serious influence. Figure 2.16 presents the results for the 3-classes tasks. The most noticeable difference regarding the previous case is the rejection of the Friedman’s null hypothesis for the Accuracy and Precision metrics. Furthermore, the average rankings of the classification models are clearly outperforming the regression ones in terms of Accuracy. In terms of Precision, even though the Friedman’s test was verified, implying that the average rankings are significantly not equal, it is not easy to pinpoint the success of one approach over the other. Indeed, the first three places are taken by classification modelling variants and with some visible distance from all the other models, but the next two models
38 Actionable Forecasting Figure 2.14: Best classification variant (red) against the best regression variant (blue) for the Utility score, where the presence of asterisk implies that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05. belong to the regression approach. Therefore, we claim that both modelling approaches were quite even in this metric, with some slight advantage for the classification approach. Concerning the last two metrics, Recall and Utility, no significant results were obtained. Even though the top five modelling variants were all classification ones, there is not enough evidence to support any advantage of this approach. Actually, it is quite interesting to note that the regression approach is able to be competitive against models that were optimised towards the Utility score. Indeed, the process of forecasting the numeric variable is helping to reduce the severity of the errors on the predicted decisions. Figure 2.17 shows the results for 5-classes tasks. They are quite similar to the previous case. The classification models outperform the regression ones in terms of Accuracy, but the results are more similar on the remaining metrics, with a slight advantage of the former approach in the Precision metric. Just like in the previous cases, the classification models dominate the first five positions in all the metrics, but in terms of Recall and Utility the differences are smaller. Finally, the results on the 10-classes tasks are shown in Figure 2.18. The gap between the Accuracy score of both approaches is similar while the small advantage of the classification models on the Precision metric observed so far is strengthened. Curiously, the results regarding the Utility and Recall scores become even more balanced.
2.6 Comparison of Regression and Classification modelling approaches 39 Figure 2.15: The top five average ranking model variants of both modelling approaches (Classification and Regression) are forming a new set of variants, and their average rankings are re-calculated within this new set of models for tasks with 2classes only. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. Overall, the increase in the number of classes on the response variable is having little influence on the results. Actually, it seems that there is a striking difference, depending on whether the number of classes is higher than 2or not. If it is not, then both modelling approaches seem to produce quite similar results where it is hard to pinpoint any outperformance of one over the other. This could be explained by the fact that the usage of cost-benefits has little room to significantly improve the performance of the standard
40 Actionable Forecasting Figure 2.16: The top five average ranking model variants of both modelling approaches (Classification and Regression) are forming a new set of variants, and their average rankings are re-calculated within this new set of models for tasks with 3classes only. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. models. With 3or more classes, a gap between the Accuracy of both approaches appears, and is kept constant regardless of the number of classes. In terms of Precision, a slight advantage for the classification models was observed, that is also kept fairly constant across the different number of classes. However, when evaluating the Recall and Utility of both approaches, it seems that the increase of the number of classes (decisions) makes both
2.6 Comparison of Regression and Classification modelling approaches 41 Figure 2.17: The top five average ranking model variants of both modelling approaches (Classification and Regression) are forming a new set of variants, and their average rankings are re-calculated within this new set of models for tasks with 5classes only. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. approaches more and more balanced. The results from both analysis (top modelling variant per task and this last one) are coherent. The first has produced strong results suggesting an advantage of the classification modelling approach. In the latter, even though the regression approach was able to become more competitive in some specific cases (for instance when considering the Recall or the
42 Actionable Forecasting Figure 2.18: The top five average ranking model variants of both modelling approaches (Classification and Regression) are forming a new set of variants, and their average rankings are re-calculated within this new set of models for tasks with 10 classes only. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. Utility for tasks with more classes), the best average ranking model variants were always belonging to the classification approach. Both studies point in the same direction from different perspectives, strengthening the validity of our conclusions.
2.7 Conclusions 43 2.7 Conclusions This chapter has presented a comparative study of two different approaches to deal with the actionable forecasting problem. The first, and more conventional approach, uses regression tools to forecast the unknown numeric value and then uses some decision rules to choose the "correct" decision based on these predictions. The second approach tries to directly forecast the "correct" decision. We have compared both approaches on a diverse set of generic tasks. We have also used an extensive set of modelling tools, considering a large set of parameter variants for each one, in order to guarantee some robustness of our conclusions. Overall, for these generic actionable forecasting tasks, we have observed some consistent results that we now summarise. If the user is willing to perform an extensive search of the best model types and parameter variants for each specific task, then for most tasks the best results were obtained using the classification modelling approach. If extensive model tuning is not an hypothesis for the user then our experiments lead to two main conclusions. If the goal consists solely on reaching the best possible ratio of correct decisions (i.e. Accuracy), then a classification tool should be selected, mainly in tasks whose response variable domain has at least 3classes (decisions). On the other hand, one may want to consider the severity of the mistakes associated to some prediction, or even consider that the possible actions/decisions to make are not all equally important. In that case, both modelling approaches seem to be quite competitive against each other, regardless of the number of classes of the response variable. Given the large set of classification and regression models that were considered, as well as different approaches to the learning task, we claim that there is significant experimental evidence to support these conclusions. The experiments carried out in this chapter have also allowed us to draw some other conclusions in terms of the use of cost-benefit matrices. In a generic context, as in this analysis, the use of cost-benefit matrices as an effort to maximise the utility of the predictions of the models, has clearly enhanced the performance on some metrics for a high percentage of modelling variants, without seriously compromising any other metric. Knowing that the models with costs obtained several significant wins over the models not using them, mainly on tasks with more classes, and considering that the regression modelling approach has managed to be competitive across all the tasks, we could say that the regression approach should outperform the classification approach, mainly on tasks with more classes, if classification models were not used with costs.
50 An Application to Financial Trading analysis will be similar to the one carried out in Section 2.5 of the previous chapter, but with three cases instead of one: i) Usage of Smote on classification models; ii) Usage of Smote on regression models; iii) Usage of cost-benefit matrices on classification models. 3.3.1 Hypothesis 1: Re-sampling the data sets - Classification In this section we test our first hypothesis that states that re-sampling the data sets with the goal of balancing the response variable will enhance the performance of classification models. We start by comparing the top modelling variant obtained without using resampling against the best one using Smote, company by company, applying a Wilcoxon test in each case. In Figure 3.1 we analyse recall and precision for the buy and sell signals combined. The results were somewhat expected. The models applied to the data set with re-sampling achieved significantly better results in terms of recall and worse in terms of precision. This means that this implementation is making the models more capable of detecting the rare classes, which makes sense since the respective training data has more buy and sell signals. However, this increase of recall comes at the cost of making more mistakes, leading thus to worse results in terms of precision. Since we are in a trading market context, we should prefer safer decisions rather than riskier ones. Without checking the impact in terms of financial results, we can not yet conclude if the usage of Smote is advantageous in our problem, but these results could still be interesting in a more general problem, where detecting the rare class may be of an extreme importance, such as diagnosis problems. Figure 3.1: Best classification variant without Smote against the best classification with Smote for the macro versions of Precision and Recall metrics (asterisks denote that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05).
3.3 Hypothesis testing 51 Figure 3.2 allows us to analyse the financial consequences of the re-sampling procedure. We can observe that in terms of Total Return and Sharpe Ratio the models applied to data sets without any re-sampling achieved significantly better results and, in several cases, by a large margin. This means that the results of pre-processing the data using re-sampling is leading to models that make more risky decisions and that these decisions are often wrong, leading to catastrophic losses. Figure 3.2: Best classification variant without Smote against the best classification with Smote for the Total Return and Annualised Sharpe Ratio metrics (asterisks denote that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05). Considering all the plots at the same time, our results provide evidence that using Smote on the training data will make the models predict significantly more trading signals (Buy and Sell). This leads to higher recall but unfortunately also to much lower precision because these signals are frequently wrong, with serious financial consequences. However, we should not forget we are looking at the best variant of each type (with and without Smote). Although we may be tempted to think that the behaviour may happen with all the models overall, we have no solid evidence to state that. Therefore, we will test this hypothesis in another way. Similarly to the analysis of the usage of cost-benefit matrices for generic tasks, we will group the variants according to the type of model and analyse the impact of using Smote per type of model. In other words, we will compare the SVM variants applied to data sets without re-sampling against the same SVM variants where the Smote was applied. The procedure is then repeated for the other type of models. We will compare the top modelling variant of each group as well as the median score (per group).
52 An Application to Financial Trading Regarding the top modelling variant of each type, the results are shown in Figure 3.3. The first thing to notice is that the recall was significantly better every single time across all the type of models where re-sampling was applied. On the other way, the precision was worse almost every time, except on SVM’s. There is no doubt that the use of the Smote algorithm is making the models more capable of detecting the buy and sell classes at the cost of riskier decisions. Concerning the financial metrics, we observe a clear advantage of the standard models (i.e. no re-sampling applied), with SVM and AdaBoost being the only models benefiting from the use of re-sampling on some tasks. Figure 3.3: Segmented by type of model and by metric, a Wilcoxon test is performed between the best model variant of each modelling tool (Classification without Smote vs Classification with Smote). Since there are 12 data sets in each segment, then there are 12 results per segment. Each bar shows the results for each segment, where each one of the five colours is associated to a type of win (significant/non-significant win without Smote - strong/light blue, draw - yellow, non-significant/significant win with Smote - light/strong green) The length of each colour describes the number of times that type of win occurred. Overall, the conclusions from this study are similar to the ones from the first analysis. However, in both cases (analysing all types together or each individual type), we have looked to the top modelling variant, meaning that this behaviour may not be representative of all model variants for each type of model. In order to study this spectrum of variants,
3.3 Hypothesis testing 53 we will do a very similar study, but instead of considering the best variant of each type of model, we will consider the median score of all the variants inside that same group. Note that we no longer can use statistical tests, because we are not selecting a variant to be the representative of a group, but rather the median of the score of all the respective model variants. The results are presented in Appendix B.1. The conclusions are very similar to the ones using the best variant. There are some subtle differences though, namely on the financial metrics. The standard models were better almost every single time. This is an interesting observation, since there were twotypes of models that were benefiting from the re-sampling on some tasks when considering their top modelling variant. It seems that if one is not willing to do extensive model tuning in this type of trading tasks, then the usage of Smote is not to be recommended. Finally, we have also calculated, for every individual metric, the average rankings of each model across all the companies, making use of a post-hoc Nemenyi’s statistical test. By doing so, we can compare the most robust and adaptive modelling variants with and without the usage of Smote, since the best models from this study will be the ones that could perform well across all the companies. The results are shown in Figure 3.4. The model names with a “_1” are related to the standard versions where no Smote was applied, while the ones with “_3” were obtained using this re-sampling method. The conclusions are consistent with the previous parts of this analysis. While in terms of Macro-Recall (of the buy and sell signals), the models created using Smote can statistically outperform the others, the opposite happens in a clear way for all the remaining metrics. To sum up, we have collected sufficient evidence to conclude that re-sampling the data sets for the classification models will have a negative impact on their performance, namely when considering trading evaluation metrics. This behaviour was observed both when evaluating the top variant per task, the top variant per type model, the median score within each type of model, or the average rankings across all the metrics. Only the SVM, TREE and AdaBoost types of models have benefited from the re-sampling method on a small set of tasks. We will now carry out a similar study on the impact of re-sampling regarding the regression modelling approach. 3.3.2 Hypothesis 1: Re-sampling the data sets - Regression Once again we will start by analysing the top modelling variant of all the models without re-sampling against the top one using Smote. The results are shown in Figure 3.5, where we see a very similar behaviour as in the classification case regarding the recall and precision of the buy and sell signals. As before, when the re-sampling method was used, the models obtained a significantly better recall but again at the cost of making more mistakes thus leading to poor precision results.
54 An Application to Financial Trading Figure 3.4: The top five average ranking model variants of both modelling alternatives (Classification with and without the usage of Smote) are forming a new set of variants, and their average rankings are re-calculated. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. In Figure 3.6 we can check the corresponding financial results in terms of Total Return and Sharpe Ratio scores for regression models with and without the usage of Smote. Even though we observe that not using the re-sampling method leads to better results in every single case, the differences are definitely smaller than in the classification case.
3.3 Hypothesis testing 55 Figure 3.5: Best regression variant without Smote against the best classification with Smote for the macro versions of Precision and Recall metrics (asterisks denote that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05). This may be explained by the way we create artificial examples of the minority classes. In the regression data sets, when a case is created it is not grouped into a class. This may be a considerable advantage. Every single artificial observation (created by the Smote algorithm) is composed by a set of the artificially assigned values of the predictors and the respective value of the response variable. It is possible that for the given set of those explanatory values, the respective true response value would be different from the one created by the Smote method. While in the regression approach, this error is numeric and it is expected to be small, in the classification approach it may lead to a whole different class, thus eventually confusing more seriously the respective models. Proceeding as before, we will now look at the top modelling variants of each type of model. The results are given in Figure 3.7. In terms of Recall of the combined minority classes, the models with Smote were always better (or at least not worse), except one single time. However, there are several cases that are not statistically significant according to the Wilcoxon statistical test. Regarding Precision, the standard models are overall better, though the Neural Networks and Tree types did not verify this trend. Finally, considering the financial metrics all the models but the NNET, TREE and KNN were in-arguably better without the usage of re-sampling, while the others presented even results. Finally, we also considered the median score of each group in Appendix B.2. Similar behaviours are observed, with the main difference being the fact that the median score of the standard KNN and NNET models on the financial metrics out-performed the respective variants with Smote. If one is not willing to make an exhaustive search for the ideal
56 An Application to Financial Trading Figure 3.6: Best regression variant without Smote against the best regression with Smote for the Total Return and Annualised Sharpe Ratio metrics (asterisks denote that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05). parameters, then the usage of Smote is not advisable at all. Similarly to the testing of Smote on the classification models, we also consider the average rankings across all the companies making use of a post-hoc Nemenyi’s statistical test. The results are given in Figure 3.8. The model names with a “_1” are related to the standard versions where no Smote was applied, while the ones with “_2” were constructed using this re-sampling method. The observed scores in terms of Recall and Precision are coherent with what we have seen before. The usage of resampling significantly boost the Recall level at the cost of a severe decrease in terms of Precision. Regarding the trading metrics, the results are surprisingly more even, though the standard modelling variants were always superior. In summary, our experiments provide strong empirical evidence that using Smote to try to overcome the limitation of few existing extreme variations on the prices is not improving the results of the resulting models. Equivalently to the classification case, using the re-sampling algorithm increases the number of buy and sell signals forecasted by the regression models, but at the cost of incorporating too much risk. Therefore, in terms of the trading metrics, the best results across all the types of analysis are obtained when not using Smote. However, the TREE and KNN types of model could actually benefit from this feature on some tasks.
3.3 Hypothesis testing 57 Figure 3.7: Segmented by type of model and by metric, a Wilcoxon test is performed between the best model variant of each modelling tool (Regression without Smote vs Regression with Smote). Since there are 12 data sets in each segment, then there are 12 results per segment. Each bar shows the results for each segment, where each one of the five colours is associated to a type of win (significant/non-significant win without Smote - strong/light blue, draw - yellow, non-significant/significant win with Smote - light/strong green) The length of each colour describes the number of times that type of win occurred. 3.3.3 Hypothesis 2: Adding cost-benefits This section considers the second hypothesis we have put forward, namely that the usage of cost-benefits matrices will boost the performance of the classification models, since the information on the implicit ordering among the classes will be passed to the models. We have used the following procedure to obtain the cost-benefit matrices for our tasks. Correctly predicted buy/sell signals have a positive benefit estimated as the average return of the buy/sell signals in the training set. On the other hand, in the case of incorrectly predicting a true hold signal as buy (or sell), we assign it minus the average return of the buy (or sell) signals. Basically, the benefit associated to correctly predicting one rare signal is entirely lost when the model suggests an investment when the correct action would be
58 An Application to Financial Trading Figure 3.8: The top five average ranking model variants of both modelling alternatives (Regression with and without the usage of Smote) are forming a new set of variants, and their average rankings are re-calculated. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. doing nothing. In the extreme case of confusing the buy and sell signals, the penalty will be minus the sum of the average return of each signal. Choosing such a high penalty for these cases will eventually change the model to be less likely to make this type of very dangerous mistakes. Considering the case of incorrectly predicting a true sell (or buy) signal as hold, we also charge for it, but in a less severe way. Therefore, the average of
3.3 Hypothesis testing 59 the sell (or buy) signal is considered, but divided by two. This division was our way of “teaching” the model that it is preferable to miss an opportunity to earn money rather than making the investor lose money. Finally, correctly predicting a hold signal gives no penalty nor reward, since no money is either won or lost. Table 3.2 shows an example of such cost-benefit matrix that was obtained with the data from 1981-01-05 to 2000-10-13 of Apple. According to this particular evolution of the Apples’ shares, correctly forecasting a sell signals leads to an average profit of 4.9%, while correctly predicting a correct buy signals leads to a profit of 3.3%. Trues s h b s 4.9 -4.9 -8.2 Pred h -2.4 0.00 -1.7 b -8.2 -3.3 3.3 Table 3.2: Example cost-benefit matrix for Apple shares. Using this matrix and the probabilities of observation belonging to any of the classes that is produced by the used model, we can decide the final predicted class. Basically, for each possible class c, we multiply the probabilities vector by the line of the cost-benefit matrix corresponding to the class c. For each class, we obtain the expected reward if we forecast c. Therefore, the prediction will be the class with the highest expected reward. In Table 3.2, we can see how using this matrix may change the final predictions for 4 illustrative test cases. The first three columns are the obtained probabilities for each class by some classification model. The Original Pred is the class prediction of the model without using the cost-benefit matrix, which merely consists of selecting the class with highest probability. The True Class column is the correct signal while the New Pred is the class predicted if using the cost-benefit matrix together with the probabilities. Prob of s Prob of b Prob of h True Class Original Pred New Pred 0.21 0.42 0.37 s b h 0.09 0.21 0.70 b h h 0.09 0.44 0.47 h h b 0.17 0.44 0.39 h b h Table 3.3: Illustration of the usage of the cost-benefit matrix into the original predictions of a classification model. At the first line we have successfully prevented the most harmful mistake possible, the confusion between the sell and buy signals. Note that in this case, the original prediction was a buy signal, while the correct one was the sell. By using cost-benefit matrix, we could change the outcome to a hold signal, which despite being still wrong, it is definitely preferable. In the second observation, no changes have occurred in the model outcome. In
66 An Application to Financial Trading Figure 3.14: Best classification variant against the best regression one for the Total Return and Sharpe Ratio metrics (asterisks denote that the respective variant is significantly better, according to a Wilcoxon test with α= 0.05). From an economical perspective we have observed some contradictory results. For instance, for some companies (e.g. Meg) it was possible to achieve a very high level of Total Return (above 60% return), but the maximum Sharpe Ratio that was achieved was very low. This means that the best model for the first metric was taking enormous amounts of risk and that the high level of return achieved was probably due to pure luck. On the other hand, there are some companies (e.g. Exas) for which both high values of Total Return and of Sharpe Ratio were reached, though we are not sure if these were achieved by the same model. Given the high variability of the results across companies, taking conclusions solely based on the analysis of the best variant per model and per metric may be misleading. This establishes the motivation for the second part of our experiments. 3.4.2 Post-hoc Nemenyi test: Average Ranks In this second part of our experiments, instead of grouping by metric and company, we will just group by metric and study the average rank of each approach across all the companies (top 5 of each approach are considered). With the use of the Friedman test followed by the post-hoc Nemenyi test, we check whether there are statistically significant differences among these average rankings of the top 5 variants of each approach. This way, if an approach obtains a very good result for one company but poor for all the others (meaning that it was lucky in that specific company), its average ranking will be low allowing the
3.4 Comparison of Classification and Regression modelling approaches 67 top average rankings to be populated by the true top models that perform well across most companies. Figure 3.15 shows the results of this new comparison in terms of Precision and Recall. The results are consistent with the previous comparison. Across all the tasks simultaneously, the classification models are achieving on average better results when it comes to capture the highest amount of sell and buy signals while the regression models are making correct decisions more often when they forecast one of those two classes. However the difference in terms of Recall seems to be stronger (Nemenyi’s test is statistically significant for some cases) than in terms of Precision. One should also note the following: the top 5average-ranking models of either approach in the Recall metric is solely composed by models that were constructed using the re-sampling method Smote. This is theoretically expected, since the training set of these models was modified until all the classes were equally represented. Therefore, the model is more likely to predict the buy and sell signals, leading to higher values of recall, but also with the potential of making more mistakes (this behaviour was also observed when studying the results of applying Smote). Figure 3.15: The top five average ranking model variants of both modelling approaches (Classification and Regression) are forming a new set of variants, and their average rankings are re-calculated. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. Figure 3.16 summarises the results in terms of Total Return and Sharpe Ratio. In both metrics, since we could not reject the Friedman’s null hypothesis, the post-hoc Nemenyi’s test was not performed. This means that we can not say with 95% confidence that there
68 An Application to Financial Trading is some difference in terms of Total Return or Sharpe Ratio between these modelling approaches. Nevertheless, there are some observations to remark. Regarding the first sub-image, i.e. the Total Return metric, the model with the best average ranking is a classification model using cost-benefit matrices. All the remaining classification variants are in their original form (without using cost-benefit matrices) and occupying mostly the last positions. Moreover, not a single variant obtained using Smote appears in this top 5 for each approach, which means that we confirm that re-sampling does not seem to pay off for this class of applications due to the economic costs of making more risky decisions. Furthermore, another very interesting remark is that all the top models are using SVMs as the base learning algorithm. Overall, we can not say that any of the two approaches to actionable forecasting is better than the other in terms of Total Return. Figure 3.16: The top five average ranking model variants of both modelling approaches (Classification and Regression) are forming a new set of variants, and their average rankings are re-calculated. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test. We analyse now the second sub-image, which shows the results of the same experiment in terms of Sharpe Ratio, i.e. the risk exposure of the alternatives. The conclusions are quite similar to the Total Return metric. Once again, no significant differences were observed. Still, one should note that the first 5 places are dominated by the classification approaches. The best variant for the Total Return is also the best variant for the Sharpe Ratio, which makes this variant unarguably the best one of our study when considering the 12 different companies. Hence, ultimately we can state that the most solid model belongs to the classification approach using an SVM with cost-benefit matrices, since it obtained the highest returns with lowest associated risk. Finally, unlike the results for Total Return, in this case we observe other learning algorithms appearing in the top 5 best results.
3.5 Conclusions 69 3.5 Conclusions This chapter presented a study of two different approaches to financial trading decisions based on forecasting models. The first, and more conventional approach, uses regression tools to forecast the future evolution of prices and then uses some decision rules to choose the "correct" trading decision based on these predictions. The second approach tries to directly forecast the "correct" trading decision. This study is a specific instance of the more general problem of making decisions based on numerical forecasts detailed on the second chapter of the thesis, that we have named actionable forecasting. We have focused on financial trading decisions because this is a specific domain that requires specific trade-offs in terms of economic results. Overall, the main conclusion of this study is that, we can not state that one approach performs significantly better than the other in the context of financial trading decisions. The scientific community typically puts more effort into the regression models, but this study strongly suggests that both have at least the same potential. Actually, the most consistent model we could obtain is a classification approach. Another interesting conclusion is that, of a considerably large set of different types of models, SVMs achieved better results both when considering classification or regression tasks. Given the large set of classification and regression models that were considered, as well as different approaches to the learning task, we claim that this conclusion is supported by significant experimental evidence. The experiments carried out in this section have also allowed us to draw some other conclusions in terms of the applicability of re-sampling and cost-benefit matrices in the context of financial forecasting. Namely, we have observed that the application of re-sampling, although increasing the number of trading decisions made by the models, would typically bring additional financial risks that would make the models unattractive to traders. On the other hand, the use of cost-benefit matrices in an effort to maximise the utility of the predictions of the models, did bring some advantages to several modelling variants. It is interesting to compare the conclusions of this specific study with the comparisons in a general setting described in the previous chapter. Were the conclusions from both studies consistent? The following are some of the most interesting observations from both studies: •The usage of cost-benefit matrices was beneficial in both cases (generic and trading problems); •While in the generic study the models with costs were better almost every time, in the trading problem the standard versions were more frequently better; •In the generic study, the RandomForest type was unarguably the best modelling type, regardless of the modelling approach, while in the trading problem problem the SVM has occupied that place;
70 An Application to Financial Trading •The classification modelling approach had some considerable advantages in the generic study. In the trading problem, both approaches seem to be equally competitive. The tasks considered in the trading problem present several distinct features from the ones in the generic problem, such as the temporal property as well as a strong unbalance of the response variable. These differences may explain some of the observed differences between the results of both chapters.
Chapter 4 Optimal Trading Signals 4.1 Introduction In this thesis we have started by presenting an extensive comparative study between two possible approaches to Actionable Forecasting. This initial study aimed general tasks where decisions must be made based on numeric forecasts. We then focused on a particular instance of these applications: financial trading. This application has several particularities and it is sufficiently important to deserve this special treatment. In this chapter we continue our study of financial trading, but we now turn our attention to the key issue of how to evaluate the trading decisions. Given a historic record of trading signals/decisions for the time-series of the prices of some company assets, it is not trivial to evaluate the quality of those signals. Merely looking at the total return obtained is not recommended since high returns may be obtained with large periods of time in which the investor would see their possessed assets with decreased value, i.e. the high return may be achieved at the cost of high variability of the returns (higher risk from the investor’s perspective). In order to deal with this problem, there are other tools that can incorporate certain notions of risk, as the Sharpe Ratio , but typically, one can never look at a single metric and draw definitive conclusions. An analysis of several trading metrics must be conducted in order to build up some confidence regarding a trading system. There is another important drawback of existing metrics. The investors have different levels of risk aversion as well as different target returns, thus leading to the preference of different trading policies. To the best of our knowledge, there are no trading metrics that may be adapted to distinct trading policies, which means that each investor will have to perform a very subjective analysis from the scores of several standard trading metrics. In this chapter, based on the work of Torgo and Dhar (2004) ,we propose a solution to this problem, that allows the establishment of which are the optimal trading actions given a certain target trading profile. These ideal actions can be used as a new form of evaluating trading systems by benchmarking their actions against this optimal performance. 71
72 Optimal Trading Signals We claim that the feedback provided by this benchmark is an interesting decision tool when evaluating trading performance. We present illustrative examples of the use of this benchmark for evaluating trading records. 4.1.1 Relationship with Activity Monitoring The proposals described in this chapter are motivated by the concept of Activity Monitoring (Fawcett and Provost,1999), more specifically by trying to see trading as an instance of these tasks. In activity monitoring data mining tasks one tries to find the right timings for issuing alarms for the so called positive activity that can be seen as target events one wants to signal timely. A standard example consists of detecting the presence of an intruder into someone’s house. There is a period of positive activity, which is while the intruder is inside, and the goal is to set up an alarm the closest possible to the beginning of this positive activity. In this example, special attention must be paid to false alarms which are certainly unwanted. In this chapter, we show that financial trading can be described in a similar way with positive activity being the time windows where successful trades are possible, with a varying investor’s definition of success. Given that high profits with minimal risks are the key issues for a successful trading record, we have used two metrics to capture these two properties as the means for defining a trader’s definition of success. More specifically, we ask the trader to indicate the minimum wanted return for each individual trade (for instance to cover the transaction costs) and the maximum draw dawn she/he is willing to take (as a way to specify the maximum period of successive losses the trader is willing to accept). In this context, a search for what would be the optimal timings regarding trading actions can be conducted. This search would be done maximising the total profit and making sure that the investor’s trading policy is verified (in terms of the minimum return per trade as well as not surpassing the defined level of maximum draw down). It is also possible to find all the timings where a long/short position could be opened knowing that it would be possible to successfully close it (according to the same trading policy). Several uses may be given to this new perspective of trading. First of all, the performance of a real trading system may be compared against the performance of the optimal trading record (where these optimal signals would depend on a trading policy). Secondly, scoring functions may be used to measure the quality of the timings produced by the real trading system. This score could, for instance, also be used as the target variable of a financial forecasting system. In the thesis we focus on the first of these uses: traders benchmarking. We describe how to use this formalisation to obtain a benchmark score that we claim it is highly interpretable in the sense that it provides information on how distant is any trader from the optimal trading record given a certain target trading profile.
4.2 Definitions and Concepts 73 4.2 Definitions and Concepts One of the main applications of our proposal is the creation of an adaptive benchmark that consists of the optimal trading positions according to the trader’s preference criteria. This “adaptive benchmark” depends on the trading policy of the user, that will be defined according to three criteria: •Minimum wanted return - The minimum profit percentage that the investor would like to obtain per trade; •Maximum Draw Down Allowed - It measures the largest peak-to-trough decline in the value of a portfolio (before a new peak is achieved). Suppose a long position is opened with a price p. At every point in time we consider the maximum value pmax that was obtained between opening the position and the present time. If at any moment, the present value of the asset is lower that the maximum draw down allowed times pmax, then the maximum draw down allowed level is breached. In other words, it describes how much devaluation we are willing to take before closing the position in order to prevent bigger losses. This criteria describes somehow the risk aversion of the client; •Transaction Cost - A percentage cost that is assigned to a trade that closes a long or short position. For instance, consider the opening of a long position with a price equal to 90 and the closing with a price of 100. With a transaction cost of 1%, the true return would be ((1 −0.01) ×100) −90 = 99 −90 = 9. In order to illustrate the influence of the minimum wanted return and the maximum draw down allowed, we show in Figure 4.1 an example that would lead to a forced sell. The yaxis represents the values of the assets, with dalpha being the opening price of a long position. Since the growth of the price was not enough to surpass the target minimum return(rg), we could not close the long position. When the asset’s price reached the dmax value, a fall was observed. This fall was high enough to surpass the maximum draw down allowed (mxDD), forcing the close of the long position in order to prevent more losses. Different values of minimum wanted return and maximum draw down allowed could have made this unwanted situation into a more desirable one. Decreasing the value of the first criteria (target return) could have let the position to be closed before the fall of the price or, on the other hand, different values for the maximum draw down allowed could have lead to fewer losses (by having a lower mxDD which would force the sell sooner) or allowed the asset’s value to increase once again (with a higher mxDD, since the investor would not sell at that timing) . We believe that with these variables we can make a very realistic approximation of the many possible investor’s trading policies. In this example, no transaction costs were used. In order to apply the methodology being described in this chapter, it is important to know the timings of all the trading actions settled by a trading system. In our study we
74 Optimal Trading Signals Figure 4.1: A long position opened by a false alarm. This image was taken out from Torgo and Dhar (2004) consider two types of trading positions: long and short. The first is opened by issuing a Buy order to obtain assets/securities and it is closed when a Sell order is issued. The goal is to sell those assets/securities at some time in the future at a higher value, thus leading to a profit. On the other hand, a short position is opened by issuing a Sell order (even though the investor may not possess any asset, which is possible through a borrowing system) and it is closed when the investor buys the securities. In other words, the investor sells a security at the price of the current time, with a promise to buy the respective assets at some point in the future. If the security’s value decreases during that time, then a profit is obtained. The key for successful trading is to open and close these positions at the right timings. 4.3 Optimal Positions We will describe the procedure to construct the optimal benchmark according to a certain trading policy. For a certain evolution of the prices of some assets, the goal is to construct the set of long and short positions (the moments of opening and closing) that would be ideal for a certain trading policy. We will describe the procedure for obtaining these ideal timings separately for each type of position. Please note that this concept of optimal positions can only be implemented a-posteriori, i.e. the optimal positions can only be calculated for a past period, after we know what happened to the market prices, which means that there is no forecasting involved in this benchmarking procedure. 4.3.1 Long Positions In order to obtain the optimal timings for executing long positions, one must look at the periods where there are no draw downs larger that the one specified in the trading policy,
4.3 Optimal Positions 75 mxdd, and search for the opening and closing timings within that period that lead to the largest profit, as long as that profit is higher than the target return rg. Algorithm 1, that was firstly conceptualised in the work of Torgo and Dhar (2004), returns a vector with the optimal trading record regarding long positions. For every unit of time, the returned vector will have one of three following values: {“hold”,“open”,“close”}. The first means that no trading actions should have occurred at that timing, i.e. no positions should be opened nor any active position should be closed. The second means that a long position should have been opened on that timing while the latter means that the current long position should be closed. This algorithm is very efficient with a computational complexity of the order of O(n). We know briefly describe the intuition behind the algorithm. Variables “mn” and “ idx.mn” are related to the minimum price observed during a certain period. While the first is the value itself, the latter is the timing where the minimum has occurred. Similarly, “mx” and “ idx.mx” are related to the maximum observed in the same period, where the first is the value and the second is the timing where that price has occurred. As detailed before, we need to look for periods between draw downs larger than mxDD, where a return larger than rgoccurs. This search consists in finding the maximum and minimum prices within those periods. The variables just described will contain the information regarding those extreme prices during each potential period to trade. From line 3to 9, the algorithm tests if the current timing corresponds to a new maximum. If so, we immediately update the variables that store the information regarding the maximum. It is also tested whether the difference between the new maximum and minimum value stored in “mn” is high enough to cover that target return plus the transaction cost. In that case, we update the variable “gain” to have the value “True”. At this moment, opening a long position at the timing “idx.mn” and closing at the timing “idx.mx” would correspond to a successful trade. However, it may not be an optimal one, since the market may still go up in the following timings. While we wait to check if this happens, we say that there is a standby long position. Lines 10 to 15 test if the current timings corresponds to a new minimum. If they are, then we are now sure that any eventual standby long position (produced in the previous lines) may now be considered as an optimal one. This happens because if a new minimum occurs, it would not make sense to have a long position starting before the timing of the minimum and ending after it. Therefore, if the variable “gain” is equal to True, we add a new long position to our final vector of signals, using the timings of minimum and highest prices observed before. Moreover, we update our minimum and maximum temporary variables to the current timing, somehow “restarting” the process of finding the next optimal long position. Finally, the maximum draw down component is taken care between lines 17 and 21. The first condition tests if the maximum draw down requirement was breached. If that is the case, then we can no longer wait for the market to go up. Therefore, we consider any eventual standby long position as an optimal one. Similarly, when a new minimum is
82 Optimal Trading Signals Figure 4.4: Positive Activity of S&P 500 between July of 2012 and April of 2013 for two different trading policies. or not, and if they do, how close would that timing be to the optimal time for the signal. Using that score, one may evaluate a trading system by seeing the scores associated to the timings that the system has opened a position. This would be a whole new perspective of evaluating a trading system. Not only would it be quite adapted to a trading policy, but also it would focus on the proper timings rather than merely at the distribution of the returns. We list some scoring functions that may be considered for long positions, starting by a trivial example:. •Function a) - Positive Activity,
4.4 Positive Activity for Trading 83 f:Tl→ S ={0,1} t7→ 1, t ∈Pl 0, t /∈Pl . with Tlrepresenting the set of timings that the system has opened a long position and Plthe activity period of a long position. This score function says whether a long position was opened in a time of positive activity or not (and therefore if it would be possible to obtain a profit or if a forced sell would inevitably occur). •Function b) - Temporal distance to the closest optimal signal, f:Tl→ S ={0,1} t7→ 1−d(t, topt), t ∈Pl 0, t /∈Pl . where d(t, topt)would be a normalised distance function between the timing tand the closest optimal timing for opening a position. The score of each point in the positive activity period would be decreased as higher the distance to the closest optimal timing for opening a long position, or zero if a forced sell would have to inevitably occur. •Function c) - Ratio between the potential return of a time twith the closest optimal signal, f:Tl→ S ={0,1} t7→ r(t) r(topt), t ∈Pl 0, t /∈Pl . where ris a function that measures the potential return that may be obtained by opening a long position at the time t. This potential return is measured by considering the closing position that would lead to the largest return possible without compromising the trading policy used. The score of each point within the positive activity period would increase as higher as the potential return of that time compared to the return associated to closest optimal position. •Function d) - Ratio between the potential return of a time twith the closest optimal signal also considering the losses,
84 Optimal Trading Signals f:Tl→ S ={0,1} t7→ r(t) r(topt), t ∈Pl −rloss(t) + rg, t /∈Pl . where rgis the target return and the function rloss(t)calculates the maximum value of the security’s prices between the timing tand the moment of forced sell. In other words, the score assigned to timings that do not belong to a positive activity period can be interpreted as how much more positive variation of the prices was needed in order for the timing tto be successful according to the used trading policy. Any of these functions can be used to assign a score to the timings for opening long positions issued by any trading system. Analogous functions may be created for short positions. Using the set of scores obtained by a trading system along time, one may conduct a study regarding the distribution of the scores of this trading system. Depending on the chosen function, one may carry out statistical tests to check if that distribution of the scores of the trading system has some desired properties. For instance, suppose we are considering the trivial function described above (0or 1depending if the timing belongs to the positive activity period). In this case, we could be interested in carrying out a statistical test to check if the average would be significantly higher than 0.5. We claim that the proposed methodology based on looking at trading as activity monitoring may open several interesting new avenues regarding ways of evaluating a trading system performance, particularly by allowing to check this performance against some user preference bias (i.e. the user trading policy). Another potential interest of these scoring functions is that they may enable new ways of constructing trading systems. Given an historical record of some security’s prices, we may calculate the score for opening a position on each point in time according to the trading policy of the investor. Then, those scores could be used as the target variable of a machine learning model. The obtained model could forecast the score for some future time stamp and based on this score trading decisions can be made. 4.5 Optimal Signals as a Trading Benchmark We have described how to obtain the optimal signals for some past period of time and given the trader’s preference biases. The motivation for finding these signals would be to create a benchmark that an investor could use to compare several alternative trading systems and choose the most adequate according to his trading policy. In this section we will explain how to use the optimal signals as a benchmark, list some of the existing
4.5 Optimal Signals as a Trading Benchmark 85 evaluation metrics, and describe a practical application of the proposed benchmark to the signals of a real trading system. 4.5.1 How to use the Optimal Benchmark We now describe how the optimal trading signals can be used to evaluate a trading system. We assume that this trading system was applied to some (past) testing period. Using some specific target trading policy (defined by the target return and maximum allowed draw down) we calculate the optimal trading systems using the algorithms described before. Having these optimal signals and the signals produced by a concrete trading system, we propose the following methodology to evaluate this trading system according to the optimal portfolio: - For any trading evaluation metric (e.g. total return, profit factor, sharpe ratio, etc.), we calculate two values: (i) the value obtained by the signals of the system being evaluated; and (ii) the value obtained by the optimal trading signals. Using these two values we calculate the ratio between them. We claim that these ratios are highly interpretable for a trader because they reflect how far is the system being evaluated to the ideal signals according to the trader target policy. For instance, assume we are considering some risk measure, such as the well known Sharpe Ratio. Suppose the ratio of the Sharpe score of the trading system over the Sharpe obtained by the optimal signals is 0.1. This means that the trading system has 90% more risk exposure than the score that would be possible to obtain on the testing period, assuming the trader targets in terms of return and maximum draw down. We claim that this proposed evaluation methodology is quite informative, highly interpretable and easy to implement. Our proposed benchmark may also open doors for new metrics. To the best of our knowledge, there are very little metrics that try to measure in some way, how distant is a concrete trading record from what would be ideal performance by an investor. Usually, it is the opposite way, where the trading records are evaluated according to how better they are against a very simple benchmark. 4.5.2 Standard Metrics for Evaluating Trading Systems We have suggested the use of several standard trading evaluation metrics to obtain ratios over the score of the optimal trading actions. In this section we describe some financial metrics that could be used in this context. We will divide our metrics into two groups: (i) metrics that only look at the final return upon closing a position; and (ii) metrics that consider the risk of the portfolio of the investor. Namely, this second group of metrics considers the daily variation the investor’s capital, depending on if a long/short position is active. If no positions are active at a certain time, then the daily variation on those periods is zero (theoretically, there is a discount rate that decreases the “value of money” over time, that we will ignore for simplicity purposes).
86 Optimal Trading Signals The following list corresponds to a small set of the existing standard metrics within the first group (i.e. the return of each trade). We consider the return as the profit percentage of a position calculated using the opening and closing prices, after discounting the transaction costs.For simplicity purposes, we assume that the amount of assets traded per position is always equal. •Profit Factor: It is defined as the ratio of the sum of the returns of the profitable trades over the sum of the losses. Values over one are desirable. •Payoff Ratio: The absolute ratio between the average return of a winning trade over the average loss of a loosing trade. However this metric has a limitation, since with a very reduced number of winning or losing trades, any of the respective averages may not be significant and lead to misleading results. Suppose a trading system has obtained several successful trades and only three not successful ones (with large losses). The average loss would thus be high leading to a poor payoff ratio, while the trading system was actually producing “good” signals. Nevertheless, this problem is only present for a small number of trades. Scores higher than 1 are also desirable. •Average Trade: The average of the returns obtained when closing a position. Naturally, a value higher than zero is a good indication, though zero would already imply that the system is at least being able to at least cover the transaction costs. •Winning Trades: The percentage of winning trades. •Expectancy: A metric described in Tharp et al. (2007). It tells you how much you should expect to make on average per currency unit at risk. Its formula is given by the average trade return over the absolute average loss of the portfolio. Therefore, a portfolio that achieved a 0.6expectancy score (which is considered to be a very good value), can be interpreted as “for every unit of currency risked, this system will generate on average 0.6units of currency”. •Total Return: It is the sum of all the returns (wins and losses included). A total return of 0.05 over a certain period means that the investor would have reached the end of this period with a 5% increase on his capital (assuming the hypothesis that every position had the same number of assets involved) All these metrics fail to penalise “successful” trades where quite high draw downs during the period of the position were achieved. In order to evaluate the risk of a trading system, one should look to daily (or some other unit of time) variation of the investor’s capital during the given period. In this sense, a return is now the variation of the capital from one unit of time to the next one. If no positions are active at a certain time, the daily variation will be zero. If a long position is active, then a positive return will occur when the asset’s price increases while if a short position is active, a positive return will be considered when that price decreases.
4.5 Optimal Signals as a Trading Benchmark 87 We list some metrics that evaluate the risk of a portfolio obtained using a certain trading system. Once again, we consider that the amount of assets bought and sold per position is always the same. For a more detailed description on portfolio evaluation metrics, we suggest the reading of the survey of Le Sourd (2007). •Average Variation: It is the average of the daily variation of the investor’s capital (returns); •Value at Risk: It measures the potential loss in value of a risky portfolio over the given period for a given confidence interval. We consider a 95% confidence level, meaning that this value will be calculated as the 0.05 percentile of the returns of this portfolio. Thus, if the VaR is, for instance, equal to −3%, then with 95% confidence level, this portfolio will not incur in a loss over 3%. It can also be interpreted in the opposite way, as a 5% probability of having a loss over 3%. Having a Value at Risk close to zero is a very good indication regarding the risk of the portfolio. For a very formal, precise and theoretical reading on this metric, we advise the reading of Peng (2009) •Tail Value at Risk: Also described in a very formal way by Peng (2009), this metric measures the expected loss knowing that a loss larger than an amount X has occurred. According to the author, this metric is a more robust and meaningful metric than the Value at Risk. We will choose the Value at Risk as X, meaning that this metric will be the same as the Expected Shortfall. A TVaR of −10%, states that, knowing an extreme loss will occur1, then in average that loss will be equal to 10%. •Sharpe Ratio: The most generic definition is given by E[R−Rb]/V [R−Rb], where Ris the vector of the returns of the portfolio under evaluation while Rbis the respective vector for a benchmark portfolio. Usually, the latter is considered to be a constant risk-free return, thus making V[R−Rb] = V[R]. This metric measures the excess return (over that constant benchmark portfolio) per unit of deviation in an investment asset or a trading strategy. For more information regarding the Sharpe Ratio, we recommend the reading of Lo (2002). In our application, we will use a zero constant risk-free return portfolio. In our application of our proposals to evaluate real trading signals, we will calculate ratios of all these metrics over the respective scores achieved by the optimal signals for several trading policies (note that there are some metrics, mainly regarding the returns of the trades, in which we can not calculate the score for the optimal signals, since they need to account for unsuccessful long/short positions, which will not exist in the optimal benchmark). 1considering an extreme loss as belonging to the 5% largest losses of the portfolio returns distribution
88 Optimal Trading Signals 4.5.3 An Application to Evaluate Real Trading Signals A set of real trading signals was obtained from an experienced trader and we shall use that opportunity to test the proposed benchmark. By using the benchmark the investor may adapt most evaluation tools to his own trading policy, thus allowing a better analysis of any trading system. Figure 4.5 shows the trading decisions of a specific trader for the S&P500 index, from the beginning of the year 1990 till 1991, where each decision (an arrow in the graph) corresponds to either selling or buying S&P500 securities. Figure 4.5: Real Trading Signals for one year period on the S&P500 index. Red arrows correspond to the timings in which assets were sold, while the green ones correspond to the timings in which assets were bought. In order to know the returns associated to his trade, we need to know the exact timings were every short or long position was opened and closed. Unfortunately, one can not state that every pair of consecutive signals in this figure forms either a long or short position. The reason is the fact that this trader may decide to carry out several successive orders of the same type (e.g. issue a Buy order at time tand issue another Buy at time t+xbefore the position opened by the first order is closed). Moreover, the investor has traded varying amounts of assets on each order, meaning that the total assets obtained by opening a long position at a certain time t, could be sold separately at different times t+h1,t+h2, etc. This means we had to do some pre-processing of the original data we were given by the trader, by considering at every point in time, a queue of assets that were owned by the investor but not traded yet. The need for this pre-processing results from the fact that our proposed method consists on comparing the performance of a set of positions (created by open and close orders), against the set of ideal positions (defined by the ideal signals).
4.5 Optimal Signals as a Trading Benchmark 89 As an illustrative example, suppose that a short position was opened in the first red arrow of Figure 4.5 issuing a Sell order of 100 assets. This amount of assets needs to be bought at sometime in the future in order to close the short position. Suppose now that in the upcoming green arrow, that correspond to a Buy order, the investor only buys 80 assets. Then, we consider that a short position was opened at the red arrow and closed on the green arrow with an amount of 80 assets. However, there are still 20 assets that need to be bought, meaning that we still need the next buy signal for those 20 assets and close a second short position. This means that in effect we have more than one short position opened at the first red arrow (two if the trader buys the remaining 20 assets all at the same time). We keep moving in our time line, but another red line appears (Sell signal). There are 20 assets that still need to be bought yet the investor wants to reinforce his position by opening another short position with, let’s say, 30 assets. After this moment, the investor has still 20 assets from the first red line to buy and now he has added more 30 to that amount. At last, suppose in the upcoming green arrow, the investor buys 90 assets. This amount is enough to fully close the short position opened at the first red arrow and also to close the short position opened at the second red arrow. This means that, after this trade, we consider two more short positions closed, leading to a total of 3short positions executed during this period. Note that after closing these positions, the trader has now 90−50 = 40 assets in his possession that need to be sold at some point in the future. That means the next red arrows will be used to close long positions opened at the current time until all the 40 assets are sold. Summarizing, the positions of this illustrative example are: •Short position opened at the first red arrow and closed at the first green arrow. Amount of 80 assets; •Short position opened at the first red arrow and closed at the second green arrow. Amount of 20 assets; •Short position opened at the second red arrow and closed at the second green arrow. Amount of 30 assets Note that after closing the above positions the trader holds 40 assets, so one or more long positions will still appear in his trading record. Once we have pre-processed the signals given by the trader, we can check what would be the optimal signals for this interval of time. In order to do that, we have to specify some trading policies. In all the cases, we will consider a transaction cost of 0.5%. We present four possible trading policies in Table 4.1 and the respective optimal trading signals for each one of them in Figure 4.6. In the trading policy A, the number of positions is the highest in contrast to their duration which are the shortest of all the four trading policies. The first is explained by the usage of a lower minimum wanted return, that allows more opportunities to successful
90 Optimal Trading Signals Minimum Wanted Return Maximum Draw Down Allowed A1% 0.5% B2% 1% C4% 2% D4% 4% Table 4.1: List of trading policies. trades to occur. However, the investor will not be able to keep a position for a very long time, since the maximum draw down allowed is quite low. This trading policy tends to favour traders who search for a high frequency trade. The trading policy B seeks for higher returns but also accepts more risk. This will lead to less trades, but longer and more profitable ones. Comparing to the previous one, we can see consecutive short (or long) positions fused, such as the first two short positions in the trading policy A that now form a single one in the new policy. We can also see periods of time now absent of any trading contrarily to the former policy. We are stressing this out, to reinforce the importance of the trading policy of an investor. It certainly leads to different records of optimal trades. In the last two trading policies, we have increased the minimum wanted return to 4% in both cases, but with two different levels of risk associated. While the first feels more natural, since the investor is looking for high returns assuming a certain and somewhat safe risk, the latter allows a maximum draw down equal to the minimum wanted return. The point here is to stress out the influence of the maximum draw down. Assuming a higher level of risk (i.e. higher maximum draw down allowed), and for the same target return, the investor will not mind waiting a considerable duration of time until he can achieve his target return. Therefore, we see longer positions, including some fusion of consecutive positions. With these four trading policies we think we cover a considerable set of real world trading scenarios used by traders and investors. Having the given signals (provided by a real trader) and the optimal signals created with different policies, we can now advance to the next stage. At first, we will analyse the given signals using the standard trading evaluation metrics. In a second step, we will analyse the ratios of the scores obtained by the original signals over the ones obtained by the optimal signals. Table 4.2 shows the scores of all the metrics described before, for the provided signals as well as for the optimal signals obtained for different trading policies, over one year period. The last four pairs of columns pay respect to the scores of the optimal signals for each trading policy as well as the ratio between the score of the trading system being tested over the optimal one. Let us first analyse the real trading signals (column “Signals”). Overall, it seems to be a decent trading system. The profit factor score indicates the returns of the trades
4.5 Optimal Signals as a Trading Benchmark 91 Figure 4.6: Optimal Trading Signals for one year period on the SP500 index. were 30% larger than the losses. However, the pay-off ratio of this trading system says that a winning trade of this trading system is on average equal to 90% of an absolute loss. It may seem that there is some discrepancy with the previous metric, but if there are more winning trades than losing ones, then both scores are compatible. The average trade return as well as the percentage of winning trades do not have impressive values, though they surpass the minimum values desired (0and 0.5respectively). Note that these values include the transaction costs. The result obtained in the Expectancy metric, suggests that for every unit of currency risked, it is expected that the system will generate 0.13 units of currency. In terms of the Total Return, assuming the investor would trade the same number of assets on very position, at the end of the testing period, the investor would have
98 Conclusions and Future Work Regarding the direct comparison of the two alternative approaches to actionable forecasting we have carried out different types of studies in order to enhance the robustness of our conclusions. We believe we have gathered enough evidence to state that overall the classification modelling approach can outperform the more frequently used approach based on regression tools. Whether the user is willing to make an extensive search for the optimal parameters to model a certain task or not, the results point in the same direction. The methods based on the classification approach tend to be better. The only setup where the regression approach was more competitive was on tasks with a high number of classes/decisions and where the user is not merely interested in the accuracy of the decisions and wants to consider different grades of severity of the decision errors. Having defined the general concept of actionable forecasting we have then moved into the analysis of a specific instance of these problems: financial trading decisions. More specifically, we have studied applications where a forecasting model tries to anticipate the future evolution of the prices of some financial assets and then a trading decision needs to be made, based on these predictions. This instance of actionable forecasting has some characteristics that make it quite different from the generic tasks studied before. Namely, we now are addressing a task based on data that is ordered by time (time series data), and our decisions are rather unbalanced, with the more important decisions being rare. Given these differences, we could not assume that the conclusions of the generic study would hold on the trading problem. Similarly to the generic problem, we have analysed the possible limitations of each modelling approach, but now for trading tasks. It is well known that most the modelling tools struggle to model unbalanced data sets. In order to deal with this problem, we have considered re-sampling our data sets using the Smote algorithm. We have tested this feature and observed that, overall, even though the models (classification and regression) could properly detect more often the important trading actions, they have also incorporated too much risk leading to serious economic losses. A few percentage of all the models (classification and regression) were able to improve their results by using this re-sampling strategy. We have also tested the usage of cost-benefit matrices to teach the classification models the danger of confusing a selling order with a buying one. We have observed once again, that using this feature improves the performance of several classification models, though in a less evident way than in the generic tasks. After the analysis of these additional methods that could improve the performance of the models we have finally compared the two approaches (classification and regression) on this particular instance of actionable forecasting. We have used 12 different companies to increase the robustness of our comparisons, and have again considered a large set of modelling tools and parameter variants. We have gathered enough information to conclude that there is no statistically significant difference between both approaches in the context of these financial trading decision problems. We have presented a third contribution on the thesis. Again in the context of financial
5.2 Future Work 99 trading, we have proposed a new perspective for evaluating trading systems. We have proposed a formalisation of the financial trading based on the existing framework of activity monitoring. Using this formalisation we have shown how to find the optimal signals with respect to a user-defined trading policy. These ideal signals allowed us to propose an optimal benchmark that is adapted to the user preferences (in terms of target return per trade and the maximum draw down allowed). We claim that this benchmark can be very useful to compare any trading system against it. We have described how to do it and presented a real application (with real data and real trading signals). We claim that this form of evaluating and comparing the performance of some set of trading systems is more informative to potential investors as it allows them to match these systems against the ideal performance according to their own trading preferences. To the best of our knowledge this is a novel way of evaluating trading systems. Finally, based on the proposed formalisation of financial trading, we have defined the concept of positive activity within these tasks. These are periods of time where issuing a signal is rewarding according to the trader’s preferences. This concept provides the bases for defining scoring functions that can be used to characterise/evaluate the signals of any trading system in terms of how far they are from the ideal timings. Moreover, these scoring functions also have potential to be used in other concepts like for instance in terms of trying to forecast their future value and use these predictions for trading. 5.2 Future Work Regarding the comparison of both the classification and regression modelling approaches to Actionable Forecasting tasks, we claim we have covered most of the key parts: (i) we have considered the limitations of each approach, by considering re-sampling and cost-benefit matrices; (ii) we have considered a set of metrics that cover any potential interest of the user (Accuracy, Utility Score, Recall, Precision); (iii) we have conducted distinct types of analysis that altogether lead to extensive and robust conclusions; (iv) we have considered generic non-temporal tasks and temporal tasks (trading only). This last point is, perhaps, the only one that can be strengthened. As some future work, we could open our analysis for generic temporal tasks or generic non-temporal but heavily unbalanced for instance. This would cover almost any actionable forecasting task setup. Still, an exception that was not covered in this thesis are problems where the decision, given a numeric forecast, is non-deterministic. Whether a similar study could be carried out for this other type of decision problems is another topic for future research. Regarding our proposal for an optimal benchmark and positive activity for evaluating trading systems, there is some future work to be considered. Concerning the optimal benchmark, the main path to follow from this moment is to focus on the creation of new metrics. The same way that metrics such as the Sharpe Ratio analyse the excess return over a null benchmark at the cost of an increase in terms of risk, one can also think of
100 Conclusions and Future Work metrics that analyse how close is the return to the optimal one and at which costs in terms of excess risk over the optimal one. Moreover, these decisions would depend on the trading policy of the investor, making these new metrics potentially more interesting to investors. In terms of the positive activity periods, we believe that some research should be carried out in terms of finding the most informative scoring functions and the characteristics of the distribution of the respective scores. The possibility of using these scoring functions in the context of prediction models is also a potentially interesting avenue for future research.
Appendices 101
Appendix A Actionable Forecasting - Generic Tasks Figure A.1: Segmented by type of model and by metric, the median score of each group is directly compared (Classification without costs vs Classification with costs). Since there are 8data sets in each segment, then there are 8results per segment. Each bar shows the results for each segment, where each one of the three colours is associated to a type of win (without costs - light blue, draw - yellow, with costs - light green) The length of each colour describes the number of times that type of win occurred. 103
104 Actionable Forecasting - Generic Tasks Figure A.2: The top five average ranking model variants of both modelling approaches (Classification without costs and Classification with costs) are forming a new set of variants, and their average rankings are re-calculated within this new set of models for tasks with 2classes only. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test.
Actionable Forecasting - Generic Tasks 105 Figure A.3: The top five average ranking model variants of both modelling approaches (Classification without costs and Classification with costs) are forming a new set of variants, and their average rankings are re-calculated within this new set of models for tasks with 5classes only. Each model is thus given an average ranking (x axis). The presence of at least one black line implies that the Friedman’s null hypothesis of all averages being equal was rejected. In that case, every pair of model variants not connected by any black line is considered to have their average rankings significantly different according to a Nemenyi’s statistical test.
106 Actionable Forecasting - Generic Tasks
Appendix B Actionable Forecasting - Trading Tasks Figure B.1: Segmented by type of model and by metric, the median score of each group is directly compared (Classification without Smote vs Classification with Smote). Since there are 12 data sets in each segment, then there are 12 results per segment. Each bar shows the results for each segment, where each one of the three colours is associated to a type of win (without Smote - light blue, draw - yellow, with Smote - light green) The length of each colour describes the number of times that type of win occurred. 107
114 BIBLIOGRAPHY Venables, W. N. and Ripley, B. D. (2002). Modern Applied Statistics with S. New York, fourth edition. ISBN 0-387-95457-0. Weiss, G. M. (2004). Mining with rarity: a unifying framework. ACM SIGKDD Explorations Newsletter, 6(1):7–19. Weiss, G. M. (2005). Mining with rare cases. In Data Mining and Knowledge Discovery Handbook, pages 765–776. Springer. Wilder, J. (1978). New Concepts in Technical Trading Systems. Trend Research.