scieee AI-readable full text Open interactive document viewer

Energy consumption flexibility analysis : data driven approach

Nogueira Sousa, Andre

Abstract

Demand Response can be defined as the ability of electricity suppliers to shift the consumption pattern of their clients aiming to balance the supply and demand of energy, especially during periods of high demand. The expected result is that the aggregated effect of compliant clients will add flexibility capacity to the power grid to avoid overloading the generation or distribution energy systems. One of the main challenges in a Demand Response event is to efficiently define the set of clients that are likely to react positively to a signal from the supplier requiring them to shift their consumption in a specific time range. With this goal, this project proposes a data driven approach for measuring the potential for flexibility of individual customers. A metric is calculated using historical consumption data and is based on the consistency of the load pattern of each client. This flexibility score is associated with a probability of positive response to a shift consumption request. Finally, given a requested consumption reduction in a time range of a specific date, the flexibility score and a forecasting model are used to predict the customers hourly consumption and therefore, the consequent possible reduction. This process is used to simulate the selection of a set of customers for a Demand Response event using real data from about one thousand customers from Estabanell y Pahisa, S.A. in Catalonia.

Full text

Master Thesis Energy Consumption Flexibility Analysis - Data Driven Approach Author: Andre Nogueira Sousa Advisor: Esteve Rodr´ıguez Delgado - Estabanell y Pahisa, S.A. Co-Advisor: Bernat Coma-Puig - Department of Computer Science Tutor: Jose Luis Balcazar Navarro - Department of Computer Science MASTER IN INNOVATION AND RESEARCH IN INFORMATICS DATA SCIENCE SPECIALTY Defense date: 28/06/2022 i Abstract Demand Response (DR) can be defined as the ability of electricity suppliers to shift the consumption pattern of their clients aiming to balance the supply and demand of energy, especially during periods of high demand. The expected result is that the aggregated effect of compliant clients will add flexibility capacity to the power grid to avoid overloading the generation or distribution energy systems. One of the main challenges in a Demand Response event is to efficiently define the set of clients that are likely to react positively to a signal from the supplier requiring them to shift their consumption in a specific time range. With this goal, this project proposes a data driven approach for measuring the potential for flexibility of individual customers. A metric is calculated using historical consumption data and is based on the consistency of the load pattern of each client. This flexibility score is associated with a probability of positive response to a shift consumption request. Finally, given a requested consumption reduction in a time range of a specific date, the flexibility score and a forecasting model are used to predict the customers hourly consumption and therefore, the consequent possible reduction. This process is used to simulate the selection of a set of customers for a Demand Response event using real data from about one thousand customers from Estabanell y Pahisa, S.A. in Catalonia. Table of Contents Abstract i Table of Contents ii List of Figures v List of Tables vi List of Abbreviations vii 1 Introduction 1 2 Literature Review 3 2.1 Flexibility Identification in households . . . . . . . . . . . . . . . . . . . . 3 2.2 Time Series Forecasting . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 3 Methodology 8 3.1 Datadescription ................................ 8 3.2 FlexibilityScore................................. 9 3.2.1 Frequency of Operation . . . . . . . . . . . . . . . . . . . . . . . . 9 3.2.2 Peak Time Operation . . . . . . . . . . . . . . . . . . . . . . . . . . 10 3.2.3 Approximate Entropy . . . . . . . . . . . . . . . . . . . . . . . . . 10 3.2.4 FinalScore:............................... 11 3.3 Forecasting ................................... 11 3.3.1 Algorithm Description . . . . . . . . . . . . . . . . . . . . . . . . . 11 3.3.2 Hyperparameters ............................ 15 3.3.3 Resampling Protocol and Hyperparameters Tunning . . . . . . . . . 17 3.4 Client Selection Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 4 Analysis of Results 21 4.1 Exploratory Data Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 4.1.1 MissingData .............................. 21 4.1.2 Correlation Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . 23 4.1.3 Time Series Decomposition Analysis . . . . . . . . . . . . . . . . . 24 4.1.4 Features Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . 26 4.1.5 External Data - Meteoblue . . . . . . . . . . . . . . . . . . . . . . . 27 4.1.6 DataVisualization ........................... 28 4.2 FlexibilityScore................................. 32 4.3 Forecasting ................................... 35 4.3.1 Inclusion of Meteorological Data in the model . . . . . . . . . . . . 46 ii Table of Contents iii 4.3.2 Comparison with different methods . . . . . . . . . . . . . . . . . . 47 4.4 ClientSelection................................. 48 5 Conclusions and Future Work 52 References 54 A Appendix 58 A.1 Temporal Fusion Transformers - Results Comparison . . . . . . . . . . . . 58 A.2 Implementation of the analysis . . . . . . . . . . . . . . . . . . . . . . . . . 59 List of Figures Fig. 1: Decision-tree generic example . . . . . . . . . . . . . . . . . . . . . . . . 12 Fig. 2: Intuition for Boosting algorithm . . . . . . . . . . . . . . . . . . . . . . 12 Fig. 3: XGBoost advantages against Gradient Boosting . . . . . . . . . . . . . . 13 Fig. 4: Sliding window for XGBoost time-series prediction . . . . . . . . . . . . 14 Fig. 5: Reconfiguration of the time-series . . . . . . . . . . . . . . . . . . . . . . 15 Fig. 6: Training - Testing split . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 Fig. 7: Data extraction and missing data handling . . . . . . . . . . . . . . . . . 22 Fig. 8: Density distribution of missing data . . . . . . . . . . . . . . . . . . . . 22 Fig. 9: Autocorrelation for total power consumption - up to 7 days . . . . . . . 23 Fig. 10: Autocorrelation for total power consumption - up to 60 days . . . . . . . 24 Fig. 11: Time series decomposition - additive . . . . . . . . . . . . . . . . . . . . 25 Fig. 12: Time series decomposition - multiplicative . . . . . . . . . . . . . . . . . 26 Fig. 13: Feature Engineering - Mean consumption by hour and month . . . . . . 27 Fig. 14: Feature Engineering - Rolling average of 24 hours . . . . . . . . . . . . . 27 Fig. 15: Temperature and Energy consumption Pearson’s correlation . . . . . . . 28 Fig. 16: Average consumption per hour - grouped by day of the week . . . . . . . 29 Fig. 17: Average consumption per hour - grouped by day of the week . . . . . . . 29 Fig. 18: Average consumption per hour - grouped by weekend or not . . . . . . . 29 Fig. 19: Average consumption per hour - grouped by month . . . . . . . . . . . . 30 Fig. 20: Average consumption per hour - grouped by month . . . . . . . . . . . . 30 Fig. 21: Average consumption per day - grouped by month . . . . . . . . . . . . 31 Fig. 22: Average temperature per month . . . . . . . . . . . . . . . . . . . . . . 31 Fig. 23: Peak Time Operation - Histogram . . . . . . . . . . . . . . . . . . . . . 32 Fig. 24: Frequency of Operation - Histogram . . . . . . . . . . . . . . . . . . . . 33 Fig. 25: Approximate entropy - Histogram . . . . . . . . . . . . . . . . . . . . . . 33 Fig. 26: Flexibility Score - Histogram . . . . . . . . . . . . . . . . . . . . . . . . 34 Fig. 27: Correlation between metrics . . . . . . . . . . . . . . . . . . . . . . . . . 35 Fig. 28: Metrics for XGBoost - Window = 1 up to 10 day — Clients = All — Months=All................................. 36 Fig. 29: Metrics for XGBoost - Window = 1 up to 10 day — Clients = All — Months=All................................. 37 Fig. 30: Metrics for XGBoost - Window = 1 up to 10 day — Clients = All — Months=All................................. 38 Fig. 31: Feature Importance - Covariates . . . . . . . . . . . . . . . . . . . . . . 40 Fig. 32: Predicted versus Real Consumption . . . . . . . . . . . . . . . . . . . . . 41 Fig. 33: Predicted versus Real Consumption . . . . . . . . . . . . . . . . . . . . . 42 Fig. 34: Predicted versus Real Consumption . . . . . . . . . . . . . . . . . . . . . 43 iv List of Figures v Fig. 35: Predicted versus Real Consumption . . . . . . . . . . . . . . . . . . . . . 44 Fig. 36: Predicted versus Real Consumption . . . . . . . . . . . . . . . . . . . . . 45 Fig. 37: Feature Importance Meteorological - Covariates . . . . . . . . . . . . . . 46 Fig. 38: MAPE comparison between XGBoost and TFT . . . . . . . . . . . . . . 47 Fig. 39: Total energy consumption of target day . . . . . . . . . . . . . . . . . . 48 Fig. 40: DR effect in the total consumption of the selected group of clients . . . . 49 Fig. 41: DR effect in the total consumption of the whole set of clients . . . . . . 50 Fig. 42: Flexibility Score effect in the number of clients selected . . . . . . . . . . 51 Fig. A.1: MAPE comparison between XGBoost and TFT . . . . . . . . . . . . . . 59 List of Tables vi List of Tables Tab. 1: Extract of Client Power Consumption Data . . . . . . . . . . . . . . . . 8 Tab. 2: Metrics for XGBoost - Window = 1 day — Clients = All — Months = All 36 Tab. 3: Metrics for XGBoost - Window = 1 up to 10 day — Clients = All — Months=All................................. 37 Tab. 4: Metrics for XGBoost - Window = 1 — Clients = All — Months = 3 up to7 ...................................... 38 Tab. 5: Metrics for XGBoost - Window = 1 — Clients = 50 up to 754 — Months =All ..................................... 39 Tab. 6: Metrics for XGBoost - Window = 6 - Clients = All - Months = All - Meteorological ................................ 46 Tab. 7: MAPE Comparison - from literature review . . . . . . . . . . . . . . . . 48 List of Abbreviations vii List of Abbreviations AppEn Approximate Entropy ARIMA Auto Regressive Moving Average BPNN back propagation neural networks CNN Convolutional Neural Network DAQFF Deep Air Quality Forecasting Framework DARNN Dual-Stage Attention-Based RNN DeepAR auto-regressive probabilistic Recurrent Neural Network DeepGlo Deep Global Local Forecaster DeepState Deep State Space Model DNF Disjunctive Normal Formulas DR Demand Response FO Frequency of Operation FS Flexibility Score GBRT Gradient Boosting Regression Tree GRASP Greedy randomized adaptive search procedure kNN k Nearest Neighbors LSTM Long Short-Term Memory MAPE Mean Absolute Percentage Error MIP mixed integer programming MLP Multi-Layer Perceptron NLP Neuro Linguistic Programming NODE Neural Oblivious Decision Ensembles PO Peak Time Operation RMSE Root Mean Squared Error SVM Support Vector Machine SVR Super Vector Regressor SWT stationary wavelet transform TFT Temporal Fusion Transformer TRMF Temporal Regularized Matrix Factorization UK United Kingdom XGBoost Extreme Gradient Boosting 1 Introduction 1 1 Introduction Energy distribution and generation systems are constantly monitored to guarantee balance between supply and demand. This task involves complex analysis and investments from governmental and private institutions. Generation and distribution must be adequate to guarantee the energy supply during the consumption peaks while not wasting energy in low demand hours. Several factors collaborate to make this problem even more complex. One example is the constant change in demand due to technological advancements and consequent electrification of houses and industrial environments, for instance, can the current power system adapt to the increased electricity demand from the shift to electric vehicles? [19] Another important variable is the inclusion of renewable sources of energy that have their generation related to weather conditions making the supply less certain, [3]. In this context, flexible energy systems are those where the supply and the demand can be modified to a certain extent in order to maintain an efficient equilibrium in the power system. The work developed in this thesis is focused on the demand side of the balance, more specifically in the household level of consumption. In 2021, the residential sector corresponded to 26% of the total amount of energy consumed in the EU [3], a significant share of the consumption with a potential for flexibility in the demand. Household flexibility can be achieved through demand response, where the flexible behavior of the clients is activated by a signal from the electricity suppliers requiring them to shift their consumption from a specific time in exchange for rewards, such as a discount, gift cards and so on. The success of a demand response event relies on the efficient selection of customers to be signaled. Signaling the whole customer portfolio of a company may lead to an over reduction of demand in a moment followed by a rebound with an increased demand, only shifting the peak of consumption. On the other hand, signaling fewer consumers may lead to not achieving the flexibility goals, and still need to deal with the unbalance in the grid. To achieve the goal, the proper number of clients must be selected and signaled to participate in the demand response event. To select these customers, several variables must be considered, such as consumer hourly profile, willing to participate and so on. Therefore, the first part of the present work is centered on the task of ranking the potential for flexibility of the customers, the idea is to create a method to score each customer in terms of their impact in the demand response event. The second part of the project is focused on the development of a model to forecast the energy consumption of the selected group of clients, aiming to evaluate if the consumption reduction predicted to the demand response event time will be achieved. 3 Methodology 8 3 Methodology The methodology section is divided in four subsections: data description, flexibility score, forecasting method and finally client selection process. 3.1 Data description The data source used in this project is provided by Estabanell y Pahisa S.A., a local energy provider from Catalonia. The main data source is the energy consumption information from 866 residential clients, consisting of a time series with 1 hour time step for the power consumption in Watts of each client. The data is provided for a range of 7 months from August of 2021 to February of 2022. Information that could lead to the identification of the clients was properly anonymized. Table 1 presents an extract of the data source, where each column is identified by an unique integer number, to anonymously distinguish clients, and the values on each row are the consumption in Watts of the hour for each client. For instance, in Table 1, at 4 am on July 23, 2021, client ’10’ consumed 21 Watts, while client ’98’ consumed 1082 Watts. Time 0 1 10 ... 97 98 99 23/07/2021 1:00 38.0 4.0 35.0 ... 82.0 649.0 295.0 23/07/2021 2:00 38.0 5.0 30.0 ... 105.0 1715.0 281.0 23/07/2021 3:00 44.0 4.0 23.0 ... 97.0 1668.0 253.0 23/07/2021 4:00 49.0 5.0 21.0 ... 79.0 1082.0 249.0 23/07/2021 5:00 34.0 4.0 22.0 ... 111.0 248.0 227.0 Table 1: Extract of Client Power Consumption Data Weather conditions are expected to affect the consumption patterns of household clients, once heating and air conditioning may represent a significant share of the energy demand in a house. Therefore, an additional data source with meteorological information from the region where the clients are located is used to support the forecasting task. The data is presented as a time series with the same range and time step from the clients’ database. The following meteorological information is provided: •Sea level pressure in hPa •Temperature in Celsius •Precipitation in mm •Wind direction in degree 3 Methodology 9 •Wind speed in m/s •Relative humidity in percentual The data was collected through an API from meteoblue [25] service provided by Estabanell to support the research. 3.2 Flexibility Score Aiming to select clients with a higher relevance for the demand response event, a flexibility score (FS) was defined. The score was based in the work developed in [1], considering modifications to better adequate to the current data and objectives. Three attributes have been considered in the definition of the FS: •Frequency of Operation •Peak Time Operation •Approximate Entropy 3.2.1 Frequency of Operation Frequency of Operation is defined as the frequency in which the client has a consumption above a threshold in the target hours of the DR. The idea behind this metric is to verify that the client has a minimum frequent consumption in the target hours that might represent an appliance which use could be delayed or advanced. The Frequency of Operation (FO) of a client iis calculated as follows: FOi=Pδi[d] N(3.2.1) Where: δi[d] =    1,if Pte t=tsPowerid[t]> τ. 0,otherwise. (3.2.2) 3 Methodology 10 Nis the total number of days, tsand tethe starting and end of the target interval and d is the index for the different days in the data. So, Powerid[t] is the energy consumption of client iat day din the hour t. The threshold was defined as τ= 500 Watts based on the average consumption of typical home appliances in operation. 3.2.2 Peak Time Operation This metric aims to rank users by their total consumption during the target hours of the DR event, to identify clients that together represent a significant effect in the peak hours total grid demand. The Peak Time Operation (PO) of a client iis defined as: POi= N X d=1 te X t=ts Powerid[t] (3.2.3) Again, the results are normalized across all clients to stay in the range [0,1]. 3.2.3 Approximate Entropy Approximate Entropy (AppEn) is a statistic metric used to measure the complexity of a time-series [9], indicating how predictable are the fluctuations on it. A time-series with many repetitive patterns will have a small AppEn when compared with a complex one. The pseudocode for the algorithm is presented in algorithm 1. The parameter mdefines the length of each window run of data that is compared, while rspecifies a filtering level. Values of m= 2 and requal to 20% of the standard deviation of the consumption of the client were defined as recommended by [9]. The AppEn was calculated for each client in the database and the score was defined as: AppEnscorei= 1 −AppEni(3.2.4) This adjustment is made in order to assign a higher score to clients with a more predictable consumption, the desired characteristic of the analysis. 3 Methodology 11 Algorithm 1 Approximate Entropy algorithm Require: u={u(1), u(2), u(3), ..., u(N)}Time series with Nsteps Require: {m∈Z|N > m > 0} Require: {r∈R|r > 0} for l in {m, m + 1}do for i in {1,2, ..., N −l+ 1}do x(i) = {u(i), u(i+ 1), u(i+ 2), ..., u(i+l−2)} end for for i in {1,2, ..., N −l+ 1}do count = 0 for j in {1,2, ..., N −l+ 1}do if maxk[|x(i)−x(j)|]≤rthen count =count + 1 end if end for C(i) = count/(N−l+ 1.0) end for Φ(l) = 1 N−l+1 Plog(C) end for AppEn = Φ(m)−Φ(m+ 1) 3.2.4 Final Score: The final flexibility score (FS) of each client is calculated as the average of the 3 metrics: FSi=FOi+POi+AppEnscorei 3(3.2.5) 3.3 Forecasting 3.3.1 Algorithm Description Extensive research can be found about time series forecasting, ranging from traditional models that rely on statistical metrics until complex deep learning algorithms. In this project it was decided to apply a Gradient Boosting Regression Tree (GBRT) approach, more specifically the Extreme Gradient Boosting Regression Tree (XGBoost). This supervised learning algorithm has presented levels of accuracy similar to those obtained by Deep Neural Network algorithms for structured data, [11], with the advantages of being easily scalable, present a fast training and allows to understand feature importance. 3 Methodology 12 XGBoost is an ensemble algorithm based on decision-trees that uses gradient boosting. Decision-trees is a tree-like structure where each internal node denotes an attribute in the data while the branches represent a decision/rule, Figure 1. The leaf-nodes are the outcome. This simple method can be used for regression or classification. Figure 1: Decision-tree generic example Boosting is an ensemble technique, where different prediction models are combined to obtain a unique answer. This is achieved by training a new model in the error of the previous one and applying different weights based on performance metrics, Figure 2. Finally, the gradient boosting applies gradient descent to minimize the loss function of the predictors. Figure 2: Intuition for Boosting algorithm So, the XGBoost adopts an ensemble of decision-trees as weak learners for the boosting using gradient descent to minimize the loss function. XGBoost differs from Gradient 3 Methodology 13 Boosting by the addition of system optimizations and improvements in the algorithm, Figure 3 presents a summary of the optimizations implemented in XGBoost. The XGBoost was designed to allow parallelization of the tree building, and to execute tree pruning using a maximum depth parameter to prune the trees backward. Additionally, it allows Lasso and Ridge regularization and comes with built in cross-validation in each iteration, [26]. Details of the method are found in [4]. Figure 3: XGBoost advantages against Gradient Boosting - Figure from [26] To allow exploring the flexibility of the algorithm it was proposed to adopt the windowbased feature-engineered approach presented in [11]. Elsayed et al., [11], argue that using a na¨ıve implementation of XGBoost would not be the best option for sequential data, such as the time-series used in this project. Therefore, they propose to re-configure the time series into windowed inputs/outputs and train the algorithm within this several training windows instead of using the time-series as a unique sequence of data to predict the testing data. This way the time-series is transformed to a tabular format with each row being a sliding window of fixed size, Figure 4. 3 Methodology 14 Figure 4: Sliding window for XGBoost time-series prediction - figure based on [22] Considering a window size of w∈Zfor the input and h∈Zfor the output/target, Figure 5 describes the reconfiguration of the time-series into the tabular format for a single training set, starting at point ifrom the time-series. This reconfiguration is applied for i={0,1·h, 2·h, 3·h, . . . , (n·h)−w−h}, generating ntraining sets, where n·h is the total number of points in the time-series. In the current project the window size for the output is one day, so h= 24, since the time-step is 1 hour. The window size for the input, w, is a parameter to be tuned in the analysis. This method allows to consider covariates appending the covariate-vector in the input window. This input vector is passed to the algorithm using a multi-output XGBoost regressor to predict the target, with the mean squared error as the loss function. The first image of Figure 5 shows the structure of a regular time series, the region marked in blue is flattened to became the different attributes of one observation in the table, while the region marked in orange is flattened to become the target variables of the observation, generating the line of the table presented in the lower part of Figure 5. This process is repeated sliding the window by 24 hours at a time. 3 Methodology 15 Figure 5: Reconfiguration of a single window of the time-series - figure based on [11] 3.3.2 Hyperparameters A key aspect to the performance of the algorithm is the hyperparameters. According to the XGBoost documentation, [6], there are three types of parameters in XGBoost: •General •Booster •Learning task General parameters are responsible to define the global operation of the algorithm. The main parameters of this type are: •booster: allows to choose between tree-based models (gbtree and dart) or linear functions (gblinear) as the booster. •nthreads: number of parallel threads to be used when running the algorithm. Booster parameters are the ones related to the chosen booster. For the current project the focus in on gbtree. The main parameters to be tuned for this booster are: 3 Methodology 16 •n estimators: the number of decision trees in the model. The value ranges from [0,∞]. •learning rate: is the size of the shrinkage used after each boosting step to avoid overfitting. This parameter shrinks the feature weights to make the process more conservative. It ranges from [0,1]. •gamma: sets the minimum loss reduction required to execute a new partition on a leaf node of the tree. The value depends on the loss function defined and can range from [0,∞]. Higher values lead to a more conservative model. •max depth: sets the maximum dept of each tree. Also used to avoid overfitting, once trees with high depth allow the model to learn very specific patterns in the sample. The value ranges from [0,∞]. •min child weight: also aiming to control overfitting this parameter sets the minimum sum of weights required in a child leaf. While executing the tree partition if a leaf node results in a sum of weights smaller than min child weight, the algorithm will stop partitioning. The value ranges from [0,∞], with larger values being more conservative. •subsample: defines the fraction of observations to be randomly sampled for each tree. Subsampling occurs in every boosting iteration. The goal is also to avoid overfitting. It ranges from [0,1]. •colsample bytree: defines the fraction of columns to be randomly sampled when constructing each tree. This subsampling occurs once for each tree. It ranges from [0,1]. •lambda: L2 regularization term on the weights. •alpha: L1 regularization on the weights. •scale pos weight: used to control the balance between positive and negative wights. Finally, the learning task parameters are the ones that define the optimization goals that are calculated each step. The parameters of this type are: •objective: defines the loss function being minimized. In the current project this parameter is set to use the squared error function. •eval metric: defines the metric to be used in the validation data. More than one metric can be defined. For regression analysis the available metrics are RMSE and MAE. 3 Methodology 17 3.3.3 Resampling Protocol and Hyperparameters Tunning The data set is composed of 7 months of data with a time step of 1 hour, resulting in 5208 observations. First the data set is split in training and testing data to avoid overfitting, with the first 90% of the data being used for training and the last 10% for testing, Figure 6, presents the complete time series with the sum of the energy consumption of all clients at the vertical axis in Watts. In Figure 6 the blue line is the part of the data used to train the model, while the orange line is the extract of the data used to test the model. Figure 6: Training - Testing split The hyperparameter tunning is executed in the training data using a k-fold cross validation strategy with 3 folds. This strategy was adopted to avoid overfitting while maintaining the training data as large as possible. A set of options is defined for each hyperparameter presented and a Random Search strategy is adopted. The Random Search was selected instead of a Grid Search once the number of hyperparameters to tune is large, which could lead to a high training time. In Random Search not all parameters’ values are tried out, but a fixed number is defined and is randomly sampled from the specified distributions. In the current project the number of random samples was set to 10. 3.4 Client Selection Process The client selection process intends to define a set of clients that will receive a signal requiring them to reduce their consumption in a specific time interval, with the goal of reducing the total consumption by a fixed input value in kW. The selection process starts by ranking the clients by the Flexibility Score calculated in section 3.2. Then, a binomial probability distribution with a success probability of p= 0.7 is used to simulate the acceptance or not of each client in the data frame. With this information, nclients are selected and a XGBoost model is trained to predict the 4 Analysis of Results 24 Figure 10: Autocorrelation for total power consumption - up to 60 days 4.1.3 Time Series Decomposition Analysis Time series decomposition considers that a series is a composition of the following elements: •Level: the average value of the series for a rolling window defined by the seasonality analysis. •Seasonality: the repeated short-term cycle. •Noise: the residual of the previous two. First, an additive model is considered, which assumes that the three components are summed to obtain the time-series. Figure 11 presents the decomposition for the total power consumption in Watts, with the first graph in the figure presenting the actual time series, and the next three graphs preseting, respectively, the level, seasonality and noise of the decomposed time series. 4 Analysis of Results 25 Figure 11: Time series decomposition - additive The multiplicative model assumes that the 3 components are multiplied to obtain the time series. This method is usefull to understand the proportional effect of the seasonality and the noise over the trend. As shown in Figure 11, Figure 12 allow us to identify a 24-hour cycle in the seasonality, however, we can easily infer that this 24-hour cycle leads to a variation from 70% up to 130% in the level of the trend. We can also see that the noise leads to a variation from 80% up to 120%. 4 Analysis of Results 26 Figure 12: Time series decomposition - multiplicative 4.1.4 Features Engineering One of the main advantages of XGBoost is that it automatically performs feature selection. This characteristic allows us to introduce additional features when training the model and let the algorithm deal with selecting the relevant ones for the predictive task, of course being careful with the limitations of the method. The first step on this matter was to perform a feature extraction, generating the following additional features from the time stamp: •Hour of the day •Day of the week •Day of the year •Weekend or not •Week of the Year •Month 4 Analysis of Results 27 •Year Next, transformation methods were applied to generate new features from the power consumption, as follows: •Average consumption by: –hour and month –hour and day of the week –hour and week of the year •Rolling average from 2 up to 24 hours Figure 13 and Figure 14 present examples of the resulting features generated. With the first being the mean total energy consumption by hour of each month in the data and the second a rolling average with a 24 hours window. In both graphs the blue line is the total energy consumption in Watts while the orange line is the engineered feature. Figure 13: Feature Engineering - Mean consumption by hour and month Figure 14: Feature Engineering - Rolling average of 24 hours 4.1.5 External Data - Meteoblue As explained in section 3.1, an additional data source with meteorological information from the region where the clients are located is used to support the forecasting task. This data source was collected through an API from meteoblue [22] service. 4 Analysis of Results 28 A correlation study was performed to evaluate the bond between the power consumption and the temperature of the region. This method is applied considering the whole time series and applying a rolling window of 24 hours. The Pearson’s correlation for the two time series is of −0.46, indicating a moderate correlation. Figure 15 shows the Pearson correlation considering a rolling window of 24 hours, where we can notice a unstable behavior, indicating that the correlation varies a lot depending on the position on the time series. The horizontal axis in Figure 15 is a counter of the 24 hours frames in the data, for instance, point 1 in the horizontal axis is the first 24 hours frame in the data, so the vertical axis value would be the Pearson’s correlation between the temperature and the total energy consumption for the first 24 hours in the time series. Figure 15: Temperature and Energy consumption Pearson’s correlation 4.1.6 Data Visualization Data visualization beyond the ones already presented in this section are provided to support a better understanding of the database trends and average values. First, Figure 16 and Figure 17 present comparisons between the average hourly energy consumption grouped by day of the week, with days 0 to 6 representing Monday to Sunday. In both figures the horizontal axis is the 24 hours of the day, starting from 0 up to 23 and the vertical axis is the mean total energy consumption of each hour grouped by the day of the week. In Figure 17 the shadded area along each curve represents half of the standart deviation in the data. The results indicate an expected pattern of reducing the energy consumption during late night, when most of the people are sleeping, increasing in the morning, with a stabilization around 10 am, to start increasing again at 18 hours leading to the peak of consumption around 21 hours. 4 Analysis of Results 29 Figure 16: Average consumption per hour - grouped by day of the week Figure 17: Average consumption per hour with deviation - grouped by day of the week Additionally, week-days have a slightly different average consumption profile when compared with weekends. To confirm this trend, the hourly average consumption is also plotted grouping weekends and not weekends, Figure 18, where we can notice that the average consumption on weekends is lower from 6 am to 10 am, we can infer that this effect is related to the fact that on weekends most people do not need to wake up early to work. Figure 18: Average consumption per hour with deviation - grouped by weekend or not Next, the average hourly consumption is evaluated grouped by month and the results 4 Analysis of Results 30 are presented in Figure 19 and Figure 20. In both figures the horizontal axis is the 24 hours of the day, starting from 0 up to 23 and the vertical axis is the mean total energy consumption of each hour grouped by the month. In Figure 20 the shadded area along each curve represents half of the standart deviation in the data. These results show a tendency of increase in the consumption as a whole along the day from December on. Figure 19: Average consumption per hour - grouped by month Figure 20: Average consumption per hour with deviation - grouped by month Figure 21 presents the mean total energy consumption in a day for each month, where we can confirm that the consumption increases in November, December and January. 4 Analysis of Results 31 Figure 21: Average consumption per day - grouped by month To evaluate the relation between the total consumption in a day and the temperature of the region, Figure 22 presents the mean temperature of each month in the data set, where we can notice that the coldest months (November, December and January) are the ones with the highest energy consumption in Figure 21. Figure 22: Average temperature per month 4 Analysis of Results 32 4.2 Flexibility Score The different metrics used to define the Flexibility Score were calculated in accordance with the methodology presented in section 3.2. The peak time operation was calculated considering the time interval from 20h to 21h. Figure 23 presents an histogram of the metric. For this parameter we can notice that most of the clients present a value up to 0.2 and it is clear the existance of outliers. However, once the goal of this parameter is to rank clients, it was decided to keep the outliers, since they may be good candidates for demand response events, specially in this case where a high value indicates a client with high consumption in the peak hour. Figure 23: Peak Time Operation - Histogram The frequency of operation was calculated considering the time interval from 20h to 21h and a threshold of 500 Watts, this value was tunned to allow distinguishing clients and set a minimum value that represent a domestic appliance usage. Figure 24 presents an histogram of the metric. For this parameter there is a descending number of clients per value. 4 Analysis of Results 33 Figure 24: Frequency of Operation - Histogram The approximate entropy was calculated adopting the methodology presented in section 3.2. Figure 25 presents an histogram of the metric. We can notice that most of the clients present values between 0.2 and 0.6 for this metric, with a distribution similar to normal, although the Shapiro-Wilk test does not confirm the normality of the data. Figure 25: Approximate entropy - Histogram 4 Analysis of Results 40 Figure 31: Feature Importance - Covariates The covariates with more than 0.05 of importance are listed bellow: •Energy Consumption: 0.082 •Average consumption by hour and week of the year: 0.095 •Rolling average consumption for 2 hours: 0.08 •Rolling average consumption for 3 hours: 0.072 •Rolling average consumption for 9 hours: 0.078 •Rolling average consumption for 11 hours: 0.068 •Rolling average consumption for 17 hours: 0.059 We can notice that some of the engineered features are as important as the actual target variable for the prediction. Also, the variables extracted from the time stamp did not affect the prediction and could be removed from the training without any effect in the accuracy of the method, with the exception of the Month. 4 Analysis of Results 41 Figure 32 presents the average error per hour in the testing data, and Figures 33, 34, 35, 36 presents the comparison between the predicted and actual consumption for the different days in the testing data. The vertical bars in the plot are the percentual difference in the consumption for each hour of the day, while the blue and red curves are the actual and predicted energy consumption, respectively. It is clear that the highest errors happen in the interval from 10 to 17 hours, while in the peak hours the errors are considerably lower. Figure 32: Predicted versus Real Consumption - Mean by hour 4 Analysis of Results 42 Figure 33: Predicted versus Real Consumption - part 1 4 Analysis of Results 43 Figure 34: Predicted versus Real Consumption - part 2 4 Analysis of Results 44 Figure 35: Predicted versus Real Consumption - part 3 4 Analysis of Results 45 Figure 36: Predicted versus Real Consumption - part 4 4 Analysis of Results 46 4.3.1 Inclusion of Meteorological Data in the model Lastly, it was evaluated the effect of including the meteorological information in the data as covariates. The model was obtained using 7 months of data, all the client’s consumption and an input window of 1 day. The results are presented in Table 6 along with the results for the same model however without including the meteorological covariates. The results are quite similar with the model including metereological information presenting a slightly better accuracy, an improvement of 0.08%. For this reason it was decided to not include this information in the model, once it considerably increases the training time. Meteo Covariates Window [days] Months Clients MAPE train WAPE train MAPE test WAPE test No 1 7 875 0.2% 0.2% 5.84% 5.85% Yes 1 7 875 0.9% 0.9% 5.76% 5.74% Table 6: Metrics for XGBoost - Window = 6 - Clients = All - Months = All - Meteorological Figure 37 presents the importance of each covariate in the model considering the meteorological information, with the importance of all covariates summing to 1 and the meteorological features highlighted. We can notice that the temperature achieved more than 0.05 of importance. Figure 37: Feature Importance Meteorological - Covariates 4 Analysis of Results 47 4.3.2 Comparison with different methods Additionally, a comparison of the results obtained using the window based XGBoost and Temporal Fusion Transformers is presented in the attachments, section A.1. Figure 38 shows the MAPE obtained for each of the methods for the days in the training data, February. Figure 38: MAPE comparison between XGBoost and TFT From the results we can conclude that the methods tested present similar accuracy in the prediction task, however, the training time for the TFT is considerably higher than the window based XGBoost, which motivated to maintain the XGBoost. TFT takes in average 60 minutes to be trained while the window based XGBoost is trained in average in 5 minutes, using the same machine with the maximum parallel processing available locally. However, future studies are encouraged to evaluate the accuracy gains of an ensemble of the two methods. To complement the analysis Table 7 presents the summary of the results from the publications evaluated in the literature review for the 24 hours energy consumption forecasting task, where we can notice values ranging from 3.4% up to 10% of MAPE for different methods and data sources. Therefore, the MAPE of 5.8% obtained in the current research are considered satisfactory. 4 Analysis of Results 48 Method Data Description Institute Reference MAPE Incremental Time Series Clustering Three months of energy consumption from 84 clients in Los Angeles (USA) University of Southern California [32] from 5% up to 8% ISCOA-LSTM Two years of energy consumption data from academic building in Mumbai (India) Indian Institute of Technology [33] 5.3% XGBoost/SVR/kNN ensemble One year of energy consumption data from the Jeju Island (South-Korea) Jeju National University [15] 3.4% LTSM One year if air conditioning energy consumption data from academic building in Guangzhou University (China) Guangzhou University [36] from 8% up to 10% XGBoost One year of energy consumption from 370 houses in Portugal University of Hildesheim [11] 9.9% Table 7: MAPE Comparison - from literature review 4.4 Client Selection The Demand Response event was simulated considering target reductions of 10000 W, 15000 W, 20000 W, 25000 W and 30000 W in the peak consumption of the day 24/02/2022. It was defined a threshold of 5% in this value so the algorithm will accept values from target plus or minus 5%. The algorithm described in section 3.3 was executed. First, the total consumption of the day is predicted using the XGBoost model trained with the sum of the consumption of all clients, achieving the results presented in Figure 39, where the horizontal axis is the hour of the day from 0 up to 23 hours and the vertical axis the sum of the energy consumption of all clients. Figure 39: Total energy consumption of target day Then, the algorithm iterates over the number of clients selected for the DR event, estimating the consumption of this group using a XGBoost model trained with the sum of the consumption of the selected clients with a simulated acceptance to the request to reduce consumption in the peak hours. It is considered that these clients will shift 10% of their consumption from the peak hours to the hours right after the DR. In the current 4 Analysis of Results 49 scenario, 10% of the consumption from 20h to 21h will be shifted to 22h and 23h. Figure 40 presents the effect in the group of customers selected with a simulated acceptance to the request, where the blue curve is the predicted total energy consumption of the group if no DR request was executed while the red curve is the predicted total energy consumption of the group considering that the DR request was executed. We can notice that part of the consumption of the target hours (20 and 21 hours) is shifted to the next two hours (22 and 23 hours). Figure 40: DR effect in the total consumption of the selected group of clients References 56 [20] Bryan Lima et al. “Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting”. In: (Sept. 27, 2020). url: https://doi.org/10.48550/ arXiv.1912.09363. [21] C. Liu, S. C. H. Hoi, and P. Zhao. “Online ARIMA algorithms for time series prediction”. In: Proc. 13th AAAI Conference of Artificial Intelligence (2016). [22] Jie Liu. Time Series Forecasting 101. 2022. url: https://www.esri.com/arcgisblog/products/arcgis-pro/analytics/time-series-forecasting-101-part-2-forecast- covid-19-daily-new-confirmed-cases-with-exponential-smoothing-forecast-and- forest-based-forecast/ (visited on 05/09/2022). [23] Alicia Troncoso Lora et al. “Time-Series Prediction: Application to the Short-Term Electric Energy Demand”. In: Current Topics in Artificial Intelligence. TTIA 2003. Lecture Notes in Computer Science(), vol 3040. Springer, Berlin, Heidelberg. (2004). url: https://doi.org/10.1007/978-3-540-25945-9 57. [24] Priyanka Mary Mammen et al. “Want to Reduce Energy Consumption, Whom should we call ?” In: e-Energy 18: Proceedings of the Ninth International Conference on Future Energy Systems (June 15, 2018). url: https://doi.org/10.1145/3208903. 3208941. [25] meteoblue. meteoblue. 2006. url: https://content.meteoblue.com/en/about-us (visited on 04/28/2022). [26] Vishal Morde. XGBoost Algorithm: Long May She Reign! 2022. url: https : / / towardsdatascience.com/https-medium-com-vishalmorde-xgboost-algorithm-long- she-may-rein-edd9f99be63d (visited on 05/09/2022). [27] F.J. Nogales et al. “Forecasting next-day electricity prices by time series models”. In: IEEE Transactions on Power Systems, VOL. 17, NO. 2 (2002). url: https: //doi.org/10.1109/TPWRS.2002.1007902. [28] A.D. Papalexopoulos and T.C. Hesterberg. “A regression-based approach to shortterm system load forecasting”. In: IEEE Transactions on Power Systems, VOL. 5, NO. 4 (1990). url: https://doi.org/10.1109/59.99410. [29] Lyes Saad Saoud, Hasan Al-Marzouqi, and Ramy Hussein. “Household Energy Consumption Prediction Using the Stationary Wavelet Transform and Transformers”. In: IEEE Access, VOL. 10 (Jan. 14, 2022). url: https://doi.org/10.1109/ACCESS. 2022.3140818. References 57 [30] Rajat Sen, Hsiang-Fu Yu, and Inderjit Dhillon. “Think Globally, Act Locally: A Deep Neural Network Approach to High-Dimensional Time Series Forecasting”. In: Energies, VOL. 12, NO.10 (2019). [31] Ravid Shwartz-Ziv and Amitai Armon. “Tabular data: deep learning is not all you need”. In: Information Fusion, VOL. 81 (2022). url: https://doi.org/10.1016/j. inffus.2021.11.011. [32] Yogesh Simmhan and Muhammad Usman Noor. “Scalable prediction of energy consumption using incremental time series clustering”. In: IEEE International Conference on Big Data (). url: https://doi.org/10.1109/BigData.2013.6691774. [33] Nivethitha Somua, Gauthama Raman, and Krithi Ramamrithama. “A hybrid model for building energy consumption forecasting using long short term memory networks”. In: Applied Energy, VOL. 261 (Nov. 11, 2019). url: https://doi.org/10. 1016/j.apenergy.2019.114131. [34] Ashish Vaswani et al. “Attention Is All You Need”. In: arXiv (Dec. 6, 2017). url: https://doi.org/10.48550/arXiv.1706.03762. [35] Qingsong Wen et al. “Transformers in Time Series: A Survey”. In: arXiv (Mar. 7, 2022). url: https://doi.org/10.48550/arXiv.2202.07125. [36] Chonggang Zhou et al. “Using long short-term memory networks to predict energy consumption of air-conditioning systems”. In: Sustainable Cities and Society, VOL. 55 (2020). url: https://doi.org/10.1016/j.scs.2019.102000. A Appendix 58 A Appendix A.1 Temporal Fusion Transformers - Results Comparison One of the state-of-the-art methods in time series forecasting is the Temporal Fusion Transformers (TFT). This method applies a neural network architecture integrating LSTM with Transformers attention heads. This algorithm was introduced by Lim et al [20], and studies such as the one presented by Arik and Pfister [2] have shown the capacity of the algorithm to achieve great results when compared with alternative methods. Additionally, the TFT trains the model using a quantile loss function, therefore, allowing to generate a probabilistic forecast. For these reasons it was decided to include in the project a comparison between the TFT and the window based XGBoost presented in section 3.3 for the prediction task. The same train/test split adopted for the XGBoost was applied, and the results compared in terms of the MAPE for each day in the training data (February). For the TFT an hyperparameter tunning was also executed in the training data using the k fold cross validations with 3 folds, exactly the same approach adopted in the XGBoost. The following hyperparameters were considered in the tunning process: •input chunk length = [24,48,72] - Size of the input layer •hidden size = [32,64] - Size of the Hidden layer •lstm layers = [1,2,4] - Number of layers in the LSTM •num attention heads = [2,4] - Number of attention heads •dropout = [0.1,0.2] - Dropout rate of the nodes •batch size = [32,64] - Number of observations to train before updating the weights Achieving the following results: •input chunk length = 24 •hidden size = 32 A Appendix 59 •lstm layers = 1 •num attention heads = 4 •dropout = 0.1 •batch size = 32 The training and tunning processes were executed with 300 epochs. Figure A.1 presents the comparison of the MAPE for each day in the training data. Figure A.1: MAPE comparison between XGBoost and TFT The TFT achieved a mean MAPE of 7.60% in the testing data, while the window based XGBoost achieved a mean MAPE of 5.84% for the same data. However, it is expected that with a more extensive hyperparameter tunning, longer training and inclusion of covariates the TFT can achieve even better results, which was outside the scope of the current project. A.2 Implementation of the analysis All the analysis was implemented in Python using Jupyter Notebooks, and the different project files generated are listed bellow: •01 loading data.ipynb: Read the data from Estabanell Data Lake, prepares the data to analysis and store it locally. A Appendix 60 •02 EDA and FS.ipynb: Executes the Exploratory Data Analysis and calculates the Flexibility Score of all clients. •03 XGBoost.ipynb: Trains the models and execute hyperparameter tunning for the different configurations. •04 Flex analysis opt.ipynb: Simulation of the Demand Response events. •05 Tranformer.ipynb: Trains the models using Temporal Fusion Transformer and execute hyperparameter tunning for the different configurations. •06 Methods comparison.ipynb: Compares the accuracy of the XGBoost with the Temporal Fusion Transformer •requirement.txt: List of requirements/dependencies of Python libraries used. •toolkit timeseries.py: Functions implemented to deal with the time-series (outliers, feature engineering and etc.). All the Jupyter Notebooks listed are presented in the attachments, with the exception of 01 loading data.ipynb, once this code is used to anonymize the data, therefore, containing sensitive information that could be used to identify clients. The source code is also provided as Python files in the attachments for each Jupyter Notebook.