Load Forecasting on Special Days
Full text
FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO Load Forecasting on Special Days Daniel Carlos Silva Costa e Sá Mestrado Integrado em Engenharia Electrotécnica e de Computadores Supervisor: Prof. Dr. Vladimiro Miranda Second Supervisor: Dr. Jean Sumaili Porto, July 2013
© Daniel Carlos Silva Costa e Sá, 2013
Resumo Hoje em dia, com a desregulamentação do sistema de energia, a necessidade de maior eficiência e estabelecimento de novos padrões de preservação do meio ambiente foram introduzidas restrições mais duras sobre o planeamento, gestão e controlo do sistema de energia. O sucesso comercial das empresas de energia depende da capacidade de apresentar propostas competitivas, assim sendo, alcançar melhorias na previsão de carga pode levar a um aumento substancial dos lucros comerciais. Deste modo, existem muitos métodos de previsão que têm sido publicados na literatura científica, cada um deles com especificações diferentes, dependendo dos seus objetivos. Nesta dissertação será levado em conta a repercussão com dias especiais, como feriados. A média dos erros de previsão de carga para os feriados é muito mais elevada em comparação com os dias normais, devido ao facto de nestas situações não existir quantidade suficiente de informação histórica para representar as suas características. Várias técnicas de previsão têm sido aplicada a este tipo de previsão, a maioria das abordagens baseiam-se em técnicas de redes neuronais. Muitos investigadores têm apresentado bons resultados e melhorias visíveis em novas metodologias em comparação com métodos tradicionais, mas nenhum deles tem sido capaz de resolver o problema da falta de informação histórica em dias especiais. Portanto, esta tese tem como objetivo principal a resolução deste problema. Nesta dissertação será feito o estudo de uma nova abordagem técnica para este problema, com base em Redes Neuronais Autoassociativas / Autoencoders como um estimador de dados em falta, em que serão considerados os dias especiais como os dados em falta. O algoritmo Information Theoretical Learning Mean Shift é utilizado para um processo denotado truque de densificação, ou seja, preencher com dados virtuais um conjunto escasso de dados relacionados com o consumo de energia diária em dias especiais. Isto permite um treino adequado das redes neuronais com os dados virtuais, reservando-se todos os dados reais (escassos) para fins de validação. Este método foi aplicado num problema de previsão de demanda com dados reais de uma concessionária de distribuição no Brasil, onde a previsão para dias especiais foi difícil devido à falta de dados em registos históricos. Palavras-chave: Mean shift, Information Theoretic Learning, Rede Neuronal Autoassociativa, Autoencoder, previsão de carga, dias especiais, feriado. i
ii
Abstract Nowadays with the deregulation of the power system, requirement of higher efficiency and establishment of new standards on environmental preservation, were introduced harder constraints on the planning, management and control of the power system. Commercial success of the energy companies depends on the ability to submit competitive bids, and improvements in forecasting the load can lead to substantial increases in trading profits. Therefore, exist many forecasting methods that have been published in scientific literature, each of them with different specifications, depending on its objectives. In this dissertation the repercussion of some special days will be taken in consideration, such as holidays. Average load forecasting errors on holidays is much higher than those for normal days because in this situation there is not enough historical information to represent their characteristics. Several forecasting techniques have been applied to this kind of forecasting, the majority of the approaches are based on neural network techniques. Many researchers have presented good results and visible improvements on new methodologies compared with the traditional methods but none of them has been able to solve the problem of the lack of historic information on special days. Therefore, this thesis have as its main purpose the resolution of this problem. In this dissertation the studying of a new technical approach to this problem will be made, based in Autoassociative Neural Network (AANN) / Autoencoder as a missing data estimator, in which the special days will be considered as the input missing data. The Information Theoretical Learning Mean Shift algorithm is used to a process nominated densification trick, i.e., populate, with virtual data, a scarce set related to daily energy consumption in special days. This allows the proper training of neural networks with the virtual data, reserving all the scarce real data for validation purposes. This method was applied in a demand forecasting problem with real data of a Brazilian distribution utility, where the prediction for special days was difficult to be achieved due to the lack of data in historical records. Keywords: Mean shift, Information Theoretic Learning, Autoassociative Neural Networks, Autoencoder, load forecasting, special days, holiday. iii
iv
Agradecimentos Agradeço ao meu orientador Prof. Vladimiro Miranda por todas as ideias inspiradoras, orientação e confiança, todos os conselhos e sugestões. Foi um privilégio trabalhar com ele no INESC Porto. Quero agradecer toda a disponibilidade e prontidão dos colaboradores INESC Porto, ao Dr. Jean Sumaili que para mim foi desde logo um orientador e à Joana Hora e Vera Palma pela sua ajuda na compreensão de algumas das ferramentas necessárias para o desenvolvimento da tese. Aos meus colegas e amigos que me apoiaram ao longo da minha formação académica, em especial ao Rodrigo e ao Rui. Os meus agradecimentos à minha família, especialmente aos meus pais, António e Rita e aos meus irmãos, João e Francisco, por todo o apoio, incentivo e confiança em mim depositados ao longo desta jornada. Aos meus amigos escuteiros do agrupamento 94, que complementaram a minha educação e transmitiram muitos dos valores que levo para a vida. Uma palavra amável e gentil é dirigida à Rita, minha namorada e melhor amiga, por todo o amor e carinho, juntos, somos “aprendizes de viajante”. Agradeço a Deus pela vida, pela saúde e pelas bênçãos recebidas. Daniel Sá v
xii LIST OF FIGURES A.1 Monday holidays, cluster with 11 patterns. . . . . . . . . . . . . . . . . . . . . . 61 A.2 Tuesday holidays, cluster with 19 patterns. . . . . . . . . . . . . . . . . . . . . . 61 A.3 Wednesday holidays, cluster with 14 patterns. . . . . . . . . . . . . . . . . . . . 62 A.4 Thursday holidays, cluster with 22 patterns. . . . . . . . . . . . . . . . . . . . . 62 A.5 Friday holidays (1), cluster with 10 patterns. . . . . . . . . . . . . . . . . . . . . 62 A.6 Friday holidays (2), cluster with 6 patterns. . . . . . . . . . . . . . . . . . . . . 63 A.7 Monday holidays, cluster with 11 patterns. . . . . . . . . . . . . . . . . . . . . . 63 A.8 Tuesday holidays, cluster with 22 patterns. . . . . . . . . . . . . . . . . . . . . . 63 A.9 Wednesday holidays, cluster with 14 patterns. . . . . . . . . . . . . . . . . . . . 64 A.10 Thursday holidays, cluster with 22 patterns. . . . . . . . . . . . . . . . . . . . . 64 A.11 Friday holidays, cluster with 19 patterns. . . . . . . . . . . . . . . . . . . . . . . 64 A.12 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Monday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 A.13 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Tuesday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 A.14 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Wednesday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . 66 A.15 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Thursday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 A.16 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Friday holidays (1). . . . . . . . . . . . . . . . . . . . . . . . . . . 67 A.17 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Friday holidays (2). . . . . . . . . . . . . . . . . . . . . . . . . . . 67 A.18 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Monday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 A.19 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Tuesday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 A.20 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Wednesday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . 69 A.21 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Thursday holidays. . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 A.22 Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Friday holidays (1). . . . . . . . . . . . . . . . . . . . . . . . . . . 70
List of Tables 4.1 National holidays in Brazil. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 4.2 Database complete description. . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 4.3 Database complete description. . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 5.1 Network training functions parameters. . . . . . . . . . . . . . . . . . . . . . . 45 5.2 Forecasting models, complete description. . . . . . . . . . . . . . . . . . . . . . 49 5.3 Parameters initialization. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 5.4 Performance metrics and their calculations. . . . . . . . . . . . . . . . . . . . . 53 5.5 Prediction results summary of all forecasting models (Thursday). . . . . . . . . . 53 5.6 Prediction results summary of the Model 18. . . . . . . . . . . . . . . . . . . . . 55 5.7 Prediction results summary. Comparison between two different approachs of densification of data (A1 and A2). . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 5.8 Characteristics of two different forecasting methods. Proposed in this work and proposed in the paper [4]............................... 57 5.9 Prediction results summary. Comparison with the results achieved in [4]..... 58 B.1 Forecasting models, complete description. . . . . . . . . . . . . . . . . . . . . . 71 B.2 Prediction results summary of forecasting models with 7 days preceding the holiday(Monday)..................................... 72 B.3 Prediction results summary of forecasting models with 7 days preceding the holiday(Tuesday). ................................... 72 B.4 Prediction results summary of forecasting models with 7 days preceding the holiday(Wednesday)................................... 72 B.5 Prediction results summary of forecasting models with 7 days preceding the holiday(Thursday).................................... 72 B.6 Prediction results summary of forecasting models with 7 days preceding the holiday(Friday(1)).................................... 73 B.7 Prediction results summary of forecasting models with 7 days preceding the holiday(Friday(2)).................................... 73 xiii
xiv LIST OF TABLES
Abbreviations and Symbols List of abbreviations: AANN Auto Associative Neural Network ANN Artificial Neural Network ARIMA Autoregressive Integrated Moving Average BNN Bayesian Neural Network EA Evolutionary Algorithm EC Evolutionary Computation EP Evolutionary Programming EPSO Evolutionary Particle Swarm Optimization ES Evolutionary Strategy FEUP Faculdade de Engenharia da Universidade do Porto GA Genetic Algorithms GBMS Gaussian Blurring Mean Shift GDP Gross Domestic Product GMS Gaussian Mean Shift INESC Instituto de Engenharia de Sistemas e Computadores ITL Information Theoretic Learning ITLMS Information Theoretic Learning Mean Shift LMA Levenberg-Marquardt Algorithm LP Linear Programming MAE Mean Absolute Error MAPE Mean Absolute Percentage Error MCC Maximum Correntropy Criterion MEE Minimum Entropy Error MEEF Minimum Error Entropy with Fiducial Points MLP Multi Layer Perceptron MSE Mean Square Error NMAE Normalised Mean Absolute Error NMSE Normalised Mean Square Error NN Neural Network PCA Principal Component Analysis pdf probability density function PSO Particle Swarm Optimization purelin linear transfer function xv
xvi ABREVIATURAS E SÍMBOLOS RWN Recurrent Wavelet Network SCADA Supervisory Control and Data Acquisition SSE Sum Square Error std standard deviation STLF Short Term Load Forecasting tansig hyperbolic tangent sigmoid transfer function trainbr bayesian regulation backpropagation trainlm Levenberg-Marquardt backpropagation List of symbols: σParzen Window Size λLagrange multiplier ϕ(·)Activation Function GGaussian kernel HRenyi’s quadratic entropy V information potencial DCS Cauchy-Schwartz Distance min minimum
Chapter 1 Introduction This master thesis was developed in INESC Porto, integrated in the Master Degree in Electrical and Computer Engineering at the Faculty of Engineering of the University of Porto (FEUP). A new concept of load forecasting on special days is presented in this work, using the Information Theoretical Learning Mean Shift (ITLMS) algorithm in a process of densification of the real data set, resulting in the criation of virtual data to train an Autoassociative Neural Network (AANN) / Autoencoder. The main concern is to resolve the problem of not enough amount of historical information to represent special days, such as holidays. This approach is based on Autoencoders as a missing data estimator, in which will be considered the special days as the input missing data. In order to predict the holiday demand, it is used the Evolutionary Particle Swarm Optimization (EPSO) as an optimization algorithm. 1.1 Background and Context Load forecasting became an essential instrument in power system planning, management and operation. The reasons for its growing importance are related to the deregulation of the power system. The energy market, demands higher efficiency and establishment of new standards on environmental preservation. This introduced harder constraints on power system management and control. These changes require more sophisticate tools of planning and operation and, therefore, more accurate predictions of load is necessary [7]. Nowadays, basic operating functions such as unit commitment, economic dispatch, hydrothermal coordination, transaction evaluation, fuel scheduling, unit maintenance, transaction evaluation and system security analysis can be performed efficiently with an accurate and robust forecast, improving the security of the power system and reducing the generation and operation costs [8]. Commercial success depends on the ability to submit competitive bids, and improvements in 1
2Introduction forecasting the load can lead to substantial increases in trading profits. The forecasting tools are interesting not only to power system operators, but also to load serving entities, merchant plants or generators and other market participants [9]. The quest for top-quality forecasting involves a broad variety of investigation fields, including several areas of engineering, economy, meteorology, and others. This explains the considerable forecasting methods that have been published in scientific literature. The practical details of each particular load forecasting implementation differ from case to case, depending on the objectives. In this dissertation the repercussion with special days, such as holidays will be taken into account. Average load forecasting errors for the holidays are much higher than those for normal days. In fact, their rate of occurrence may be considerably higher when dealing with real data. Besides, these kinds of events may change the general forecasting operations, channeling the performance to unacceptable levels [7,10]. Special days also make the load forecast more difficult to treat because in these situations there is not enough amount of historical information to represents their characteristics. Various forecasting techniques have been applied to this type of forecasting and the majority of the recently reported approaches are based on neural network techniques. Many researchers have presented good results. The attraction for these methods lies in the assumption that neural networks are able to learn properties of the load, which would otherwise require careful analysis to discover [11]. In short, the motivation in this dissertation thesis lies in fact that special days have been a recurring problem in load forecasting and this new approach based on Autoencoders and ITLMS algorithm may come to represent a successful tool to resolve the problem which none of forecasting technique has been able to resolve, the scarce existence of historical information on these day’s type. 1.2 Objectives Considering the promising results given in the recent paper [4], Sumaili, Miranda et al. have applied a densification trick using Information Theoretical Learning Mean Shift algorithm to allow demand forecasting in special days with scarce data. This work will be explained in greater detail in the chapter 2. In this dissertation will be studied this tool for scarce data treatment. • Will the ITLMS algorithm be able to identify distinct clusters in consumption data using a process of clustering associating the holidays in distinct days of the week? • Will the ITLMS algorithm be usefull to allow virtual data collecting of each distinct specific clusters?
1.2 Objectives 3 • Will be needed the criation of more or less groups of virtual data in order to perform a suitable neural network training set? • Will this tool be able to resolve the problem of lack of historical data on special days? Consulting the relevant literature have been observed great results for Autoencoders used as recognition machine, with this powerful tool can be estimated missing data in a database. • If it is considered the special days as a missing data, might this tool adequately predict these kind of days? • Might Autoencoder be an usefull tool in a load forecasting? • Comparing the achieved results with Autoencoders and the results of the work [4], which of them is the best? Therefore this dissertation looking for achieve these objectives set out.
4Introduction
Chapter 2 State of the Art Several research centers and companies invested in research and development of methods / models of load forecasting, which led to a large number of forecasting systems, some of which are already under operation and marketing. The prediction systems are essentially characterized by the forecast horizon (minutes, hours, days), computational complexity and value of the forecast error. This chapter does not aim to present a detailed study of all forecasting methodologies in the literature, but rather to show that there are numerous applications developed with the aim of making load forecasting on special days. The main factors influencing the load will also be presented in more detail, how the classification of load forecasting is performed and the kind of division of days per year. 2.1 Factors which influence the load behaviour An electric network is formed by the random uniting of different consumers. Changes in the consumption of different groups makes up the future load different from the previous circumstances. It is therefore crucial to study and understand the factors which influence the load in order to present methods to minimize the difference between the actual load and the forecasted load. There are many factors which influence the load behaviour. These factors all differ in terms of time of onset, duration and effect on the electricity consumption. According to [12], these factors may be divided into two groups, namely special events and ever-present factors. Sports events and strikes are examples of special events, while weather and human behavioural patterns are examples of ever-present factors. The most significant external factor that influences the load is probably the weather [11,13, 14,15]. The factors relating to the weather that are usually taken into account are temperature, rainfall, humidity, wind speed and cloud cover, as it is logical that all these factors have an effect on the use of electricity. Other factors, such as the psychological effect of hot, sunny weather, air conditioners and television audience behavior have also been suggested. Others permanent factors 5
12 State of the Art by Tanaka et al. [64]. The fuzzy regression approach showed usefulness to problems of load forecasting and load estimation in power distribution systems [65,66]. In the paper [10], a new fuzzy linear regression method for the short-term load forecasting of the holidays was proposed. An improved Tanaka’s fuzzy regression model [63], and the fuzzy regression approach [65] by introducing fuzzy input-output data using shape-preserving fuzzy arithmetic operations. Coefficients and both input and output data are considered as fuzzy numbers. The new fuzzy regression model improves the prediction accuracy for the short-term load forecasting of the holidays falling on any type of day. The maximum average percentage error obtained was 3.57% in the short-term 24 hourly loads forecasting of the holidays for the years of 1996–1997. Other researchers like [29,67,68,69,70,71,72,73,74] use also the fuzzy approach in its works. 2.5 A densification trick using ITL Mean Shift to allow demand forecasting in special days In the recent paper [4], Sumaili, Miranda et al. proposed a new method to resolve the problem of the lack of historical data in special days. They were inspired by the results of the Information Theoretical Learning Mean Shift algorithm aplied in a process denoted densification trick successfully applied in a problem of incipient fault diagnosis in power transformers [75], where scarce data on failures existed. Thus, the ITLMS algorithm was used to populate, with virtual data, a scarce set related to daily energy consumption in special days. This allows the proper training of neuronal networks with the virtual data, reserving all the scarce real data for validation purposes. The networks are then used to predict consumption in special days. An example with real data from a Brazilian distribution utility was used in order to illustrate this technique. The remarkable accuracy achieved in forecasting for holidays confirmed the correctness of this new approach. With the division of the data set by the five work days of the week (from Monday to Friday) was obtained the following forecasting indicators: the NMAE varied from 1.85% to 3.92%, while the variation range of the std was from 1.66% to 2.50%. It is important to note that this results was obtained using a simple neural network for each cluster cluster of special days. More sophisticated arrangements of neural networks are likely to allow further improvement with a narrower accuracy.
2.6 Conclusion 13 2.6 Conclusion Through the analysis of recent research in the area of load forecasting on special days, it was verified that good results have been achieved and visible improvements in new methodologies was achieved compared with the traditional methods. It has been proven that the methods based on ANN are good approximators with the capability of modeling every nonlinear system. However, conventional ANN based short-term load forecasting techniques have limitations in their use on special days. Therefore, some researchers suggest new techniques with integrated models that reduce the uncertainty and the nonlinear property of the load, in this way it was possible to minimize the forecast errors. Despite the success of these methods none of them has been able to solve the problem of the lack of historical information on special days. So this thesis has as main purpose the resolution of this problem based in the recent work of Sumaili, Miranda et al. [4]. The method implemented will be described below in more detail.
14 State of the Art
Chapter 3 Tools In this chapter the used tools in this thesis work will be presented. These main tools include the Information Theoretic Learning Mean Shift (ITLMN) algorithm, Autoassociative Neural Networks (AANN) also known as Autoencoders and Metaheuristics methods, with focus in Evolutionary Computation (EC) algorithms (Evolutionary Algorithm (EA), Particle Swarm Optimization (PSO) and Evolutionary Particle Swarm Optimization (EPSO)). The research where the autoencoder structure to implement the new method to load forecasting on special days was inspired, will also be addressed. Therefore, the following chapter seeks to inform more about these tools, their advantages and how they are currently applied. 3.1 Information Theoretic Learning Mean Shift The Information Theoretic Learning Mean Shift (ITLMS) algorithm was introduced by Rao, Principe and Martins [76,1] as a means to capture the dominant structures in the data set, as embedded in its estimated probability density function (pdf) [75]. Figure 3.1: pdf estimated, from [1]. 15
16 Tools In this subchapter this algorithm, as well as its potentialities will be presented. The Mean Shift algorithm was firstly proposed by Fukunaga and Hostler in 1975 [77]. In this paper they showed that this algorithm is a steepest descent technique where the points of a new dataset are moving in each iteration towards the modes of the original dataset. Considering a dataset X0= (Xi)N i=1εRD, using the nonparametric method of parzen window technique [78] and a gaussian kernel given by G(t) = e−t 2with bandwidth σ>0, the pdf can be estimated by: p(x,σ) = 1 N N ∑ i=1 G x−xi σ 2!(3.1) The objective of this algorithm is to find the modes of the dataset where ∇p(x) = 0. With that in mind, the iterative stationary point equation is: m(x) = ∑N i=1G x−xi σ 2·xi ∑N i=1G x−xi σ 2(3.2) The difference m(x)−xis known as mean shift. In literature, this first algorithm is known as Gaussian Blurring Mean Shift (GBMS) indicating the successive blurring of the dataset towards its respective modes due the actual solution being a single point that minimizes the overall entropy of the data set. In spite of this important development, the Mean Shift idea was forgotten until 1995, when Cheng [79] introduced a slight change in the algorithm. While in Fukunaga’s algorithm the original dataset is forgotten after the first iteration, X(0)=X0, the Cheng’s algorithm keeps this dataset in memory. This initial dataset is used in every iteration to be compared with the new dataset Y. However Yis initialized the same way, Y(0)=X0. This also introduces a small change in the iterative equation: m(x) = ∑N i=1G x−x0i σ 2·x0i ∑N i=1G x−x0i σ 2(3.3) In literature, this algorithm changed is known as Gaussian Mean Shift (GMS) Algorithm. Mean Shift algorithms have been shown a very versatile and robust tool in feature space analysis [80] and is often used in image segmentation [81,82], denoising, tracking objects [83] and several other computer vision tasks [84,85]. Recently, in 2006, Rao, Principe and Martins [76,1] introduced a new formulation of mean shift known as Information Theoretic Learning Mean Shift (ITLMS) and showed that GBMS and GMS are special cases of this one.
3.1 Information Theoretic Learning Mean Shift 17 The idea in this algorithm was to create a cost function that minimizes the cross entropy of the data while the Cauchy-Schwartz distance is kept at a given value. Knowing that a gaussian kernel (with bandwidth σ>0) is given by: Gσ=e−x2 2·σ2(3.4) an estimation of a pdf, using the parzen window technique [78], is: p(X) = 1 N N ∑ i=1 Gσ(x−xi)(3.5) Renyi’s quadratic entropy [86] for a pdf can be calculated using: H(X) = −log +∞ Z −∞ p2(x)dx (3.6) Therefore, replacing 3.5 into 3.6, H(X) = −logV(X)(3.7) with V(X) = 1 N2 N ∑ i=1 N ∑ j=1 Gσ0(xi−xj)(3.8) where σ0=√2σ.V(x)is known as the information potential of the pdf p(X). The derivative of this expression with respect to a single point xigives a quantity denoted information force exerted by all data particles on xi. To measure the cross entropy between two pdf, one has H(X,X0) = −logV(X,X0)(3.9) with V(X,X0) = 1 N2 N ∑ i=1 N ∑ j=1 Gσ0(xi−x0j)(3.10) The Cauchy-Schwartz distance between two pdfs (pand q) can be calculated using: DCS(X,X0) = log(Rp2(x)dx)·(Rq2(x)dx) (Zp(x)·q(x)dx)2(3.11) DCS(X,X0) = −[H(X) + H(X0)−2H(X,X0)] (3.12)
18 Tools The ITLMS algorithm aims at finding data sets Xthat capture structural information from a set X0. This is achieved by a double criteria optimization, minimizing the entropy of Xwhile keeping the Cauchy–Schwartz distance at some value k. An unconstrained optimization formulation, under a parameter λ(Lagrange multiplier) that represents the tradeoff between the two objectives is given by J(X) = min H(X) + λ·[DCS(X,X0)−k](3.13) Differentiating J(X)with respect to each xigives an algorithmic rule that allows the transformation of X0into another set at iteration t+1, making use of the information contained in the pdf of Xat iteration t, estimated by 3.5: xt+1 i=c1·S1+c2·S2 c1·S3+c2·S4(3.14) where c1=1−λ V(X),c2=1−λ V(X,X0)(3.15) and S1= N ∑ j=1 Gσ kxt i−xtjk2 σ0!×xt j(3.16) S2= N ∑ j=1 Gσ kxt i−xt 0jk2 σ0!×x0j(3.17) S3= N ∑ j=1 Gσ kxt i−xtjk2 σ0!(3.18) S4= N ∑ j=1 Gσ kxt i−xt 0jk2 σ0!(3.19) As shown in [76], adjusting the λparameter changes the data properties sought by the algorithm: λ=0– the algorithm minimizes the data entropy, returning a single point. This is the GBMS algorithm; λ=1– the algorithm is a mode seeking method. The particles converge to the modes of the pdf p(X), the same as GMS;
3.2 Autoassociative Neural Networks 19 λ>1– the principal curve of the data is returned (1 <λ<2). A higher value of λmakes the algorithm seek to represent all the characteristics of the pdf. Each generation of points xt idescribe a pdf p(Xt)that retains information from p(X0). Each point xt ialong the iterations tdescribes a path from xi0toward a mode of the pdf p(X0), or to a principal curve of the data cluster, or to a region of higher density, depending on the value of λ adopted. By path, one means a succession of points X0,X1,...,Xt,... that may be driven toward or away from the mode, depending on allowing points to follow the direction of the information force (∂V/∂Xas in 3.8) or the reverse direction. The set XV=X1∪X2... ∪Xtis the set of virtual data generated by the ITLMS algorithm. It forms a dense cluster that shares properties with the original X0. This use of XVis called the densification trick. This property was successfully applied in a problem of incipient fault diagnosis in power transformers [75] by Miranda et al., where scarce data on failures existed. Therefore, this densification trick becomes especially useful when data is scarce or often insufficient to a neural network training practice. The insufficient number of samples is a difficulty present in many works reported. In particular, the solidity of models whose validation rests on such a low number of test samples may be questioned. With the use of the ITLMS, the training set may be composed of only virtual points, keeping the totality of the real data to be used in the testing phase. This largely increases the robustness of the testing procedure and the confidence in the results it will provide. The densification trick using ITLMS demonstrates to be a powerful tool to resolve the problem of the typical lack of historical data on special days, as demonstrated in [4] by Sumaili, Miranda et al.. Its application in the construction of a neural network system for the 1 day-ahead prediction of electric energy consumption in special days was suggested, for a Brazilian distribution utility. In chapter 4the results of the densification of data sets which will be applied as pratical example in this thesis will be presented. 3.2 Autoassociative Neural Networks For a better understanding of which is an Autoassociative Neural Networks (AANN), the next general contextualization of Artificial Neural Networks will be provided. Artificial Neuronal Networks, or just Neuronal Networks (NN) are machines designed the way human brain performs/learns a particular task or function of interest [2]. NN provide a principled framework for learning linear and non-linear mappings from an input to an output space, corresponds to a connectionist paradigm of information processing, including a massive paralel process of numerical computacions [6], through a process of learning. The basic processing element of a NN is the neuron [87]. Neurons are composed of several inputs, one output and an activation function which executes the internal processing, transforming the inputs into the output. Usually,
20 Tools neurons are organized in layers with unidirectional links always in a forward direction, from the input to output of the NN(feedforward networks). Figure 3.2: Nonlinear Model of a Neuron, from [2]. Connections between neurons are associated with synaptic weights wk j, such that a signal emitted by a neuron is multiplied by the weight of the conection before entering a next neuron [6]. This process is schematized in figure 3.2 and the equations demonstrated are: The weighted sum of inputs xj: uk= m ∑ j=1 wk j.xj(3.20) The summing junction of the bias bkto the uk: vk=uk+bk(3.21) Finally, the output ykis the result of vkthrough activation function ϕ(·). yk=ϕ(vk)(3.22)
3.2 Autoassociative Neural Networks 21 Autoassociative Neural Networks (AANN), also known as autoencoders, are feedforward neural networks with a middle hidden layer that intends to reconstruct the output equal the input. Thereby, the size of the output layer is always the same as the size of the input layer. The simplest autoencoder architecture has only one middle hidden layer, once the use of more hidden layers makes the training tedious and furthermore, it will also not give good results [88]. The optimal number of the hidden neurons, thought dependent on the type of application, must be smaller than that of the input and output. In the figure 3.3, a typical diagram of an autoencoder is shown. Figure 3.3: The structure of a eigth-input, eight-output autoencoder The autoencoders perform two main operations: a forward compressing operation transforming from data space to code space at the hidden layer called “encoding”, and reverse transformation from code space to data space at the output layer called “decoding”. If linear activation functions are used, autoencoder will be performing similar to Principal component analysis (PCA) method [89,90] that is, reduce the dimensionality of data. With nonlinear activation functions, autoencoders chart the input space on a nonlinear manifold in such a way that an approximate reconstruction is possible with less error [91]. Plus, PCA does not easily show how to do the inverse reconstruction, which is straightforward with autoencoders [92]. The goal is to train the network such that the composed operation is as close as possible to the identity mapping. By defining a network structure where inputs and outputs are tied to training samples, appropriate network parameters (weights and biases) can be trained by on criterion optimization, the classical function adopted is the minimization of the Mean Square Error (MSE) between the input and the outputs. If Xis the input vector and Ythe output vector, then for Nsamples: MSE :min ε=min 1 N N ∑ k=1kXk−Ykk2(3.23) A good interpretation of the MSE criterion is that it represents the minimization of the variance of the pdf of error distribution. However, this criterion is optimal only if this distribution is
28 Tools bgbest position found by the swarm of particles in their past life; Rnd random numbers sampled from a uniform distribution in [0,1]. The following figure illustrates this concept. Figure 3.7: Ilustrating the movement of a particle iin PSO, influenced by the three terms: Inertia, Memory and Cooperation [3]. The weights affecting the various terms are affected in each iteration by multiplying random numbers, which causes a disturbance in the trajectory of each particle which has been shown to be beneficial for space exploration and discovery of the optimal solution. The weights in this simple model are defined initially and externally. This raises a tuning problem of these weights in order to reach the convergence. In fact, this is the main disadvantage of PSO algorithm, it is not self-adaptative. In order to achieve better results we can apply mechanisms in the movement rule that, although none of them solve the problem of the lack of self-adaptivity, it can bring improvements to the PSO algorithm [3]. There exists then two principal mechanisms, one of them can be described in the following manner: Vnew i=Dec(t)·Wii·Vi+Rnd ·Wmi·(bi−Xi) + Rnd ·Wci·(bg−Xi)(3.29) I.e. apply in the inertia term a Dec(t)function whose value decreasing with the progress of iterations, reducing progressively the importance of this term [112], and also apply a new weight Wii. The other mechanism (proposed by Maurice Clerc [113]) consists in the multiplication of the movement rule by a constriction factor K. This factor consists of a diagonal matrix of constriction factors of dimension k.
3.4 Metaheuristic Methods 29 Kk=2 |2−Wk−qW2 k−4·Wk| ,Wk=WmK+W ck,Wk>4 (3.30) 3.4.2 Evolutionary Particle Swarm Optimization (EPSO) As described above the EPSO algorithm can be seen as a hybrid method of ES/EP and PSO techniques. As an ES, an EPSO algorithm may be described (as Miranda in [3]) by the following general scheme: Replication each particle is replicated ntimes; each particle has its strategic parameters mutated; Reproduction each mutated particle generates an offspring through recombination, according to the particle movement rule, described below; Evaluation each offspring has its fitness evaluated; Selection by stochastic tournament or other selection procedure, the best particles survive to form a new generation, composed of a selected descendant from every individual in the previous generation. The EPSO reproduction rule can be described by the following expression where given a particle Xi, a new particle Xnew iwill be: Xnew i=Xi+Vnew i(3.31) The movement rule of the EPSO is given by Vnew i=Wi∗ i·Vi+Wm∗ i·(bi−Xi) + Wc∗ i·(b∗ g−Xi)·P(3.32) where Wiiweight conditioning the inertia term; Wmiweight conditioning the memory term; Wciweight conditioning the cooperation term; bibest position found by the particle in its past life; bgbest position found by the swarm of particles in their past life;
30 Tools Pcommunication factor, assumes binary variables of value 1 with probability p and value 0 with probability (1-p); the p value, set as an external parameter, controls the passage of information within the swarm and is considered 1 in classical formulations. EPSO algorithms include the adoption of a stochastic star communication topology, instead of the deterministic scheme usually adopted in PSO. This has the advantage of sharing the knowledge of each particle’s knowledge of the global best position, controlled by a communication probability P, which is self-adaptive throughout the algorithm run and also externally defined. The effect produced by the adoption of a stochastic star communication topology is that a particle will ignore the global best on some iterations and include it in other iterations. This not only allows more local search by each particle, but also allows the elimination of disturbing noise, by allowing the dynamics of particle movement to be more stable and avoiding premature convergence [109]. The symbol ∗indicates that these parameters will undergo evolution under a mutation process. The difference for the particle swarm PSO is that the evolution does not only occur in the behavior of particles, but also on the weights that affect the movement of these in the search space. One of the main features is that it is a self-adaptive method, that is, automatically adjusts the swarm behavior in order to enhance efficiency in the search and on the other hand, prevent the divergence of the swarm. This characteristic lies in the fact that at a given instant, there is a particle which has the best position in the search space, and the population of particles have to move in this direction. In addition, each particle is also attracted to its previous best position. This process of the EPSO is illustrated in the following figure 3.8. Figure 3.8: Illustration of EPSO particle reproduction: a particle Xigenerates an offspring at a location commanded by the movement rule [3].
3.4 Metaheuristic Methods 31 The approximate basic mutation rule for the strategic parameters is the following: Wk∗ i=Wki·[1+τ·N(0,1) ] (3.33) where N(0,1)random variable with Gaussian distribution (0 mean and variance 1); τlearning parameter, fixed externally, controlling the amplitude of the mutations – smaller values of τlead to higher probability of having values close to 1. As for the global best bg, it is randomly disturbed to give b∗ g=bg+Wb∗ i·N(0,1)(3.34) where wbiis the strategic weight parameter associated with particle i. It controls the size of the interval of bg where it is more likely to find the real global best solution. This weight wbiis mutated (denoted by *) according to the general mutation rule. In several papers it is possible to verify the advantages of this optimization tool in electrical power applications [111,114,115,116,117,118,119,120].
32 Tools
Chapter 4 Data Treatment The historical data which were taken into account in this thesis are the same that were applied in [4] by Sumaili, Miranda et al., the real data from a Brazilian distribution utility. The historical data refer to about 10 years of consumption (from January 2002 to September 2012). In Brazil, public holidays may be legislated at the federal, statewide and municipal levels. Most holidays are observed nationwide, but each state and city may have its own holidays as well. Apart from the yearly official holidays (listed below), the Constitution of Brazil also establishes that election days are to be considered national holidays as well. General elections are held on the first Sunday of October, in the first round, and on the last Sunday of October, in the second round, of every even year. Table 4.1: National holidays in Brazil. Date Holiday name Holiday type January 1 New Year’s Day Fixed by date 47 days before Easter Carnival/Shrove Tuesday Fixed by day (Tuesday) Day after Carnival Carnival end (until 14 hrs) Fixed by day (Wednesday) Friday before Easter Good Friday Fixed by day (Friday) Computus1Easter Day Fixed by day (Sunday) April 21 Tiradentes Day Fixed by date May 1 Labour Day Fixed by date Thursday after Trinity Sunday2Corpus Christi Fixed by day (Thursday) September 7 Independence Day Fixed by date October 12 Our Lady of Aparecida Fixed by date November 2 All Souls Day Fixed by date November 15 Republic Proclamation Day Fixed by date December 25 Christmas Day Fixed by date December 31 New Year’s Eve (from 14 hrs) Fixed by date 1The Computus (Latin for “computation”) is the calculation of the date of Easter, the first Sunday after the first ecclesiastical full moon (that follows the Northern spring equinox) falling on or after 21 March 2Trinity Sunday is the first Sunday after Pentecost. Pentecost is celebrated seven weeks (50 days) after Easter Sunday, hence its name. 33
34 Data Treatment Other days can also be considered as special days, like the days preceding the Carnival, the Christmas Eve, the Valentine’s Day or even the Fridays or Mondays in extended weekends, among other. For more details the following website may be consulted: www.timeanddate.com/holidays/ brazil/ In this chapter the treatment of the historical data will be described taking into account the following criteria: • In this study of load forecasting on special days, holidays which occur at Saturdays and Sundays will not be analysed, nor consecutive holidays with frequency inferior to eight days; • The load forecasting will be based on the daily energy consumption of the days preceding the holiday • The demand of the same special days are dissimilar each year due to the system load growth/decline trend. If this yearly growth is ignored, the general shapes of same days become similar. Therefore, in this study, like in many others, the load forecasting will be performed based on historical data of holidays with the same behavior (e.g. special days that occur on Wednesdays and which have the same weekly behavior); • As mentioned earlier, to resolve the problem of the lack of historical data on special days the ITLMS algorithm will be used to make the densification of data set. 4.1 Normalization of Data Set and Its Classification Using ITLMS As mentioned above, the demand of the same special days are dissimilar each year due to the system load growth/decline trend. If this yearly growth is ignored, the general shapes of same days become similar. Therefore, the load forecasting can be performed based on historical data of holidays which occur at the same day and with the same weekly behavior. Therefore, the first step to data treatment of the holidays and their previous days is to make a correction in all of them in order to obtain their similarity. The normalization was made with respect to the consumption of the previous week. The next images illustrates this method.
4.1 Normalization of Data Set and Its Classification Using ITLMS 35 Figure 4.1: Demand correction method. (a) Set of consumption patterns. (b) Normalized special day patterns. Figure 4.2: Normalization of the data set, from [4]. It is important to refer which of the special days correspond to the last represented day. After applying this method, ITLMS algorithm was used to understand the similarity between the patterns shown in figure 4.2b. Therefore, it was possible to identify distinct patterns for special days, and cluster them in similar classes. This way, each special day / holiday was associated to a particular pattern (cluster). With setting λ=0.9 in 3.15 the identification of thirteen different modes was possible. The patterns converging to a common mode were grouped in individual clusters. It was thus possible to form six clusters corresponding to the five days of the week (from Monday to Friday, two clusters on Friday) and others seven groups with the remaining outliers that were not taken into consideration in this study. In a second approach with λ=0.1 in 3.15 ten different modes were identified. Five clusters corresponding to the five days of the week were formed, and other five groups with the remaining outliers that were also not taken into consideration in this study. The reason for not considering the remaining outliers is because some patterns correspond
36 Data Treatment to a very special cases which should deserve individual analysis. Some of these cases possibly correspond to blackouts, which severely reduced the daily consumption, or also, holidays that do not have a fixed week day distort the observed pattern or even the occurence of two holidays in the same week. The following figure shows two examples of clusters organized. In the same cluster there are holiday with occurrence on the same day and same weekly behavior. (a) Monday holidays, cluster with 11 patterns. (b) Tuesday holidays, cluster with 19 patterns. Figure 4.3: Holidays grouped by the same weekly behavior and which occurred at the same day (Approach 1). In these two graphs it is possible to observe which different special days (properly standardized) was grouped in the same cluster in order to perform the load forecasting. The representation of all clusters of patterns can be seen in appendix A.1. In table 4.2 the number of patterns obtained on database using the properly correction and the ITLMS for classification is shown. Table 4.2: Database complete description. Approach 1 (λ=0.9) Approach 2 (λ=0.1) Special Day Real Data Monday 11 Tuesday 19 Wednesday 14 Thursday 22 Friday 10 Friday 6 Special Day Real Data Monday 11 Tuesday 22 Wednesday 14 Thursday 22 Friday 19 The reason of the database considering two different groups of special days which occur on Friday in the approach 1 is because these holidays, even falling on same weekday, have a different weekly behavior and the settings given to ITLMS led to consider these days in different clusters (See figure 4.4).
4.2 Densification of Data Set 37 (a) Friday holidays (1), cluster with 10 patterns. (b) Friday holidays (2), cluster with 6 patterns. Figure 4.4: Holidays grouped by the same weekly behavior and which occurred at the same day. It is interesting to observe that the lowest value of consumption in each diagram corresponds to the consumption on Sunday, so, it is simple to identify each diagram in association to each day of the week. The phenomenon of "extended weekend" is easily detected in the Tuesday cluster where the consumption on Monday is on average smaller that on the other working days. 4.2 Densification of Data Set As stated before in section 3.1, the densification trick using ITLMS algorithm will be applied as a way of resolving the problem of lack of historical data, insufficient to an AANN training practice. It is thus intended that the training set is composed of only virtual points, keeping the totality of the real data to be used in the validation phase. This largely increases the robustness of the validation procedure and the confidence in the results it will provide. The ITLMS algorithm for each cluster was run in order to create a dense cluster of virtual data. In table 4.3 the full description of the database obtained is shown. The different values of virtual data can be justified with the number of the original real data and the number of iterations needed by the mean shift algorithm to converge to a single mode. In this study the convergence was reached after 21 iterations, generating 21 virtual patterns for each original data point in approach 1. In approach 2 the convergence was reached after 59 iterations, generating 59 virtual patterns for each original data point. I.e., the total number of virtual patterns in each cluster is the number of original real points times the number of iterations performed.
44 Load Forecasting Models The following picture illustrate the general autoencoder structure. Figure 5.3: Autoencoder structure, from MATLAB. As shown by the picture, and was presented in subchapter 3.2, the autoencoder needs two transfer functions for “encoding” and “decoding”, i.e., a compressing operation transforming from data space to code space at the hidden layer, and reverse transformation from code space to data space at the output layer. Transfer functions or activation functions, ϕat 3.22, are used for limiting the amplitude of the output of a neuron. The “encoding” transfer function was performed by a hyperbolic tangent sigmoid (tansig) and is given by a=tansig(n) = 2/1+e(−2·n)−1 (5.7) where nrepresents input data and aoutput data. Figure 5.4: Hyperbolic tangent sigmoid transfer function, from [5]. The “decoding” was performed by a linear transfer function (purelin) and is given by a=purelin(n) = n(5.8) Figure 5.5: Linear transfer function, from [5].
5.1 Training with MATLAB NN Toolbox 45 The following figure illustrate an example of the window of the results of training the autoencoder with the MATLAB NN Toolbox. In this example the cluster of Thursday holidays of the first approach of the densification trick using ITLMS algorithm and the function Levenberg-Marquardt backpropagation (trainlm) as the network training function were used. Figure 5.6: Neural network training window, from MATLAB. In this figure it is possible to observe the principal parameters which were taken into account to make the autoencoder training: Table 5.1: Network training functions parameters. Maximum number of epochs 5000 Performance goal 0 Maximum validation failures 2000 Minimum performance gradient 1×10−14 Note: These parameters were chosen after many tests and simulations, is possible achieve better results with different parameters, but this would require an exhaustive study. Besides of these parameters, to perform the diverse training models the selection of the training function (trainlm or trainbr) and the structure of the autoencoder (8-7-8, 7-6-7 or 5-4-5) were also taken into account. In the example of the figure 5.6 can be seen how the choice of all these parameters was made.
46 Load Forecasting Models Another important observation in this example is the NN training which was stopped when the maximum number of epochs were reached. The MATLAB NN Toolbox gives also more results to evaluate the performance of training models. Then the comparative results of the two training functions (trainlm and trainbr) with the autoencoder structure 8-7-8 under the same conditions will be presented. (a) trainlm (training performance MSE=5.47×10−17)(b) trainbr (training performance MSE=1.64×10−16) Figure 5.7: Performance plots, from MATLAB. (a) trainlm (b) trainbr Figure 5.8: Training state plots, from MATLAB.
5.1 Training with MATLAB NN Toolbox 47 Analysing the results is always important evaluate if an overfitting problem occurs at the process of training the NN. A model is typically trained by maximizing its performance on some set of training data. By the other hand, the model efficacy is determined not by its performance on the training data but by its ability to perform well on unseen data. When a model begins to memorize training data rather than learning to generalize from trend this is called overfitting. This corresponds a test data error much higher than the train data error and means that the neural system is over determined [6]. A properly trained system should correspond with the same order of magnitude for the error measures to both training and testing data. To avoid overfitting, this point must be identified and the training must be stoped. The following figure ilustrates the point where the optimal learning and generalization are achieved, that is close to the global minimum of test error. Figure 5.9: Overfitting, overfitting, from the training epoch t∗, from [6]. Therefore, through the results of the figure 5.7 is possible verify that the problem of overfitting does not occur. Despite the results indicate that the trainlm is the training function with minor training performance MSE, the evaluation of the best method of training autoencoders is only possible to verify further along once that only the results of the complete model can demonstrate the correct validation of the predictions.
48 Load Forecasting Models 5.2 Description of the Forecasting Models In this section the complete models which were used to perform the forecasts will be described. Taking into account the base models presented in the section 3.3 the following three models will be considered: Model 1 – Unconstrained search model - an optimization algorithm searches for the input values that minimize the input/output error on the signal that correspond to the special day; Model 2 – Unconstrained search model - an optimization algorithm searches for the input values that minimize the input/output error on all the signals except one that correspond to the special day; Model 3 – Constrained search model - an optimization algorithm searches for the input values that minimize the input/output error on all the signals. The following pictures illustrate well these three forecasting models. Figure 5.10: Model 1 – Autoencoder and optimization algorithm - Unconstrained Search model. Figure 5.11: Model 2 – Autoencoder and optimization algorithm - Unconstrained Search model.
5.2 Description of the Forecasting Models 49 Figure 5.12: Model 3 – Autoencoder and optimization algorithm - Constrained Search model. Therefore, it is concluded that, complete forecasting models can be divided as follows. Table 5.2: Forecasting models, complete description. Model Optimization search model Days preceding the holiday Training functions 14trainlm 2 trainbr 36trainlm 4 trainbr 57trainlm 6 trainbr 74trainlm 8 trainbr 96trainlm 10 trainbr 11 7trainlm 12 trainbr 13 4trainlm 14 trainbr 15 6trainlm 16 trainbr 17 7trainlm 18 trainbr
50 Load Forecasting Models 5.3 Results analysis In this section the results of the models under analysis will be presented, as well as the relating discussion of them. Before being assessed all forecasting models, the metaheuristics PSO and EPSO were evaluated in order to find which is the most robust. 5.3.1 PSO vs EPSO In order to evaluate the most robust metaheuristic method to implement in forecasting model, some tests were performed. In these tests only one forecasting model was considered and the results given by PSO and EPSO in the same conditions. The forecasting model was • Training function – trainbr; • Horizon of days preceding the holiday – 7 days; • Optimization search model – model 3. The initialization parameters of the two metaheristics are Table 5.3: Parameters initialization. PSO EPSO Parameters Value Wm0.03 Wc0.06 Iterations 200 Parameters Value Wi0.03 Wm0.03 Wc0.06 Wb0.06 τ0.01 P0.8 Iterations 200 Note: These parameters initialization were chosen after many tests and simulations, is possible achieve better results with different parameters, but this would require an exhaustive study.
5.3 Results analysis 51 Therefore, the results after ten forecasts of the first real data of Thursday holidays cluster (normalized daily energy consumption = 0.126935) are illustrated in the next figures. (a) PSO (b) EPSO Figure 5.13: Ten forecasts of the same day using PSO and EPSO as optimization algorithm. The following pictures illustrate an example of the MSE minimization which were performed by PSO and EPSO algorithms. (a) PSO (b) EPSO Figure 5.14: MSE minimization using PSO and EPSO as optimization algorithm. Given these results, it can be concluded that EPSO proved to be a metaheuristic more robust than the PSO. The algorithm robustness, which has to do with the warranty (probability) that, regardless of the initialization, the algorithm will converge to the optimal or your vicinity. It is not expected that the algorithm is executed several times on the same problem. It is expected that it gives only one trust result. It is precisely this confidence in the result which is measured by the concept of robustness, it is expected that the algorithm when, executed several times, always finds good results
52 Load Forecasting Models with very small deviations of the optimal solution. Analyzing the results in detail, it was found that the PSO, in many cases, gets stuck in local optima and that this influences the quality of results. Furthermore, EPSO has a better accuracy for the same computational effort. Therefore, the optimization method chosen is the EPSO. 5.4 Prediction results The tests of the all models were performed in four stages: 1. With the first approach of the densification of data and the application of only one cluster (e.g. cluster of Thursday holidays) all 18 forecasting models were tested. The intention with this study is verify the influence of more or less days in the forecasting model; 2. After the evaluation the results of the first stage the best models (consideration of more or less days preceding the holiday) were considered to perform the forecasting on all clusters of the first approach of data densification; 3. For the best model found in the stage 2, the influence of the consideration of two different approachs of densification of data was evaluated (Approach 1: 21 virtual data for each real data; Approach 2: 59 virtual data for each real data). 4. In the end, the comparison with prediction results given in [4] will be made. The results analisys will be evaluated by some forecasting indicators such as the variation range std (standard deviation), MSE (mean square error), MAE (mean absolute error), NMSE (normalized mean square error), NMAE (normalized mean absolute error) and MAPE (mean absolute percentage error). These performance metrics and their calculations are shown in the following table.
5.4 Prediction results 53 Table 5.4: Performance metrics and their calculations. Metrics Calculation std std (xi) = n ∑ i=1(xi−x) n−1 MSE MSE = n ∑ i=1(xi−ˆxi)2 n MAE MAE = n ∑ i=1|xi−ˆxi| n NMSE NMSE = n ∑ i=1(xi−ˆxi)2 std(ˆxi)·n NMAE NMAE = n ∑ i=1|xi−ˆxi| n ∑ i=1|xi| MAPE MAPE = n ∑ i=1|xi−ˆxi|/xi n×100% ∗xiand ˆxiare the real values and predicted values. 5.4.1 Stage 1 In this first stage, it is intended to make the evaluation of all forecasting models. After all tests the following results were reached. Table 5.5: Prediction results summary of all forecasting models (Thursday). Model std MSE MAE NMSE NMAE MAPE 1 1.058% 1.0E-04 0.871% 2.865% 6.675% 6.719% 2 0.483% 2.5E-05 0.372% 0.699% 2.850% 2.841% 3 0.944% 8.2E-05 0.783% 2.261% 6.006% 6.036% 4 0.384% 1.0E-05 0.267% 0.288% 2.047% 2.037% 5 0.138% 1.3E-05 0.298% 0.350% 2.282% 2.282% 6 0.277% 8.6E-06 0.221% 0.238% 1.712% 1.712% 7 1.929% 3.4E-04 1.433% 9.512% 10.991% 11.093% 8 0.984% 9.0E-05 0.800% 2.487% 6.134% 6.161% 9 2.669% 6.7E-04 2.122% 18.532% 16.272% 16.438% 10 0.625% 3.0E-05 0.470% 0.823% 3.607% 3.592% 11 0.453% 1.1E-05 0.288% 0.308% 2.208% 2.209% 12 0.306% 1.7E-06 0.090% 0.048% 0.691% 0.681% 13 1.757% 2.9E-04 1.352% 7.880% 10.368% 10.462% 14 0.724% 5.1E-05 0.590% 1.404% 4.523% 4.535% 15 1.708% 2.7E-04 1.476% 7.490% 11.317% 11.406% 16 0.537% 2.3E-05 0.415% 0.636% 3.183% 3.168% 17 0.373% 7.3E-06 0.233% 0.202% 1.788% 1.786% 18 0.304% 1.1E-06 0.071% 0.031% 0.544% 0.540%
60 Conclusion 6.1 Future Work In spite of the advances done in this work, much work remains to be done in the area of load forecasting on special days. It would be important verify the performance of this new forecasting model, with historical data from other power distribution utilities. It would be also interesting perform the load forecasting not just on holidays but also on normal days. Besides the several tests that were made, more tests with other parameters in each tool of the model can allow the achievement of better results. The development a more efficient algorithm to the autoencoder training can bring better results. There are several ones using evolutionary algorithms in the literature, for example in [128] in order to perform the wind power forecasting, the NN training with metaheuristic EPSO was implemented. In the same work as in others [129,130], were considered others optimization criteria adopting entropy concepts to train the NN, based in mutual information principle [131]. Renyi’s Entropy is combined with a Parzen Windows estimation of the error pdf to form the basis of three criteria (MEE, MCC and MEEF) under which neural networks are trained. In some researchs, the results was favourably compared with the traditional MSE criterion.
Appendix A Results of the data treatment A.1 Results of the load correction method A.1.1 Approach 1 Figure A.1: Monday holidays, cluster with 11 patterns. Figure A.2: Tuesday holidays, cluster with 19 patterns. 61
62 Results of the data treatment Figure A.3: Wednesday holidays, cluster with 14 patterns. Figure A.4: Thursday holidays, cluster with 22 patterns. Figure A.5: Friday holidays (1), cluster with 10 patterns.
A.1 Results of the load correction method 63 Figure A.6: Friday holidays (2), cluster with 6 patterns. A.1.2 Approach 2 Figure A.7: Monday holidays, cluster with 11 patterns. Figure A.8: Tuesday holidays, cluster with 22 patterns.
64 Results of the data treatment Figure A.9: Wednesday holidays, cluster with 14 patterns. Figure A.10: Thursday holidays, cluster with 22 patterns. Figure A.11: Friday holidays, cluster with 19 patterns.
A.2 Results of the densification of data sets 65 A.2 Results of the densification of data sets A.2.1 Approach 1 (a) Real data of Monday holidays. (b) Virtual data of Monday holidays. Figure A.12: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Monday holidays. (a) Real data of Tuesday holidays. (b) Virtual data of Tuesday holidays. Figure A.13: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Tuesday holidays.
66 Results of the data treatment (a) Real data of Wednesday holidays. (b) Virtual data of Wednesday holidays. Figure A.14: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Wednesday holidays. (a) Real data of Thursday holidays. (b) Virtual data of Thursday holidays. Figure A.15: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Thursday holidays.
A.2 Results of the densification of data sets 67 (a) Real data of Friday holidays (1). (b) Virtual data of Friday holidays (1). Figure A.16: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Friday holidays (1). (a) Real data of Friday holidays (2). (b) Virtual data of Friday holidays (2). Figure A.17: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Friday holidays (2).
68 Results of the data treatment A.2.2 Approach 2 (a) Real data of Monday holidays. (b) Virtual data of Monday holidays. Figure A.18: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Monday holidays. (a) Real data of Tuesday holidays. (b) Virtual data of Tuesday holidays. Figure A.19: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Tuesday holidays.
A.2 Results of the densification of data sets 69 (a) Real data of Wednesday holidays. (b) Virtual data of Wednesday holidays. Figure A.20: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Wednesday holidays. (a) Real data of Thursday holidays. (b) Virtual data of Thursday holidays. Figure A.21: Box plots represent the evolution of the densification trick using ITLMS algorithm on group of Thursday holidays.
3 because these holidays, even falling on same weekday, have a different weekly behavior and the settings given to ITLMS led to consider these days in different clusters. B. Prediction results 1) The influence of the consideration of two different approachs of densification of data: As the clusters of real data are not equal for the two approaches, this comparison only was made to the three clusters with the same real data. The results achieved were: Table II PREDICTION RESULTS SUMMARY. COMPARISON BETWEEN TWO DIFFERENT APPROACHS OF DENSIFICATION OF DATA (A1 AND A2). Special day Real days std NMAE A B A B Monday 11 0.349% 0.349% 0.053% 0.053% Wednesday 14 0.466% 0.463% 0.234% 0.254% Thursday 22 0.312% 0.312% 0.544% 0.686% It is possible conclude that this two approaches are equivalents. For the Mondays predictions the same forecasting indicators were obtained. In the other two cluster the results were also very close. 2) Comparison with prediction results given in [3]: The results achieved in the paper [3] will be compared with the results that were given by the model found in this work, considering, logically, the same clusters of data, data sets of the Approach B of the densification of data performed in this paper. In order to provide a fuller understanding of the differences between this two forecasting methods: Table III CHARACTERISTICS OF TWO DIFFERENT FORECASTING METHODS. PROPOSED IN THIS WORK AND PROPOSED IN THE PAPER [3]. Based on Autoencoders Based on Feedforward NN [3] Training functions Bayesian Regulation Bp. Simple Backpropagation Forecasting Model Autoencoder 8-7-8, based on Simple Feedforward NN 7-3-1, missing data estimation, after the proper NN training, optimization performed by the the daily consumption forecasting EPSO in order to predict is made considering the seven the special day (missing data) inputs as the seven days that taking in consideration the seven procede the holidays and days that precede the holiday. the output the holiday predicted. Table IV PREDICTION RESULTS SUMMARY. COMPARISON WITH THE RESULTS ACHIEVED IN [3] Special day Real days std std [3] NMAE NMAE [3] Monday 11 0.35% 1.76% 0.05% 1.85% Tuesday 22 0.45% 2.50% 0.91% 2.82% Wednesday 14 0.46% 1.90% 0.25% 3.41% Thursday 22 0.31% 1.66% 0.69% 2.55% Friday 19 0.50% 2.46% 1.08% 3.92% The Normalized Mean Absolute Error (NMAE) varies from 0.05% (Monday) to 1.08% (Friday) for the model with Autoencoder, while the results in [3] were in the order of 1.85% (Monday) to 3.92% (Friday). The variation range of the corresponding standard deviation is from 0.31% to 0.50% better than in [3] with 1.66% to 2.50%. As can be demonstrated, with this new forecasting model into analysis, better results were achieved. An improvement of std prediction parameter in the order of 76% (Wednesday) to 82% (Tuesday) was achieved, and for NMAE in the order of 68% (Tuesday) to 97% (Monday). To all results was verified that the best results were achieved to the clusters with less days in real data. It is simple verify that this occur because with less data is more simple to find the correct pattern while for a cluster with more patterns becomes more difficult to interpolate the correct pattern. IV. CONCLUSION This paper demonstrates the advantage of using the Information Theoretic Learning Mean Shift algorithm, in the form of a desnsification trick. With this tool became possible the neuronal networks training even when faced with scarce data sets, the problem of special days. The ITLMS can be used to identify distinct clusters in the load data, associating in a easy way holidays that occurs in distinct days of the week and then allow the virtual data collected representing these specific clusters. It has been proven that the use of the virtual data as training set can be applied as if they were training with real data. The alliance of this powerful tool with an Auto Associative Neural Network, demonstrated to be a robust model in the load forecasting on special days. This Autoencoder based on missing data estimation, use an optimization performed by the metaheuristic EPSO in order to predict the special day (missing data) taking in consideration the days that precede the holiday. The high accuracy achieved by this method confirms that this tools can bring improvements in the performance of the load forecasting methodologies, specially, on days with occurence of scarce historical data to represent their behavior. Through the obtained results, it can be concluded that Autoencoders based on missing data estimation gives better results than a simple Feedforward Neural Network. This topology could be an important step in load forecasting on special days and even on normal days. REFERENCES [1] J. Fidalgo and J. Lopes, “Load forecasting performance enhancement when facing anomalous events,” Power Systems, IEEE Transactions on, vol. 20, no. 1, pp. 408–415, 2005. [2] A. Laukkanen, “The use of special day information in a demand forecasting model for nordic power market,” 2004. [3] J. Sumaili, V. Miranda, L. Rego, ´ A. Santana, and C. Francˆ es, “A densification trick using mean shift to allow demand forecasting in special days with scarce data,” 17th International Conference on Intelligent System Applications to Power Systems, pp. 1–5, Tokyo, Japan, 2013. [4] V. Miranda, A. Castro, and S. Lima, “Diagnosing faults in power transformers with autoassociative neural networks and mean shift,” Power Delivery, IEEE Transactions on, vol. 27, no. 3, pp. 1350–1357, 2012. [5] V. Miranda, J. Krstulovic, H. Keko, C. Moreira, and J. Pereira, “Reconstructing missing data in state estimation with autoencoders,” Power Systems, IEEE Transactions on, vol. 27, no. 2, pp. 604–611, 2012.
References [1] Sudhir Rao, Allan de Medeiros Martins, and J. C. Príncipe. Mean shift: An information theoretic perspective. Pattern Recogn. Lett., 30(3):222–230, February 2009. [2] Simon Haykin. Neural Networks: A Comprehensive Foundation. Prentice Hall PTR, Upper Saddle River, NJ, USA, 2nd edition, 1998. [3] V. Miranda. Evolutionary algorithms with particle swarm movements. In Intelligent Systems Application to Power Systems, 2005. Proceedings of the 13th International Conference on, pages 6–21, 2005. [4] J. Sumaili, V. Miranda, L. Rego, Á. Santana, and C. Francês. A densification trick using mean shift to allow demand forecasting in special days with scarce data. 17th International Conference on Intelligent System Applications to Power Systems, pages 1–5, Tokyo, Japan, 2013. [5] Neural Network Toolbox User’s Guide. [6] V. Miranda. Redes neuronais - treino por retropropagação. In Texto de apoio à disciplina de Controlo Difuso e Redes Neuronais no 5º ano da LEEC, FEUP, pages 207–212, Porto, 2007. [7] J.N. Fidalgo and J.A.P. Lopes. Load forecasting performance enhancement when facing anomalous events. Power Systems, IEEE Transactions on, 20(1):408–415, 2005. [8] M. Farhadi and M. Farshad. A fuzzy inference self-organizing-map based model for short term load forecasting. In Electrical Power Distribution Networks (EPDC), 2012 Proceedings of 17th Conference on, pages 1–9, 2012. [9] Chin Yen Tee, J.B. Cardell, and G.W. Ellis. Short-term load forecasting using artificial neural networks. In North American Power Symposium (NAPS), 2009, pages 1–6, 2009. [10] Kyung-Bin Song, Young-Sik Baek, Dug Hun Hong, and Gilsoo Jang. Short-term load forecasting for the holidays using fuzzy linear regression method. In Power Engineering Society General Meeting, 2005. IEEE, pages 1338 Vol. 2–, 2005. [11] Pauli Murto. Neural network models for short-term load forecasting. Master’s thesis, Helsinki University of Technology, 1998. [12] Kieran Richard Godden. Electric load forecasting for holiday periods. Master’s thesis, Faculty of Science at the Rand Afrikaans University, 1997. [13] Ching-Lai Hor, S.J. Watson, and S. Majithia. Analyzing the impact of weather variables on monthly electricity demand. Power Systems, IEEE Transactions on, 20(4):2078–2085, 2005. 77
78 REFERENCES [14] S. Ruzic, A. Vuckovic, and N. Nikolic. Weather sensitive method for short term load forecasting in electric power utility of serbia. Power Systems, IEEE Transactions on, 18(4):1581–1586, 2003. [15] Antti Laukkanen. The use of special day information in a demand forecasting model for nordic power market, 2004. [16] G. Gross and F.D. Galiana. Short-term load forecasting. Proceedings of the IEEE, 75(12):1558–1573, 1987. [17] H.S. Hippert, C.E. Pedreira, and R.C. Souza. Neural networks for short-term load forecasting: a review and evaluation. Power Systems, IEEE Transactions on, 16(1):44–55, 2001. [18] Yige Zhao, P.B. Luh, C. Bomgardner, and G.H. Beerel. Short-term load forecasting: Multilevel wavelet neural networks with holiday corrections. In Power Energy Society General Meeting, 2009. PES ’09. IEEE, pages 1–7, 2009. [19] L.A.D. de Luca, C.M. de Oliveira, and R.S. Wazlawick. Load behavior changes after holidays on thursdays. In Computational Science and Engineering Workshops, 2008. CSEWORKSHOPS ’08. 11th IEEE International Conference on, pages 101–106, 2008. [20] Manoj Kumar. Short-term load forecasting using artificial neural network techniques. Master’s thesis, National Institute of Technology Rourkela, 2009. [21] Wesin Ribeiro Alves. Modelos para previsão de carga a curto prazo através de redes neurais artificiais com treinamento baseado na teoria da informação. Master’s thesis, Universidade Federal do Pará Instituto de Tecnologia, 2011. [22] G. Chicco, Roberto Napoli, and Federico Piglione. Load pattern clustering for short-term load forecasting of anomalous days. In Power Tech Proceedings, 2001 IEEE Porto, volume 2, pages 6 pp. vol.2–, 2001. [23] R. Lamedica, A. Prudenzi, M. Sforna, M. Caciotta, and V.O. Cencellli. A neural network based technique for short-term forecasting of anomalous load periods. Power Systems, IEEE Transactions on, 11(4):1749–1756, 1996. [24] S. Ahmmed, M.A.A. Khan, M.K. Hasan, A.Y. Saber, M.N. Huda, and M.Z. Rahman. Stlf using neural networks and fuzzy for anomalous load scenarios - a case study for hajj. In Electrical and Computer Engineering (ICECE), 2010 International Conference on, pages 722–725, 2010. [25] Dou Quansheng, Pan Guanyu, Shi Zhongzhi, and Yang Bin. Knowledge extraction model for power load characteristics of special days and extreme weather. In Intelligent Computing and Intelligent Systems, 2009. ICIS 2009. IEEE International Conference on, volume 1, pages 496–499, 2009. [26] Qia Ding, Hui Zhang, Tao Huang, and Junyi Zhang. A holiday short term load forecasting considering weather information. In Power Engineering Conference, 2005. IPEC 2005. The 7th International, pages 1–61, 2005. [27] Kwang-Ho Kim, Hyoung sun Youn, and Yong-Cheol Kang. Short-term load forecasting for special days in anomalous load conditions using neural networks and fuzzy inference method. Power Systems, IEEE Transactions on, 15(2):559–565, 2000.
REFERENCES 79 [28] M. Farhadi and S. M. Moghaddas-Tafreshi. A novel model for short term load forecasting of iran power network by using kohonen neural networks. In Industrial Electronics, 2006 IEEE International Symposium on, volume 3, pages 1726–1731, 2006. [29] Young-Min Wi, Sung-Kwan Joo, and Kyung-Bin Song. Holiday load forecasting using fuzzy polynomial regression with weather feature selection and adjustment. Power Systems, IEEE Transactions on, 27(2):596–603, 2012. [30] L.A.D. de Luca. Previsao de carga em sistemas de potencia durante feriados prolongados: Efeito do feriado na quinta-feira sobre a carga da sexta-feira. Master’s thesis, Universidade Federal de Santa Catarina, 2008. [31] I. Aquino, C. Perez, J. K. Chavez, and S. Oporto. Daily load forecasting using quick propagation neural network with a special holiday encoding. In Neural Networks, 2007. IJCNN 2007. International Joint Conference on, pages 1935–1940, 2007. [32] Yuan-Yih Hsu and Chien-Chuen Yang. Design of artificial neural networks for short-term load forecasting. i. self-organising feature maps for day type identification. Generation, Transmission and Distribution, IEE Proceedings C, 138(5):407–413, 1991. [33] D. Srinivasan, C.S. Chang, and A.C. Liew. Demand forecasting using fuzzy neural computation, with special emphasis on weekend and public holiday forecasting. Power Systems, IEEE Transactions on, 10(4):1897–1903, 1995. [34] R. Barzamini, M.-B. Menhaj, A. Khosravi, and S. H. Kamalvand. Short term load forecasting for iran national power system and its regions using multi layer perceptron and fuzzy inference systems. In Neural Networks, 2005. IJCNN ’05. Proceedings. 2005 IEEE International Joint Conference on, volume 4, pages 2619–2624 vol. 4, 2005. [35] N. Mahdavi, M.-B. Menhaj, and S. Barghinia. Short-term load forecasting for special days using bayesian neural networks. In Power Systems Conference and Exposition, 2006. PSCE ’06. 2006 IEEE PES, pages 1518–1522, 2006. [36] S. Barghinia, S. Kamankesh, N. Mahdavi, A. H. Vahabie, and A.A. Gorji. A combination method for short term load forecasting used in iran electricity market by neurofuzzy, bayesian and finding similar days methods. In Electricity Market, 2008. EEM 2008. 5th International Conference on European, pages 1–6, 2008. [37] Qingqing Mu, Yonggang Wu, Xiaoqiang Pan, Liangyi Huang, and Xian Li. Short-term load forecasting using improved similar days method. In Power and Energy Engineering Conference (APPEEC), 2010 Asia-Pacific, pages 1–4, 2010. [38] Ying Chen, P.B. Luh, and S.J. Rourke. Short-term load forecasting: Similar day-based wavelet neural networks. In Intelligent Control and Automation, 2008. WCICA 2008. 7th World Congress on, pages 3353–3358, 2008. [39] Ying Chen, P.B. Luh, Che Guan, Yige Zhao, L.D. Michel, M.A. Coolbeth, P.B. Friedland, and S.J. Rourke. Short-term load forecasting: Similar day-based wavelet neural networks. Power Systems, IEEE Transactions on, 25(1):322–330, 2010. [40] D.C. Park, M.A. El-Sharkawi, II Marks, R.J., L.E. Atlas, and M.J. Damborg. Electric load forecasting using an artificial neural network. Power Systems, IEEE Transactions on, 6(2):442–449, 1991.
80 REFERENCES [41] Yun Lu, Xin Lin, and Weifu Qi. The method of short-term load forecasting based on the rbf neural network. In Electricity Distribution, 2005. CIRED 2005. 18th International Conference and Exhibition on, pages 1–4, 2005. [42] D. Srinivasan, A.C. Liew, and C.S. Chang. Forecasting daily load curves using a hybrid fuzzy-neural approach. Generation, Transmission and Distribution, IEE Proceedings-, 141(6):561–567, 1994. [43] S.S. Sharif and J.H. Taylor. Real-time load forecasting by artificial neural networks. In Power Engineering Society Summer Meeting, 2000. IEEE, volume 1, pages 496–501 vol. 1, 2000. [44] M.A. Aboul-Magd and E.E.-D.E.-S. Ahmed. An artificial neural network model for electrical daily peak load forecasting with an adjustment for holidays. In Power Engineering, 2001. LESCOPE ’01. 2001 Large Engineering Systems Conference on, pages 105–113, 2001. [45] A.P. Alves da Silva, U. P. Rodrigues, A.J.R. Reis, and L.S. Moulin. Neurodem-a neural network based short term demand forecaster. In Power Tech Proceedings, 2001 IEEE Porto, volume 2, pages 6 pp. vol.2–, 2001. [46] S. Barghinia, P. Ansarimehr, H. Habibi, and N. Vafadar. Short term load forecasting of iran national power system using artificial neural network. In Power Tech Proceedings, 2001 IEEE Porto, volume 3, pages 5 pp. vol.3–, 2001. [47] J.W. Taylor and R. Buizza. Neural network load forecasting with weather ensemble predictions. Power Systems, IEEE Transactions on, 17(3):626–632, 2002. [48] T. W S Chow and C.T. Leung. Neural network based short-term load forecasting using weather compensation. Power Systems, IEEE Transactions on, 11(4):1736–1742, 1996. [49] A. Piras, A. Germond, B. Buchenel, K. Imhof, and Y. Jaccard. Heterogeneous artificial neural network for short term electrical load forecasting. Power Systems, IEEE Transactions on, 11(1):397–402, 1996. [50] A. Khotanzad, R. Afkhami-Rohani, T.-L. Lu, A. Abaye, M. Davis, and D.J. Maratukulam. Annstlf-a neural-network-based electric load forecasting system. Neural Networks, IEEE Transactions on, 8(4):835–846, 1997. [51] H. Yoo and R.L. Pimmely. Short term load forecasting using a self-supervised adaptive neural network. Power Systems, IEEE Transactions on, 14(2):779–784, 1999. [52] Ku-Long Ho, Yuan-Yih Hsu, Chuan-Fu Chen, Tzong-En Lee, Chih-Chien Liang, TsauShin Lai, and Kung-Keng Chen. Short term load forecasting of taiwan power system using a knowledge-based expert system. Power Systems, IEEE Transactions on, 5(4):1214–1221, 1990. [53] C.N. Lu, H.-T. Wu, and S. Vemuri. Neural network based short term load forecasting. Power Systems, IEEE Transactions on, 8(1):336–342, 1993. [54] H. B. Gooi, C.Y. Teo, L. Chin, S.Y. Ang, and E. K. Khor. Adaptive short-term load forecasting using artificial neural networks. In TENCON ’93. Proceedings. Computer, Communication, Control and Power Engineering.1993 IEEE Region 10 Conference on, number 0, pages 787–790 vol.2, 1993.
REFERENCES 81 [55] Hui-Feng Shi and Yan-Xia Lu. Bayesian neural networks for short term load forecasting. In Wavelet Analysis and Pattern Recognition, 2009. ICWAPR 2009. International Conference on, pages 160–165, 2009. [56] Hui-Feng Shi and Yanxia Lu. Short-term load forecasting based on bayesian neural networks learned by hybrid monte carlo method. In Machine Learning and Cybernetics (ICMLC), 2010 International Conference on, volume 3, pages 1494–1499, 2010. [57] E.H. Tito, G. Zaverucha, M. Vellasco, and M. Pacheco. Bayesian neural networks for electric load forecasting. In Neural Information Processing, 1999. Proceedings. ICONIP ’99. 6th International Conference on, volume 1, pages 407–411 vol.1, 1999. [58] Yuan Ning, Yufeng Liu, and Qiang Ji. Bayesian - bp neural network based short-term load forecasting for power system. In Advanced Computer Theory and Engineering (ICACTE), 2010 3rd International Conference on, volume 2, pages V2–89–V2–93, 2010. [59] A. Baniamerian, M. Asadi, and E. Yavari. Recurrent wavelet network with new initialization and its application on short-term load forecasting. In Computer Modeling and Simulation, 2009. EMS ’09. Third UKSim European Symposium on, pages 379–383, 2009. [60] S.S. Rao and B. Kumthekar. Recurrent wavelet networks. In Neural Networks, 1994. IEEE World Congress on Computational Intelligence., 1994 IEEE International Conference on, volume 5, pages 3143–3147 vol.5, 1994. [61] Weijian Ren, Zhenghui Zhang, Yubo Duan, Qiong Wang, and Hongli Dong. An adaptive diagonal recurrent wavelet neural network based on compact wavelet frame and its application. In Control and Automation, 2005. ICCA ’05. International Conference on, volume 1, pages 599–604 Vol. 1, 2005. [62] S.M. Kelo and S.V. Dudul. Short-term load prediction with a special emphasis on weather compensation using a novel committee of wavelet recurrent neural networks and regression methods. In Power Electronics, Drives and Energy Systems (PEDES) 2010 Power India, 2010 Joint International Conference on, pages 1–6, 2010. [63] S. Uejima H. Tanaka and K. Asai. Linear regression analysis with fuzzy model. In IEEE Trans. Syst. Man Cybern, pages vol. 12, pp. 1291–1294, Dec. 1982. [64] H. Tanaka and J. Watada. Possibilistic linear systems and their application to linear regression model. In Fuzzy Sets and Syst., pages vol. 27, pp. 275–289, 1988. [65] J. Nazarko and W. Zalewski. The fuzzy regression approach to peak load estimation in power distribution systems. Power Systems, IEEE Transactions on, 14(3):809–814, 1999. [66] J. Nazarko and W. Zalewski. An application of the fuzzy regression analysis to the electrical load estimation. In Electrotechnical Conference, 1996. MELECON ’96., 8th Mediterranean, volume 3, pages 1563–1566 vol.3, 1996. [67] S.E. Papadakis, J.B. Theocharis, S. J. Kiartzis, and A.G. Bakirtzis. A novel approach to short-term load forecasting using fuzzy neural networks. Power Systems, IEEE Transactions on, 13(2):480–492, 1998. [68] B. Ye, N. N. Yan, C.X. Guo, and Y.J. Cao. Identification of fuzzy model for short-term load forecasting using evolutionary programming and orthogonal least squares. In Power Engineering Society General Meeting, 2006. IEEE, pages 8 pp.–, 2006.
82 REFERENCES [69] P.K. Dash, S. Dash, and S. Rahman. A fuzzy adaptive correction scheme for short term load forecasting using fuzzy layered neural network. In Neural Networks to Power Systems, 1993. ANNPS ’93., Proceedings of the Second International Forum on Applications of, pages 432–437, 1993. [70] P.K. Dash, A.C. Liew, and S. Rahman. Fuzzy neural network and fuzzy expert system for load forecasting. Generation, Transmission and Distribution, IEE Proceedings-, 143(1):106–114, 1996. [71] Ma-WenXiao, Bai-XiaoMin, and Mu-LianShun. Short-term load forecasting with artificial neural network and fuzzy logic. In Power System Technology, 2002. Proceedings. PowerCon 2002. International Conference on, volume 2, pages 1101–1104 vol.2, 2002. [72] P.K. Dash, G. Ramakrishna, A.C. Liew, and S. Rahman. Fuzzy neural networks for time-series forecasting of electric load. Generation, Transmission and Distribution, IEE Proceedings-, 142(5):535–544, 1995. [73] Kwang-Ho Kim, Jong-Keun Park, Kab-Ju Hwang, and Sung-Hak Kim. Implementation of hybrid short-term load forecasting system using artificial neural networks and fuzzy expert systems. Power Systems, IEEE Transactions on, 10(3):1534–1539, 1995. [74] C. Jaipradidtham. Next day load demand forecasting of future in electrical power generation on distribution networks using adaptive neuro-fuzzy inference. In Power and Energy Conference, 2006. PECon ’06. IEEE International, pages 64–67, 2006. [75] V. Miranda, A.R.G. Castro, and S. Lima. Diagnosing faults in power transformers with autoassociative neural networks and mean shift. Power Delivery, IEEE Transactions on, 27(3):1350–1357, 2012. [76] Sudhir Rao, Weifeng Liu, J.C. Principe, and A. de Medeiros Martins. Information theoretic mean shift algorithm. In Machine Learning for Signal Processing, 2006. Proceedings of the 2006 16th IEEE Signal Processing Society Workshop on, pages 155–160, 2006. [77] K. Fukunaga and L. Hostetler. The estimation of the gradient of a density function, with applications in pattern recognition. Information Theory, IEEE Transactions on, 21(1):32– 40, 1975. [78] Emanuel Parzen. On estimation of a probability density function and mode. The Annals of Mathematical Statistics, 33(3):pp. 1065–1076, 1962. [79] Yizong Cheng. Mean shift, mode seeking, and clustering. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 17(8):790–799, 1995. [80] D. Comaniciu and P. Meer. Mean shift: a robust approach toward feature space analysis. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 24(5):603–619, 2002. [81] Richard Szeliski. Computer Vision: Algorithms and Applications. Springer-Verlag New York, Inc., New York, NY, USA, 1st edition, 2010. [82] A. Pooransingh, C.-A. Radix, and A. Kokaram. The path assigned mean shift algorithm: A new fast mean shift implementation for colour image segmentation. In Image Processing, 2008. ICIP 2008. 15th IEEE International Conference on, pages 597–600, 2008.
REFERENCES 83 [83] Zhi-Qiang Wen and Zi xing Cai. Mean shift algorithm and its application in tracking of objects. In Machine Learning and Cybernetics, 2006 International Conference on, pages 4024–4028, 2006. [84] K.A. Shah, H.K. Kapadia, V.A. Shah, and M.N. Shah. Application of mean-shift algorithm for license plate localization. In Engineering (NUiCONE), 2011 Nirma University International Conference on, pages 1–5, 2011. [85] Pengfei Li, Shaoru Wang, and Junfeng Jing. The segmentation in textile printing image based on mean shift. In Computer-Aided Industrial Design Conceptual Design, 2009. CAID CD 2009. IEEE 10th International Conference on, pages 1528–1532, 2009. [86] A. Renyi. Some fundamental questions of information theory. In Selected Papers of Alfred Renyi, Akademia Kiado, Budapest, volume 2, page 526–552, 1976. [87] J. C. Principe, N. R. Euliano, and W. C. Lefebvre. Neural and adaptive systems: fundamentals through simulations. Wiley New York, 2000. [88] G. E. Hinton and R. R. Salakhutdinov. Reducing the dimensionality of data with neural networks. In Science, pages vol. 313, no. 5786, pp. 504–507, Jul. 2006. [89] Terence D. Sanger. Optimal unsupervised learning in a single-layer linear feedforward neural network, 1989. [90] I. T. Jolliffe. Principal Component Analysis. Springer, second edition, October 2002. [91] Nathalie Japkowicz, Stephen Jose Hanson, and Mark A. Gluck. Nonlinear autoassociation is not equivalent to pca. Neural Comput., 12(3):531–545, March 2000. [92] V. Miranda, J. Krstulovic, H. Keko, C. Moreira, and J. Pereira. Reconstructing missing data in state estimation with autoencoders. Power Systems, IEEE Transactions on, 27(2):604– 611, 2012. [93] Mussa Abdella and T. Marwala. The use of genetic algorithms and neural networks to approximate missing data in database. In Computational Cybernetics, 2005. ICCC 2005. IEEE 3rd International Conference on, pages 207–212, 2005. [94] Fulufhelo Vincent Nelwamondo, Dan Golding, and Tshilidzi Marwala. A dynamic programming approach to missing data estimation using neural networks. Inf. Sci., 237:49–58, 2013. [95] Mohamed S. Nelwamondo, F. V. and T. Marwala. Missing data: A comparison of neural networks and expectation maximization techniques. Current Science, 93:pp1514, December 2007. [96] S. Mohamed and T. Marwala. Neural network based techniques for estimating missing data in databases. The 16th Annual Symposium of the Pattern Recognition Association of South Africa, Langebaan, South Africa, pages pp. 27–32, 2005. [97] Geoffrey Hinton and Ruslan Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313(5786):504 – 507, 2006. [98] P. Munro G. W. Cottrell and D. Zipser. Learning internal representations from gray-scale images: An example of extensional programming. in Proc. 9th Annu. Conf. Cognitive Science Society, Seattle,WA, 1987.
84 REFERENCES [99] M.K. Fleming and G.W. Cottrell. Categorization of faces using unsupervised feature extraction. In Neural Networks, 1990., 1990 IJCNN International Joint Conference on, pages 65–70 vol.2, 1990. [100] B. Golomb, T. Sejnowski, and Howard Hughes. Sex recognition from faces using neural networks. In Applications of Neural Networks, pages 71–92. Editor), Kluwer Academic Publishers, 1995. [101] S. Narayanan, II Marks, R. J., J.L. Vian, J. J. Choi, M. A. El-Sharkawi, and B.B. Thompson. Set constraint discovery: missing sensor data restoration using autoassociative regression machines. In Neural Networks, 2002. IJCNN ’02. Proceedings of the 2002 International Joint Conference on, volume 3, pages 2872–2877, 2002. [102] B.B. Thompson, R.J. Marks, and M.A. El-Sharkawi. On the contractive nature of autoencoders: application to missing sensor restoration. In Neural Networks, 2003. Proceedings of the International Joint Conference on, volume 4, pages 3011–3016 vol.4, 2003. [103] Wei Qiao, Zhi Gao, Ronald G. Harley, and Ganesh K. Venayagamoorthy. Robust neuroidentification of nonlinear plants in electric power systems with missing sensor measurements. Eng. Appl. Artif. Intell., 21(4):604–618, June 2008. [104] Marwala T Leke-Betechuoh B and Tettey T. Autoencoder networks for hiv classification. Current Science, 91, No. 11:pp1467–1473, December 2006. [105] Tshilidzi Marwala Sizwe M. Dhlamini, Fulufhelo V. Nelwamondo. Sensor failure compensation techniques for hv bushing monitoring using evolutionary computing. Tenerife, Spain, pages 430–435, December 16-18, 2005. [106] S. Mohagheghi, G.K. Venayagamoorthy, and R.G. Harley. Optimal wide area controller and state predictor for a power system. Power Systems, IEEE Transactions on, 22(2):693–705, 2007. [107] S. Narayanan, J.L. Vian, J.J. Choi, II Marks, R.J., M.A. El-Sharkawi, and B.B. Thompson. Missing sensor data restoration for vibration sensors on a jet aircraft engine. In Neural Networks, 2003. Proceedings of the International Joint Conference on, volume 4, pages 3007–3010 vol.4, 2003. [108] V. Miranda. Computação evolucionária: uma introdução. Technical report, FEUP, 2005. [109] Keko H. Miranda, V. and Á. J. Duque. Stochastic star communication topology in evolutionary particle swarms (epso). International Journal of Computational Intelligence Research (IJCIR), 4:105–116, 2008. [110] J. Kennedy and R. Eberhart. Particle swarm optimization. In Neural Networks, 1995. Proceedings., IEEE International Conference on, volume 4, pages 1942–1948 vol.4, 1995. [111] V. Miranda and N. Fonseca. Epso - best-of-two-worlds meta-heuristic applied to power system problems. In Proceedings of the Evolutionary Computation on 2002. CEC ’02. Proceedings of the 2002 Congress - Volume 02, CEC ’02, pages 1080–1085, Washington, DC, USA, 2002. IEEE Computer Society. [112] Yuhui Shi and Russell C. Eberhart. Parameter selection in particle swarm optimization. In Proceedings of the 7th International Conference on Evolutionary Programming VII, EP ’98, pages 591–600, London, UK, UK, 1998. Springer-Verlag.
REFERENCES 85 [113] M. Clerc. The swarm and the queen: towards a deterministic and adaptive particle swarm optimization. In Evolutionary Computation, 1999. CEC 99. Proceedings of the 1999 Congress on, volume 3, pages –1957 Vol. 3, 1999. [114] V. Miranda, C. Cerqueira, and C. Monteiro. Training a fis with epso under an entropy criterion for wind power prediction. In Probabilistic Methods Applied to Power Systems, 2006. PMAPS 2006. International Conference on, pages 1–8, 2006. [115] H. Leite, J. Barros, and V. Miranda. Evolutionary algorithm epso helping doubly-fed induction generators in ride-through-fault. In PowerTech, 2009 IEEE Bucharest, pages 1–8, 2009. [116] Naing Win Oo and V. Miranda. Evolving agents in a market simulation platform - a test for distinct meta-heuristics. In Intelligent Systems Application to Power Systems, 2005. Proceedings of the 13th International Conference on, pages 6 pp.–, 2005. [117] V. Miranda and N. Fonseca. Epso-evolutionary particle swarm optimization, a new algorithm with applications in power systems. In Transmission and Distribution Conference and Exhibition 2002: Asia Pacific. IEEE/PES, volume 2, pages 745–750 vol.2, 2002. [118] H. Leite, J. Barros, and V. Miranda. The evolutionary algorithm epso to coordinate directional overcurrent relays. In Developments in Power System Protection (DPSP 2010). Managing the Change, 10th IET International Conference on, pages 1–5, 2010. [119] V. Miranda, L. de Magalhaes Carvalho, M.A. da Rosa, A.M. Leite Da Silva, and C. Singh. Improving power system reliability calculation efficiency with epso variants. Power Systems, IEEE Transactions on, 24(4):1772–1779, 2009. [120] H. Keko, A.J. Duque, and V. Miranda. A multiple scenario security constrained reactive power planning tool using epso. In Intelligent Systems Applications to Power Systems, 2007. ISAP 2007. International Conference on, pages 1–6, 2007. [121] K. Levenberg. A method for the solution of certain nonlinear problems in least squares. The Quarterly of Applied Mathematics, 2, pages 164–168, 1944. [122] Donald W. Marquardt. An algorithm for least-squares estimation of nonlinear parameters. SIAM Journal on Applied Mathematics, 11(2):431–441, 1963. [123] Henri Gavin. The levenberg-marquardt method for nonlinear least squares curve-fitting problems. September 28 2011. [124] M.T. Hagan and M.-B. Menhaj. Training feedforward networks with the marquardt algorithm. Neural Networks, IEEE Transactions on, 5(6):989–993, 1994. [125] David J.C. MacKay. Bayesian interpolation. Neural Computation, 4:415–447, 1991. [126] David J.C. MacKay. A practical bayesian framework for backprop networks. Neural Computation, 4:448–472, 1991. [127] F. Dan Foresee and M.T. Hagan. Gauss-newton approximation to bayesian learning. In Neural Networks,1997., International Conference on, volume 3, pages 1930–1935 vol.3, 1997.