Sensory fusion applied to power system state estimation considering information theory concepts
Abstract
With the increasing integration of synchronized phasor measurement units (PMU) in power grids appears the necessity to create methods capable to merge the information obtained from different classes of sensors, namely PMU and the conventional sensors already integrated in the SCADA system. Thus, this dissertation proposes a sensory fusion that guarantees the previous requirement. Beyond that, this thesis proposes an application of concepts related with Information theory to the state estimation problem. This application aims to propose a robust state estimator without the necessity of previous treatment of the acquired data.
Full text
Faculdade de Engenharia da Universidade do Porto Sensory fusion applied to power system state estimation considering information theory concepts Bruna Tavares Final Version Dissertation conducted under the Integrated Master in Electrical and Computers Engineering Major Energy Supervisor: Professor Doutor Vladimiro Henrique Barrosa Pinto de Miranda Second Supervisor: Professor Doutor Jorge Pereira 26/7/2016
ii © Bruna Tavares, 2016
iii Resumo Com o aparecimento de sincrofasores e com o crescente interesse na sua integração nas redes elétricas torna-se imperioso o desenvolvimento de métodos capazes de fundir a informação obtida através de diferentes classes de sensores, nomeadamente, informação obtida através de sensores convencionais integrados no sistema SCADA e sincrofasores, também conhecidos como Phase Meaurement Units - PMU. Assim sendo, esta tese apresenta uma proposta de fusão sensorial que permite a integração de diferentes classes de sensores na estimação de estado podendo ser dada mais confiança a uma classe ou a outra. Para além da fusão sensorial, este trabalho foca-se também na integração de conceitos relacionados com teoria de informação à estimação de estado, propondo uma alternativa à convencional estimação de estado resolvida através de mínimos quadrados. A substituição do critério de mínimos quadrados por conceitos relacionados com teoria de informação pretende propor uma nova vertente para um estimador de estado robusto sem necessidade de pré-tratamento da informação recolhida pelos sensores.
iv
v Abstract With the increasing interest on the integration of synchronized phasor measurement units (PMU) in electric network appears the necessity to create methods capable to merge the information obtained from different classes of sensors, namely, the conventional sensors already integrated in the SCADA system and synchronized phasor measurements units, also known as PMU. Thus, this dissertation proposes a sensory fusion method that guarantees the previous requirement, allowing the integration of different classes of sensors for the state estimation and assignment of different level of trust to each sensory system. Beyond that, this thesis focuses on the appliance of information theory related concepts on the state estimation, proposing an alternative to the conventional state estimation solved by least square error. The substitution of the least square criterion of the state estimator by information theory concepts aimed at proposing a robust state estimator without the necessity of previous treatment of the acquired data. Key words: Correntropy , Information theory, Parzen window, Phasor Measurement Units, Renyi’s quadratic entropy, sensory fusion, state estimation, weighted-least square error.
vi
vii Acknowledgments First of all, I would like to thank my thesis advisor Professor Vladimiro Miranda of the Electrical and Computer Engineering Department at Porto’s University. From the great and innovative ideas to the urge motivation, Prof. Miranda was always a sustainer of this work. Without his support and guidance, this thesis wouldn’t have followed such a great path. I also want to address an acknowledgment to my co-supervisor, Professor Jorge Pereira. The door to Prof. Pereira’s office was always open whenever I needed some help with challenges on my research or with problem resolution. I want also to thank INESC Porto for the good reception and opportunity to integrate the CPES department. A special acknowledgment to CPES team, for all the support and good moments inside and outside INESC Porto. I was not expecting such a great environment, support and friendship between all the members in a company team. That was a good surprise and help my daily work with a good mood. Finally, I must express my gratitude to my family, specially to my parents, for all the unfailing support and continuous motivation through all my life and my study years. Without their support I would never be the person I am today.
viii
ix “Pleasure in the job puts perfection in the work” -Aristotle
xvi Figure 4-5Residual errors (p.u.) for the group one using Correntropy, varying α. On the right side the point α=0 was ignored. ............................................................. 42 Figure 4-6Residual errors (p.u.) for the group one using Mean Squared Error, varying α. On the right side the point α=0 was ignored. ........................................................ 42 Figure 4-7 - Residual errors (p.u.) for the group one using Renyi’s Quadratic entropy(case one), varying α. On the right side the point α=0 was ignored. ............................... 43 Figure 4-8 – Residual errors (p.u.) for the group one using Renyi's Quadratic Entropy (case 2), varying α. .......................................................................................... 43 Figure 4-9Residual errors (p.u.) for the group two, using Renyi’s Quadratic entropy(case one), varying α. On the right side the point α=1 was ignored. ............................... 43 Figure 4-10-Residual errors (p.u.) for the group two, using MCC on the left and MSE on the right, varying α. ....................................................................................... 44 Figure 4-11-Residual errors (p.u.) for the group two, using Renyi’s Quadratic Entropy (case 2), varying α ........................................................................................... 44 Figure 5-1Typical Pareto-front obtained by the multiple fusion points, depending on α and β, in black. Ideal point (maxfit1, maxfit2) represented by the grey point. ................ 47 Figure 5-2 – Residual errors for each estimation of power injections and lines power flows, using L1 as the fusion metric and each one of the evaluation metrics. ..................... 51 Figure 5-3 - Residual errors for each estimation of Voltage phase, using L1 as the fusion metric and each one of the evaluation metrics. ................................................ 51 Figure 5-4 - Residual errors for each estimation of power injections and lines power flows, using L2 as the fusion metric and each one of the evaluation metrics. ..................... 51 Figure 5-5 - Residual errors for each estimation of Voltage phase, using L2 as the fusion metric and each one of the evaluation metrics. ................................................ 52 Figure 5-6 - Residual errors in buses power injections and line power flows, using CIM as the fusion metric (with different Parzen window’s size )and MSE as evaluation metric. .... 53 Figure 5-7Residual errors in voltage phases, using CIM as the fusion metric (with different Parzen window’s size )and MSE as evaluation metric. ......................................... 53 Figure 5-8 - Residual errors in buses power injections and line power flows, using CIM as the fusion metric (with different Parzen window’s size )and MCC as evaluation metric. ... 54 Figure 5-9 - Residual errors in voltage phases, using CIM as the fusion metric (with different Parzen window’s size )and MCC as evaluation metric. ......................................... 54 Figure 5-10 - Residual errors in buses power injections and line power flows, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 1) as evaluation metric. .................................................................................................. 54 Figure 5-11 - Residual errors in voltage phases, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 1) as evaluation metric. ............................... 55 Figure 5-12 - Residual errors in buses power injections and line power flows, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 2) as evaluation metric. .................................................................................................. 55
xvii Figure 5-13 - Residual errors in voltage phases, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 2) as evaluation metric. ............................... 55 Figure 5-14 - Residual errors in buses power injections and line power flows using CIM as the fusion metric, with Parzen size 0.0001. ...................................................... 56 Figure 5-15 - Residual errors in voltage phases using CIM as the fusion metric, with Parzen size 0.0001. ............................................................................................ 56 Figure 5-16 - Residual errors in buses power injections and lines power flows using L1 as the fusion metric, with σ= 2 and σ=0.1. .......................................................... 58 Figure 5-17 - Residual errors in voltage phases using L1 as the fusion metric, with σ=2 and σ=0.1. ................................................................................................... 59 Figure 5-18Residual errors in buses power injections and lines power flows using L1 as the fusion metric, with σ=0.1, and ignoring the residual error of the component Pinj3. ........................................................................................................... 59 Figure 5-19 - Residual errors in voltage phases using L1 as the fusion metric , with σ=0.1. ... 60 Figure 5-20 - Residual errors in buses power injection and lines power flow, using CIM as the fusion metric, with σ=1. Evaluation metrics with σ=2 and σ=0.1. ...................... 60 Figure 5-21 - Residual errors in voltage phases, using CIM as the fusion metric, with σ=1. Evaluation metrics with σ=2 and σ=0.1. .......................................................... 61 Figure 5-22 - Residual errors in buses power injections and lines power flows, using CIM as the fusion metric (with σ=1) and ignoring the residual error of the component Pinj3. Evaluation metrics MCC and RQE (2) with σ=0.1, RQE (1) with σ=0.5. ...................... 61 Figure 5-23 - Residual errors in voltage phases, using CIM as the fusion metric (with σ=1). Evaluation metrics MCC and RQE (2) with σ=0.1, RQE (1) with σ=0.5. ...................... 62 Figure 5-24 - Power injections residual errors obtained with different fusion methods. MCC, with σ=0.1, and MSE as evaluation metrics. ...................................................... 64 Figure 5-25 – Voltage phases residual errors obtained with different fusion methods. MCC, with σ=0.1, and MSE as evaluation metrics. ...................................................... 64 Figure 5-26 - Power injections residual errors obtained with different fusion methods with RQE as evaluation metric, with σ=0.1. ............................................................ 65 Figure 5-27 – Voltage phases residual errors obtained with different fusion methods with RQE as evaluation metric, with σ=0.1. ............................................................ 66 Figure 6-1 - Absolute residual errors in power injection, obtained without PMU. ............... 72 Figure 6-2 – Absolute residual errors in voltage magnitude, obtained without PMU............. 72 Figure 6-3Absolute residual errors in power and current injections and lines power flow, obtained with PMU. ................................................................................... 73 Figure 6-4 – Absolute residual errors in voltage magnitudes, on the left, and on voltage phases, on the right. Errors obtained with PMU. ................................................ 73 Figure 6-5 - Errors between estimations and original values of injected power. ................ 74
xviii Figure 6-6 - Errors between estimations and original values of injected currents and lines power flows. ........................................................................................... 74 Figure 6-7 - Errors between estimations and original values of voltages phases. ............... 74 Figure 6-8 - Errors between estimations and original values of voltages magnitudes. .......... 75 Figure 6-9 - Absolute value of residual errors in injected power and currents and lines power flow for the AC model, without gross errors. .................................................... 77 Figure 6-10Absolute value of residual errors in voltage phases for the AC model, without gross errors. ............................................................................................ 77 Figure 6-11Absolute value of residual errors in voltage magnitudes for the AC model, without gross errors. ................................................................................. 77 Figure 6-12 - pdf of the errors for the case without outliers, using MSE, on the left, and MCC on the right. ........................................................................................... 78 Figure 6-13pdf of the errors for the case without outliers, using RQE. .......................... 78 Figure 6-14 – Absolute value of residual errors in injected power and currents and lines power flow for the AC model with an missing measurement. ................................ 79 Figure 6-15 - Absolute value of residual errors in voltage phases for the AC model with an missing measurement. ............................................................................... 79 Figure 6-16Absolute value of residual errors in voltage magnitudes for the AC model with an missing measurement. ............................................................................ 80 Figure 6-17 - pdf of the residual errors for the case of an missing measurement, using RQE, on the left, and MCC on the right. ................................................................. 80 Figure 6-18 - pdf of the residual errors for the case of an missing measurement, using MSE. ........................................................................................................... 81 Figure 6-19 - Absolute value of residual errors in injected power and currents and lines power flow for the AC model with inversion on one line power flow. ...................... 81 Figure 6-20 - Absolute value of residual errors in injected power and currents and lines power flow for the AC model with inversion on one line power flow. ...................... 82 Figure 6-21 - Absolute value of residual errors in injected power and currents and lines power flow for the AC model with inversion on one line power flow. ...................... 82 Figure 6-22 - pdf of the residual errors for the case of an inversion of a line power flow, using RQE, on the left, and MCC on the right. ................................................... 83 Figure 6-23 - pdf of the residual errors for the case of an inversion of a line power flow, using MSE. .............................................................................................. 83 Figure A-1-Network topology. ............................................................................. 94 Figure B-1 - Residual errors in buses power injections and lines power flows, using L2 as the fusion metric, with σ=2 and σ=0.1. ............................................................... 105
xix Figure B-2 - Residual errors in voltage phases using L2 as the fusion metric, with σ=2 and σ=0.1. .................................................................................................. 105 Figure B-3 - Residual errors in buses power injections and lines power flows using L2 as the fusion metric, with σ=0.1 and ignoring the residual error of the component Pinj3. ..... 106 Figure B-4Residual errors in voltage phases using L2 as the fusion metric, with σ=0.1. ..... 106 Figure B-5 - Residual errors in buses power injections and lines power flows using CIM as the fusion metric, with σ=1 and ignoring the residual error of the component Pinj3. Evaluation metrics with σ=0.1. .................................................................... 106 Figure B-6 - Residual errors for each estimation of voltage phases using CIM as the fusion metric, with σ=1. Evaluation metrics with σ=0.1. ............................................. 107 Figure B-7 - Residual errors in buses power injections and lines power flows, using CIM as the fusion metric with σ=1. Evaluation metric: RQE (case 1) with σ=0.1, σ=0.5 and σ=0.05. ................................................................................................. 107 Figure B-8 - Residual errors in voltage phases, using CIM as the fusion metric with σ=1. Evaluation metric: RQE (case 1) with σ=0.1, σ=0.5 and σ=0.05. ............................ 107
xx
xxi List of tables Table 3-1Different objective functions. ............................................................... 26 Table 3-2 - Measurements used to compare the different criteria . ............................... 27 Table 3-3State estimations provided by each criterion. ............................................ 28 Table 3-4Residual errors obtained from each criterion. ............................................ 28 Table 3-5State estimations provided by each criteria, with different PW sizes. Gross error in Pinj9.................................................................................................... 31 Table 3-6Residual errors for in the different estimations, when Pinj9 has a gross error. ...... 31 Table 3-7 - State estimations provided by each criteria, with different PW sizes. Gross error in Pinj3.................................................................................................... 32 Table 3-8 – Residual errors in the different estimations when Pinj3 has a gross error. .......... 32 Table 3-9Estimations and residual errors for a too small Parzen Window. ...................... 33 Table 4-1 - Measurements origin .......................................................................... 38 Table 5-1Objective function using the metric L1 to find the closest point to the ideal one. ........................................................................................................... 48 Table 5-2 - Objective function using the Euclidian metric to find the closest point to the ideal one. ............................................................................................... 49 Table 5-3 - Objective function using the Correntropy Induced Metric to find the closest point to the ideal one. ............................................................................... 49 Table 5-4 - Measurements used to find the optimal point (with and without gross errors). Part one: active power. .............................................................................. 49 Table 5-5 - Measurements used to find the optimal point (with and without gross errors). Part two: Voltage phases and lines power flows. ............................................... 50 Table 5-6Ideal point to the DC model without gross errors. ....................................... 50 Table 5-7 - Maximum and minimum values each fitness evaluation reaches for each group of sensors using the different metrics as objective function of the estimation. .......... 52
xxii Table 5-8 - Ideal point, calculated with σ=2, to the measurements affected with gross error . .......................................................................................................... 57 Table 5-9 - Ideal point calculated with σ=0.1 to the measurements affected with gross error . ................................................................................................... 58 Table 5-10 - Number of fitness evaluation each combination took to find the optimal point. ........................................................................................................... 62 Table 5-11 – Duration of the process each combination took to find the optimal point. ....... 62 Table 5-12 – Simplified objective function using the metric L1 to find the closest point to the ideal one. .......................................................................................... 65 Table 5-13Errors’ entropy and mean absolute error for the results obtained with simple fusion. ................................................................................................... 66 Table 5-14 – Errors’ entropy and mean absolute percentage error for the results obtained with fusion metrics L1 and L2....................................................................... 66 Table 6-1Measurements used in the AC model without gross errors. ............................ 70 Table 6-2 - Measurements used in the AC model with gross errors. ................................ 71 Table 6-3 - Mean percentage error and entropy calculated to each sensor system ............. 75 Table 6-4 - Objective function to find the optimal point by the L1 metric. ...................... 76 Table 6-5 – Simplification for the SE objective function, given by the L1 metric. ............... 76 Table A-1 Line parameters in ohm........................................................................ 94 Table A-2Line parameters in p.u. ....................................................................... 95 Table A-3 - Apparent power and power factor for the loads in each bus. ........................ 95 Table A-4 - Maximum power generated by each DER. ................................................ 96 Table A-5Power distribution according time. ......................................................... 96 Table A-6Loads in hour 17 a.m. ......................................................................... 97 Table A-7Generated Power by the DER in the at 17a.m. .......................................... 97 Table A-8Injected Power in MW and Mvar at 17 a.m. ............................................... 97 Table A-9Injected Power in p.u. at 17 a.m. .......................................................... 98 Table A-10Voltage profile at 17 a.m. .................................................................. 98 Table A-11-Power flow results and generated measurements without PMU. ..................... 99 Table A-12 -Power flow results and generated measurements with PMU, for injected power. .......................................................................................................... 100 Table A-13 - Power flow results and generated measurements with PMU, for injected current and bus voltages............................................................................ 101 Table A-14Measurements of lines power flows to the AC model ................................. 101
xxiii Table A-15PSS/E results considering R=0 ............................................................. 102 Table A-16-Power flow results and generated measurements, for active power, considering R=0. ..................................................................................................... 102 Table A-17 - Power flow results and generated measurements, for reactive power and voltages, considering R=0. .......................................................................... 103 Table A-18DC measurements considering the adapted network. R=0. .......................... 103 Table A-19Power Flow measurements in lines 3-4 and 6-7 assuming DC model. .............. 104
xxiv
xxv List of acronyms and symbols List of acronyms CIM Correntropy Induced Metric CS Cauchy-Schwarz CPES Center for Power Energy Systems DER Distributed Energy Resources EMS Energy Management System INESC TEC INstituto de Engenharia de Sistemas e ComputadoresTecnologia e Ciência LSE Least Square Error MAPE Mean Absolute Percentage Error MCC Maximum Correntropy Criterion MSE Mean Squared Error p.u. Per Unit pdf Probability Density Function PSS/E Power Transmission System Planning Software PW Parzen Window RQE Renyi’s Quadratic Entropy SCADA Supervisory Control And Data Acquisition SE State Estimation WLS Weighted Least Square List of symbols σ Standard deviation, also used in this work as PW’s size ϴ Voltage phase α Trust assign to the measurements obtained by sensors from group one. β Trust assign to the measurements obtained by sensors from group two.
6 State of art the message rate is too high for the channel, it is necessary to limit the massage rate. That limitation leads to a need of optimally messages compression. The compression is done as established by the rate distortion theory, minimizing the mutual information between the original message and its compressed version. After that, there is the phase of error correction. It’s necessary to encode the source-compressed data for error-free transmission in order to withstand the channel noise in the transmission. This step is achieved by maximizing the mutual information between the source-compressed message and the received message, as established by the channel capacity theorem. The last step, consists on the decompression. It aims to decode and decompress the received message to recover the original message. Basically, first it minimizes redundancy for efficiency and then it adds redundancy to mitigate the noise effects in the channel. The point here is that the same statistical descriptor, mutual information, is specifying the two compromises for error-free communication: the data compression minimum limit (rate distortion) and the data transmission maximum limit (channel capacity) [5]. The mathematical formulation for Shannon’s Entropy is explored in the next sub-section. 2.1.2 - Shannon’s entropy Let’s assume a random variable X, with xiϵIRD and a random sample A with n pairs {xi,p(xi)}, assuming that the probability density function (pdf) of the variable x is described by P={(xi,pi(xi)), i=1,2…n } or simply P={(xi,pi), i=1,2…n }, the entropy or uncertainty through Shannon is given by : 𝐻1(𝑥)=𝐸[−log(𝑃)]=−∑𝑝𝑖𝑙𝑜𝑔𝑝𝑖 𝑛𝑖=1 , (2.1) where ∑𝑝𝑖 𝑛𝑖=1 =1 and pi≥0. Therefore, Shannon defined the uncertainty of the variable X as the sum across the set of uncertainty in each component weighted by the probability of each component. The uncertainty of each component is given by the 𝑙𝑜𝑔𝑝𝑖 . For continuous pdf distributions we have: 𝐻1(𝑥)=−∫𝑝(𝑥)log(𝑝(𝑥))𝑑𝑥 . (2.2) The Shannon’s entropy is seen as a measure of uncertainly in the information content of the pdf p(x). If the logarithm basis is 2, it obtains the information quantity is measured in bits. 2.1.3 - Renyi’s entropy In the mid-1950s, Alfred Renyi introduced the parametric family of entropies as a mathematical generalization of Shannon’s Entropy. Renyi’s entropy appeared when he tried to find the most general class of information measure. That measure should allow to sum statistically independent systems and should be compatible with Kolmogorov’s probability axioms [5]. For the same conditions of the previous subsection the Renyi’s entropy is defined as a family of functions related to a parameter α as follows 𝐻𝛼=1 1−𝛼𝑙𝑜𝑔∑𝑝𝑖𝛼𝑛𝑖=1 , (2.3) with α>0 and α≠1. When α=1 the expression losses its meaning. However, it’s possible to demonstrate that when 𝛼→1− 𝑎𝑛𝑑 𝛼→1+ Hα tends to H1 -Shannon’s entropy. Thus, lim 𝛼→1𝐻𝛼=𝐻1.
Information Theoretic Learning 7 Therefore, it can be assumed that the Renyi’s entropy converges to Shannon entropy bilaterally. Also, Shannon’s entropy is between the possible values of Renyi’s definition, with real parameters α and β. One has 𝐻𝛼≥𝐻1≥𝐻𝛽,𝑤here 0<α<1<β. Comparing Shannon’s and Renyi’s entropies, equations (2.2) and (2.3), the main difference is the placement of the logarithm in the expression. While in Shannon’s entropy the log(pi) term is weighted, in Renyi’s entropy the log is outside the term that involves the α power of the probability mass function. The parameter α makes the Renyi’s entropy much more flexible enabling some different measurements of dissimilarity within a given distribution. Usually, Hα(x) is called the spectrum of Renyi information, seen as a function of α, and its graphical plot is useful in statistical inference. In [5] is described some important axioms, as: 1. Renyi’s Entropy measure, Hα(p1, p2, … pn), is a continuous function of all the probabilities pk. In this way, a small change in the probability distribution results in a small variation of the entropy. 2. Hα(p1, p2, … pn) is permutationally symmetric. In other words, if the order within the vector p changes, the entropy value remains the same. The permutation of any pk in the distribution doesn’t affect the uncertainty or disorder of the distribution and, thus, it should not affect the entropy; 3. H(1/n, …, 1/n) is a monotonic increasing function of n. In an equiprobable distribution, the increase number of choices leads to an increase of uncertainty or disorder, increasing also the entropy measure. 4. Recursivity. In a recursive function it should possible to observe the follow condition 𝐻α(𝑝1,𝑝2,…,𝑝𝑛)=𝐻α(𝑝1+ 𝑝2,𝑝3,…,𝑝𝑛)+(𝑝1+ 𝑝2) 𝐻α(𝑝1 𝑝1+ 𝑝2,𝑝2 𝑝1+ 𝑝2) , (2.4) which means that the entropy of N outcomes can be expressed in terms of the entropy of N − 1 outcomes plus the weighted entropy of the combined two outcomes; 5. Additivity. Considering two independent probability distributions, p = (𝑝1,𝑝2,…,𝑝𝑛) and q= (𝑞1,𝑞2,…,𝑞𝑛) , the joint probability distribution is denoted by p⋅q , the property Hα(p.q)= Hα(p)+ Hα(q) is called additivity . 2.1.4 - Renyi’s quadratic entropy Renyi’s quadratic entropy, α=2, gets a special attention by its practical applications. It is represented by the follow expression 𝐻2=−𝑙𝑜𝑔∑𝑝𝑖2𝑛𝑖=1 . (2.5) For continuous pdf distributions we have: 𝐻1(𝑥)=−1 1−2∫(𝑝(𝑥))2𝑑𝑥 . (2.6) This particular Renyi’s entropy has an interesting geometric interpretation. Assuming a discrete random variable with a probability density function in the form of P={(xi,pi(xi)), i=1,2…n }, that pdf can be seen as a singular point in a probability space of n dimensions, where the axis I correspond to the pi(xi) of each xi. That point, corresponding to the coordinates pi(xi), is over the hyperplane given by: ∑𝑝𝑖=1 𝑛 𝑖=1
8 State of art Therefore, the metric H2 is just the (negative) logarithm of the Euclidian distance between that point and the origin. 2.1.5 - Parzen window The Parzen window method is a powerful tool to interact with some concepts of information theory. Most of the times we assume a pdf as “real”, “exact” or “hidden but assumed real”; however, many times what we have is a discrete set of points which are taken as a representative sample of the assumed unknown pdf. To provide an analytic description for an estimate 𝑓 of the unknown pdf that would explain the present data in a sample of n points yi ϵ RM in a M-dimensional space, the Parzen windows method is quite useful. This method consists in apply a Kernel function centered on each sample point. If a Gaussian Kernel is used, the estimated pdf 𝑓 is obtained by the summation of the individual contributions of the Kernel applied on every point, as 𝑓𝑌(𝑧)=1𝑛∑𝐺(𝑧−𝑦𝑖,𝜎2𝐼) 𝑛1 , (2.7) where G(𝑧−𝑦𝑖,𝜎2𝐼) is the Gaussian Kernel function and 𝜎2𝐼 is the covariance matrix . It’s possible to prove that, when σ→ 0 and n→ ∞ the estimated pdf tends to the real pdf, 𝑓→ f [6]. It’s easy to understand that the Parzen window’s size, characterized by the standard deviation σ, conditions the shape of the estimate for the true pdf distribution. Thereby, a higher standard deviation provides a smoother form, while a low σ generates a more irregular pdf. This influence is possible to observe in the following charts, figures 2-1 to 2-3. Figure 2-1 - Graphic illustration of a sample of 8 points D={-1.70 -1.40 -0.98 -0.45 -0.30 0 0.20 0.30 0.60 1.10 1.40}, represented by Dirac pulses at each location. Figure 2-2 – pdf estimated by PW method for a set of 11 points, with σ=0.1. The blue line represents the sum of the Gaussians. The estimated pdf corresponds to the blue line divided by 11, in way that its integral is equal to 1.
Information Theoretic Learning 9 Figure 2-3 - pdf estimated by PW method for a set of 11 points, with σ=0.4. The blue line represents the sum of the Gaussians. The estimated pdf corresponds to the blue line divided by 11, in way that its integral is equal to 1. As it is possible to observe for a higher value of σ, figure 2-3 the shape of the estimated pdf is smoother than the one provided by a lower standard deviation, figure 2-2. In Chapter 3.1 -is possible to observe the application of the PW method to estimate an error distribution, based in the work [7]. The error distribution is need to the calculation of the RQE. However, only a discreet sample of points is known, the residual errors between the estimation and the measured values. Besides the estimation of the error distribution, the application of the PW method to the RQE concept also intends to introduce a parameter able to manage the expected error distribution. Thus, the proper definition of the parameter σ must allow the identification and natural elimination of large errors. 2.1.6 - Correntropy Correntropy is related with the probability of two random variables being similar. Their similitude is measured in the neighborhood of the combined space controlled by a kernel function. That is, the window width defined by the kernel acts like a visor controlling the observation window where the similitude is evaluated. Actually, it can be said that Correntropy uses the Parzen window method controlled by the kernel function. This window has the effect of mitigate or eliminate the effect of outliers. Correntropy is a proper measure for the similarity between two random variables X and Y, and can be expressed as 𝑉𝜎(𝑋,𝑌)=𝐸[𝐺(𝑋−𝑌,𝜎2𝐼)] , (2.8) where the 𝐺(𝑋−𝑌,𝜎2𝐼) is the Gaussian Kernel with variance 𝜎2 [8]. Truly speaking, the true pdf is unknown and only an estimative obtained from a data sample is possible. Thus, the estimated value of Correntropy can be calculated by 𝑉(𝑋,𝑌)=1𝑛∑𝐺(𝑥𝑖−𝑦𝑖,𝜎2𝐼)=1𝑛∑𝐺(𝜀𝑖,𝜎2𝐼) 𝑛𝑗=1 𝑛𝑗=1 , (2.9) where ε is the difference between the two variables. Through [9] there are some important Correntropy properties, among which: 1. Correntropy is symmetric V(X,Y)=V(Y,X); 2. Correntropy is positive and bounded: 0<V(X,Y) ≤1/𝜎√2𝜋. The maximum 1/𝜎√2𝜋 is achived only when X=Y; 3. Correntropy involves all the even moments of the random variable ε=X-Y.
10 State of art Applying the Taylor series development on the Gaussian Kernels it’s possible to rewrite the Correntropy expression as 𝑉(𝑋,𝑌)= 1𝑛∑1 𝜎√2𝜋 𝑛𝑗=1 ∑(−1)𝑘 (2𝜎2)𝑘𝐾!𝐸[(𝑥−𝑦)2𝑘] ∞ 𝑘=0 . (2.10) Correntropy considers all the even moments of the random variable ε. Also, its easily understood that the size σ of the kernel controls the impact of higher moments. When σ increases, the only relevant term becomes the second order moment, which is the expectation of the square difference between the variablesthe variance of the error. Thus, it’s fair to assume that Correntropy is a generalization of correlation, that is, it can be used as a measure of dependency between two random variables. The advantage is that, by considering higher order moments, Correntropy can express non-linear dependences between random variables, and not only linear dependences. The same does not happen with correlation, where only second order moments are considered. Thus, in correlation, only linear dependences between random variables are considered. 4. Correntropy 𝑉𝜎(𝑋,𝑌) is the value of the pdf 𝑓𝜀,𝜎(𝜀) evaluated at ε=0 , 𝑉𝜎(𝑋,𝑌)=𝑓𝜀,𝜎(0) , (2.11) which means that any process that maximizes the error frequency around 0 will increase the Correntropy value. 5. Correntropy corresponds to the probability of X=Y. Assuming the joint pdf of the two random variables, X and Y, denoted as fx,y, we have 𝑓𝜀(0)=∫𝑓𝑋,𝑌(𝑥,𝑥)𝑑𝑥=𝑃(𝑋=𝑌) . (2.12) In other words, the estimate of Correntropy 𝑉𝜎(𝑋,𝑌) with kernel size σ is the integral of estimated 𝑓𝑋,𝑌,𝜎(𝑥,𝑦) along the line x=y; Let’s focus now on the CIM, Correntropy Induced Metric. Assuming two vectors, X=(x1,x2,…,xn) and Y=(y1,y2,…,yn), CIM is represented by 𝐶𝐼𝑀(𝑋,𝑌)=(𝐺(0,𝜎2)−𝑉(𝑋,𝑌))1/2. (2.13) Assumed as a metric by showing the follow properties: 1. Non-negativity. CIM≥0 ; 2. Identity of indiscernibles. CIM(X,Y)=0 X ≡Y; 3. Symmetry. CIM(X,Y)= CIM(Y,X); 4. Triangular inequality. CIM(X,z) ≤ CIM(X,Y)+CIM(Y,Z) CIM has the following property: because of it non-homogeneous it doesn’t induce a norm in the space. However, it behaves as some other metrics depending on the distance to the origin. For instance, in the zero neighborhood the isolines are close to circles, such as the L2 metric (Euclidian space). far from that neighborhood, the isolines have the square shape, acting like the L1 metric (sum of coordinates). And finally, in more distant zones, the metric reaches a saturation level and acts like L0 metric, meaning that it’s indifferent to the distance. Such behavior is quite visible in figure 2-4.
Information Theoretic Learning 11 Figure 2-4Distance contours of CIM(X, 0) in 2-D sample space (kernel size is set to 1). Figure obtained from [9]. It’s really interesting to compare the Euclidian metric with the CIM. For that comparison, let’s assume two vectors X and Y with N components and each metrics expressions Euclidian metric 𝑑𝐿2(𝑋,𝑌)=(∑ (𝑥𝑖−𝑦𝑖)2 𝑛𝑗=1 )1/2. (2.14) CIM metric 𝑑𝐶𝐼𝑀(𝑋,𝑌)=(𝐺(0,𝜎2)−1𝑛∑𝐺(𝑥𝑖−𝑦𝑖,𝜎2𝐼) 𝑛𝑗=1 )1/2 . (2.15) In the Euclidean metric, the component that represent a higher influence in its value is the higher difference between X and Y. However, on the CIM, when that difference is too high relating to the parameter σ, it tends to lose its influence on the CIM result. This happens because the Gaussian of a high value (relating to the standard deviation) has a small result. Therefore, the influence of a gross error on the sum of the Gaussians won’t be notable. In this way, for the CIM, the distance between two vectors X and Y will be described by the components with close values, ignoring the components with higher differences. It’s important to get clearly that this event depends always on the kernel size, σ. Thus, if the difference between xi and yi is already high, it’s possible to change the value of the component xi without affect the value of the mutual CIM. That property is interesting when we don’t want to just measure the distance between two vectors, but, when applied to state estimator or other appliances, to exclude outliers without previous treatment. While the CIM has the advantage to ignore big differences between X and Y, that is, outliers, the L2 metric has exactly the opposite effect. The L2 result gets affected by that outlier and, as the difference between X and Y increases, the distance value increases also. 2.1.7 - Cauchy-Schwarz inequality Another interesting measure metric is the Cauchy-Schwarz Inequality. Assuming two probability density functions, p1(x) and p2(x), the distance of Cauchy-Schwarz can be given by
12 State of art 𝑑𝐶𝑆(𝑋,𝑌)=−𝑙𝑜𝑔 ∫𝑝1(𝑥)𝑝2(𝑥) +∞ −∞ √∫𝑝1(𝑥) +∞ −∞ ∫𝑝2(𝑥) +∞ −∞ . (2.16) This metric has the following properties: Symmetric. 𝑑𝐶𝑆(𝑋,𝑌)=𝑑𝐶𝑆(𝑌,𝑋); Is bounded. 0≤𝑑𝐶𝑆(𝑋,𝑌)≤1; Only achieves the value 0 when the two functions are exactly the same. Min Dcs=0 p1(x)=p2(x). Applying the expression (2.16) to discrete distributions p1 and p2, with a set of n points, it turns into 𝐷𝐶𝑆(𝑋,𝑌)=−𝑙𝑜𝑔 ∑𝑝1𝑖𝑝2𝑖 𝑛 𝑖=1 √∑𝑝1𝑖2 𝑛 𝑖=1 ∑𝑝2𝑖2 𝑛 𝑖=1 . (2.17) It is possible to observe that in (2.17) the nominator is represented by the calculation for the inner product of the vectors p1i and p2i, and in the denominator are the Euclidean norm of both vectors, as represented in (2.18). Thus, (2.17) can be simplified by (2.19) that is easily convert in (2.20): 𝐷𝐶𝑆(𝑋,𝑌)=−𝑙𝑜𝑔 ∑𝑝1𝑖𝑝2𝑖 𝑛 𝑖=1 √∑𝑝1𝑖 𝑛 𝑖=1 √∑𝑝12𝑖 𝑛 𝑖=1 ; (2.18) 𝐷𝐶𝑆(𝑋,𝑌)=−𝑙𝑜𝑔 𝑃1.𝑃2 ‖𝑃1‖‖𝑃2‖ ; (2.19) 𝐷𝐶𝑆(𝑋,𝑌)=−𝑙𝑜𝑔‖𝑃1‖‖𝑃2‖ 𝐶𝑜𝑠𝜃12 ‖𝑃1‖‖𝑃2‖=−log (𝐶𝑜𝑠𝜃12) . (2.20) By the interpretation of (2.20), one must say that Cauchy-Schwarz Inequality depends on the angle between the two vectors (distributions), not depending on the length of the vector. 2.2 -State estimation This section presents the main aspects related with power system state estimation, discussing fundamental concepts for the developed work. First, the formulation of the state estimation problem is presented, and then, different metrics already applied to its objective function are also exposed. Those different metrics are: WLS to the conventional state estimation process and maximum Correntropy criterion and M-estimators as innovative approaches. The modifications introduced in the state estimation problem caused by the inclusion of PMU is also explored. 2.2.1 - Global concept The actual electric power system is provided with a Supervisory Control And Data Acquisition system, SCADA, and a EMS, Energy Management System, having a set of monitoring and controlling functionalities. One of these important features is the State estimation [10, 11]. The measured values are usually corrupted with errors and a set of simultaneous measurements can be incompatible due the electrical laws.
State estimation 13 State Estimation, SE, aims to process acquired measurements in order to find an estimated valid operational point, with a low error regarding the measured values. In other words, its objective is to find a set of variables (states) that adjusts in the most adequate way to a set of network values (measurements) [12, 13]. The state estimation problem can be formulated mathematically using the following constrained optimization problem: min 𝐽(𝑥) , (2.21) Subject to 𝑐(𝑥)=0 , (2.22) 𝑓(𝑥)≤0 , (2.23) Where: x is the vector of state variables; J(x) is a scalar function of errors representing the objective function; c(x) equality-constraint vector; f(x) inequality-constraint vector. The function J(x) definition is based on a desirable criterion, and it measures the error between the measured values and the estimated values, being that error the parameter desirable to be minimized. The metric mostly used as the conventional state estimation objective function is denominated weighted least square error [14]. The application of new metrics as SE objective function has been theme of discussion. These new metrics includes information theory related concepts, as is was in [1], or m-estimators [15]. The state estimation nonlinear measurement model is formulated using a set of measurement equations. These measurement equations, h(x), relate a set of state variables, x, with a set of measured values, z, and a measurement noise, e, as: 𝑧=ℎ(𝑥)+𝑒, (2.24) where: h(x) is the non-linear m-dimensional vector function relating states to measurements; x is the n-dimensional state vector composed by bus voltage magnitude and phase angles; z is the m-dimensional measurement vector. These measurements consist in bus voltage, line power flows and power injection values; e is the m-dimensional error vector; with n<m in order to have an observable system. The difference between the measured and the estimated value represents the residual error value, represented by 𝑟=𝑧−𝑧=𝑧ℎ(𝑥) . (2.25) The true value of the measurement error, e, consists in the difference between the measured value and the unknown true value of that component and is calculated by 𝑒=𝑧−𝑧𝑟𝑒𝑎𝑙=𝑧−ℎ(𝑥𝑡𝑟𝑢𝑒). (2.26) The true value of the measured component or the state are theoretical and unknown, so it’s the error, e, what we know is an estimative to that error, the residual, r.
14 State of art 2.2.2 -– Weighted least square error objective function As the name of the method suggests, this method’s objective is to minimize the square error, assigning weights to each error between the measured and the estimated value. Thus, its objective function is 𝑚𝑖𝑛 𝐽(𝑥)=[𝑧−ℎ(𝑥)]𝑇𝑅−1[𝑧−ℎ(𝑥)] , (2.27) where R is the m-m-dimension weight matrix, corresponding to the variance of the error and it’s obtained by Cov(e) = E[e.eT]=diag{σ1,σ2, …σm} [16]. The expected accuracy of the sensor i is used to obtain the standard deviation σi of the corresponding measurement i. Therefore, the weight matrix, R, is constructed under the assumption that each sensor measurement has an expected error distribution. More information about the weight matrix can be found in [16]. To solve the objective function represented in (2.27), one needs to find the minimum value of J(x). Therefore, the first-order differential equality has to be observed: 𝑔(𝑥)=𝑑𝐽(𝑥) 𝑑𝑥 =−𝐻𝑇(𝑥)𝑅−1[𝑧−ℎ(𝑥)]=0 , (2.28) where H(x) is the Jacobian matrix with m*n dimension. 𝐻(𝑥)=[𝑑ℎ1(𝑥) 𝑑𝑥1…𝑑ℎ1(𝑥) 𝑑𝑥𝑛 ……… 𝑑ℎ𝑚(𝑥) 𝑑𝑥1…𝑑ℎ𝑚(𝑥) 𝑑𝑥𝑛] . (2.29) Expanding the non-linear function g(x) into its Taylor series around the state vector, xk, at iteration k, results 𝑔(𝑥)=𝑔(𝑥𝑘)+𝐺(𝑥𝑘)(𝑥−𝑥𝑘)+⋯=0 , (2.30) It’s possible to get into an iterative solution scheme known as The Gauss-Newton method if the higher order terms are neglected. The iterative function of this method is 𝑥𝑘+1=𝑥𝑘−[𝐺(𝑥𝑘)]−1.𝑔(𝑥𝑘) , (2.31) where: k is the iteration order xk is the state vector at iteration k. 𝐺(𝑥𝑘)=𝑑𝑔(𝑥𝑘) 𝑑𝑥 =𝐻𝑇(𝑥𝑘).𝑅−1.𝐻(𝑥𝑘) g(𝑥𝑘)=−𝐻𝑇(𝑥𝑘).𝑅−1(𝑧−ℎ(𝑥𝑘)) The matrix G(x), denominated gain matrix, is a sparse, positive definite and symmetric provided that the system is fully observable [16]. The matrix G(x) is typically not inverted, the state estimate 𝑥 is the obtained by the forward/back substitutions at each iteration k, by [𝐺(𝑥𝑘)]𝛥𝑥𝑘+1=−𝐻𝑇(𝑥𝑘).𝑅−1(𝑧−ℎ(𝑥𝑘)), (2.32) where 𝛥𝑥𝑘+1=𝑥𝑘+1−𝑥𝑘.
State estimation 15 The algorithm adopted for WLS state estimation is iterative and terminates when either a maximum number of iterations has been reached or the residual is lower than a defined limit. The iterative process is described in the next points [12, 16]. 1. initialize process setting index k=0; 2. Initialize the state vector xk, a good initialization is important to a quick convergence, typically corresponds to a flat voltage profile, where all bus voltages are assumed to be 1.0 p.u. and all in phase with each other; 3. Calculate the measurement Jacobian, H(xk); 4. Calculate the gain matrix, G(xk); 5. Calculate the measurement function , h(xk); 6. Calculate the correction term, Δxk= xkxk-1; 7. Test for convergence, |Δxk|<ε? 8. If it didn’t converge yet update xk+1=xk+ Δxk and go to step 3. If it already converged stop. The described procedure assumes that we have some little deviations on measured values and that other information as grid topology, or systems data, are known and well defined, what is not always truth. A set of additional functions are necessary to have a robust state estimator. Some of these functions are: topology processing; observability analysis; bad data processing. 2.2.3 - Correntropy as an alternative estimator As it was said before, WLS state estimators is the most common criterion and mostly used in the industries. However, new criteria based in theoretic information and M-estimators have been studied, in [1] and [15], respectively. Correntropy is a proper measure of the similarity between two random variables X and Y, and can be expressed as in [8] 𝑉𝜎(𝑋,𝑌)=𝐸[𝐺(𝑋−𝑌,𝜎2𝐼)] . (2.33) Applying this criterion to the state estimation process, one must see X as the measured value and Y as de estimated value, having the residuals R= X-Y. The number of measurements is limited and discreet, thus, also it is the number of estimation. In this way, Correntropy has to be calculated based on those discreet points, by 𝑉(𝑋,𝑌)=1𝑛∑𝐺(𝑥𝑖−𝑦𝑖,𝜎2𝐼)=1𝑛∑𝐺(𝜀𝑖,𝜎2𝐼) 𝑛𝑗=1 𝑛𝑗=1 . (2.34) It’s remarkable an affinity between the Correntropy and the Parzen window method. Actually, by adjusting the parameter σ, the window that defines the visible spectrum of errors is adjusted. Therefore, if εi is much bigger than σ, that error is seen as an outlier and the component i is ignored. The convenient selection of the PW size allows to naturally identify and ignore components whose measurement contains a gross error. This is the big difference between the Correntropy and the WLS. While on the WLS a gross error pulls the estimations of all the other components in the wrong direction, in the Correntropy the gross error is ignored, allowing all the other elements to conduct to the right estimation [1].
22 State of art 2.3.4 - Sensor data fusion process Sensory fusion has different process according to the area where it is inserted. In [27] there is a report about aircraft sensor data fusion. It describes the sensory fusion process in that field. The sensory fusion process is divided in two states. The first one is the tracking stage. In this stage all the sensors measurements referring to a platform (fighter aircraft engaged in an air mission) are brought together over time. Thus, it gives the most complete and accurate estimate of the platform’s position, motion and status that the measurements allow. The desirable would be to gather all the sensor measurements into a single tracking process in a timely and reliable way. Thus, a single track database would be produced and the sensory fusion would be complete at this stage. However, there are a lot of constraints on the systems and this doesn’t happen. In the end of this first stage several tracking process have produced them own database, each with a different perspective or image of the observed system, these different images must be combined in the next stage. The second stage, Track-to-Track Fusion, aims to fusion all these track databases into only one fused database. To fusion all the databases it’s necessary to perform the three next tasks: Align the tracks to the same spatial axis set and time; Deduce the number of targets that originated the acquired data and the platform from which each track arose; Form joint tracks and joint identity statements for targets reported by more than one source. In [28] is proposed a distribution state estimator based on the autoencoders for three-phase state estimator. The proposed method presents 3 main steps: Constructing the historical data set; Training the autoencoder (offline procedure); Building the distribution state estimator algorithm. Briefly speaking, only after define a set of historical data and train the autoencoder(as a neural network), a state estimation with the actual measurements is provided. In a certain way, can be said that there is a fusion across time, since historical data is influencing the estimation process for the actual measurements. However, all the actual measurements are gathered together, and the state estimation is provided without any kind of consideration or distinction about the sensor that provided the measurement. In other works, as [12-13], even existing distinct sensor types, the measured data is collected all together. The only distinction is through the weight matrix from the WLS estimator. There is no special attention or distinguish about the sensory system that provided the measure, and other parameters as sensory system’s reliability.
23 Chapter 3 Comparison between different state estimation optimization functions The mostly used criterion as objective function of power system state estimation is the weighted least square error [14]. The application of the WLS minimizes the variance of probability density function (pdf) of the residual error distribution. A good performance of the state estimator under the WLS criterion requires a Gaussian error distribution, what’s not always true. Sometimes the measurements contain gross errors. Such bad data affect the WLS solution defiling the estimative of components with healthy measurements. Once WLS cannot recognize and discard these large errors, it cannot be considered a robust estimator. Therefore, in this chapter, the traditional WLS will be compared with information theory related concepts. This chapter is dedicated to the application of those concepts as state estimation (SE) optimization function. Basically, the idea is to find a criterion able to maximize the information extracted from the available measurements. In other words, it aims at maximizing the information on the estimations by minimizing the informational content of the error distribution. If we achieve an error distribution with no information, the error pdf corresponds to a Dirac function centered in zero, meaning that there is only one solution for the state variables, which explain all measured values perfectly. That’s what would happen with a minimum entropy. In this way, a criterion based on theoretic information minimizes the informational error distribution, while the WLS minimizes the variance of the error distribution. The advantage of information theory related criteria is that, unlike the WLS that only consider the second order moments, it uses all the moments of the error distribution, providing more reliable results, as it was already explain in chapter 2.1.6 -. Correntropy, one of the alternative criteria, shows really interesting properties that can be easily applied in state estimation. MCC has the particularity of changing the perception of similarity according with the distance between the measured and estimated value. That is, depending on that distance, it behaves as different norms, namely: Euclidian-L2, absolute distance-L1 and then a saturation can be observed approaching L0. Thus, MCC gives more attention to close values and, then, gradually ignore the components as the residual increases. Therefore, MCC is able to ignore measurements corrupted with large errors. In this way, Correntropy is proposed as robust method to measure the similarity between measured and
24 Comparison between different state estimation optimization functions estimated values. Actually, this criterion was already tested in power system state estimation, in [1], promising really interesting results as a robust state estimator. In this chapter others criteria are discussed beyond the Correntropy, such as Renyi’s quadratic entropy and Cauchy-Schwarz distance. In order to evaluate the characteristics of the different metrics, two different scenarios are considered: Measurements without gross errors; Measurements with one gross error. For the scenario with gross error, two different cases were compared. A gross error in injected power in bus 3 and bus 9, respectively. 3.1 - Application of the alternative criteria in state estimation problem The first experiment in this thesis consists in comparing some different criteria as state estimation objective function. The evaluated criteria in this chapter are: Correntropy Induced Metric (also used as maximum Correntropy criterion) (3.1), Renyi’s quadratic entropy (3.4), Cauchy-Schwarz distance (3.10) and, as a comparison, the traditional least square error. The main objective of the state estimation is to minimize the distance between measured values and its estimations, being possible to measure that distance by different metrics. Two vectors are considered in the following expressions: the measurement vector Z with n components zi, with iϵ{1,2….n}, and the estimation vector 𝑍 with the n components 𝑧i , where iϵ{1,2….n}. The difference between this two vectors is the residual vector, that is, the deviation of the estimation relating to the measured value, 𝑟𝑖=𝑧𝑖−𝑧𝑖. Using the CIM, the objective is to minimize the distance between the estimation and the measured value, given by 𝑑𝐶𝐼𝑀(𝑍,𝑍)=(𝐺(0,𝜎2)−𝑉(𝑍,𝑍)1/2=(𝐺(0,𝜎2)−1𝑛∑𝐺(𝑧𝑖−𝑧𝑖,𝜎2𝐼) 𝑛𝑖=1 )1/2 . (3.1) However, it is possible to see that the expression (3.1) has a constant inside. The constant remains always the same, what matters is the variable term. Thus, the term 𝐺(0,𝜎2) can totally be ignored, just as the fraction term 1/n. Is still possible to make another simplification: To minimize the square root of a positive defined function is exactly the same as minimize that function. Thus, (3.1) can be simplified by 𝑑𝐶𝐼𝑀′(𝑍,𝑍)=−∑𝐺(𝑧−𝑧𝑖,𝜎2𝐼) 𝑛𝑗=1 =−∑1 𝜎√2𝜋𝑒−(𝑧𝑖−𝑧𝑖)2 2𝜎2 𝑛𝑗=1 . (3.2) Now is possible to convert the minimization problem into a maximization problem, just by convert the negative function into a positive. It visible that the term 1 𝜎√2𝜋 is the same in all the components inside the summation and it can also be ignored, leading to 𝑑𝐶𝐼𝑀(𝑍,𝑍)=∑𝑒−(𝑧𝑖−𝑧𝑖)2 2𝜎2 𝑛𝑗=1 . (3.3) In this way, applying Correntropy to the state estimation problem, the objective function is the maximization of (3.3). From now on, the Correntropy used as the SE criterion will be
Application of the alternative criteria in the state estimation problem 25 denoted MCC, Maximum Correntropy Criterion. In chapter 5, Correntropy will be used to measure the distance between two points, in that case, will be referred as CIM. The reason for this choice was to avoid confusion between the two simultaneous employments of this metric. Now, applying the Renyi’s quadratic entropy, the objective function is the minimization of the error distribution information. Once the true error is unknown, the objective is to minimize the residual error information by 𝐻2=−𝑙𝑜𝑔∫𝑓(𝑟)2𝑑𝑟 ∞ −∞ , (3.4) or the maximization of the symmetric function of (3.4), where f(r) represents the pdf of the errors. In effect, the f(r) true distribution is unknown, what is available is a set of discrete residual errors. The residual error distribution, f(r), can be estimated by the PW method from a discreet sample of error values. It consists in applying in each residual a Gaussian kernel as 𝑓(𝑟)=1𝑛∑𝐺(𝑟−𝑟𝑖,𝜎2𝐼) 𝑛𝑖=1 , (3.5) where n is the number of samples and σ is the size of the Parzen windows, given by the standard deviation of the adopted Gaussian Kernel. Besides the estimation of the error distribution, the application of the PW method to the RQE concept also intends to introduce a parameter able to manage the expected error distribution. Thus, the proper definition of the parameter σ must allow the identification and natural elimination of large errors. Applying the estimate 𝑓(𝑟) to (3.4) results 𝐻2=−𝑙𝑜𝑔∫1 𝑛2∑ ∑ 𝐺(𝑟−𝑟𝑖,𝜎2𝐼) 𝑛𝑗=1 𝐺(𝑟−𝑟𝑗,𝜎2𝐼) 𝑛𝑖=1 ∞ −∞ , (3.6) that is the same as 𝐻2=−𝑙𝑜𝑔1 𝑛2∑ ∑ ∫𝐺(𝑟−𝑟𝑖,𝜎2𝐼)𝐺(𝑟−𝑟𝑗,𝜎2𝐼) ∞ −∞ 𝑛𝑗=1 𝑛𝑖=1 . (3.7) In (3.7) we came across the integral of the product of Gaussians, which may be assimilated to the convolution of Gaussians, with a well-known result: The Gaussian of the difference [7]: 𝐻2=−𝑙𝑜𝑔1 𝑛2∑ ∑ 𝐺(𝑟𝑖−𝑟𝑗,𝜎2𝐼) 𝑛𝑗=1 𝑛𝑖=1 , (3.8) 𝐻2=−𝑙𝑜𝑔1 𝑛2∑ ∑ 1 𝜎√2𝜋𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑛𝑗=1 𝑛𝑖=1 . (3.9) In (3.9) we have the final form of RQE applied to state estimation problem. The constants outside the summation don’t interfere in the optimization process and the maximization of a logarithm is the same as maximize the term inside the logarithm. Therefore, it’s important to mention that minimize that expression is the same as maximize only the double sum of the exponential term. Thus, the objective function is the one presented in table 3-1. As explained in chapter 2.1.7 - the Cauchy-Schwarz inequality can be calculated by 𝐷𝐶𝑆(𝑍,𝑍)=−𝑙𝑜𝑔 ∑𝑧𝑖𝑧𝑖 𝑛𝑖=1 √∑𝑧𝑖2 𝑛𝑖=1 ∑𝑧𝑖2 𝑛𝑖=1 (3.10)
26 Comparison between different state estimation optimization functions and can be interpreted as a measure the collinearity between two vectors, by the simplification 𝐷𝐶𝑆(𝑍,𝑍)=−log (𝐶𝑜𝑠𝜃𝑧𝑧), (3.11) depending only on the angle between the two compared distributions. Once again, the exclusion of the signal minus from the expression (3.10) turns the minimization problem into a maximization problem. The distributions we want to compare, or to approach, are the measurement distribution and its estimation, Z and 𝑍, respectively. Summing up, the set of different objective functions that are going to be applied to the SE process can be consulted in table 3-1. Table 3-1Different objective functions. Weighted Least Square Error 𝐦𝐢𝐧 [𝒁−𝒁]𝑻𝑹−𝟏[𝒁−𝒁] Correntropy (MCC) 𝑚𝑎𝑥∑𝑒−(𝑧𝑖−𝑧𝑖)2 2𝜎2 𝑛 𝑗=1 Renyi’s Quadratic Entropy (RQE) 𝑚𝑎𝑥∑∑𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑛 𝑗=1 𝑛 𝑖=1 Cauchy-Schwarz inequality max ∑𝑧𝑖𝑧𝑖 𝑛𝑖=1 √∑𝑧𝑖2 𝑛𝑖=1 ∑𝑧𝑖2 𝑛𝑖=1 It’s important to mention that, the construction of the weight matrix (from WLS) requires information about the error covariance, for the multiple components. For a matter of simplification, in this work was assumed that all the errors had the same variance. Thus, the matrix R can be ignored. Actually this simplification will not interfere with the following studies. Assuming this simplification, the tested criterion was the least square error, or minimal mean squared error (MSE). Because in the also in MCC and RQE could be integrated a weight matrix, and at the moment it is not considered, the comparison with MSE approach is fairer than WLS. From now on, the minimal MSE will be used to detonate the traditional SE optimization function. 3.2 -Network and measurements for the DC model state estimator The different criteria were tested in the network presented in the annexes (A). This adaptation is based in a typical European medium voltage network adapted from [29]. In figure 3-1 is possible to see the network topology and, in table 3-2, the measurement set for the DC model. For more information, consult Annex A.
Network and measurements for the DC model state estimator 27 Figure 3-1 - Network used to compare the different criteria. Table 3-2 - Measurements used to compare the different criteria . Bus/line Power Flow result Measurement affected with noise Pinj (p.u.) 100 0.111326 0.111469 1 0.004130 0.004325 2 -0.080749 -0.080625 3 -0.056928 -0.056763 4 -0.004777 -0.004955 5 -0.004907 -0.005229 6 0.144789 0.144945 7 -0.029052 -0.028692 8 -0.005973 -0.006022 9 -0.042991 -0.042779 10 -0.031513 -0.031764 11 -0.003354 -0.003434 ϴ (rad) 2 -0.013365 -0.013383 5 -0.014911 -0.014913 8 -0.015388 -0.015403 Pij (p.u.) 3-4 -0.017071 -0.016890 6-7 0.088141 0.087880
28 Comparison between different state estimation optimization functions 3.3 - State Estimation without gross errors The state estimation problem was solved with the assistance of EPSO. The state variables were the voltage phases. The tested objective functions were the ones previously presented in the table 3-1. Table 3-3State estimations provided by each criterion. Component Z 𝒁𝐌𝐒𝐄 𝒁𝐌𝐂𝐂 𝒁𝐑𝐐𝐄 𝒁𝐂𝐒 Pinj (p.u.) 100 0.111469 0.111375 0.111375 0.111376 0.000341 1 0.004325 0.004241 0.004241 0.004240 0.000013 2 -0.080625 -0.080704 -0.080703 -0.080706 -0.000247 3 -0.056763 -0.056836 -0.056836 -0.056840 -0.000174 4 -0.004955 -0.004976 -0.004977 -0.004975 -0.000015 5 -0.005229 -0.005294 -0.005294 -0.005294 -0.000016 6 0.144945 0.144769 0.144769 0.144765 0.000443 7 -0.028692 -0.028560 -0.028559 -0.028556 -0.000087 8 -0.006022 -0.006016 -0.006016 -0.006016 -0.000018 9 -0.042779 -0.042777 -0.042777 -0.042776 -0.000131 10 -0.031764 -0.031773 -0.031774 -0.031772 -0.000097 11 -0.003434 -0.003448 -0.003449 -0.003446 -0.000011 ϴ (rad) 2 -0.013383 -0.013406 -0.013406 -0.013406 -0.000041 5 -0.014913 -0.014972 -0.014972 -0.014972 -0.000046 8 -0.015403 -0.015447 -0.015447 -0.015447 -0.000047 Pij (p.u.) 3-4 -0.016890 -0.016803 -0.016802 -0.016807 -0.000051 6-7 0.087880 0.088203 0.088203 0.088199 0.000270 Table 3-4Residual errors obtained from each criterion. Component 𝑹𝐌𝐒𝐄 𝑹𝑴𝑪𝑪 𝑹𝐑𝐐𝐄 𝑹𝐂𝐒 Pinj (p.u.) 100 -0.000094 -0.000094 -0.000093 -0.111128 1 -0.000084 -0.000083 -0.000085 -0.004312 2 -0.000079 -0.000079 -0.000081 0.080378 3 -0.000073 -0.000073 -0.000077 0.056589 4 -0.000021 -0.000021 -0.000019 0.004940 5 -0.000065 -0.000065 -0.000065 0.005213 6 -0.000176 -0.000176 -0.000180 -0.144502 7 0.000132 0.000133 0.000136 0.028604 8 0.000006 0.000006 0.000006 0.006004 9 0.000002 0.000002 0.000003 0.042648 10 -0.000010 -0.000011 -0.000008 0.031666 11 -0.000014 -0.000015 -0.000012 0.003423 ϴ (rad) 2 -0.000023 -0.000023 -0.000024 0.013342 5 -0.000059 -0.000059 -0.000059 0.014867 8 -0.000044 -0.000044 -0.000044 0.015356 Pij (p.u.) 3-4 0.000087 0.000088 0.000083 0.016839 6-7 0.000323 0.000323 0.000319 -0.087610
State estimation without gross errors 29 The estimated values obtained using the different criteria can be consulted in table 3-3 and its residual error in table 3-4. For the MCC and RQE criteria it was utilized an PW size 5. For a better illustration of the results, in the following charts is possible to see the residuals distribution obtained though the PW method. That distribution was obtained by placing a Gaussian kernel, with a standard deviation of 0.01, in each of the residual values. Summing all those terms and divide by the number of summed terms was obtained the residual error pdf’s presented in figures 3-2 and 3-3. Figure 3-2 – Residual errors distribution for MSE, on the left , and RQE , on the right. Figure 3-3 – Residual errors distribution for MCC, on the left , and CS , on the right. 3.3.1 - Results analysis The first big conclusion is about the Cauchy-Schwarz distance. The error of the estimation provided by this metric are at the same order as the measured values, revealing to be much higher than the errors obtained with the other tested metrics. It is also possible to observe that it’s distribution, figure 3-3, doesn’t follow a Gaussian distribution as expected and as the other metrics follow. In the Cauchy-Schwarz is possible to see peaks and a distorted distribution caused by the bad estimation. In fact, as explained before, this metric measures the distance between two distributions according with them collinearity. Thus, the application of CS metric in a state estimation problem leads to a trouble situation: it will try to find an estimation vector collinear with the measurement vector, which can lead to a solution far from the optimal one. That is what happened in the previous example. The found estimation vector is collinear with the measurement vector; However, the “optimal” point found has a high error comparing to the
30 Comparison between different state estimation optimization functions residual errors obtained by the other metrics. Despite the other metrics provided estimation vectors not collinear with measurement vector, the estimations were closer to the measured values. For now on, the Cauchy-Schwarz metric will be excluded once it was proved that collinearity is not a good criterion to the SE problem. When the measurements do not contain gross errors, Correntropy and Renyi’s quadratic entropy criterions conduct to an optimal estimation, near to the measurement vector. Also, a similarity between Correntropy and MSE results is notable. In fact, when there are no gross errors the Correntropy metric presents the same result as (or really close to) MSE, once MSE minimizes (zi-𝑧𝑖 )2 and, in a certain way the Correntropy minimizes also that component. Maximize the exponential of the negative square error (divided by a constant) is the same as minimizing the square error. The big difference is that the Correntropy has a division term affecting the square error, and to make this effect even more notable, also an exponential. These additional terms in Correntropy have effect in situations with gross errors, as we will see in the next section. 3.4 - State estimation with gross errors One problem with the MSE state estimator is its reaction to bad data. When facing outliers, the component that contains a big error infects also the other components estimation, deallocating all the estimations in its direction. For that reason, in order to avoid bad data and to perform a proper estimation, the Least Square Error state estimator requires a pretreatment of the acquired data. Therefore, it would be interested to evaluate the other metrics performance under the presence of large errors. An important parameter that requires special attention is the Parzen Window size for the identification of outliers. The explore the behavior of the different criteria, it was used the measurements presented in 3.2 with a gross error in Buss 9 and other case with a gross error in bus 3. The reason to compare two different cases with large errors in two different buses is to assess its influence in adjacent line power flows estimation. In the first study case a gross error was introduced in a bus without measurements power flow in adjacent lines, bus 9, and, in the second study case, a large error was introduced in bus 3, where an adjacent line power flow is available. From the analysis of tables 3-5 and 3-6, it has become evident that, facing a gross error, the MSE results are corrupted and all the estimations are moving towards that error. The same happens with Correntropy and Renyi’s quadratic Entropy when the Parzen Window size is high comparing to the error magnitude (in Correntropy) or to error distribution variance (in Renyi’s entropy), case σ=5. In Correntropy, when σ=5, the Parzen Window size is too large and the error introduced in 9 “fits” inside it. In other words, because the introduced error is too small comparing to the PW, it is not seen as a gross error but as an available measurement, as all the others. When the Parzen Window size decreases, σ=0.1, the error introduced in bus 9 is seen as an outlier. In fact, the value of the error is intensified dividing by 0.1, then applying the negative exponential, results a really small component compared to the other components, calculated from smaller errors. Thus, because the influence of the measured value 𝑃𝑖𝑛𝑗9 doesn’t have an influence to the summation of all the terms of the Correntropy expression, that component is ignored. This is the reason why, for σ=0.1, all the errors in the other components are really
State estimation with gross errors 31 small. The only big error (between the estimation and the measured value) is the one in injected power in bus 9. In this way, that measurement was ignored and the Correntropy provided a proper estimation ignoring the outlier. Table 3-5State estimations provided by each criteria, with different PW sizes. Gross error in Pinj9. Component Z 𝒁MSE 𝒁MCC (σ =5) 𝒁RQE (σ =5) 𝒁MCC (σ =0.1) 𝒁RQE (σ =0.1) Pinj (p.u.) 100 0.111469 0.044074 0.044051 0.062641 0.111374 0.111352 1 0.004325 -0.112388 -0.112369 -0.124765 0.004239 0.004198 2 -0.080625 -0.221311 -0.221282 -0.248725 -0.080705 -0.080748 3 -0.056763 -0.234357 -0.234333 -0.277623 -0.056837 -0.056889 4 -0.004955 -0.190753 -0.190738 -0.171085 -0.004978 -0.005016 5 -0.005229 -0.174114 -0.174084 -0.171543 -0.005296 -0.005336 6 0.144945 0.018992 0.019035 -0.018493 0.144768 0.144726 7 -0.028692 -0.285607 -0.285730 -0.238693 -0.028562 -0.028607 8 -0.006022 -0.214281 -0.214295 -0.212761 -0.006018 -0.006063 9 2.000000 1.793422 1.793411 1.797805 -0.042760 -0.042312 10 -0.031764 -0.227952 -0.227948 -0.215925 -0.031775 -0.031816 11 -0.003434 -0.195726 -0.195717 -0.180833 -0.003450 -0.003489 ϴ (rad) 2 -0.013383 -0.000869 -0.000867 -0.002595 -0.013406 -0.013402 5 -0.014913 0.016229 0.016229 0.016201 -0.014972 -0.014961 8 -0.015403 0.025673 0.025673 0.025623 -0.015447 -0.015434 Pij (p.u.) 3-4 -0.016890 -0.034213 -0.034199 -0.077728 -0.016803 -0.016830 6-7 0.087880 -0.048806 -0.048718 -0.088328 0.088202 0.088155 Table 3-6Residual errors for in the different estimations, when Pinj9 has a gross error. Component RMSE RMCC (σ=5) RRQE (σ=5) RMCC (σ=0.1) RRQE (σ=5) Pinj (p.u.) 100 -0.067395 -0.067418 -0.048828 -0.000095 -0.000117 1 -0.116712 -0.116693 -0.129089 -0.000085 -0.000127 2 -0.140686 -0.140658 -0.168100 -0.000080 -0.000124 3 -0.177594 -0.177570 -0.220860 -0.000074 -0.000126 4 -0.185797 -0.185783 -0.166129 -0.000023 -0.000061 5 -0.168884 -0.168855 -0.166314 -0.000067 -0.000106 6 -0.125953 -0.125910 -0.163438 -0.000177 -0.000219 7 -0.256915 -0.257038 -0.210001 0.000130 0.000085 8 -0.208259 -0.208273 -0.206739 0.000004 -0.000041 9 -0.206578 -0.206589 -0.202195 -2.042760 -2.042312 10 -0.196188 -0.196185 -0.184161 -0.000011 -0.000052 11 -0.192292 -0.192283 -0.177399 -0.000016 -0.000055 ϴ (rad) 2 0.012513 0.012516 0.010787 -0.000023 -0.000019 5 0.031142 0.031143 0.031114 -0.000059 -0.000048 8 0.041076 0.041076 0.041026 -0.000043 -0.000030 Pij (p.u.) 3-4 -0.017323 -0.017309 -0.060838 0.000087 0.000060 6-7 -0.136686 -0.136598 -0.176208 0.000322 0.000275
38 Sensory fusion considering two sensor systems Table 4-1 - Measurements origin Bus/line Sensor Class Measured value Pinj (p.u.) 100 Conventional (1) 0.111469 1 Conventional (1) 0.004325 2 PMU (2) -0.080625 3 Conventional (1) -0.056763 4 Conventional (1) -0.004955 5 PMU (2) -0.005229 6 Conventional (1) 0.144945 7 Conventional (1) -0.028692 8 PMU (2) -0.006022 9 Conventional (1) -0.042779 10 Conventional (1) -0.031764 11 Conventional (1) -0.000044 ϴ (rad) 2 PMU (2) -0.013383 5 PMU (2) -0.014913 8 PMU (2) -0.015403 Pij (p.u.) 3-4 Conventional (1) -0.016290 6-7 Conventional (1) 0.087880 The DC model will be used considering that there are PMU sensors on buses 2, 5 and 8. PMU measures voltage and current phasors. For the simplifications of DC model is considered that the voltage magnitude is 1 p.u. causing the injected power to assume the same value as the injected current. Thus, for the DC model will be consider that the PMU measurements are voltages phases and power injections (once for this model the power injection value is the same as current injection). In other to obtain graphics with a better representation of the curve we want to expose, the measurements used were the ones already presented in chapter 3, but with a little deviation in one of the measurements, and it can be found in table 4-1. 4.2 - Trust variation among the different classes of sensors The main idea of this section is to assign different weights to the fitness evaluations of each set of estimations, according to the sensory system that provided the measurement. For that, the measurement set was divided in two groups. Each group estimation errors were evaluated according to an evaluation metric, leading to two fitness evaluations. This evaluation metric is correspondent to the objective function of the state estimation. The total evaluation is obtained assigning an weight to each one of those evaluations as 𝑓𝑖𝑡=𝛼∗𝑓𝑖𝑡1+𝛽∗𝑓𝑖𝑡2 , (4.1)
Trust variation among the different classes of sensors 39 where: fit1 and fit2 corresponds to the fitness evaluation of residual errors, from group one and two, respectively. It can be calculated trough WLS, Correntropy or Renyi’s quadratic Entropy; α and β are the weight assign to the evaluation of each set of residual errors, where β=α-1; The optimal values assign to α and β depend on multiple factors. Not only it depends on the confidence inherent to each sensory system, but also on the observability that the quantity of sensors and its position in the network give about the system. In order to inspect the variation of α and β influence, on the estimations, the state estimation was solved with assistance of EPSO algorithm. This time, the objective function is the maximization of (4.1) (expect using MSE, were was computed a minimization). The terms fit1 and fit2 were calculated according to each metric: least square, Correntropy and Renyi’s quadratic entropy. For the cases where the criterion of the state estimation is the MSE or MCC it’s easy to decide how will be computed either fit1 and fit2, because each term are related with only one measurement and its estimation. The same doesn’t happen with Renyi’s Quadratic Entropy. For this measure each term is related with two residuals of two different measurements. As explain in chapter 3 the expression utilized to RQE is 𝐻2=∑ ∑ 𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑛𝑗=1 𝑛𝑖=1 . (4.2) In this case, there are two possible ways to compute the fitness functions. The simplest way is to calculate each fitness function fit1 and fit2, relating each residual only with residuals from the same sensory system, respectively, 𝑓𝑖𝑡1=∑ ∑ 𝑒−(𝜀𝑖−𝜀𝑗)2 2𝜎2 𝑗𝜖𝑛1𝑖𝜖𝑛1 and (4.3) 𝑓𝑖𝑡2=∑ ∑ 𝑒−(𝜀𝑖−𝜀𝑗)2 2𝜎2 𝑗𝜖𝑛2𝑖𝜖𝑛2 . (4.4) The other way is to relate each residual with all the other residuals, independent of its origin, this option allows the sum of fit1 with fit2 to be equal with the fitness value of all the group together, through the following expressions 𝑓𝑖𝑡1=∑ ∑ 𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑗𝜖𝑛1⋀𝑛2𝑖𝜖𝑛1 and (4.5) 𝑓𝑖𝑡2=∑ ∑ 𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑗𝜖𝑛1⋀𝑛2𝑖𝜖𝑛2 . (4.6) In the last situation, what will introduce a difference between the optimal point found by the fusion process or the simple process (estimation considering only one set with all the measurements) is the values attributed to the parameters α and β. For the following tests, this cases will be denominated as Renyi’s quadratic entropy 1 and 2, respectively.
40 Sensory fusion considering two sensor systems 4.2.1 - Influence of α and β in both fitness functions In order to ascertain the variation of parameters α and β, tests were performed for α from 0 to 1 in intervals of 0.01, with β=α-1. In the following graphics. 4.1 to 4.4, is possible to see the Pareto-front of non-dominated points, found by EPSO, for fit1 and fit2, obtained by each of the different criteria. These graphics illustrate the variation in both fitness functions, fit1 and fit2, according to the trust attributed to each group of measurements. In this way, when α assumes the maximum value, 1, only fit1 has influence in the final fitness function. Thus, it aims to maximize fit1 without paying any attention to the fit2 value, in fact, β assumes the value 0. Hence, it’s easy to understand that fit1 reaches its maximum when α is equal to 1. The same stands to fit2, it reaches its maximum when β=1 and α=0, once only fit2 has influence in the final fitness function. There is a particular situation that needs special attention. When using the Mean Squared Error and Correntropy, figures 4-1 and 4-2, the fitness function is evaluated depending on the residual error of each component estimation separately. Therefore, when β=1 and α=0, the state estimator will perform estimations for all the 17 components bases only in 6 measurements, 3 of injected power and 3 of voltage phases. Obviously the system has no observability and, despite fit2 reaches its maximum, fit1 will decrease drastically. This happens because all the measurements will be estimated having in consideration only few measurements. Hence, the estimation for the group two will adjust perfectly with a minimum error, but the error in the estimations of group one will be high. This is the reason why fit1 decreases so much. Taking this fact in consideration, both figures 4-1 and 4-2 show a graphic where the point (α, β) = (0,1) appear, in the left, and other graphic without that point, in the right. As it its visible in figures 4-1 and 4-2, on the right, for α is different than 0, is possible to observe gradual increase of fit1 and, at the same time, a decrease of fit2 as α increases. for RQE as SE criterion (case 1), figure 4-3 on the left, is possible to observe two isolated points. In fact, it considers the relation of all the residuals within each group of sensors separately. For the comparison of image 4-3 with 4-1 and 4-2, can be concluded that RQE (case 1) is more susceptible to lose the system observability. In the left is possible to see those isolated points, while in the right these points ((α, β) = (0,1) and (α, β) = (1,0)) were excluded allowing a better perspective of the evolution of the fitness evaluations of each sensory system varying α and β. Figure 4-1 - Fitness 1 and 2 variation with α, assuming mean squared error. Representation of the simple fusion by the red dot
Trust variation among the different classes of sensors 41 Figure 4-2 - Fitness 1 and 2 variation with α, assuming Correntropy. Representation of the simple fusion by the red dot Figure 4-3 - Fitness 1 and 2 variation with α, assuming RQE case 1. Representation of the simple fusion by the red dot Figure 4-4Fitness 1 and 2 variation with α, assuming RQE case 2.Representation of the simple fusion by the red dot. The same distant point (α, β) = (0, 1) for Renyi’s quadratic entropy (case 2) is not visible and it’s easy to justify why. In Renyi’s quadratic entropy (case 2) the fit evaluation doesn’t depend on the error of each measurement separately, but in its dimension comparing to all the other errors, as is can be seen by the equation (4.5) and (4.6). In figures 4-1 to 4-4 is also possible to see the point equivalent to the simple fusionthe red dot. That is, the estimation considering only one set with all the measurements, where all
42 Sensory fusion considering two sensor systems the measurements have the same weight to the final fitness function. As expected, the optimal point obtained with the simple fusion is situated among the curve of non-dominated points. 4.2.2 - Residual analysis In the figures 4-5 to 4-11 is possible to observe the evolution of the residual errors, from the diverse estimations, in order to α. As explained before, for MSE and MCC criterions, when α=0 only the measurements of class two are considered and that number of measurements are insufficient to guarantee the system observability. In figures 4-1 and 4-2 is possible to observe an isolated point for α=0, now the same point is also observed in 4-5 and 4-6 represented by high errors in comparison with the errors for other values of α. In fact, these errors on estimations of group one is at least 100 times bigger in comparison with the errors for the other values of α. For a better observation of the evolution of the error, through the variation of α, the point α=0 was ignored. It is possible to observe a new graphic with representation of the evolution of the errors in 4-5(on the right side) to MCC and 4-6(on the right side) for MSE. Also in figure 4-3 is possible to observe isolated points for the REQ (case 1). In figures 4-7 and 4-9(left side) is also possible to see these same points represented now, not by the deviation in the fitness functions, but by the deviation in the errors. The point with lower value for the fitness evaluation of the group one is represented by the high error presented in figure 4-7, for α=0. The point represented in figure 4-3 with lower value for the fitness evaluation of the group 2 is now represented in figure 4-12 for α=1, where is possible to observe errors extremely higher than the others obtain for different values of α. Figure 4-5Residual errors (p.u.) for the group one using Correntropy, varying α. On the right side the point α=0 was ignored. Figure 4-6Residual errors (p.u.) for the group one using Mean Squared Error, varying α. On the right side the point α=0 was ignored.
Trust variation among the different classes of sensors 43 Figure 4-7 - Residual errors (p.u.) for the group one using Renyi’s Quadratic entropy(case one), varying α. On the right side the point α=0 was ignored. . Figure 4-8 – Residual errors (p.u.) for the group one using Renyi's Quadratic Entropy (case 2), varying α. For the second group of sensors only in the RQE case 1 were observed a case of an isolated point, with a residual error much higher than the others for different values of α. That situation is represented in figure 4-9. In the same figure but in the right side is represented the variation of the errors of the estimations for the group 2 in function of α, but ignoring the point α=1, or β=0. Figure 4-9Residual errors (p.u.) for the group two, using Renyi’s Quadratic entropy(case one), varying α. On the right side the point α=1 was ignored.
44 Sensory fusion considering two sensor systems Figure 4-10-Residual errors (p.u.) for the group two, using MCC on the left and MSE on the right, varying α. Figure 4-11-Residual errors (p.u.) for the group two, using Renyi’s Quadratic Entropy (case 2), varying α Is important to mention that, for MCC and MSE, despite what happens when α=0 to the group one, the same doesn’t happen to the group two when β=0. The errors in the estimations of group two increase as β decreases (α increases), but there is not an isolated point with an superior dimension error, as it happens to the estimations of the group one when α=0. That’s easily explain, because the group one has much more measurements and, even assuming that the measurements of the group 2 are in fault (that’s what happen when it’s assumed that β=0, to MCC and MSE) the system is still observable. Hence, despite the growing of estimation errors, it doesn’t grow so much as in the previous case, as it is visible by the point α=1 in figures 4-10. As it was expected, in a general mode, the estimation errors of the group one, figures 4-5 to 4.8, decreases as α increases. Which is obvious, α determines the weight of that components to the objective function, once it grows the trust in that measurements increases and the estimations tends to approach their values. The same can be said to the measurements of the group two, but when α decreases, that is, β increases, figures 4-9 to 4-11. Using the minimization of MSE or Correntropy maximization as SE objective function, it’s possible to see a big slope of the residuals through the parameter α. However, when using the Renyi’s Quadratic Entropy (case 2) that slope isn’t that much accentuated, figure 4-11. In fact, the computation of the Renyi’s Quadratic Entropy (case 2) was applying the expression (4.5) and (4.6). Thus, the fit1 is given by the sum of all the components from group one interacting with all the other errors, even with components from the group two. In this way, even when α=0 the system doesn’t loss its observability, because the components from the group one are still taken in consideration, even that the relations with the errors inherent to the group two are prevalent.
Trust variation among the different classes of sensors 45 In RQE (case 1) the fitness function is calculated considering only residuals referring to the same group. As it correlates all the residuals of the same group, to calculate each fitness evaluation, it promotes much more interaction by the errors obtained by each sensory system. In this way, this criterion is more susceptible to lose the system observability, while in case 2, measurements provided by both systems interact with each other maintaining the system observability. 4.3 - Chapter conclusions The definition of α and β parameters conducts the state estimation either closer to a measurement set or to the other, depending on the trust attributed to each sensory system. It’s important to mention that, excluding the cases where the system loses its observability, each one of the non-dominated points, present in figure 4-1 to 4-4, are possible. The question is: which point should be assumed as the real state? One proposal is to find stochastic weights for the parameters α and β. This can be done in a future work, requiring lot of research and tests. Such investigation must evaluate the probability of each sensory system to conduct the state estimation to the real one, depending on: sensor type, quantity, location, synchronism and other possible factors. As that information is not yet available, another way to find an optimal point has to be explored. The following chapter is dedicated to that search, in order to find an optimal fusion point. The sensory fusion method presented in this chapter combines conventional sensors with PMU. Both types of sensors measure different properties, but their measurements complement each other, providing a better perception of the power system state. Thus, this sensory fusion method can be classified as complementary, by the Durrant-Whyte definition, and fusion across attributes, by Boudjemaa. As it is possible to observe in the previous section, the variation of each sensory system weight originates a Pareto-front, resulting in a decision problem. Through a proper definition about these weights, the decision of the optimal fusion state is done automatically. That is, the fusion of the input features generates an output decision. Thus, the proposed stochastic fusion is, through Dasarathy, a fusion with input/output characteristics of the type feature input/ decision output, FeI-DeO.
46 Sensory fusion considering two sensor systems
47 Chapter 5 Selecting the optimal fusion point In chapter 4, was concluded that the variation of each sensory system weight, on the SE objective function, originates a Pareto-front with the same form as 5.1. The extremal points conduct to situations where only one class of the sensors was considered, and it can be seen as less reliable situation since the system can losses its observability. However, approaching to a situation of α close to β and close to 0.5 both fit1 and fit2 decrease but, all the measurements are used to perform a better result. That situation is somewhere represented in the graphic where both fit1 and fit2 assume high values. Figure 5-1Typical Pareto-front obtained by the multiple fusion points, depending on α and β, in black. Ideal point (maxfit1, maxfit2) represented by the grey point. There isn’t just a single way to decide which point represents better the true state of the power system. Still, one simple way to choose an optimal point can be select the one closest to the ideal point given by (Maxfit1, Maxfit2). The values for Maxfit1 and Maxfit2 are easy to obtain with no need to draw all the curve. In fact, when only one class of sensor is considered, α=1 or β=1, only its fitness function is maximized, ignoring the measurements provided by the other sensor system. Hence, when α is equal to 1 the maximum value of fit1 is reached, and when β is equal to 1 the maximum point of fit2 is achieved. Another issue is how to measure that distance. Which is the metric that leads to a minimal error? In order to choose a metric that better approaches the optimal to the ideal point, some tests were provided, evaluating the effect of different metrics on that step of the sensory fusion. The explored metrics, applied to the search of the optimal point were: -2 0 2 4 6 0 1 2 3 4 5 6 7 8 fit1 fit2 Ideal point Fusion points
54 Selecting the optimal fusion point Figure 5-8 - Residual errors in buses power injections and line power flows, using CIM as the fusion metric (with different Parzen window’s size )and MCC as evaluation metric. Figure 5-9 - Residual errors in voltage phases, using CIM as the fusion metric (with different Parzen window’s size )and MCC as evaluation metric. Figure 5-10 - Residual errors in buses power injections and line power flows, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 1) as evaluation metric. 0 0.0005 0.001 0.0015 0.002 0.0025 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric CIM, Estimation : MCC σ=100 σ=20 σ=5 σ=1 σ=0.1 σ=0.05 σ=0.01 0 0.00005 0.0001 0.00015 0.0002 0.00025 0.0003 0.00035 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric CIM, Estimation : MCC σ=100 σ=20 σ=5 σ=1 σ=0.1 σ=0.05 σ=0.01 0 0.0002 0.0004 0.0006 0.0008 0.001 0.0012 0.0014 0.0016 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric CIM, Estimation : RQE(case1) σ=100 σ=20 σ=5 σ=1 σ=0.1 σ=0.05
Sensory fusion optimal point - Measurements without gross errors 55 Figure 5-11 - Residual errors in voltage phases, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 1) as evaluation metric. Figure 5-12 - Residual errors in buses power injections and line power flows, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 2) as evaluation metric. Figure 5-13 - Residual errors in voltage phases, using CIM as the fusion metric (with different Parzen window’s size )and RQE (case 2) as evaluation metric. 0 0.00002 0.00004 0.00006 0.00008 0.0001 0.00012 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric CIM, Estimation : RQE(case1) σ=100 σ=20 σ=5 σ=1 σ=0.1 σ=0.05 0 0.0001 0.0002 0.0003 0.0004 0.0005 0.0006 0.0007 0.0008 0.0009 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric MCC, Estimation : RQE(case2) σ=100 σ=20 σ=5 σ=1 σ=0.1 σ=0.05 σ=0.01 0 0.00002 0.00004 0.00006 0.00008 0.0001 0.00012 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric CIM, Estimation : RQE(case2) σ=100 σ=20 σ=5 σ=1 σ=0.1 σ=0.05 σ=0.01
56 Selecting the optimal fusion point Despite PW with high size can lead to bad results, it is necessary to pay attention when the Parzen Window’s size is too small. In fact, when it is too low it starts to ignored even the optimal point. This was what happened for Parzen Window’s size 0.0001, where, for all the evaluation metrics the estimation errors reached really high values, figure 5-14. What happened in this case was that the fitness evaluation ranges of group one and group two are not the same, thus, by decreasing the Parzen Windows’s size it starts to ignore the measurements of one group of the sensors and consider only the other, where the variation of the fitness evaluation is smaller. That’s why some of the elements present estimations with really big errors and others with really small, for σ=0.0001. Figure 5-14 - Residual errors in buses power injections and line power flows using CIM as the fusion metric, with Parzen size 0.0001. Figure 5-15 - Residual errors in voltage phases using CIM as the fusion metric, with Parzen size 0.0001. Considering the errors represented in figures 5-6 to 5-13 can be assumed that the Parzen Window’s size 1 generate good results, to this study case. Therefore, in the following section, 5.3, the PW size applied to CIM as fusion metric will be 1. 0 2 4 6 8 10 12 14 16 18 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric: CIM Parzen size =0.0001 MSE MCC RQE2 RQE1 0 0.000002 0.000004 0.000006 0.000008 0.00001 0.000012 0.000014 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric: CIM Parzen size =0.0001 MSE MCC RQE2 RQE1
Sensory fusion optimal point - Measurements without gross errors 57 5.2.1 - Results analysis / relevant conclusions. Comparing the results obtained using the absolute and Euclidean norms as fusion metrics, the most relevant conclusion is the behavior of the Renyi’s quadratic entropy (case 2). For most of the components, its residual errors are higher than the errors obtained by the other evaluation metrics. The RQE (case 2) estimation errors are also higher than the noise introduced to the original values (measurements without noise). Still analyzing the fusion metrics L1 and L2, the results obtained by the evaluation metrics MSE and MCC are practically the same, as expected, because there are no outliers and thus both criteria conduct to the same state estimation. Between MSE and MCC metrics and RQE (case 1) there are no obvious benefits, by the comparison of the results. Even that RQE1 leads to a different state, its residual errors are within the noise introduced to the original values. Thus, it can be considered that the state estimation provided by RQE1 is a reliable state. As it was already mentioned, using the RQE (case 2), the residual errors reach much higher values than the other evaluation metrics and also higher than the noise introduced to the original values of the measurements (without noise). Also, in RQE case 2, the sensory fusion is not done considering the two groups of measurements separately, in fact, even the evaluation of each system considers interaction with errors from the other group. For that reason, and also considering that the results were worse than the ones obtained by other metrics, the RQE case 2 can be abandoned for these studies. The CIM as fusion metric is a more delicate situation. Actually, the Parzen Window will interfere with the result, and that PW will consider the deviance between the ideal point and the optimal point found. That distance is unknown and also can have really different ranges to each of its components (variation of fit1 and fit2 can assume really different values), so it’s risky to use that metric in the fusion step. 5.3 - Sensory fusion optimal point - Measurements with gross errors In the previous chapter, was already observed the influence of the different evaluation metrics as SE objective function in the presence of gross errors. Now, it was included another step, the fusion step, where the measurements are divided in two groups provided from two different sensory systems. This sub-chapter aims to evaluate the influence the fusion step has in the treatment of the outliers. For this study were used the same measurements as in the previous section. However, now the measurement of the injected power in bus 3 contains a large error, table 5-4. As the measurements were modified also the ideal point was deallocated, in this way, the new ideal point to each evaluation metric can be consulted in table 5-8, considering the PW applied to these metrics with size 2. Table 5-8 - Ideal point, calculated with σ=2, to the measurements affected with gross error . maxfit1 maxfit2 MSE -0.682033996077471 -0.165376488179705 MCC 10.915560350232500 5.979774529346460 RQE 1 119.805597386184000 35.489865315248300 RQE 2 184.751933517149000 101.200230478727000
58 Selecting the optimal fusion point To the fusion metric CIM, the Parzen Window’s size was assumed to be 1, because of what was already tested in the previous section. However, also the metrics MCC and RQE, used as evaluation metrics, have the PW concept. Thus, that metrics were tested assuming two values for the PW size, 2 and 0.1. One large window, that allow gross errors to be assumed as part of the measurements, and a thinner window, that ignore the gross errors presented in the measurements. It's important to highlight that the variation of the Parzen Window in the fusion step (for the CIM) doesn’t affect at all the ideal point, once it is not even used to its obtainment. However, varying the PW of the evaluation metrics will change the ideal point, because the fitness values are calculated considering a fixed PW size. Thus, the new ideal points are presented in table 5-9. The MSE doesn’t depend on the parameter σ, for that reason it remains the same. Table 5-9 - Ideal point calculated with σ=0.1 to the measurements affected with gross error . maxfit1 maxfit2 MSE -0.682033996077471 -0.165376488179705 MCC 9.999992456509830 5.999999141260640 RQE 1 100.999901583335000 35.999983319472600 RQE 2 160.999799009255000 95.999933402185100 In the following figures is possible to see four graphics representing the absolute value of the residual errors. Two of them represent the errors on power injections and lines power flows: one with all the PW sizes and the measurement affected with the gross error, and other ignoring the results for the component affected with a gross error and including only the PW sizes 0.01. The same stands for the graphics of voltage phases. Figure 5-16 - Residual errors in buses power injections and lines power flows using L1 as the fusion metric, with σ= 2 and σ=0.1. 0.0 0.5 1.0 1.5 2.0 2.5 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric : L1 MSE MCC σ=2 MCC σ=0.1 RQE 2 σ=2 RQE 2 σ=0.1 RQE 1 σ=2 RQE 1 σ=0.1
Sensory fusion optimal point - Measurements with gross errors 59 Figure 5-17 - Residual errors in voltage phases using L1 as the fusion metric, with σ=2 and σ=0.1. It is possible to observe that for MSE, MCC and REQ, with big PW size, all the results get affected by the error in Pinj3. While, for a smaller PW size, 0.1, the MCC and RQE metric ignore that measurement, and only its error is high, all the other component’s errors approach to zero. In the previous image is possible to observe that the decrease of the PW size leads to a better result that MSE, or MCC and RQE with a big PW. However, it’s not possible to observe the different effect of each evaluation metric in the errors, due the big error presented in the Pinj3. For that reason, the following image represent only the measurements for PW size 0.1, and ignore the residual error in Pinj3. Figure 5-18Residual errors in buses power injections and lines power flows using L1 as the fusion metric, with σ=0.1, and ignoring the residual error of the component Pinj3. 0.000 0.005 0.010 0.015 0.020 0.025 0.030 0.035 0.040 0.045 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric : L1 MSE MCC σ=2 MCC σ=0.1 RQE 2 σ=2 RQE 2 σ=0.1 RQE 1 σ=2 RQE 1 σ=0.1 0 0.00005 0.0001 0.00015 0.0002 0.00025 0.0003 0.00035 Pinj 100 Pinj 1 Pinj 2 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric: L1 MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.1
60 Selecting the optimal fusion point Figure 5-19 - Residual errors in voltage phases using L1 as the fusion metric , with σ=0.1. Despite what happened in the previous experiment (measurements without gross error), where the RQE case 2 generate residuals far from the ones obtained with MCC and RQE case 1, here, when the PW size decrease the RQE 2 errors approach to same obtained by the MCC. In figures 5-18 and 5-19 the residuals generated by MCC and RQE case 2 are practically the same. The results obtained with the Euclidean norm, L2, as fusion metric are really similar to the ones obtained with the absolute norm, L1. For that reason, the L2 results are in the annex B. The behavior of the fusion metrics L1 and L2 are similar with no obvious benefits in choose one or another metric. Figure 5-20 - Residual errors in buses power injection and lines power flow, using CIM as the fusion metric, with σ=1. Evaluation metrics with σ=2 and σ=0.1. 0 0.00002 0.00004 0.00006 0.00008 0.0001 0.00012 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric : L1 MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.1 0.0 1.0 2.0 3.0 4.0 5.0 6.0 7.0 8.0 9.0 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric : CIM MSE MCC σ=2 MCC σ=0.1 RQE 2 σ=2 RQE 2 σ=0.1 RQE 1 σ=2 RQE 1 σ=0.1
Sensory fusion optimal point - Measurements with gross errors 61 Figure 5-21 - Residual errors in voltage phases, using CIM as the fusion metric, with σ=1. Evaluation metrics with σ=2 and σ=0.1. Using the CIM as the fusion metric, the residual errors show a different behavior than the ones obtained by L1 and L2 metrics. In particular, the estimations obtained by the evaluation metric RQE (case one) present a big deterioration when applying the CIM as the fusion metric, either for high PW size 2, or small, size 0.1. In annex B is possible to observe the same results presented in figures 5-20 and 5-21, but only for MCC and RQE (case one and two) with σ=0.1, for a better observation of the behavior of RQE (case one) with the fusion metric CIM. In order to obtain better results, different PW size were tested with the evaluation metric RQE (case 1), an improvement of the results was observed for PW size 0.5, figure 5-22 and 5-23. Some other results obtained with a different PW for the RQE (one) can be found in annex B. Applying a convenient PW size to the evaluation metrics MCC and RQE (case 2), with the fusion metric CIM, the results are reliable and similar to the ones obtained with L1 and L2, figures 5-22 and 5-23. However, the behavior of the evaluation metric RQE (case 1) when combined with fusion metric CIM is not convenient, even for a PW size 0.5. Figure 5-22 - Residual errors in buses power injections and lines power flows, using CIM as the fusion metric (with σ=1) and ignoring the residual error of the component Pinj3. Evaluation metrics MCC and RQE (2) with σ=0.1, RQE (1) with σ=0.5. 0.00 0.02 0.04 0.06 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric : CIM MSE MCC σ=2 MCC σ=0.1 RQE 2 σ=2 RQE 2 σ=0.1 0 0.0002 0.0004 0.0006 0.0008 0.001 0.0012 Pinj 100 Pinj 1 Pinj 2 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric: CIM MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.5
62 Selecting the optimal fusion point Figure 5-23 - Residual errors in voltage phases, using CIM as the fusion metric (with σ=1). Evaluation metrics MCC and RQE (2) with σ=0.1, RQE (1) with σ=0.5. From the observation of figures 5-20 to 5-23, it’s evident that it’s complicated to apply CIM as the fusion metric in the most adequate way, because it is necessary to adjust not only the PW size of the fusion metric but also the PW of the evaluation metrics. Thereby, it is not such easy to identify that parameters as it is when using the absolute norm (L1) or Euclidean norm (L2). The residual errors are the most important thing to guarantee with the state estimation. However, the iterations or the duration of the process is also an important parameter, taking in account that the problem has been solved by a meta-heuristic methodEPSO [30]. Thus, in tables 5-10 and 5-11 can be found the amount of fitness evaluation each metrics combination took until reach the optimal point, and the time all that process took. Table 5-10 - Number of fitness evaluation each combination took to find the optimal point. fitness evaluations MSE MCC (σ=2) RQE1 (σ=2) RQE2 (σ=2) MCC (σ=0.1) RQE1 (σ=0.1) RQE2 (σ=0.1) L1 4.82E+06 2.70E+06 2.16E+06 1.78E+06 1.46E+06 2.98E+06 2.31E+06 L2 7.65E+06 6.46E+06 1.53E+07 1.31E+07 8.61E+05 2.13E+06 3.25E+06 CIM (σ=1) 1.85E+06 2.49E+06 9.41E+04 3.46E+06 2.41E+06 3.19E+05 2.89E+06 Table 5-11 – Duration of the process each combination took to find the optimal point. Duration (s) MSE MCC (σ=2) RQE1 (σ=2) RQE2 (σ=2) MCC (σ=0.1) RQE1 (σ=0.1) RQE2 (σ=0.1) L1 831 505 422 364 232 494 408 L2 1370 1002 2516 2347 134 354 564 CIM (σ=1) 295 389 16 596 376 54 504 The evaluation metric RQE (case 1), with PW size 0.1 and combined with the fusion metric CIM, doesn’t lead to a good result, so it’s irrelevant that it took just few fitness evaluations and less time than the others. The absolute norm, L1, took most of the times less time to converge than the Euclidean metric, L2, and that’s not surprising since the expression is simpler. However, the results are not concluding in which norm will always perform a better time. Also, it is important to mention that, for now, the problem has been solved by EPSO because the purpose of this thesis is to 0 0.00002 0.00004 0.00006 0.00008 0.0001 0.00012 0.00014 0.00016 0.00018 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric : CIM MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.5
Sensory fusion optimal point - Measurements with gross errors 63 propose new perspectives about the sensory fusion. For that reason, a meta-heuristic method is enough to solve the proposed problem, just to show that the new methods and metrics are applicable, convenient and to show them advantages. The need for a further work capable to solve the problem by a differential method, consuming much less time is evident. 5.3.1 - Results analysis / relevant conclusions. The application of the developed fusion method to the SE process doesn’t interfere with the capability of criteria as minimum Renyi’s quadratic error or maximum Correntropy criterion ignore outliers. From the tested fusion metrics, the CIM is the most complicated to manage, because is necessary to assign values to the PW size of the fusion metric and also of the evaluation metric. As it was observed previously, it’s not so simple to identify the right PW size as it is when using the absolute and the Euclidean norms, L1 and L2, respectively. Besides that, the CIM doesn’t show any advantage on its application. Moreover, even without gross errors it’s complicated to find a PW size that shows totally convenience to both fitness evaluation functions (fitness evaluation of group one and two). In fact, the highness of the distance between the optimal and the ideal point is unknown, so doesn’t make much sense to limit it through a PW method. With the introduction of an outlier, the evaluation metric RQE (case 2) generated errors close to the ones obtained with MCC. Despite the good results obtained in this section, the RQE2 meaning is not what has been sought in this thesis. The idea of the fusion step is to fusion information descendant from both sensory systems separately. RQE (case 2) promotes the interaction of each component’s error with all the errors from its own group and also with the other one. So it’s not separately. Moreover, it doesn’t present any advantage in comparison with the other metrics, actually, in the previous section it presented results even worse than the other criteria. For that reasons this approach to the RQE will be abandoned in the next chapter. 5.4 - Comparison between simple fusion and the developed fusion method As it was already mentioned, the inclusion of PMU in the monitoring system can be done in different ways. One possibility is the one that has been discussed on the previous sections of this chapter. Another way, the simplest and most common, is to include the measurements obtained by PMU in the same vector of SCADA measurements, computing the state estimation with all the measurements together, without any distinction based on its origin, that is, the proceed in chapter 3. This section aims to compare the results obtained by both sensory fusion methods. For that, the results of chapter 3 are compared with the result of the first section of this chapter, for the case with no gross errors. The objective here is not to compare the effect of the different metrics in the presence of gross errors, that was already explored in chapters 3.5 and 5.1.2., but to compare the two different approaches of sensory fusion. In the following figures is possible to observe the results for the simple fusion (without any distinguish by the measurements’ origin) and the developed fusion method in this chapter, with fusion metrics L1 and L2. This three cases were compared under the different evaluation metrics: MSE, MCC and RQE, with a PW size 0.1.
70 Application of the fusion method to the AC model Table 6-1Measurements used in the AC model without gross errors. Bus Power Flow result Measurement affected with noise (without PMU) Pinj (p.u.) 100 1.121140 1.122810 1 0.041322 0.039534 2 -0.807506 -0.809074 3 -0.569125 -0.572176 4 -0.048778 -0.045104 5 -0.048648 -0.043448 6 1.449080 1.451090 7 -0.291667 -0.291622 8 -0.059664 -0.062184 9 -0.430091 -0.429711 10 -0.315175 -0.318894 11 -0.032993 -0.033324 Qinj (p.u.) 100 -0.238982 -0.236943 1 0.455981 0.457423 2 -0.265455 -0.267485 3 -0.182535 -0.182715 4 -0.010337 -0.011612 5 -0.009668 -0.009689 6 0.590705 0.588743 7 -0.096752 -0.094147 8 -0.011773 -0.010499 9 -0.109040 -0.109329 10 -0.096427 -0.098370 11 -0.006464 -0.006126 ϴ (rad) 2 -0.012750 - 5 -0.014990 - 8 -0.015320 - V (p.u.) 2 0.997079 0.998354 5 0.996383 0.999383 8 0.995781 0.992141 Pij (p.u.) 3 4 0.169774 0.166874 6 7 -0.886705 -0.885805 Qij (p.u.) 3 4 0.158326 0.159126 6 7 -0.344520 -0.349320 Iinj(real) (p.u.) 2 -0.806411 - 5 -0.048673 - 8 -0.059729 - Iinj(im) (p.u.) 2 -0.276537 - 5 -0.010434 - 8 -0.012739 -
Measurements 71 The measurement sets that contain gross errors are presented in the table 6-2. These sets will be used to test the behavior of the evaluation metrics in the presence of outliers, in a real based network. Table 6-2 - Measurements used in the AC model with gross errors. Bus Measurements (missing measurement) Measurement (Power flow inversion) (with PMU) (with PMU) Pinj (p.u.) 100 1.122810 1.122810 1 0.039534 0.039534 3 -0.572176 -0.572176 4 -0.045104 -0.045104 6 0.000000 1.451090 7 -0.291622 -0.291622 9 -0.429711 -0.429711 10 -0.318894 -0.318894 11 -0.033324 -0.033324 Qinj (p.u.) 100 -0.236943 -0.236943 1 0.457423 0.457423 3 -0.182715 -0.182715 4 -0.011612 -0.011612 6 0.588743 0.588743 7 -0.094147 -0.094147 9 -0.109329 -0.109329 10 -0.098370 -0.098370 11 -0.006126 -0.006126 ϴ (rad) 2 -0.012826 -0.012826 5 -0.015083 -0.015083 8 -0.014795 -0.014795 V (p.u.) 2 0.998354 0.998354 5 0.999383 0.999383 8 0.992141 0.992141 Pij (p.u.) 3 4 0.166874 0.166874 6 7 -0.885805 0.885805 Qij (p.u.) 3 4 0.159126 0.159126 6 7 -0.349320 0.349320 Iinj(real) (p.u.) 2 -0.806061 -0.806061 5 -0.048913 -0.048913 8 -0.059329 -0.059329 Iinj(im) (p.u.) 2 -0.276589 -0.276589 5 -0.010506 -0.010506 8 -0.012667 -0.012667 6.2 - The inclusion of the PMU measurements in state estimation PMU has been theme of several discussions, many research about its integrations in the network’s monitoring system have been conducted [32]. Until now, this work has been developing a fusion method capable to integrate PMU measurements in the measurement system. However, it’s important to prove that the inclusion
72 Application of the fusion method to the AC model of that measurements are really favorable. Thus, a first case was tested, where all the buses have power injection sensors, table 6-1-measurements without PMU. Then a second case was studied, where the injected power and voltage magnitude sensors on buses 2, 5 and 8 where replaced by PMU’s. In this way, measurements of these buses were replaced by injected currents and voltages (magnitude and phases) measurements, table 6-1-Measurements with PMU. The residual errors obtained in each case, without and with PMU’s, is represented in the following figures 6-1 to 6-5. These results, when the evaluation metrics were RQE or Correntropy, were obtained for a PW with size 1. As it was already seen in previous chapters, when there are no outliers, MCC and MSE generate the same results. In following graphics, the MSE is not represented just for a matter of simplification of the results, once the state estimation obtained by the minimization of MSE or maximum Correntropy were practically equal. Figure 6-1 - Absolute residual errors in power injection, obtained without PMU. Figure 6-2 – Absolute residual errors in voltage magnitude, obtained without PMU. 0.0000 0.0005 0.0010 0.0015 0.0020 0.0025 0.0030 0.0035 0.0040 0.0045 0.0050 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 Qinj 100 Qinj 1 Qinj 2 Qinj 3 Qinj 4 Qinj 5 Qinj 6 Qinj 7 Qinj 8 Qinj 9 Qinj 10 Qinj 11 P34 P67 Q34 Q67 (p.u.) Without PMU MCC RQE 0.0000 0.0005 0.0010 0.0015 0.0020 0.0025 0.0030 0.0035 0.0040 0.0045 V2 V8 V5 (p.u.) Without PMU MCC RQE
The inclusion of the PMU measurements in state estimation 73 Figure 6-3Absolute residual errors in power and current injections and lines power flow, obtained with PMU. Figure 6-4 – Absolute residual errors in voltage magnitudes, on the left, and on voltage phases, on the right. Errors obtained with PMU. From the analysis of the residual errors represented on the previous graphics, it can be said that, in a general mode, the errors have decreased with the introduction of PMU; However, the measured components are not the same in both cases. For instance, the currents injections only appear with the introduction of the PMU, and the power injection, in buses 2,5 and 8, were only measured when PMU where not connected to that buses. Thus, to analyze the PMU integration effect on the results, it’s necessary to observe not only the residual errors (difference between estimations and measurements) but also the true estimation errors. In this case is possible to observe the real error, once the measurements were generated from real power flow values, presented in annex A.3.1. It’s important to mention that the estimations were computed based on measurements affected with noise. The state variables, voltage phases and magnitude, obtained in each case, with and without PMU, were used to calculate estimates to all the components present in both cases (figures 6-1 to 6-5). In this case, the estimated values were not compared with the measurements (that are already affected with errors) but with the original values. The error between those state estimations and the original values can be observed in figures 6-6 to 6-9. 0.0000 0.0005 0.0010 0.0015 0.0020 0.0025 0.0030 0.0035 0.0040 0.0045 0.0050 Pinj 100 Pinj 1 Pinj 3 Pinj 4 Pinj 6 Pinj 7 Pinj 9 Pinj 10 Pinj 11 Qinj 100 Qinj 1 Qinj 3 Qinj 4 Qinj 6 Qinj 7 Qinj 9 Qinj 10 Qinj 11 P34 P67 Q34 Q67 I2 (re) I5 (re) I8I8 (re) I2(i) I5(i) I8(i) (p.u.) With PMU MCC RQE 0.0000 0.0002 0.0004 0.0006 0.0008 0.0010 0.0012 V2 V5 V8 (p.u.) With PMU MCC RQE 0.0000 0.0010 0.0020 0.0030 0.0040 0.0050 ϴ2 ϴ5 ϴ8 (p.u.) With PMU MCC RQE
74 Application of the fusion method to the AC model Figure 6-5 - Errors between estimations and original values of injected power. Figure 6-6 - Errors between estimations and original values of injected currents and lines power flows. Figure 6-7 - Errors between estimations and original values of voltages phases. 0 0.0005 0.001 0.0015 0.002 0.0025 0.003 0.0035 0.004 0.0045 0.005 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 Qinj 100 Qinj 1 Qinj 2 Qinj 3 Qinj 4 Qinj 5 Qinj 6 Qinj 7 Qinj 8 Qinj 9 Qinj 10 Qinj 11 (p.u.) Errors on power injections MCC (without PMU) RQE (without PMU) MCC (with PMU) RQE (with PMU) 0 0.0005 0.001 0.0015 0.002 0.0025 0.003 0.0035 0.004 P34 P67 Q34 Q67 I2 (re) I5 (re) I8I8 (re) I2(i) I5(i) I8(i) (p.u.) currents injections and power flows MCC (without PMU) RQE (without PMU) MCC (with PMU) RQE (with PMU) 0 0.00001 0.00002 0.00003 0.00004 ϴ2 ϴ5 ϴ8 (p.u.) Errors on voltage phases MCC (without PMU) RQE (without PMU) MCC (with PMU) RQE (with PMU)
The inclusion of the PMU measurements in state estimation 75 Figure 6-8 - Errors between estimations and original values of voltages magnitudes. From the observation of the figures 6-5 to 6-7, it’s easy to conclude that the replacement of conventional sensors by PMU brigs better results, with lower errors and more reliable state estimations. For a better general perspective of the errors, obtained by the different sensor systems, the mean absolute percentage error (MAPE) was calculated to each case. Once the results obtained by the minimization of MSE were really similar to the ones obtained by MCC, in the previous graphics, the representations MSE errors were ignored, in order to save some space. In table 6-2 is possible to see the MAPE to each sensor system, obtained using different SE criteria. Once both criteria lead to the same solution, as expected, the MAPE obtained for the MSE and MCC criterions were the same. The expression used to calculate the MAPE was 𝑀𝐴𝑃𝐸=1𝑛∑|𝑧𝑘−𝑧𝑘 𝑍𝑘| 𝑛𝑘=1 , (6.1) where 𝑍𝑘 is the true value of the component k, and 𝑍𝑘 is its estimated value. Also the RQE was calculated for the diverse results, by the expression already presented in chapter 3. Table 6-3 - Mean percentage error and entropy calculated to each sensor system Without PMU With PMU MSE MCC RQE MSE MCC RQE MAPE 2.27% 2.27% 2.29% 1.83% 1.83% 1.83% RQE -3.09472 -3.09472 -3.09472 -3.09483 -3.09483 -3.09483 If by the observation of the previous graphics could leave some doubt, about the advantage or disadvantage of PMU inclusion in measure system, the data in table 6-2 clarify it. The inclusion of PMU in the sensor system decreases the error of the estimations, for both evaluation metrics, MCC and RQE, and it also decreases de error’s entropy. Thus, one must say that, PMU increases the reliability of the measurement system. 6.3 - State estimation using the fusion method In chapter 4 were proposed a fusion method with capability to assigning stochastic values to each sensory system, according to the probability of each system to conduct to the real state. As the information about those stochastic values is not yet available, in chapter 5, was 0 0.0001 0.0002 0.0003 0.0004 0.0005 V2 V5 V8 (p.u.) Errors on voltage magnitude MCC (without PMU) RQE (without PMU) MCC (with PMU) RQE (with PMU)
76 Application of the fusion method to the AC model proposed some fusion metrics to find an optimal fusion point. This point is over the Paretofront resulted from the trust variation on each sensory system. That optimal point can be found by minimizing the absolute distance between the ideal and the optimal point we are searching for. For this case, the SE objective functions, depending on different evaluation metrics, are the ones presented in the following table. Table 6-4 - Objective function to find the optimal point by the L1 metric. Objective function Fusion metric: L1 Evaluation metric MSE min |𝑚𝑎𝑥𝑓𝑖𝑡1−∑(𝑧𝑖−𝑧𝑖)2 𝑖𝜖𝑛1|+|𝑚𝑎𝑥𝑓𝑖𝑡2−∑(𝑧𝑖−𝑧𝑖)2 𝑖𝜖𝑛2| MCC min |𝑚𝑎𝑥𝑓𝑖𝑡1−∑𝑒−(𝑧𝑖−𝑧𝑖)2 2𝜎2 𝑖𝜖𝑛1|+|𝑚𝑎𝑥𝑓𝑖𝑡2−∑𝑒−(𝑧𝑖−𝑧𝑖)2 2𝜎2 𝑖𝜖𝑛2| RQE min |𝑚𝑎𝑥𝑓𝑖𝑡1−∑ ∑ 𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑗𝜖𝑛1𝑖𝜖𝑛1|+|𝑚𝑎𝑥𝑓𝑖𝑡2−∑ ∑ 𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑗𝜖𝑛2𝑖𝜖𝑛2| As it was seen in chapter 5, those expressions can be simplified by the ones presented in table 6-5. In both tables the parameter σ refers to the PW size and na is the set of measurements provided by the sensory system a. N consists in the aggregation of both measurement system, 1 and 2. N=n1Un2. Table 6-5 – Simplification for the SE objective function, given by the L1 metric. Objective function Fusion metric: L1 Evaluation metric MSE min ∑(𝑧𝑖−𝑧𝑖)2 𝑖𝜖𝑛1+∑(𝑧𝑖−𝑧𝑖)2 𝑖𝜖𝑛2 MCC max ∑𝑒−(𝑧𝑖−𝑧𝑖)2 2𝜎2 𝑖𝜖𝑛1+∑𝑒−(𝑧𝑖−𝑧𝑖)2 2𝜎2 𝑖𝜖𝑛2 RQE max ∑ ∑ 𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑗𝜖𝑛1𝑖𝜖𝑛1+∑ ∑ 𝑒−(𝑟𝑖−𝑟𝑗)2 2𝜎2 𝑗𝜖𝑛2𝑖𝜖𝑛2 When using the evaluation metrics MSE and MCC, the objective functions are mathematically equivalent to the simple fusion. For the RQE, the simplification doesn’t lead to the same mathematical expression as the simple fusion, because of the following reason: in the simple fusion, the RQE promotes the interaction between all the errors, independent to the sensory system each error belongs, while, in the developed fusion method, the sensory systems are separated. That is, for RQE, each error interacts only with the errors within the same sensory system. The main advantage, of the fusion method compared to the simple fusion, is its ability to easily convert into a stochastic method, by assigning weights to each one of the sensory systems term. In order to find proper ways to calculate those stochastic values, new research and studies must be performed.
State estimation using the fusion method 77 In this way, the developed fusion method was tested with the measurements presented in table 6-1measurements with PMU. The residual errors obtained with each evaluation metric can be found in the following graphics. Figure 6-9 - Absolute value of residual errors in injected power and currents and lines power flow for the AC model, without gross errors. Figure 6-10Absolute value of residual errors in voltage phases for the AC model, without gross errors. Figure 6-11Absolute value of residual errors in voltage magnitudes for the AC model, without gross errors. 0.0000 0.0005 0.0010 0.0015 0.0020 0.0025 0.0030 0.0035 0.0040 0.0045 0.0050 Pinj 100 Pinj 1 Pinj 3 Pinj 4 Pinj 6 Pinj 7 Pinj 9 Pinj 10 Pinj 11 Qinj 100 Qinj 1 Qinj 3 Qinj 4 Qinj 6 Qinj 7 Qinj 9 Qinj 10 Qinj 11 P34 P67 Q34 Q67 I2 (re) I5 (re) I8I8 (re) I2(i) I5(i) I8(i) (p.u.) Power and currents MSE MCC RQE 0.0000 0.0001 0.0002 0.0003 0.0004 0.0005 0.0006 0.0007 ϴ2 ϴ5 ϴ8 (p.u.) Voltage magnitude MSE MCC RQE 0.0000 0.0010 0.0020 0.0030 0.0040 0.0050 V2 V5 V8 (p.u.) Voltage phase MSE MCC RQE
78 Application of the fusion method to the AC model As it is possible to observe, the same conclusions of the DC model are transposed here. The MSE and MCC approach the same result, when there are no gross errors. The RQE leads to a different state with different residual errors. Despite of the difference, once the errors are always lower than the noise introduced to the original values, that state is still acceptable and reliable. Is still important to mention that, the point obtained with RQE is not worse than the one obtained by MCC or MSE. It’s different, and in some components leads to a lower error, and in others to a higher error. In fact, it is the state obtained considering the criterion of less entropy, calculated by Renyi’s Quadratic Entropy, and from that point of view it conducts to a better point, comparing to the ones obtained with the other criteria. In the following chart is possible to see the errors distribution, calculated though the PW method with a standard deviation of 0.1. Figure 6-12 - pdf of the errors for the case without outliers, using MSE, on the left, and MCC on the right. Figure 6-13pdf of the errors for the case without outliers, using RQE. In the previous tests, all the curves are similar of even equal, thus, the representation of each curve was made in independent charts. Therefore, a better observation of each curve is possible, without overlapping. As there are no gross errors, and the noise introduced to the original measurements was a Gaussian cantered in zero, the residuals distribution assumed also that Gaussian form, for each one of the different metrics.
State estimation using the fusion method without gross errors 79 6.4 - State estimation using the fusion method and considering typical errors In this section, two different cases of bad data are studied, in order to compare the information theory related concepts with the classic MSE. The cases are: one missing measurement and an inversion of line power flow, where both active and reactive power are inverted. For both cases the PW applied for MCC and RQE criteria had size 0.1. 6.4.1 - Missing measurement in 𝑃𝑖𝑛𝑗6 In this subsection is assumed an outlier, a typical error of a missing measurement. For that, the active injected power in bus 6 was assumed as zero. The state estimation was computed for the 3 different criteria, MSE, MCC and RQE. The measurement set is the same of the table 6-2-measurements with gross errors (missing measurement). Figure 6-14 – Absolute value of residual errors in injected power and currents and lines power flow for the AC model with an missing measurement. Figure 6-15 - Absolute value of residual errors in voltage phases for the AC model with an missing measurement. 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 Pinj 100 Pinj 1 Pinj 3 Pinj 4 Pinj 6 Pinj 7 Pinj 9 Pinj 10 Pinj 11 Qinj 100 Qinj 1 Qinj 3 Qinj 4 Qinj 6 Qinj 7 Qinj 9 Qinj 10 Qinj 11 P34 P67 Q34 Q67 I2 (re) I5 (re) I8 (re) I2(i) I5(i) I8(i) (p.u.) Power and currents MSE MCC RQE 0.0000 0.0010 0.0020 0.0030 0.0040 ϴ2 ϴ5 ϴ8 (p.u.) Voltage magnitude MSE MCC RQE
86 Conclusions and future work selected Parzen window size, thus, a wider window allows the assumed error distribution to be more embracing than a narrower window. A thinner window is more restrict, allowing the outliers to be identified and ignored. In Chapter 6 is possible to observe the SE’s results in the presence of two usual cases of outliers, namely missing measurement and inversion of a line power flow. While these outliers notably influenced the state estimation provided by the minimization of MSE, leading to high deviations for most of the components, MCC and RQE managed to provide good estimations for all of the components including those affected by the error. The study about the inclusion of the PMU in the measurement system proved to be important as it leads to a better state estimation, by minimizing its residual errors. Such results can be explained by the high trust in PMU, and also, by the inclusion of voltage phases to the measurement set. PMU has other interesting advantages to the power system, among which: measurements synchronized by GPs-clock and lower sampling periods. The big progress provided by this thesis is the perspective assigned to the fusion method of two separately sensory systems. A perspective of stochastic fusion was born here. However, further research and tests must be done in order to find the stochastic values to assign to each sensory system. For now, such information is not available yet, hence, this work used another way to choose the optimal fusion point. This point was selected by its minimal distance to an ideal point. The ideal point is the point with coordinates defined by both best fitness evaluations of each sensory system. The chosen metric to evaluate that distance was the absolute norm. The reason for this choice was its simplicity and its capability to be easily converted into a stochastic fusion. Such transition will be possible when there’s a larger amount of information on the probability of each sensory system to provide the estimation corresponding to the real state. The fusion method doesn’t interfere with the properties of the evaluation functions MCC and RQE. In fact, in Chapter 6 it is possible to behold the application of the fusion method using those evaluation metrics in a real parameters based network. The results are convenient and show good properties when facing outliers. Summing up: in this thesis an innovative fusion method, integrating also alternative state estimation criteria, was developed and proposed. The developed method was tested in different network circumstances presenting always advantages in comparison with the traditional process. 7.2 - Future work The evident objective of this thesis was to expose promising properties of the developed method and to confirm its advantages related to the conventional process. For that proof of concept, a meta-heuristic algorithm – EPSO - was used. As such properties were presented and its convenient application was proven, now is desirable to adapt the developed method to real time problems. The EPSO algorithm takes a lot of time to find the optimal fusion point in real networks. Thus, the need for a development of a real-time resolution method emerges. One proposal for a future work is to resolve the proposed optimization problems by a differential method such as the Gauss-Newton. Another imperative work is about the stochastic fusion. As it was shown in Chapter 4, the state estimations’ results differ a lot with the variation of the weights assign to each one of
Future work 87 the sensory system. To choose the state closer to the real one, a stochastic fusion may be used, where the parameters α and β, described in Chapter 4, are determined by the probability of each sensory system to lead to the real system’s state. To obtain such information, studies must be done considering: reliability of each sensory type, the number of sensors in each system, their location and other possible issues as errors by asynchronism. One determining factor for the outlier’s identification is the Parzen window size. In the previous studies it was possible to find a parameter able to provide a proper identification and treatment of bad data. However, the selection of that parameter has to take in consideration the magnitude of the errors. More studies are required to elaborate a process able to determine the ideal value for the Parzen window’s size according to the environment where it will be inserted.
88
89 References 1. V. Miranda, A. Santos, and J. Pereira, “State Estimation Based on Correntropy: A Proof of Concept,” in IEEE Trans. Power Syst., vol.24, no.4, pp.1888,1889, Nov. 2009. 2. Costa, Antonio Simões, André Albuquerque, and Daniel Bez. "An estimation fusion method for including phasor measurements into power system real-time modeling", in IEEE Transactions on Power Systems, vol.28, no.2 , pp.1910-1920, 2003. 3. Zeleny, M. " Multiple Criteria Decision Making, McGraw-Hill", New York,1982. 4. C.E. Shannon, "A mathematical theory of communication", Bell System Technical Journal, Vol. 27, no.3, pp. 379–423, 1948. 5. J. C. Principe, " Information theoretic learning: Renyi's entropy and kernel perspectives", Springer Information and Science, Springer-Verlag New York, 2010, Chapters 1 and 2. 6. E. Parzen, “On the estimation of a probability density function and the mode”, Annals Math. Statistics, vol. 33, 1962, p. 1065. 7. V. Miranda, "Information Theoretic Learning Principles a short tutorial", in ISAP Conference and debate, 2015, pp.12-15. 8. W.F.Liu, P.P. Pokharel, and J.C. Principe, "Correntropy: A localized similarity measure", in 2006 IEEE International Joint Conference on Neural Network Proceedings, 2006: pp. 4919-4924. 9. Liu, W.F., P.P. Pokharel, and J.C. Principe," Correntropy: properties and applications in non-gaussian signal processing", in IEEE Transactions on Signal Processing, 2007. vol.55, no.11, pp. 5286-5298. 10. Clara Gouveia, A.S., Helder Leite, "The Application Of Distribution State Estimation To Support A Real-Time Voltage Control Algorithm: A Path To Increase The Integration Of Distributed Generation", in 21st International Conference on Electricity Distribution, 2011. 11. André Madureira, R.B., Luís Seca, "Advanced System Architecture And Algorithms For Smart Distribution Grids: The Sustainable Approach", in 23rd International Conference on Electricity Distribution, 2015. 12. Paula S. Castro Vide, FP Maciel Barbosa, and Isabel M. Ferreira, "State estimation model including synchronized phasor measurements." in Universities' Power Engineering Conference (UPEC), Proceedings of 2011 46th International. VDE, 2011. 13. Paula S. Castro Vide, FP Maciel Barbosa, and Isabel M. Ferreira, "Combined use of SCADA and PMU measurements for power system state estimator performance
90 enhancement." in Energetics (IYCE), Proceedings of the 2011 3rd International Youth Conference on. IEEE, 2011. 14. A. Monticelli, "State Estimation In Electric Power Systems, A Generalized Approach", Kluwer Academic Publishers , Chapter 2, 1999. 15. W. Wu, Y.G., B. Zhang, A. Bose, S. Hongbin," Robust state estimation method based on maximum exponential square", in IET Generation, Transmission & Distribution, 2011. 16. Abur, Ali, and Antonio Gomez Exposito, "Power system state estimation: theory and implementation", CRC press, 2004, Chapter 2. 17. M-estimators by Zhengyou Zhang , 1996. Available on http://research.microsoft.com/en-us/um/people/zhang/INRIA/Publis/TutorialEstim/node24.html , last acess: 19th June 2016 18. Green, Peter J. "Iteratively reweighted least squares for maximum likelihood estimation, and some robust and resistant alternatives", in Journal of the Royal Statistical Society. Series B (Methodological) , pp:149-192,1984. 19. W.J.J. Rey, "Introduction to Robust and Quasi-Robust Statistical Methods", ed. Springer. Berlin, Heidelberg 1983, Chapter 8. 20. Mitchell, H.B., "Multi-sensor data fusion an introduction", New York: Springer, 2007, Chapter 1. 21. K. Faceli, A.C.P.L.F. De Carvalho, and S.O. Rezende, "Combining intelligent techniques for sensor fusion", Applied Intelligence, vol.20, no.3: pp. 199-213, 2004. 22. G., K., "Data Fusion: From primary metrology to process measurement", in IEEE Instrumentation and measurement, vol.3 pp:1325-1329, 1999. 23. Karthik Nandakumar, "Integration of multiple cues in biometric systems", Diss. Michigan State University, pp:1-11, 2005. 24. R. Boudjemaa, F.A.B.," Parameter estimation methods for data fusion", National physical Laboratory Report No. CMSC 38-04, 2004. 25. F., D.-W.H., "Sensor models and multisensor integration", in International Journal of Robotics Research no.7, pp.97-113. 1988. 26. Belur V. Dasarathy, "Decision fusion". Vol. 1994. Los Alamitos, CA: IEEE Computer Society Press, 1994. 27. C. A. Noonan, and M. Pywell. "Aircraft sensor data fusion: An improved process and the impact of ESM enhancements", in Multi-Sensor Systems and Data Fusion for Telecommunications, Remote Sensing and Radar (1998). 28. P. N. Pereira Barbeiro, et al. "Exploiting autoencoders for three-phase state estimation in unbalanced distributions grids" Electric Power Systems Research 123 pp:108-118, 2005 29. Kai Strunz, N. Hatziargyriou, and C. Andrieu. "Benchmark systems for network integration of renewable and distributed energy resources", Cigre Task Force C6 (2014): 04-02, pp:33-37 30. V. Miranda, H. Keko, and A. J. Duque, “Stochastic star communication topology in evolutionary particle swarms (EPSO)”, Int. J. Comput. Intell. Res., vol. 4, no. 2, pp.105116, 2008. 31. Opara, Jakov Krstulovic. "Information Theoretic State Estimation in Power Systems." (2014).
91 32. Aminifar, Farrokh, et al. "Synchrophasor measurement technology in power systems: Panorama and state-of-the-art", in IEEE Access 2 (2014): 1607-1628.
92
93 Annex A 12 busses system The network used to test the different fusion methods, and state estimation criteria, is presented in this chapter. This network consists in an adaptation from a Medium Voltage Distribution Network Benchmark, European configuration, from Cigre reports [29]. A typical European MV distribution network is three-phase and its structure can either be meshed or radial, being that in rural environment it tends to be radial. There is a big effort to keep the European networks balanced. Once those networks are assumed to be symmetric and balanced, in the study cases, the neutral wires are going to be ignored. Therefore, the presented line parameters will be already the equivalent of the three phases. As it is possible to see in figure A-1 there are some distributed energy resources (DER) along the network. More specifically, in buses 1, 6 and 9, represented by generators in the scheme of figure A-1. In the next sessions is possible to see the network topology, line parameters, the DER productions, load characteristics, injected power, the true values power flow values and “measurements” affected by Gaussian errors with small dimension. A.1 - Topology The topology of the utilized network is shown in figure A-1. In the connection from bus 100 to 1, is possible to see a transformer with voltage 63/15 KV with a leakage reactance of 8% and a nominal power 25MVA. The branches impedances can be consulted in the following table, A-1, measured in Ω, and in table A-2, measured in p.u. To convert into p.u. system it was used the base power Sb=25 MVA to the DC model and Sb=2.5 MVA to the AC model. The voltage base is Vb=15KV for the medium voltage part of the network (buses 1 to 11) and Vb=63KV for the bus 100.
94 12 buses sytem Figure A-1-Network topology. Table A-1 Line parameters in ohm. Branch From To R X (Ω) (Ω) 1 1 2 0.31640 0.35000 2 2 3 0.49720 0.55000 3 3 4 0.06780 0.07500 4 4 5 0.06780 0.07500 5 5 6 0.16950 0.18750 6 6 7 0.02260 0.02500 7 7 8 0.19210 0.21250 8 8 9 0.03390 0.03750 9 9 10 0.09040 0.10000 10 10 11 0.03390 0.03750 11 11 4 0.05650 0.06250 12 3 8 0.14690 0.16250
Topology 95 Table A-2Line parameters in p.u. Branch From To R X R X (p.u.) with Sb=25MVA (p.u.) with Sb=25MVA (p.u.) with Sb=2.5MVA (p.u.) with Sb=2.5MVA 1 1 2 0.035156 0.038889 0.003516 0.003889 2 2 3 0.055244 0.061111 0.005524 0.006111 3 3 4 0.007533 0.008333 0.000753 0.000833 4 4 5 0.007533 0.008333 0.000753 0.000833 5 5 6 0.018833 0.020833 0.001883 0.002083 6 6 7 0.002511 0.002778 0.000251 0.000278 7 7 8 0.021344 0.023611 0.002134 0.002361 8 8 9 0.003767 0.004167 0.000377 0.000417 9 9 10 0.010044 0.011111 0.001004 0.001111 10 10 11 0.003767 0.004167 0.000377 0.000417 11 11 4 0.006278 0.006944 0.000628 0.000694 12 3 8 0.016322 0.018056 0.001632 0.001806 A.2 - Load and DER data The network loads are dived into two groups, “residential” and “commercial/industrial”. In the next table, A-3, is possible to see the maximum value for the apparent power of the loads in each bus, as well as the power factor of each installation type. Table A-3 - Apparent power and power factor for the loads in each bus. Bus Apparent Power, S (kVA) Power factor Residential Commercial/Industrial Residential Commercial/Industrial 1 8700 4500 0,98 0,95 2 2500 0,95 3 185 1650 0,98 0,95 4 245 0,98 5 250 0,98 6 265 0,98 7 900 0,95 8 305 0,98 9 2750 0,95 10 290 800 0,98 0,95 11 170 0,98 In figure A-1 is represented some DER in buses 1, 6 and 9. Such distributed resources are cogeneration, mini-hydric and solar photovoltaic as represented in table A-4.
102 12 buses sytem Table A-15PSS/E results considering R=0 Bus Number Voltage (p.u.) Angle (deg) Angle (rad) 1 1.001954 -0.509292 -0.0088888 2 1.001163 -0.765749 -0.0133648 3 1.001558 -0.886939 -0.0154800 4 1.001690 -0.878788 -0.0153377 5 1.001886 -0.854360 -0.0149114 6 1.002397 -0.787507 -0.0137446 7 1.002302 -0.801535 -0.0139894 8 1.001719 -0.881691 -0.0153884 9 1.001659 -0.893212 -0.0155895 10 1.001619 -0.896662 -0.0156497 11 1.001644 -0.890455 -0.0155414 100 1.000000 0.000000 0.0000000 For the PSS/E results, once again, it was calculated the buses injected power, table A-15, That values were contaminated with noise just like in the previous subsection. The contamination was: Gaussian error with the width [-0.0004; 0.0004] to the power injections and lines power flows; Gaussian error with the width [-0.0005; 0.0005] to the voltage magnitude; Gaussian error with the width [-0.00002; 0.00002] to the voltage phase, considering the voltage phase error in radians. The generated measurements can be found in table A-16 and A-17. Table A-16-Power flow results and generated measurements, for active power, considering R=0. Bus Power Flow result Measurement affected with noise Pinj (p.u.) 100 0.111326 0.111469 1 0.00413 0.004325 2 -0.080749 -0.080625 3 -0.056928 -0.056763 4 -0.004777 -0.004955 5 -0.004907 -0.005229 6 0.144789 0.144945 7 -0.029052 -0.028692 8 -0.005973 -0.006022 9 -0.042991 -0.042779 10 -0.031513 -0.031764 11 -0.003354 -0.003434
Power flow results and measurements 103 Table A-17 - Power flow results and generated measurements, for reactive power and voltages, considering R=0. Bus Power Flow result Measurement affected with noise Qinj (p.u.) 100 -0.02393 -0.024226 1 0.045606 0.045582 2 -0.02654 -0.02667 3 -0.018284 -0.018049 4 -0.001043 -0.00102 5 -0.000966 -0.000884 6 0.058912 0.059035 7 -0.009478 -0.009279 8 -0.00133 -0.001663 9 -0.010813 -0.010482 10 -0.009614 -0.009353 11 -0.000621 -0.000693 ϴ (rad) 2 -0.013365 -0.013383 5 -0.015338 -0.014913 8 -0.015388 -0.015403 V (p.u.) 2 1.001163 1.00148 5 1.001886 1.00200 8 1.001719 1.00214 To solve the state estimation by the DC model, it’s necessary only few of the data presented in the previous table. Once the voltage magnitudes are assumed to be equal to 1 in all the buses, that components can be ignored. The same happens with the reactive injected power, once the voltage magnitudes are assumed to be close to one and the phase of contiguous buses are assumed to be nearly equal. Thereby, in table A-18 are presented the measurements used in the DC model. Table A-18DC measurements considering the adapted network. R=0. Bus/line Power Flow result Measurement affected with noise Pinj (p.u.) 100 0.111326 0.111469 1 0.004130 0.004325 2 -0.080749 -0.080625 3 -0.056928 -0.056763 4 -0.004777 -0.004955 5 -0.004907 -0.005229 6 0.144789 0.144945 7 -0.029052 -0.028692 8 -0.005973 -0.006022 9 -0.042991 -0.042779 10 -0.031513 -0.031764 11 -0.003354 -0.003434 ϴ (rad) 2 -0.013365 -0.013383 5 -0.014911 -0.014913 8 -0.015388 -0.015403
104 12 buses sytem To guarantee observability of the network it was also introduced power flow measurements in lines 3-4 and 6-7. For the DC model, those components are easily calculated by the following expression 𝑃𝑖𝑗=𝛳𝑖−𝛳𝑗 𝑥𝑖𝑗 . (X.8) where 𝑥𝑖𝑗 is the line ij reactance, and 𝛳𝑖 is the voltage phase at bus i. Thus, to the measurements included in table A-16 can be added the power flow measurements, presented in table A-19. The calculated values were affected with a Gaussian error equivalent to the one induced in the injected power. Table A-19Power Flow measurements in lines 3-4 and 6-7 assuming DC model. Line Power Flow result Measurement affected with noise Pij (p.u.) 3-4 -0.017071 -0.016290 6-7 0.088141 0.087880 It’s important to mention that the voltage phases were obtained by PMU. In fact, the DC model will be used considering that there are PMU sensors on buses 2, 5 and 8. The PMU measures voltages and current phasors. For the simplifications of DC model is considered that the voltage magnitude is 1 p.u., thus, the injected power has the same value as the injected current. Hence, for the DC model will be consider that the PMU measurements are voltages phases and power injections (once for this model the power injections values are the same as current injections).
105 Annex B Some results of the optimal fusion point In this annex are present some graphics about the Chapter 5, section 3. This graphics allow a better observation of some results and also show some experimentations done until find the desirable result. Figure B-1Residual errors in buses power injections and lines power flows, using L2 as the fusion metric, with σ=2 and σ=0.1. Figure B-2 - Residual errors in voltage phases using L2 as the fusion metric, with σ=2 and σ=0.1. 0.0 0.5 1.0 1.5 2.0 2.5 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric : L2 MSE MCC σ=2 MCC σ=0.1 RQE 2 σ=2 RQE 2 σ=0.1 RQE 1 σ=2 RQE 1 σ=0.1 0.000 0.010 0.020 0.030 0.040 0.050 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric : L2 MSE MCC σ=2 MCC σ=0.1 RQE 2 σ=2 RQE 2 σ=0.1
106 Some results of the optimal fusion point Figure B-3 - Residual errors in buses power injections and lines power flows using L2 as the fusion metric, with σ=0.1 and ignoring the residual error of the component Pinj3. Figure B-4Residual errors in voltage phases using L2 as the fusion metric, with σ=0.1. Figure B-5 - Residual errors in buses power injections and lines power flows using CIM as the fusion metric, with σ=1 and ignoring the residual error of the component Pinj3. Evaluation metrics with σ=0.1. 0 0.00005 0.0001 0.00015 0.0002 0.00025 0.0003 0.00035 Pinj 100 Pinj 1 Pinj 2 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10Pinj 11 P34 P67 (p.u.) Fusion metric: L2 MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.1 0 0.00002 0.00004 0.00006 0.00008 0.0001 0.00012 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric : L2 MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.1 0 0.2 0.4 0.6 0.8 1 Pinj 100 Pinj 1 Pinj 2 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10 Pinj 11 P34 P67 (p.u.) Fusion metric: CIM MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.1
107 Figure B-6 - Residual errors for each estimation of voltage phases using CIM as the fusion metric, with σ=1. Evaluation metrics with σ=0.1. Figure B-7 - Residual errors in buses power injections and lines power flows, using CIM as the fusion metric with σ=1. Evaluation metric: RQE (case 1) with σ=0.1, σ=0.5 and σ=0.05. Figure B-8 - Residual errors in voltage phases, using CIM as the fusion metric with σ=1. Evaluation metric: RQE (case 1) with σ=0.1, σ=0.5 and σ=0.05. 0 0.002 0.004 0.006 0.008 0.01 0.012 0.014 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric : CIM MCC σ=0.1 RQE 2 σ=0.1 RQE 1 σ=0.1 0 1 2 3 4 5 6 7 8 Pinj 100 Pinj 1 Pinj 2 Pinj 3 Pinj 4 Pinj 5 Pinj 6 Pinj 7 Pinj 8 Pinj 9 Pinj 10Pinj 11 P34 P67 (p.u.) Fusion metric: CIM RQE 1 σ=0.1 RQE 1 σ=0.5 RQE 1 σ=0.05 0 0.005 0.01 0.015 0.02 ϴ2 ϴ5 ϴ8 (p.u.) Fusion metric: CIM RQE 1 σ=0.1 RQE 1 σ=0.5 RQE 1 σ=0.05
108 Some results of the optimal fusion point
109 Annex C Article for submission Considering the fusion method developed in this thesis, as well as the studies about state estimation concepts related with information theory, it was decided to write a paper. In this paper, the main ideas about a new sensory fusion perspective are presented. Also some experiences about the integration of concepts related with information theory in the developed fusion method are exposed. In this annex, a long abstract of this thesis is shown.
110 Abstract –This letter proposes an innovative perspective about two topics of the state estimation process. The first is related to the fusion of two distinct sensory systems, where the attention paid to each of them can vary. The second aspect is about the replacement of the conventional optimization function (Weighted Least Square) by information theory related criterions, namely Correntropy and Renyi’s Quadratic Entropy. These criteria aim to propose a novel way to identify and correct large errors. Index Terms — Correntropy, Information Theory, Renyi’s quadratic entropy, Sensory fusion, State estimation. I. INTRODUCTION HIS paper opens the discussion on a new perspective about the power system State Estimation (SE). The inclusion of PMU in the measurement system has been having a growing importance. However, the data fusion from conventional sensors and PMU has been proposed in a simple way: all the measurements are collected together with no distinction on which type of sensor are involved. In this paper an innovative vision about the data fusion in state estimation is presented. Another aspect about the actual SE is its criterion. The Least Square Error (LSE) error is an optimal approach to estimate parameters only if the underlying distribution of errors is Gaussian. Thus, a gross error will cause disturbance in the estimation and, for that reason, should be identified and removed. Once LSE is not capable to ignore the gross errors, additional features are needed in the state estimator in order to filter possible outliers. The new paradigm proposed in [1] is based in maximizing the information that one can extract from the available measurements, minimizing the information on the residuals. In this work, it is extended to the sensory fusion context. This paper presents a proof of concept in terms of a theoretical model and examples about a new perspective of sensory fusion and on how the adoption of information theory related concepts allow the natural identification and correction of gross errors, also in the sensory fusion context. 1 e-mail: brunacostatavaresmail.com 2 e-mail: [email protected] II. PARZEN WINDOW, RENYI’S QUDRATIC ENTROPY AND CORRENTROPY The Parzen Window is a useful method to estimate a pdf distribution when the available data are discreet samples. It consists of applying in each sample a Gaussian kernel with a determined standard deviation [2]. The estimated distribution is given by: 𝑓 𝑌(𝑧)=1 𝑛∑𝐺(𝑧 − 𝑦𝑖, 𝜎2𝐼) 𝑛 1. (1) An ideal estimation would lead to a residual error distribution with no information, being represented by a Dirac function. If that distribution is centered at zero, it means that all errors are zero. In this way, a situation with minimum Entropy is found. One interested measure of entropy is Renyi’s Quadratic Entropy (RQE) and is given by the following expression (2) [3]. Applying the PW method one gets (3). 𝐻2= −𝑙𝑜𝑔 ∑𝑝𝜀𝑖 2 𝑛 𝑖=1 , (2) 𝐻2= −𝑙𝑜𝑔 1 𝑛2∑ ∑ 𝐺(𝜀𝑖− 𝜀𝑖, 𝜎2𝐼) 𝑛 𝑗=1 𝑛 𝑖=1 . (3) Correntropy is a measure of similarity between the two distributions [4]. In SE problem these distributions are estimated and measured values. Its expression is 𝑉 (𝑋, 𝑌)=1 𝑛∑𝐺(𝑥𝑖− 𝑦𝑖,𝜎2𝐼)=1 𝑛∑𝐺(𝜀𝑖,𝜎2𝐼) 𝑛 𝑗=1 𝑛 𝑗=1 , (4) where n is the size of the distributions, εi is the residuals for the component i and σ can be interpreted as the PW size. III. SENSORY FUSION The SE problem has the aim of minimize the estimation error. Once true values of the parameters are unknown, also the true value of the error is. What is available is the residual, the difference between the estimation and the measured values, given by 𝜀𝑖= 𝑧𝑖− 𝑧𝑖. The optimization of a function of the residual can be done through the minimization of its square error or RQE or by the maximization of the Correntropy, maximum Correntropy criterion (MCC). Let’s assume an expression that evaluates a set of measures called fitness. The fitness evaluation can be done, among other options, by one of the three presented criteria, LSE, RQE or MCC. In measurement systems with two sensory systems, it is possible to pay more attention to one system or to the other. 3 email: [email protected] Sensory fusion applied to power system state estimation considering information theory concepts Bruna Tavares1, Vladimiro Miranda2, Fellow, IEEE, Jorge Pereira3 T
111 This can be done assigning trust levels to the fitness function of each sensory system, fit1 and fit2. The fitness evaluation of the global system can be given by 𝑓𝑖𝑡 = 𝛼𝑓𝑖𝑡1+ 𝛽𝑓𝑖𝑡2, (4) where α represents the trust assigned to System 1 and β=1-α. The variation of the parameters α and β will influence the compromise reached. Thus, when α is 1 only System 1 is taken in consideration and its fitness evaluation reaches its maximum; at same time, the System 2 reaches its minimum. As α varies, the fitness varies gradually. Figure 1Pareto front obtained by α variation, represented by black circles. The red dot represent the point obtained by the simple fusion. The fitness function utilized was MCC, for the other metrics the form is equivalent. As figure 1 evidences, there are a lot of non-dominated points different from the one obtained with the simple fusion. The question is: Which point is more close to the real state of the system? A way to obtain that point would by assigning stochastic values to the parameters α and β. To find those values, more studies are required, with research about the probability of each sensory system in leading to the real point. Another way would be to choose the point closest to the ideal. The ideal point is the given by the coordinates (max fit1, max fit2). The metric found most convenient to measure the distance between both points was the absolute norm. L1 norm shows interesting properties: it’s easily converted to the stochastic fusion and also its simplification doesn’t require a pre-calculation of the point (max fit1, max fit2). The fusion objective functions for the SE are following presented, for LSE, MCC and RQE, respectively: min (∑ (𝑧𝑖−𝑧𝑖)2 𝑖𝜖𝑛1 +∑ (𝑧𝑖−𝑧𝑖)2 𝑖𝜖𝑛2 ), (5) max (∑𝑒−(𝑧𝑖−𝑧 𝑖)2 2𝜎2 𝑖𝜖𝑛1 +∑𝑒−(𝑧𝑖−𝑧 𝑖)2 2𝜎2 𝑖𝜖𝑛2 ), (6) max( ∑ ∑ 𝑒−(𝜀 𝑖−𝜀 𝑗)2 2𝜎2+ 𝑗𝜖𝑛1𝑖𝜖𝑛1 ∑ ∑ 𝑒−(𝜀 𝑖−𝜀 𝑗)2 2𝜎2 𝑗𝜖𝑛2𝑖𝜖𝑛2 ). (7) IV. APPLICATION IN AC SYSTEM A test has been conducted with a typical European medium voltage network, adapted from [5]. A measurement set was provided from a power flow solution and a gross error was introduced in bus 6 by eliminating its active power injection from 1.45 p.u. to 0 p.u.. The SE problem was solved with assistance of the EPSO algorithm. In this way, all the different objective functions (5), (6) and (7) were tested, with PW size 0.1. The residuals obtained, as well as their distribution, are exhibited in figure 2. While the LSE criterion led to high residuals in a large number of components, the MCC and RQE criteria provided an estimation with no significant residuals, except the residuals in Pinj6, reaching 1.45 p.u.. That residual shows that MCC and RQE criteria were able to identify and correct the outlier, finding the true value of that measurement. The residual pdf for MCC and RQE approaches a Gaussian function centered at 0 and a peak on the value of the gross error. Thus, the outlier does not influence the estimation. The same is not true to LSE were the Gaussian gets distorted in the direction of the outlier, which means that the gross error not only was not identify but corrupted the estimation of other components. Figure 2Top: abs. residuals (in p.u.) obtained by each method. The x-axis represents residual values in p.u.. Bottom: pdf of residuals obtained by LSE, MCC and RQE. The pdf was obtained with assistance of Parzen windows method with size σ = 0.1, Eq.(1). V. CONCLUSION This paper provides experimental results as proof of concept to the adoption of a new perspective in SE. First, it proposes a multi-criteria framework to achieve sensory fusion when in the presence of conventional and PMU devices. Second, it proposes to replace the traditional LSE by information theory based criteria, MCC and RQE – in this way the outliers are naturally isolated and ignored. The work shows that such adoption can be made, with positive results, in the context of sensory fusion. The challenge now is to develop an efficient algorithm for realtime use, taking advantage of these properties REFERENCES [1] V. Miranda, A. Santos, and J. Pereira, “State Estimation Based on Correntropy: A Proof of Concept,” IEEE Trans. Power Syst., vol.24, no.4, pp.1888,1889, Nov. 2009. [2] E. Parzen, “On the estimation of a probability density function and the mode”, Annals Math. Statistics, v. 33, 1962, p. 1065. [3] Principe, Jose C. Information theoretic learning: Renyi's entropy and kernel perspectives. Springer Science & Business Media, 2010. Chapter 2. [4] Liu, W.F., P.P. Pokharel, and J.C. Principe, Correntropy: A localized similarity measure. 2006 IEEE International Joint Conference on Neural Network Proceedings, Vols 1-10, 2006: p. 4919-4924. [5] Strunz, Kai, N. Hatziargyriou, and C. Andrieu. "Benchmark systems for network integration of renewable and distributed energy resources." Cigre Task Force C6 (2014): 04-02, pp33-37. 0.0 0.5 1.0 1.5 Pinj 100 Pinj 1 Pinj 3 Pinj 4 Pinj 6 Pinj 7 Pinj 9 Pinj 10 Pinj 11 Qinj 100 Qinj 1 Qinj 3 Qinj 4 Qinj 6 Qinj 7 Qinj 9 Qinj 10 Qinj 11 P34 P67 Q34 Q67 I2 (re) I5 (re) I8 (re) I2(i) I5(i) I8(i) LSE MCC RQE