Full text
A data-driven topology identification method for low-voltage distribution networks based on the wavelet transform Sebasti´ an García a,* , Matteo Fresia b , J.M. Mora-Merch´ an a , Alejandro Carrasco a , Enrique Personal a , Carlos Le´ on a a Electronic Technology Department, Escuela Polit´ ecnica Superior, University of Seville, Seville 41011, Spain b Department of Electrical, Electronic, Telecommunications Engineering and Naval Architecture, University of Genoa, Genova 16145, Italy ARTICLE INFO Keywords: Distribution networks Smart meters Data analytics Topology Identification ABSTRACT A comprehensive knowledge of topology is of great importance for the effective operation and maintenance of distribution networks. This paper contributes with a novel data-driven topology identification method for lowvoltage distribution networks based on the wavelet transform. The method uses only energy measurements from smart meters, being compatible with the current European smart meter capabilities. The method identifies the feeder and phase topology of single and three-phase customers, even in unbalanced situations. A computationally-efficient methodology to link customers’ time-frequency features with their network connection is proposed. The performance of the method is assessed on eleven non-synthetic networks, with a robustness assessment of factors such as network observability, dataset size, measurement errors, and Renewable Energy Sources (RES) penetration. Accuracy rates exceeding 95 % are obtained in most cases, outperforming an energyconservation approach. A 98 % accuracy can be achieved with a 30-day hourly dataset if at least 80 % of network observability is provided. For lower observability levels, 45 or 60 days of data are needed to reach similar rates. The sensitivity analysis of measurement error demonstrated that it had a negligible influence on the results. The method showed favorable results even in scenarios with high-RES penetration, with accuracy values exceeding 95 %. 1. Introduction The growing integration of Renewable Energy Sources (RES) into low-voltage Secondary Distribution Networks (SDN) at customer level offers a potential path towards the reduction of energy costs while enabling consumer’s self-sufficiency and enhancing grid resilience, contingent upon effective management [1]. However, while this paradigm offers new opportunities, it also poses complex management challenges for Distribution System Operators (DSO) [2]. In this scenario, it is crucial for DSOs to have a comprehensive awareness of the precise location of low-voltage customers within the distribution network, with a clear knowledge of their topological connection [3]. Knowing the topology connection of customers is of great importance for the effective operation and maintenance of distribution networks, as it enables the mitigation of feeders unbalance [4], the location and classification of faults [5], as well as for state estimation and network reconfiguration to reduce losses [6]. However, secondary distribution have traditionally suffered a lack of documentation compared to transmission and primary distribution networks [7], which frequently results in DSOs being unaware of customer topology location [8]. Although this information is typically recorded upon customer’s connection to the grid, databases are not always reliable due to a number of factors [9]. These include errors when the connection was registered, grid reconfigurations following faults or maintenance operations where changes have been made in users’ connections but have not been documented, unnotified changes or even grids that have never been documented [7,9]. In terms of network observability, it is noteworthy that SDNs remain less monitored than transmission and primary distribution networks. However, European DSOs have made a significant effort over the last decade to increase grid observability in SDN [10], with the deployment of millions of smart meters at customer level [11,12] and at the feeder cables at the head of the Secondary Distribution Substation [13]. This deployment has facilitated the development of novel applications, * Corresponding author. E-mail address: [email protected] (S. García). Contents lists available at ScienceDirect Electric Power Systems Research journal homepage: www.elsevier.com/locate/epsr https://doi.org/10.1016/j.epsr.2025.111517 Received 4 December 2024; Received in revised form 4 February 2025; Accepted 5 February 2025 Electric Power Systems Research 243 (2025) 111517 Available online 16 February 2025 0378-7796/© 2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/bync-nd/4.0/ ).
techniques, and methodologies based on data analytics [14], which have been made possible by the data generated by smart meters [15]. However, despite the considerable enhancement that smart meters have brought to European SDN, they do possess certain shortcomings compared to those at transmission network or at primary distribution. In this regard, the primary drawback of SDN smart meters is their lack of synchronization [16], which is attributed to the discrepancies in the internal real-time clocks of devices and the inherent delays [17], packet losses, and signal interferences associated with the Power Line Communication (PLC) channel employed by DSOs [18,19]. In this context, European smart meters are not designed for continuous synchronized voltage sampling, the main purpose of these devices is the measurement of energy consumption for billing purposes [20], which is typically measured in hourly intervals [21], with any minor discrepancies in timing considered sufficiently insignificant to be disregarded. In this sense, smart meters are continuously measuring energy consumption, storing the readings in their internal memory, and generally transmit all the readings stored since the previous transmission once or twice per day [16]. Considering the current situation, this paper proposes a novel datadriven topology identification method using the data provided by smart meters. The method is based on the wavelet transform and it is intended for the identification of feeder and phase topology of customers in European-like low-voltage distribution networks. 1.1. Literature review In the existing literature, several papers propose the use of datadriven methods to identify distribution network topology. In [22] and [23], a probabilistic graphical-based model using bus voltage measurements is proposed. A method employing an unsupervised graph-based theory model based on voltage data is presented in [24]. Supervised machine learning techniques using voltage data are used in [25] to detect the status of switching devices and to identify the network topology. In [26], a method based on the analysis of active power and voltage profiles obtained from end-customers smart meters and from intermediate nodes. Correlation between voltage profiles collected by IoT devices are used in [27] to identify the topology of a low-voltage distribution network. In reference [28], a linear regression topology identification method is proposed that utilizes voltage magnitude, active power, and reactive power as inputs. A heuristic-based method is proposed in [29], which requires voltage and current profiles as well as power factor measured by end-user meters. In [30], a Gaussian mixture model with voltage data is used to identify transformer-customer relationship in low-voltage distribution networks. While some of the previously referenced papers presented methods for identifying the topology of low-voltage SDNs, these proposals are not currently viable given the limitations of the European smart meter deployment. This is due to the fact that they are unable to provide continuous voltage sampling. In practice, the functionality of these meters in terms of voltage measurement is typically limited to the storage of specific voltage events, which are then used in order to assess the quality of the service provided by the DSO [19]. In this context, some alternative methodologies rely exclusively on the energy consumption profiles obtained from smart meters. For example, [31] proposes a stepwise regression method to identify phase connection using the energy conservation law. A similar concept but using LASSO regression is proposed in [32] and [33]. Reference [34] also employs the energy conservation law, formulating the problem as a weighted least-square based optimization. An approach using the Kalman filters to identify phase topology is proposed in [35]. Despite these methods use the data provided by European smart meters, their application of the energy conservation law results in a limitation: their performance decreases as the unmeasured consumptions increases. These unmeasured consumptions may be due to technical and non-technical losses (i.e. fraud), measurement errors or even customers with legacy meters. To overcome this situation, other methods use spectral information. For example, [36] and [37] use an spectral and saliency analysis to identify phase topology. Similarly, the authors of [38] utilize a modified k-means clustering algorithm with high-frequency features. In [39] the authors propose a method based on Bayesian inference using variations in energy consumption. However, these approaches are focused on the identification of user phase connectivity, rather than the discernment of feeder topology relations. 1.2. Contributions A review of the literature reveals that numerous methods for identifying topology in low-voltage networks have been proposed. However, the input requirements for some of these methods make their actual application in European low-voltage networks technically infeasible. Furthermore, some of these methods do not consider the possibility of single and three-phase connections to the grid, using single-phase equivalent models. Other methods that are compatible with the current technical capabilities of European smart meters are limited to the identification of phase connectivity, rather than discernment of feeder topology. In this context, spectral analysis has proved to be a promising tool for establishing a correlation between a customer’s consumption profile and the connection to which they are affiliated. However, the wavelet transform provides more detailed information, comprising both time and frequency and, to the best of the authors’ knowledge, it has not been used before for topology identification purposes. Furthermore, the papers on this topic lack a comprehensive examination of the robustness of the proposed methodologies under conditions of grid observability, dataset size, measurement error, and DER penetration. In this context, this paper proposes a novel data-driven topology identification method for low-voltage distribution networks, employing a multi-resolution analysis technique: the wavelet transform. The wavelet transform is used to exploit the inherent relation in timefrequency characteristics between the consumption profile measured at the feeder to which a user is connected and their own consumption profile. The proposed approach relies exclusively on energy measurement data extracted from customers’ smart meters and at the feeder head of the substation. Using the wavelet transform, the time-frequency characteristics of the energy consumption profiles are extracted. These features are then used by the method to identify the topology of both single and three-phase customers. The methodology employs an exhaustive procedure to link customers’ features with their point of connection, identifying even those customers with less representative features. The effectiveness of the method is evaluated using a total of eleven non-synthetic European SDNs. In this sense, the method is validated through a robustness analysis conducted under different levels of grid observability, dataset size, measurement errors, and RES presence. The main contributions of the present paper are represented by: •Proposal of a novel wavelet-based topology identification method for low-voltage distribution networks that is compatible with the current measurement capabilities of the European smart meter deployment. •A comprehensive methodology for processing the multi-resolution time-frequency features extracted by means of the wavelet transform is presented. The features of customers are used to determine the topology location to both single and three-phase customers, even in unbalanced situations. A similarity score and a dedicated processing sequence are defined to enhance the identification of even those with less representative features while reducing the computational time. •Evaluation of the proposed method with eleven non-synthetic SDN using real smart meter data. The method outperforms the energy conservation approach while demonstrating robustness to low observability of the network (unmetered loads), network size, measurement errors and high DER penetration. Moreover, the method proved to be effective in all scenarios even when using limited S. García et al. Electric Power Systems Research 243 (2025) 111517 2
datasets (a maximum of 60 days of hourly data are needed in a worstcase scenario). The rest of the paper is organized as follows: Section 2 offers a mathematical definition of the problem. In Section 3, the topology identification method based on the wavelet transform is presented. Section 4 presents and discusses the results. Finally, Section 5 is dedicated to conclusions and also provides some possible further developments of the study. 2. Problem definition Let us consider a SDN (low-voltage) that is supplied by a substation with NF four-wire three-phase feeders, as illustrated in Fig. 1. Let us suppose that the grid fed by the substation has a total of NG customers, of which NS have a single-phase connection (phase to neutral) and NT have a three-phase connection to the grid. Thus, NG=NS+NT. The problem consists in identifying the topology connection of users on the grid. In other words, to determine which feeder and phase a user is connected to (in the case of three-phase users, only the feeder connection is necessary). Therefore, it is a complex combinatorial classification problem, with ((3NF)NS+NNT F)potential solutions. The complexity of the problem can be illustrated by reference to a typical SDN with five feeders, supplying 300 users, of whom 50 have three-phase connections. For this network, the number of potential solutions is 1.05 ×10294. In order to facilitate the subsequent analysis, it is necessary to define the sets F and P, as illustrated in (1) and (2), which represent the indexes of the feeders and the phases respectively. Furthermore, the sets CS and CT, as shown in (3) and (4), represent the indexes of the singlephase and three-phase users, respectively. Finally, the set CG presented in (5) represents the indexes of all customers connected to grid fed by the substation. F= {1,2,⋯,NF}(1) P= {R,S,T}(2) CS= {1,2,⋯,NS}(3) CT= {NS+1,NS+2,⋯,NS+NT}(4) CG=CS∪CT= {1,…,NS,NS+1,…,NG}(5) Let the vector Cn of length M, as defined in Eq. (6), be the energy consumption profile measured by the smart meter of customer n, where cn,m is the incremental energy consumption of the customer between instants m and m−1. As three-phase smart meters typically report the total sum of customer consumption without providing a disaggregation by phase, the vector Cn of a three-phase user (i.e., n∈CT) represents the total aggregated customer’s consumption across all phases for that user. Cn=[cn,1,cn,2, ..., cn,m, ..., cn,M],∀n∈CG(6) For better management, the consumption profiles of single-phase and three-phase customers are grouped together in matrices CS and CT, respectively. Both matrices are shown in Eqs. (7) and (8). The two matrices can be stacked together to form the matrix CG, as shown in (9), which represents the consumption profiles of all the customers of the grid. CS=⎡ ⎣ C1 ⋮ CNS ⎤ ⎦(NS×M) (7) CT=⎡ ⎣ CNS+1 ⋮ CNG ⎤ ⎦(NT×M) (8) CG=[CS CT](NG×M) (9) In a similar way, let the vector Fkp of length M shown in Eq. (10) represent the energy consumption profile of feeder k and phase p measured at the head of the secondary distribution by the feeder supervisor (see Fig. 1). In this context, the element fkp,m of vector Fkp is the energy consumption of feeder k and phase p between instants m and m− 1. Fkp =[fkp,1,fkp,2, ..., fkp,m, ..., fkp,M],∀k∈F,∀p∈P(10) The consumption profiles of all the phases of the k-th feeder are encoded in the matrix Fk shown in (11). The matrix FG illustrated in (12) represents the complete set of consumption profiles observed among the grid’s feeders. Fk=⎡ ⎢ ⎢ ⎣ FkR FkS FkT ⎤ ⎥ ⎥ ⎦(3×M) ,∀k∈F(11) FG=⎡ ⎢ ⎢ ⎣ F1 ⋮ FNF ⎤ ⎥ ⎥ ⎦(3NF×M) (12) Let xn,kp be a binary variable indicating that user n is connected to feeder k and phase p. In the case of single-phase users, equality (13) must be satisfied, indicating that the user is connected at a single point on the grid. For three-phase users, equalities (14) to (16) must hold, guaranteeing that users are connected to the three phases of the same feeder. ∑ k∈F∑ p∈P xn,kp =1,∀n∈CS(13) ∑ k∈F∑ p∈P xn,kp =3,∀n∈CT(14) xn,kR =xn,kS,∀k∈F,∀n∈CT(15) xn,kS =xn,kT,∀k∈F,∀n∈CT(16) Fig. 1. Schematic representation of a Secondary Distribution Substation with NF four-wire feeders. S. García et al. Electric Power Systems Research 243 (2025) 111517 3
For an easy management, let Xk be defined as the connectivity matrix of customers to feeder k, as illustrated in (17). If the matrices of all feeders are grouped together as shown in Eq. (18), the resulting matrix XG represents the connection of all customers to the grid. Thus, the objective is to determine this matrix. Xk=⎡ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ x1,kR x1,kS x1,kT x2,kR x2,kS x2,kT ⋮ ⋮ ⋮ xNG,kR xNG,kS xNG,kT ⎤ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦(NG×3) ,∀k∈F(17) XG= [X1X2⋯XNF]T (3NF×NG)(18) In conclusion, the aim is to determine the connectivity matrix XG, given matrices CG and FG, which are derived from the data obtained through the Advanced Metering Infrastructure of the DSO. In the existing literature, the classical approach to identifying the topology of low-voltage users is to consider the energy-conservation law and formulate the problem as (19). This approach yields a connectivity matrix XG which minimizes the discrepancy between the energy delivered by the substation’s feeders and the energy exchanged with the network by the users. In Eq. (19), the symbol ⊙represents the Hadamard product and W is a matrix of size 3NF×NG (same dimensions and structure as XG). The first NF columns of W correspond to single-phase users and are set to 1 in this matrix, while the remaining columns, which correspond to three-phase users, are set to 1/3 (if a balanced distribution among phases is assumed). min XG‖FG− (XG⊙W)CG‖1s.t.(13)− (16)(19) The main disadvantage of this approach is that its performance decreases as the unmeasured consumption increases. In other words, the assumption underlying (19), whereby the sum of the individual energy measurements of the users is equal to the energy supplied by the substations’ feeders, inevitably gives rise to a problem due to the presence of multiple sources of unmeasured consumption within the network. In this context, there are several sources of unmeasured consumption in secondary distribution networks: losses in the cables, non-technical losses (i.e., fraud), measurement errors (e.g., due to instrument tolerance or failures), customers with legacy energy meters that are not monitored, etc. In addition to this disadvantage, three-phase users in low-voltage networks usually have an unbalanced consumption. As a result, the columns corresponding to these users in W are unknown, thereby increasing the complexity of the problem. Thus, the method proposed in this paper does not use (19) or any of its variations. To address these limitations, this paper proposes a novel waveletbased method that exploits the inherent relation in time-frequency characteristics between the consumption profile measured at the feeder to which a user is connected and their own consumption profile. 3. Wavelet-based topology identification approach This section introduces a brief description of the wavelet transform, after which the proposed topology identification method is presented. 3.1. Wavelet transform The Wavelet Transform (WT) is a multi-resolution analysis technique that provides a time-frequency representation of a given signal [40]. In contrast to classical spectral transformations, such the Fourier Transform (FT), which provides a high-resolution representation of signals in the frequency domain, the WT is capable of simultaneously obtaining a representation of local spectral and temporal information of a given signal. In other words, the WT provides information about spectral components at a given instant of time. The rationale behind the use of a wavelet-based analysis for the topology identification problem is that the consumption pattern of a specific customer has a direct impact on the feeder to which it is connected. In other words, the time-frequency features of the customer are also present on the feeder to which it is connected. In this context, where often times interesting features of signals are non-stationary or transitory characteristics, multi-resolution time-frequency analysis as the WT becomes an ideal tool. The WT expresses a given signal into a linear combination of a particular set of functions obtained by translating and scaling another function called mother wavelet. A wavelet is an oscillatory function of finite length. Some examples of wavelets functions are the Haar, Meyer and the Daubechies [40]. In this context, the Continuous Wavelet Transform (CWT) can be defined as (20): CWT ψ ( τ ,s) x=1 |s| √∫x(t) ψ *(t− τ s)dt (20) where x(t)is the input signal, ψ *(t)is the complex conjugate of the mother wavelet ψ (t), and τ and s are the translation and scale parameters, respectively. The translation parameter τ serves to modify the location of the wavelet along the time axis while the scale parameter s dilates (if s>1) or comprises (if s<1) the wavelet signal. In this context, larger values of the scale parameter facilitate the analysis of the low-frequency response of the signal, whereas smaller values are utilized to analyze the high-frequencies. The CWT may be understood as a means of quantifying the correlation between a given wavelet (at different scales) and a given time-varying signal. For the practical implementation, the discretized CWT is computed on a discrete grid of scales and translations. In this sense, a signal x of length M can be described into a set of MNsc coefficients, being Nsc the number of scales used. As can be seen, the CWT yields to a high number of coefficients, being a highly redundant transformation [41,42]. To overcome this situation, the Discrete Wavelet Transform (DWT) implements this transformation by means of a filter banks. In this sense, the DWT is computed by passing a discrete-time input signal x= [x1,x2, ⋯,xm,⋯,xM]through a bank of high-pass and low-pass half-band filters with impulse responses g and h, respectively. These two filters are designed to represent the mother wavelet [43]. Moreover, they are related each other and must satisfy the standard quadrature mirror filter condition. The passing of the input signal through these filters corresponds to the convolution operation, as shown in (21) and (22). yhigh m= (x*g)m=∑ k=M k=− M xkgm−k(21) ylow m= (x*h)m=∑ k=M k=− M xkhm−k(22) After the filtering of the signal, in which half of the frequencies have been removed, and in accordance with Nyquist’s sampling theorem, the signal can be subsampled by 2 without losing information. This subsampling also serves to double the scale. From a mathematical perspective, this operation can be described as (23) and (24): am= (x*g)m↓2=∑ k=M k=− M xkg2m−k(23) dm= (x*h)m↓2=∑ k=M k=− M xkh2m−k(24) where ↓ denotes the subsampling operation and am and dm are the outputs of the high-pass and low-pass filters, respectively, after subsampling by 2. In the context of the DWT, the vectors of coefficients cA = [a1,a2,⋯,am,⋯,aNDWT ]and cD = [d1,d2,⋯,dm,⋯,dNDWT ]are referred to as the approximation and detail coefficients, respectively. The length of these vectors are given by NDWT =⌈(M+Nfilter −1)/2⌉, where Nfilter S. García et al. Electric Power Systems Research 243 (2025) 111517 4
is the length of the filters and M the length of the input signal. This operation of obtaining the approximation and detail coefficients, also known as sub-band coding, can be repeated for further decomposition of the signal by applying the same procedure to the approximation coefficients, thereby increasing the frequency resolution, as illustrated in Fig. 2. Thus, a signal can be decomposed into a maximum of LDWT max =⌊log2(M/(Nfilter −1))⌋levels, to ensure that at least one coefficient remains uncorrupted by the edge effects caused by signal extension during filtering. In other words, the decomposition stops when the length of the input signal is shorter than the filter length. In summary, after passing the input signal through the filter bank, there are LDWT max detail coefficients vectors, which represent the high timefrequency features of the input signal, and one vector of approximation coefficients, which represent the low time-frequency features. In most cases, the approximation coefficients are disregarded. Thus, following a DWT analysis of a given signal, a vector of features is obtained comprising the concatenation of all detail coefficients levels, as shown in (25). It should be noted that in certain circumstances, not all the possible levels are employed for the decomposition of the input signal, as at certain decomposition level the additional information gained is minimal. In regard to the topology identification, we have found that the performance of the method does not exhibit a significant increase when more than two levels of decomposition are employed. DWT(x) = [cD1cD2⋯cDLDWT max ](25) Coming back to the topology identification problem, and as previously stated: the consumption pattern of a customer directly affects the feeder to which it is connected. Thus, if the features of the consumption profile of a user are extracted by means of the DWT, it is possible to determine the customer’s position within the network through an examination of the feeder consumption profile that exhibits comparable attributes on their DWT features. To illustrate the rationale behind the use of a wavelet-based analysis, consider the following simple example. If a user experiences a significant change in their consumption profile, which can be differentiated from the profiles of other users, that change will also be present in the feeder consumption profile to which the user is connected. In this context, the advantage of employing a multiresolution technique, as the wavelet-based analysis, is that it not only identifies the primary features of the consumption profile, but also the specific point in time at which these features emerged. This allows for differentiation of these features in the feeders’ consumption profile from those of other users. Of course, this analysis is not a straightforward process; rather, it requires a complex method to identify even those customers whose features are particularly diluted in the feeder consumption profile. In this context, the following section provides a comprehensive description of the wavelet-based topology identification method. 3.2. Topology identification method The topology identification method is described in this subsection. A flowchart summarizing the method’s steps is provided in Fig. 3. As previously stated, the method just requires the energy measurements from smart meters and the energy measurements at each of the feeders of the secondary distribution network from the AMI infrastructure of the DSO. With these measurements, the CG (9) and FG (12) matrixes are constructed. The connectivity matrix XG is completely unknown and its entries are all zero. Moreover, two empty sets IS and IT are created. These sets will comprise the indexes of single-phase customers whose topology have been identified. As stated along the text, three-phase smart meters report the aggregated energy consumption of the user among all the phases, without providing disaggregation by phase. Additionally, in low-voltage networks, these connections are typically unbalanced. Thus, it is imperative to distinguish between the analysis of three-phase and single-phase customers in the method. In this context, the method analyzes first single-phase users, identifying their topology connection. Thereafter, the consumption at the feeders is aggregated, summing the consumption of all phases, to analyze the three-phase users. The left-hand side of the flowchart of Fig. 3 is dedicated to the processing of single-phase customers, while the right side (starting from the asterisk node) corresponds to the processing of three-phase customers. The first step of the method is to extract the features from the consumption profiles of users and feeders by means of the DWT. In this paper, the Haar wavelet has been employed. The rows of the CG matrix are associated with the consumption profiles of individual users, while the rows of the FG matrix are linked to the consumption profiles of specific phases within a given feeder of the network. Thus, the DWT is computed row-wise over these two matrices, as shown in Eqs. (26)-(30), resulting in the matrixes CDWT G (28) and FDWT G (30). These two matrixes comprise the features of customers and feeders respectively. A copy of the matrix FDWT G and its content is stored in a new matrix FCDWT G, which will be the matrix used by the method, while the original matrix FDWT G is not modified. The purpose of this last operation will be described later. CDWT n=DWT(Cn),∀n∈CG(26) FDWT kp =DWT(Fkp),∀k∈F,∀p∈P(27) CDWT G=⎡ ⎢ ⎢ ⎣ CDWT 1 ⋮ CDWT NG ⎤ ⎥ ⎥ ⎦(NG×MDWT) (28) FDWT k=⎡ ⎢ ⎢ ⎣ FDWT kR FDWT kS FDWT kT ⎤ ⎥ ⎥ ⎦(3×MDWT) ,∀k∈F(29) FDWT G=⎡ ⎢ ⎢ ⎣ FDWT 1 ⋮ FDWT NF ⎤ ⎥ ⎥ ⎦(3NF×MDWT) (30) After extracting the features, the next step is to compute the similarity level between the features of single-phase customers and the feeders of the secondary distribution network. The relationship between Fig. 2. Obtaining the approximation (cA) and detail (cD) coefficients of the Discrete Wavelet Transform using sub-band coding. S. García et al. Electric Power Systems Research 243 (2025) 111517 5
the customer’s features and those of the feeder to which it is connected is suggested to be linear. This is based on the premise that the measured consumption at the feeder is the summation of all customers connected to that feeder, and that the DWT is linear and time-invariant operation. In this context, the Pearson Correlation Coefficient (PCC), which is a measure of linear correlation, has been selected to compute the similarity level between the customers’ features and those at the feeders. There are other time-series metrics to measure similarity, such as Dynamic Time Warping (DTW), which is an effective similarity metric when time scales between signals are not well adjusted, or the sampling periods of signals are different or irregular. However, the applicability of the DTW is limited in our application since our method uses hourly energy data collected by smart meters, where seconds’ discrepancies are negligible. Furthermore, the algorithm of the DTW has a complexity of O(M1M2), where M1 and M2 are the length of the signals to be evaluated. In contrast, the PCC is a much more computationally efficient method, as it is just a simple algebraic operation. Thus, the similarity level is obtained by computing the Pearson Correlation Coefficient (PCC) between the CDWT n features of a user n∈CS and the features at each of the feeders FCDWT kp ∀k∈F,∀p∈P. The PCC between two signals x and y is shown in (31), where M is the number of measurements in the signals. This process yields a similarity vector Ωn of size 3NF for each customer, as shown in Eq. (32). The PCC values can range from 1 to −1, with 1 indicating perfect correlation and −1 indicating an inverse correlation. r(x,y) = M∑xiyi−∑xi∑yi M∑x2 i− (∑xi)2 √ M∑y2 i− (∑yi)2 √(31) Fig. 3. Flowchart illustrating the proposed topology identification method. S. García et al. Electric Power Systems Research 243 (2025) 111517 6
One straightforward methodology for determining the location of customers would be to ascertain which feeder exhibits the greatest degree of similarity. However, this approach could potentially result in misclassification, as the features of the customers may be very diluted in the feeder consumption profile, rendering them indistinguishable. To overcome this situation, customers that are more likely to be correctly identified are analyzed first. In order to identify the customers that are more likely to be correctly identified, a similarity score metric has been defined. This score scn for a customer n is computed as the maximum value obtained in the correlation analysis in (32) plus the Euclidean distance with the second highest value obtained, as shown in shown in (33): scn=max(Ωn)+ (max(Ωn)− max2(Ωn))2 √(33) where max2(Ωn)denotes the second-highest value in the Ωn vector. Therefore, the similarity score serves to indicate whether the customer features exhibit a strong similarity with those of the feeder, with an award conferred based on the distance to the feeder exhibiting the second-highest similarity. Thus, if the similarity score is computed for all single-phase users, the customer n* S obtained as (34), is the one that is more likely to be correctly identified. n* S=argmax n{scn:n∈CS}(34) At this moment, the topology location of user n* S is updated in the connectivity matrix XG. As the rows of this matrix represent feederphase locations and the columns represent users, the [ξS,n* S]element of matrix XG is set to one, denoting the user’s connection, with ξS calculated as follows: ξS=3⌊(argmax x(Ωn* S[x]))/3⌋+⌊(argmax x(Ωn* S[x]))% 3⌋(35) where Ωn* S[x]denotes the x-th position in vector (32) of user n* S and % denotes modulus operation in this case. In addition, the user n* S is included into a set IS which comprises the indexes of single-phase customers whose connections have been located. This can be expressed as IS←IS∪{n* S}. After identifying the customer, the following operation is performed to overcome the possible situation in which some customers’ features may be diluted in the feeder consumption profile, rendering them indistinguishable. Thus, to suppress this influence while reducing the dilution of the remaining customers’ features in the feeders, the consumption profile of the already located customer is subtracted from its feeder. As the DWT coefficients are obtained by means of the convolution operation, as it was shown in (23) and (24), which is a linear and time-invariant (LTI) operation, it is possible to subtract the already located customers’ features from its feeder. Thereby avoiding the need of calculating again the DWT over the feeders’ consumption profiles. From a mathematical perspective, this operation is described as follows: FCDWT G=FDWT G−XGCDWT G(36) As previously stated, this operation serves to reduce the dilution of features in the feeders of customers that have not yet been located, thereby facilitating their subsequent identification. Now, the purpose of storing a copy of the FDWT G matrix can be elucidated: the FCDWT G serves as a working copy for the method in which the already identified customers are removed. At this point, the described process is repeated until all single-phase users are identified within the network topology (see the flowchart in Fig. 3). This is, the vectors Ωn∀n∈ {CS\IS}are updated to consider the new FCDWT G features, the similarity score is then recalculated scn∀n∈ {CS\IS}and the customer with the highest similarity score is identified and located within the network topology. Then, the connection is updated in the XG, matrix and the new features FCDWT G are obtained. This cycle is repeated until all single-phase users are located. As the cycle progresses, the features of customers become increasingly distinct from those of the feeders, thereby facilitating the process of identification. Once all single-phase users have been identified, a change in the method is introduced to enable the identification of three-phase users, as they typically provide the total aggregated consumption, without disaggregation by phase. Thus, for three-phase users, the similarity vector Ωn∀n∈CT is obtained in a manner that differs from single-phase users. Specifically, the vector Ωn is obtained as illustrated in (37). The distinction from the vector for single-phase users presented in (29) is that, for each feeder, the features of all phases are aggregated. Thus, the vector Ωn for three-phase users have a reduced length of NF values. This aggregation of features is possible due to the LTI nature of the DWT implementation. The subsequent process for identifying the topology of three-phase customers is similar to that previously described for single-phase customers. In this sense, the similarity score scn∀n∈CT is obtained as was described in (33). The customer n* T with the highest similarity score is identified, as shown in (38). The topology connection of customer n* T is obtained as illustrated in (39), in which ξT represent the feeder in which the three-phase user is located. The connection of customer n* T is updated in the connectivity matrix XG, setting to one the elements [ξT+i,n* T]∀i∈ {0,1,2}(i.e. the three phases of the feeder) of this matrix. The user is also added to a set IT representing the already found three-phase customers: IT←IT∪{n* T}. Once the customer’s topology has been identified, its features can be removed from those of the feeders, as illustrated in (40) and recalling that W is a matrix of equal dimensions as XG in which the first NF columns are ones and the remaining are set to 1/3. For three phase users it is necessary to include the W matrix in order to distribute the consumption of the user among the phases of the feeder. In this case, it is irrelevant to assume a balanced Ωn=[r(CDWT n,FCCWT 1R),r(CDWT n,FCCWT 1S),r(CDWT n,FCCWT 1T),⋯,r(CDWT n,FCCWT NFR),r(CDWT n,FCCWT NFS),r(CDWT n,FCCWT NFT)],∀n∈CS(32) Ωn=[r(CDWT n,∑ p∈P FCDWT 1p),r(CDWT n,∑ p∈P FCDWT 2p),⋯,r(CDWT n,∑ p∈P FCDWT NFp)],∀n∈CT(37) S. García et al. Electric Power Systems Research 243 (2025) 111517 7
situation in W (i.e., the columns for three-phase users are set to 1/3 in this matrix), as the consumptions of all phases of feeders will then be regrouped to obtain the similarity vector Ωn (see Eq. (37)). Again, this regrouping is possible due to the LTI characteristics of the DWT implementation. n* T=argmax n{scn:n∈CT}(38) ξT=argmax x(Ωn* T[x])(39) FCDWT G=FDWT G− (XG⊙W)CDWT G(40) This process is repeated until all three-phase customers are identified within the network topology (see Fig. 3). As well as with single-phase users, those three-phase users that have already been identified are not considered further. Thus, the next similarity vector and similarity scores are obtained just for the remaining yet-to-be-identified customers (i.e. ∀n∈ {CT\IT}), as shown in the flowchart of Fig. 3. The process ends when all the three-phase customers are topologically identified. 4. Results This section describes the SDNs used to evaluate the proposed topology identification method and the results after applying the method in different test scenarios. The method robustness is evaluated under different sizes of dataset, different levels of grid observability, measurement errors and RES penetration. Moreover, the method is compared with an energy conservation approach. 4.1. Use case European SDNs are four-wire networks with a voltage of 400/230 V [44]. These networks utilize a delta-wye power transformer with the neutral grounded and usually supplies between 200 and 400 customers [45]. In order to evaluate the method in such an environment, a real (non-synthetic) European distribution network has been used. The network is described in [46], in which the authors provide a comprehensive details of a real distribution network from northern Spain. The grid is composed by 30 SDNs fed by a primary distribution network (medium-voltage). The authors of the previously cited paper also provide the model of the network in [47] in OpenDSS format [48]. Of this network, eleven secondary distribution networks are used to evaluate the proposed topology identification method. A summary of the characteristics of the distribution networks is presented in Table 1: total number of customers (NG), the ratio of three-phase users to total customers (NT/NG), the number of feeders (NF) and the mean, median, minimum and maximum number of customers in feeders. As illustrated, the networks are arranged in ascending order based on the total number of customers. Additionally, a column with an ID is included to facilitate subsequent reference to each of the SDNs. The customer consumption profiles used are from smart meters belonging to real customers (anonymized). Specifically, the data used for the experiments has been extracted from the Advanced Metering Infrastructure of Medina Garvey, a Spanish DSO. The data from the smart meters has hourly periodicity. A total of 60 days have been used for the analysis. The power flow in the networks was simulated using the OpenDSS model provided in reference [47] to obtain the active energy consumption profiles at the head of the feeders at the substation, which were then used to construct the FG matrix. Unless otherwise stated the data does not include the presence of RES. The impact of RES is assessed in a dedicated subsection of the paper. For the DWT, the impulse filters implement the Haar wavelet, and two levels of decomposition have been used. 4.2. Performance of the proposed topology identification method The proposed method has been used to identify the topology location of the customers of the SDNs described in Table 1. The results are compared with those yielded by an energy-conservation approach, as described at the end of Section 2 by Eq. (19), since this method has the same inputs and provides the same outputs as the wavelet-based method proposed in this paper. Just 60 days of hourly data have been used to perform the experiment. Different levels of network observability have been considered. In this regard, ten network observability levels have been evaluated, spanning from a scenario where 100 % of customers have a smart meter to one where only 10 % of them do (in steps of 10 %), in order to account for the possibility that measurements may not be available for some customers. Fig. 4 presents the accuracy results in the topology identification for the different observability scenarios. Lines with circle-shaped markers correspond to the results of the proposed wavelet-based method while those with triangle-shaped markers correspond to the result of the Table 1 Characteristics of the Secondary Distribution Networks used to evaluate the proposed method. Secondary Substation Name ID NGNT/NGNFMean Customers per feeder Median Customers per feeder Minimum Customers in feeder Maximum Customers in feeders CELLERUELO SDN1 152 0.14 3 50.6 31 29 92 PLAZA CUBIERTA SDN2 237 0.20 5 47.4 27 12 139 LA_SOLEDAD SDN3 259 0.13 4 64.7 66 1 125 EDIFICIO EL MARQUES SDN4 273 0.12 3 91.0 87 84 102 CALVO SOTELO SDN5 298 0.12 5 59.6 67 30 90 LA ISLA SDN6 362 0.07 5 72.4 80 28 107 PLAZA ARGUELLES SDN7 365 0.15 5 73.0 80 44 87 MARQUESILLA DE CANILLEJAS SDN8 458 0.10 5 91.6 106 37 120 VALERIANO LEON SDN9 554 0.10 7 79.1 81 17 175 LA_GUAXA SDN10 604 0.11 7 86.2 76 68 123 PARQUE LUZ SDN11 620 0.09 9 68.9 63 28 141 Fig. 4. Accuracy evaluation for the proposed wavelet-based method (with circle-shaped makers) and for the energy-conservation approach (with triangleshaped markers) for different observability scenarios. S. García et al. Electric Power Systems Research 243 (2025) 111517 8
classical energy-conservation approach. The color represents the SDN. It is important to note that the accuracy results presented here are limited to customers with a measurable consumption pattern (for observabilities <100 %), as it is assumed that there is no available data for the remainder of customers. It can be observed that the proposed method (circle-shaped markers) can identify the topology of all SDN in all observability scenarios, with accuracies higher than 95 % in all cases. Just a slight decrease in the performance is observed when network observability is poor (<30 %). The conventional approach (triangleshaped markers) yields favorable results only when network observability is high (≥90 %). For this last method, it is clear that a reduction in observability leads to an overall decline in accuracy. In this context, it is important to note that while penetration levels of smart meters in European low-voltage networks are high nowadays, resulting in good observability levels of SDNs, this evaluation is not only relevant for cases in which some customers do not have smart meters (or they have legacy meters without connectivity). This evaluation is also of interest in that it supposes the existence of non-measured consumption within the network. In this sense, it is not uncommon for non-measured consumption to occur in secondary distribution networks, which can be attributed to an increase in technical or non-technical losses. In order to evaluate the computational requirements, Table 2 presents the execution time for the proposed method and for the energybased approach. The experiment has been carried out on an Intel Xeon CPU E5–2650 v4 @ 2.20 GHz machine. It can be observed that the execution time of the proposed method increases in proportion to the number of clients in the network. Nevertheless, this time remains within an acceptable range even for large networks. In contrast, the energybased approach exhibits a significantly higher execution time compared with the proposed method, with a particularly pronounced increase observed in the case of large SDNs (SDN9 to SND11). To compare our proposal with other types of time-series similarity techniques, we also implemented a version of the method that used the DTW instead of the PCC. The results of this comparison indicated that the execution time for SDN1 (with only 130 clients) was 2610.49 s when using DTW (compared to 1.11 s of our implementation). Furthermore, the accuracy was only 71.3 %, even considering 100 % observability of the network. To further assess the scalability of the proposed method with other kinds of networks, with a higher number of customers and feeders, several synthetic networks have been constructed. Table 3 shows the execution time (in seconds) of the proposed method for networks ranging from 10 to 20 feeders and from 700 to 1000 customers. The findings indicate a negligible impact on the execution time with the quantity of feeders and a consistent increase with the number of customers. However, given that the proposed method is not intended for execution in real time, but rather for use with historic datasets, the execution times remain within reasonable limits. It is worth mentioning that the accuracy in the identification in these networks has always been higher than 97 %, proving the robustness of the method even with synthetic networks. In conclusion, the proposed method has demonstrated its efficacy in identifying the topology location of customers, even in circumstances where the observability of the network is poor (i.e., a high level of unmeasured consumption). Furthermore, it has been shown to be computationally efficient, requiring <24 s in the most unfavorable European-like scenario (when the network size is high). 4.3. Performance under different dataset sizes and observability levels To further assess the performance of the proposed topology identification method, the impact of varying dataset size has been evaluated. In this context, dataset sizes ranging from just 15 days of data to 60 days, with an increment of 15, days have been used. Moreover, to also consider the influence of observability with low dataset sizes, ten network observability levels have been evaluated (from a 100 % to a 10 %), in a similar way as in the previous subsection. The proposed method has been applied to the eleven SDN described in Table 2, resulting in a total of 440 scenarios. Fig. 5 summarizes the accuracy results for all the possible cases. Each heatmap of Fig. 5 aggregates the result of a specific observability level (indicated in the title of the heatmap). The x-axis identifies the size of the dataset used for topological identification, expressed in days. In the y-axis, each row identifies one SDN. As can be observed, as the observability of the grid decreases, the dataset size needed to identify customers in the network also increases. When network observability is at a sufficiently high level (≥80 %), high accuracy values (≥98 %) can be obtained with just 30 days of data. For medium or low observability levels (<80 % and ≥40 %), 45 days are necessary to obtain high accuracy values. However, when observability is very poor (≤30 %), at least a dataset with 60 days of data is necessary. Furthermore, it can be observed that the number of customers in the network has a negligible impact on the accuracy obtained in the topology identification. This is evidenced by the fact that there is no significant difference in the result between small-size and large-size networks. In conclusion, the proposed method has proved to be effective in identifying network topology, even in circumstances where limited-size datasets are used (requiring just 60 days of hourly data in the most unfavorable scenario) and when the observability of the network is poor (i.e., a high level of unmeasured consumption). 4.4. Influence of measurement error In order to assess the impact of measurement error introduced by Table 2 Execution time for the proposed wavelet-based method and for the energy-based approach for the eleven SDN described in Table 1, using 60 days of hourly data. Network Proposed method execution time (s) Energy-based method execution time (s) SDN1 1.11 23.84 SDN2 2.77 73.97 SDN3 3.18 67.56 SDN4 3.37 46.05 SDN5 4.43 91.51 SDN6 6.57 121.80 SDN7 6.83 175.44 SDN8 11.20 192.72 SDN9 17.87 577.14 SDN10 22.17 1068.43 SDN11 23.80 1169.00 Table 3 Execution time (in seconds) for the proposed wavelet-based method with synthetic networks considering higher number of feeders and customers. Feeders Customers 700 800 900 1000 10 31.59 44.66 59.78 75.15 15 32.86 45.84 61.90 76.50 20 33.85 47.56 62.83 78.06 S. García et al. Electric Power Systems Research 243 (2025) 111517 9