Full text
Electrical Power and Energy Systems 156 (2024) 109695 Available online 5 December 2023 0142-0615/© 2023 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/). Contents lists available at ScienceDirect International Journal of Electrical Power and Energy Systems journal homepage: www.elsevier.com/locate/ijepes Congestion forecast framework based on probabilistic power flow and machine learning for smart distribution grids Alejandro Hernandez-Matheusa,∗, Kjersti Bergb, Vinicius Gadelhaa, Mònica Aragüés-Peñalbaa, Eduard Bullich-Massaguéa, Samuel Galceran-Arellanoa aCITCEA Energy, Universidad Politecnica de Catalunya (UPC), Barcelona, Spain bNorwegian University of Science and Technology (NTNU), Trondheim, Norway ARTICLE INFO Keywords: Line congestions Probabilistic power flow Machine learning Distribution system operators Demand forecasting ABSTRACT The increase in renewable energy sources and new technologies such as electric vehicles and storage can generate uncertainties in distribution grid operations, increasing the likelihood of congestions in power lines. Distribution system operators (DSOs) face several challenges while operating their grids in such conditions. These congestions deteriorate the electrical equipment in the long term, reducing its life span. This work proposes a framework to predict grid asset congestions on a daily basis. A congestion forecast framework is proposed by combining probabilistic power flows and machine learning algorithms to support DSOs in their daily decision-making. The framework is tested on a modified IEEE-33 bus system and CINELDI MV Reference system with hourly synthetic data. The results showed that the framework is able to closely predict the congestions on the lines. Computational capabilities are reported and discussed. The study indicates that the framework is a suitable tool for day-to-day congestion predictions in smart distribution grids yielding low error in expected values. 1. Introduction Technology advances and cost reductions have accelerated renewable distributed generation deployment exponentially in the last years. At the same time, there has been an increasing trend for electrification, which means more power demand is required from electricity instead of fossil fuels. This results in higher power demand. The presence of these tendencies is a step towards achieving the green energy transition. However, both distributed energy generation and electrification of demand induce variability and uncertainty in the distribution grid. The consequences of these uncertainties can lead to congestion in the grid, making it difficult to operate and manage the distribution grids [1]. Congestions occur when the power flows in the grid exceed grid constraints, such as line thermal loading capabilities. These events can happen under different operation conditions, for example, an abrupt increase in the demand, an excess of generation from renewable energy sources, as well as a sudden imbalance between demand and generation [2,3]. These congestion events can result in consequences in short-term and long-term. In short-term, thermal overloading in power lines and transformers [4], which can lead to voltage drops. In the longterm, these elements present cable degradation and undesirable lifetime reduction [5], precipitating large investments from the distribution system operators (DSOs). ∗Corresponding author. E-mail address: [email protected] (A. Hernandez-Matheus). Congestion management techniques have been greatly explored in the literature. Some of these include increasing hosting capacity by traditional methods such as voltage regulation with tap-changing transformer and grid reinforcement [2,6], however, these methods also represent a need for investment to the DSO. Grid operation mechanisms such as load shedding, local generation curtailment and market-based schemes, such as dynamic tariffs or electricity services, [7–10] are other practices DSOs employ to manage congestions. Although effective, some of these solutions require load shedding and renewable generation curtailment, both undesirable techniques given economic compensation payouts and green energy reduction, respectively. Methods and solutions for DSOs to be able to prepare for these events in advance can help to better plan the operational and economic performance of distribution networks, as reported in [11]. Methods for day-ahead congestion management have been explored by demand response [12] and optimisation of dynamic tariffs [13]. However, these studies assume the ability of the DSO to know the grid values in dayahead, thus, they do not provide any methodology to forecast such congestions. The increased implementation of information and communication technologies has transformed distribution grids into smart grids, enabling DSOs to monitor their grids more effectively by receiving and https://doi.org/10.1016/j.ijepes.2023.109695 Received 4 April 2023; Received in revised form 3 November 2023; Accepted 30 November 2023
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 2 A. Hernandez-Matheus et al. storing large volumes of information and grid variables [14,15]. Moreover, the extensive availability of such data creates opportunities to develop energy predictive tools and services that support DSOs in their day-to-day operations and decision-making, employing data-driven methods such as machine learning (ML) [16]. This paper explores a methodology that enables the forecasting of congestion for a smart distribution grid with data-driven methods by utilising the data available for DSOs. 1.1. Related literature Several approaches exist to study and analyse the grid under different scenarios. Probabilistic power flow (PPF), also named probabilistic load flow, is one technique researchers and engineers have been analysing more in-depth in the last few years. Initially formulated by [17], this technique relies on considering variables in the power system as uncertain values following a probability distribution function [18], hence, evaluating and analysing the grid under a large number of operation scenarios. Variables such as load demand, photovoltaic (PV) generation, wind generation, and electric vehicle charging are considered uncertain variables in the operation of the power system [19]. In some works, the topology of the grid is also considered as an uncertain variable [20], although this variable is out of the scope of this work. The goal of PPF is to study the state variables of the system, such as bus voltages, current of the lines, and transformer loading, under the analysed operation scenarios [21]. In contrast to deterministic load flow, PPF analysis evaluates the power system’s current and future operation by studying the uncertain variables in the grid [22]. PPF is a well-known technique to analyse congestion in the grid because of the exploration of a great number of scenarios following the uncertain variables and studying the effect of the behaviour of these variables in the grid. Therefore, there are currently many studies applying this technique to model system uncertainties and obtain system state variables under these scenarios. Ref. [21] reviewed several methods for probability distribution modelling of PV generation and electric vehicle charging and classifies PPF-solving into numerical methods, analytical methods, and approximate methods. A relevant share of the current studies uses numerical approaches, especially Monte-Carlo methods [23,24]. These methods, although accurate and easy to implement, as they do not require any simplification of the original power flow equations, demand expensive computational power by solving a large number of scenarios explored [18]. Additionally, the size of the power system under study is proportional to the computation time and data storage needed [25]. Alternative methods to Monte-Carlo have also been explored. For example, importance sampling, Latin hypercube sampling [26], Principal Component Analysis [27] and Gauss Quadrature [28]. However, these sampling methods can reduce the accuracy of the probabilistic model [20]. While PPF is a suitable technique for analysing power systems under various conditions, including congestion events, its computational cost remains a drawback for the practical implementation of an operational tool for DSOs, even with reduced scenarios. Therefore, integrating this technique with a datadriven approach could be beneficial by leveraging the data generated during the Monte-Carlo process. The drawbacks of PPF, including computational cost, and large storage needs, open the possibility of exploring other methods as solutions for practical implementation. Data-driven algorithms, such as ML methods, have emerged as one solution to overcome the computational challenges required by PPF [29]. ML is a sub-field of artificial intelligence (AI) that, involves the processing of large quantities of data, to learn an approximate function through statistical algorithms with the goal of developing predictive models [30]. In the case of power systems, a great deal of research has been done regarding ML and AI for practical applications [14,31,32]. Refs. [33,34] evaluated several ML regressors to predict voltage in distribution grids and analyse extreme voltage scenarios, obtaining accurate results compared to the power flow solution benchmark. The authors employ different techniques for linear scenario generation. This suggests ML methods have the potential to complement PPF by leveraging the extensive uncertainty scenarios generated from the sampling method. By harnessing ML algorithms, it becomes feasible to efficiently learn from these scenarios and rapidly compute the desired output variables of the grid. This approach provides a viable solution to mitigate the computational burden and practical infeasibility associated with PPF analysis. Some studies have explored the combination of PPF and ML. For example, [1,20,29] implemented neural networks to predict voltages and power flow through the grid’s lines. Other data-driven methods have been combined with PPF such as polynomial chaos [35,36] and Gaussian process emulator [26]. Ref. [37] performed the combination of PPF with a graph-aware neural network. The authors tested this combination also with fully connected neural networks. Similarly, [38] combined a PPF, with the Quasi-MC method, with a convolutional neural network. Both studies achieved an improvement in the computation efficiency of the PFF with the combination of a deep learning-based algorithm. In a similar approach, researchers have made use of ML methods to predict large scenarios of optimal power flow (OPF). Ref. [39] leveraged a stacked extreme ML algorithm to predict the OPF solutions of a great number of scenarios. Similarly, [40] compared neural networks and ML regressors to predict bus voltages and angles from results from OPF. Although these studies have shown the benefits of combining ML with PPF, and similar, there is a lack of literature on the combination of these methods to forecast line congestion in a grid. Regarding practical applications for similar approaches, [41] develops and implements a tool for DSOs to forecast congestion by developing a PV generation and load forecast and combining it with PPF to analyse the voltages, lines, and transformers. However, the authors remark on the high computational toll from the PPF based on the forecast horizon. This paper aims to improve the methods available in the current literature by employing a combination of PPF and ML, exploiting the benefits of both techniques. The methodology avoids the need of performing PPF several times, by substituting the PPF model with a trained ML model, able to provide the possible congestions in the lines given load and PV generation forecast day-ahead. The capabilities of the framework are presented and explored in different study cases. 1.2. Contributions Currently, the literature lacks solutions based on a PPF and ML combination for a practical implementation of grid congestion forecast. To cover this research gap, this work proposes a framework that combines PPF and ML to forecast and analyse line congestion in a smart distribution grid for a day-ahead operation. The framework’s output serve as a tool for DSOs to forecast line congestions and support the decision-making in the management of such events. Hence, the contributions of this paper can be summarised as follows: •Development of a model that combines PPF and ML techniques to return the line currents of a grid with known topology day-ahead, with acceptable accuracy. •Development of a framework, based on the previous model, able to forecast congestions day-ahead for distribution grids with large penetration of renewable generation. This enables DSOs to take day-ahead decisions to improve the management of their grid to avoid congestion and, ultimately, expensive grid reinforcement investments. The rest of the article is organised as follows: Section 2depicts the methodology for this framework, Section 3explains the study case and simulation setup, Section 4presents the results and evaluation of the application of the proposed framework, while Section 5discusses the results and further work.
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 3 A. Hernandez-Matheus et al. Fig. 1. Proposed framework for congestion forecasting. 2. Methodology This section introduces the underlying techniques that shape the congestion forecast framework. The scheme aims to be a tool to warn the grid operator about day-ahead operation issues by forecasting congestions. The framework is divided into two main layers: Training and Execution, as depicted in Fig. 1. 2.1. Training layer The training layer builds the congestion forecast model (see Fig. 1). It follows the following steps: Given input data (1), a PPF is executed with the Monte-Carlo method. Power flows are executed for each generated sample (2), and line currents are stored. Next, the sampled historical consumption and PV generation are fed into an ML algorithm (3), with line currents as the target variable. The algorithm is trained to map the relationships of these variables for the distribution grid. 2.1.1. Input data In the first step, the necessary information for the framework is gathered. The data includes electrical information about the distribution grid, such as topology and lines, and transformer impedances. The grid parameters of the transformers and lines are considered in steady-state conditions. The change of such parameters during highload conditions is not relevant to the scope of this study. This information is analysed and modelled in a power systems solver, in this case, pandapower [42]. Additionally, the time-series data, such as historical electricity consumption and PV generation, is cleaned for further steps. The data requirements for the development of the proposed framework is as follows: •Grid Topology •Electrical parameters of lines and transformers •Historical measurements of electricity consumption and PV generation 2.1.2. Probabilistic power flow The Monte-Carlo method is used to perform the PPF. This numerical method relies on generating and solving a large number of scenarios based on input probability distributions for uncertain variables considered, in this case: electricity consumption and PV generation. The steps to execute this technique are as follows: 1. The net active and reactive power (𝑃𝑛𝑒𝑡 and 𝑄𝑛𝑒𝑡, respectively), denotes the consumption minus PV generation for each bus and is sampled from the input data, see (1) and (2). The samples are stacked together, generating a dataset of many operation scenarios for the distribution grid. 𝑃𝑛𝑒𝑡 =𝑃𝐿𝑜𝑎𝑑 −𝑃𝑃 𝑉 𝑔𝑒𝑛 (1) 𝑄𝑛𝑒𝑡 =𝑄𝐿𝑜𝑎𝑑 (2) 2. Each time step of the historical electricity consumption is considered independent of each other. Each operation scenario is referred to as c. For each scenario, a normal sampling 𝑋∼ (𝜇, 𝜎2)of a fixed number of iterations using the hourly active power value as the mean and the 10% of this value as standard deviation, following previous studies [20,29]. The sampling holds the form of (3). 𝑐∼(𝜇=𝑃𝑛𝑒𝑡, 𝜎2= 10%𝑃𝑛𝑒𝑡)(3) 3. For every scenario of the dataset created, a power flow calculation is executed in pandapower, solved with the Backward– Forward Sweep algorithm. The line currents results are stored for further analysis. 4. Finally, a statistical analysis of the output values is performed, studying the mean and standard deviation to investigate and visualise the probability distribution of the resulting line currents.
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 4 A. Hernandez-Matheus et al. Table 1 Hyperparameters for ML models. Model Hyperparameters Values Random Forest Number of estimators 50, 500, 1000 Features at every split Auto, square root Minimum samples to split a node 2, 5, 10 Minimum samples at each leaf node 1, 2, 4 Bootstrap Binary Support Vector Machine Loss function Polynomial, linear Regularisation parameter 2, 5, 10 Elastic Net Penalty term multiplier 1e−5, 1e−4, 1e−3, 1e−2, 1e−1, 0, 1, 10, 100 Penalty mixing parameter 0 to 1 in 0.01 steps As described in Section 1, given a large number of power flow solutions, the PPF demands a great deal of computational power and storage. Therefore, it is not suitable for day-to-day application. For this reason, the method is coupled with an ML model described in the next subsection. 2.1.3. Machine learning The ML trained model aims to replace the computationally intensive PPF by providing accurate predictions of line currents. The main use of ML in this framework is to generate a model able to predict the output line currents to substitute the extensive results matrices generated in the PPF. In this way, the framework is able to obtain accurate results of line currents without the ample storage space needed. Following previous literature, several ML algorithms are compared in the context of power flow prediction. The performance of an ML algorithm is highly dependent on the specific problem it is applied to, for such, in this step we aim to identify the most suitable ML algorithm for each of the study cases examined in the following section. To achieve this, we compare the training and performance of different ML algorithms. The algorithms chosen for this section have been previously analysed in Section 1, where their ability to accurately capture power flows was demonstrated. Specifically, the algorithms to be compared are: Elastic Net [43], which is a linear regression model with hybrid 𝐿1and 𝐿2regularisation; Random Forest (RF), an ensemble model of decision trees; and Support Vector Machine (SVM) [44]. A brief explanation of the mathematical foundation of these algorithms can be found in [30]. The input for the ML algorithm is composed of the net bus active and reactive power injections, while the output variables are the line currents. Hence, mapping the power flow results yield from the PPF. The power injections shape the input feature matrix and the values are normalised to the maximum and minimum values. Additionally, the training and testing sets are split 70% and 30%, respectively. The hyperparameters for each model are optimised following the grid search algorithm, and results are cross-validated by 𝐾= 4 fold to find the best set of hyperparameter combinations for each model. Ref. [45] provides an overview of hyperparameter optimisation. Table 1 shows the hyperparameter search space for each model. The ML framework is implemented via Scikit-Learn [46]. The models trained are compared by calculating the mean absolute error (MAE) metric, as shown in (4), where 𝑦𝑖are the prediction values, 𝑥𝑖the true values, and 𝑛the number of lines: 𝑀𝐴𝐸 =∑𝑛 𝑖=1 |𝑦𝑖−𝑥𝑖| 𝑛(4) Additionally, the 𝑅2value is calculated from (5) where 𝑆𝑆𝑟𝑒𝑠 is the sum of squared residuals and 𝑆𝑆𝑡𝑜𝑡 is the total sum of the square: 𝑅2= 1 − 𝑆𝑆𝑟𝑒𝑠 𝑆𝑆𝑡𝑜𝑡 (5) This value is also used to show the goodness of fit for each model. The best-performing model is chosen considering the lowest prediction errors and computational KPIs such as training time and storage. Fig. 2. Illustration of calculation of congestion probability. The computer environment is Intel(R) Core(TM) i7-10750H CPU @ 2.60 GHz 16 GB RAM. The trained ML model constitutes the congestion forecast model and the capability of yielding accurate results easily makes it suitable for the development of an online tool able to handle data and execute this framework. As long as the grid maintains its topology, the training for the ML algorithm will only need to be performed once. 2.2. Execution layer In the execution layer, the framework operates on a daily basis and it is possible to observe how the congestion model performs with forecasted load and PV generation as input. The output of the congestion model is then post-processed to calculate the probability of congestion for each line for every hour of the next day. As depicted in Fig. 1, the demand forecast model (4) represents a trained model able to provide a forecast of the loads and PV generation of the grid. The forecast model for day-ahead electricity consumption and PV generation is out of the scope of this paper, and therefore, a profile symbolising a forecast is used to demonstrate the performance of the framework. Subsequently, the forecast is given as input, and the congestion model will then calculate the expected value of the line currents. The calculation involves determining the cumulative probability of a line exceeding its limit threshold, as denoted by (6) and (7).𝜇′and 𝜙′ represent the predicted mean and standard deviation of the congestion model for each line 𝑖, respectively. 𝑧corresponds to z-value of the distribution, and 𝑇 ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑 is chosen by the DSO to manage congestion risk on the line. The probability result is the value then assessed by the DSO for the post-analysis and decision-making. An illustration of the calculation is displayed in Fig. 2. 𝑧=(𝑇 ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑 −𝜇′) 𝜙′(6) 𝑃(𝐼𝑖> 𝑇 ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑) = 1 − 𝑃(𝑍≤𝑧)(7)
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 5 A. Hernandez-Matheus et al. Fig. 3. Modified IEEE 33 Bus System. 2.3. Framework evaluation To evaluate the performance of this framework, a comparative assessment of the output of the framework and deterministic results is carried out. Given a synthetic forecast for the loads and the PV generation, the framework is executed, and the expected output is compared with the results from a deterministic power flow. The absolute errors, with respect to the deterministic power flow result, are then calculated to obtain the performance of the framework: 𝐸𝑟𝑟𝑜𝑟𝑖=𝐼𝑖,𝐷𝑃 𝐹 −𝐼𝑖,𝐹 𝑊 𝐼𝑖,𝐷𝑃 𝐹 (8) where 𝐼𝑖,𝐷𝑃 𝐹 is the current in line 𝑖from the deterministic power flow and 𝐼𝑖,𝐹 𝑊 is the current in line 𝑖from the ML framework. 3. Case study This section depicts the case studies employed in this work for evaluating the framework proposed. The configuration and assumptions of the distribution grids analysed are explained. 3.1. IEEE test case 33 bus system The IEEE 33 Bus System [47] is used to test the framework. The test case represents a distribution grid of 33 nodes in 12.66 kV. Loads and PV generation nodes of the distribution grid are shown in Fig. 3. The line characteristics are the reference values available in pandapower library [42], however the maximum line current (max_i_kA) is modified to represent congestions in the grid. The electricity consumption profiles used as input are synthetic hourly consumption profiles for January and February, and March of 2020 data from a Spanish DSO. The PV generation profile has been generated from [48] considering the location in Spain. The size and location of the PV generation in the grid are shown in Table 2. This dataset is used for the training layer of the framework. An algorithm has been developed to generate random load peaks in the historical consumption mentioned previously. This algorithm adds random noise to the load profile at any given hour. This modification generates a historical electricity consumption with spikes on the load, generating congestions in the grid. Fig. 4 shows the modified aggregated load profile and net active power (𝑃𝑛𝑒𝑡)for two weeks, where the Table 2 IEEE test case 33 bus PV system capacities. PV system Max. PV generation (MW) Bus 25 0.275 Bus 26 0.249 Bus 27 0.370 Bus 28 0.236 Bus 29 0.250 Bus 30 0.244 Bus 31 0.211 Bus 32 0.232 Fig. 4. Aggregated hourly demand and net active power profiles for 2 weeks of the training time. spikes in consumption can be seen. When including the PV generation, the load is lowered in most hours, but there are also some hours where the line flows are reversed. Fig. 5 shows a boxplot of the behaviour of the distribution grid line current loading for the analysed time period of two months. Note that line loading is given in percentage of 𝐼𝑚𝑎𝑥, therefore the plot never shows negative values even though there are reverse power flows in some hours. For this case study, the threshold selected for line congestion is 70% of the thermal limit. Fig. 5 shows that lines 2 and 5 experience the highest line currents, with respect to the maximum, during normal operation. Additionally, to test the framework, the expected demand and PV generation for the next day are given by a synthetic forecast generated
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 6 A. Hernandez-Matheus et al. Fig. 5. IEEE Test Case 33 Bus line currents for the analysed time period of three months. Table 3 CINELDI MV set of most congested lines. Line Bus from Bus to Line 4 5 7 Line 5 7 9 Line 6 9 12 Line 7 12 26 by sampling several values of the modified load profile and PV generation. The modified profiles and PV generation used for this work can be found in [49]. 3.2. CINELDI MV reference system extended with new loads The CINELDI MV Reference System is also used to test the framework (see Fig. 6). This is a representative Norwegian radial, medium voltage (22 kV) distribution system with 124 buses and 123 lines. The detailed description of the dataset is available at [50]. The reference version of the grid is extended with new loads such as Local Energy Communities (LECs), installed in buses 30, 38 and 65, to represent future distribution grids. These LECs represent an aggregated behaviour of residential loads, and technologies such as PV generation, local battery storage and electric vehicle charging. Additionally, two PV power plants of 5 MWp are installed in buses 89 and 104. The available data includes load time series for a full year for all 54 load points and the 3 LECs. The PV generation profile of the power plants is retrieved from [48] within a southern Norwegian location. The lines have been modified to show the representation of the congestion in the system, with the maximum current of each line being reduced 10%, to add more congestion scenarios. Fig. 7 shows the most affected set of lines in the distribution grid affected by the extended conditions. Table 3 shows the topology of the lines shown. 4. Results and discussion This section presents the results of the framework for the training layer and execution layer for the two study cases to showcase the capabilities and scalability of the framework. 4.1. IEEE test case 33 bus system This sub-section showcases the main results of the framework evaluated in the IEEE Test Case 33 Bus System. Table 4 Prediction metrics for line currents for the ML algorithms. Model MAE 𝑅2 Random Forest (RF) 0.0013 0.759 ElasticNet (EN) 0.0007 0.913 Support Vector Machine (SVM) 0.0007 0.913 Table 5 Computational KPIs comparison. Method Iterations Time (min) Storage (kB) PPF 1000 363.28 607 RF N/A 107.55 476 EN N/A 69.07 57 SVM N/A 3.90 19 4.1.1. Congestions model The input data for the ML models are generated during the PPF process by sampling the net active and reactive power injections at each bus. The label to predict for the ML models is the line current measured in kilo-Ampere (kA). The performance of the trained models, given the test dataset, is compared against the results of the PPF. The results are shown in Table 4. The results show that EN and SVM models perform better than the RF model, with an MAE of 0.0007 compared to 0.0013. The highest 𝑅2value is 0.913, which is the same for EN and SVM, while RF has an 𝑅2value of 0.759. Hence, SVM and EN obtain better results than a more robust algorithm such as RF, regardless of its linear nature. Additionally, the lower result for RF is an indication of overfitting. These results suggest that the regression task of predicting line currents has a linear component, yielding a more accurate performance from a simpler model such as SVM and EN. Regarding computational KPIs, the training times, that include PPF and ML training, and storage are shown in Table 5. The results show that RF takes the longest time to train of the three methods with 108 min. SVM uses the shortest time, requiring 4 min. It should, however, be noted that this training time is proportional to the hyperparameters for the grid search, and in this case study, both EN and SVM have fewer hyperparameters to tune than RF. In this case, the storage related to PPF Monte-Carlo-based is taking into consideration the large matrices of power flow results. As for storing the trained models, it can be observed that both EN and SVM require significantly less space than RF. Considering MAE and 𝑅2as discriminant metrics, SVM is the model chosen as the congestion forecast model. Fig. 8 shows two hours reflecting high demand in the grid for the three models. The figure shows the comparison of the ML models probability distribution fit with the PPF results for line 2 and line 5, which are the main congested lines given the configuration of the grid. It can be observed that EN and SVM can closely predict the mean of the line currents, as confirmed by the metrics shown in Table 4. However, the variance is not as accurately predicted, making the result a more conservative prediction. In other words, the framework could generate false positive predictions for line congestions. Nonetheless, with a 70% line congestion threshold, this conservative prediction could be acceptable for the DSO. In general, the models have more accurate predictions for line 5 (with low currents), than line 2 (with higher currents). SVM predicts a mean value of 0.1742 kA for line 2, while the PPF results in 0.1743 kA. For line 5, SVM predicts a mean value of 0.1113 kA, compared to 0.11069 kA for PPF. 4.1.2. Execution As explained in Section 2, a synthetic forecast is created for the execution of the framework. This forecast is then preprocessed as explained in Section 2.1.1 and fed to the congestion model. That is, the values of the net power injection for each bus have been normalised to training values. The framework is executed for all the hours of the
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 7 A. Hernandez-Matheus et al. Fig. 6. CINELDI MV reference system [50]. Fig. 7. CINELDI Reference System line currents for the analysed time period of three months. Table 6 Probability (%) of line congestions day-ahead. Line h =15 h =16 h =17 h =18 h =19 Line 2 45.78 0 0 33.5 0 Line 5 94.18 51.81 41.43 10.42 0 forecast and the execution time is reported. However, for the sake of readability, the results are shown for the most demanding hours of the day: from 15 h to 19 h. The total execution time for the 24 h forecast amounts to 0.0827 s. Fig. 9 shows the expected line current values for specified hours given the forecast. Following the radial nature of the grid, lines 2 and 5 present a higher risk of congestion. The final result of the framework is presented in Table 6, where the probability of line congestions for the specified hours is given. Hour 15 has the highest probability of having line congestions the next day, with a 46% and 94% probability of reaching the threshold for lines 2 and 5, respectively. This information can be used by the DSO to plan the operation for the next day, making it possible to take into consideration the uncertainty related to demand and generation forecasts. 4.1.3. Framework evaluation To evaluate the framework, the forecasted demand is fed to a deterministic power flow solver to determine the line current for 24 h of the given day. These results are further compared with the expected value results of the framework for different times of the day and normalised to the results of the power flows. Hence, the results are absolute errors with respect to the deterministic result, calculated as shown in (8). The compared results of the expected values of lines 2 and 5 for each model can be observed in Table 7. The results of the framework using EN and RF as congestion models are shown for the sake of comparison. It can be seen from the table that the expected values from the framework are close to the deterministic power flow result; however, in the hours with the highest demand, h =15 and h =18, the framework results in a higher error for both lines in comparison to the other hours of the day. This indicates that when there is a high demand, uncertainty is also higher. However, with a maximum error of 8.782%, in comparison with a deterministic power flow, the results are considered acceptable for the DSO to use this tool as support for decision-making. 4.1.4. Sensitivity analysis on uncertainty in net demand forecast The framework is also tested to evaluate its performance towards the uncertainty in the forecasted net demand. This net demand forecast follows the same principle as (1). This synthetic forecast exemplifies a new operation scenario of the grid. To test the sensitivity of the model, the forecast is extended by adding uncertainty in the net demand forecast, in the form of ±10% of the mean for each load point, as shown in Fig. 10. This value of uncertainty is considered a way to introduce variability into the forecast. It represents a simplified assumption employed to assess the performance of our framework in the presence of these variations. It is important to emphasise that in forecast analysis, varying standard deviation values are often observed across different hours and environmental conditions. The labels for this uncertainty evaluation in the net demand corresponding to Upper Bound Case (UBC), Expected Value (EV) and Lower Bound Case (LBC). The absolute error with respect to the line currents from the deterministic power flow is shown in Fig. 11. It can be observed that for UBC the framework returns a higher value than the benchmark result from the deterministic power flow solution. In this sense, the framework is more conservative in the UBC scenario forecast. On the other hand, it can be seen that for lower demand, in the LBC, the model yields lower values for the current in the lines, with −16.79% for line 5 at hour 19.
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 8 A. Hernandez-Matheus et al. Fig. 8. PPF vs. ML probability density results in comparison for congested lines 2 and 5. Table 7 Absolute errors of line currents (%), comparing ML framework with deterministic power flow. Model h =15 h =16 h =17 h =18 h =19 Line 2 Line 5 Line 2 Line 5 Line 2 Line 5 Line 2 Line 5 Line 2 Line 5 SVM 6.494 3.416 1.168 1.104 −1.244 −1.049 8.782 1.2038 −1.907 −4.871 EN 6.750 4.222 1.176 1.078 −1.267 0.136 9.035 2.170 −1.9602 −5.058 RF 35.48 26.7199 23.808 21.52 16.790 13.914 37.287 8.9503 17.190 1.634 Fig. 9. Expected line currents from 15 h to 19 h. Fig. 10. Aggregated net demand (demand minus production) forecast with uncertainty. Although it is expected for the DSOs to not have issues in their grid in a lower-demand scenario, it is still not ideal for the model to return lower values as for this specific application the model should have a higher sensitivity or higher ‘‘false-alarm rate’’ when returning the results. Fig. 11. Line current prediction error for uncertainty in demand, 15 h to 19 h. Table 8 PPF execution time and storage comparison. Case study Period of data Time (min) Storage (kB) IEEE test case (33 bus) 3 months 363.28 607 CINELDI MV system (127 bus) 3 months 596.78 2017 CINELDI MV system (127 bus) 1 month 183.24 691.9 These results also reflect the susceptibility of the model performance if the average values of the loads or generation either increase or decrease over time. This would generate inaccurate results over time, and therefore a need for the model to be re-trained with new operational data gathered by the DSO. 4.2. CINELDI MV reference system This sub-section showcases the main scalability of the framework proposed and the main results in the CINELDI MV Reference System. Table 8 shows the comparison of computation time for PPF for the two study cases. From the table, it can be appreciated the increase in the amount of computation time that the larger study case demands, with an execution of 596 min. This increase in computation time is given by the large number of variables in the CINELDI MV grid. The storage of the results of these variables can take up to 2 GB. From this
International Journal of Electrical Power and Energy Systems 156 (2024) 109695 9 A. Hernandez-Matheus et al. Table 9 CINELDI MV system SVM results. Model MAE R2 Time (min) SVM 0.0006 0.664 8.367 Fig. 12. Line current prediction error for congestion hours. comparison, it can be shown how much the number of variables affects the execution time of the PPF. To evaluate the framework on this distribution grid, one month of available data has been used and the framework has only been executed with SVM as congestion model. The results of the trained model are shown in Table 9, where it can be observed that SVM can predict correctly the line currents for this case study. In comparison to the previous study case, the SVM trained performs with similar accuracy, however, the training time is higher. Similar to the PPF, the increase in training time is correlated to the amount of input and output variables of each study case. The CINELDI MV system has over 3 times the number of lines to predict than the IEEE test case 33 bus. Following the same methodology to generate the synthetic forecast for day-ahead as in the previous case study, the results of the framework executed from 9 h to 18 h are shown in Fig. 12. The congestion forecast execution time for the CINELDI distribution grid is 0.793 s. It can be observed that the framework returns a lower value for the current in all cases. The maximum error reported is −11.45% in hour =10 for line 7. On the other hand, for hours h =13 and h =14, the framework returns values close to 0%. Note that for this study case, only the expected value of the forecast is taken into consideration. After the analysis of the results, it can be concluded that the model performs satisfactory when applied to large distribution grids, proving the scalability of the framework. 5. Conclusions In this work, a framework based on the combination of PPF technique and ML has been developed to generate a congestion forecast tool to be used in day-to-day operations. Such a tool is aimed to be a decision support for DSOs. The results showed that the Support Vector Machine algorithm could reproduce the results of the PPF with adequate MAE and 𝑅2values. This allows the repetition of the large computation of the Monte-Carlo technique to be avoided when exploring the grid under study in similar operation scenarios. By executing the Monte-Carlo simulations once, the predictive model can emulate many different scenarios with acceptable accuracy. The framework capabilities are tested in the IEEE 33 bus test case and its scalability of the framework is also tested in the CINELDI MV reference system of 124 buses. The PPF computation times are compared, showing high computation times for both cases, 362.28 mins and 596 mins, respectively. However, the performance of the trained congestion models is similar for both cases, with an MAE of 0.0007 and 0.0006, respectively. For the execution of the 24 h line congestion forecast, the trained models take 0.0827 s and 0.793 s, for IEEE 33 bus and CINELDI MV, respectively. The computational advantages and the accurate performance of the framework proposed generate an operational tool for congestion forecast and grid management able to be implemented in online platforms such as cloud services However, an inherent drawback of the proposed methodology is that it is not flexible in terms of topology, since the power flow results change when the topology changes. Therefore, the framework is not suitable for distribution grids with constant changes in topology. Also, the sensitivity of the framework was demonstrated under different forecast scenarios, showing that the performance fluctuates under conditions on high uncertainty scenarios, with the framework returning higher errors when predicting line currents in hours with high demand. For future work, the study could be extended with a real implementation of the framework to evaluate the performance of the trained models in a real operation environment. This would provide valuable information on how the framework could be used to support DSOs in their daily decision-making. Furthermore, this framework could be used in combination with a flexibility allocation forecast, a combination that yields a holistic congestion management scheme. CRediT authorship contribution statement Alejandro Hernandez-Matheus: Conceptualisation, Investigation, Methodology, Writing – original draft, Writing – review & editing, Visualisation. Kjersti Berg: Investigation, Methodology, Writing – original draft, Writing – review & editing, Visualisation. Vinicius Gadelha: Writing – review & editing. Mònica Aragüés-Peñalba: Conceptualisation, Supervision. Eduard Bullich-Massagué: Conceptualisation, Supervision. Samuel Galceran-Arellano: Supervision. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Data availability The data has been shared in an open repository. Acknowledgements The authors would like to thank Rubi Rana for discussion regarding the CINELDI Reference System. This work was supported by the project consortium of the research project FINE (Flexible Integration of Local Energy Communities into the Norwegian Electricity Distribution System), financed by the Research Council of Norway [project number 308833]. This work has also been supported by the BD4OPEM H2020 project, which has received funding from the European Union’s Horizon 2020 research and innovation program under Grant Agreement No. 872525. Mònica Aragüés is Associate Professor and Eduard Bullich-Massagué is a lecturer of the Serra Húnter programme. References [1] Wang D, Zheng K, Chen Q, Zhang X, Luo G. A data-driven probabilistic power flow method based on convolutional neural networks. Int Trans Electr Energy Syst 2020;30(7):1–15. http://dx.doi.org/10.1002/2050-7038.12367. [2] Bach Andersen P, Hu J, Heussen K. Coordination strategies for distribution grid congestion management in a multi-actor, multi-objective setting. In: IEEE PES innovative smart grid technologies conference Europe. IEEE; 2012, p. 1–8. http://dx.doi.org/10.1109/ISGTEurope.2012.6465853. [3] Esmat A, Usaola J, Moreno MÁ. Distribution-level flexibility market for congestion management. Energies 2018;11(5). http://dx.doi.org/10.3390/en11051056. [4] Pillay A, Prabhakar Karthikeyan S, Kothari DP. Congestion management in power systems - A review. Int J Electr Power Energy Syst 2015;70:83–90. http://dx.doi.org/10.1016/j.ijepes.2015.01.022.