Full text
Treball de Fi de Màster Master’s in Renewable Energy Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs MEMORIA Escola Tècnica Superior d’Enginyeria Industrial de Barcelona Author: Kevin Binz Varghese Director: Mònica Aragüés Peñalba Co-director: Marc Jené Vinuesa Call: July 17th 2024
Page 2 July 17th 2024
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 3 Abstract This thesis explores machine learning-based baseline estimation methods for residential households, particularly focusing on their application in demand response strategies. By compiling and analyzing baseline estimation methods, this work proposes optimal demand response strategies for a sample of 300 households in Australia taken from electricity provider Ausgrid. These households, equipped with behind-the-meter photovoltaic systems, present a unique opportunity to study the impacts of photovoltaic intermittency on smart meter readings and overall energy management. An Incentive Based Demand Response (IBDR) mechanism is formulated to show the load reduction during peak hours, profits retained by the retailer, and benefits gained by the customer for shifting the load during peak electricity consumption periods. Various baseline estimation methods are thoroughly studied to quantify the incentives for residential customers, including averaging methods and machine learning models, evaluated for customers with different consumption levels. Different demand response event days are selected and evaluated to ensure a broad range of results. Key findings include the significant load reduction potential of the IBDR mechanism, particularly among high-consumption customers. The Random Forest Regression model consistently provides the best fit for demand response event day load curves, highlighting the superiority of machine learning models over simple averaging methods in baseline estimation. The analysis emphasizes the need for adaptive approaches, as per the No Free Lunch theory, suggesting future exploration of hybrid or ensemble models. Effective consumer education on peak-hour consumption is crucial for enhancing grid stability and promoting sustainable energy practices. These insights underscore the importance of targeting high-consumption customers with demand response incentives and the variability in model performance, paving the way for more efficient demand response programs. In the end, the code used for this thesis is provided for the evaluation of baseline models to quantify the fair amount of compensation that should be provided to the residential customers for participating in the demand response program.
Page 4 July 17th 2024 Content ABSTRACT ________________________________________________ 3 CONTENT __________________________________________________ 4 NOMENCLATURE ___________________________________________ 7 LIST OF FIGURES __________________________________________ 10 LIST OF TABLES ___________________________________________ 12 1. INTRODUCTION _______________________________________ 14 1.1. Motivation ........................................................................................................ 15 1.2. Defining the Scope ......................................................................................... 16 1.3. Objectives of the thesis .................................................................................. 17 2. THEORETICAL BACKGROUND __________________________ 19 2.1. Smart Metering ............................................................................................... 19 2.2. Grid System Operators and Aggregators: Their roles and status ............. 20 2.2.1. Distribution System Operators (DSOs) ............................................................. 20 2.2.2. Transmission System Operators (TSOs) ......................................................... 21 2.2.3. Aggregators ......................................................................................................... 23 2.3. Electricity Tariffs .............................................................................................. 24 3. DEMAND SIDE MANAGEMENT ___________________________ 26 3.1. Types of Demand Side Management Practices. ........................................ 26 3.1.1. Energy Efficiency ................................................................................................ 27 3.1.2. Load Shedding .................................................................................................... 27 3.1.3. Load Shifting ........................................................................................................ 27 3.1.4. Power Generation ............................................................................................... 28 3.2. Demand Response......................................................................................... 30 3.2.1. Demand Response Schemes ............................................................................ 31 4. STATE-OF-THE-ART ALGORITHMS _______________________ 33 4.1. Data Science Approach ................................................................................. 33 4.2. Incentive-Based Demand Response Model (IBDR) ................................... 34 4.3. Price Elasticity of Demand ............................................................................. 36 4.4. Customer Baseline Estimation ...................................................................... 37 4.4.1. Customer Baseline Load Estimation Methods. ................................................ 38 4.4.2. Averaging Methods ............................................................................................. 39 4.4.3. Regression Methods ........................................................................................... 39 4.4.4. Control Group ...................................................................................................... 40 4.4.5. Machine Learning................................................................................................ 40
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 5 4.5. Evaluation Metrics for Baseline Evaluation ................................................. 40 5. EXPERIMENTAL FRAMEWORK __________________________ 42 5.1. Experimental Setup and Data Analysis Methodology ................................ 42 5.2. Data Origin, Collection and Preparation ...................................................... 42 5.3. Customer Selection Based on Consumption Levels .................................. 46 5.4. DR Event Simulation with IBDR Mechanism ............................................... 47 5.5. Baseline Estimation: Averaging Methods .................................................... 49 5.5.1. High X of Y ........................................................................................................... 49 5.5.2. Low X of Y ............................................................................................................ 49 5.5.3. Mid X of Y............................................................................................................. 50 5.5.4. Nearest X of Y ..................................................................................................... 50 5.6. Baseline Estimation: Machine Learning Models ......................................... 50 5.6.1. Splitting the Dataset ............................................................................................ 51 5.6.2. Feature Description............................................................................................. 52 5.6.3. Ridge Regression................................................................................................ 54 5.6.4. Extreme Gradient Boosting Regression ........................................................... 55 5.6.5. Random Forrest Regression.............................................................................. 56 5.6.6. Support Vector Machines ................................................................................... 57 6. EXPERIMENTAL RESULTS ______________________________ 58 6.1. IBDR Mechanism Results.............................................................................. 58 6.2. Baseline Method Comparison ....................................................................... 63 7. PLANNING ____________________________________________ 76 7.1. Gantt Chart ...................................................................................................... 77 8. ECONOMIC, ENVIRONMENTAL AND SOCIAL IMPACT _______ 79 8.1. Economic Impact: Project Development Cost ............................................. 79 8.2. Environmental Impact .................................................................................... 80 8.2.1. Equipment and Carbon Footprint ...................................................................... 80 8.2.2. Renewable Energy Context ............................................................................... 80 8.3. Social impact and gender equality ................................................................ 81 9. CONCLUSIONS ________________________________________ 83 9.1. IBDR Mechanism ............................................................................................ 83 9.2. Performance of Baseline Models. ................................................................. 84 9.3. Implications and Future Directives ............................................................... 85 10. ACKNOWLEDGMENTS _________________________________ 87 11. BIBLIOGRAPHIC REFERENCES __________________________ 88
Page 6 July 17th 2024
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 7 Nomenclature Abbreviations RES: Renewable Energy Systems EF: Energy Flexibility IEA: International Energy Agency DSM: Demand Side Management DM: Demand Response BTM: Behind The Meter PV : Photovoltaic IBDR: Incentive-Based Demand Response PBDR: Price-Based Demand Response BL: Baseline Load SMS: Smart Metering Systems MMD: Modern Measuring Systems ToU: Time-of-Use DSO: Distribution System Operator TSO: Transmission System Operator BRP: Balance Responsible Parties EU: European Union NSW: New South Wales FiT: Feed-in-Tarrif GG: Gross Generation FERC: Federation Energy Regulatory Commission
Page 8 July 17th 2024 CBL: Customer Baseline Load LMP: Locational Marginal Pricing PED: Price Elasticity of Demand PEM: Price Elasticity Model PJM: Pensylvenia New Jersey Maryland NYISO: New York Independendant System Operator AI: Artificial Intelligence ML: Machine Learning TC: Total Consumption NC: Net Consumption GC: General Consumption CL: Controlled Load RRP: Recommended Retail Price LR: Load Reduction RP: Retailer Profit CB: Customer Baseline I: Incentive rates AEMO: Australian Electricity Market Operator XGBoost: eXtreme Gradient Boosting RFR: Random Forrest Regression SVM: Support Vector Machine SVR: Support Vector Regression SDGs: Sustainable Development Goals
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 9 Nomenclature AUD: Australian Dollars cAUD: cent Australian Dollars MWh: Megawatt-hour kWh: Kilowatt-hour ℇ∶ Price Elasticity of Demand Q : Electricity Demand P: Electricity Price 𝐹 (𝑙𝑖(𝑑𝑖,𝑡𝑖))∶ Function to categorize customers according to consumption 𝑙𝑖(𝑑𝑖,𝑡𝑖)∶ Actual Consumption of the customer at i-th day and i-th time. 𝑏(𝑑𝑖,𝑡𝑖)∶ Baseline Load Consumption fo the customer at i-th day and i-th time.
Page 16 July 17th 2024 to manage the new dynamics of energy supply and demand, making the evaluation and improvement of DSM a priority. In this thesis, the focus will be centered on IBDR for residential consumers as an effective demand response mechanism that can mitigate the mismatches between supply and demand. By incentivizing consumers to adjust their electricity usage, IBDR programs help maintain grid stability and prevent economic losses and damages to distribution system operators (DSOs), transmission system operators (TSOs), and energy utilities [5]. This formed the main research questions that is addressed in this thesis: 1. How can residential customers be incentivized to shift their electricity consumption away from peak periods? 2. Which is the best method to estimate what the residential customer's load would have been in the absence of a Demand Response (DR) event? Moreover, accurate baseline estimation is crucial for the success of demand response programs. Baseline estimation allows for the assessment of the benefits derived from demand response, ensures fair compensation for participants, and provides utilities with reliable predictions of demand-side flexibility [5]. This thesis addresses these issues by leveraging data analytical approach to DR programs and compare simple averaging as well machine learning-based baseline estimation methods tailored to the unique conditions of behind-the-meter PV installations. Furthermore, there has never been a better time to explore the integration of artificial intelligence (AI) and machine learning (ML) in the energy sector. AI and ML have reached a critical point and are set to transform every industry, impacting economies and societies in the coming decades. AI enables an unprecedented ability to analyze enormous datasets and discover complex patterns and relationships, which is pivotal for advancing demand response strategies and enhancing the efficiency and reliability of the power grid. 1.2. Defining the Scope The scope of this thesis encompasses the development of a DR mechanism that incentivizes customers to shift their load from peak hours to those of PV generation nd the comparison of the performance of baseline estimation methods to quantify the fair compensation that needs to be given to the residential customer. The thesis is performed with a specific focus on households equipped with PV systems in New South Wales, Australia. The study includes the following key elements:
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 17 1. Temporal Scope: The analysis covers a defined period during which data on household energy consumption, PV generation and various features for modelling will be collected and analyzed. The methods developed will be applicable to both current and future demand response scenarios. 2. Spatial Scope: The research is limited to a sample of 300 households in Australia taken from an open-source dataset provided by Ausgrid, representing a range of climatic and geographic conditions. This allows for a comprehensive assessment of baseline estimation methods across different environments. 3. Technological Scope: The thesis focuses on averaging and machine learningbased approaches to baseline estimation, leveraging advanced algorithms to improve the accuracy and reliability of demand response predictions. These methods will be compared to highlight their advantages and potential limitations. 4. Environmental Scope: The integration of renewable energy sources and the promotion of energy flexibility in buildings contribute to broader environmental goals, such as reducing greenhouse gas emissions and promoting sustainable energy use. By addressing these dimensions, this thesis aims to provide a comprehensive understanding of the challenges and opportunities associated with baseline estimation for demand response in the context of increasing renewable energy integration. 1.3. Objectives of the thesis The primary objective is to illustrate the financial benefits for both customers and retailers who participate in the DR program and this mechanism could be identified as a baseline to help energy planners inform the residential sector about the benefits of load shifting. Secondly, this thesis aims to evaluate, measure, and validate various baseline estimation methods for households equipped with BTM households with PV systems. By analyzing the consumption patterns and behavior of these households, the thesis aims to identify the most feasible and accurate baseline estimation methods. These methods are crucial for designing effective demand response strategies tailored to households with BTM PV installations. Moreover, the baseline estimation models developed and validated in this thesis have broader applications. They can be used in future scenarios to disaggregate the generation of BTM energy resources, providing insights and incentives to TSOs, DSOs, and energy utility companies. This, in turn, will support the integration of renewable energy sources into the grid, enhance grid stability, and promote efficient energy use.
Page 18 July 17th 2024 The particular objectives would be: 1. Conduct a literature review regarding smart metering, electricity tarrifs and demand response. 2. Perform a comprehensive data analysis and preprocessing of the Ausgrid dataset to make it fit for the study. 3. Develop an IBDR mechanism to provide insights into customer participation during DR events. 4. Formulate and compare the profits gained by the energy retailer and the benefits to customers participating in the demand response across different incentive rates. 5. Evaluate nine baseline methods to estimate the expected consumption and provide a benchmark against the actual recorded consumption. 6. Perform a comparative analysis using error metrics and the IBDR mechanism to assess the forecasting capabilities of the ML models and the average baseline estimation methods, especially during scheduled demand response events. .
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 19 2. Theoretical Background 2.1. Smart Metering Smart Metering Systems (SMS) and Modern Measuring Devices (MMD) are advanced electricity meters that employ digital technology instead of the traditional electromechanical mechanisms. These systems are pivotal in the digitization of the energy industry, enabling more accurate and efficient data collection and transmission [6]. These modern meters cater to various categories of energy consumers. Traditional end-users benefit from more accurate billing and real-time consumption data, enabling better energy management and cost savings. Additionally, SMS and MMDs are essential for active energy consumers, often referred to as prosumers. Prosumers, who both produce and consume energy (such as those with solar panels), can optimize their energy usage and feed surplus energy back into the grid, thus contributing to a more sustainable energy system [7]. SMS are installed at the point where a building or residence connects to the electrical grid, allowing them to collect data on energy consumption and generation throughout the building. This location enables the meters to capture information on the flows of power in both directions, as illustrated in Figure 2 [8]. Figure 1: Typical location of a Smart Metering System in the electrical power grid [8] Compared with traditional meters, SMS provide a significant advancement by measuring electricity usage at more frequent intervals, such as every 5 or 15 minutes [8]. These meters automatically transmit this data to the utility, enabling real-time monitoring and analysis. This continuous data flow offers numerous benefits to both consumers and utilities [9]. For consumers, smart meters offer detailed insights into their energy consumption patterns, which can help them better manage their electricity usage. This data helps consumers in taking advantage of time-of-use (ToU) tariff rates, reducing their energy costs by shifting consumption to off-peak times. Additionally, the granular data provided
Page 20 July 17th 2024 by SMS supports the adoption of energy-efficient practices such as DR programs helping consumers to reduce their consumption as well as carbon footprint [66]. Energy utilities benefit significantly from smart meters through improved detection and restoration of power outages. Smart meters can instantly report outages, allowing utilities to quickly identify and address the problem areas, thus enhancing the reliability and efficiency of the power supply. The detailed consumption data also enables utilities to optimize grid operations, minimize energy losses, and make informed decisions regarding infrastructure investments [9]. In the context of DR and DSM, smart meters enable precise monitoring of energy use, allowing utilities to verify customer participation and effectiveness during DR events [67]. This capability is vital for providing accurate incentives to consumers who adjust their usage in response to grid needs. Furthermore, the data from smart meters supports the development of robust baseline methodologies, which are used to estimate what energy consumption would have been in the absence of a DR event. Accurate baselines are critical for assessing the impact of DR activities and ensuring fair compensation [41]. Without the granularity and accuracy offered by smart meters, utilities would struggle to measure and validate the performance of DR programs effectively. The integration of smart meters thus enhances the reliability of DR initiatives and supports the transition to more dynamic and responsive energy systems [66]. 2.2. Grid System Operators and Aggregators: Their roles and status 2.2.1. Distribution System Operators (DSOs) DSOs manage the medium and low-voltage distribution networks that deliver electricity from the transmission system to end-users, such as homes and businesses. They play a crucial role in the final delivery of electricity and ensuring its quality and reliability. Power quality and loss reduction are critical for DSOs to ensure the efficient operation of the distribution network. DSOs must maintain the balance of loads among phases to reduce losses, increase network capacity, minimize the risk of failures, and improve voltage profiles [27]. In addition, DSOs need to support resilience during extreme events. Services such as islanding, black start, and emergency load management are essential to enhance the resilience of distribution networks during natural disasters and extreme weather conditions. These services ensure that the network can quickly recover and continue to operate under adverse conditions [31].
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 21 Another significant need for DSOs is deferring network investments. By using voltage control and congestion management services, DSOs can address current or forecasted physical congestions, thereby reducing the immediate need for infrastructure expansion investments. This approach allows for more efficient capital allocation and extends the lifespan of existing assets [26]. Managing physical congestion is also a critical task for DSOs. This involves real-time management, intraday adjustments, or planning months ahead to ensure sufficient power is provided despite network limitations. Effective congestion management ensures that the distribution network operates smoothly without overloading any part of the system. Voltage violations control is another essential function for DSOs. Maintaining voltages within specific limits and restoring nominal values after disturbances are crucial for minimizing reactive power flows, reducing technical losses, [27] and avoiding costly investments in additional infrastructure. By controlling voltage violations, DSOs ensure the stability and reliability of the distribution network. 2.2.2. Transmission System Operators (TSOs) TSOs are responsible for the high-voltage transmission network that transports electricity over long distances from power plants to distribution networks. Their primary role is to ensure the reliable and efficient transmission of electricity, maintaining the balance between electricity supply and demand in real-time. TSOs have specific needs that differ from those of DSOs but are equally critical for the overall stability of the power grid. One of the primary needs of TSOs is balancing requirements. TSOs require frequency response services to maintain grid frequency, which is vital for the stability of the entire power system. Innovative services such as ramp control, smoothed production, and damping power system oscillations are necessary to manage the dynamic behavior of the grid effectively [27]. Congestion management is another crucial need for TSOs. They manage both intraregional and cross-border congestion through real-time operations, short-term planning, and re-dispatch. Effective congestion management ensures that electricity can flow freely across regions and borders, preventing bottlenecks and maintaining a stable supply. Non-frequency ancillary services are also essential for TSOs. These services include reactive power and voltage control, which are necessary to maintain the quality of power and ensure the stability of the voltage levels throughout the transmission network. System restoration capabilities, such as black start and islanding, are critical for recovering the grid after significant outages and ensuring continuous power supply [31].
Page 22 July 17th 2024 Lastly, TSOs need to ensure the adequacy of the power system. They achieve this through capacity remuneration mechanisms that maintain system reliability by ensuring there is enough capacity available to meet demand. These mechanisms incentivize the availability of sufficient resources to handle peak loads and unexpected surges in demand, thereby maintaining the overall stability and reliability of the power system. Need Description DSO TSO Power Quality and Loss Reduction [26, 31] Phase Balancing: Maintain load balance among phases to reduce losses and improve voltage profiles. Extreme Events’ Support [26, 27,31] Islanding: Operate isolated grid sections. Black start: Restore power after blackout. Emergency Load Management: Manage loads during emergencies. Network Investments’ Deferring [26, 31] Voltage Control (Power Based): Use flexible loads to manage voltage levels. Congestion Management (Capacity Based): Use flexibility to manage physical congestions. Physical Congestion Control [26, 27, 31] Congestion Management: Manage realtime and predictive congestion. Voltage Violations Control [26] Voltage Control: Maintain and restore voltage levels after disturbances. Balancing Requirements [27] Frequency Response Services: Maintain grid frequency through reserves. Innovative Frequency Response: Implement advanced frequency control services. Congestion Management [27] Intra-Regional and Cross-Border: Manage congestion within and across regions. Non-Frequency Ancillary Services for Voltage Control and Restoration [27] Reactive Power and Voltage Control: Provide reactive power services and faultride-through capabilities.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 23 System Restoration [27, 31] Black Start and Islanding Operations: Support grid restoration. Adequacy Requirement [27] Capacity Remuneration Mechanisms: Ensure adequate capacity availability. Table 1: Summary of flexibility needs for TSOs and DSOs 2.2.3. Aggregators Aggregators play a vital role by combining the flexible load of multiple consumers and representing them in the electricity market [31]. Household customers' involvement in DR or DSM programs, whether they are solely energy consumers or also producers (e.g., through photovoltaic systems, becoming prosumers), necessitates pooling or aggregation, which in turn requires some level of coordination. They help smaller consumers participate in DR programs by offering services like demand forecasting, energy management, and coordination with grid operators. The entities or organizations that will facilitate this aggregation, known as "aggregators," can either be based within existing market participants such as DSOs and retailers, or operate as independent entities [31]. Aggregators must comply with national regulations and standards, ensuring effective participation in DR programs. They support DSOs and TSOs by providing aggregated demand-side flexibility to enhance grid stability and integrate renewable energy sources [27]. When the power system is under stress, a system operator typically requests DR services. In response, the DR aggregator sends a DR signal to customers who have previously signed incentive-based DR contracts with the aggregator. After the DR event, the system operator needs to pay financial compensations to the DR aggregator according to the load reduction/increment amount achieved by the DR aggregator. Figure 2: Main Traditional Stakeholders in the European Electricity Market, their functions and revenues [31]
Page 24 July 17th 2024 As a summary of this section , figure 2 illustrates the main stake holders in the European electricity markets and how they are co-related to each other. System and grid operators, including TSOs and DSOs, are responsible for the secure transmission and distribution of energy, as well as the maintenance and expansion of the grid. With DSM, DSOs can effectively push down market prices by enabling peak shaving and in turn reduce their own operation and maintenance cost. Whereas TSOs, who have to ensure energy balance and grid stability, can benefit from EF in the sense that there is extra time to start up the ancillary services. Furthermore, through DSM, TSOs can also reduce other costs such as operational cost of power plants that provide frequency reserve services. Retailers buy and sell electricity, earning competitive prices from trade. Major consumers, such as energy aggregators, provide flexibility and plan their consumption, generating revenue by selling flexibility and energy savings. All these players may act as Balance Responsible Parties (BRPs), managing their imbalances in the electricity market in accordance with EU regulations [31]. This thesis primarily focuses on both retailers and consumers, aiming to demonstrate a model that can generate profits for all stakeholders involved. 2.3. Electricity Tariffs Before introducing the topics of DSM and DR, it is important to first form a foundation on the concept of electricity tariffs. While this topic will not be the main focus of the thesis, a brief introduction to electricity tariffs and pricing strategies is essential. Electricity tariffs are an important component of the electricity retail market, significantly influencing the decision-making of end-users and fostering competition within the market [71]. These tariffs not only serve as a basis for expense calculations but also impact load energy management. They are shaped by various factors including the costs of electricity generation, transmission, distribution, and governmental taxation. As a result, electricity pricing can vary widely between different countries and even within regions of the same country. Effective tariff structures are essential for optimizing energy use, reducing power losses, minimizing unnecessary investments, and promoting the use of sustainable energy resources. Conversely, improper tariff methodologies can lead to high power losses, increased operational expenses, and environmental pollution. To effectively incentivize and activate the flexibility potential of the residential sector, it is necessary to talk about certain electricity tariffs that are not only commonly contracted but also relavant to the case of New South Wales, Australia which is the origin of the dataset used for ideating the DR program and Baseline Load (BL) estmation methods. A few types of electricity tarrifs that could be categoried as the following were given in [72] as:
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 25 1. Flat Rate or Single Rate Tariff 2. Block Tariff 3. Time of Use (ToU) Tariff 4. Controlled Load Tariff 5. Demand Charge Tariffs 6. Feed-In Tariffs (FiT) 7. Variable Feed-In Tariffs Among the various tariff types, the ToU tariff is the most common. [71] This static pricebased contract sets rates according to the time of day electricity is consumed. Electricity used during peak periods is charged at a higher rate (c/kWh) compared to off-peak or valley periods. Furthermore, ToU tariffs can vary based on whether the day is a weekday or weekend, and can also differ across seasons. Another interesting type of tariff presented here is the FiT. These tariffs are set at a premium price for the electricity generated and fed back into the grid, making it economically attractive for individuals and businesses to invest in renewable energy technologies such as solar panels, wind turbines, and other forms of renewable energy [71]. Due to the lack of specific information about the tariff types used by the households in the used dataset, it is be assumed that a dynamic pricing contract incorporating both ToU and FiT structures is in place [72]. But as mentioned earlier, The tariff structure for electricity varies significantly between countries, demonstrating a wide range of approaches to billing consumers. Some countries implement tariffs that are almost entirely based on the energy component, while others rely heavily on fixed and capacitybased components as well as combinations of the two in other countries [71].
Page 32 July 17th 2024 In price-based DR (PBDR) programs, participants adjust their electricity usage from typical consumption patterns in response to fluctuating electricity prices. Prices are elevated during peak demand periods, encouraging building operators to either reduce their consumption (i.e., curtail demand) or shift their usage to off-peak times when electricity is cheaper and available in larger quantities. PBDR programs, although devoid of customer privacy and scalability issues, present fairness concerns when a uniform price rate is applied to customers with varying consumption levels, disadvantaging those with naturally lower usage (externality problem) [18]. Additionally, customers may need help keeping track of fluctuating prices over different periods. To effectively manage their load, customers require a scheduling mechanism, whether manual or automated. Explicit DR (Incentive Based DR) Consumers receive financial incentives to reduce or shift their energy usage upon request [18]. Programs such as direct load control, demand bidding, and interruptible load programs can implement this. IBDR programs involve utilities offering financial incentives to customers who allow utility providers to control their electrical loads directly through a contractual agreement. This agreement between the utility provider and the customer provides some degree of authorization to directly schedule, reduce, or disconnect to save costs [21]. Recently, these programs have been extended to residential customers and are managed using centralized controllers, which control both the control decisions and actions [22]. For instance, during a DR event, utility providers may increase the temperature set points of air-conditioning systems to reduce consumption, subsequently returning them to their original settings once the event concludes [23]. This thesis leverages the reward-based nature of IBDR programs to develop a model that examines the impact of various incentive rates on household behavior during peak hours.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 33 4. State-of-the-art Algorithms To estimate the baseline load for residential customers and examine the risks associated with overand under-estimation for both clients and energy aggregators, this thesis leverages Data Science and ML principles, IBDR, and CBL Estimation Methods. This section provides a comprehensive explanation of the technologies utilized in this project, along with an overview of the current approaches employed in the demand response sector. 4.1. Data Science Approach The rapidly expanding discipline of data science [24] enables the effective analysis of vast volumes of data from various sources, facilitating informed decision-making to address issues with digitally recorded data. Data science typically involves the application of advanced statistical methods and scientific principles to extract valuable business insights from data [24]. In the context of this project, the focus shifts to the digitalization aspect of the energy transition. The energy sector is experiencing major transformations, motivated by the necessity for sustainable and efficient energy systems. The digitalization of this sector involves integrating advanced technologies to manage and analyze the massive amounts of data generated by energy production, distribution, and consumption processes [25]. The vast amount of data generated in the energy sector presents an unprecedented opportunity to leverage data science and machine learning. By applying these technologies, we can develop sophisticated models to analyze energy consumption patterns, predict future demands, and optimize energy distribution. These models can provide insights into how energy systems operate, identify inefficiencies, and suggest improvements to enhance overall performance. For instance, ML algorithms offer a practical solution for predicting peak energy usage times, thereby enabling better load management and reducing the risk of blackouts. Similarly, data analytics can help identify trends and patterns in energy usage, which can inform policy decisions and the development of more efficient energy practices. Furthermore, integrating RES, such as wind, solar, and geothermal energy, into the grid necessitates advanced modeling and prediction capabilities to ensure a stable and reliable energy supply. Data science plays a critical role in this by providing accurate forecasts of energy production from renewable sources, which are inherently variable. This thesis leverages the principles of data science and ML and presents a study on determining customer baseline models using various averaging methods such as High
Page 34 July 17th 2024 X of Y days, Low X of Y days, and Mid X of Y days. These methods are widely employed in DR programs to estimate customer load behavior. The study aims to compare these traditional methods with current state-of-the-art machine learning models, including Ridge Regularization, Extreme Gradient Boosting Regression, Random Forest Regression, and Support Vector Regression. The following section provides a detailed description of the baseline estimation methods. These methods will be evaluated through an IBDR mechanism to assess changes in load during a DR event, considering the presence of PV generation.The following section provides a detailed description of the baseline estimation methods. These methods will be evaluated through an IBDR mechanism to assess changes in load during a DR event, considering the presence of PV generation. 4.2. Incentive-Based Demand Response Model (IBDR) As mentioned in section 3, IBDR programs already contribute to flexible demands in the industrial sector, but much less in the residential sector [22]. This section discusses the literature that shows how advanced statistical modeling, microeconomics, and reinforced learning were utilized as the current state-of-the-art to measure customer participation in IBDR systems. While the focus lies on BL estimation of residential households with PV generation for demand response, in this thesis, an IBDR model is developed to assess the load reduction from peak periods. Customers can either choose to reduce their kWh consumption in two formats; one is when they are obliged to do so, as in they fill a power purchasing agreement with the aggregator, or they can choose to reduce their load when signals are sent through. The IBDR model used in this paper would follow the former method to pertain to the scope of the project. To start with, an interesting question that arises would be, "Why would the customers participate in DR programs?". This question was studied by Aman. S & Steven. V in their paper for the willingness of customers to participate in DR programs [33]. They work with a data set of 186 respondents and focus on Belgium's winter electricity peaks, where the central question relates to the readiness to accept limits on the use of home appliances in return for compensation. The findings indicate that the willingness to enroll in a program increases with age, environmental consciousness, home ownership, and lower privacy concerns. Additionally, the analysis predicts that 95% of the surveyed sample would likely enroll in a daily load control program for annual household compensation. The authors conclude that while an initial rollout targeting older, environmentally conscious homeowners could be successful, broader implementation necessitates
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 35 educating the population on the program's environmental and financial benefits and addressing their data privacy concerns. Another study [34] suggests that IBDR policy is heavily required to guide the customer's electricity consumption. The paper suggests that an IBDR policy should consider the implementation-side management factors such as incentives (monetary subsidies) and publicity and response-side household attributes such as demographic characteristics of households, household electricity use scenarios, and household electricity habits and willingness. A policy implementation path was analyzed here, and it showed that residents who participated in DR programs showed a 0.09 kWh reduction in their electricity consumption for a 1.5-hour period compared to the households that did not participate [34]. There are multiple studies contributing reinforced machine learning methods [35, 36, 37, 38] to discern appliance type, customer participation rate, and incentive-based optimization to obtain maximum customer benefits and retailer profits. All these papers helped conclude that IBDR programs are considered reward-wise programs, whereas price-based programs are considered punishment-wise programs. The voluntary nature of reward-wise programs makes people more positive and responsive in the long term. In contrast, the obligatory nature of the punishment-wise program makes people nervous, and responses are more transient. In their paper, "Costs of Demand Response from Residential Customers' Perspective," Safdar, Hussain, and Lehtonen (2019) [39] explore the economic implications of DR programs from the standpoint of residential customers. The study emphasizes the need to balance the benefits of DR with the potential discomfort and inconvenience experienced by customers. Through a customer survey, the authors assess the willingness to accept compensation for shifting energy usage, revealing that customer participation in DR is heavily influenced by perceived compensation and loss of comfort. The paper proposes a linear mathematical model to calculate DR costs, highlighting that these costs are significantly lower than interruption costs, thus advocating for the broader implementation of DR programs [39]. These papers have contributed to understanding the implementation of an IBDR program, which formulates the benefits obtained by the customers as well as the retailer profits. The next section covers the literature on how the IBDR model is formulated.
Page 36 July 17th 2024 4.3. Price Elasticity of Demand A study on price elasticity of demand [78] states that DR is an appealing concept that encourages active customer engagement in the distribution sector through the mechanism of price elasticity of demand (PED). Recent studies support the use of a simple price elasticity model for creating a demand response strategy [78]. highlights that the Price Elasticity Model (PEM) is an appealing and straightforward model for assessing flexible demand in DR. It effectively measures customer demand sensitivity to price variations, which is fundamental for simulating DR events and understanding how changes in price influence electricity consumption. This simplicity and effectiveness make PEM a practical choice for initial DR strategy development. Another related study on the appplication of dynamic price elasticity of demand [79] indicates that PEM is extensively used to model load responses in PBDR, showing its reliability and effectiveness in different contexts. The integration of flexibility through self and cross-elasticity within PEM underscores its capability to handle varying customer responses, reinforcing its utility in DR strategies [79]. Thus, the empirical support from these studies justifies the use of a simple price elasticity of demand response model, providing a robust and adaptable framework for developing effective demand response strategies. With the above context, to create the simulation of a DR event, it is essential to know about the impact of PED also known as Demand Price Elasticity (denoted by 𝜀). It is defined as the percentage change in the demand of a commodity (here, electricity) to 1% change in the percentage change of price of the commodity [64]. The formula for PED is given as: 𝜀=∆𝑄 𝑄 ∆𝑃 𝑃 (1) Here, ∆𝑄 𝑄 represents the percentage change in electricity consumption and ∆𝑃 𝑃 represents the change in retail price of electricity [64]. As shown in Equation 1, PED can have positive or negative values, indicating the correlation between influencing factors and demand. A negative elasticity means the factor is inversely related to demand, while a positive elasticity means they are directly related. • When | 𝜀 | > 1, demand is elastic, indicating high sensitivity to changes in characteristics. • When | 𝜀 | = 1, demand has unit elasticity, meaning changes in characteristics lead to proportional changes in demand.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 37 • When | 𝜀 | < 1, demand is inelastic, showing low sensitivity to changes in characteristics. • When 𝜀 = 0, demand is completely inelastic, unaffected by changes in characteristics. For the IBDR model in this thesis, the evaluation of 𝜀 is crutial as it sets the reference to the amount of load that is to be reduced from the peak hours of the customer load profile. To falicitate profits for the customers and retailers, the time period for the DR event was chosen according to the durations in which | 𝜀 | > 1, which denotes a high sensitivity of the energy load to the sudden change in price, here, Recommended Retail Price (RRP). 4.4. Customer Baseline Estimation In this section, the concept of BL estimation will be thoroughly studied. One of the major elements of evaluating the impact of the DR program is the estimation of the demand shifts or sheds. These estimations are made with Baseline models, which are used to predict the residential load curve if a DR event had not occurred [40]. Although baseline estimations of energy consumption are typically just one component of broader DSM strategies, they are foundational for assessing and measuring DR. They should be prioritized by policymakers and DR program planners. The vast amount of literature on this subject is, therefore, unsurprising [43]. To delve deeper, the importance of accurate baseline calculations cannot be overstated. They enable system operators and utility companies to gauge the success of DR initiatives with precision. By having a clear and accurate baseline, these entities can quantify the actual reduction in energy usage during DR events [5]. This quantification is essential for ensuring that participants in DR programs are compensated appropriately for their efforts to reduce or shift their electricity usage. Such fair compensation is not just a financial matter, but a key factor in maintaining the trust and engagement of participants in these programs, which is a crucial element in their long-term success. Additionally, it offers a predictive measure of the potential demand-side flexibility during DR events. Accurate baseline calculations enable system operators and utility companies to gauge the success of DR initiatives. With a clear baseline, they can quantify the actual reduction in energy usage during DR events and ensure participants receive appropriate compensation [43]. Moreover, dependable baseline data plays a crucial role in predicting demand-side flexibility, which is not only essential for grid stability and planning but also for the integration of RES and the optimization of grid operations. It helps operators
Page 38 July 17th 2024 forecast the amount of load that can be shed or shifted during future DR events, thereby contributing to a more flexible and sustainable energy system. The assessment of baseline consumption is a key factor in evaluating demand response programs. While these baseline estimations are part of more extensive DSM solutions, they are indispensable for the evaluation and measurement of DR, highlighting their critical importance for policymakers and program designers. The following section will provide a description of different types of baseline estimation technologies. 4.4.1. Customer Baseline Load Estimation Methods. This section emphasizes the baseline estimation methods used in the thesis. An important aspect of customer baseline load estimation is the simplicity of the baseline model [41]. The methodologies used should not be intimidating that the customers are concerned over the usage of their load consumption data. On the other hand, complex statistical modeling through regression methods was helpful in determining the accuracy of the baseline models. At the same time, it was much more difficult for stakeholders to understand compared to simple averaging methods [43]. The CBL is an estimate of the electricity consumption that would have occurred if there had been no DR event. Accurate CBL estimation is crucial for evaluating the actual demand reduction achieved during DR events and ensuring fair compensation for participants [41]. The study by O. Valentini et al. [5] on different demand response methods for evaluating customer baseline load proved crucial to the development of this thesis. The baseline models researched in this paper proved a reference point for selecting the nine baseline methods. There are different types of performance evaluation methodologies used to estimate the demand reduction potential for a demand resource, as stated in [43]. These were proposed by the National American Energy Standard Board (NAESB) and included five types of evaluation methodologies: Baseline Type 1, Baseline Type 2, Maximum Baseline Load, Meter Before/Meter After, and Metering Generator Output [5]. To pertain within the scope of the thesis, Baseline Type 1 is selected as it utilizes the historical interval meter data of the demand response and the possibility of using historical weather and electricity price data to generate a basic load profile, which will be explained in section 6. Through the literature reviewed in this paper, baseline estimation technologies can be segregated into the following types: i. Averaging Methods ii. Regression Methods iii. Control Group Methods
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 39 iv. Machine Learning Methods In the sections below, different methodologies studied on behalf of the baseline estimation techniques are discussed. 4.4.2. Averaging Methods These methods consist of simple averaging calculations which are straightforward and can be understood by their naming. These methods were studied due to their vast use cases and simplicity in implementation. The methods use historical data over short durations preceding a DR event and calculate the CBL using non-DR days. Usually, it starts off with the selection of reference days (denoted as Y) from which non-DR event days (referred to as X) are selected according to the respective criteria of the averaging method. From the different papers studied, the X and Y values or days differ according to different application, which can create a varied range of estimations for the CBL [45]. Different approaches to the same averaging method can be found used by different ISOs which are all given in [5]. For example, PJM (Pennsylvania New-Jersey Maryland) Economic uses the High 4 of 5 averaging method whereas NYISO (New York Independent System Operator) chooses to follow High 5 of 10 methods for the same averaging calculation [45, 43, 42]. This provides a larger variety of results that span over different durations. In lieu of this scenario, this thesis showcases how advanced machine learning methodologies would compare against simple averaging methods. 4.4.3. Regression Methods Regression models are essential tools for statistical analysis, offering simplicity and ease of interpretation for linear relationships. These methods employ advanced models to predict baseline load patterns during DR events. These methods aim to determine the most accurate non-linear relationships between selected features—such as temperature, historical load, and other relevant variables—and the baseline load, which acts as the dependent variable. Unlike simple averaging or similar-day methods, regression techniques can incorporate data from the event day itself, making them more adaptable and potentially more accurate [41, 42]. By integrating real-time event day information, regression models can better account for the specific conditions that influence load patterns, resulting in more precise baseline estimations. This capability is crucial for effectively measuring the impact of DR programs and ensuring fair compensation for load reductions. The complexity of regression models, however, lies in accurately capturing and modeling the intricate relationships between multiple factors and the baseline load, which often requires sophisticated algorithms and robust data sets. However, DR events usually take place during extreme weather conditions which can be difficult to replicate by just simple linear and polynomial regression methods [46].
Page 40 July 17th 2024 4.4.4. Control Group Another widely used technique focuses on selecting the customers with the similar load profile as the customer who underwent DR. These methods are known as control group or cluster-based techniques. A widely used method for clustering is the K-Means clustering to link DR event patterns for participant customers with non-event customer load curves. [5, 8, 42]. While there is an absence of historical load data [47] suggests synchronous load matching as a possible way for probabilistic customer clustering. 4.4.5. Machine Learning As the dependency on AI and ML draws closer with each passing day with different techniques being studied and average customer being aware of the benefits that ML brings to the table [48], this thesis tries to incorporate advanced machine learning models like Ridge Regression, Random Forest Regression (RFR), Extreme Gradient Boosting (XGBoost), and Support Vector Regression (SVR) along with simple averaging models for baseline load estimation. Other advanced methods include using more sophisticated tools such as neural networks combined with machine learning, handling stochastic uncertainty (which is key for probabilistic estimation), and more. Some authors [46] even try to exploit both preevent (historical) and post-event (“future”) dependencies. Probabilistic estimation is another well-explored group of methods. According to a report by Valles et al. [49], utilizing a probabilistic method to understand consumer responsiveness greatly improves the precision and dependability of DR programs. This approach considers the entire spectrum of consumer behavior variability, offering DR providers a solid framework to effectively measure and manage residential consumer flexibility. Furthermore, [47] explores the probabilistic Gaussian process regressions as a way to use the control group and combine (averaging) quantile regression methods for CBL estimation. The main drawback, however, is the fact that there is a need for various dependent data apart from the historical load data, which can be a challenge to gather. 4.5. Evaluation Metrics for Baseline Evaluation The evaluation of the CBL models generally require performance evaluation metrics. The key performace indicators for the BL models are talked about in this section. The evaluation of the CBL models generally require performance evaluation metrics. One of the most common performance metric when estimating the performance of a CBL calculation method is to calculate the mean absolute error (MAE) [5, 41]. It is described as follows: 𝑀𝐴𝐸= ∑|𝑏𝑖(𝑑𝑖,𝑡𝑖)− 𝑙𝑖(𝑑𝑖,𝑡𝑖)| 𝑡 ∈𝑇 |𝐶||𝐷||𝑇| (2)
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 41 where 𝑙𝑖(𝑑𝑖,𝑡𝑖) represents actual consumption and 𝑏𝑖(𝑑𝑖,𝑡𝑖) represents the baseline value at time t. C is the set of all customers, D is the set of all the the days included in the evaluation, and T is the set of all time slots in a day d. The lower the value for MAE is, the higher the accuracy for the method. In addition to this, [41] also describes the selection of a different method to calculate the bias of the CBL calculation methods. A tweak in the formula for MAE that allows for nonabsolute values for calculating the bias of the baseline method. The MAE formula is changed to be the following: 𝐵𝑖𝑎𝑠= ∑(𝑏𝑖(𝑑𝑖,𝑡𝑖)− 𝑙𝑖(𝑑𝑖,𝑡𝑖)) 𝑡 ∈𝑇 |𝐶||𝐷||𝑇| (3) If the bias is positive the baseline method overestimates the consumers’ actual load whereas negative values indicate the method is underestimating the consumers’ actual load. Additionally, Mean Squared Error is also utilised to evaluate the performannce of the different baseline models, as in energy load prediction, occasional large deviations from the actual load can lead to incorrect incentive calculations, potentially undermining the effectiveness and fairness of the demand response program. 𝑀𝑆𝐸=∑(𝑏𝑖(𝑑𝑖,𝑡𝑖)− 𝑙𝑖(𝑑𝑖,𝑡𝑖))2 𝑡 ∈𝑇 |𝐶||𝐷||𝑇| (4) MSE’s emphasis on outliers ensures that these significant deviations are minimized, leading to more reliable and consistent baseline estimations. MSE provides a different perspective on model performance compared to MAE. While MAE gives a straightforward average error, MSE provides insight into the variance of the errors [86]. Evaluating both metrics together can offer a more comprehensive understanding of the model's performance. For instance, if a model has a low MAE but a high MSE, it indicates that while the average error is small, there are some instances of very large errors that could be problematic [86]. All three performance metrics will be used on the DR Event days and the model with the least MAE will be selected as the best fit for that particular event day.
Page 48 July 17th 2024 electricity reduction amount during DR events [32]. Furthermore, this evaluation overlooks the benefits gained by retailers from governing bodies for conducting the IBDR mechanism. This simulation is designed to aid energy planners in raising customer awareness regarding the reliance on non-renewable energy sources and their environmental impact. This can lead to a more informed and sustainable energy consumption behavior among consumers. In this section, the description of the methodology and implementation for simulating a DR event on the selected days is presented. The simulation involves reducing electricity load during peak hours with the aim of evaluating the impact of different incentive rates on load management, retailer profits, and customer benefits. The methodology to implement the IBDR mechanism involved several key steps. Initially, data preparation and filtering were performed by defining the event date and specific hours for peak electricity consumption. The dataset was filtered to focus on the relevant timeframe for the DR event date. Next, incentive rates were established at 30%, 20%, and 10% of the average RRP for the DR event date, and data for the DR period was filtered accordingly. These incentives aimed to encourage consumers to reduce their electricity usage during peak hours. Finally, the price elasticity of demand response was used to determine how responsive demand was to price changes during the DR period. This involved analyzing the percentage change in quantity and price to ensure sufficient variation in the data for a meaningful elasticity value. The simulation of the DR event iterated through different incentive rates, adjusting load based on the calculated elasticity. For each incentive rate, the following steps were performed: incentives were set for peak hours to encourage load reduction; load reduction was calculated based on elasticity and incentives using the formula: 𝐿𝑅=𝑙𝑖∗ ℇ ∗ 𝐼/𝑅𝑅𝑃𝑖 (12) Where, LR is the load reduction, 𝑙𝑖 is the original load at time i, ℇ is the price elasticity of demand, I is the incentive rate, and 𝑅𝑅𝑃𝑖 is the recommended electricity price for time i. The original cost without DR and the reduced cost with DR were computed. Cost savings (CS) were determined as the difference between these costs. Retailer profits (RP) were calculated by subtracting incentive costs from cost savings using the equation: 𝑅𝑃=(∑( 𝑙𝑖∗𝑅𝑅𝑃)−∑(𝑏𝑖∗𝑅𝑅𝑃))−∑(𝐼∗𝐿𝑅𝑖) (13) where, 𝑙𝑖 is the original load at time i, and 𝑏𝑖 is the BL at time iare the electricity consumption paterns before and after the DR event. Customer Benefits (CB) were computed as the sum of customer incentives and savings:
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 49 𝐶𝐵=∑(𝐼∗𝐿𝑅𝑖) + ∑(𝑅𝑅𝑃∗𝐿𝑅𝑖) (14) Where the first half of the equation talks about the received incentives by the customers and the second half explains the savings earned by participating in the DR program. The results for each incentive rate included the load during peak hours, retailer profits, and customer benefits. These results were displayed and analyzed to evaluate the effectiveness of different incentive rates in reducing peak electricity demand and the associated financial impacts on both retailers and customers. By following this methodology, the simulation provided a comprehensive understanding of how incentivebased demand response mechanisms could influence electricity consumption patterns and the economic outcomes for both consumers and retailers. 5.5. Baseline Estimation: Averaging Methods Description of the different CBL estimation methods are given in this section. 5.5.1 to 5.5.4 talk about the simple averaging methods used and 5.4.5 describes the machine learning principles and modelling used for the experiments in this thesis. 5.5.1. High X of Y The High X of Y method involve a few steps starting with choosing the highest X number of days over a period of Y reference days. Once these days are chosen, the average of the load at that timeslot over these X highest days are chosen. For better visualization, the study conducted in [41] defines HighXofY method as: 𝑏𝑖(𝑑,𝑡) = 1𝑋 ∑𝑙𝑖(𝑑,𝑡) 𝑑∈𝐻𝑖𝑔ℎ(𝑋,𝑌,𝑑) (15) Here, 𝑏𝑖(𝑑,𝑡) is the calculated baseline for the defined 𝐻𝑖𝑔ℎ(𝑋,𝑌,𝑑) ⊆ D (Y, d) for a customer i ⊆ C for timeslot t ⊆ T on day d. Most common methods are High5of10 method used by NYISO for weekdays and High4of5 method proposed by PJM Economic for weekdays. These methods have been utilized in this thesis project [41]. 5.5.2. Low X of Y LowXofY is the exact opposite of HighXofY method as it involves choosing the lowest X number of days over a period of Y reference days. But the similarity between these methods involves choosing the average of load at that time slot over these X number of days. 𝐿𝑜𝑤(𝑋,𝑌,𝑑) is defined as: 𝑏𝑖(𝑑,𝑡) = 1𝑋 ∑𝑙𝑖(𝑑,𝑡) 𝑑∈𝐿𝑜𝑤(𝑋,𝑌,𝑑) (16)
Page 50 July 17th 2024 Here, 𝑏𝑖(𝑑,𝑡) is the calculated baseline for the defined 𝐿𝑜𝑤(𝑋,𝑌,𝑑) ⊆ D (Y, d) for a customer i ⊆ C for timeslot t ⊆ T on day d. In this case, Low4of5 and Low5of10 methods were used, and results were studied for different customers in this thesis [42]. 5.5.3. Mid X of Y MidXofY method involves calculating the middle X number of days by dropping the (XY)/2 days with the highest consumption and (X-Y)/2 days with the lowest consumption, whilst retaining the days with the median level of electricity consumption. If MidXofY is defined as Mid(X,Y,d) ⊆ D(Y,d) then: 𝑏𝑖(𝑑,𝑡) = 1𝑋 ∑𝑙𝑖(𝑑,𝑡) 𝑑∈𝑀𝑖𝑑(𝑋,𝑌,𝑑) (17) Here, 𝑏𝑖(𝑑,𝑡) is the calculated baseline for the defined 𝑀𝑖𝑑(𝑋,𝑌,𝑑) ⊆ D (Y, d) for a customer i ⊆ C for timeslot t ⊆ T on day d. For the t, Mid4of6 were used, and results were studied for different customers in this thesis [42]. 5.5.4. Nearest X of Y This method is more straightforward than the previous three methods. To determine nearness, the total electricity consumption throughout the day outside the event window is calculated. From the Y non-demand response days preceding a demand response day, we select the X days with total energy consumption most like that of the demand response day. This method uses the average of the nearest X days to calculate the baseline on the DR Event date. While HighXofY, MidXofY, and LowXofY provide the baseline according to the consumption levels of a non-demand response day by summing the load over a day. The NearestXofY method only considers the total electricity consumption outside of the demand response event window [42]. 5.6. Baseline Estimation: Machine Learning Models This section explains about the different aspects considered while experimenting on the machine learning models. The models used to compare against the baseline methods for their effectiveness are: I. Ridge Regression II. Extreme Gradient Boosting Regression III. Random Forrest Regression IV. Support Vector Regression These models and how they were evaluated are explained in more detail in the subsequent sections.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 51 5.6.1. Splitting the Dataset Splitting the dataset into training and test sets is a fundamental practice in machine learning to ensure effective model training and reliable evaluation. Initially, I considered using a separate validation set from the training set, but it was more efficient to use a time series split with cross-validation [55]. The training set, typically the largest portion (70%), is used to teach the model the patterns and relationships in the data. This large portion ensures the model has sufficient data to learn from, improving its ability to generalize. The test set, comprising 30% of the data, is reserved for evaluating the final model's performance on unseen data. This separation ensures an unbiased assessment of how well the model can generalize to new, real-world data. For instance, with the dataset split after March 14, 2013, the model is trained on earlier data and tested on later data. With 17,520 data points, a 70-30 split results in substantial sets for training (12,264 rows) and testing (5,256 rows). Instead of using a separate validation set, cross validation was employed within the training set to iteratively validate the model during the training process. In a time series context, where future data is unknown during training, it's crucial to simulate real-world conditions. Therefore time series cross-validation was utilized [73], which sequentially splits the training set into different folds while respecting the temporal order of the data as seen in figure 6. Figure 6: Graphical representation of a Time Series Cross Validation Split [74]
Page 52 July 17th 2024 By using the `TimeSeriesSplit` from the `sklearn.model_selection` library [73], the validation process considers the temporal nature of the data. This method performs 10fold cross-validation on the training set, providing average validation metrics such as MAE, MSE, and Bias values. These metrics offer a reliable measure of model performance and help fine-tune hyperparameters without influencing the test set. A con of this method is that future data might leak into the training and validation set that might train the This approach provides a robust and trustworthy evaluation of the machine learning model's results, ensuring they are reliable and comparable against averaging baseline models. 5.6.2. Feature Description Feature selection is a crucial process in machine learning models, as it helps to train the target data effectively by selecting the most relevant features. In my approach to predict net electricity consumption, several features were identified and engineered to provide insights into the factors influencing electricity usage. The weather-related data was sourced from the Open-Meteo Historical Weather API [81], and the RRP of electricity was obtained from the Australian Energy Market Operator (AEMO) [52]. Feature Description Unit temperature Air temperature °C relative_humidity Percentage of humidity in the air % dewpoint Temperature at which air is saturated °C apparent_temp Perceived temperature °C precipitation Amount of precipitation mm rain Amount of rain mm weather_code Weather condition code - pressure_msl Mean sea level pressure hPa surface_pressure Surface level pressure hPa cloud_cover Percentage of cloud cover % wind_speed10 Wind speed at 10 meters m/s wind_speed100 Wind speed at 100 meters m/s isday Indicator for day (1) or night (0) - sun_dur Duration of sunshine s direct_rad Direct solar radiation W/m² diffuse_rad Diffuse solar radiation W/m² direct_nor Direct normal radiation W/m² direct_inst Direct instantaneous radiation W/m²
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 53 diffuse_inst Diffuse instantaneous radiation W/m² nor_instant Normal instantaneous radiation W/m² RRP Recommended Retail Price cAUD/kWh Table 3: Features and Description To enhance the predictive capabilities of the dataset, several additional features have been included. These features are designed to capture temporal patterns, cyclic behaviors, and specific characteristics of the days. Below is a description of these features: Temporal Lagged Variables 1. lagged_48h: The value of the target variable lagged by 48 half hourly values. This captures the value of target a day ago. 2. lagged_96h: The value of the target variable lagged by 96 half hourly values. This captures the value of target two days ago. 3. lagged_week: The value of the target variable lagged by 168 hours. This captures the value of target one week ago. 4. target_diff: The difference between the maximum and minimum values of target over the past week. This captures the variability in target within a week. Cyclic Encoding for Temporal Variables 1. hour_decimal: Represents the hour of the day as a decimal value, where 0 represents midnight and 23.50 represents just before the next midnight. 2. sin_hour: The sine of the hour_decimal, used to capture the cyclic nature of hours in a day. 3. cos_hour: The cosine of the hour_decimal, used to capture the cyclic nature of hours in a day. 4. week_sin: The sine of the day of the week (e.g., Monday = 0, Sunday = 6), used to capture weekly cyclic patterns. 5. week_cos: The cosine of the day of the week, used to capture weekly cyclic patterns.
Page 54 July 17th 2024 6. month_sin: The sine of the month of the year (e.g., January = 0, December = 11), used to capture annual cyclic patterns. 7. month_cos: The cosine of the month of the year, used to capture annual cyclic patterns. Type of Day and Other Features 1. weekend: A binary variable indicating whether the day is a weekend (1) or a weekday (0). This helps capture the difference in patterns between weekends and weekdays. These additional features are included to better capture the temporal dependencies, cyclic behaviors, and specific characteristics of different days, which can significantly improve the accuracy and robustness of predictive models. To further enhance the predictive accuracy of the model, the SelectKBest method combined with f_regression from the sklearn.feature_selection library was employed [54]. This feature selection technique evaluates each feature's statistical significance in predicting the target variable, net electricity consumption. By ranking and selecting the top features, only the most impactful variables were included in the final model. 5.6.3. Ridge Regression Ridge regression, also refered to as L2 Regularization, is a crucial technique in machine learning, essential for developing robust models in scenarios susceptible to overfitting and multicollinearity. This method modifies standard linear regression by incorporating a penalty term based on the squared values of the coefficients, which is particularly beneficial when handling highly correlated independent variables [53]. The function for Ridge regression is defined as: min 𝑤(||𝑋𝑤−𝑦||22+𝛼||𝑤||22) (18) Here, ||𝑋𝑤−𝑦||22 refers to the residual sum of squares (RSS) which is the squared difference between the residual value y and predicted values Xw. To understand this better, X is the matrix of inputs, w is the vector of coefficients, including the intercepts, and y is the vector of observed values. The regularization part of this model is 𝛼||𝑤||22, which adds the penalty for large coefficients where 𝛼 (sometimes referred to as ‘λ’) is the regularization parameter, which regulates the amount of shrinkage to the parameters. This penalty term prevents overfitting [53].
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 55 The elliptical contours (red circles) seen in Figure 5 represent the cost functions. The constraints introduced by the penalty factor form a circular region (green circle) around the origin. Relaxing these constraints increases the constrained region. The coefficients are determined by finding the first point where the elliptical contours intersect the circular constraint region. This results in the coefficients being shrunk towards zero but never reaching exactly zero, unlike in lasso regression. This approach mitigates overfitting by imposing a penalty on the size of the coefficients, leading to more robust models [65]. Among its primary advantages, ridge regression significantly reduces overfitting by imposing complexity penalties, manages multicollinearity by balancing effects among correlated variables, and improves model generalization to enhance performance on unseen data. By shrinking the regression coefficients through a regularization term, ridge regression minimizes the residual sum of squares while introducing a bias-variance tradeoff that stabilizes the model [53]. This makes it especially effective in highdimensional datasets where the number of predictors exceeds the number of observations. Additionally, the optimal regularization parameter ‘λ’ can be determined through cross-validation, ensuring that the model achieves the best possible balance between bias and variance. 5.6.4. Extreme Gradient Boosting Regression The concept of gradient boosting was introduced by Jerome Friedman in 1999 and has since become a fundamental tool in the data scientist's toolkit. [56] Whereas, the relatively new approach of Extreme Gradient Boosting (XGBoost) was more recently Figure 7: Visualizing Ridge Regression [65]
Page 56 July 17th 2024 introduced by Chen and Guestrin in 2016 [57]. XGBoost is an avant machine learning algorithm that enhances the performance of predictive models through the technique of gradient boosting. It is an optimized version of gradient boosting that includes additional features like regularization, parallel processing, and tree pruning. These enhancements make XGBoost faster and more efficient while improving model performance and robustness. To understand XGBoost further, there is a need to understand the concept of boosting and how gradient boosting works. Boosting is an ensemble learning stratergy that combines the strengths of multiple weak learners (usually decision trees) to create a robust predictive model. Gradient boosting is a specific type of boosting that optimizes the model by reducing the residual errors (gradients) of the previous models. This is done by fitting new models to the gradient of the loss function relative to the predictions. The function for XGBoost is used to minimize regularized objectives, as shown [57]: ℒ(𝜙)= ∑𝑙(𝑦𝑖𝑦𝑖) 𝑖+ ∑Ω(𝑓𝑘) 𝑘 (19) Here, 𝑙(𝑦𝑖𝑦𝑖) is the loss function for each observation, i [56] and Ω(𝑓𝑘) is the regularization term for each tree, k. Ω penalizes the complexity of the model [57]. Applying regularization helps to control the complexity of the trees and prune branches that do not contribute significantly to the model, therefore enhancing generalization. 5.6.5. Random Forrest Regression Random Forest Regression (RFR) is an advanced ensemble learning technique that enhances the basic decision tree algorithm to improve prediction accuracy and robustness. Developed by Leo Breiman in 2001, [83] this method constructs numerous decision trees during training, with the final prediction being the mean of the individual tree predictions. At the heart of RFR lies the decision tree algorithm [80]. Decision trees split data based on input feature values, creating branches until a prediction is made. However, they can overfit, modeling noise instead of the true data distribution. RFR plays a crucial role in mitigating this issue by using bagging (Bootstrap Aggregating), where multiple models (decision trees) are trained on different data subsets, and their predictions are aggregated. The overall prediction for a given input x is the average of the predictions from all individual trees in the forest. Mathematically, this can be expressed as: 𝑦=1 𝑁∑𝑇𝑖(𝑥) 𝑁 𝑖=1 (20) where, 𝑦 is the predicted value, N is the total number of decision trees, 𝑇𝑖(𝑥) is the prediction from the i-th decision tree for the input x [83]. RFR is highly accurate and robust, reducing overfitting by averaging multiple decision trees. It can handle numerous
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 57 input variables, provides insights into feature importance, and is versatile for both classification and regression tasks. However, it can be computationally intensive and requires significant memory and processing power. The model's interpretability is lower compared to single decision trees, and it may struggle with highly imbalanced datasets [80]. 5.6.6. Support Vector Machines Support Vector Regression (SVR) is a supervised machine learning algorithm that is used for regression tasks. It is an extension of the Support Vector Machine (SVM), which is typically used for classification. SVR applies the same principles as SVM but in the context of predicting continuous values rather than discrete class labels [84]. The primary goal of SVR is to find a function that approximates the relationship between the input features x and the output variable y, within a certain error margin. SVR aims to fit the best line within a threshold value, ensuring that the deviations of the predictions are within an acceptable range, defined by the parameter 𝜖. SVR can handle non-linear relationships by using the kernel trick. Kernels transform the input data into a higherdimensional space where a linear separation (or regression) is more feasible. Common kernels include linear, polynomial, and radial basis function (RBF) [85]. The final model generated by SVR is represented as: 𝑓(𝑥)=∑ (𝛼𝑖− 𝛼𝑖∗)𝐾(𝑥𝑖,𝑥)+𝑏 𝑁 𝑖=1 (21) Where 𝛼𝑖,𝛼𝑖∗ are the Lagrange multipliers, 𝐾(𝑥𝑖,𝑥) is the kernel function and b is the bias term [85]. SVR is effective in high-dimensional spaces and works well with both linear and non-linear data using kernel functions. It maximizes the margin, enhancing generalization, and is robust to outliers with its ϵ-insensitive loss function. However, SVR can be computationally intensive, especially for large datasets, and selecting appropriate kernels and hyperparameters can be complex. It is also sensitive to the regularization parameter C and kernel parameters [84].
Page 64 July 17th 2024 From the visualization of error metrics in Figure 13, it can be suggested that, RFR model consistently demonstrates strong performance across various dates and metrics. For June 24th, SVR has the lowest MAE at 0.60, indicating high accuracy, while XGBoost excels in minimizing MSE with a value of 1.07, and Ridge has the lowest Bias at 0.06. On June 20, RFR stands out with the lowest MAE (1.29) and MSE (4.19), though XGBoost has the least Bias (-1.03). For May 22, RFR again leads with the lowest MAE (0.85), MSE (1.75), and Bias (-0.63), making it the most reliable model overall. Therefore, the SVR model is the best for June 24, while the RFR model is the best for both June 20 and May 22. Figure 13: Error Metrics visualization for Customer 65 for selected DR event dates: (i) MAE (ii) MSE (ii) Bias
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 65 Figure 14: Scatter Plot for customer 65: 2013-06-24 Figure 15: Scatter Plot for customer 65: 2013-06-20 Figure 16: Scatter Plot for Customer 65: 2013-05-22
Page 66 July 17th 2024 Figures 14 to 16 highlight that for the dates June 20th and May 22nd, RFR outperforms other models. However, for June 24th, the trend shifts towards SVR showing superior performance. To assess the evaluation metrics for NC and TC for resident 65, the MAE, MSE, and Bias were avergaed over the three DR days. NC TC Model MAE MSE Bias MAE MSE Bias Ridge 1.192 3.083 -0.745 1.165 2.96 -0.717 XGBoost 1.072 2.65 -0.531 0.998 2.305 -0.428 Random Forest 0.985 2.391 -0.523 0.954 2.301 -0.493 SVR 0.989 2.777 -0.789 1.068 3.083 -0.838 PJM High 4x5 1.31 3.549 -0.383 1.297 3.477 -0.283 NYISO High 5x10 1.249 3.327 -0.387 1.245 3.219 -0.313 Low 5x10 1.383 4.786 -1.139 1.356 4.574 -1.051 Mid 4x6 1.327 4.248 -0.871 1.313 4.069 -0.779 Nearest 5x10 1.29 3.527 -0.502 1.275 3.418 -0.402 Table 7: Comparing NC vs TC for Customer 65 This comparison reveals that overall, the models perform better with TC compared to NC. This is evidenced by generally lower values of MAE and MSE and less bias in the TC scenario. While the RFR model consistently performs well in both scenarios, it exhibits slightly better accuracy and error minimization with TC. Now that the best metrics for each DR event day is found, lets visualize the estimated CBL curves along with the actual consumption. Figure 17: Visualizing the Best Model for 2013-06-24
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 67 Figure 18: Visualizing the best model for 2013-06-20 Figure 19: Visualizing the best model for 2013-05-22 Shifting the attention to second customer, Customer 209, the same methodology as above is followed here as well.
Page 68 July 17th 2024 Figure 20: Error Metrics visualization for Customer 209 for selected DR event days(i) MAE (ii) MSE (ii) Bias Following a similar approach as customer 65, customer 209 is evaluated based on MAE, MSE and Bias over the three selected DR event days. On June 24, XGBoost outperformed all other models in terms of MAE (0.3), MSE (0.15), and Bias (0.05), showcasing its robustness across all evaluation metrics. On June 20, the RFR model demonstrated the lowest MAE (0.37), while XGBoost excelled in both MSE (0.37) and Bias (-0.11). Nevertheless, RFR showed better evalauation metrics indicating superior overall performance. For May 22, RFR again led in MAE, Bias, as well as achieving the lowest MSE. Overall, RFR consistently proved to be the best model for this customer, exhibiting lower prediction errors and balanced estimations across different dates and metrics. XGBoost also performed well but generally ranked second to RFR, making RFR the most reliable choice for minimizing prediction errors for resident 209.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 69 The scatter plots shown in figure 21 to 23 reiterate the above analysis of random forest regression performing better for June 20th and May 22nd. Figure 21: Scatter Plot for Customer 209: 2013-06-24 Figure 22: Scatter Plot for Customer 209: 2013-06-20
Page 70 July 17th 2024 Figure 23: Scatter Plot for Customer 209: 2013-06-25 Now comparing the NC vs TC of customer 209, the same methodology as before is followed again. From Table 8, it can be deduced that the models tend to perform better with TC compared to NC. This is indicated by generally lower MAE and MSE values and less bias in the TC scenario. The RFR model consistently performs well in both NC and TC but shows slightly better accuracy and error minimization with TC. XGBoost also performs well, particularly in reducing Bias. Therefore, for the new customer, TC provides more reliable and accurate model performance. NC TC Model MAE MSE Bias MAE MSE Bias Ridge 0.655 0.778 -0.180 0.692 0.843 -0.187 XGBoost 0.396 0.402 -0.142 0.370 0.340 -0.149 Random Forest 0.381 0.359 -0.105 0.358 0.306 -0.095 SVR 0.462 0.491 -0.192 0.451 0.507 -0.229 PJM High 4x5 0.696 0.992 -0.152 0.675 0.938 -0.107 NYISO High 5x10 0.718 1.106 -0.115 0.716 1.092 -0.082 Low 5x10 0.727 1.173 -0.396 0.685 1.094 -0.357 Mid 4x6 0.696 1.098 -0.288 0.676 1.037 -0.247 Nearest 5x10 0.655 0.906 -0.202 0.637 0.854 -0.156 Table 8: Comparison of Evaluation metrics for NC vs TC for Customer 209 Again, similar methodology to the previous customer is used to visualize the CBL estimation against the original load will be presented.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 71 Figure 24: Visualizing the best model for 2013-06-24 Figure 25: Visualizing the best model for 2013-06-20
Page 72 July 17th 2024 Figure 26: Visualizing the best model for 2013-05-22 Now the result for customers 65 and 209, who where in the high consumption category, are compared with customer 174. For customer 174, the error metrics show significantly better results compared to the previous higher consumption customers. This improvement can be attributed to the lower load curve of this customer, which simplifies the CBL estimation and reduces the need for complex models. For this low consumption household, the errors (MAE, MSE, and Bias) are generally lower, indicating that predicting energy consumption is easier and more accurate. Models like SVR and RFR perform exceptionally well, underscoring their reliability in low consumption scenarios. This suggests that simpler models may suffice for accurately predicting consumption in such households.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 73 Figure 27: Visualizing the error metrics for customer 174 (i) MAE (ii) MSE (iii) Bias For this low consumption household, as seen in figure 27, the errors (MAE, MSE, and Bias) are generally lower as compared to the other customers, indicating that predicting energy consumption is easier and more accurate. Models like SVR and RFR perform exceptionally well, underscoring their reliability in low consumption scenarios. Rigth behind these complex machine learning models, simple averaging methods such as Mid4of5 and Nearest 5of10 also show good results regarding the MAE. This suggests that simpler models may suffice for accurately predicting consumption in such households [41]. The load curves for 174 are visualised below in Figures 28-30
Page 80 July 17th 2024 8.2. Environmental Impact The environmental impact of the project is discussed in this section. The project was conducted in Barcelona, Spain, primarily at the UPC ETSEIB Library during business hours and at the author’s residence during other times. Notably, there was no need for travel as part of the project work, thus minimizing additional carbon emissions. The carbon intensity for Spain, as published by the National Commission on Markets and Competition (CNMC) on April 25th, 2024, is 260 gCO2e/kWh [59]. This figure will be used to evaluate the carbon footprint of the thesis project. 8.2.1. Equipment and Carbon Footprint The equipment used for the fulfillment of the master’s thesis, along with their energy consumption and carbon footprint, are summarized in the following table: Equipment Nominal Power (W) Hours Used (hrs) Total Energy Consumption (kWh) Carbon Emmisions (kgCO2e) Laptop (fully charged) 60 600 36 9.360 Lighting (LED) 10 600 6 1.560 WiFi Router 20 600 12 3.120 Total 54 14.04 Table 9: Environmental Impact Assessment From table, the total energy consumed as well as the carbon emissions associated to utilizing electricity has been deduced to 54 kWh and 14.04 kgCO2e, respectively. 8.2.2. Renewable Energy Context In 2023, 54% of Spain's energy was produced from renewable sources, with wind energy accounting for 24.2% and coal usage at just 1.6%. This positions Spain as one of the leading countries in renewable energy production, demonstrating a significant and ongoing shift towards incorporating renewable energy sources into its energy mix [59]. Although Spain's carbon intensity is higher compared to a country like Sweden, the thesis project carried out in Spain did not result in a substantial carbon footprint.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 81 Conducting this research in Spain, as opposed to my home country, India, resulted in a lower environmental impact. Spain's substantial use of renewable energy sources significantly mitigated the project's carbon emissions. In contrast, had this project been conducted in India, where the energy mix relies more heavily on fossil fuels, the carbon footprint would likely have been much higher. This comparative advantage underscores the environmental benefits of conducting research in regions with a higher proportion of renewable energy in their power generation mix, leading to more sustainable research practices. 8.3. Social impact and gender equality This project is committed to fostering a positive social and environmental impact, ensuring inclusivity, and avoiding any form of discrimination. From the outset, a concerted effort was made to maintain an equitable approach, ensuring that no bias or discrimination based on gender, race, or social status occurred. The project's primary goal has been to highlight and enhance the positive effects of DR models. By focusing on accurate CBL estimation, the project aims to ensure fair compensation for residential customers, promoting equitable energy solutions for all. The Ausgrid dataset, which is openly available to everyone, was utilized in this research. This open-access data is crucial for transparency and inclusivity, allowing companies and researchers from around the world to engage with and build upon the findings. The employment of baseline estimation methods and Incentive-Based Demand Response IBDR mechanisms within this project emphasizes energy efficiency and effective resource management. By leveraging these methods, the project demonstrates how energy consumption can be optimized, leading to significant reductions in electricity usage and costs. This approach not only promotes energy efficiency but also supports the creation of a more sustainable and resilient energy system. Moreover, this research aligns with the United Nations Sustainable Development Goals (SDGs), as illustrated in Figure [60]. Specifically, it addresses: • SDG 7: Affordable and Clean Energy: By promoting energy efficiency and incentivizing reductions in energy use, the project contributes to making energy more affordable and cleaner [60]. • SDG 11: Sustainable Cities and Communities: The project fosters sustainable urban environments by advocating for energy-saving measures and efficient energy use [60].
Page 82 July 17th 2024 • SDG 13: Climate Action: By enhancing energy efficiency and reducing carbon emissions, the project directly supports efforts to combat climate change [60]. Figure 32: SDG Goals [88] Overall, this thesis aims to have a positive impact on society and the environment by addressing key sustainability goals and promoting a fair and efficient energy system.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 83 9. Conclusions The experiments conducted for this thesis aimed to evaluate the effectiveness of the proposed IBDR mechanism and compare various CBL estimation methods. This study offers a comprehensive overview of state-of-the-art algorithms tailored to different DR programs, encompassing both simple and complex baseline estimation methods, such as averaging techniques and machine learning regression methods. The primary dataset used in this research was sourced from the Ausgrid network, providing detailed electricity consumption data for residential customers with BTM PV installations. Additionally, to identify relevant features for the machine learning models, weather and electricity price data were obtained from Open Meteo [81] as well as AEMO [52]. This combination of data sources enabled a thorough analysis and comparison of the different CBL estimation techniques within the context of the proposed IBDR mechanism. 9.1. IBDR Mechanism Addressing the first research question mentioned in the motivation, How can residential customers be incentivized to shift their electricity consumption away from peak periods? The proposed IBDR mechanism demonstrated a significant potential for load reduction, especially among customers with higher consumption levels. Customer 65, in particular, benefited the most from the provided incentives Althought not as high, it aligned with the common use of IBDR mechanisms in the industrial sector [22]. This finding underscores the importance of incentivizing higher consumption customers to maximize the benefits of demand response programs. The lowest change in consumption was found in customer 174 which can be attributed to the fact that this customer already has low consumption. This could highlight one of the dependencies of the IBDR model to the price elasticity of demand. his highlights one of the dependencies of the IBDR model on the price elasticity of demand. Due to the large bias towards customers who have higher demand during peak periods, the effective capture of load reduction using a simple price elasticity model may not be as representative. However one of the important use case of price elasticity of demand could be the ability to be able to peak hous of electricty durign the day. The challenge lies in the significant data requirements, which include electricity prices, tariffs, and predicted customer behavior. Reflecting on the IBDR mechanism, it becomes clear that energy planners must prioritize educating consumers about peak hours. By effectively shifting consumption patterns away from these periods, significant benefits can be realized. Enhanced grid stability and reduced reliance on fossil fuels are two major advantages, contributing to a more sustainable energy system. The current imperative is to foster greater customer
Page 84 July 17th 2024 participation in programs related to DSM and DR. Encouraging such involvement will not only help balance the grid more efficiently but also pave the way for broader adoption of renewable energy sources and smarter energy consumption habits. 9.2. Performance of Baseline Models. To answer the second research question introduced in the thesis— " Which is the best method to estimate what the residential customer's load would have been in the absence of a Demand Response (DR) event? " —several experiments and analyses were conducted. The experiments revealed that the RFR model consistently provided the best fit for a range of DR event day load curves. It was deemed the best model in 8 out of 18 experimental results, proving to be a robust method for baseline estimation in demand response contexts. This leads to a comparison between the performance of simple averaging methods and machine learning models. The superior performance of machine learning models is not surprising since these models are trained on various features, making them heavily dependent on external factors such as weather conditions and electricity prices. In contrast, simple averaging methods show promise due to their minimal requirement for additional data beyond the historical consumption of households. However, averaging methods are not designed to predict the underlying trends and patterns of customers with high stochasticity. This conclusion is evident from the final results for customer 174, where the evaluation metrics were significantly better compared to customers 65 and 209. There were instances where XGBoost, RFR and SVR models achieved the best MAE values. This variability aligns with the No Free Lunch (NFL) theory [81], which posits that no single optimization algorithm is universally superior across all problem types. This underscores the need for adaptive approaches that leverage the strengths of various models under different conditions. Comparisons between NC and TC revealed significant insights into baseline estimation and model performance. For customers 65 and 209, conducting CBL estimation using TC consistently yielded better MAE, MSE, and Bias metrics. Including TC, which accounts for both general consumption and controlled loads, enhances the accuracy of baseline models. The intermittent nature of PV generation introduces significant variability in energy data, making NC, which nets general consumption minus gross generation, less reliable. In contrast, TC, including all energy consumption data without subtracting PV generation, provides a more stable and comprehensive dataset for model training and evaluation. For households with BTM PV installations, this underscores the importance of PV disaggregation.
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 85 Specific Case Analysis: The analysis of the DR event on June 24th, 2013, for customer 274, who had lower consumption compared to the other two customers, showed distinct results. This case highlighted the effectiveness of baseline models in handling lower consumption profiles, demonstrating that the IBDR mechanism could still be beneficial even for customers with lower overall energy usage. The choice of June 24th, 2013, as a later date in the dataset provided more accurate baseline estimations due to the availability of similar days for comparison. This further validated the models’ ability to predict consumption patterns effectively, enhancing the reliability of the DR programs. 9.3. Implications and Future Directives This section presents several key implications and future directives based on the research findings. These points are intended to guide future studies and practical applications in demand response and energy management. • The findings highlight the importance of focusing on high consumption customers with demand response incentives to maximize load reduction and financial benefits. Since IBDR mechanisms have been commonly used in the industrial sector, targeting high consumption residential customers can ensure that efforts are directed towards areas with the highest potential impact. • The variability in model performance, as indicated by the No Free Lunch (NFL) theory, suggests that future research should investigate hybrid or ensemble approaches. Combining multiple models could adapt better to varying conditions and improve overall performance. • Energy planners must prioritize effective communication and education about peak hour consumption to enhance grid stability and promote sustainable energy practices. By focusing on these aspects, they can design more effective demand response programs, which in turn fosters better consumer behavior and grid management. • The need for PV disaggregation and the preference for using TC in baseline estimations highlight the complexities involved in integrating renewable energy sources into the grid. Further studies can cater towards dissaggregating the energy generationg from “prosumer” households for better BL estimations during peak hours.
Page 86 July 17th 2024 The code developed for the IBDR mechanism and for the CBL estimation methods can be found in the GitHub repository below: https://github.com/kevinbinzv/Master_Thesis.git
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 87 10. Acknowledgments To Dad, Mom, and Karen, your unwavering support over the past four months has been the cornerstone of my academic journey. Your belief in my abilities and your constant encouragement have been invaluable, and for this, I am deeply grateful. This thesis represents a significant milestone in my pursuit of the EIT Innoenergy Double Degree Master Program in Renewable Energy at Universitat Politècnica de Catalunya and KTH Royal Institute of Technology. I would like to extend my heartfelt gratitude to Dr. Mònica Aragüés Peñalba and Marc Jené Vinuesa. Thank you for placing your trust in me and allowing me to conduct my Master Thesis Research under your guidance. As someone newly captivated by Data Science and Artificial Intelligence, I hope I have met your expectations with this opportunity. My journey into Big Data and Data Science began with Dr. Mònica Aragüés Peñalba's course, Data Science for Electrical Engineering at UPC. The insights and knowledge I gained from this course inspired me to delve deeper into the field. Marc, your invaluable insights, support, and guidance throughout this undertaking have been indispensable. I am profoundly grateful to both Mònica and Marc for this incredible opportunity. Lastly, to my friends, your constant check-ins and unwavering support have meant the world to me. Your encouragement has kept me going through the toughest times. To all of you, I extend my deepest thanks.
Page 88 July 17th 2024 11. Bibliographic References [1] Rongling Li, Andrew Satchwell, Donal Finn, Toke Haunstrup Christensen, Michaël Kummert, et al.. Ten questions concerning energy flexibility in buildings. Building and Environment, 2022, 223, pp.109461. ⟨10.1016/j.buildenv.2022.109461⟩. ⟨hal03839586⟩ [Accessed: April. 24, 2024]. [2] "Annex 67," Annex 67, [Online]. Available: https://www.annex67.org. [Accessed: April. 24, 2024]. [3] O. Zinaman, M. Miller, A. Adil, D. Arent, J. Cochran, R. Vora, S. Aggarwal, M. Bipath, C. Linvill, A. David, M. Futch, R. Kaufman, E. V. Arcos, J. M. Valenzuela, E. Martinot, D. Noll, M. Bazilian, and R. K. Pillai, Power Systems of the Future. The Electricity Journal, vol.28, no.2, pp.113-126, 2015. https://doi.org/10.1016/j.tej.2015.02.006. [Accessed: April. 24, 2024]. [4] Minou, Marilena & Thanos, George & Vasirani, Matteo & Ganu, Tanuja & Jain, Mohit & Gylling, Arne. (2014). Evaluating Demand Response Programs: Getting the Key Performance Indicators Right. [Accessed: April. 24, 2024]. [5] Valentini, O.; Andreadou, N.; Bertoldi, P.; Lucas, A.; Saviuc, I.; Kotsakis, E. Demand Response Impact Evaluation: A Review of Methods for Estimating the Customer Baseline Load. Energies 2022, 15, 5259. https://doi.org/10.3390/en15145259 [Accessed: May 26, 2024.] [6] Knayer, T., & Kryvinska, N. (2022). An analysis of smart meter technologies for efficient energy management in households and organizations. Energy Reports, 8, 785-799. https://doi.org/10.1016/j.egyr.2022.03.041 [Accessed: April. 26, 2024]. [7] B. Völker, A. Reinhardt, A. Faustine, and L. Pereira, “Watt’s up at home? Smart meter data analytics from a consumer-centric perspective,” Energies, vol. 14, no. 3. MDPI AG, Feb. 01, 2021. doi: 10.3390/en14030719. [Accessed: April. 26, 2024]. [8] S. Hasan, "Beyond Smart Meters," KAPSARC, Jul. 11, 2021. [Online]. Available: https://www.kapsarc.org/research/publications/beyond-smart-meters/. [Accessed: April. 26, 2024]. [9] Advanced Metering Infrastructure (AMI): Smart Meters and New Technologies. from: https://www.researchgate.net/publication/332111353_Advanced_Metering_Infr astructure_AMI_Smart_Meters_and_New_Technologies [Accessed: April. 26, 2024]. [10] Driivz. "Demand Side Management (DSM)." Driivz https://driivz.com/glossary/demand-side-
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 89 management/#:~:text=Demand%20Side%20Management%20(DSM)%20is,monetar y%20incentive%20to%20reduce%20demand. [Accessed: May. 6, 2024]. [11] U.S. Energy Information Administration, "Demand-side management programs save energy and reduce peak demand," EIA, [Online]. Available: https://www.eia.gov/electricity/data/eia861/dsm/. [Accessed: May. 6, 2024]. [12] Neukomm M, Nubbe V, Fares R. Grid-interactive efficient buildings technical report series: overview of research challenges and gaps. U.S. Department of Energy; 2019. [Accessed: May. 26, 2024]. [13] Exro. (n.d.). Peak Shaving: Optimize Power Consumption with Battery Energy Storage Systems. Retrieved from https://www.exro.com/peak-shaving-optimizepower-consumption [Accessed: May. 26, 2024]. [14] Rwegasira, Diana & Ben Dhaou, Imed & Kondoro, Aron & Kelati, Amleset & Tenhunen, Hannu & Mvungi, Nerey. (2018). Load-shedding techniques: A comprehensive review. International Journal of Smart Grid and Clean Energy. 8. 10.12720/sgce.8.3.341-353. [Accessed: April. 6, 2024]. [15] Sinha, Alpana & De, Mala. (2016). Load shifting technique for reduction of peak generation capacity requirement in smart grid. 1-5. 10.1109/ICPEICES.2016.7853528. [Accessed: May. 26, 2024]. [16] Ibrahim, Ismail & Rigoni, Valentin & O'Loughlin, Cathal & O'Donnell, Terence. (2017). Real-Time Simulation Platform for Evaluation of Frequency Support from Distributed Demand Response. [Accessed: May. 26, 2024]. [17] Li H, Wang Z, Hong T, Piette MA. Energy flexibility of residential buildings: a systematic review of characterization and quantification methods and applications. Adv Appl Energy 2021;3: 00054. Available from: https://doi.org/10.1016/J.ADAPEN.2021.100054. [Accessed: April. 03, 2024]. [18] Farshid Shariatzadeh, Paras Mandal, Anurag K. Srivastava, Demand response for sustainable energy systems: A review, application and implementation strategy, Renewable and Sustainable Energy Reviews, Volume 45, 2015, Pages 343-350, ISSN 1364-0321, https://doi.org/10.1016/j.rser.2015.01.062. [Accessed: March. 26, 2024]. [19] Yoshiki Shimomura, Yutaro Nemoto, Fumiya Akasaka, Ryosuke Chiba, Koji Kimita, A method for designing customer-oriented demand response aggregation service, CIRP
Page 96 July 17th 2024 [70] Empire State Building, "Empire State Building Retrofit White Paper," 2010. [Online]. Available: https://www.esbnyc.com/sites/default/files/esb_retrofit_white_paper_090910_1.pdf. [Accessed: 24-March-2024]. [71] D. A. Zaki and M. Hamdy, "A Review of Electricity Tariffs and Enabling Solutions for Optimal Energy Management," Energies, vol. 15, no. 22, p. 8527, 2022. doi: 10.3390/en15228527. [Accessed: May. 15, 2024]. [72] Next Kraftwerke, "What is Peak Shaving?" [Online]. Available: https://www.nextkraftwerke.com/knowledge/what-is-peak-shaving. [Accessed: May. 15, 2024]. [73] "sklearn.model_selection.TimeSeriesSplit," scikit-learn: Machine Learning in Python, [Online]. Available: https://scikitlearn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html. [Accessed: March. 7, 2024]. [74] S. Das, "Cross-Validation in Time Series," Medium, Mar. 2021. [Online]. Available: https://medium.com/@soumyachess1496/cross-validation-in-time-series566ae4981ce4. [Accessed: March. 7, 2024]. [75] P. Haessig, "ausgrid-solar-data," GitHub repository, 2021. [Online]. Available: https://github.com/pierre-haessig/ausgrid-solar-data. [Accessed: Jun. 15, 2024]. [76] NSW Department of Planning and Environment, "Controlled Load Profiles and Sample Meters: Position Paper," Apr. 2024. [Online]. Available: https://www.energy.nsw.gov.au/sites/default/files/2024-04/202404-DCCEEWposition-paper-controlled-load-profiles-and-sample-meters.pdf. [Accessed: Jun. 15, 2024]. [77] M. -X. Zhu and Y. -H. Shao, "Classification by Estimating the Cumulative Distribution Function for Small Data," in IEEE Access, vol. 11, pp. 41142-41157, 2023, doi: 10.1109/ACCESS.2023.3269504. [Accessed: Jun. 16, 2024]. [78] V. C. Pandey, N. Gupta, K. R. Niazi, A. Swarnkar, and R. A. Thokar, "An adaptive demand response framework using price elasticity model in distribution networks," Electric Power Systems Research, vol. 202, 2022, Art. no. 107597. Available: https://doi.org/10.1016/j.epsr.2021.107597 [Accessed: Jun. 16, 2024]. [79] G. Kansal and R. Tiwari, "Elasticity modelling of price-based demand response
Baseline Load Estimation of Residential Customers for Incentive Based Demand Response Programs Page. 97 programs considering customer’s different behavioural patterns," Electric Power Systems Research. [Accessed: Jun. 16, 2024]. [80] "sklearn.ensemble.RandomForestRegressor," scikit-learn: Machine Learning in Python, [Online]. Available: https://scikitlearn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.htm l. [Accessed: March. 7, 2024]. [81] "Historical Weather API," Open-Meteo, 2024. [Online]. Available: https://openmeteo.com/ [Accessed: March. 7, 2024]. [82] "No free lunch theorem," Wikipedia, The Free Encyclopedia, Apr. 2024. [Online]. Available: https://en.wikipedia.org/wiki/No_free_lunch_theorem. [Accessed: Jun. 24, 2024]. [83] L. Breiman, "Random Forests," Machine Learning, vol. 45, no. 1, pp. 5-32, 2001. [Online]. Available: https://link.springer.com/article/10.1023/A:1010933404324 [Accessed: March. 7, 2024]. [84] C. Cortes and V. Vapnik, "Support-Vector Networks," Machine Learning, vol. 20, no. 3, pp. 273-297, 1995. [Online]. Available: https://link.springer.com/article/10.1007/BF00994018. [Accessed: March. 7, 2024]. [85] "sklearn.svm.SVR," scikit-learn: Machine Learning in Python, [Online]. Available: https://scikit-learn.org/stable/modules/generated/sklearn.svm.SVR.html. [Accessed: March. 7, 2024]. [86] T. Hodson, T. Over, and S. Foks, "Mean Squared Error, Deconstructed," Journal of Advances in Modeling Earth Systems, vol. 13, 2021. doi: 10.1029/2021MS002681. [Accessed: Jun. 24, 2024]. [87] Gabaldon, Antonio & Garcia Garre, Ana & Ruiz-Abellón, M.C. & Guillamón, Antonio & Alvarez, C. & Fernandez-Jimenez, Luis. (2021). Improvement of customer baselines for the evaluation of demand response through the use of physically-based load models. Utilities Policy. 70. 101213. 10.1016/j.jup.2021.101213. [Accessed: Jun. 24, 2024]. [88] "Dürr Group Sustainability Approach: Sustainable Development Goals," Dürr Group, 2024. [Online]. Available: https://www.durr-group.com/en/sustainability/sustainabilityapproach/sustainable-development-goals. [Accessed: Jun. 26, 2024].
Page 98 July 17th 2024