scieee AI-readable full text Open interactive document viewer

Hierarchical Reinforcement Learning for Optimal EV Charging: A Multi-Level Framework for Dynamic Pricing and Load Scheduling

Vamvakas, Dimitrios; Korkas, Christos; Tsaknakis, Christos; Kosmatopoulos, Elias

Abstract

This paper explores the application of Hierarchical Reinforcement Learning (HRL) in optimizing electric vehicle (EV) charging, addressing challenges in load scheduling, energy cost management, and real-time dynamic pricing. We propose a hierarchical environment framework with two interconnected levels designed to efficiently manage charging demands across multiple Charging Stations (Chargym). The upper level focuses on dynamic pricing optimization, while the lower level handles load distribution among EVs. The Deep Deterministic Policy Gradient (DDPG) algorithm is implemented within this framework and evaluated against a baseline to assess its performance. Experimental results demonstrate that HRL effectively decomposes the complex EV charging problem into manageable subtasks, achieving improved efficiency in scheduling and pricing while ensuring cost-effective energy distribution. The upper-level DDPG agent in our formulation has shown a 35.58% improvement, and the overall DSO profit has shown a 101.04% improvement, when compared to the RBC baseline.

Full text

Hierarchical Reinforcement Learning for Optimal EV Charging: A Multi-Level Framework for Dynamic Pricing and Load Scheduling Dimitrios G. Vamvakas1,3∗, Christos D. Korkas1,2, Christos D. Tsaknakis1,3, and Elias B. Kosmatopoulos1,3 Abstract— This paper explores the application of Hierarchical Reinforcement Learning (HRL) in optimizing electric vehicle (EV) charging, addressing challenges in load scheduling, energy cost management, and real-time dynamic pricing. We propose a hierarchical environment framework with two interconnected levels designed to efficiently manage charging demands across multiple Charging Stations (Chargym). The upper level focuses on dynamic pricing optimization, while the lower level handles load distribution among EVs. The Deep Deterministic Policy Gradient (DDPG) algorithm is implemented within this framework and evaluated against a baseline to assess its performance. Experimental results demonstrate that HRL effectively decomposes the complex EV charging problem into manageable subtasks, achieving improved efficiency in scheduling and pricing while ensuring cost-effective energy distribution. The upper-level DDPG agent in our formulation has shown a 35.58% improvement, and the overall DSO profit has shown a 101.04% improvement, when compared to the RBC baseline. Hierarchical Reinforcement Learning, Electric Vehicle charging, dynamic pricing, load scheduling I. INTRODUCTION The problem of optimizing energy efficiency and maintaining grid stability has been widely studied in recent years. Key concerns in the energy sector include the depletion of traditional energy sources, the rapid growth of global energy demand—projected to reach 17,487 Mtoe (tonnes of oil equivalent) by 2040 [1]—and the rising carbon emissions. Additional challenges such as sharp energy peaks, fluctuations in consumption, and resource waste pose significant risks to grid stability, demand-supply reliability, and cost efficiency. These factors not only increase consumer costs but also reduce utility company profits and complicate overall grid management. Modern technologies, such as Electric Vehicles (EVs), Energy Storage Systems (ESSs), and Renewable Energy Sources (RESs), further increase the complexity of grid control and energy management. While these technologies can contribute to grid strain due to changing energy dynamics and higher consumption, they also have the potential to support grid stability through Vehicle-to-Grid (V2G) mechanisms in EVs and the efficient utilization of *Corresponding author 1Dimitrios G. Vamvakas, Christos D. Korkas, Christos D. Tsaknakis, and Elias B. Kosmatopoulos are with the Informatics & Telematics Institute (I.T.I.), Center for Research and Technology Hellas (CERTH), Greece. [email protected], [email protected], [email protected], [email protected] 2Christos D. Korkas is also with the Electrical and Computer Engineering Dept. (ECE), University of Western Macedonia (UOWM), Greece. 3Dimitrios G. Vamvakas, Christos D. Tsaknakis, and Elias B. Kosmatopoulos are also with the Electrical and Computer Engineering Dept. (ECE), Democritus University of Thrace (DUTH), Greece. stored or generated energy in ESSs and RESs. The challenge, therefore, lies in integrating these technologies optimally within the grid to maximize efficiency and stability [2]. Various approaches have been explored to address this multi-dimensional problem. Model Predictive Control (MPC) has been widely used [3] due to its ability to optimize energy scheduling and pricing decisions. However, MPC relies on accurate system models, making it less effective in highdimensional, uncertain, and dynamic environments. Gametheoretic methods, such as Stackelberg games, have been employed to model interactions between utility providers and consumers, enabling leader-follower optimization in dynamic pricing and load scheduling [4], [5]. Other hierarchical control strategies [6], [7] have also demonstrated promising results but may become computationally expensive in largescale implementations. More recently, reinforcement learning (RL) and deep learning-based approaches have gained attention for their ability to learn optimal policies in dynamic environments without requiring detailed system models [8]– [10]. One of the most commonly applied strategies in energy optimization is Demand Side Management (DSM), which aims to balance power demand and influence consumption patterns through direct or indirect control. Direct control DSM employs explicit control over Distributed Energy Resources (DERs) and loads [11], [12], modifying user consumption based on predefined policies. Indirect control DSM, on the other hand, relies on control signals that influence, but do not directly dictate, DER operations. A. Main Contributions In this work, we focus on indirect control DSM by implementing a Hierarchical Reinforcement Learning (HRL) framework for EV charging optimization. The proposed framework consists of two hierarchical levels: •Upper Level (Distribution System Operator - DSO): The DSO issues 24-hour dynamic price signals at the end of each day, adapting pricing based on the hourly consumption patterns and satisfaction levels of lowerlevel entities. It aims to maximize profits while ensuring grid stability by reducing demand peaks. Given the electricity requested must be available at all times during the day, we consider the amount of electricity bought by the DSO as a quantity directly dependent on the electricity requested from the low levels. If there is no request, the DSO will not have any cost, as it will not buy any amount of energy. •Lower Level (Load Aggregators - LAs / EV Charging Stations): Each LA represents an EV charging station 2025 33rd Mediterranean Conference on Control and Automation (MED) June 10 - 13, 2025. Tangier,Morocco 979-8-3315-7719-3/25/$31.00 ©2025 IEEE 405 2025 33rd Mediterranean Conference on Control and Automation (MED) | 979-8-3315-7719-3/25/$31.00 ©2025 IEEE | DOI: 10.1109/MED64031.2025.11073233 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:39 UTC from IEEE Xplore. Restrictions apply. that makes charging and discharging decisions based on real-time pricing and customer energy demands. The goal of the LAs is to minimize their costs while maximizing user satisfaction. For the lower-level LAs, we utilize the Chargym environment [8], with each LA controlled by a Deep Deterministic Policy Gradient (DDPG) agent [13]. The upper-level DSO is modeled as an environment designed to represent grid-level decision-making dynamics. As a baseline, we employ RuleBased Control (RBC) approaches for the upper-level and evaluate performance by comparing RBC-DDPG (baseline) against DDPG-DDPG (HRL) formulation to evaluate the system’s performance. The rest of this paper is as follows: Section II presents the Theoretical Approach on HRL. Section III presents the description of the System Architecture, its characteristics and the explanation of the hierarchical structure of our approach. Section IV presents the main results of utilizing the RBCDDPG and DDPG-DDPG agent formulations. Section V summarizes the main ideas and results on our approach and addresses the possibilities for future work. II. THEORETICAL APPROACH Reinforcement Learning (RL) is a subfield of machine learning (ML) that focuses on solving complex problems through trial and error. In RL, autonomous agents, comprised of state-of-the-art computational algorithms, interact with an environment to learn optimal decision-making strategies. This interaction occurs at discrete time-steps, where the agent selects an action at time step t, which in turn transitions the environment to a new state. The agent then receives feedback in the form of reward signals, which evaluate the effectiveness of the action during their interaction with the environment [14]–[16]. This is an iterative process that allows the agent to learn how to refine its policy and converge toward taking optimal actions according to the environment dynamics. However, in the case of multidimensional tasks, where the environment transitions are defined by a larger number of possible actions, the complexity increases and standard RL algorithms have been proven to perform poorly [17]. Hierarchical Reinforcement Learning (HRL) aims to solve such tasks with advanced complexity, by decomposing the learning process into structured levels in a hierarchical manner [18], enabling more efficient decision-making. Each of these tasks is governed by its own policy. This allows the agents to operate at different levels of temporal abstraction [19]–[21], dividing the initial problem into simpler subproblems. This technique has been proven to both reduce training speed, and improve the convergence rate towards the optimal solution [17], [22]. For that reason, HRL stands out as a promising approach, primarily due to its communication and coordination between different levels of the hierarchy. In HRL, higher-level policies govern decision-making by defining subproblems as high-level actions, while maximizing the overall task’s reward. These subproblems serve as subtasks or subgoals, guiding lower-level policies, which treat them as independent RL challenges. In turn, the lowerlevel agents execute primitive actions, driving state transitions within their respective environments, while using their own rewards and optionally incorporating the main task reward. This hierarchical structure enables more efficient learning and decision-making in complex multi-step tasks. In our formulation, we implemented two levels in the hierarchy, one upper level, which acts as a DSO agent and issues electricity pricing signals, and one lower level, which is made up of multiple agents in the form of LAs, each of which aims to achieve its personal goal and reach its maximum reward. The LAs manage their energy consumption by defining the charging/discharging rate of EVs and simultaneously distributing the requested energy in different chargers, depending on energy demand. The high-level pricing actions, which define the cost of the LAs, affect the individual rewards and function as low level subgoals. In this case, agents do not share information with each other, as they only communicate with the upper level in the hierarchy. The upper level therefore takes the role of a centralized entity that has access to the observations and the general behavior of each individual agent. Each agent acts based on its own observations of the environmental conditions and its policies, and is indirectly guided from the upper-level agent through pricing signals to adjust their actions. Our formulation aims to strike an optimal balance between the profitability and satisfaction of the DSO, which is influenced by grid stability, and the costs and satisfaction of the LAs. This balance ensures efficient energy distribution while maintaining system reliability and economic viability. III. FRAMEWORK AND MATHEMATICAL DESCRIPTION A. Markov Decision Processes The real-time EVCS scheduling problem in our formulation, is modeled as a Markov Decision Process (MDP), which is characterized by a transition function governing state evolution. An MDP is defined by the tuple (S, A, P, R), where: •Sdenotes the state space, defining the system’s possible configurations and environmental conditions. •Arepresents the action space, defining the agent’s available decisions. •Pis the state transition probability function, given by P:S×A×S→[0,1], which describes the likelihood of transitioning from state sto s′at time t+ 1 after taking action aat time t. •Ris the reward function, defined as R:S×A×S→R, which describes the agent’s performance based on the transition. The transition function for the MDP is then expressed as: P(s′, r |s, a) where s′represents the next state and ris the corresponding reward, given the current state sand action a. 406 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:39 UTC from IEEE Xplore. Restrictions apply. B. Semi-Markov Decision Processes Most HRL problems are formulated as Semi-Markov Decision Processes (SMDPs) due to the hierarchical structure and variable-duration subtasks. High-level decisions take multiple time steps to execute, making the environment naturally semi-Markovian. SMDPs allow modeling state transitions for extended time periods, which makes them suitable for hierarchical learning. An SMDP is represented by a tuple: M= (S, A, P, R, F ) where: •Sis the state space, •Ais the action space, •P(s′|s, a, τ)is the state transition probability, which depends not only on the current state sand action a, but also on the time duration τfor which the action persists, •R(s, a, s′, τ)is the reward function, dependent on both the transition and duration, •F(τ|s, a, s′)is the probability distribution over transition durations. Although SMDPs typically allow stochastic transition durations, we consider a simplified setting with fixed τper transition for the sake of tractability and alignment with real-world control intervals in building and charging station operations. Our framework aligns with an SMDP due to the following characteristics: 1. The high-level DDPG agent takes dynamic price decisions, that persist for a period of one hour, influencing the lower-level optimization, as represented by: aupper t=πupper(supper t) where aupper tis the price set for the next hour, based on the upper-level state supper t. 2. The price decision, which includes a pricing list of 24 values (one value per hour), remains fixed for a duration τ= 24 hours, as the low-levels are optimizing their energy scheduling actions hourly, based on these pricing signals. The environment follows semi-Markovian transitions: P(supper t+τ|supper t, aupper t, τ) indicating that the next state depends on the high-level action and the duration it remains active. 3. Meanwhile, the low-level agents operate at a finer time scale, optimizing their actions alow tbased on the pricing signal: alow t=πlow(slow t, aupper t) where slow tis the local state of a low-level agent. 4. The upper-level agent receives its reward based on aggregated metrics (profit, satisfaction and fluctuation values) after the lower-level agents act: Rupper =f(profit,satisfaction,fluctuation) while the lower-level agents receive their individual rewards at each time step: Rlow t=g(cost,local satisfaction) 5. The hierarchical nature introduces temporal abstraction, allowing the upper-level decisions to influence multiple lower-level actions and, in turn, receive feedback. C. System Architecture The current architecture is as follows: The DSO as the upper level sends the pricing list, p, to the LA agents (DDPG 1, DDPG 2, DDPG 3). The agents begin optimizing their strategies based on these prices and output primitive actions, an, which alter the Charging Stations’ states, sn. After applying their actions, the agents receive back reward signals rn. At the end of a 24-hour period, when the low-level agents finish their tasks, the DSO receives their satisfaction, satn, as defined by the second term of (6), and the energy consumption, ln, and aggregates them accordingly. In Fig. 1 the proposed system architecture is presented. Fig. 1. Hierarchical Reinforcement Learning framework. D. State and Action spaces The two levels operate in different time domains, with different state and action spaces for each low-level agent. Since there is no direct communication or intervention between agents, there are also no dependencies. Therefore, for the Chargym environments [8], the action space is alow t= [−1,1] and the state space is presented in (1). The charging/discharging power for each vehicle iat timestep t is found in (2) and its constraints are found in (3) and (4). slow t= (Gt, prt, Gt+1, Gt+2, Gt+3, prt+1, prt+2, prt+3, ..., SoC1 t, SoC2 t, ..., SoC10 t, Tleave1 t, Tleave2 t, ..., T leave10 t)(1) MaxEnergyi t=(1 −SoCi t)∗Bmax actioni t≥0 (SoCi t)∗Bmax actioni t<0(2) 407 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:39 UTC from IEEE Xplore. Restrictions apply. MaxEnergyi t≤Pch,max ·nch (3) Pdem,i t=actioni t·MaxEnergyi t(4) Values G, Gt+1, ... refer to the solar radiation produced from the photovoltaic panels (PV) of the charging station, at time step t, t + 1, and so on, prt, prt+1, ... refer to the prices, SoCN trefer to the State of Charge of the battery of EV N, at time step t, and the TleaveN tvalues indicate the time EV Nremains at the station. Bmax is the maximum EV battery capacity. Pch,max is the maximum charging output for a charger, nch is the charging/discharging efficiency in %, and Pdem,i tis the demanded power from EV Nat time step t. The action space for the upper-level DSO is the 24-hour pricing signals, with aupper t= [0.05,0.3], and the state space, found in (5), holds 96 values: supper t= (prt−24, prt−23, ..., prt−1, lt−24, lt−23, ..., lt−1, satt−24, satt−23..., satt−1, ft−24, ft−23, ..., ft−1)(5) These state values refer to 24 historical pricing values, prt, 24 total load consumption values, lt, 24 total low-level satisfaction values, satt, and 24 total fluctuation values, ft. E. Rewards The reward for the low-level agents [8] and the reward for the upper level are found and presented in (6) and (7): rlow t=X i∈Ωt (pt·lt) + X i∈Ψt [2 ·(1 −SoCi t)]2(6) rupper t=profit −satisfaction −fluctuation (7) The profit term (8) of the upper-level reward function is the difference of the pricing signal and the marginal cost of electricity generation, multiplied by the total amount of energy that was requested at time step t. The satisfaction term (9) is the aggregated satisfaction of the lower-level agents and the fluctuation term (10) is the fluctuation function for sharp peak reduction: profit =X i∈Ωt [(pt−ct)·lt](8) satisfaction = N X n=1 X i∈Ωt [2 ·(1 −SoCi t)]2(9) fluctuation =γ·X i∈Ωt (lt−lavg)2(10) In the aforementioned, ctis the marginal cost of electricity generation for the DSO, ltis the requested energy at time step t,lavg is the daily average energy requested, and γ is the weight which controls the sensitivity to sharp peaks. A smaller γreduces the impact of the fluctuation, that we consider as a penalty alongside aggregated satisfaction, while a larger γincreases the impact of the penalty, making its contribution larger to the final reward, depending on the DSO’s objectives. In our results, γis set to 0.01, scaling the fluctuation penalty to a comparable range with the other reward terms. To ensure the upper-level rewards are meaningful for training, each term is scaled between 0 and 1 using the running maximum approach, seen in (11). This normalization divides each value by its maximum within a window. Since units cancel out, the equation’s terms becomes unitless. The method maintains relative magnitudes across datasets with varying scales. ˆxt=xt maxt−N+1,...,t(x)(11) Upper-Level Objective: The objective of the upper-level is to maximize the profit and minimize the satisfaction and fluctuation penalties (hence the minus signs on these terms). By selling more electricity, the DSO makes more profit, however it needs to achieve a balance between profits, dissatisfaction and grid instability. IV. RESULTS AND PERFORMANCE EVALUATION In this section, we present the experimental evaluation of the HRL environment. The evaluation is conducted across two distinct experimental setups, each designed to assess different configurations of agents at both the upper and lower levels: •Baseline Setup: One RBC agent at the upper level, and three DDPG agents at the lower level, used for baseline performance comparison. •Full HRL Setup: Both the upper and lower levels are controlled by DDPG agents. The experiments aim to solve the load scheduling problem in the Chargym environments through dynamic pricing by the upper-level agent. The baseline uses RBCs to simulate human-level decision-making for comparison against HRL methods. While current tests use a fixed RBC, future work will assess a family of RBC variants for robustness. Training the three DDPG low-level agents took 6 hours and the highlevel agent 2 hours on an Intel i5-11400 CPU with 16GB RAM. On an Apple Silicon M3 Max with 36GB RAM, runtimes dropped to 2.5 hours and 40 minutes, respectively. A. Implementation details Using the Chargym environment from [8] and the DDPG algorithm from the Stable-Baselines3 framework [23], our approach builds on insights from recent work [17], [22], [24]. This framework was selected as it provides a robust and efficient DDPG implementation. Our work aims to validate HRL’s potential and explore its hierarchical aspects and lay the foundation for future research with buildings, power plants, and DERs to study complex, real-world dynamics. The upper-level agent outputs a daily pricing list, with one price per hour, and learns to optimize its policies through a 30-day period. The lower-level agents operate in 24-hour 408 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:39 UTC from IEEE Xplore. Restrictions apply. time steps. At each hour, they output a charging/discharging signal based on the EV dynamics. The upper level actions are based on low level satisfaction, marginal cost of electricity production and energy fluctuations present in the grid. B. Rule Based Control agents for the two levels Rule Based Control (RBC) agents have been created to serve as a baseline for the hierarchy. The RBC in our paper is comprised of a set of rules that simplify the decisionmaking process. The upper-level agent provides high-level actions based on observations received from the low-level environments. These observations refer to the low-level agents’ satisfaction values. The RBC serves as a useful baseline, closely resembling human-like actions without complex computations or exploration. Though suboptimal and inflexible compared to learning agents, it helps compare human-level performance to machine learning policies. Here, actions are rule-based, relying on low-level observations, with no learning involved. The RBC’s rule are easy to interpret and are as seen in (12): pt=ct+θ·0.1(12) The pricing strategy of the RBC allows the DSO to increase the prices slightly to surpass the marginal cost and ensure it has profit. θis equal to 1 when the satisfaction of the low levels is maximum, and 0.5 otherwise. C. Performance 1) Case 1 — RBC - DDPG agents: In this setup, the low-level agents are DDPG-based and capable of learning, while the upper-level agent follows a simple RBC strategy. The DDPG low-level agents adjust their actions based on learned policies, influencing both their own costs and the DSO’s profit, as shown in Fig. 2, since the DSO seeks to maximize profit, which inherently increases costs for the lowlevel agents, while they aim to minimize their own expenses. The fluctuation penalty and aggregated satisfaction penalty metrics are presented in Figs. 3 and 4, respectively. Fig. 2. DSO profit Fig. 3. DSO fluctuation penalty Fig. 4. Aggregated satisfaction penalty 2) Case 2 — DDPG - DDPG agents: In the full HRL setup, both upperand lower-level agents employ learningbased strategies. The upper-level rewards along with profit values for both cases are summarized in Table I. Results indicate a substantial increase in DSO profit of 101.04%, compared to the RBC case (Fig. 5). The upper-level agent also achieves higher rewards, demonstrating improved performance. Specifically, in Case 1, the RBC agent stabilizes at a mean reward of -267, shown in Fig. 5, whereas in Case 2, the DDPG reaches an average of -172, indicating a 35.58% improvement. This improvement in profit, however, comes with tradeoffs. In Case 1, although the reduction in DSO profit occurs, it is accompanied by smaller fluctuation values compared to Case 2, indicating a more stable system (Fig. 3). The Fig. 5. DSO Rewards 409 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:39 UTC from IEEE Xplore. Restrictions apply. DDPG-powered DSO reaches higher fluctuation penalties, indicating increased variability in system behavior. Despite this, the HRL approach not only leads to improved overall profitability, but additionally still maintains a slightly improved level of the aggregated satisfaction penalty, with the RBC and DDPG reaching an average of 3702.63 and 3675.04 respectively. This suggests that the learned policies optimize profit at the expense of slightly higher fluctuations, while ensuring user satisfaction remains nearly unchanged. TABLE I MEAN EVALUATION REWARD Approach Mean Reward Profit RBC DSO -267 1129.95 DDPG DSO -172 2251.25 V. CONCLUSIONS AND FUTURE WORK This paper presented a hierarchical reinforcement learning (HRL) framework for optimizing electric vehicle (EV) charging, focusing on dynamic pricing and load scheduling. By integrating a multi-level structure with Load Aggregators (LAs) and a Distribution System Operator (DSO), the approach effectively balances cost minimization, grid stability, and user satisfaction. Through the Deep Deterministic Policy Gradient (DDPG) algorithm, it demonstrated improved efficiency in decision-making, with agents successfully converging to optimal actions. This method extends standard HRL by using dynamic pricing incentives instead of rigid subgoals, allowing decentralized, user-centric responses from low-level agents. This soft influence approach better reflects realworld DSO behavior. With decoupled temporal scales—daily for the DSO, hourly for nodes—and reward normalization via running profit maxima, stability, scalability and realism are enhanced, offering a more adaptable and efficient HRL solution. The results highlight HRL’s potential in addressing complexities of EV charging management, paving the way for future enhancements, such as customer-specific pricing and more realistic modeling of market dynamics. ACKNOWLEDGMENT We acknowledge partial support of this work by the European Commission Horizon Europe - REHOUSE : Renovation packagEs for HOlistic improvement of EU’s bUildingS Efficiency, maximizing RES generation and cost-effectiveness (Grant agreement ID: 101079951) and Horizon Europe - Harmonise : Hierarchical and Agile Resource Management Optimization for Networks in Smart Energy Communities (Grant agreement ID: 101138595) REFERENCES [1] T. Ahmad and D. Zhang, “A critical review of comparative global historical energy consumption and future demand: The story told so far,” Energy Reports, vol. 6, pp. 1973–1991, 2020. [2] D. Vamvakas, P. Michailidis, C. Korkas, and E. Kosmatopoulos, “Review and evaluation of reinforcement learning frameworks on smart grid applications,” Energies, vol. 16, no. 14, p. 5326, 2023. [3] Y. Yao and D. K. Shekhar, “State of the art review on model predictive control (mpc) in heating ventilation and air-conditioning (hvac) field,” Building and Environment, vol. 200, p. 107952, 2021. [4] A. Di Giorgio, F. Liberati, and S. Canale, “Optimal electric vehicles to grid power control for active demand services in distribution grids,” in 2012 20th Mediterranean Conference on Control & Automation (MED), pp. 1309–1315, IEEE, 2012. [5] K. Amasyali, Y. Chen, B. Telsang, M. Olama, and S. M. Djouadi, “Hierarchical model-free transactional control of building loads to support grid services,” IEEE Access, vol. 8, pp. 219367–219377, 2020. [6] H. N. Nguyen, C. Zhang, J. Zhang, and L. B. Le, “Hierarchical control for electric vehicles in smart grid with renewables,” in 2017 13th IEEE International Conference on Control & Automation (ICCA), pp. 898– 903, IEEE, 2017. [7] Y. Wu, Z. Wang, Y. Huangfu, A. Ravey, D. Chrenko, and F. Gao, “Hierarchical operation of electric vehicle charging station in smart grid integration applications—an overview,” International Journal of Electrical Power & Energy Systems, vol. 139, p. 108005, 2022. [8] G. Karatzinis, C. Korkas, M. Terzopoulos, C. Tsaknakis, A. Stefanopoulou, I. Michailidis, and E. Kosmatopoulos, “Chargym: An ev charging station model for controller benchmarking,” in IFIP International Conference on Artificial Intelligence Applications and Innovations, pp. 241–252, Springer, 2022. [9] L. Zhang, Y. Gao, H. Zhu, and L. Tao, “A distributed real-time pricing strategy based on reinforcement learning approach for smart grid,” Expert systems with applications, vol. 191, p. 116285, 2022. [10] C. D. Korkas, S. Baldi, P. Michailidis, and E. B. Kosmatopoulos, “A cognitive stochastic approximation approach to optimal charging schedule in electric vehicle stations,” in 2017 25th Mediterranean Conference on Control and Automation (MED), pp. 484–489, IEEE, 2017. [11] A. M. Kosek, G. T. Costanzo, H. W. Bindner, and O. Gehrke, “An overview of demand side management control schemes for buildings in smart grids,” in 2013 IEEE international conference on smart energy grid engineering (SEGE), pp. 1–9, IEEE, 2013. [12] C. D. Korkas, C. D. Tsaknakis, A. C. Kapoutsis, and E. Kosmatopoulos, “Distributed and multi-agent reinforcement learning framework for optimal electric vehicle charging scheduling.,” Energies (19961073), no. 15, 2024. [13] T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” 2019. [14] R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction. MIT Press, 2nd ed., 2018. [15] M. A. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization, vol. 12, no. 3, p. 729, 2012. [16] V. Casagrande, M. Ferianc, M. Rodrigues, and F. Boem, “An online learning method for microgrid energy management control,” in 2023 31st Mediterranean Conference on Control and Automation (MED), pp. 263–268, IEEE, 2023. [17] N. G¨ urtler, D. B¨ uchler, and G. Martius, “Hierarchical reinforcement learning with timed subgoals,” Advances in Neural Information Processing Systems, vol. 34, pp. 21732–21743, 2021. [18] S. Pateria, B. Subagdja, A.-h. Tan, and C. Quek, “Hierarchical reinforcement learning: A comprehensive survey,” ACM Computing Surveys (CSUR), vol. 54, no. 5, pp. 1–35, 2021. [19] P.-L. Bacon, J. Harb, and D. Precup, “The option-critic architecture,” in Proceedings of the AAAI conference on artificial intelligence, vol. 31, 2017. [20] T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum, “Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation,” Advances in neural information processing systems, vol. 29, 2016. [21] O. Nachum, S. S. Gu, H. Lee, and S. Levine, “Data-efficient hierarchical reinforcement learning,” Advances in neural information processing systems, vol. 31, 2018. [22] K. Frans, J. Ho, X. Chen, P. Abbeel, and J. Schulman, “Meta learning shared hierarchies,” arXiv preprint arXiv:1710.09767, 2017. [23] A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-baselines3: Reliable reinforcement learning implementations,” Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021. [24] A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu, “FeUdal networks for hierarchical reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning (D. Precup and Y. W. Teh, eds.), vol. 70 of Proceedings of Machine Learning Research, pp. 3540–3549, PMLR, 06–11 Aug 2017. 410 Authorized licensed use limited to: Tallinn University of Technology. Downloaded on December 22,2025 at 10:09:39 UTC from IEEE Xplore. Restrictions apply.