scieee AI-readable full text Open interactive document viewer

GraphOpticon: A Global proactive horizontal autoscaler for improved service performance & resource consumption

Theodoropoulos, Theodoros; Patel, Yashwant Singh; Townend, Paul; Korontanis, Ioannis; Makris, Antonios; Tserpes, Konstantinos

Full text

Contents lists available at ScienceDirect Future Generation Computer Systems journal homepage: www.elsevier.com/locate/fgcs GraphOpticon: A Global proactive horizontal autoscaler for improved service performance & resource consumptionI Theodoros Theodoropoulos a,c,∗, Yashwant Singh Patel b, Uwe Zdun a, Paul Townend b, Ioannis Korontanis c,d, Antonios Makris d, Konstantinos Tserpes d aUniversity of Vienna, Software Architecture Research Group, Vienna, 1090, Austria bUmeå University, Dept. Computing Science, Umeå, 90187, Sweden cHarokopio University of Athens, Omirou 9, Athens, 17778, Greece dNational Technical University of Athens, Heroon Polytechniou 9, Athens, 15780, Greece A R T I C L E I N F O Keywords: Cloud computing Green computing Graph neural networks Deep learning Resource usage forecasting Resource consumption Service performance A B S T R A C T The increasing complexity of distributed computing environments necessitates efficient resource management strategies to optimize performance and minimize resource consumption. Although proactive horizontal autoscaling dynamically adjusts computational resources based on workload predictions, existing approaches primarily focus on improving workload resource consumption, often neglecting the overhead introduced by the autoscaling system itself. This could have dire ramifications on resource efficiency, since many prior solutions rely on multiple forecasting models per compute node or group of pods, leading to significant resource consumption associated with the autoscaling system. To address this, we propose GraphOpticon, a novel proactive horizontal autoscaling framework that leverages a singular global forecasting model based on Spatio-temporal Graph Neural Networks. The experimental results demonstrate that GraphOpticon is capable of providing improved service performance, and resource consumption (caused by the workloads involved and the autoscaling system itself). As a matter of fact, GraphOpticon manages to consistently outperform other contemporary horizontal autoscaling solutions, such as Kubernetes’ Horizontal Pod Autoscaler, with improvements of 6.62% in median execution time, 7.62% in tail latency, and 6.77% in resource consumption, among others. 1. Introduction The growing complexity of distributed computing environments introduces variability in workload demands, latency issues, and data management challenges, necessitating automated resource management strategies. Resource scaling [1] is the process of dynamically adjusting computational resource allocation, such as CPU and memory, based on workload fluctuations to optimize performance and cost efficiency [2]. Horizontal scaling refers to the process of increasing or decreasing the number of instances in response to demand. Kubernetes,1 a widely adopted container orchestration framework, plays a crucial role in managing these scaling processes through mechanisms IThis research was funded in whole or in part by the Austrian Science Fund (FWF) project CQ4CD, Grant -DOI: 10.55776/I6510. For open access purposes, the author has applied a CC BY public copyright license to any author accepted manuscript version arising from this submission. Furthermore, this work was funded by the Horizon Europe Framework Programme under grant agreements No 101135775 (PANDORA), No 101120990 (SOPRANO), and No 101092711 (SovereignEdge.Cognit). Finally, this work is partially funded by the Kempe Foundations (Kempestiftelserna) grant ‘‘De facto Center of Excellence in Autonomous Distributed Systems’’. ∗Corresponding author at: University of Vienna, Software Architecture Research Group, Vienna, 1090, Austria. E-mail addresses: [email protected] (T. Theodoropoulos), [email protected] (Y.S. Patel), [email protected] (U. Zdun), [email protected] (P. Townend), [email protected] (I. Korontanis), [email protected] (A. Makris), [email protected] (K. Tserpes). 1https://kubernetes.io/. such as the Horizontal Pod Autoscaler (HPA) [3], ensuring efficient allocation of computational resources via pod replication. In large-scale systems, resource utilization metrics typically exhibit time-series patterns with non-linear behaviors. Recurrent Neural Networks (RNNs) [4] have demonstrated effectiveness in modeling such data distributions. Advanced RNN-based architectures, such as Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU) [5], can predict resource utilization trends, enabling more efficient system orchestration. While conventional time-series forecasting models focus on single-step-ahead predictions, multi-step-ahead forecasting strategies provide a sequence of future values, allowing for more granular https://doi.org/10.1016/j.future.2025.107926 Received 23 November 2024; Received in revised form 26 March 2025; Accepted 13 May 2025 Future Generation Computer Systems 174 (2026) 107926 Available online 31 May 2025 0167-739X/© 2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ ). T. Theodoropoulos et al. resource management. This enables the implementation of proactive resource allocation and scaling techniques that mitigate bottlenecks and improve overall efficiency. Encoder-Decoder (ED) Deep Learning (DL) architectures have shown superior performance in multi-step forecasting compared to traditional prediction models [6]. The encoder processes a variable-length sequence, transforming it into a structured representation that the decoder then utilizes to generate predictions. Recent advancements in handling non-Euclidean data have also led to the emergence of Graph Neural Networks (GNNs) [7], which excel in solving problems with spatial components. This capability stems from their inherent ability to leverage and utilize the spatial characteristics of data related to a given problem. Spatio-temporal GNNs [8] have been particularly successful in forecasting problems, as their architecture allows them to simultaneously capture both spatial and temporal dependencies. This is achieved by using graph convolutions to model spatial dependencies and RNNs to capture temporal dependencies in alignment with the Encoder-Decoder paradigm. Proactive horizontal autoscaling strategies involve leveraging such models to enhance resource allocation efficiency by dynamically adjusting the number of processing nodes in anticipation of workload fluctuations. A review of the scientific literature reveals that prior works on proactive horizontal autoscaling focus on improving resource consumption solely from the perspective of the workloads being executed. Hence, they do not consider the resource consumption of the autoscaling system itself. In proactive horizontal pod autoscaling, workload resource consumption refers to the resources utilized by the running application workloads within a Kubernetes cluster. This metric is essential for assessing system load and making informed scaling decisions. Conversely, autoscaling system resource consumption accounts for the resources used by the autoscaler itself. An efficient proactive autoscaling system should minimize its own resource footprint while effectively scaling workloads to maintain performance and optimize resource utilization. Many of the prior works are designed to employ multiple forecasting models (one per pod or group of pods), which would introduce significant overhead if integrated into the scaling decision-making process. Employing multiple DL models in decision-making can improve resource management by enabling task specialization. However, this approach also increases resource usage, as each model requires computational resources for both training and inference. Training multiple models is particularly resource-intensive, and frequent updates can further amplify resource demands. Moreover, managing the outputs of multiple models introduces computational overhead and potential latency, affecting performance in real-time applications. Therefore, it is essential to minimize the number of forecasting mechanisms deployed in a system, ensuring efficient resource scaling and management while mitigating the subsequent resource consumption. This overhead could potentially offset the benefits of proactive resource allocation and deallocation. These observations motivated us to propose GraphOpticon. GraphOpticon is a proactive horizontal autoscaling solution designed to enhance service performance while minimizing autoscaling system resource consumption by leveraging only a singular global forecasting model instead of numerous local specialized ones. Towards achieving this goal, GraphOpticon is based on the implementation of information fusion and distillation processes, driven by the characteristics of the deployed services, as well as by the ontological relations that these services form with each other in order to further refine the generalization capabilities of the employed forecasting model that is based on Spatio-temporal GNNs. The key contributions of our research are: •We advocate for the use of a singular global forecasting model (instead of numerous local ones) in order to establish resourceefficient proactive horizontal autoscaling. •We propose GraphOpticon, a novel global horizontal autoscaling solution that leverages the information fusion & distillation processes to improve both service performance and workload resource consumption while minimizing the underlying autoscaling system resource consumption. •We extensively analyze the architectural decisions, results, and potential ramifications that are intertwined with the use of a global proactive horizontal autoscaling solution. The remaining sections of this work are organized as follows. Section 2 provides the current research status of relevant proactive horizontal autoscaling solutions. Section 3 establishes the motivation behind this study. Section 4 presents the problem formulation. Section 5 describes the proposed ’GraphOpticon’ solution. Section 6 discusses several implementation aspects that are intertwined with GraphOpticon. Section 7 focuses on the experimental results to evaluate the efficiency of the proposed solution. Finally, Section 8 concludes this study and discusses potential directions for future work. 2. Related work This section explores recent advancements in proactive horizontal autoscaling techniques. To achieve this, we meticulously analyze numerous scientific works in the frame of various aspects, such as forecasting models, workflows, datasets, evaluation tools, and evaluation metrics (as explained in Table 2). An overview of this analysis is presented in Table 1. Aside from showcasing the fact that GraphOpticon constitutes an advancement towards more efficient resource orchestration solutions, the purpose of this section is to guarantee that the subsequent experimental evaluation will be conducted in a manner that is aligned with the corresponding scientific literature. Forecasting models are widely explored in the literature for resource orchestration. These models play a crucial role in proactive autoscaling, where they predict future resource needs and scale infrastructure accordingly. To improve the forecasting performance of container-based load prediction models, Tang et al. [22] design ’Fisher’, which consists of metric selection and neural network training components. The metric selection component identifies relevant metrics using a novel shapebased time series clustering technique. Subsequently, a Bidirectional LSTM (BiLSTM) is applied to predict the one-step-ahead workload. Patel et al. [23] propose a dynamic consolidation technique for cloud systems. They present a clustering-based stacked bidirectional LSTM model to forecast the future CPU and memory usage of machines. Utilizing the prediction results, they design different consolidation approaches to show improvements in energy, migrations, and SLA violations. Radhika et al. [24] present ARIMA and LSTM algorithms to predict future CPU usage from a 3-tier architecture of web applications hosted on a private cloud platform. The LSTM approach notably shows high accuracy and effectiveness in forecasting future web application demands. Theodoropoulos et al. [6] introduced a novel Encoder-Decoder architecture that utilizes stacked LSTM and BiLSTM layers at both the encoder and the decoder parts of the model. This approach managed to provide superior results in terms of service demand forecasting accuracy against a plethora of contemporary forecasting solutions, and is referred to as the Hybrid LSTM Encoder-Decoder. Following works [25] have also demonstrated the superiority of the Hybrid LSTM Encoder-Decoder in the frame of resource consumption forecasting, surpassing competitors such as LTMS, Bi-LSTMs, GRUs, the CNN-LSTMs, as well as various other ED architectures, providing slightly worse results only when compared against Spatio-temporal GNNs that are based on Discrete-Time Dynamic Graphs (DTDG) [26] paradigm. Proactive horizontal autoscaling strategies leverage machine learning and predictive models to enhance resource allocation efficiency by dynamically adjusting the number of processing nodes in anticipation of workload fluctuations. Various approaches have been proposed, Future Generation Computer Systems 174 (2026) 107926 2 T. Theodoropoulos et al. Table 1 Analysis of horizontal pod autoscaling approaches. Contributors & Year Forecasting model Workflow Datasets Evaluation tools Evaluation metrics MarieMagdelaine et al. [9] (2020) LSTM Applying LSTM model to dynamically adjust the number of replicas and pool of resources Simulated traffic dataset Digital Ocean cloud provider’s VMs Availability, Latency, Number of Pod Replicas Imdoukh et al. [10] (2020) LSTM Utilizing proactive machine learning method for auto-scaling of Docker containers in response to dynamic workload HTTP request (logs of Worldcup98) Python MSE, 𝑅2, RMSE, MAE, Resource Underprovisioning & Overprovisioning, Inference Time Goli et al. [11] (2021) Linear regression, random forest, support vector regression Apply ML models to forecast the required number of replicas for each microservice and considers the impact of scaling one microservice on others within a given workload Generated custom dataset using Teastore application Compute Canada Arbutus Cloud platform Response Time, Number of Pod Replicas Ju et al. [12] (2021) ARMA, LSTM Incorporates multiple metrics for workload forecasting and later utilizes these predictions to dynamically scale the applications Random access, NASA datasets Edge computing cluster MSE, Response Time, Resource Underutilization Dang-Quang et al. [13] (2021) BiLSTM Applying a proactive custom autoscaler using BiLSTM model for handling the dynamic workload NASA and FIFA World Cup 98 log traces Python, Google Colab environment MSE, RMSE, 𝑅2, MAE, Resource Underprovisioning & Overprovisioning, Elastic Speedup Yadav et al. [14] (2021) SVR Applying SVR model to perform horizontal elasticity for Docker containers Web service logs of the Complutense University of Madrid Python, Computing cluster RMSE, MAE, MSE, Resource Consumption, Elastic Speedup Nguyen [15] (2022) LSTM, GNN Applying LSTM for workload prediction and then Graph Convolution Networks for relationship modeling between workload and resource consumption Microsoft’s Azure traces AWS EC2 instances MSE, MAE, Number of Pod Replicas Violos et al. [16] (2022) Feedforward+RNN Applying horizontal proactive autoscaling to provide scale up and scale down decisions Alibaba cluster traces Python, CloudSim Plus RMSE, MSE, MAE, Execution Time, Tail Latency Dogani et al. [17] (2022) Attention+GRU Adapted proactive autoscaling method to predict the multi-step resource usage based on cooldown time FIFA Worldcup Dataset, NASA Log Python, Computing cluster MAE, MAPE, RMSE, Resource Underprovisioning & Overprovisioning, Elastic Speedup Kakade et al. [18] (2023) Bi-LSTM Predict future demands and automatically scale containers accordingly Generated custom dataset Computing cluster MAE, RMSE, Latency Theodoropoulos et al. [19] (2023) GNN-LSTM Uses Spatio-temporal GNNs to forecast CPU consumption, and then leverage said forecasts to conduct proactive horizontal autoscaling Generated custom dataset Python, CloudSim Plus MAE, RMSE, Latency, Execution Time, Number of Pod Replicas Dogani et al. [20] (2024) FedAVGBiGRU Design proactive auto-scaling for a multi-step prediction model FIFA World Cup 98 Web Server Computing cluster MAE, MAPE, RMSE, Resource Underprovisioning & Overprovisioning, Elastic Speedup Ahmad et al. [21] (2025) Prophet introduce ProSmart HPA, a resource-efficient horizontal pod auto-scaler, which utilizes machine learning for proactive scaling to mitigate resource mismanagement. FIFA World Cup 98 Web Server AWS EC2 instances Resource Underutilization, Resource Underprovisioning & Overprovisioning, Elastic Speedup Proposed: GraphOpticon (2025) GNN-LSTM, Information Fusion & Distillation Algorithms Utilizes Spatio-temporal GNNs and information fusion & distillation algorithms to improve service performance and resource consumption Google cluster traces Python, CloudSim Plus RMSE, MAE, Execution Time, Latency, Number of Pod Replicas, Training Time, Inference Time, Number of Parameters each aiming to improve certain aspects of the autoscaling process. For instance, some of them focus on improving performance in the frame of evaluation metrics such as Inference Time, Latency, Availability, Elastic Speedup, Execution Time, and Response Time. Violos et al. [16] introduce an ’Intelligent Horizontal Proactive Autoscaling’ strategy that utilizes resource usage metrics (CPU) of processing edge nodes to make timely and efficient scale-up and scale-down decisions. Their approach is based on a double-tower deep learning-driven topology, which simultaneously analyzes and distinguishes the time-series resource usage metrics for local processing nodes and the edge infrastructure’s aggregated resource metrics. The framework shows improvements in latency and execution time. Kakade et al. [18] designed a Bi-LSTM modelbased proactive autoscaler to predict future demands and automatically scale containers. Their experiments with a three-node Kubernetes setup indicate that Bi-LSTM outperforms stacked LSTM, while the proactive autoscaler achieves better performance than the default Kubernetes autoscaler. Marie-Magdelaine et al. [9] develop an LSTM-based proactive auto-scaling approach that dynamically adjusts the resource pool horizontally and vertically to optimize availability and minimize latency in cloud-native applications. Aside from improving performance, other works aim to also enhance resource efficiency in the frame of Inference Time, Training Time, over-provisioning, under-provisioning, and Resource Underutilization. For instance, Imdoukh et al. [10] design an LSTM-based adaptive forecasting model that predicts future HTTP demands to determine the required number of containers, minimizing delays caused by starting and stopping active containers and eliminating oscillations during scaling operations. Dang-Quang et al. [13] propose a BiLSTM-based proactive autoscaler to predict future HTTP workloads, incorporating a 1 min ’Cool Down Time’ (CDT) interval to mitigate Future Generation Computer Systems 174 (2026) 107926 3 T. Theodoropoulos et al. Table 2 Evaluation metrics used in contemporary works on proactive horizontal autoscaling. Evaluation metrics Descriptions Mean Absolute Error (MAE) Average magnitude of errors between predicted and actual values Root Mean Squared Error (RMSE) Evaluates the standard deviation of the prediction errors Mean Absolute Percentage Error (MAPE) Average percentage error between predicted and actual values R-squared (𝑅2) Square of the multiple correlation coefficient between the observed outcomes and the predicted values Mean Squared Error (MSE) Average squared difference between the forecast and the observed values Availability Percentage of successfully processed requests out of the total user requests issued Training time Time required to train the model on a given dataset Inference time Time required for a trained model to make predictions on unseen data Number of parameters Total count of trainable weights and biases in the model Resource underprovisioning Resources that a microservice needs but is unavailable Resource overutilization Resource utilization exceeding a predefined threshold value Resource overprovisioning The residual resources which are not utilized Elastic speedup The rate at which a system dynamically adjusts resources to meet workload demands Workload Resource Consumption Resource consumption caused by the running application workloads Autoscaling System Resource Consumption Resource consumption caused by the autoscaler oscillations and a resource removal approach to handle underutilized resources. Similarly, Dogani et al. [17] introduce an attention-based GRU encoder–decoder (K-AGRUED) model for proactive autoscaling in Kubernetes, reducing scaling operations and under-provisioning compared to the Kubernetes horizontal pod autoscaler. Dogani et al. [20] devise a proactive auto-scaling approach for edge environments using FedAvg and multi-step workload forecasting with a Bidirectional Gated Recurrent Unit (BiGRU). Their results show improvements in resource overprovisioning, underprovisioning, and elastic speedup while reducing data transmission between edge nodes and cloud servers. Ahmad et al. [21] propose Smart HPA and ProSmart HPA, two resourceefficient horizontal pod autoscalers; Smart HPA applies a reactive scaling policy, while ProSmart HPA leverages Prophet-based machine learning for proactive scaling. Their results demonstrate improvements in resource utilization by mitigating underutilization, overprovisioning, and underprovisioning. Aside from the aforementioned resource efficiency evaluation metrics, the number of pod replicas plays a significant role in the overall resource consumption. The number of pod replicas directly impacts resource consumption by determining how many instances of a specific application or service run simultaneously within a Kubernetes cluster. Each pod replica consumes CPU, memory, and other resources based on the application’s requirements. Increasing the number of replicas spreads the workload, leading to higher resource utilization across the cluster, which can improve performance and availability but also increase overall resource consumption. Conversely, decreasing replicas reduces resource usage but may also impact the application’s ability to handle traffic or provide high availability. Thus, managing pod replicas helps balance resource efficiency with performance needs. Following this line of thought, various works have emerged. Yadav et al. [14] design an autoscaler that enables horizontal scaling for Docker containers using a Support Vector Regression (SVR)-based proactive method, leveraging the IBM MAPE-K computing platform to optimize resource allocation in terms of the number of deployed replicas. Nguyen et al. [15] introduce a graph-based proactive horizontal pod autoscaling strategy for microservices, employing an LSTM-GNN hybrid model that first predicts upcoming workloads and then determines the optimal number of pods required for efficient resource allocation, leading to significant resource savings. Goli et al. [11] propose ’Waterfall,’ a predictive autoscaler using machine learning models such as linear regression, random forest, and SVR to estimate the required replicas for each microservice while considering interdependencies between services. Ju et al. [12] introduce the Proactive Pod Autoscaler for Kubernetes, utilizing multiple user-defined and customizable metrics for workload forecasting, dynamically scaling applications, and outperforming the default Kubernetes pod autoscaler in resource utilization efficiency. Theodoropoulos et al. [19] propose a GNN-LSTM-based approach that uses Spatio-temporal GNNs to forecast CPU consumption and then leverages these forecasts to conduct proactive horizontal autoscaling, achieving performance gains in execution time and latency, while requiring fewer pod replicas. Despite their various scientific contributions, all aforementioned works examine resource consumption only from the perspective of workloads. However, they do not account for the potential autoscaling system resource consumption. In proactive horizontal pod autoscaling, workload resource consumption refers to the resource consumption caused by the running application workloads within the Kubernetes cluster. This metric helps determine how much load the system is handling and is crucial for making scaling decisions. On the other hand, autoscaling system resource consumption pertains to the resource consumption caused by the autoscaling system itself. An optimal proactive autoscaling system should minimize its own resource footprint while effectively scaling workloads to maintain performance and resource efficiency. Since various of these works leverage numerous forecasting models (one for each pod or set of pods), the overhead that would derive from their incorporation into the scaling decision-making process would be significant and more than likely negate the improvements in terms of workload resource consumption caused by proactively allocating, and de-allocating compute resources. To the best of our knowledge, no work exists in the corresponding scientific literature that aims at constructing a proactive horizontal autoscaling solution that improves both service performance and workload resource consumption while minimizing the underlying autoscaling system resource consumption. Our work is dedicated to mitigating this research gap. 3. Motivation The use of numerous DL models comes in stark contrast with our goal to construct a proactive horizontal autoscaling solution that improves both service performance and workload resource consumption and minimizes the underlying autoscaling system resource consumption. Having many local models in a system allows for task specialization, which can enhance forecasting accuracy by tailoring each model to capture the unique characteristics of individual compute nodes [27]. Local models are designed with architectures, features, and training datasets that optimize them for particular computing environments. However, a critical aspect of enabling reduced resource consumption in the frame of autonomic computing lies in the need to keep the number of forecasting models to a minimum. Each model requires resources to operate, meaning that a system with many local models will generally consume more resources than one with a single, consolidated global model. This demand can be substantial when the models are run frequently, as with real-time CPU forecasting in autonomic systems. Training multiple local models can be very resource-intensive, especially if each model is regularly updated to ensure optimal accuracy. Training typically consumes far more resources than inference, Future Generation Computer Systems 174 (2026) 107926 4 T. Theodoropoulos et al. and multiple training cycles for different models can lead to significantly higher energy costs, impacting the sustainability of the system. Balancing the outputs of multiple local models requires additional computational overhead, as the system must aggregate, reconcile, or prioritize forecasts from each model. This integration process can introduce latency, particularly in real-time scenarios, where autonomic responses are critical for maintaining system stability or optimizing resource allocation. Motivated by the aforementioned drawbacks of leveraging numerous local DL models, we propose the use of a singular global forecasting model for proactive horizontal autoscaling in order to improve service performance & resource consumption. Our main assumption is that one of the key advantages of using a global forecasting model in distributed computing environments is its ability to leverage shared underlying patterns across multiple pods. In many containerized applications, pods exhibit similar usage trends due to common workload types, synchronized traffic patterns, or shared infrastructure constraints. By recognizing these patterns, a global model can generalize across pods, leading to improved forecasting accuracy compared to individual pod-specific models that may overfit to transient fluctuations or noise. Furthermore, since pods operate within a cluster to host services, their resource consumption is often influenced by their service type and cluster membership. By incorporating these factors, a global forecasting model can effectively capture workload correlations across pods, leading to improved accuracy compared to isolated, pod-specific models that may overfit to transient variations. Pods hosting the same service tend to exhibit similar CPU usage trends due to synchronized request patterns, common workload dependencies, and shared execution environments. For instance, multiple pods serving the same API requests will likely experience correlated resource consumption trends, with peaks during high-traffic periods and lower usage during off-hours. A global forecasting model trained on aggregated data from such pods can identify these recurring patterns and seasonal fluctuations, allowing for more reliable predictions even when dealing with newly deployed or short-lived pods. The impact of shared underlying patterns is particularly valuable in autoscaling scenarios, where Kubernetes dynamically adjusts the number of pods in response to workload changes [28]. When a new pod is spawned, it lacks historical resource utilization data, making forecasting difficult for local models. However, a global model can instantly predict its expected resource consumption by leveraging data from existing pods running the same service. This reduces the lag time in making accurate forecasts, ensuring that resource allocation remains efficient and preventing overprovisioning and underprovisioning. 4. Problem formulation Multi-cluster deployments involve managing and orchestrating applications across multiple independent clusters of pods. These clusters are denoted by the set 𝐶=𝑐1, 𝑐2,…, 𝑐𝑐, with 𝑐𝑐 indicating the 𝑐th cluster, where 1≤𝑐≤|𝐶|. In the context of clusters, pods are the fundamental units that encapsulate and run one or more containers. They provide isolation, resource management, and a consistent deployment model across different clusters in container orchestration systems like Kubernetes. A pod is the smallest deployable unit in Kubernetes and can host one or more containers. Pods provide a layer of abstraction and encapsulation for their containers. These pods are denoted by the set 𝑃=𝑝1, 𝑝2,…, 𝑝𝑝, with 𝑝𝑝 indicating the 𝑝th pod, where 1≤𝑝≤|𝑃|. Furthermore, they are used to host services of various types. These types of services are denoted by the set 𝑆=𝑠1, 𝑠2,…, 𝑠𝑠, with 𝑠𝑠 indicating the 𝑠th type of service, where 1≤𝑠≤|𝑆|. Much like the service types mentioned above, each pod is intertwined with a specific replication time that highly depends on the service this pod hosts. The replication time for pods refers to the duration it takes for a pod in a container orchestration system (such as Kubernetes) to be replicated. This process involves pulling container images, setting up the environment, and initializing the application. Faster start-up times are desirable for efficient scaling and responsiveness in dynamic environments. These replication times are denoted by the set 𝑇= 𝑡1, 𝑡2,…, 𝑡𝑡, with 𝑡𝑡 indicating the 𝑡th replication time, where 1≤𝑡≤|𝑇|. Each pod 𝑝 is associated with a service type 𝑆 and a replication time 𝑇. Finally, each pod exhibits a certain type of resource consumption. In the frame of this work, resource consumption corresponds to the ongoing percentage of CPU utilization, as it is the most widely applied resource metric for CPU-intensive applications in Kubernetes [29]. The ongoing resource consumption at pod 𝑝 at time 𝑡 is denoted as 𝑅𝑝 𝑡. Table 3 summarizes the notations used in this work. Our work aims to introduce an advanced forecasting model that, through information refinement & fusion, is capable of producing more accurate resource consumption predictions regarding multiple pods. In time-series analysis, the multi-step formulation involves predicting future values of a time series by forecasting multiple time steps ahead. This approach contrasts with the single-step approach, which only estimates the next point in time. In the present challenge’s context, the output vector’s dimensional space is denoted as 𝑅|𝑃|∗𝑀, where |𝑃| represents the number of pods for which traffic predictions are intended at time point 𝑡, and 𝑀 represents the number of future steps for these projections. Similarly, the input vector’s dimensional space is defined as 𝑅|𝑃′|∗𝑀′, with |𝑃′| corresponding to pods whose resource consumption variations depend on those of 𝑃, and 𝑀′ indicating the number of preceding time steps contributing to the retrospective observation window (look-back window). It is essential to note that in the frame of this work, the value of 𝑀′ is equivalent to that of 𝑀. To delve further into our analysis, we concentrate on a specific time point 𝑡𝑖 and define the input vector 𝑋 as: 𝑋= {𝑥𝑖−𝑀′+1,…, 𝑥𝑖−𝑧′,…, 𝑥𝑖}, 𝑧′∈𝑀′,(1) where 𝑥𝑖−𝑧′=𝑅1 𝑡𝑖−𝑧′, 𝑅2 𝑡𝑖−𝑧′,…, 𝑅|𝑃′| 𝑡𝑖−𝑧′ represents the resource consumption of each pod 𝑝∈𝑃′ at time 𝑡𝑖−𝑧′. Similarly, we model the output vector 𝑌 as: 𝑌= {𝑦𝑖+1,…, 𝑦𝑖+𝑧,…, 𝑦𝑀}, 𝑧 ∈𝑀, (2) where 𝑦𝑖+𝑧=𝑅1 𝑡𝑖+𝑧, 𝑅2 𝑡𝑖+𝑧,…, 𝑅|𝑃| 𝑡𝑖+𝑧 represents resource consumption at pod 𝑝∈𝑃 at time 𝑡𝑖+𝑧. As our proposed solution is based on the information fusion & distillation properties of graph neural networks, it is crucial to transform the previous problem formulation into a graph format. There are various graph types, with a significant distinction being whether the considered graph structures are static or dynamic. Dynamic graphs can be categorized into Discrete-Time Dynamic Graphs (DTDG) [26] and Continuous-Time Dynamic Graphs (CTDG) [30]. This work adopts the DTDG approach to represent resource consumption across pods dynamically, since its ability to capture spatio-temporal dependencies in the frame of resource consumption forecasting has been documented in the corresponding scientific literature, as discussed previously in the Related Work section of this work. In the DTDG paradigm, a dynamic graph is defined as a sequence of snapshots of a static graph, each corresponding to a specific time-step 𝑡, with the duration between consecutive time-steps termed as 𝑡𝑤𝑖𝑛𝑑𝑜𝑤. These snapshots create a temporal continuum, enabling the emergence of temporal patterns. Each static graph consists of multiple nodes and edges representing spatial relations. In our context, each graph corresponds to a computational infrastructure that consists of |𝑃| Pods that exhibit a certain resource consumption behavior at each time-step 𝑡. Given an undirected graph 𝐺 with |𝑃| nodes and 𝐸 edges, nodes correspond to pods, and edges represent the correlations in resource consumption that emerge within the context of the various pods. This graph can be described by a weighted Adjacency Matrix 𝐀∈R|𝑃|×|𝑃| incorporating edge weights 𝑤𝑖𝑗 and a Feature Matrix 𝐅∈R|𝑃|×𝑉, where 𝑉 is the dimension of each feature vector. The Feature Matrix represents the collective of Feature Vectors. Each of the |𝑃| rows in the Feature Matrix corresponds to a Feature Future Generation Computer Systems 174 (2026) 107926 5 T. Theodoropoulos et al. Fig. 1. The pipeline of the proposed solution. Table 3 Notations used in this paper. Notations Descriptions R𝑛n-dimensional euclidean space 𝐶Set of clusters 𝑃Set of pods 𝑅𝑡 𝑝Resource consumption of 𝑃 𝑜𝑑𝑝 at time-step 𝑡 𝑆Set of services 𝑆𝑝Type of service deployed in 𝑃 𝑜𝑑𝑝 𝑇𝑝Replication time for 𝑃 𝑜𝑑𝑝 𝑀Number of input time-steps 𝑀′Number of prediction time-steps 𝑡𝑤𝑖𝑛𝑑𝑜𝑤 Duration between consecutive time-steps 𝑉Dimension of Feature Vector 𝐹Feature Matrix 𝑋Input Matrix 𝐺Undirected Graph 𝐴Adjacency Matrix with edge weights 𝑤𝑖𝑗 𝐼Identity Matrix 𝑊Learnable weight matrix |𝐸|The number of elements in a given set 𝐸 Vector describing node-specific attributes. In this study, a graph node (pod) is characterized by a Feature Vector with a dimension equal to 𝑀′, representing resource consumption values recorded at the respective pod over the last 𝑀′ time intervals, hence 𝑉=𝑀′. Moreover, each of the 𝑀′ columns in the Feature Matrix 𝐹 represents a distinct time interval 𝑡 within the input sequence. This setup allows instances of the Feature Matrix to be treated as time-series data. The Adjacency Matrix 𝐴 remains constant, representing the resource consumption correlation among the various Pods. Conversely, the Feature Matrix 𝐹 is dynamic and varies for each time interval 𝑡. Consequently, snapshots of the Feature Matrix 𝐹 are conceptually analogous to the aforementioned input vector 𝑋∈R|𝑃|×𝑀′. 5. GraphOpticon This work aims to establish a proactive horizontal autoscaling solution that is capable of simultaneously achieving better service performance and reduced resource consumption. The proposed solution operates based on accurate resource consumption predictions that involve numerous pods across numerous time steps. To that end, the authors of this work propose GraphOpticon, a proactive horizontal autoscaling solution that leverages the information fusion & distillation properties of graph neural networks. GraphOpticon consists of 4 components. These components are: •Monitoring Component •Input Construction Component •Forecasting Component •Output Distillation Component These components are designed to operate as parts of a singular pipeline. This pipeline is depicted in Fig. 1. 5.1. Monitoring component This component is built upon the functionalities of modern monitoring frameworks, such as Prometheus. It is designed to perform three distinct operations related to information retrieval. First, it periodically scrapes CPU utilization values for each pod at every time-step 𝑡 to construct the Input Matrix 𝑋. The second operations involve the construction of the weighted Adjacency Matrix 𝐴. While the Monitoring Component does not directly compute the weights of 𝐴, it is responsible for gathering the necessary data and transmitting it to the Input Construction Component, which performs the weight calculations. Towards assisting in the construction of 𝐴, the Monitoring Component retrieves information about the cluster 𝐶 each pod 𝑝 belongs to and the service 𝑆 it hosts. This process occurs once during the initialization of GraphOpticon. The third and final operation of the Monitoring Component involves conveying the estimated replication times 𝑇 of all examined services to the Output Distillation Component. 5.2. Input construction component This component uses the name of the cluster 𝐶 that a pod 𝑝 belongs to, along with the name of the service 𝑆 the pod hosts, to create the Adjacency Matrix 𝐴. The aforementioned input information is provided by the Monitoring Component. Graph edges are crucial for illustrating the relationships among pods and determining which nodes shall be involved in the feature aggregation process that shall be performed by the encoder part of the Forecasting Component. The encoder part of the Forecasting Component is based on Graph Convolutional Networks (GCNs) [31] . Feature aggregation, in the frame of GCNs, is the process of combining node features from a node’s neighbors (and sometimes itself) using weighted summation or averaging to capture local graph structure and propagate information. In this paper, we aim to leverage the relations that are present among the various pods in order to construct the corresponding Adjacency Matrix 𝐴. According to this approach, the Adjacency Matrix 𝐴 is formed based on the service 𝑆 and cluster 𝐶 similarities between each pair of pods (nodes). The algorithm for calculating the corresponding weights is detailed in Algorithm 1. If two pods do not facilitate the same type of service 𝑆𝑝, then the corresponding weight associated with edge 𝑒𝑖𝑗 equals zero. If two Pods facilitate the same type of service 𝑆𝑝, but they do not belong to the same cluster 𝐶, then the corresponding weight associated with edge 𝑤𝑖𝑗 equals 0.5. Finally, if two Pods facilitate the same type of service 𝑆𝑝, and they belong to the same cluster 𝐶, then the corresponding weight associated with edge 𝑒𝑖𝑗 equals 1. An example of this approach is showcased in Fig. 2. This figure depicts 4 clusters (1,2,3,4) with 16 pods that host 4 different types of services (red, green, blue, yellow). According to the proposed approach, pods that belong to the same cluster 𝐶 and host the same type of service 𝑆 (inside black rectangles) shall have edge weights 𝑤 equal to 1. Pods that host the same service 𝑆 and belong to different clusters 𝐶 (connected by a colored edge) shall have edge weights 𝑤 equal to 0. The rest shall have edge weights equal to 0. Future Generation Computer Systems 174 (2026) 107926 6 T. Theodoropoulos et al. Fig. 2. Input construction process. The underlying rationale for leveraging the Input Construction Component is to streamline the feature aggregation process by restricting it to pods hosting the same service type. This simplifies the model’s complexity by focusing on distilled spatial correlations inherent in the input structure. Moreover, the forecasting model can capture more refined dependencies by introducing the weight differentiation between pods that inhabit the same cluster and pods that do not. The proposed weight assignment derives from the assumption that pods that host the same type of service are expected to present more similarities in terms of their resource (CPU) consumption patterns when compared with pods that do not host the same type of service. Furthermore, pods that host the same type of service and belong to the same cluster are expected to present more similarities in terms of their resource (CPU) consumption patterns when compared with pods that host the same type of service but do not belong to the same cluster. This assumption serves as an extension to our line of thought that was presented in Section 3. Algorithm 1 Input Construction Algorithm. Input: The |𝑃| Pods, alongside their corresponding 𝐶𝑝 and 𝑆𝑝 attributes, which describe the cluster that each pod belongs to and the type of service each pod facilitates, respectively. Output: The weighted Adjacency Matrix 𝐴. Begin algorithm 1. For each pair of pods 𝑖, 𝑗 in 𝑃: 2. If 𝑆𝑖≠𝑆𝑗: 3. Then 𝐴𝑖, 𝑗 ←0. 4. If (𝑆𝑖=𝑆𝑗) ∧ (𝐶𝑖≠𝐶𝑗): 5. Then 𝐴𝑖, 𝑗 ←0.5. 6. If (𝑆𝑖=𝑆𝑗) ∧ (𝐶𝑖=𝐶𝑗) : 7. Then 𝐴𝑖, 𝑗 ←1. 8. Return 𝐴 End algorithm 5.3. Forecasting component We have combined a weighted Graph Convolutional Network (GCN) [31] and LSTM layers to construct an encoder–decoder architecture capable of predicting resource consumption. Encoder–decoder architectures for time-series forecasting involve an encoder that processes input sequences into a fixed representation and a decoder that uses this representation to predict future sequences. The weighted GCNs serve as the encoder to extract structural features from the input sequence to generate a consolidated representation. This operation is conducted as follows: 𝐻encoder =Weighted GCNencoder(𝑋, 𝐴)(3) Here, ℎencoder denotes the consolidated representation post the application of weighted stacked graph convolution, where 𝐴 signifies the weighted Adjacency Matrix of the graph, and 𝑋 indicates the Feature Matrix, which is the input of the forecasting entity. The weighted Adjacency Matrix 𝐴 is provided by the Input Construction Component, while the Feature Matrix 𝑋 is provided by the Monitoring Component. Subsequently, the constructed representation is channeled into the LSTM component of the model, enabling the capture of temporal patterns at the level of graph snapshots. Functioning as a decoder, the LSTM component generates the desired forecasts through the following process: 𝑌=LSTMdecoder(𝐻encoder)(4) The term LSTMdecoder denotes the LSTM network, which accepts the aggregated representation from the encoder to produce an output that is then passed through a dense layer for multi-step prediction generation. In the context of multi-step time-series prediction, the Forecasting Component takes a sequence of graph signals as input, where each signal represents a different time-step and is depicted as a graph signal on a consistent graph. The objective is to forecast future time-series values based on the graph signals from preceding time steps (see Fig. 3). 5.3.1. Encoder (Weighted GCN layer) The encoder part of the Forecasting Component is based on the use of weighted Graph Convolutional Networks (GCNs). Weighted GCNs are a powerful framework for learning representations of nodes in graphs, considering both the graph structure and features associated with nodes. Let 𝐗 denote the feature matrix of size |𝑃|×𝑀, where 𝑁 is the number of nodes and 𝑀 is the number of features per node. The weighted adjacency matrix 𝐀 represents the connections between nodes in the graph. The propagation rule in a weighted GCN is defined as: 𝐇(𝑙+1) =𝜎( 𝐃−1 2 𝐀 𝐃−1 2𝐇(𝑙)𝐖(𝑙))(5) where: •𝐇(𝑙) is the feature representation matrix at layer 𝑙, • 𝐀 is the weighted adjacency matrix with added self-connections, • 𝐃 is the degree matrix of  𝐀, •𝐖(𝑙) is the weight matrix of layer 𝑙, •𝜎 is the ReLU activation function. The input feature matrix 𝐻(0) is typically initialized as 𝑋, which is the input sequence provided by the Monitoring Component. Furthermore, in the frame of this work, 𝑙= 1. 5.3.2. Decoder (LSTM layer) The decoder part of the Forecasting Component is based on Long Short-Term Memory (LSTM) networks. LSTM networks employ the Hidden State mechanism to capture dynamic temporal patterns. What sets LSTM networks apart is their utilization of the Cell State structure, introducing Cell State manipulation through regulatory mechanisms known as Gates. Each LSTM node comprises three gate-related elements, all incorporating sigmoid layers to ensure differentiation within the range of 0 − 𝑡𝑜 − 1. The sigmoid activation function scales values to facilitate the assessment of data importance and decision-making regarding retention or omission. Gate structures include two sets of weight matrices, labeled 𝑊 and 𝑈, associated with Hidden State and input and additional matrices for Cell State. The input 𝑋𝑡 corresponds to timestamp 𝑡. Gates utilize these matrices, input, and prior Hidden State (ℎ𝑖𝑑𝑑𝑒𝑛𝑡−1). Future Generation Computer Systems 174 (2026) 107926 7 T. Theodoropoulos et al. Fig. 3. Architecture of the forecasting component. The Forget Gate determines which historical information from past timestamps to exclude from the Cell State. Its output is computed using Eq. (6). The Input Gate assesses the significance of recent input, updating the Cell State using Eq. (7). Cell State calculation employs the C vector, generated as per Eq. (8), with the tanh activation function mitigating gradient issues. The Cell State update process is described in Eq. (9), combining the output of the Forget Gate and the Input Gate with C. The Output Gate computes the subsequent hidden state using Eq. (10). The new Hidden State is calculated according to Eq. (11). Updated Cell State and Hidden State are then propagated to subsequent LSTM nodes for the next time-step [32]. forget𝑡=sigmoid(𝑋𝑡⋅𝑊𝑓+hidden𝑡−1 ⋅𝑈𝑓)(6) input𝑡=sigmoid(𝑋𝑡⋅𝑊𝑖+hidden𝑡−1 ⋅𝑈𝑖)(7) C=tanh(𝑋𝑡⋅𝑊𝑐+hidden𝑡−1 ⋅𝑈𝑐)(8) 𝐶𝑡=forget𝑡⋅𝐶𝑡−1 +input𝑡⋅C𝑡(9) output𝑡=sigmoid(𝑋𝑡⋅𝑊𝑜+hidden𝑡−1 ⋅𝑈𝑜)(10) hidden𝑡=output𝑡⋅tanh(𝐶𝑡)(11) In the frame of the decoder, 𝑜𝑢𝑡𝑝𝑢𝑡𝑡 is passed through a dense layer to construct the desired 2-dimensional output 𝑌 shape. 5.4. Output distillation component The Forecasting Component produces a multi-step output 𝑌 corresponding to 𝑃 pods. The Output Distillation Component is designed to receive the multi-step prediction 𝑌 generated by the Forecasting Component as input, along with the replication time 𝑇 corresponding to each pod 𝑝. Upon receiving these inputs, the Output Distillation Component determines, based on each pod, which step from the multi-step prediction should be retained to produce a single-step prediction. This functionality is conducted using a dedicated list (𝑁), which associates each replication time 𝑇 with a corresponding time-step of the multi-step prediction. Calculating 𝑁 is showcased using Algorithm 2. This decision is guided by the necessity of selecting a time step further into the future than the anticipated completion time of the deployment process. Additionally, it prioritizes a time step closer to the expected completion moment of the deployment process, as looking too far ahead compromises prediction accuracy. This process tailors predictions to each pod. Fig. 4 depicts an example of how the proposed component selects the prediction step to retain. 6. System implementation Before proceeding to the experimental evaluation section of this work, it is of paramount importance to address any potential limitations Fig. 4. Output distillation process. that the architecture of GraphOpticon may present. These potential limitations involve aspects such as the compatibility of GraphOpticon with contemporary frameworks, as well as its scalability, and the underlying formalism that is used in the frame of this work. This section is dedicated to addressing these limitations. 6.1. Integration The aforementioned components that were described in the previous section of this work are designed to be integrated into cloud and edge computing frameworks such as Kubernetes or K3s,2 inheriting foundational orchestration and management mechanisms. GraphOpticon does not seek to replace these established frameworks but aims to enhance them, specifically regarding improved service performance and resource consumption. By limiting the required changes to the scaling mechanisms, GraphOpticon maintains the overall stability and 2https://k3s.io/. Future Generation Computer Systems 174 (2026) 107926 8 T. Theodoropoulos et al. Algorithm 2 Output Distillation Algorithm Input: List of replication times 𝑇= [𝑡1, 𝑡2,…, 𝑡|𝑃|], and list of prediction time-steps 𝑀= [𝑚1, 𝑚2,…, 𝑚|𝑀|]. The temporal difference between each time step equals 𝑡𝑤𝑖𝑛𝑑𝑜𝑤. Output: List of singular time-steps 𝑁 corresponding to each replication time 𝑇. Begin algorithm 1.Sort 𝑀 in ascending order 2.Initialize 𝑁 as an empty list 3.For each 𝑡 in 𝑇: 4. Initialize 𝑛𝑒𝑥𝑡_𝑠𝑡𝑒𝑝 ←𝑁𝑜𝑛𝑒 5. For each 𝑚 in 𝑀: 6. If 𝑚 > 𝑡: 7. Then 𝑛𝑒𝑥𝑡_𝑠𝑡𝑒𝑝 ←𝑚 8. break 9. Append 𝑛𝑒𝑥𝑡_𝑠𝑡𝑒𝑝 to 𝑁 10.Return 𝑁 End algorithm reliability of the existing orchestration frameworks. This strategy minimizes disruption and complexity, allowing for a smoother integration while still achieving the desired improvements. Enhancing this aspect without altering other parts of the framework ensures that GraphOpticon can offer performance improvements without the need for major changes or overhauls, thus preserving the integrity and usability of the underlying orchestration system. The same principles have also been applied in regard to its compatibility with the leveraged monitoring framework. The Monitoring Component is designed in a manner that enables it to seamlessly integrate with the Prometheus3 monitoring system. Its primary functions include monitoring application components, whether deployed as pods or virtual machines, and monitoring hosts, whether physical or virtual, on Kubernetes or K3s clusters located in the cloud or at the edge. For our case, it is necessary to retrieve the CPU utilization of pods. In the Monitoring Component’s setup, Prometheus periodically pulls metrics from monitoring agents named Prometheus exporters. Particularly for pod monitoring, the Monitoring Component utilizes kube-state-metrics4 exporter to perform pod monitoring. This exporter provides metrics by interfacing with the Kubernetes API server and incorporating data from both Kubelet and cAdvisor. The cAdvisor exporter collects resource usage statistics, including CPU utilization, for all running pods. Kubelet is an agent on each node in a Kubernetes cluster, working with the Kubernetes API server to collect information on node events, pod statuses, and resource usage. The leveraged architectural design [33] enables the Monitoring Component to detect the replication time of pods by utilizing the kubestate-metrics exporter and its above-mentioned capability to monitor pod statuses. The detection of a running pod produces an alert that includes the deployment time as a timestamp, the number of replicas, the namespace to which the pod belongs, the host node, node IP, pod IP, pod status, external IPs, and the applied pod labels. The above alert was configured to meet our case requirements by including the master node IP indicating the cluster’s IP and a master’s label indicating the cluster name. This feature can match pods with the cluster they are running in. This addition was also achieved based on metrics retrieved by kubestate-metrics. A naming convention is essential to accurately identify the name of the service provided by a pod. Specifically, the pod name should be associated with the service name it provides, ensuring that the monitoring mechanism generates meaningful alerts for the entire system. 3https://prometheus.io/. 4https://github.com/kubernetes/kube-state-metrics. 6.2. Scalability analysis The centralized architecture of GraphOpticon raises several scalability concerns. Chief among them is the potential for performance bottlenecks, as the autoscaler must process vast amounts of data from numerous pods, which could lead to increased latency and delayed scaling decisions. There is the risk that as the system scales, the overhead associated with data collection and processing could become a significant drain on resources, further exacerbating performance issues. Calculating the complexity of the proposed autoscaler provides a structured method to address these concerns. By analyzing the corresponding complexity, one can quantify how the autoscaler’s performance degrades as the number of pods and metrics increases. This quantitative insight allows for the identification of potential bottlenecks and critical thresholds, guiding the design towards more efficient algorithms and data handling methods. Towards examining the scalability of the proposed solution, we analyze the computational complexity of the data collection and processing algorithms that are leveraged in the frame of GraphOpticon. These include the Monitoring process, the Input Construction process, and the Output Distillation process. Since scalability is investigated in relation to the number of pods, complexity shall be calculated on the basis of the number of pods 𝑃. In order to construct the Input Matrix 𝑋, the Monitoring component periodically retrieves CPU utilization metrics from the Metrics Server or an external monitoring system. Given that the Monitoring component queries metrics for each running pod individually, the time complexity of this operation is 𝑂(𝑃), where 𝑃 is the number of pods being monitored. The same complexity applies to the retrieval of the required information that is later used by the Input Construction Component to create the Adjacency Matrix 𝐴. The Input Construction Component is responsible for generating the Adjacency Matrix 𝐴 based on relationships between pods. This process involves iterating over all 𝑃 pods and performing pairwise comparisons to establish edge weights. In the worst case, each pod is compared against every other pod, leading to a complexity of 𝑂(𝑃2). The quadratic growth arises due to the necessity of examining all potential connections within the system. However, it is of paramount importance to point out that this process is fully carried out only when GraphOpticon is instantiated. The Output Distillation component involves selecting the most relevant prediction step for each pod. Given 𝑀 sorted prediction timesteps, the nearest step can be identified using a binary search operation with complexity 𝑂(log |𝑀|). Since this operation is performed for each of the 𝑃 pods, the total complexity is given by 𝑂((𝑃+|𝑀|) log |𝑀|). Based on these results, the overall complexity of GraphOpticon shall be equal to 𝑂(𝑃2) for its instantiation and equal to 𝑂((𝑃+|𝑀|) log |𝑀|) during its operation. In other words, when the Adjacency Matrix 𝐴 is calculated, the complexity is dominated by the quadratic term, yielding 𝑂(𝑃2). However, after the initial matrix is constructed, the subsequent complexity is equal to 𝑂((𝑃+|𝑀|) log |𝑀|). Since we expect that 𝑀 < 𝑃 , the term 𝑃+|𝑀| is dominated by 𝑃. So it simplifies to 𝑂(𝑃). As a result, the complexity becomes approximately 𝑂(𝑃log |𝑀|). In this case, 𝑙𝑜𝑔|𝑀| grows slower than 𝑃, and the overall complexity will still be influenced by 𝑃, with a logarithmic factor that does not change the linear dependence on 𝑃. So, this complexity can be characterized as linear with a logarithmic correction. Thus, it is safe to conclude that GraphOpticon constitutes a scalable solution, the complexity of which grows in an almost linear manner as the number of pods increases. 6.3. Formalism In the frame of Kubernetes’ autoscaling operation, the number of pods dynamically increases and decreases. As a result, the size and structure of the Adjacency Matrix 𝐴 would constantly change over time. However, this would contradict the CTDG paradigm on which GraphOpticon is based. Towards resolving this issue, in the frame of Future Generation Computer Systems 174 (2026) 107926 9 T. Theodoropoulos et al. resource consumption, skewness, and kurtosis of latency, indicating greater inconsistency and more frequent outliers in task execution times. This discrepancy can be attributed to the different methods used for horizontal autoscaling in Kubernetes and the Standard approach. The Standard approach utilizes two distinct thresholds for scaling in and out, providing a more controlled and predictable mechanism for resource allocation. When the system load reaches the upper threshold, additional resources are allocated to manage the increased demand, and when the load drops below the lower threshold, excess resources are released. This binary threshold system ensures that resources are scaled in a predictable manner, which helps maintain a more consistent performance across varying loads. In contrast, Kubernetes employs HPA, which uses a more dynamic formula for scaling. The HPA adjusts the number of running pods based on the observed CPU utilization against a target utilization defined by the user. The formula calculates the desired number of pods as a ratio of the current CPU utilization to the target utilization. While this method allows for more fine-tuned and responsive scaling, it can also lead to greater variability in execution times due to the rapid and frequent adjustments in resource allocation. This can result in a system that is more susceptible to sudden spikes in demand, leading to higher maximum execution times, larger ranges, and increased skewness and kurtosis in the distribution of latency. The proactive horizontal autoscaling approach employed by GraphOpticon is a key factor in its superior performance. Unlike the Standard and Kubernetes approaches, which react to changes in demand, GraphOpticon anticipates demand fluctuations and scales resources accordingly. This proactive approach leads to lower execution time and latency by scaling resources ahead of anticipated demand spikes, minimizing the occurrence of high-latency events, and maintaining a high quality of service. In addition, proactive scaling ensures that resources are allocated precisely when needed, avoiding both underutilization and overprovisioning. This efficient resource management reduces the average number of compute nodes used, lowering operational costs, and improving system efficiency. GraphOpticon’s efficiency can be attributed to two factors. The first is its ability to encapsulate temporal patterns throughout the temporal continuum. The second is its ability to further refine the encapsulation of these temporal patterns through information fusion and distillation. GraphOpticon, contrary to the other examined approaches, takes into consideration the ongoing and past CPU consumption values across numerous compute nodes to produce the values upon which the horizontal autoscaling process shall take place. These produced values provide valuable information regarding the number of pods that will be deployed in the near future. These insights can be leveraged to conduct the horizontal autoscaling process in an advanced manner that is associated with enhanced service performance and reduced resource consumption. A sudden burst in resource consumption can be followed by: (i) a further increase in resource consumption and (ii) a decrease in resource consumption. GraphOpticon is capable of leveraging various temporal patterns to calculate which one of the scenarios is more likely to play out and formulate its predictions accordingly. In case a gradual increase follows the burst in resource consumption, GraphOpticon will proactively allocate additional computational resources to mitigate the ramifications on service performance that would derive from the insufficiency of resources. In case a decrease follows the burst in resource consumption, GraphOpticon will either de-allocate a compute node or maintain the ongoing compute node configuration as is. This is essential for achieving reduced resource consumption. Aside from reduced resource consumption, this approach can achieve better service performance. In the scenario of a sudden burst in resource consumption, the reactive approaches would allocate more compute nodes. Newly produced tasks would then be assigned to these newly added compute nodes. However, if the sudden burst is just a random event, which is followed by a decrease in resource consumption, then resource overprovisioning would emerge, and subsequently, the newly added compute nodes would be de-allocated. The tasks that were assigned to the de-allocated nodes shall have to be re-assigned to other nodes, thus increasing latency and the overall execution time. Furthermore, when encountering a sudden drop in resource demand, GraphOpticon can effectively manage two possible scenarios: (i) a further decrease in resource usage or (ii) a subsequent increase. By analyzing various temporal patterns, GraphOpticon predicts which scenario is more likely and adjusts its resource allocation strategy accordingly. If the drop is expected to be followed by an additional decrease, GraphOpticon proactively de-allocates compute nodes or maintains the current configuration without adding more resources. This approach prevents resource overprovisioning, reduces unnecessary costs, and enhances resource efficiency. By scaling down resources only when necessary, GraphOpticon ensures a lean operational state, which is crucial for minimizing resource consumption. On the other hand, if the drop is anticipated to be temporary and followed by a gradual increase in resource consumption, GraphOpticon retains or slightly increases resource allocation. This proactive strategy helps mitigate potential performance issues arising from resource shortages when demand rises again. These performance issues also include the fact that the tasks of the previously de-allocated nodes would have to be re-assigned to alternative compute nodes, thus increasing latency and the overall execution time. By forecasting the need for more resources in advance, GraphOpticon maintains stable and efficient service performance, avoiding the drawbacks of reactive scaling. This constant reassignment process increases latency and overall execution time, negatively impacting service performance. Additionally, the rapid scaling up and down of nodes can introduce instability and additional overhead, further exacerbating performance issues. This phenomenon is referred to as resource oscillation. As showcased in the frame of this work, GraphOpticon is capable of mitigating resource oscillation by leveraging predictive analytics. The Proactive manages to consistently outperform the Standard and Kubernetes approaches in terms of service performance, yet it performs slightly worse than GraphOpticon. This is due to the fact that although it is also a proactive approach and thus can outperform reactive approaches, as is evident by the prior discussion, its forecasting process is not as accurate as that of GraphOpticon. Proactive exhibits the highest resource consumption among all approaches when accounting for the computational resources that are required to host the various forecasting models. In the event that we do not account for these additional resources, Proactive manages to outperform the Standard and Kubernetes approaches. Similarly to GraphOpticon, it is able to de-allocate resources proactively, which can lead to reduced resource consumption compared to reactive approaches. However, the additional resource overhead that is required in order to facilitate the 5 forecasting models negates the benefits of proactive horizontal autoscaling in terms of resource consumption. These experimental results, alongside the fact that GraphOpticon is considerably more lightweight than even a singular instance of the Proactive approach, enable us to safely conclude that GraphOpticon is the superior option in terms of improving service performance while reducing the resource consumption that is associated with workloads and the autoscaling system itself. 8. Conclusions & Future research directions In this work, we introduced GraphOpticon, a global proactive horizontal pod autoscaling solution aimed at optimizing service performance and reducing resource consumption. It integrates four key components: Monitoring, Input Construction, Forecasting, and Output Distillation, utilizing GNNs for information fusion and distillation. The Input Construction component refines input sequence representations, enabling accurate resource consumption predictions across multiple Future Generation Computer Systems 174 (2026) 107926 16 T. Theodoropoulos et al. pods and time steps through a global Forecasting mechanism. Output Distillation then leverages these predictions to generate tailored insights specific to pod characteristics. By leveraging predictive analytics, GraphOpticon maintains efficient resource usage and enhanced service performance, avoiding the pitfalls of overprovisioning and reassignment delays seen in reactive approaches. This proactive management ensures stability and efficiency, delivering optimal results in varying demand scenarios. These findings highlight the advantages of proactive scaling strategies in contemporary distributed systems. GraphOpticon not only reduces latency and execution time but also improves workload source consumption while minimizing autoscaling system resource consumption, ensuring a more reliable and efficient distributed computing environment. The aforementioned decrease in workload resource consumption and minimization of autoscaling system resource consumption is expected to result in enhanced cost-savings and energy efficiency. As organizations increasingly rely on cloud-based infrastructure and distributed systems, the need for more sustainable and cost-effective solutions becomes critical. As such, future work shall focus on developing more detailed models to quantify the financial and environmental benefits of GraphOpticon’s proactive autoscaling approach. This will involve conducting large-scale evaluations across various industries and workloads to measure the real-world impact of improved resource utilization. CRediT authorship contribution statement Theodoros Theodoropoulos: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Yashwant Singh Patel: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Resources, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Uwe Zdun: Writing – review & editing, Writing – original draft, Supervision, Project administration, Investigation, Formal analysis, Conceptualization. Paul Townend: Writing – review & editing, Writing – original draft, Supervision, Investigation, Formal analysis, Conceptualization. Ioannis Korontanis: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Antonios Makris: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Resources, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Konstantinos Tserpes: Writing – review & editing, Writing – original draft, Supervision, Project administration, Investigation, Formal analysis, Conceptualization. Declaration of competing interest The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Theodoros Theodoropoulos reports financial support was provided by Harokopio University of Athens. Theodoros Theodoropoulos reports financial support was provided by University of Vienna. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Data availability Data will be made available on request. References [1] E.F. Coutinho, F.R. de Carvalho Sousa, P.A.L. Rego, D.G. Gomes, J.N. de Souza, Elasticity in cloud computing: a survey, Ann. Telecommun. Ann. Télécommun. 70 (2015) 289–309. [2] N. Roy, A. Dubey, A. Gokhale, Efficient autoscaling in the cloud using predictive models for workload forecasting, in: 2011 IEEE 4th International Conference on Cloud Computing, IEEE, 2011, pp. 500–507. [3] Horizontal pod autoscaling, section: docs. [online], 2024, Available: https:// kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/. [4] B. Shiva Prakash, K. Sanjeev, R. Prakash, K. Chandrasekaran, A survey on recurrent neural network architectures for sequential learning, in: Soft Computing for Problem Solving: SocProS 2017, Volume 2, Springer, 2019, pp. 57–66. [5] T. Theodoropoulos, J. Violos, S. Tsanakas, A. Leivadeas, K. Tserpes, T. Varvarigou, Intelligent proactive fault tolerance at the edge through resource usage prediction, 2023, arXiv preprint arXiv:2302.05336. [6] T. Theodoropoulos, A.-C. Maroudis, J. Violos, K. Tserpes, An encoder-decoder deep learning approach for multistep service traffic prediction, in: 2021 IEEE Seventh International Conference on Big Data Computing Service and Applications, BigDataService, IEEE, 2021, pp. 33–40. [7] Y. Zhou, H. Zheng, X. Huang, S. Hao, D. Li, J. Zhao, Graph neural networks: Taxonomy, advances, and trends, ACM Trans. Intell. Syst. Technol. (TIST) 13 (1) (2022) 1–54. [8] Z.A. Sahili, M. Awad, Spatio-temporal graph neural networks: A survey, 2023, arXiv preprint arXiv:2301.10569. [9] N. Marie-Magdelaine, T. Ahmed, Proactive autoscaling for cloud-native applications using machine learning, in: GLOBECOM 2020-2020 IEEE Global Communications Conference, IEEE, 2020, pp. 1–7. [10] M. Imdoukh, I. Ahmad, M.G. Alfailakawi, Machine learning-based auto-scaling for containerized applications, Neural Comput. Appl. 32 (13) (2020) 9745–9760. [11] A. Goli, N. Mahmoudi, H. Khazaei, O. Ardakanian, A holistic machine learningbased autoscaling approach for microservice applications, Closer 1 (2021) 190–198. [12] L. Ju, P. Singh, S. Toor, Proactive autoscaling for edge computing systems with kubernetes, in: Proceedings of the 14th IEEE/ACM International Conference on Utility and Cloud Computing Companion, 2021, pp. 1–8. [13] N.-M. Dang-Quang, M. Yoo, Deep learning-based autoscaling using bidirectional long short-term memory for kubernetes, Appl. Sci. 11 (9) (2021) 3835. [14] M.P. Yadav, Rohit, D.K. Yadav, Maintaining container sustainability through machine learning, Clust. Comput. 24 (4) (2021) 3725–3750. [15] H.X. Nguyen, S. Zhu, M. Liu, Graph-PHPA: graph-based proactive horizontal pod autoscaling for microservices using LSTM-GNN, in: 2022 IEEE 11th International Conference on Cloud Networking, CloudNet, IEEE, 2022, pp. 237–241. [16] J. Violos, S. Tsanakas, T. Theodoropoulos, A. Leivadeas, K. Tserpes, T. Varvarigou, Intelligent horizontal autoscaling in edge computing using a double tower neural network, Comput. Netw. 217 (2022) 109339. [17] J. Dogani, F. Khunjush, M. Seydali, K-agrued: a container autoscaling technique for cloud-based web applications in kubernetes using attention-based gru encoder-decoder, J. Grid Comput. 20 (4) (2022) 40. [18] S. Kakade, G. Abbigeri, O. Prabhu, A. Dalwayi, S.P. Patil, B. Sunag, et al., Proactive horizontal pod autoscaling in kubernetes using bi-lstm, in: 2023 IEEE International Conference on Contemporary Computing and Communications, InC4, vol. 1, IEEE, 2023, pp. 1–5. [19] T. Theodoropoulos, A. Makris, I. Kontopoulos, J. Violos, P. Tarkowski, Z. Ledwoń, P. Dazzi, K. Tserpes, Graph neural networks for representing multivariate resource usage: A multiplayer mobile gaming case-study, Int. J. Inf. Manag. Data Insights 3 (1) (2023) 100158. [20] J. Dogani, F. Khunjush, Proactive auto-scaling technique for web applications in container-based edge computing using federated learning model, J. Parallel Distrib. Comput. 187 (2024) 104837. [21] H. Ahmad, C. Treude, M. Wagner, C. Szabo, Towards resource-efficient reactive and proactive auto-scaling for microservice architectures, J. Syst. Softw. (2025) 112390. [22] X. Tang, Q. Liu, Y. Dong, J. Han, Z. Zhang, Fisher: An efficient container load prediction model with deep neural network in clouds, in: 2018 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Ubiquitous Computing & Communications, Big Data & Cloud Computing, Social Computing & Networking, Sustainable Computing & Communications, ISPA/IUCC/BDCloud/SocialCom/SustainCom, IEEE, 2018, pp. 199–206. [23] Y.S. Patel, R. Jaiswal, R. Misra, Deep learning-based multivariate resource utilization prediction for hotspots and coldspots mitigation in green cloud data centers, J. Supercomput. 78 (4) (2022) 5806–5855. [24] E. Radhika, G.S. Sadasivam, J.F. Naomi, An efficient predictive technique to autoscale the resources for web applications in private cloud, in: 2018 Fourth International Conference on Advances in Electrical, Electronics, Information, Communication and Bio-Informatics, AEEICB, IEEE, 2018, pp. 1–7. [25] T. Theodoropoulos, A. Makris, I. Kontopoulos, A.-C. Maroudis, K. Tserpes, Multi-service demand forecasting using graph neural networks, in: 2023 IEEE International Conference on Service-Oriented System Engineering, SOSE, IEEE, 2023, pp. 218–226. Future Generation Computer Systems 174 (2026) 107926 17 T. Theodoropoulos et al. [26] B. Yu, H. Yin, Z. Zhu, Spatio-temporal graph convolutional networks: A deep learning framework for traffic forecasting, in: Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, International Joint Conferences on Artificial Intelligence Organization, 2018, http://dx.doi.org/10. 24963/ijcai.2018/505. [27] D. Jakubovitz, R. Giryes, M.R. Rodrigues, Generalization error in deep learning, in: Compressed Sensing and Its Applications: Third International MATHEON Conference 2017, Springer, 2019, pp. 153–193. [28] M.A. Tamiru, J. Tordsson, E. Elmroth, G. Pierre, An experimental evaluation of the kubernetes cluster autoscaler in the cloud, in: 2020 IEEE International Conference on Cloud Computing Technology and Science, CloudCom, IEEE, 2020, pp. 17–24. [29] E. Casalicchio, V. Perciballi, Auto-scaling of containers: The impact of relative and absolute metrics, in: 2017 IEEE 2nd International Workshops on Foundations and Applications of Self* Systems, FAS* W, IEEE, 2017, pp. 207–214. [30] E. Rossi, B. Chamberlain, F. Frasca, D. Eynard, F. Monti, M. Bronstein, Temporal graph networks for deep learning on dynamic graphs, 2020, http://dx.doi.org/ 10.48550/ARXIV.2006.10637, URL https://arxiv.org/abs/2006.10637. [31] Y. Zhao, J. Qi, Q. Liu, R. Zhang, WGCN: graph convolutional networks with weighted structural features, in: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021, pp. 624–633. [32] J. Chung, C. Gulcehre, K. Cho, Y. Bengio, Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014, arXiv preprint arXiv:1412.3555. [33] I. Korontanis, A. Makris, T. Theodoropoulos, K. Tserpes, Real-time monitoring and analysis of edge and cloud resources, in: Proceedings of the 3rd Workshop on Flexible Resource and Application Management on the Edge, FRAME ’23, Association for Computing Machinery, New York, NY, USA, ISBN: 9798400701641, 2023, pp. 13–18, http://dx.doi.org/10.1145/3589010.3594892. [34] M.C. Silva Filho, R.L. Oliveira, C.C. Monteiro, P.R.M. Inácio, M.M. Freire, CloudSim plus: A cloud computing simulation framework pursuing software engineering principles for improved modularity, extensibility and correctness, in: 2017 IFIP/IEEE Symposium on Integrated Network and Service Management, IM, 2017, pp. 400–406. [35] M. Zakarya, L. Gillam, A.A. Khan, I.U. Rahman, Perficientcloudsim: a tool to simulate large-scale computation in heterogeneous clouds, J. Supercomput. 77 (4) (2021) 3959–4013. [36] M. Zakarya, L. Gillam, K. Salah, O. Rana, S. Tirunagari, R. Buyya, CoLocateMe: Aggregation-based, energy, performance and cost aware VM placement and consolidation in heterogeneous IaaS clouds, IEEE Trans. Serv. Comput. 16 (2) (2022) 1023–1038. [37] A. Verma, L. Pedrosa, M. Korupolu, D. Oppenheimer, E. Tune, J. Wilkes, Largescale cluster management at google with borg, in: Proceedings of the Tenth European Conference on Computer Systems, in: EuroSys ’15, Association for Computing Machinery, New York, NY, USA, 2015. [38] M. Tirmazi, A. Barker, N. Deng, M.E. Haque, Z.G. Qin, S. Hand, M. HarcholBalter, J. Wilkes, Borg: the next generation, in: Proceedings of the Fifteenth European Conference on Computer Systems, in: EuroSys ’20, Association for Computing Machinery, New York, NY, USA, 2020. [39] M. Zakarya, L. Gillam, H. Ali, I.U. Rahman, K. Salah, R. Khan, O. Rana, R. Buyya, epcAware: A game-based, energy, performance and cost-efficient resource management technique for multi-access edge computing, IEEE Trans. Serv. Comput. 15 (3) (2020) 1634–1648. [40] M. Zakarya, L. Gillam, Modelling resource heterogeneities in cloud simulations and quantifying their accuracy, Simul. Model. Pr. Theory 94 (2019) 43–65. Theodoros Theodoropoulos is a research and teaching assistant at the Faculty of Computer Science, University of Vienna. He received his Engineering Diploma from the School of Electrical and Computer Engineering at the National Technical University of Athens. He is currently pursuing a Ph.D. at the Department of Informatics and Telematics, Harokopio University of Athens. He has participated in several EUand nationally funded projects. His main research interests include efficient computing, deep learning, distributed systems, and software architecture. Yashwant Singh Patel is a Staff Scientist in the Department of Computing Science at Umeå University, Sweden. He received his Ph.D. in Computer Science and Engineering from the Indian Institute of Technology Patna, India. He has authored more than 40 peer-reviewed papers and articles in high-impact journals and internationally reputed conferences. His research interests include Cloud Edge Systems, Edge Intelligence, Serverless Computing, Energy Efficient Computing, and Distributed Algorithms. Uwe Zdun is a full professor of software architecture at the Faculty of Computer Science, University of Vienna. Before that, he worked as an assistant professor at the Vienna University of Technology and the Vienna University of Economics. He received his doctoral degree from the University of Essen in 2002. His research focuses on software design and architecture, distributed systems engineering (microservices, service-based, cloud, APIs, IoT, and blockchain-based systems), DevOps and continuous delivery, SW Engineering for ML, and ML for SW Engineering, software patterns, software modeling, model-driven development, and empirical software engineering. Uwe has published more than 300 articles in peer-reviewed journals, conferences, book chapters, and workshops. He is co-author of the books "Patterns for API Design - Simplifying Integration with Loosely Coupled Message Exchanges," "Remoting Patterns – Foundations of Enterprise, Internet, and Realtime Distributed Object Middleware," "Process-Driven SOA – Proven Patterns for Business-IT Alignment," and "SoftwareArchitektur." He has participated in 36 research and development projects and served as PI on 31. Uwe was editor-in-chief of the journal Transactions on Pattern Languages of Programming (TPLoP) published by Springer and is associate editor of the Journal of Systems and Software (JSS) published by Elsevier, the Computing journal published by Springer, and the IEEE Software magazine. Paul Townend is an Associate Professor at Umeå University, Sweden, where he is the founder of the Green Distributed Computing research group and a research leader within the Autonomous Distributed Systems Lab (ADSLab). He is PI for almost $6M in active research projects related to energyand carbonaware distributed systems, and has served as General Chair for 15 IEEE international conferences. Paul is Co-PI and Scientific Advisor of the WARA-Ops national data infrastructure (wara-ops.org) and is a member of the WASP Graduate School Management board – a six member committee responsible for the education of over 500 active PhD students within Sweden. Ioannis Korontanis is currently a Ph.D. candidate at the Department of Informatics and Telematics of Harokopio University of Athens. He received his BSc from the Department of Computer Science and Engineering of Technological Education Institute of Thessaly. He received his MSc ‘‘Web Technologies and Applications’’ at the Department of Informatics and Telematics of Harokopio University of Athens. The topic of his PhD thesis is ‘‘Computer Infrastructure Models and Optimal Resource Orchestration’’. His main research interests include distributed and real-time processing, cloud modelling languages & application models in Cloud & Edge, Cloud & Edge monitoring systems, CI/CD, Dev Ops and Big Data Analysis. He currently participates in the EU funded project named PANDORA (HORIZON Research and Innovation Actions). In the past has also participated in numerous EU funded projects. Antonios Makris is currently a Senior Researcher at the School of Electrical and Computer Engineering of the National Technical University of Athens. He received his BSc Degree in Computer Science in 2013 and MSc Degree in Web Engineering in 2015, both from Harokopio University of Athens. In 2022, he received his PhD in the area of Distributed Systems from the same department. His main research interests include Distributed Computing, Edge and Cloud Computing, Machine/Deep learning, AI Robustness, Efficient Computing, Big Data Management and Analysis, and Spatiotemporal and Trajectory Analysis. He has participated in numerous EU funded and national projects. Konstantinos Tserpes is an Assistant Professor at the School of Electrical and Computer Engineering, National Technical University of Athens (NTUA). Prior to 2024, he was a faculty member in the Department of Informatics and Telematics at Harokopio University of Athens. He holds a PhD in Distributed Systems from NTUA. His research focuses on distributed software systems, machine learning, and data analytics. In these fields, he has made significant contributions through numerous scientific publications, student mentorship, university teaching, and research projects. He has authored more than 100 papers, supervised dozens of undergraduate and postgraduate students’ theses, taught over 10 courses across two academic institutions, and served as the supervisor for seven PhD students, two of whom have successfully defended their dissertations. In addition to his academic work, he has played key roles in numerous EU and nationally funded projects, serving as both a scientific coordinator and a general coordinator in several initiatives. Future Generation Computer Systems 174 (2026) 107926 18