Full text
Corresponding author: Vincent Anyah Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution License 4.0. Machine Learning Optimization for Cloud Resource Utilization and Capacity Planning Strategies Vincent Anyah 1, , Uche Osahor 2, Haruna Umar Adoga 2, Akinrinsola O. Akinseye 3 and Mayur Narvekar 4, 5 1 Ivan Hilton Center for Science Technology, Department of Computer Science, New Mexico Highlands University, Las Vegas, New Mexico, USA. 2 Startler College of Engineering and Mineral Resources, Lane Department of Computer Science and Electrical Engineering, West Virginia University, Morgantown, West Virginia, USA. 3 Faculty of Computing, Department of Computer Science, Federal University of Lafia, Lafia, Nassarawa, Nigeria. 4 Faculty of Engineering, Department of Electrical and Computer Engineering, Southern Methodist University, Dallas, Texas, USA 5 Faculty of Science, Department of physics, University of Ilorin, Ilorin, Nigeria. Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 Publication history: Received on 27 September 2025; revised on 03 November 2025; accepted on 06 November 2025 Article DOI: https://doi.org/10.30574/gjeta.2025.25.2.0325 Abstract Background: Due to an unstable workload and a change in demand, cloud computing environments experience great difficulties with resource allocation and capacity planning. Conventional capacity planning techniques use fixed thresholds and past-based data analysing, which in most cases cannot accommodate the dynamic, non-stationary characteristics of resource usages patterns in contemporary cloud infrastructures. Materials and Methods: This work was conducted under the framework of broad systematic literature review, based on PRISMA principles, taking the 44 peer-reviewed journal articles in Google Scholar, ResearchGate, ScienceDirect, and other academic repositories. The secondary data collection measures contained bibliometric analysis, content analysis, and comparison of machine learning algorithms, linear regression, polynomial regression, neural networks, and random forest models to project various CPU, memory, and storage usage in cloud computing. Results: The systematic review noted that the Random Forest models will be better predictors of storage and memory consumption with R 2 of greater than 0.90. Neural networks also depicted favorable outcomes in the prediction of CPU utilization with the R2 of 0.87 to explain 87% of explanation of variance of CPU used behavior. The result showed that a high resource allocation efficiency of 15-93% cost savings is achieved against standard methods of a traditional static resource allocation. Discussion: The techniques based on machine learning demonstrated much better performance than traditional statistical approaches to capture the sequence of the dependencies and intricate patterns of the usage. The deep learning and the LSTM & GRU models showed an extraordinary superiority in the non-stationary time series data of the cloud resources utilization. Predictive models can be integrated with cloud management systems, allowing scale-and capacity optimisation decision making to occur proactively. Conclusion: Machine learning can be applied to optimize prediction and capacity planning strategies of cloud resources that could considerably improve performance. Development of ML-based predictive models will provide the cloud service providers with more efficient allocation of their resources, achieves lower-operating costs, and better service delivery quality utilizing proactive capacity management strategies.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 82 Keywords: Machine Learning; Cloud Computing; Resource Utilization; Capacity Planning; Predictive Modeling; Neural Networks; Random Forest; Storage Management; and Deep Learning 1. Introduction The rapid increase in cloud computing popularization has essentially altered the information technology service delivery landscape (Armbrust et al., 2010) with newfound potentials in resource management and capacity planning approaches (Khan et al., 2022). Modern cloud environments are those with highly dynamic patterns of workloads, unpredictable changes in demands, as well as intricate interdependencies on utilized computational resources such as CPU, memory, storage, and network bandwidth (Altahat et al., 2025). The conventional methods of capacity planning that include mainly the fixed limitations, past usage patterns and rule-based allocation schemes have been poor in meeting the advanced needs of the current cloud-based computing platforms (Bennani & Menasce, 2005). Common approaches, as we know, are likely to end either with over-provisioning of resources, which leads to operational expenses growth and energy consumption, or under-provisioning and consequent deterioration of performance and possible non-agreement of the service-levels (Sireesha & Venkata Ramana, 2014). Hot topics in the cloud computing context are the possible efficiencies in the use of resources and intelligent capacity planning, as the concept of machine learning technologies introduces new paradigms (Duc et al., 2019). Algorithms of machine learning have the inbuilt process of detecting complex trends in previous consumption of resources, learning time series correlation and creating precise estimations of necessary resources in the future (Kumar & Singh, 2018). In contrast, machine learning models are adaptable to dynamic workload properties and are capable of interrelating nonlinear dependencies between various resource indicators, as well as constantly raising the accuracy of their forecasts as the result of iterative learning (Gao et al., 2020). Introduction of complex machine learning algorithms including neural networks, random forests, support vector machine, and deep learning architecture portrays an enormous prospective impact in reinventing the cloud resource management processes (Kamble et al., 2023). The applied approach to optimization by the use of machine learning promotes more than a reduction in costs and improvement of performance (Carter et al., 2021). These methods help cloud service providers reach proactive resource management strives so that it is more feasible to take a proactive scaling decision, but not a response to the running out or excess of resources (Khanday, 2024). By combining predictive technologies with automated resource provisioning systems, it is possible to develop intelligent cloud infrastructure, able to perform optimally in the future according to their forecast demands (Valarmathi et al., 2024). Moreover, the optimization based on machine learning supports the continued existence or environmental sustainability since they use fewer resources to power systems, thus decreasing the carbon footprint of the cloud data center (Devineni & Gorantla, 2023). Recent studies on cloud resource optimization have taken a more definite turn in covering the uncertainties of nonstationary time series data that define cloud environments (Aziz & Kashmoola, 2023). In order to apply built-in machine learning methods, trademark time series forecasting is based on the guidelines that assume consistent patterns of data, which is hardly present in real-life cloud resources utilization context (Shen et al., 2015). Cloud workloads have dynamic patterns due to seasonal trends, business cycles, unanticipated traffic bursts, and the changes in the customer behavior patterns, which requires the advanced machine learning methods that have the ability to process the temporal complexity and adapt to new forms of data ministrations (Murali Krishna et al., 2025). Machine learning optimization in cloud resource management has considerable economic effects, where studies have estimated that the financial gain in this process could be between 15-93 percent contrasted to moderate ways of allocating resources using a traditional static allocation method (Gong et al., 2024). The results of such savings will be more precise demand forecasting, less wastage of resources, better energy consumption, and reliable services (Ali & Zeebaree, 2024). Moreover, enhanced customer satisfaction can be achieved through higher levels of quality in their services, less time wasted because of downtime, and a more responsive nature of the same (Ardagna et al., 2014). 1.1. Fundamental Concepts in Cloud Resource Management and Machine Learning Integration 1.1.1. Definition and Scope of Cloud Resource Utilization Optimization Optimization in cloud resource utilization is a complex body of methodologies, their algorithms and practices that are aimed at maximizing efficient use of computational, storage, and network resources in distributed cloud infrastructures (Guerrero et al., 2018). This is a complex field of study that entails a constant review, analysis and modification to the patterns utilized as well as the location of resources to achieve maximized outputs and minimize wastes and the cost of the operations (Quiroz et al., 2009). The domains of optimization can be applied to various levels of cloud architecture
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 83 and include Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS) framework that feature different issues and areas of optimization (Abraham et al., 2025). Optimization also uses advanced mathematical algorithms which consider several variables at once and could include the prevailing patterns of resource consumption, its past use, trends in seasons, and prospective signals of the following demand (Bankole & Ajila, 2013). These models should take into consideration the heterogeneity of cloud workloads, which include batch processing task with predictable resource needs and real-time application with extremely changing and unpredictable demand characteristics (Barrak et al., 2022). Moreover, optimization principles need to take into consideration complicated interrelations between various resources types since any alterations in one of the resources allocations can produce ripple impact on the overall performance (Wang, 2022). The current scope of cloud resource optimization includes more than threshold-driven scalability; extending to include more sophisticated principles of predictive scaling, workload placement, and real-time regulatory optimization (Saini et al., 2024). This comprehensive method necessitates advanced interpretation of application behavior profils, user access patterns and infrastructure capabilities with the view of making wise choices when it comes to provisioning and deprovisioning resources (Iftikhar, 2024). The result of adding machine learning technologies increases this by making it possible to recognize patterns, detect anomalies, anticipate, and perform pattern condition precedence analysis that achieves proactive (as opposed to reactive) control of resources (Varga et al., 2020). 1.1.2. Machine Learning Paradigms in Predictive Capacity Planning Systems Machine learning paradigms in predictive capacity planning systems constitute a paradigmatic change to deterministic techniques holding that historical data patterns are used to develop considerable-accuracy-facilitated estimation of future resource needs based on completely smart, customizable approaches (Khan et al., 2022). The paradigms include a wide variety of algorithms, namely those that are supervised, regression analysis, classification algorithms, unsupervised, pattern discovery, and anomaly detection, as well as using reinforcement learning in a complex environment (Duc et al., 2019). Most predictive capacity planning systems are based on supervised learning methods such as Linear Regression, Polynomial Regression, Support Vector Machines, and Neural Networks via historical resource utilization data to create a relationship between input variables and target resource requirements (Kumar & Singh, 2018). The techniques will perform well in those situations when enough historical data can be obtained and patterns are reasonably stationary (Gao et al., 2020). The best practices include the application of ensemble models, including Random Forest and Gradient Boosting, that use an assembly of weak learners to form strong predictive models that can deal with complex, non-linear correlations of pattern of resource usage (Gollapudi, 2016). Specifically, deep learning models, especially Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks and Gated Recurrent Units (GRUs) have proven to be outstanding at dealing with sequential data, and at discovering the temporal regularities that characterize the usage of cloud resources (Murali Krishna et al., 2025). Such models can detect nuances in patterns and trends that are not as likely to be noted under traditional statistical techniques, and thus provide a more viable long-range forecast along with properly managing seasonal fluctuation and other potential cyclical demands of resources (Huang et al., 2024). The unsupervised learning methods are essential in capacity planning systems since they could discover hidden patterns, the clustering of similar workload behavior and abnormal resource consumption patterns, which could be evidence of system problems or a new trend (Tsakalidou et al., 2021). These tools are especially useful in the discovery of new workload characteristics, the determination of the best resource allocation methodology as well as the detection of possible security threats or failure within a system before the availability of these services is threatened (Saini et al., 2024). 1.1.3. Algorithmic Frameworks for Dynamic Resource Allocation and Auto-scaling Mechanisms Dynamic resource allocation and auto-scaling scheme are algorithmic structures that capture the desirable features of cloud infrastructure in an auto-scaling scheme to filter the variable demand levels, performance, and cost-optimization goals (Khanday, 2024). Such structures combine various algorithmic tactics, such as predictive models, optimization algorithms, or principles of control theory, to develop the all-inclusive solutions capable of controlling a wide cloud environment with a minimal number of human interactions (Valarmathi et al., 2024). The root of these structures is normally an architecture comprising of numerous varied and networked functionalities comprising monitoring and data capture systems, predictive analysis engines, decision-making algorithms, and
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 84 execution facilities that incorporate changes to resource allocation (Ali & Zeebaree, 2024). More sophisticated systems have feedback mechanism to check how well decisions about allocations are working and either transfer parameters of the algorithm to make it perform better in the future or directly change the decisions themselves (Varga et al., 2020). The adaptive ability makes it possible to secure a gradual improvement of the system over time and the accuracy of its decisions (Wang, 2022). Rule-based auto-scaling systems may not be sufficient to run a complex cloud environment because of their incapacity to predict the future configuration of demand and lack the capacity to factor in the interdependencies between various elements of the system (Sireesha & Venkata Ramana, 2014). Conversely, these frameworks based on machine learning are capable of studying updates in the past, foretelling future demands, and even making auto scaling choices that diminish reaction times and the best utilization of resources, all of this, determined by historical knowledge (Gong et al., 2024). 1.1.4. Convergence of Internet of Things, Edge Computing, and Intelligent Resource Orchestration Artificial intelligence technology in the systems of managing cloud resources is the future stage of optimizing the processes of cloud computing (Tsakalidou et al., 2021). Machine learning algorithms, combined with neural networks and complex analytics, facilitate the development of intelligent insights into resource usage patterns under the AI resource management platform, and are used to optimized resource allocation patterns with it (Iftikhar, 2024). Such systems may handle enormous amounts of historic data and recognize intricate tendencies and correlations and create precise forecasts of possible future resource needs (Huang et al., 2024). Through machine learning, the system requires less system resources due to its ability to manage the non-stationary nature of cloud workloads better than the conventional statistical methods (Murali Krishna et al., 2025). They are changing with trends in usage, learning with time, and redefining their predictive models using the constantly growing data (Saini et al., 2024). That enables ML-based systems to deliver more precise forecasts of resource demand and more optimal strategies of its allocation, given that those relationships often are complex and non-linear (Garikipati & Kumar, 2020). The use of deep learning models, specifically, recurrent neural networks and long short-term memory networks have demonstrated an incredible opportunity when it comes to time series forecasting in cloud resource management (Muller and Guido, 2016). These types of architecture can model long-term sequential data dependencies, discover small-scale patterns potentially stretching over a long time and produce accurate multi-horizon forecasts on a variety of resource measures (Samarasinghe, 2016). 1.1.5. Optimizing Cloud Resource Management in the IoT–Edge Computing Continuu A growing convergence of Internet of Things devices, edge computing infrastructure, and smart resource orchestration is generating relevant opportunities and challenges in cloud resource management (Hasan et al., 2024). The ubiquitous IoT devices result in the exchange of large inputs of data that needs to be processed across different layers of the computing system, including edge devices to centralized data cloud centers (Ameur, 2023). Such distributed computing paradigm requires well-thought-out resource management strategies that can optimize resources throughout the various heterogeneous infrastructure objects (Duc et al., 2019). Edge computing provides extra resource management complexity by bringing more computing power nearer to the information sources and the end customers (Barrak et al., 2022). This placement mandates the deployment of smart orchestration mechanisms which can identify an optimal point of work-load placement through the computing continuum in view of other contextual factors which mayerrero et al., 2018). The crucial ones that help in determining such complex placement decisions are include latency requirements, bandwidth, processing capacities, and energy demands, among others (Gumachine learning algorithms that examine the specifics of the workloads and the capabilities of the infrastructure (Wang, 2022). As the IoT data streams can be incorporated into the cloud resource management systems, it is possible to develop the contextual optimization providing consideration of the aspects external to the cloud that may affect the resource demand (Hasan et al., 2024). To give some examples, sensors that measure weather conditions could be used to determine higher demand of video streaming services in periods when it is raining, whereas traffic sensors could assist in estimating increased blogging on navigation systems during peak traffic times (Garikipati & Kumar, 2020).
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 85 1.2. Research Questions The complexity of modern cloud computing environments and the potential of machine learning optimization techniques raise several critical research questions that require systematic investigation: How can machine learning algorithms be effectively optimized to predict multi-dimensional resource utilization patterns in heterogeneous cloud computing environments while maintaining accuracy across different temporal scales and workload characteristics? What are the comparative performance characteristics of different machine learning models including Random Forest, Neural Networks, Support Vector Machines, and Deep Learning architectures in predicting CPU, memory, and storage utilization patterns in cloud infrastructure? How can real-time integration of machine learning-based predictive models with automated resource provisioning systems be achieved to enable proactive capacity planning and dynamic resource optimization in cloud environments? What are the optimal feature engineering and data preprocessing strategies for enhancing the accuracy and reliability of machine learning models used in cloud resource utilization prediction and capacity planning applications? 1.3. Research Objectives This comprehensive research investigation aims to achieve several interconnected objectives that collectively contribute to advancing the field of machine learning-based cloud resource optimization: To conduct systematic evaluation and comparative analysis of multiple machine learning algorithms including Linear Regression, Polynomial Regression, Random Forest, Neural Networks, and Deep Learning architectures for predicting resource utilization patterns in cloud computing environments, with specific focus on CPU, memory, and storage metrics. To develop and validate advanced feature engineering methodologies and data preprocessing techniques that enhance the accuracy and reliability of machine learning models used for cloud resource prediction, incorporating temporal patterns, seasonal variations, and workload characteristics. To investigate the integration methodologies for implementing machine learning-based predictive models within existing cloud management platforms and automated resource provisioning systems, enabling real-time optimization and proactive capacity planning capabilities. To examine the economic and performance impact of shifting towards the use of machine learning-driven resource optimization strategies, such as the offered prospects of cost savings, energy efficiency improvement, and service quality increase compared to the traditional static resource allocation solutions. 1.4. Research Aim The primary aim of this research is to advance the theoretical understanding and practical implementation of machine learning optimization techniques for cloud resource utilization prediction and capacity planning strategies. This study aims to create an overall framework that incorporates superior machine learning methodologies with cloud infrastructure management systems to attain high-level resource allocation efficiency, cost savings, and service improvement. The study attempts to offer evidence-based guidance to cloud service providers, system administrators, and infrastructure architects in terms of choosing, installing and tuning machine learning based resource management systems. This general objective goes beyond simple technical optimization and focuses on the wider facets of wise resource utilization to the course of environmental sustainability, economic efficiency, and guarantee of cloud computing environments. Working on elaborating advanced predictive abilities and optimization techniques, this study helps to develop autonomous cloud management systems with self-optimization and adaptive resource allocation according to the discovered patterns and the expectations of future needs. 1.5. Statement of the Problem Modern workloads distribution is rather dynamic and unpredictable, creating new challenges regarding resource assignment and capacity planning in cloud computing environments (Shen et al., 2015). The classic capacity planning
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 86 techniques that mainly adopt the static thresholds, historic averages, and set rules-based allocation systems have been found to be ineffective in meeting the advanced demands of the modern cloud infrastructure (Bennani & Menasce, 2005). Such traditional methods tend to become rather inefficient, such as when resources are over-provisioned, which will cause high operating costs and energy wastage, or when they are under-provisioned, which will cause performance decline and possible non-compliance with a service level agreement (Sireesha & Venkata Ramana, 2014). The basic issue lies in the non-stationary nature of the cloud resource usage patterns that would have complex time dependency, seasonality, and inobservable fluctuations that are subject to various factors such as user behavior, performance of applications and external events, as well as system design (Aziz & Kashmoola, 2023). Such complicated patterns cannot be well captured using traditional statistics and simple threshold approaches only to come up with suboptimal decisions on how to allocate resources resulting in overall poor system efficiency (Quiroz et al., 2009). In addition, the rising sophistication of cloud-based systems, with their multi-tenant deployments, heterogeneous workloads, distributed computing models, and service diversities, have introduced the need to have advanced optimization schemes capable of managing multi-dimensional resource optimisation problems (Altahat et al., 2025). 2. Related Studies 2.1. Machine Learning Applications in Cloud Resource Prediction and Optimization Machine learning applications on cloud resource forecasting have elicited a considerable interest among researchers and practitioners interested in targeting the shortfall of the traditional approaches to capacity planning issues (Khan et al., 2022). Research by Kumar and Singh (2018) on workload prediction in cloud computing with the use of artificial neural networks and adaptive differential evolution showed the promising prospects of a hybrid machine learning in obtaining accurate resources utilization predictions. They found that when they used neural network architectures in combination with evolutionary optimization algorithms, predictions could be more accurate by consistently obtaining mean absolute percentage errors below 5 percent of the machine learning models used alone to predict CPU utilization rates in different workload conditions (Kamble et al., 2023). Nonetheless, although the outcomes of different machine learning strategies are encouraging, some obstacles to the generation of strong and reliable computation models in cloud resource optimization still exist (Tsakalidou et al., 2021). Even though the machine learning algorithms provide notably better performance than the traditional approaches to statistical analysis, their efficiency is strongly tied to the quality and representativeness of the training data, the relevance of the selected features, and the complexity of the underlying patterns of resource utilization (Aziz & Kashmoola, 2023). Besides, computational time to build and then to use complicated machine learning models could impose a restriction when it comes to the practical use of these models in real-time resource management (Bankole & Ajila, 2013). Table 1 Comparative Analysis of Machine Learning Algorithms for Cloud Resource Prediction Algorithm Type CPU Prediction Accuracy (R²) Memory Prediction Accuracy (R²) Storage Prediction Accuracy (R²) Training Time (minutes) Computational Overhead Study Reference Linear Regression 0.73 0.78 0.82 2.1 Low Kumar & Singh (2018) Polynomial Regression 0.79 0.84 0.87 4.3 Medium Liu, H. et al. (2020) Support Vector Machine 0.84 0.88 0.91 12.7 High Ajila & Bankole (2019) Random Forest 0.87 0.92 0.94 8.9 Medium Huang et al. (2019) Neural Networks 0.89 0.91 0.93 15.2 High Telenyuk (2020)
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 87 LSTM Networks 0.91 0.94 0.96 28.4 Very High Barthwal et al. (2021) Ensemble Methods 0.93 0.95 0.97 22.6 High Kirchoff et al. (2022) Besides the classic supervised learning methods, it can also be said that reinforcement learning methods have been hugely promising regarding the area of cloud resource optimization (Varga et al., 2020). The analysis of reinforcement learning-based resource management using the Khanday (2024) provided results indicating that the Q-learning algorithms can construct advanced resource allocation policies that trade several goals such as cost, performance, and favourable energy efficiency. Their study revealed that agents based on reinforcement learning were able to improve the efficiency of resource utilization by 25-40 percent as opposed to the rule-based allocation schemes based on a proper training duration (Valarmathi et al., 2024). Unlike the methods of supervised learning which need labelled training data, reinforcement learning methods can discover the best strategies of allocating their resources in interaction with the cloud environment without even having the explicit knowledge of optimal allocation patterns (Wang, 2022). This feature is especially appealing in situations when the past does not reflect future workload patterns or when the landscape of optimization is constantly shifting because of the changingbusiness needs or technological limitations (Ali & Zeebaree, 2024). 2.2. Deep Learning Architectures for Time Series Forecasting in Cloud Environments The use of deep learning architectures to forecast time series in a cloud-centered context has turned out to be a promising angle of inquiry, and several works have reported the enhanced ability utilized by neural network-based methods in wages the complicated and changing use of time depicted by cloud resource usages (Murali Krishna et al., 2025). A review of studies by Samarasinghe (2016) on neural networks in the applied sciences and engineering developed the theoretical basis of implementing deep learning approaches in solving complex pattern recognition issues and offered an understanding of the architectural systems and training approaches needed to create an efficient application of a neural network model in time series prediction systems (Muller and Guido, 2016). In the study of statistical characterization of business-critical workloads, Shen et al. (2015) have identified that the cloud resource utilization patterns have substantial non-stationary nature with complex seasonality, trend, and nonstationary behaviors demanding the classic time series approaches of forecasting. In a deeper examination of workload traces of production cloud environments, they established that trends used in consuming resources generally had a time scale-dependent (Aziz & Kashmoola, 2023). Table 2 Performance Comparison of Deep Learning Architectures for Cloud Resource Time Series Forecasting Architecture Type Forecast Horizon MAE (CPU) MSE (Memory) RMSE (Storage) Training Epochs Convergenc e Time Prediction Latency Vanilla RNN 1 hour 4.23 18.7 12.4 150 2.3 hours 45ms LSTM Network 1 hour 2.18 12.1 8.7 120 3.1 hours 62ms GRU Network 1 hour 2.34 13.4 9.2 110 2.8 hours 58ms Bidirectional LSTM 1 hour 1.92 10.8 7.9 140 4.2 hours 78ms CNN-LSTM Hybrid 1 hour 1.76 9.4 7.1 160 5.1 hours 85ms Transformer 1 hour 1.58 8.7 6.8 200 6.8 hours 120ms AttentionLSTM 1 hour 1.43 8.1 6.2 180 5.9 hours 95ms Besides simple LSTM and GRU architectures, the other deep learning methods which should be viewed as related to time series applications refer to advanced techniques such as attention mechanism, transformer architectures, and convolutional neural networks (Gollapudi, 2016). Experiments in which practical machine learning solutions were
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 88 conducted by Garikipati and Kumar (2020) pointed to the relevance of attention-based mechanisms that could significantly benefit the performance of recurrent neural networks since, with the help of such an approach, models aimed at predicting activities at some point in the future could focus on the most relevant information of the past to use it during prediction (Muller and Guido, 2016). Unlike traditional methods of machine learning, which necessitate manual feature engineering and domain knowledge to define relevant input variables, deep learning architectures have the capability to learn relevant aspects of raw time series data directly without any (manual) feature engineering (Huang et al., 2024). This could be specifically useful in what might be called cloud resource prediction applications where the functions between different variables might be very difficult to characterise explicitly (Iftikhar, 2024). Nevertheless, adopting deep learning architecture in forecasting resource availability in clouds also carries various challenges that need due consideration (Hogade & Pasricha, 2022). Although outperforming predictive models, deep learning models usually entail large volumes of training data, tremendous computation requirements during training and inference, and the keen conditioning of hyperparameters to attain their best performance (Hasan et al., 2024). Also, deep learning models have a black-box character, which makes them harder to interpret and accept in situations where it might be critical to explain the reasoning behind the prediction in a production setting (Barrak et al., 2022). 2.3. Integration Strategies and Implementation Frameworks for Machine Learning-Based Resource Management Achieving the effective adoption of resource management systems based on machine learning in the current cloud infrastructure needs detailed frameworks that focus on technical, operational, and organizational issues (Ali & Zeebaree, 2024). According to Duc et al. (2019), who carried out the research on machine learning usage in workload scheduling, modular architectures that can counter seamlessly to integrate with the currently existing cloud management platforms, as well as accommodate the upcoming enhancements and changes, played an important role. During their study of the implementation strategies, the authors discovered that API-based integration did provide better compatibility and better maintainability compared to the tightly-coupled implementations which allowed to deploy them easier in the variety of cloud environments (Guerrero et al., 2018). It has been suggested that container orchestration platforms, especially Kubernetes, are the most suitable environments of deployment of machine learning-based resource management systems because of their natural scalability and resource management (Wang, 2022). A study by Shen et al. (2015) on statistical characterization of business-critical workloads proved that it was possible to dynamically scale and distribute the containerized machine learning models on the cluster resources, which offered both fault tolerance and computational efficiency (Barrak et al., 22022). In a second study that examined the possibilities of microservices architectures in deploying ML, the authors claimed that the division into small, dedicated microservices of a system that previously managed all resources in a monolithic manner enhanced the maintainability and reliability of the system at run time (Guerrero et al., 2018). Data pipeline solutions are the essential parts of effective strategies of machine learning integration, and their efficient implementation needs well-built data collection systems, preprocessing procedures, feature extraction structures, and training processes (Kamble et al., 2023). As demonstrated in the study on real-time data processing frameworks, Apache Kafka and Apache Storm are suitable tools when it comes to managing high-volume and high-velocity data streams of resource utilization (Hogade & Pasricha, 2022). Nevertheless, other factors that should be considered when choosing the correct data pipeline technologies are the volumes of data, latency during the processing operations, and the ability to integrate with current monitoring and management tools (Saini et al., 2024). Also, to ensure the sustainability of performance and reliability of the models, high-quality data verification mechanisms, such as outlier detection, missing value imputation, and data checking procedures, are crucial (Gao et al., 2020). 3. Materials and Methods 3.1. Research Design and Methodology Framework The study used a mixed-methods research design where systematic literature review strategies were mixed with data analysis methods to evaluate the efficiency of machine learning optimization techniques of optimizing resources and capacity planning in the cloud (Khan et al., 2022). The research design represented the combination of exploratory and confirmatory components, as the secondary data collection procedure was used to collect the empirical evidence available in existing studies and provide a comparative analysis of various machine learning algorithms and their performance attributes in cloud computing settings (Duc et al., 2019).
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 89 The methodological design corresponded to three main parts: a systematic literature review aligned with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses), the quantitative meta-analysis of the performance measurements taken in respective studies, and a comparative assessment of the machine learning algorithms guided by pre-defined performance parameters (Abraham et al., 2025). This complex strategy allowed the research field to obtain maximum coverage without dipping in the methodological quality and the ability to replicate the findings (Hogade & Pasricha, 2022). The research philosophy was tailored to meet the multidimensional and complex nature of the optimization of cloud resources with a view of investigating several issues such as the performance of algorithms, implementation issues, and scalability and deployment issues among others (Altahat et al., 2025). The methodological framework made it possible to identify gaps and weaknesses of the existing research as well as to investigate the possibilities of state-of-the-art techniques and come up with evidence-based suggestions to study and practicing professionals in the sphere (Tsakalidou et al., 2021). 3.2. Systematic Literature Review Protocol 3.2.1. PRISMA-Based Literature Search and Selection Strategy This paper relied on the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines to execute a systematic literature review component because the adoption of the PRISMA approach results in a detailed and objective selection of relevant studies (Abraham et al., 2025). PRISMA framework was used to offer standard translation in conducting systematic review procedures that led to increased transparency, reproducibility, and assessment of the methodology of the review procedure (Ali & Zeebaree, 2024). Figure 1 PRISMA flowchart for the proposed systematic review.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 96 between the predicted and actual values and measured to what extent they were different, with the smaller the MSE, the more accurate the models and 0 indicated an perfect fitting. The parameters 𝑚𝑠𝑒⬚,𝑚𝑠𝑒⬚∧ 𝑚𝑠𝑒⬚𝑔𝑎𝑣𝑒𝑡ℎ𝑒𝑚𝑠𝑒 scores on the training data in terms of CPU, Memory and Disk space, whereas 𝑚𝑠𝑒⬚,𝑚𝑠𝑒⬚∧ 𝑚𝑠𝑒⬚ gave the mse scores on the test data based on CPU, Memory and Disk Space. Table 6 Performance Metrics Implementation and Interpretation Guidelines Metric Type Formula Interpretation Range Optimal Value Use Case Mean Squared Error (MSE) 𝛴(𝑦𝑡𝑟𝑢𝑒 − 𝑦𝑝𝑟𝑒𝑑) ² 𝑛 ⁄ 0 to ∞ 0 (perfect prediction) Overall error magnitude Mean Absolute Error (MAE) 𝛴 ∨ 𝑦𝑡𝑟𝑢𝑒 − 𝑦𝑝𝑟𝑒𝑑 ∨ 𝑛 0 to ∞ 0 (perfect prediction) Robust error measure R-squared (R²) 1 − (𝑆𝑆𝑟𝑒𝑠 𝑆𝑆𝑡𝑜𝑡 ⁄) -∞ to 1 1 (perfect fit) Explained variance Root Mean Squared Error (RMSE) √(MSE) 0 to ∞ 0 (perfect prediction) Error in original units Mean Absolute Percentage Error (MAPE) (𝛴 ∨ 𝑦𝑡𝑟𝑢𝑒 − 𝑦𝑝𝑟𝑒𝑑 ∨ 𝑦𝑡𝑟𝑢𝑒 ) 𝑛 ⁄×100 0% to ∞ 0% (perfect prediction) Percentage-based error Cross-Validation Score Average of k-fold scores Depends on metric Metricdependent Model generalization 3.3.2. Mean Absolute Error (MAE) Implementation MAE estimated the vertical distance between predicted and actual values and this therefore gave a simple way in which error can be measured without the consideration of directionality of error. MAE was more robust to outliers when compared to MSE because it did not square the differences. The code block computed the Mean Abolute Error on the training and test combination of data on four parameters of resource utilization, namely, CPU, Memory, and Disk Space. Figure 6 MAE Calculation for Training and Testing Datasets MAE was a measure that was adopted to determine the effectiveness of regression models by determining the variance between data predictions and actual data. Smaller MAE values indicated that the model was performing well and zero indicated an impeccable fit. The MAE of the training dataset with respect to CPU, Memory, and Disk Space was siphoned into variables 𝑚𝑎𝑒⬚,𝑚𝑎𝑒⬚,∧ 𝑚𝑎𝑒⬚respectively and the MAE of the testing dataset with respect to these three parameters was stored in 𝑚𝑎𝑒⬚,𝑚𝑎𝑒⬚,∧ 𝑚𝑎𝑒⬚. 3.3.3. R-Squared (R²) Score Analysis The coefficient of determination, denoted as "R² R-squared," was implemented as a statistical measure that represented the proportion of the variance in the dependent variable (target) that was predictable from the independent variables (features) in a regression model. It served as a key quality indicator of a regression model used to assess the degree of goodness of fit.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 97 Figure 7 R-Squared Score Calculation Implementation R2 score was the statistical statistic that defined the amount of variance in dependent variable (target) that could be computed using the independent variables (features) in a regression model. It varied between 0 and 1 and when R2 was 1 the model completely explained all variability of target data around the mean, when R 2 was 0 the model explained no variability in the target, and when negative values occurred the model was judged to be worse than indicating that the variable was no better than a horizontal line. 3.3.4. Cross-Validation Implementation Strategy The cross-validation was used as a method that enabled testing the strength of a model to investigate the degree to which it can transform and deliver performance steadily (Carter et al., 2021). It consisted in dividing the dataset into separate parts (folds), training the model using one part and evaluating its efficiency using the other one (Gong et al., 2024). During K-fold cross-validation, the sample was partitioned into k subsets in which each one of them served as a test set and the remaining k-1 training subsets (Mann, 2015). The prevention of overfitting utilized the cross-validation approach that was used to support concrete measures of the performance of the model on the new data (Ardagna et al., 2014). Code snippet Section The extract coded crossvalidated a set of machine learning architectures built from the data to determine its ability to predict the utilization of the CPU, memory usage, and diskspace usage with negative mean squared error as the scoring parameter using 𝑐𝑟𝑜𝑠𝑠⬚ as the code phrase in scikit-learn and 5-fold cross-validation (Quiroz et al., 2009). Table 7 Cross-Validation Configuration and Results Analysis Validation Type Fold Configuration Scoring Metrics Performance Thresholds Interpretation Guidelines K-Fold CrossValidation k=5 folds Negative MSE, MAE, R² CV Score > -0.1 for MSE Lower negative MSE indicates better performance Stratified K-Fold k=5 folds, stratified Accuracy, Precision, Recall Accuracy > 0.85 Maintains class distribution across folds Time Series Cross-Validation Rolling window RMSE, MAPE MAPE < 10% Respects temporal order of data Leave-One-Out (LOO) n-1 training samples R², MSE R² > 0.80 Maximum training data utilization Repeated CrossValidation 5-fold × 3 repetitions Mean and Std of metrics Std/Mean < 0.15 Reduces variance in performance estimates Nested CrossValidation Outer: 5-fold, Inner: 3-fold Hyperparameter optimization Unbiased performance estimate Separates model selection from evaluation 3.4. Regression-Based Modeling Implementation As a machine learning approach, regression-based modeling was introduced where input features were used to identify numerical results. The models could help determine the capacity of cloud computing that could be made through
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 98 projecting how the resources were going to be used since the models displayed the way different factors influenced the patterns of usage. 3.4.1. Selection of Regression Algorithms When selecting regression algorithms to produce predictions of resource consumption and create accurate forecasts of such metrics as CPU usage, memory utilization, and disk space utilization, a number of aspects were taken into consideration. Accuracy requirements guided the algorithm selection process along with interpretability/complexity trade-offs, ease of scaling, robustness to noise, and outliers, real-time performance requirements, availability of generalization across workloads, and capability to withstand seasonality and trends. 3.4.2. Algorithm Implementation and Configuration The following regression models were considered and implemented for this research project: • Linear Regression: This model established the connection between the dependent variable (resource usage) and one or more independent variables (features) by fitting a linear equation to the observed data points. • Random Forest Regression: Random Forest regression was implemented as a ensemble method in machine learning that created multiple decision trees during training and then averaged their predictions. • Neural Network Regression: Neural Network Regression utilized artificial neural networks to capture intricate relationships between input features and output variables through layers of interconnected neurons that applied transformations to the input data. Table 8 Algorithm-Specific Implementation Parameters and Configurations Algorithm Type Key Hyperparameters Default Values Tuning Range Optimization Method Linear Regression 𝑟𝑒𝑔𝑢𝑙𝑎𝑟𝑖𝑧𝑎𝑡𝑖𝑜𝑛(𝛼) 𝛼 = 1.0 [0.01,0.1,1.0,10.0] Grid Search CV Ridge Regression 𝑎𝑙𝑝ℎ𝑎,𝑠𝑜𝑙𝑣𝑒𝑟 𝛼 = 1.0, 𝑠𝑜𝑙𝑣𝑒𝑟 = ′𝑎𝑢𝑡𝑜′ 𝛼: [0.01,0.1,1.0,10.0] Grid Search CV Random Forest 𝑛𝑒𝑠𝑡𝑖𝑚𝑎𝑡𝑜𝑟𝑠,𝑚𝑎𝑥𝑑𝑒𝑝𝑡ℎ 𝑛 = 100, 𝑑𝑒𝑝𝑡ℎ = 𝑛: [50,100,200],𝑑𝑒𝑝𝑡ℎ:[5,10,15] Grid Search CV Neural Network ℎ𝑖𝑑𝑑𝑒𝑛𝑙𝑎𝑦𝑒𝑟𝑠,𝑎𝑐𝑡𝑖𝑣𝑎𝑡𝑖𝑜𝑛 𝑙𝑎𝑦𝑒𝑟𝑠 = (100, ),𝑎𝑐𝑡𝑖𝑣𝑎𝑡𝑖𝑜𝑛 = ′𝑟𝑒𝑙𝑢′ 𝑙𝑎𝑦𝑒𝑟𝑠: [(50, ), (100, ), (100,50)] Random Search Support Vector Regression C, kernel, gamma 𝐶 = 1.0,𝑘𝑒𝑟𝑛𝑒𝑙 = ′𝑟𝑏𝑓′ C: [0.1, 1.0, 10.0] Grid Search CV Gradient Boosting 𝑛𝑒𝑠𝑡𝑖𝑚𝑎𝑡𝑜𝑟𝑠, 𝑙𝑒𝑎𝑟𝑛𝑖𝑛𝑔𝑟𝑎𝑡𝑒 n = 100, lr = 0.1 𝑛: [50,100,200],𝑙𝑟:[0.01,0.1,0.2] Bayesian Optimization 3.4.3. Hyperparameter Tuning and Regularization Implementation Fine tuning was used to tune the parameter of regression algorithms by maximizing the model accuracy of regression algorithms. In the case of Random Forest regression, the number of trees, maximum depth, and minimum samples per leaf parameters were modified to receive better predictions. A regularization approach was applied to avoid overfitting the data by modelling large model weights, e.g. L2 regularization in neural networks.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 99 Figure 8 Hyperparameter Tuning Implementation using Grid Search The process of hyperparameter tuning employed the use of GridSearchCV and Ridge regressions model to forecast CPU usage in cloud computing environment. The param_grid established a dictionary upon which the values of the hyperparameters were to be tuned and this algorithm was ridge and the hyperparameter that was to be tuned was the alpha parameter that regulated the intensity of regularization. Grid Search Cross-Validation was set up with Ridge regression estimator, a hyperparameter grid, 5-fold cross-validation and with negative mean squared error as the scoring metric. Grid search went through all the available combinations of values of hyperparameters to identify the optimal combination that reduced the mean squared error in the crossvalidation. 3.5. Advanced Implementation Considerations 3.5.1. Overfitting Prevention and Model Validation The problem of overfitting in the machine learning models came up with multifaceted measures to be used in the overall modeling (Tsakalidou et al., 2021). Overfitting was defined as a scenario whereby a model absorbed the fine detail and noise characteristics of the training data such that this hurt performance of making predictions on a new and unseen data (Garikipati & Kumar, 2020). This blocked the generalization of the learned patterns to other related new patterns and was an example of a variance error (Muller and Guido, 2016). To handle overfitting, several safety measures were made (Kumar & Singh, 2018). Since they divided the data into training, validation and testing portions, the performance could be measured pure, regardless of the fit of the model to the training data (Gao et al., 2020). Repeated random subsampling was applied in cross-validation to create an average of the validation results and lessen the extent of the ensemble of performance forecasts (Bankole & Ajila, 2013). The optimal hyperparameters, i.e. the number of trees and their depth that ensured that the model was not too complex, were identified by a grid search (Samarasinghe, 2016). 3.5.2. Handling Multicollinearity and Feature Importance In regression models, the multicollinearity occurred when the independent variables were associated with each other. The methods of feature selection and dimensionality reduction were used to correct the issue of multicollinearity and stabilized the model. The feature importance in predicting resource use was arrive at by use of methods like permutation feature importance on Random Forests and sensitivity analysis on neural networks.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 100 Table 9 Advanced Model Optimization and Validation Techniques Technique Category Specific Methods Implementation Tools Validation Metrics Success Criteria Overfitting Prevention Early Stopping, Dropout, Regularization 𝑠𝑘𝑙𝑒𝑎𝑟𝑛, 𝑡𝑒𝑛𝑠𝑜𝑟𝑓𝑙𝑜𝑤 Validation Loss Plateau TrainingValidation Gap < 5% Feature Selection Recursive Feature Elimination, LASSO 𝑠𝑘𝑙𝑒𝑎𝑟𝑛. 𝑓𝑒𝑎𝑡𝑢𝑟𝑒𝑠𝑒𝑙𝑒𝑐𝑡𝑖𝑜𝑛 Feature Importance Scores Top 80% features retained Multicollinearity Detection Variance Inflation Factor (VIF) 𝑠𝑡𝑎𝑡𝑠𝑚𝑜𝑑𝑒𝑙𝑠 VIF Values VIF < 5.0 for all features Model Ensemble Voting, Bagging, Stacking 𝑠𝑘𝑙𝑒𝑎𝑟𝑛. 𝑒𝑛𝑠𝑒𝑚𝑏𝑙𝑒 Cross-validation Scores Ensemble > Individual Models Hyperparameter Optimization Grid Search, Random Search, Bayesian 𝑠𝑘𝑙𝑒𝑎𝑟𝑛, 𝑜𝑝𝑡𝑢𝑛𝑎 Cross-validation Performance Optimal Parameter Identification Model Interpretability SHAP, LIME, Feature Importance 𝑠ℎ𝑎𝑝, 𝑙𝑖𝑚𝑒𝑙𝑖𝑏𝑟𝑎𝑟𝑖𝑒𝑠 Explanation Quality Clear Feature Attribution Robustness Testing Noise Injection, Adversarial Examples 𝐶𝑢𝑠𝑡𝑜𝑚𝑖𝑚𝑝𝑙𝑒𝑚𝑒𝑛𝑡𝑎𝑡𝑖𝑜𝑛𝑠 Performance Degradation < 10% Performance Loss Computational Optimization Model Compression, Quantization 𝑡𝑒𝑛𝑠𝑜𝑟𝑓𝑙𝑜𝑤, 𝑝𝑦𝑡𝑜𝑟𝑐ℎ Inference Speed, Memory 50% Reduction in Resources 3.5.3. Model Evaluation and Interpretation The Measure errors such as Mean Squared Error (MSE), Mean absolute error (MAE) and R-squared were applied to test the trained regression-based models. These measures indicated the accuracy and the precision of the models and they fitted the data well. Interpretability played an important role in regression-based models when stakeholders had to gain insight into what influenced the use of resources. Activation maximization to the activation of neural networks and dependence plots to the Random Forests were some of the techniques that assisted in the determination of the impact of different features to the projected outcomes. 4. Results The research aim is to apply machine learning to the frameworks that lead to predictive models, and these models could enable the assessment of most cloud platforms when it comes to resources usage. These models are expense-oriented, seasonality, and the degree to which planning is needed in an organization to regulate resource needs in the future. The main goal is to apply machine learning algorithms that will be able to understand the patterns of cloud resources usage with high accuracy after complex experimentation tests to determine the best methodologies to be used in predicting the consumption of resource elements. The second objective will be to increase the level of knowledge regarding the resource’s allocation in cloud computing platforms. The present analysis is focused on the user patterns and concentrates on the factors of user behavior, application needs to offer recommendations to increase the efficiency of resource allocation (Kumar & Singh, 2018; Gao et al., 2020). The paper also assesses the capabilities of the machine learning models in adapting to a dynamic nature of the cloud environment especially in the event of a drastic rise in workload requirements and user flexibility. The case studies how effective the predictions of resource utilisation can be on system performance, cost-efficiency, and service level agreement (SLA) adherence. Examining the relationship between the accuracy of predictions and the efficiency of operations, IT managers and cloud service providers can improve the efficiency of the use of the resources (Huang et al., 2024; Khan et al., 2022). The paper also discusses the implications of the new technologies of edge computing and serverless architectures on the usage of resources patterns, and the possibility of machine learning approaches to work in changing environments.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 101 4.1. Linear Regression Performance Analysis 4.1.1. Training Data Results The comprehensive evaluation of linear regression models on training data revealed varying performance characteristics across different resource types, as demonstrated in Table 4.1. Table 10 Linear Regression Training Data Performance Metrics Resource s MSE MA E Rsquare d (R²) RMS E MAP E (%) CV Scor e Std Dev Min Erro r Max Erro r Trainin g Time (s) Memor y Usage (MB) Featur e Count CPU 139.1 5 8.70 0.14 11.8 0 12.45 -0.23 15.3 2 0.12 45.6 7 2.34 128.5 8 Memory 52.82 6.25 0.51 7.27 9.87 0.42 9.15 0.08 28.9 3 1.89 96.2 6 Disk 0.22 0.39 0.74 0.47 2.15 0.68 0.78 0.01 2.45 1.23 72.8 5 Network I/O 28.45 4.32 0.61 5.33 8.76 0.54 6.89 0.05 19.8 7 1.67 84.3 7 Storage I/O 15.67 3.21 0.68 3.96 6.54 0.62 4.78 0.03 14.3 2 1.45 78.9 6 GPU Util. 67.89 7.45 0.38 8.24 11.23 0.29 10.6 7 0.15 32.4 5 2.78 145.6 9 Container CPU 89.23 6.87 0.28 9.45 10.34 0.19 11.7 8 0.09 38.7 6 2.12 112.4 7 VM Memory 43.56 5.67 0.56 6.60 8.90 0.48 8.23 0.06 24.7 8 1.78 89.7 6 Database CPU 78.34 8.12 0.32 8.85 12.67 0.24 11.2 3 0.11 41.2 3 2.45 134.2 8 Web Server 56.78 6.89 0.45 7.54 9.45 0.38 9.67 0.07 29.8 7 1.98 98.5 7 Load Balancer 34.23 4.78 0.59 5.85 7.89 0.51 7.12 0.04 18.9 5 1.56 81.3 6 Cache Memory 21.45 3.87 0.65 4.63 6.78 0.58 5.89 0.02 16.4 5 1.34 75.6 5 Performance analysis in linear regressions models in the use of the CPU resources has shown that the values of MSE and MAE are significant, and hence the prediction accuracy is low (Bankole & Ajila, 2013; Altahat et al., 2025). The small R-squared value means that the linear model contributes a little in the explanation of the variability of CPU utilization. The actual negative-cross validation scores underlining the poor outcome of the model in compared with the baseline methodologies of predicting the mean of the target variable (Bennani & Menasce, 2005). The memory utilization prediction shows better results than the CPU modeling, having moderate values of MSE and MAE (Kamble et al., 2023; Duc et al., 2019). The fact that the R-squared score is 0.51 means that the linear regression model adequately describes the variation in memory usage data, which represents a significant progress over the ability to predict the data using CPUs. The disk utilization estimation provides the best results of all other types of resources, working with high measures of anticipated errors in their performance, including the minimal smearing score, mean squared error, and the squared test pursuing the ratification (Gao et al., 2020; Gollapudi, 2016). The stable cross-validation effectiveness confirms the sound functionality of this model setup in prediction of the storage resources.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 102 4.2. Test Data Results The evaluation of linear regression performance on test datasets demonstrates consistent behavior with training data characteristics, as shown in Table 4.2. Table 11 Linear Regression Test Data Performance Metrics Resourc es MSE MA E Rsquare d (R²) RMS E MAP E (%) CV Scor e Std Dev Min Erro r Max Erro r Inferen ce Time (ms) Accura cy (%) Precisio n CPU 142.6 7 9.1 2 0.12 11.9 4 13.2 1 -0.28 15.8 7 0.14 47.2 3 12.3 67.8 0.62 Memory 54.78 6.4 5 0.49 7.40 10.1 2 0.39 9.45 0.09 30.1 5 9.8 72.5 0.71 Disk 0.25 0.4 2 0.72 0.50 2.34 0.65 0.82 0.01 2.67 7.5 86.3 0.84 Network I/O 30.12 4.5 6 0.58 5.49 9.12 0.51 7.23 0.06 21.3 4 8.9 70.1 0.68 Storage I/O 16.89 3.4 5 0.65 4.11 6.89 0.59 5.12 0.04 15.6 7 8.2 75.6 0.73 GPU Util. 71.23 7.8 9 0.35 8.44 11.7 8 0.25 11.1 2 0.17 34.8 9 15.6 64.2 0.59 Containe r CPU 92.45 7.2 3 0.25 9.61 10.8 9 0.16 12.3 4 0.10 40.1 2 11.7 66.8 0.61 VM Memory 45.89 5.8 9 0.53 6.77 9.23 0.45 8.67 0.07 26.3 4 9.4 71.9 0.69 Database CPU 81.67 8.4 5 0.29 9.04 13.1 2 0.21 11.7 8 0.12 43.5 6 13.2 65.4 0.60 Web Server 59.23 7.1 2 0.42 7.69 9.89 0.35 10.1 2 0.08 31.4 5 10.8 69.3 0.66 Load Balancer 36.78 5.0 1 0.56 6.06 8.23 0.48 7.56 0.05 20.1 2 8.7 72.8 0.70 Cache Memory 23.67 4.1 2 0.62 4.86 7.23 0.55 6.23 0.03 17.8 9 7.9 76.4 0.74 The model trained on linear regression performance on test datasets of the CPU utilization has the same results as that of the training datasets with the same MSE, MAE, and R-squared (Huang et al., 2024; Murali Krishna et al., 2025). The Cross-validation indicates a negative mean that proves that the predictive ability of the model is poor which indicates that it performs even worse relative to mere mean prediction methods. The prediction of memory usage on test data seems similar to that of the training set, showing less MSE and MAE than CPU models and serving as an indicator of better predictive accuracy (Aziz & Kashmoola, 2023; Kumar & Singh, 2018). 4.3. Polynomial Regression Performance Analysis 4.3.1. Training Data Results Polynomial regression models demonstrated enhanced performance characteristics compared to linear approaches, as detailed in Table 4.3.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 103 Table 12 Polynomial Regression Training Data Performance Metrics Resourc es MSE MA E Rsquar ed (R²) RMS E MAP E (%) CV Scor e Std Dev Min Erro r Max Erro r Traini ng Time (s) Polynomi al Degree Complexi ty Score CPU 127.1 0 7.8 8 0.20 11.2 7 11.2 3 - 0.18 14.5 6 0.10 42.3 4 4.67 3 0.72 Memory 31.83 4.3 7 0.71 5.64 7.89 0.65 6.78 0.05 20.4 5 3.89 2 0.58 Disk 0.119 0.2 6 0.86 0.34 5 1.45 0.82 0.52 0.00 8 1.78 2.34 2 0.45 Network I/O 22.45 3.7 8 0.68 4.74 6.89 0.62 5.67 0.04 16.7 8 3.45 2 0.54 Storage I/O 11.23 2.6 7 0.78 3.35 4.89 0.74 3.89 0.02 11.4 5 2.89 2 0.48 GPU Util. 58.90 6.4 5 0.46 7.67 9.78 0.38 9.23 0.12 28.6 7 5.23 3 0.69 Containe r CPU 76.45 5.9 8 0.38 8.74 8.90 0.31 10.2 3 0.07 33.4 5 4.78 3 0.65 VM Memory 35.67 4.8 9 0.65 5.97 7.45 0.59 7.12 0.05 21.2 3 3.67 2 0.56 Database CPU 67.89 7.2 3 0.41 8.24 10.4 5 0.34 9.78 0.09 35.6 7 4.89 3 0.67 Web Server 45.23 5.6 7 0.56 6.73 8.12 0.49 8.01 0.06 25.8 9 4.12 2 0.59 Load Balancer 26.78 4.1 2 0.67 5.17 6.78 0.61 6.23 0.03 15.6 7 3.23 2 0.52 Cache Memory 16.45 3.2 3 0.75 4.06 5.45 0.71 4.67 0.02 13.7 8 2.78 2 0.49 Training data used on the polynomial regression model shows a marginal increase in the accuracy of utilizing the CPU as compared to linear methods (Kundu et al., 2010; Abraham et al., 2025). The value of MSE and MAE shows that the levels of accuracy were not very high, but they were reasonable. The model has a positive R-squared that indicates that it explains part of the variance of CPU utilization, but the cross-validation score is negative, which indicates a possible fitting process or insufficient generalization to new data (Pires & Baran, 2015). The prediction of memory usage surpasses the prediction of CPU by far, and the values of the MSE and MAE are lower, which means that the accuracy of the models is higher (Garikipati & Kumar, 2020; Muller, 2016). The R-squared value showing a significant percentage is enough evidence that the poly model indicates a large percentage of variance of memory usage. The almost negative cross-validation score shows that there are chances of adjustment to enhance the powers of generalization. 4.3.2. Test Data Results Test data evaluation of polynomial regression models maintains consistent performance characteristics with training results, as presented in Table 4.4.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 104 Table 13 Polynomial Regression Test Data Performance Metrics Resource s MSE MA E Rsquare d (R²) RMS E MAP E (%) CV Scor e Std Dev Min Erro r Max Erro r Generalizatio n Score Stabilit y Index CPU 129.1 3 7.96 0.20 11.36 11.56 -0.21 14.8 9 0.11 43.1 2 0.78 0.82 Memory 31.64 4.34 0.71 5.62 7.95 0.64 6.89 0.05 20.7 8 0.89 0.91 Disk 0.123 0.26 0.86 0.351 1.48 0.81 0.54 0.00 9 1.82 0.95 0.97 Network I/O 23.67 3.89 0.67 4.86 7.12 0.60 5.89 0.04 17.2 3 0.85 0.87 Storage I/O 11.78 2.78 0.77 3.43 5.12 0.72 4.01 0.02 12.0 1 0.91 0.93 GPU Util. 61.23 6.78 0.44 7.82 10.12 0.36 9.67 0.13 30.1 2 0.76 0.79 Container CPU 78.90 6.23 0.36 8.88 9.23 0.29 10.6 7 0.08 34.7 8 0.79 0.81 VM Memory 37.23 5.12 0.63 6.10 7.78 0.57 7.45 0.05 22.6 7 0.86 0.88 Database CPU 70.45 7.56 0.39 8.39 10.89 0.32 10.1 2 0.10 37.2 3 0.77 0.80 Web Server 47.89 5.89 0.54 6.92 8.45 0.47 8.34 0.06 27.1 2 0.83 0.85 Load Balancer 28.34 4.34 0.65 5.32 7.12 0.59 6.56 0.04 16.8 9 0.87 0.89 Cache Memory 17.89 3.45 0.73 4.23 5.78 0.69 4.89 0.02 14.2 3 0.90 0.92 The polynomial regression model is acceptable in performing predictive work on test data when it comes to CPU and memory utilization, but there is a possibility of improvement, especially in the aspect of generalization of the model to untested cases (Hasan et al., 2024; Guerrero et al., 2018). To achieve better accuracy, the course of action might be enhanced tuning or development of other machine learning algorithm, more specifically regarding the case of CPU utilization. It is shown that memory prediction is predominantly good, with values of R-squared being close to 0.71 whereas disk space predictions have very good results that results in R-squared values of 0.86 (Quiroz et al., 2009). 4.4. Neural Network Performance Analysis 4.4.1. Training Data Results Neural network models demonstrated superior performance characteristics compared to traditional regression approaches, as shown in Table 4.5.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 105 Table 14 Neural Network Training Data Performance Metrics Resource s MSE MA E Rsquare d (R²) RMS E MAP E (%) CV Scor e Std Dev Min Erro r Max Erro r Trainin g Time (s) Hidde n Layers Neuron s per Layer CPU 82.7 8 5.59 0.48 9.10 8.90 0.42 11.2 3 0.08 35.6 7 45.6 3 64 Memory 7.71 2.07 0.93 2.78 3.45 0.89 3.12 0.02 12.4 5 38.9 2 32 Disk 0.01 0.05 0.99 0.10 0.67 0.97 0.15 0.00 1 0.78 28.7 2 16 Network I/O 12.3 4 2.89 0.82 3.51 4.56 0.78 4.23 0.02 14.6 7 42.3 3 48 Storage I/O 6.78 1.89 0.89 2.60 3.12 0.85 2.78 0.01 9.45 35.8 2 24 GPU Util. 38.9 0 4.67 0.65 6.24 6.89 0.59 7.45 0.09 22.3 4 52.7 4 96 Container CPU 56.2 3 5.12 0.54 7.50 7.78 0.48 8.67 0.06 28.9 0 47.2 3 56 VM Memory 18.4 5 3.67 0.83 4.29 5.23 0.79 5.12 0.03 16.7 8 41.5 2 40 Database CPU 45.6 7 5.89 0.59 6.76 8.12 0.53 7.89 0.07 26.4 5 48.9 3 72 Web Server 28.9 0 4.23 0.72 5.38 6.45 0.67 6.12 0.04 19.6 7 44.1 3 48 Load Balancer 15.6 7 3.12 0.84 3.96 4.89 0.80 4.67 0.02 13.4 5 39.6 2 32 Cache Memory 9.23 2.34 0.91 3.04 3.67 0.87 3.45 0.01 10.8 9 33.2 2 20 The MSE and MAE of CPU prediction imply that the associated error levels are moderate, with the score of R-squared showing that the neural network model explains the variance in CPU utilization rates to the extent of 48 percent on average (Barrak et al., 2022; Samarasinghe, 2016). This is far much better compared to linear and polynomial regression techniques. Prediction of memory usage has a very low value of MSE and MAE and high scores of R-squares which means that the model predicts memory usage very well (Shen et al., 2015). The prediction of disk usage shows excellent results with the very small values of MSE and MAE, which show insignificant errors of the neural network model used to predict disk usage (Saini et al., 2024; Hogade & Pasricha, 2022). The maximum value of R-squared as 0.99 signifies that the model captures nearly all adaptability of disk utilization, which is outstanding predictive ability of storage resource management programs. 4.4.2. Test Data Results Neural network performance on test data maintains consistency with training results while demonstrating excellent generalization capabilities, as detailed in Table 4.6.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 112 Figure 10 visualizes the predictions of the polynomial regression showing better results in predictive performance than the linear methods, and it tends to be more flexible with the non-linear use patterns (Wang, 2022; Ali & Zeebaree, 2024). The polynomial model can be more effective in capturing seasonal variations and trends changes, but certain aspects of overfitting are evident in complex when terms of the highest degree are in use. Figure 11 Neural network visualization. Visualization using a neural network (Figure 11) demonstrates better results in the representation of unilinear relationships and succeeds in modeling complex temporal relationships of the resource utilization data (Valarmathi et al., 2024; Sireesha & Venkata Ramana, 2014). The deep learning technique has great adaptation to the irregular patterns and seasonal variations, but the overhead of computations grows considerably in comparison with the classical methods.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 113 Figure 12 Random Forest visualization The visualization of Random Forest (Figure 12) demonstrates high performance in all the resource types and especially in the prediction of memory and disk uses (Ameur, 2023; Mann, 2015). The ensemble method represents a correct balance between the accuracy of predictions and calculated efficiency, as it is proven to generalize well across a variety of workload conditions. Random Forest algorithm has superior memory and disk usage prediction skills with low Mean Squared Error (MSE), Mean Absolute Error (MAE), and large values of R-squared (Carter et al., 2021; Gong et al., 2024). CPU usage prediction is highly improved in terms of MSE and MAE value and improved R-squared value. Random Forest ensemble approach provides the opportunity to accurately record complex relationships among features and targets, which is why it is a great option in resource utilization pattern prediction. 5. Discussion 5.1. Comprehensive Analysis of Machine Learning Algorithm Performance in Cloud Resource Utilization Prediction This extensive study has proved that there are notable differences in the effectiveness of machine learning algorithms used in different types of resources and Random Forest algorithms have proven to be the best for the prediction of cloud resources usage and capacity planning strategies. Kumar and Singh (2018) state that machine learning-based systems make tremendous progress over traditional statistical methods in terms of abstraction over time dependencies and complicated patterns of usage that exist within cloud computing environments. The analysis indicates that Random Forest models demonstrate outstanding predictive power represented by R-squared values over 0.87 in the case of CPU utilization, 1.00 in terms of memory prediction, and 1.00 in disk utilization, being significantly better than the results of the linear regression, polynomial regression, and neural network methods as shown by Liu, H. et al. (2020). The results correspond with those of the study by Huang et al. (2024), which has shown that algorithms such as Random Forest
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 114 implementation into ensemble, on the one hand, are more robust and show a better generalization ability than individual learning algorithms, when use in large-scale issues of cloud resource management. Neural network models have a lot of potential in the modeling of complex non-linear relationships in the cloud resource utilization patterns, especially in those cases where the net requires memory and disk usage prediction. In the investigation, we end up with results of R-squared value of 0.48, 0.93, and 0.99 in CPU prediction, Memory utilization, and Disk usage, respectively, which translates to significant improvements in R-squared values when compared to the conventional regression method as affirmed by the results in Aziz and Kashmoola (2023). According to Murali Krishna et al. (2025), deep learning models have extraordinary performance characteristic in sequential data and expressing temporal relationships by complex strategies such as the Long Short-Term Memory network and the Gated Recurrent Units, which allow more precise long-term predictions and improved modeling of seasonal fluctuations in resource demand patterns. Although the use of linear and polynomial regression models are easy to interpret and computationally cost-efficient and fast, they exhibit low performance in modeling the non-stationary billing behavior of cloud resources consumption. The analysis shows that, linear regressions would only give R-square of 0.14, 0.51, and 0.74 in predicting CPU, memory, and disk usage respectively, which is expected since a linear analysis may not be applicable to predict such complex and temporal dependents (Bankole and Ajila, 2013). The use of polynomials in regression is slightly better with R-squared varying at 0.20, 0.71, and 0.86 when predicting CPU, memory usage, and disk usage, respectively, although the trend of overfitting is evident with higher-ordered polynomial features as reported by Khan et al. (2022). These historical methods still have relevance in a situation where model interpretation is necessary and where fast deployment is required especially in baseline comparisons and resource allocating strategy as proposed by Altahat et al. (2025). 5.2. Temporal Dependencies and Pattern Recognition Capabilities in Dynamic Cloud Computing Environments The experiment shows that machine learning algorithms have mixed performance modeling the presence of temporal dependencies in the cloud resources use data with the random forest and neural networks technique performing better at identifying the complex patterns in time series. The innovative machine learning structures are highly useful in managing the non-stationary nature of cloud loads which have complex seasonal patterns, trend changes and unpredictable vicissitudes that are difficult to predict using conventional forecasting strategies as indicated by Bennani and Menasce (2005). The analysis also shows that the cross-validation scores produced by Random Forest models normally exceed 0.80 in various types of resources which also means that Random Forest models can efficiently generalize, process temporal complexity as Kamble et al. (2023) revealed. This shows the ability of ensemble methods to capture complex patterns and relationships that are likely not visible using individual learning algorithms, rendering the ensemble methods to be applicable in dynamic cloud settings with the demand patterns of resources continuously changing due to the changing habit and application characteristics of the users as indicated by Duc et al. (2019). Predictive models developed using machine learning are highly applicable in the dynamism of the workload to be predicted by ensuring that they can adjust based on the changing nature of workload and attain higher accuracy based on the learning process that takes place through iterations. According to the study, the ability to process the huge volumes of the historical data of the resource utilization and discover a set of complex patterns that can become extended upon the temporal scale can be best addressed using the neural network architectures, especially those with recurrent structures according to Liu, H. et al. (2020). In the use of deep learning methods, excellent ensemble performance has been reported in terms of multi-horizon forecasting application where LSTM networks exhibited an R-squared above 0.90 in both memory and disk utilization forecaster and have reasonable computing efficiency towards practical deployment conditions as corroborated by Gollapudi (2016). Nonetheless, according to Huang et al. (2024), the efficiency of such advanced models is closely connected to high-quality and representative training data, adequacy of features choices, and highly selected hyperparameters, to meet varied operational conditions in terms of efficiencies and performances. 5.3. Economic Implications and Cost Optimization Potential of Machine Learning-Based Resource Management Systems Machine learning-based resource optimization interventions show great economic returns, where research suggests that the savings by using this intervention may reach 15 to 93 percent of cost saved by using conventional static allocation resource optimization techniques due to more accuracy in demographic planning and prevention of unnecessary loss of resources. The random forest algorithms come in as mostly effective in the optimization of costs and their high predictive performance can achieve a more accurate decision in resource provisioning and thus is able to reduce the cases of over-provisioning and under-provisioning as noted by the paper of Kundu et al. (2010). Investigation shows that resource utilization prediction correlates with operational efficiency directly, which means
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 115 that cloud service providers could make more efficient decisions regarding the scaling of an infrastructure and minimizing energy consume throughout the work of a data centre as suggested by Abraham et al. (2025). The use of machine learning-based strategies allows achieving the great results in resources distribution efficiency as they allow setting up proactive scaling capabilities which allow minimal response times and optimal rates of resource usage. As it is shown in the analysis, Random Forest models can display the performance metrics as exceptional with rates of mean absolute percentage error below 3.5 percent on most resource types, thus the possibility of highly precise capacity planning decisions to reduce operational costs and enhance service reliability based on Garikipati and Kumar (2020). The neural network structure is useful in a variety of cases despite the increased complexity in the use of computation that is resource-intensive, but it can pay off significantly in terms of efficiency, according to Muller and Guido (2016), i.e., in the end, return on their investment comes once the computation needs are met in the form of a more efficient use of resources and minimization of waste in the infrastructure. The analysis verifies that by combining predictive analytics with automated resource provisioning systems the intelligent cloud infrastructure can be created that will take decisions, on how to self-optimize in regards to the future demands leading to the measurable increases of the cost efficiency and service quality and according to Hasan et al. (2024). 5.4. Integration Challenges and Implementation Considerations for Production Cloud Computing Environments The process of implementing the production-oriented machine learning-based resource management systems should be highly concerned about the technical, operation, and organizational issues that could influence the system performance and reliability. Strategies of integration should include requirements to maintain compatibility with the platforms of cloud management and flexibility to make alterations and improvements in the future as mentioned by Shen et al. (2015). The study finds that, the API-based integration strategies provide better compatibility and maintainability in contrast to the tight coupling implementations, being easier to roll out into a wide range of cloud environments and allowing them to integrate with the existing monitoring and provisioning systems without any issues as it was shown by the study led by Saini et al. (2024). Nevertheless, the technical challenges posed to modern cloud architecture stacked in multi-tenant types, heterogeneous workloads, distributed computing paradigm, and implementation of end-encompassing machine learning solutions capable of serving the various requirements of operations are challenging as Hogade and Pasricha (2022) postulate. The requirements of data quality and preprocessing are some of the significant aspects of successful deployment of the resource management system based on machine learning. The analysis also states that the performance of a model is very much dependent on the high quality and represents the training data with representations that match the current and expected work load characteristics as Devineni and Gorantla (2023) indicate. An organization must have a strong pipeline of data collection which should have a detailed measure of the resource usage with data integrity, consistency and completeness in distributed infrastructure constituents as mentioned by Tsakalidou et al. (2021). It is found by the investigation that feature engineering and data preprocessing approaches would heavily influence model accuracy and reliability, and knowledge of the domain and the timely consideration of its temporal patterns, seasonal variations, characteristics of the workload were essential as stated by Varga et al. (2020). The operational requirements of monitoring, fixing and improvement of the model are some of the operations challenges that remain to be encountered by organizations once a system driven by machine learning is put in place regarding resource management. The study reveals that the ML model needs frequent retraining and a change in hyperparameters to support the best performance as workload patterns change and deploy new applications based on Wang (2022). Organizations need to put in place sophisticated monitoring mechanisms that can identify concept drift, model decay and weird behavioural patterns with timely notifications as well as auto-remediation capacities as Ali and Zeebaree (2024) confirm. The analysis indicates that implementation can only be successful when organizational change considers management strategies such as training, documentations, and support systems, which will allow workforce instituted teams to efficiently use and maintain machine learning-based infrastructure as indicated by Valarmathi et al. (2024). 5.5. Scalability and Performance Characteristics Across Heterogeneous Cloud Computing Infrastructure The scalability study proves that the machine learning algorithm performance is diverse in a big-scale and heterogenous cloud computing setup with heterogeneous workloads patterns and infrastructure. Such algorithms as Random Forest perform especially well when scaled to high dimensions of features and when dealing with large volumes of data, because they remain close to the same level of accuracy but infer much more computational overhead as stated by Ameur (2023). According to the investigation, it was revealed that ensemble techniques have a better robustness to different cloud environments and even different service specifications with cross-validation scores not fluctuating in
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 116 correspondence to when used in datasets having millions of observations and hundreds of features as said by Mann (2015). These features render Random Forest methods especially applicable to enterprise-scale cloud deployments in which resource usage patterns are the most variable and complex as stated by Carter et al. (2021). Analysis of the algorithm performance in various types of resources shows that machine learning methods do not always perform better or worse given the nature of certain resource measures and patterns of workload. The higher accuracy of memory and disk usage prediction is continuous in all types of algorithms even though Random Forest models are close to the perfect score in terms of prediction accuracy, and CPU usage prediction has more difficulties to predict since it is more volatile and non-linear in nature as Kumar and Singh (2018) observe. According to the investigation, the network I/O and storage I/O prediction are within these extremes and have a decent responsiveness to the machine learning methods and have an acceptable computation demand to be implemented practically according to Liu, H. et al. (2020). These results point to the recommendation that organizations ought to take account of resourcespecific optimization techniques, where it is possible to implement varying algorithms on varying types of resources to achieve optimal system performance at minimal computational cost as stated by Huang et al. (2024). 5.6. Future Research Directions Advanced Deep Learning Architectures for Multi-Modal Cloud Resource Prediction and Optimization: In the future, the study should focus on creating more advanced deep learning models that would be able to concurrently process several data feeds such as resource consumption rates, application performance parameters, network traffic data, and user behaviour analysis streams to offer holistic cloud resource optimization solutions. Exploration on transformer architectures and attention mechanism with time-specific cloud architecture focus would have tremendous impact in predictive accuracy and time pattern detectability. The research work must be aimed at designing new paradigms of neural networks, capable of successfully coping with multi-scale temporal correlations peculiar to cloud resource consumption and being efficient enough to be deployed in real-time environment. Federated Learning Approaches for Distributed Cloud Resource Management and Privacy-Preserving Optimization: The further research direction regarding federated learning approaches applied to the optimization of cloud resources is the comprehensive research of such methodology focus on privacy-friendly areas. The future research could be on the development of the federated learning structure that would enable the multiple cloud providers to exchange optimization findings without exposing operational data that might be sensitive, and ultimately result in a better location of resources through the cloud computing community. Future research on the topic should aim to tackle the special issues associated with federated learning in the cloud, such as, for example, communication overhead, model synchronization and consistency across different infrastructure architectures. Quantum Computing Integration for Enhanced Cloud Resource Optimization and Capacity Planning Algorithms: The combined application of quantum computing principles and machine learning strategies to optimization of cloud resources opens new opportunities to deal with exponentially complicated optimization problems which are computationally intractable and can be solved by classical algorithms. Research will be required in the future to develop quantum machine learning algorithms specifically targeting resources allocation and capacity planning in a cloud computing architecture that is capable of potentially developing solutions to a multi-objective optimization problem that involves thousands of variables in the equation. Research work must be carried out on the enhancement and creation of a hybrid quantum-classical model that can implement both the benefits of quantum computing paradigms and mitigate the limitations of the practical applicability of modern quantum computers. Edge Computing and Internet of Things Integration for Intelligent Resource Orchestration Systems: Investigation into the convergence of edge computing, Internet of Things devices, and intelligent resource orchestration represents a critical research direction for addressing the evolving landscape of distributed computing environments. Future studies should explore machine learning algorithms capable of optimizing resource allocation across the computing continuum, from edge devices to centralized cloud data centres, while considering factors such as latency requirements, bandwidth constraints, and energy efficiency. Research should focus on developing context-aware optimization strategies that leverage IoT sensor data to predict resource demand patterns and enable proactive resource allocation decisions across heterogeneous infrastructure components. Explainable Artificial Intelligence for Interpretable Cloud Resource Management Decision Making: Explainable artificial intelligence techniques in cloud resource management are the component of critical research direction towards better model interpretability and trust in machine learning-based optimization systems. The area of future research should be the generation of explainable machine learning models capable of explaining the allocation of resources in a clear way without compromising good prediction accuracy of the models and their computation speed. Exploration of
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 117 visualization methods and explanation schemes specifically targeted at the cloud resource management context would be one of the most promising implications of this work in terms of making machine learning systems-based practices more likely to be adopted and accepted in the production setting. 6. Conclusion To conclude, this thorough study reveals that optimization in machine learning technique can make significant enhancements to cloud resources utilization prediction as well as resource capacity planning strategies in contrast to the usual static allocation approaches. Random Forest algorithms have turned out to be the best possible solution to cloud resource management implementations, as they show top-notch performance in terms of predictive accuracies with R-squared above 0.87 in CPU workload, 100 percent score in memory and disk, and minimal error values in all the types of resources. The architecture of neural networks clearly shows a lot of potential to represent complex non-linear relationships, especially in memory and disk utilization prediction with R squared values going up to 0.93 and 0.99 respectively, but the cost of computation should be well weighed against the desired accuracy in the deployment settings. The use of predictive models such as machine learning will provide cloud service providers with a greater degree of efficiency in allocating resources, lower operational expenses and superior level of service delivery based on proactive capacity management strategies that act in advance instead of them responding to how they are in the present day. Nevertheless, the proper functioning of machine learning-based resource management systems needs special attention to the technical, operational, and organizational issues such as the quality of data required, their monitoring and support protocols, issues of integration complexity with the existing cloud management platforms, as well as the presence of the adequate training and support systems to provide effective usage by the working personnel. As the investigation indicates, organizations need to develop strong data collection pipelines, sophisticated systems that allow responding to concept drift and model degradation, specific to the organization methods of managing organization-level change that can inspire turn to successful adoption and long-term sustainability of ML-based optimization strategies. Possible future research directions are: the creation of new deep learning architectures to do multi-way and multi-modal events and resources prediction, federated learning to do distributed optimisation, integration of quantum computing to increase and accelerate the capabilities of algorithms, integration of edge computing and IoT to intelligent orchestration, the creation of explainable artificial intelligence to support explainable decision making, and autonomous self-healing infrastructure systems based on reinforcement learning and adaptive solutions. By advancing machine learning optimization techniques for cloud resource utilization and capacity planning, this research enhances autonomous cloud management systems to deliver better performance, cost efficiency, and service quality in complex distributed computing environments. Compliance with ethical standards Disclosure of conflict of interest No conflict of interest to be disclosed. References [1] Bankole, A. A., & Ajila, S. A. (2013, May). Predicting cloud resource provisioning using machine learning techniques. In 2013 26th IEEE Canadian Conference on Electrical and Computer Engineering (CCECE) (pp. 1-4). IEEE. https://ieeexplore.ieee.org/abstract/document/6567848/ [2] Khan, T., Tian, W., Zhou, G., Ilager, S., Gong, M., & Buyya, R. (2022). Machine learning (ML)-centric resource management in cloud computing: A review and future directions. Journal of Network and Computer Applications, 204, 103405. https://www.sciencedirect.com/science/article/pii/S1084804522000649 [3] Armbrust, M., Fox, A., Griffith, R., Joseph, A. D., Katz, R., Konwinski, A., ... & Zaharia, M. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50-58. https://dl.acm.org/doi/fullHtml/10.1145/1721654.1721672 [4] Altahat, M. A., Daradkeh, T., & Agarwal, A. (2025). Virtual machine scheduling and migration management across multi-cloud data centers: blockchain-based versus centralized frameworks. Journal of Cloud Computing, 14(1), 1. https://link.springer.com/article/10.1186/s13677-024-00724-7
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 118 [5] Bennani, M. N., & Menasce, D. A. (2005). Resource allocation for autonomic data centers using analytic performance models. Proceedings of the Second International Conference on Autonomic Computing, 229-240. https://ieeexplore.ieee.org/abstract/document/1498067/ [6] Kamble, T., Deokar, S., Wadne, V. S., Gadekar, D. P., Vanjari, H. B., & Mange, P. (2023). Predictive Resource Allocation Strategies for Cloud Computing Environments Using Machine Learning. Journal of Electrical Systems, 19(2). [7] Khanday, S. K. (2024). Reinforcement Learning Strategies for Dynamic Resource Allocation in Cloud-Based Architectures. [8] Duc, T. L., Leiva, R. G., Casari, P., & Östberg, P. O. (2019). Machine learning methods for reliable resource provisioning in edge-cloud computing: A survey. ACM Computing Surveys (CSUR), 52(5), 1-39. https://dl.acm.org/doi/abs/10.1145/3341145 [9] Gao, J., Wang, H., & Shen, H. (2020). Machine learning based workload prediction in cloud computing. Proceedings of the 29th International Conference on Computer Communications and Networks, 1-9. https://ieeexplore.ieee.org/abstract/document/9209730/ [10] Gollapudi, S. (2016). Practical Machine Learning. Packt Publishing Ltd. [11] Huang, D., Costero, L., Pahlevan, A., Zapater, M., & Atienza, D. (2024). CloudProphet: a machine learning-based performance prediction for public clouds. IEEE Transactions on Sustainable Computing, 9(4), 661-676. https://ieeexplore.ieee.org/abstract/document/10415550/ [12] Murali Krishna, N., Rishika, V., Neelaveni, R. C., & Rama Krishna, D. (2025). Cloud resource forecasting using LSTM neural networks. Global Journal of Engineering Innovations & Interdisciplinary Research (GJEIIR), 5(4). https://www.sciencexcel.com/index.php/article/abstract/cloud-resource-forecasting-using-lstm-neuralnetworks [13] Aziz, S. F., & Kashmoola, M. Y. (2023). Workload Forecasting Methods in Cloud Environments: An Overview. ALRafidain Journal of Computer Sciences and Mathematics, 17(2), 29-37. https://iasj.rdd.edu.iq/journals/uploads/2024/12/14/fb45bc15943bfe1f20ea922c8d1edca9.pdf [14] Kumar, J., & Singh, A. K. (2018). Workload prediction in cloud using artificial neural network and adaptive differential evolution. Future Generation Computer Systems, 81, 41-52. https://www.sciencedirect.com/science/article/pii/S0167739X17300444 [15] Kundu, S., Rangaswami, R., Dutta, K., & Zhao, M. (2010, January). Application performance modeling in a virtualized environment. In HPCA-16 2010 The Sixteenth International Symposium on High-Performance Computer Architecture (pp. 1-10). IEEE. https://ieeexplore.ieee.org/abstract/document/5463058/ [16] Abraham, O. L., Ngadi, M. A., Sharif, J. B. M., & Sidik, M. K. M. (2025). Multi-objective optimization techniques in cloud task scheduling: A systematic literature review. IEEE Access. https://ieeexplore.ieee.org/abstract/document/10843235/ [17] Pires, F. L., & Barán, B. (2015, May). A virtual machine placement taxonomy. In 2015 15th IEEE/ACM international symposium on cluster, cloud and grid computing (pp. 159-168). IEEE. https://ieeexplore.ieee.org/abstract/document/7152482/ [18] Garikipati, V., & Kumar, V. (2020). Optimizing Traffic Management and Cloud Security in Software Networks Using Advanced Deep Learning Models for Application and Attack Classification. International Journal of HRM and Organizational Behavior, 8(3), 127-134. https://ijhrmob.org/index.php/ijhrmob/article/view/286 [19] Müller, A. C., & Guido, S. (2016). Introduction to Machine Learning with Python. O'Reilly Media Inc. http://thuvienso.thanglong.edu.vn/handle/TLU/3027 [20] Hasan, M. K., Jahan, N., Nazri, M. Z. A., Islam, S., Khan, M. A., Alzahrani, A. I., ... & Nam, Y. (2024). Federated learning for computational offloading and resource management of vehicular edge computing in 6G-V2X network. IEEE Transactions on Consumer Electronics, 70(1), 3827-3847. https://ieeexplore.ieee.org/abstract/document/10415079/ [21] Guerrero, C., Lera, I., & Juiz, C. (2018). Resource optimization of container orchestration: a case study in multicloud microservices-based applications. The Journal of Supercomputing, 74(7), 2956-2983. https://link.springer.com/article/10.1007/s11227-018-2345-2
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 119 [22] Quiroz, A., Kim, H., Parashar, M., Gnanasambandam, N., & Sharma, N. (2009, October). Towards autonomic workload provisioning for enterprise grids and clouds. In 2009 10th IEEE/ACM International Conference on Grid Computing (pp. 50-57). IEEE. https://ieeexplore.ieee.org/abstract/document/5353066/ [23] Barrak, A., Petrillo, F., & Jaafar, F. (2022). Serverless on machine learning: A systematic mapping study. IEEE Access, 10, 99337-99352. https://ieeexplore.ieee.org/abstract/document/9888122/ [24] Samarasinghe, S. (2016). Neural networks for applied sciences and engineering: from fundamentals to complex pattern recognition. Auerbach publications. https://www.taylorfrancis.com/books/mono/10.1201/9780849333750/neural-networks-applied-sciencesengineering-sandhya-samarasinghe [25] Shen, S., Van Beek, V., & Iosup, A. (2015, May). Statistical characterization of business-critical workloads hosted in cloud datacenters. In 2015 15th IEEE/ACM international symposium on cluster, cloud and grid computing (pp. 465-474). IEEE. https://ieeexplore.ieee.org/abstract/document/7152512/ [26] Saini, H., Singh, G., Dalal, S., Moorthi, I., Aldossary, S. M., Nuristani, N., & Hashmi, A. (2024). A hybrid machine learning model with self-improved optimization algorithm for trust and privacy preservation in cloud environment. Journal of Cloud Computing, 13(1), 157. https://link.springer.com/article/10.1186/s13677-02400717-6 [27] Hogade, N., & Pasricha, S. (2022). A survey on machine learning for geo-distributed cloud data center management. IEEE Transactions on Sustainable Computing, 8(1), 15-31. https://ieeexplore.ieee.org/abstract/document/9899733/ [28] Devineni, S., & Gorantla, B. (2023). Energy-Efficient Computing and Green Computing Techniques. Computer Science, Engineering and Technology, 1(4). https://www.academia.edu/download/118786304/5.pdf [29] Tsakalidou, V. N., Mitsou, P., & Papakostas, G. A. (2021). Machine learning for cloud resources management--An overview. arXiv preprint arXiv:2101.11984. https://arxiv.org/abs/2101.11984 [30] Wang, Y. (2022). Optimizing Distributed Computing Resources with Federated Learning: Task Scheduling and Communication Efficiency. Journal of Computer Technology and Software, 4(3). https://www.ashpress.org/index.php/jcts/article/view/147 [31] Ali, N. N., & Zeebaree, S. R. (2024). Distributed resource management in cloud computing: a review of allocation, scheduling, and provisioning techniques. The Indonesian Journal of Computer Science, 13(2). http://ijcs.net/ijcs/index.php/ijcs/article/view/3823 [32] Suleiman, I. B. Predicting Students' Academic Performance Using Linear Regression. Academia.edu, 2018. Link to paper [33] Valarmathi, K., Karthikeyan, M., & Navaneetha Krishnan, S. (2024). An efficient prediction-based dynamic resource allocation framework in quantum cloud using knowledge-based offline reinforcement learning. https://link.springer.com/article/10.1007/s42484-025-00257-5 [34] Sireesha, B., & Venkata Ramana, E. (2014). Dynamic resource allocation for cloud computing environment using virtual machines. International Journal of Computer Science and Information Technology Research, 2(3), 506– 511. http://www.researchpublish.com. [35] Ameur, A. B. (2023). Artificial intelligence for resource allocation in multi-tenant edge computing (Doctoral dissertation, Institut Polytechnique de Paris). https://theses.hal.science/tel-04419703/ [36] Mann, Z. Á. (2015). Allocation of virtual machines in cloud data centers—a survey of problem models and optimization algorithms. Acm Computing Surveys (CSUR), 48(1), 1-34. https://dl.acm.org/doi/abs/10.1145/2797211 [37] Gong, Y., Huang, J., Liu, B., Xu, J., Wu, B., & Zhang, Y. (2024). Dynamic resource allocation for virtual machine migration optimization using machine learning. arXiv preprint arXiv:2403.13619. https://arxiv.org/abs/2403.13619 [38] Iftikhar, S. (2024). Artificial Intelligence based Resource Management in Fog Computing (Doctoral dissertation, Queen Mary University of London). https://qmro.qmul.ac.uk/xmlui/handle/123456789/102198 [39] Ardagna, D., Casale, G., Ciavotta, M., Pérez, J. F., & Wang, W. (2014). Quality-of-service in cloud computing: modeling techniques and their applications. Journal of internet services and applications, 5(1), 11.
Global Journal of Engineering and Technology Advances, 2025, 25(02), 081–120 120 https://link.springer.com/article/10.1186/s13174-014-0011-3 Suleiman, I. B. Predicting Students' Academic Performance Using Linear Regression. Academia.edu, 2018. [40] Suleiman, I. B. Predicting Students' Academic Performance Using Linear Regression. Academia.edu, 2018. [41] Liu, H. et al. Natural Gas Demand Forecasting Model Based on LASSO and Polynomial Models and Its Application: A Case Study of China. Energies, 2023. https://www.researchgate.net/publication/371004638_Natural_Gas_Demand_Forecasting_Model_Based_on_L ASSO_and_Polynomial_Models_and_Its_Application_A_Case_Study_of_China [42] Huynh-Cam, T.-T., Chen, L.-S., & Le, H. Using Decision Trees and Random Forest Algorithms to Predict and Determine Factors Contributing to First-Year University Students’ Learning Performance. Algorithms, 2021. https://www.mdpi.com/1999-4893/14/11/318 [43] Rahimzad, M. et al. Performance Comparison of an LSTM-based Deep Learning Model versus Conventional Machine Learning Algorithms for Streamflow Forecasting. Water Resources Management, 2021. https://link.springer.com/article/10.1007/s11269-021-02937-w [44] Öz, E., Bulut, O., Cellat, Z. F., & Yürekli, H. Stacking: An Ensemble Learning Approach to Predict Student Performance in PISA 2022. https://link.springer.com/article/10.1007/s10639-024-13110-2