scieee AI-readable full text Open interactive document viewer

Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite

Kozik, Kacper; Szczepanek, Natalia Diana; Ketele, Ewoud

Abstract

This report explores the challenge of detecting anomalies in the performance of Grid compute nodes within the Worldwide LHC Computing Grid (WLCG). Using the HEP Benchmark Suite (HEP-Suite), the study explores how machine learning techniques can be applied to automate the detection of performance issues, on different CPU models. The report highlights the use of SCiForest, an improved version of the Isolation Forest algorithm, which is particularly effective at identifying both global scattered and clustered anomalies. Validation metrics such as the Jaccard Index, F1 Score, and Matthews Correlation Coefficient (MCC) were used to evaluate the performance of SCiForest. The results show that SCiForest is effective in detecting clustered anomalies, though some false positives remain, indicating areas for further improvement. The report suggests continuing this research by testing additional machine learning algorithms, such as the XGBoost outlier detection algorithm, to improve detection accuracy. This ongoing work is important for ensuring the continued performance and reliability of the computing infrastructure that supports critical experiments at CERN.

Full text

Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite August 2024 AUTHOR: Kacper Kamil Kozik SUPERVISORS: Natalia Diana Szczepanek Ewoud Ketele CERN openlab Report 2 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite PROJECT SPECIFICATION Detecting anomalies in the performance of Grid compute nodes presents a formidable challenge that has been addressed through various methodologies over the past decades. A recent innovative approach, pioneered by the HEPiX Benchmarking Group in the last few months, introduces a novel dimension to this pursuit. This method leverages the power of the HEP Benchmark Suite (HEP-Suite) to scrutinize and validate the performance metrics of Grid compute nodes in relation to fabric metrics. By executing the HEPSuite tool on the grid, a substantial dataset of performance metrics has been meticulously compiled. Preliminary investigations have already indicated the promising potential of utilizing this data for constructing an anomaly detection system. Considering the extensive volume of collected metrics (features), a more profound analysis could significantly benefit from the application of advanced machine learning techniques. The objective of the summer student project is to assess the impact of these features on the computational efficiency of job slots. This involves a focused exploration of critical configurations through an in-depth analysis of the extensive dataset. The project involved deep engagement with benchmarking processes, utilizing the capabilities of HEPscore23, and developing a machine learning model specifically tailored for comprehensive parameter correlation analysis. This work provided the opportunity to make significant contributions to cutting-edge research while gaining a comprehensive understanding of high-energy physics computational frameworks. CERN openlab Report 3 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite ABSTRACT This report explores the challenge of detecting anomalies in the performance of Grid compute nodes within the Worldwide LHC Computing Grid (WLCG). Using the HEP Benchmark Suite (HEPSuite), the study explores how machine learning techniques can be applied to automate the detection of performance issues, on different CPU models. The report highlights the use of SCiForest, an improved version of the Isolation Forest algorithm, which is particularly effective at identifying both global scattered and clustered anomalies. Validation metrics such as the Jaccard Index, F1 Score, and Matthews Correlation Coefficient (MCC) were used to evaluate the performance of SCiForest. The results show that SCiForest is effective in detecting clustered anomalies, though some false positives remain, indicating areas for further improvement. The report suggests continuing this research by testing additional machine learning algorithms, such as the XGBoost outlier detection algorithm, to improve detection accuracy. This ongoing work is important for ensuring the continued performance and reliability of the computing infrastructure that supports critical experiments at CERN. CERN openlab Report 4 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite TABLE OF CONTENTS INTRODUCTION 01 PROBLEM DEFINITION AND APPROACH 02 TYPES OF ANOMALIES 03 GENERAL TYPES OF ANOMALIES TYPES OF CLUSTERED ANOMALIES IN BENCHMARKING DATA MACHINE LEARNING ALGORITHMS FOR ANOMALY DETECTION 04 ISOLATION FOREST SPLIT-SELECTION CRITERION ISOLATION FOREST FEATURE SPACE TRANSFORMATION 05 RESULTS 06 MORE SPLIT-SELECTION CRITERION ISOLATION FOREST RESULTS VALIDATION CONCLUSIONS 07 CERN openlab Report 5 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite 1. INTRODUCTION Benchmarking is a crucial process in evaluating the performance of systems or components, such as smartphones or CPU models. Just as consumers compare smartphones based on benchmark scores to find the best option for their needs, benchmarking in computing allows for the comparison of servers and CPU models using predefined tasks or workloads. This helps determine which system or component offers the best performance for specific tasks, guiding informed decision-making. In the field of high-energy physics (HEP) and the Worldwide LHC Computing Grid (WLCG) [1], benchmarking is essential for comparing CPU performance across different countries. The WLCG, which supports CERN’s large-scale experiments, relies heavily on efficient computing power. Benchmarking ensures that the computing resources allocated to these experiments are both effective and sufficient, directly influencing financial planning and resource allocation. CPU power is often quantified by multiplying the benchmark score by the number of seconds a given application is used [2]. The HEP Benchmark Suite [3] is specifically designed to meet the benchmarking needs of the WLCG. It consists of three primary packages: • HEP-Workloads: This package provides the tasks used to measure CPU performance. • HEP-Score: It calculates a benchmark score, known as HEPscore by taking the geometric mean of the results from the HEP-Workloads tasks. • HEP-Benchmark-Suite: This package coordinates the benchmarking process and compiles the results into a standardized JSON file for further analysis. One of the latest configurations, HEPScore23 [2], has been officially adopted by the WLCG. It includes seven workloads derived from five major experiments ATLAS [4], CMS [5], ALICE [6], Belle II [7], and LHCb [8]. The structure and components of these packages are illustrated in Figure 1. Figure 1. Diagram of the HEP Benchmark Suite, illustrating the workflow for configuring, running benchmarks, processing data, and publishing results [9]. Benchmarking in WLCG is not just about evaluating performance, it also plays a key role in resource allocation among countries. Each country contributes computing resources based on the size of its research community, either through direct funding or by providing equipment. Instead of focusing on cost alone, the WLCG uses the CPU power delivered by each site as a standard measure. This ensures that resource contributions are fairly evaluated and helps in planning and distributing computational resources efficiently [2]. CERN openlab Report 6 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite 2. PROBLEM DEFINITION AND APPROACH The HEP Benchmark Suite, currently in use on the Worldwide LHC Computing Grid (WLCG), has evolved to include more than just performance metrics. With the integration of various plugins, the suite can now collect user-defined metadata, such as server load, memory usage, and power consumption. Among these, the load per core is particularly significant, as it can be correlated with the HEPscore per core, a measure of a server's performance. By plotting these two metrics against each other in a 2D scatter plot, it is possible to identify patterns that reveal the performance of the computing infrastructure. Specifically, this plot can highlight anomalies, such as misconfigured servers or underperforming hardware, which may not align with expected performance levels. However, the current approach to anomaly detection within this context is manual and time-consuming. Analysts must repeatedly review these plots, searching for deviations that indicate potential issues. This repetitive task not only consumes valuable time but also introduces the possibility of human error. The goal of this project is to automate the anomaly detection process using machine learning techniques. By training models to recognize patterns within the scatter plots, it becomes possible to efficiently and accurately identify anomalies, thereby improving the overall reliability and performance of the WLCG [10]. To illustrate the concept, a scatter plot of HEPscore per core versus load per core is provided (see Figure 2), where each point represents data from a different site but using the same CPU model. By analyzing the distribution of these points, it is possible to identify anomalies - sites where the server load and performance do not align as expected, potentially indicating a misconfiguration or other issues. The correlation between HEPscore per core and load per core serves as a key feature in the unsupervised learning approach. This relationship will be leveraged in the later stages of the project to build a machine learning model capable of identifying anomalies across different sites in an automated and efficient manner, as described in the following sections. Figure 2. Scatter plot of HEPscore per Core vs. Load per Core for one CPU model (AMD EPYC 7542 32-Core Processor) across different sites, with marginal distributions shown along the top and right axes for each site. CERN openlab Report 7 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite 3. TYPES OF ANOMALIES a. GENERAL TYPES OF ANOMALIES Understanding the different types of anomalies is essential for effective anomaly detection in performance benchmarking. According to Han et al. (2022) [11], as shown in Figure 3, anomalies can be categorized into several types, each exhibiting unique characteristics: • Local Anomalies: These are deviations from the normal data that occur within a local neighborhood. They are identified by comparing data points within a small, localized area, making them distinct from the surrounding data. • Global Scattered Anomalies: These anomalies are significantly different from the normal data and are spread far apart from each other. Generated from a uniform distribution, they represent data points that are uniformly distant from the normal data distribution. • Dependency Anomalies: These occur when the dependency structure of the data is violated. In other words, the features of the data are expected to have a certain relationship, and when this relationship is not followed, a dependency anomaly is identified. • Global Clustered Anomalies: Also known as group anomalies, these are clusters of data points that deviate from the normal data distribution. They exhibit similar characteristics and are often grouped together, away from the trendline of normal data. Figure 3. Visualization of four common types of anomalies in datasets highlighting the differences in how normal data points (blue) and anomalous data points (red) are distributed [11]. In the context of the HEP Benchmark Suite and the Worldwide LHC Computing Grid (WLCG), identifying anomalies is essential for maintaining the efficiency and reliability of the computing infrastructure. Specifically, the anomalies we focus on include global clustered anomalies and global scattered anomalies: • Global Clustered Anomalies: These anomalies are characterized by a group of data points that deviate together from the expected trend. In the WLCG dataset, this is typically observed when multiple servers from the same site show abnormal behavior, such as consistently lower or higher HEPscore relative to their load per core. • Global Scattered Anomalies: These are isolated data points that deviate significantly from the expected performance but do not form a cluster. In the WLCG dataset, these scattered points usually represent servers with individual misconfigurations rather than systematic errors. For this analysis, the primary focus is on identifying global clustered anomalies, as they are indicative of larger systemic issues that could affect multiple servers or sites. The scatter plot of HEPscore per core versus load per core is a useful tool for visualizing these anomalies. In the scatter plot presented in Figure 4, points that are far from the trendline represent potential global clustered or scattered anomalies. CERN openlab Report 8 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite Figure 4. Types of Anomalies in WLCG Grid Nodes for a selected CPU model (AMD EPYC 7542 32-Core Processor), showing global clustered and scattered anomaly regions. b. TYPES OF CLUSTERED ANOMALIES IN BENCHMARKING DATA Clustered anomalies can be categorized into three main types: underperformance anomalies, overperformance anomalies, and 'other' anomalies. • Underperformance Anomalies Underperformance anomalies are identified by clusters of data points that fall below the general trendline in a scatter plot of HEPscore per core versus load per core. These anomalies indicate that certain servers or nodes are delivering less computational power than expected, given their load. This type of anomaly could be due to hardware degradation, misconfigurations, or inefficiencies in the system (example in Figure 5). • Overperformance Anomalies Conversely, overperformance anomalies are characterized by clusters of data points that fall above the general trendline. These indicate servers or nodes that are performing better than expected under a given load. While this might seem beneficial, overperformance can sometimes signal an issue, that could lead to inconsistent resource allocation or misleading performance metrics. • Other Anomalies The third type, other anomalies, are more complex as they identify a part of the site as abnormal. These anomalies are often indicated by a 'flat area' of data points on the scatter plot, where the trendline for the site differs significantly in slope compared to the general trendline for the CPU model. In such cases, the slope might be nearly horizontal, indicating a lack of expected performance scaling with increased load (example in Figure 5). CERN openlab Report 9 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite Figure 5. Types of clustered anomalies with highlighted types of anomalies. 4. MACHINE LEARNING ALGORITHMS FOR ANOMALY DETECTION Anomaly detection in large-scale datasets requires robust and efficient machine learning algorithms. Among the many available algorithms, two have shown particular promise: Isolation Forest [12] and SplitSelection Criterion Isolation Forest (SciForest) [13]. These algorithms are particularly effective at detecting both global scattered and clustered anomalies, making them well-suited for the complex data involved in HEP benchmarking, especially due to the data distribution and the anomaly regions with both types of global anomalies. a. ISOLATION FOREST Isolation Forest is a popular anomaly detection algorithm known for its effectiveness in identifying global scattered anomalies. The algorithm works by isolating observations in the feature space using random splits along hyperplanes. This process, which involves recursively partitioning the data, is particularly effective at identifying outliers that are sparsely distributed across the dataset. The key advantage of Isolation Forest lies in its ability to generalize well across different types of data distributions. By relying on random splits, it can quickly and efficiently isolate anomalies that differ significantly from the majority of the data, particularly when those anomalies are scattered rather than clustered [12]. b. SPLIT-SELECTION CRITERION ISOLATION FOREST Split-selection Criterion Isolation Forest is an enhanced version of the standard Isolation Forest algorithm. While Isolation Forest excels at detecting global scattered anomalies, SCiForest is designed to improve the detection of both global scattered and clustered anomalies. The primary innovation in SCiForest is the use of a selection split criterion based on data statistics rather than purely random splits. This method involves selecting hyperplanes that better separate local distribution peaks within the data, allowing the algorithm to more effectively isolate clusters of anomalies. By focusing on areas where data points are densely packed, SCiForest can detect subtle variations that might indicate a clustered anomaly. In the training stage of SCiForest, multiple isolation trees are generated from random sub-samples of the data. Each tree is constructed by selecting separating hyperplanes that maximize the separation of data distributions at each recursive subdivision. The evaluation stage then measures the path length of each data point within the tree, with shorter path lengths indicating anomalies. This approach encapsulates the two key properties of anomalies: they are few and significantly different from normal instances, which typically have longer path lengths in the trees [13]. SCiForest generalizes well, much like Isolation Forest, but its refined split-selection criterion makes it particularly effective at detecting clustered anomalies. CERN openlab Report 16 Anomaly Detection in Grid Compute Nodes: A Machine Learning Approach Leveraging HEP Benchmark Suite REFERENCES [1] WLCG, The Worldwide LHC Computing Grid. https://wlcg.web.cern.ch, Accessed September 10, 2024 [2] Giordano, D., et al. (2023). HEPScore: A new CPU benchmark for the WLCG. https://doi.org/10.48550/arXiv.2306.08118 [3] HEP Benchmark Suite. https://gitlab.cern.ch/hep-benchmarks/hep-benchmark-suite, Accessed September 10, 2024 [4] ATLAS, A Toroidal LHC Apparatus. https://atlas.cern, Accessed September 10, 2024 [5] CMS, Compact Muon Solenoid. https://cms.cern, Accessed September 10, 2024 [6] ALICE, A Large Ion Collider Experiment. https://alice.cern, Accessed September 10, 2024 [7] Belle II. https://www.belle2.org, Accessed September 10, 2024 [8] LHCb, Large Hadron Collider beauty. https://lhcb.web.cern.ch, Accessed September 10, 2024 [9] Giordano, D., et al. (2021). HEPiX Benchmarking Solution for WLCG Computing Resources. https://doi.org/10.1007/s41781-021-00074-y [10] Szczepanek, N., et al. (2024). HEP Benchmark Suite: Enhancing Efficiency and Sustainability in Worldwide LHC Computing Infrastructures. https://doi.org/10.48550/arXiv.2408.12445 [11] Han, S., et al. (2022). ADBench: Anomaly Detection Benchmark. arXiv:2206.09426v2 [cs.LG]. https://doi.org/10.48550/arXiv.2206.09426 [12] Liu, F. T., et al. (2008). Isolation Forest. Proceedings of the 2008 Eighth IEEE International Conference on Data Mining, 413–422. https://doi.org/10.1109/ICDM.2008.17 [13] Liu, F. T., et al. (2010). On Detecting Clustered Anomalies Using SCiForest. In Balcázar, J. L., Bonchi, F., Gionis, A., & Sebag, M. (Eds.), Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2010. Lecture Notes in Computer Science (Vol. 6322). Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-15883-4_18