scieee AI-readable full text Open interactive document viewer

Abuse and Anomaly Detection using Machine Learning for Gitlab Runners

Nersesyan, Diana; Suara, Subhashis; Posada Trobo, Ismael

Abstract

This report describes the development of a prototype machine learning (ML) system designed to detect anomalies and potential misuse within the GitLab Runner infrastructure at CERN. Two complementary approaches were applied: a classical unsupervised model (Isolation Forest) and a deep sequence model (LSTM autoencoder). The system was trained on 45 days of Prometheus metrics (CPU and memory) and evaluated using synthetic anomalies such as spikes, bursts, flatlines, and noise. Results showed that the LSTM autoencoder generalized well, effectively capturing temporal workload patterns, while Isolation Forest provided a lightweight baseline that was especially effective for spike detection. While still in an experimental stage, the prototype highlights both the promise and the challenges of applying ML and DL in large-scale CI/CD environments, and points toward future steps such as integrating OpenSearch logs, improving interpretability, and enabling real-time detection for reliable and fair resource usage across CERN’s shared computing platforms.

Full text

CERN openlab Report // 2025 PROJECT SPECIFICATION The Git Service at CERN provides a large-scale GitLab Runners infrastructure that supports CI/CD pipeline execution across the organization. Everyday, more than 50,000 jobs are executed across 10 dedicated GitLab runners clusters, each tailored to the diverse needs of the CERN community. While this shared runners environment provides flexibility and scalability, it is not protected against misuse or malicious activity. Users may unintentionally or deliberately launch resource-intensive workloads that monopolize cluster capacity, resulting in reduced availability and degraded service quality for other users. To address this challenge, we propose enhancing the system’s monitoring and security capabilities by developing a Machine Learning based anomaly and abuse detection system. 2 CERN openlab Report // 2025 ABSTRACT This report describes the development of a prototype machine learning (ML) system designed to detect anomalies and potential misuse within the GitLab Runner infrastructure at CERN. Two complementary approaches were applied: a classical unsupervised model (Isolation Forest) and a deep sequence model (LSTM autoencoder). The system was trained on 45 days of Prometheus metrics (CPU and memory) and evaluated using synthetic anomalies such as spikes, bursts, flatlines, and noise. Results showed that the LSTM autoencoder generalized well, effectively capturing temporal workload patterns, while Isolation Forest provided a lightweight baseline that was especially effective for spike detection. While still in an experimental stage, the prototype highlights both the promise and the challenges of applying ML and DL in large-scale CI/CD environments, and points toward future steps such as integrating OpenSearch logs, improving interpretability, and enabling real-time detection for reliable and fair resource usage across CERN’s shared computing platforms. 3 CERN openlab Report // 2025 TABLE OF CONTENTS 1. Introduction 06 1.1 GitLab CI/CD Infrastructure at CERN 1.2 Challenges in Ensuring Fair Usage of GitLab Runners 1.3 Motivation for Machine Learning Approaches 1.4 Scope of the Report 2. Monitoring in CI/CD: Background and Related Work 08 2.1 GitLab Runner Architecture and Usage 2.2 Existing Monitoring and Abuse Detection Approaches 2.3 Machine Learning for Anomaly Detection in CI/CD 3. Methodology 11 3.1 Data Sources (Prometheus, OpenSearch) 3.2 Data Preprocessing and Feature Engineering 3.3 LSTM Autoencoder and Isolation Forest Models 3.4 Prototype Pipeline Setup 4. Project Progress and Results 13 4.1 Dataset Preparation and Statistics 4.2 Model Training and Evaluation 4.3 Anomaly Detection Experiments 4.4 Observed Challenges and Limitations 5. Conclusions and Future Work 15 5.1 Summary of Contributions 5.2 Next Steps for Production Deployment 4 CERN openlab Report // 2025 6. References 17 7. Acknowledgments 17 5 CERN openlab Report // 2025 1. INTRODUCTION 1.1 GitLab CI/CD Infrastructure at CERN CERN’s GitLab CI/CD infrastructure operates at a significant scale within the research community. It supports more than 17,000 active users and over 180,000 projects, forming the backbone of software requirements for physics experiments and organizational services. On average, more than 20,000 pipelines are triggered monthly, resulting in over 50,000 jobs executed every day. To sustain this scale, CERN’s GitLab team provides a variety of runners, each optimized for specific workloads (see Figure 1.1). Shared runners handle most standard jobs, Spark runners support data processing with Apache Spark, GPU runners accelerate workloads requiring high-performance computation and so on. This diverse runner ecosystem ensures flexibility and efficiency in meeting the needs of the CERN community. Figure 1.1: Overview of GitLab Runner Types at CERN 1.2 Challenges in Ensuring Fair Usage of GitLab Runners Despite its robustness, the shared runner environment at CERN presents ongoing challenges in ensuring fair usage within the CERN community. Because thousands of users share the same infrastructure, individual jobs may unintentionally or deliberately consume disproportionate 6 CERN openlab Report // 2025 amounts of resources. Such workloads can degrade overall performance and, in extreme cases, render clusters unavailable for the majority of users. The current monitoring setup is effective at collecting metrics but relies heavily on manual inspection to identify cases of unfair usage. To address this gap, we implemented a custom thresholding check as part of this project. In practice, anomalies were defined not only through ML scores but also with simple rules such as CPU usage above 80% or sudden spikes in resource consumption. While this provided a baseline for comparison, it underscored the need for more adaptive approaches that can capture subtle or evolving misuse patterns. 1.3 Motivation for Machine Learning Approaches To address these challenges, there is growing interest in applying machine learning (ML) and deep learning (DL) methods for anomaly and abuse detection in CI/CD environments. Unlike static thresholds, ML-based approaches can adapt to evolving workloads and detect subtle deviations from normal behavior that may signal resource misuse or malicious activity. In this project, we applied both a classical anomaly detection method (Isolation Forest) and a deep sequence model (LSTM autoencoder) to GitLab Runner workload metrics. Isolation Forest provides a lightweight baseline for detecting unusual patterns in aggregated data, while the LSTM autoencoder leverages temporal dependencies to identify irregularities through reconstruction error. Together, these methods demonstrate the feasibility of ML-driven monitoring and mark a step toward more adaptive and intelligent systems for safeguarding shared resources and ensuring fair access within the CERN user community. 1.4 Scope of the Report This report presents the development and evaluation of a prototype anomaly detection system for GitLab Runners at CERN. The work is structured around three main aspects: ● Data Collection and Preparation: Gathering metrics such as CPU and memory usage from Prometheus. ● Prototype Model Development: Applying and evaluating both an LSTM autoencoder and an Isolation Forest model to establish baseline job behavior and detect anomalies. ● Preliminary Results and Insights: Assessing the prototype on collected data to evaluate its effectiveness in identifying resource anomalies and irregular job patterns. It should be emphasized that this project remains at the prototype stage and has not yet been integrated into production. Moreover, the evaluation relied on synthetic anomalies, since real labeled misuse cases were unavailable during development and evaluation. The current focus is on demonstrating feasibility, identifying limitations, and outlining directions for future development. 7 CERN openlab Report // 2025 2. MONITORING IN CI/CD: BACKGROUND AND RELATED WORK 2.1 GitLab Runner Architecture and Usage GitLab Runners are lightweight agents that execute CI/CD jobs defined in GitLab pipelines. Runners support several executors, such as Shell, Docker, and Kubernetes, which define how jobs are run. In the CERN environment, GitLab Runners are deployed exclusively with the Kubernetes executor. Each job submitted through GitLab is scheduled as a Kubernetes pod, providing strong workload isolation, efficient resource allocation, and seamless scalability. This approach enables the infrastructure to support a wide range of projects and users, from small-scale software builds to compute-intensive scientific workflows, while maintaining fairness and stability across the shared platform. The overall workflow can be summarized as follows: 1. A user commits code and triggers a CI/CD pipeline in GitLab. 2. The GitLab Runner receives the job and schedules it as a Kubernetes pod. 3. The pod executes the defined stages (.i.e. build or test) in an isolated containerized environment. 4. During execution, system-level metrics (e.g., CPU, memory usage) are collected by Prometheus and logs are forwarded to OpenSearch. 5. These monitoring data sources provide the basis for dashboards, anomaly detection, and root cause analysis. This pipeline ensures both operational efficiency and transparency, while laying the groundwork for applying machine learning-based anomaly detection methods. 2.2 Existing Monitoring and Abuse Detection Approaches At present, anomaly monitoring and root cause identification in the GitLab Runner infrastructure at CERN rely primarily on two tools: Grafana Dashboards and OpenSearch Logs. The GitLab team at CERN have developed a set of custom dashboards in Grafana. These dashboards visualize real-time metrics collected from Prometheus, including CPU and memory usage, job durations, job queue lengths, etc. Filters by runner type and project allow to quickly isolate suspicious activity. For example, dashboards can highlight infrastructure saturation, a growing backlog of pending jobs, or stressed CPU and memory resources (see Figures 2.1 - 2.3). OpenSearch serves as the centralized storage layer for GitLab Runner job logs. It enables searching, filtering, and correlating events across different runners and projects, making it indispensable for root cause analysis once anomalies are identified. We aimed to evaluate the 8 CERN openlab Report // 2025 existing OpenSearch Anomaly Detection plugin. To ensure a controlled environment, a dedicated cluster was set up with the required configurations, permissions, and a copy of the relevant logs. The plugin has great potential but with the logs available for GitLab runner from our end, we could not leverage the plugin's potential for anomaly detection in runners, however, it could be a great addition for anomaly detection in the GitLab application. Together, Grafana and OpenSearch form the foundation of current monitoring at CERN. While effective for visualization and root cause analysis, these tools depend heavily on manual inspection, which can be time-consuming and make it difficult to capture subtle or evolving misuse patterns. This limitation motivates the application of ML based approaches for more adaptive and proactive anomaly detection. Figure 2.1: Grafana dashboard visualization of runner infrastructure saturation 9 CERN openlab Report // 2025 Overall, the work provides a foundation for extending anomaly detection capabilities within CERN’s GitLab Runner ecosystem, complementing existing monitoring tools such as Grafana and OpenSearch. 5.2 Next Steps for Production Deployment While the prototype confirmed the feasibility of ML-based anomaly detection, several steps remain before deployment at production scale: ● Extended data coverage: Collect at least 6-12 months of Prometheus and OpenSearch data to capture seasonal usage patterns and rare anomalies. ● Hyperparameter optimization: Apply automated search methods (e.g., grid search, Bayesian optimization) to refine LSTM and Isolation Forest configurations. ● Hybrid detection strategy: Combine static threshold checks (for clear-cut cases such as CPU > 80%) with ML models for subtler anomaly detection. ● Exploration of alternative models: Extend the study to other anomaly detection methods such as Matrix Profile or more advanced deep learning architectures, to compare performance and robustness across approaches. ● Integration of OpenSearch logs for security: Incorporate job-level logs into anomaly detection to increase visibility into what commands were executed within GitLab Runners. This would allow not only resource misuse detection but also early identification of potentially malicious or suspicious command sequences. ● Alert calibration and interpretability: Define anomaly score thresholds that balance sensitivity and false-positive rates, and develop tools to explain flagged anomalies. ● Scalability and real-time inference: Move toward streaming-based preprocessing and inference, enabling near real-time alerts integrated with CERN’s monitoring stack. These steps will help transition the system from a prototype to a reliable production system capable of protecting shared CI/CD resources and ensuring fair usage across the CERN community. 16 CERN openlab Report // 2025 6. REFERENCES 1. Breunig, M. M., Kriegel, H. P., Ng, R. T., & Sander, J. (2000). LOF: Identifying density-based local outliers. Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, 29(2), 93–104. https://doi.org/10.1145/342009.335388 2. GitLab. (n.d.). GitLab Runner documentation. Retrieved September 23, 2025, from https://docs.gitlab.com/runner/ 3. Grafana Labs. (n.d.). Grafana documentation. Retrieved September 23, 2025, from https://grafana.com/docs/ 4. Liu, F. T., Ting, K. M., & Zhou, Z. H. (2008). Isolation forest. 2008 Eighth IEEE International Conference on Data Mining, 413–422. https://doi.org/10.1109/ICDM.2008.17 5. Malhotra, P., Vig, L., Shroff, G., & Agarwal, P. (2016). Long short term memory networks for anomaly detection in time series. Proceedings of the 23rd European Symposium on Artificial Neural Networks (ESANN). https://arxiv.org/abs/1607.00148 6. OpenSearch. (n.d.). OpenSearch documentation. Retrieved September 23, 2025, from https://opensearch.org/docs/ 7. Prometheus Authors. (n.d.). Prometheus documentation. Retrieved September 23, 2025, from https://prometheus.io/docs/ 8. Yeh, C. C. M., Zhu, Y., Ulanova, L., Begum, N., Ding, Y., Dau, H. A., ... & Keogh, E. (2016). Matrix profile I: All pairs similarity joins for time series: A unifying view that includes motifs, discords and shapelets. 2016 IEEE 16th International Conference on Data Mining (ICDM), 1317–1322. https://doi.org/10.1109/ICDM.2016.0179 9. Trobo, I. P., Evangelou, K., Suara, S., & Rocha, R. (2023, December 6). GitLab Runners and Kubernetes: A Powerful Duo for CI/CD. https://kubernetes.docs.cern.ch/. https://kubernetes.docs.cern.ch/blog/2023/12/06/gitlab-runners-and-kubernetes-a-powerf ul-duo-for-ci/cd/ 7. ACKNOWLEDGMENTS I would like to express my sincere gratitude to my supervisors, Subhashis Suara and Ismael Posada Trobo, for their continuous guidance and support throughout this project. Their feedback, encouragement, and expertise were invaluable in shaping the direction of this work and ensuring its successful completion. I am also grateful to Konstantinos Evangelou for his valuable input during the early stages of the project. 17