Abuse and Anomaly Detection using Machine Learning for Gitlab Runners
Authors/Creators
Description
This report describes the development of a prototype machine learning (ML) system designed to detect anomalies and potential misuse within the GitLab Runner infrastructure at CERN. Two complementary approaches were applied: a classical unsupervised model (Isolation Forest) and a deep sequence model (LSTM autoencoder). The system was trained on 45 days of Prometheus metrics (CPU and memory) and evaluated using synthetic anomalies such as spikes, bursts, flatlines, and noise. Results showed that the LSTM autoencoder generalized well, effectively capturing temporal workload patterns, while Isolation Forest provided a lightweight baseline that was especially effective for spike detection. While still in an experimental stage, the prototype highlights both the promise and the challenges of applying ML and DL in large-scale CI/CD environments, and points toward future steps such as integrating OpenSearch logs, improving interpretability, and enabling real-time detection for reliable and fair resource usage across CERN’s shared computing platforms.
Files
DianaNERSESYAN-2025SummerStudent-Report.pdf
Files
(3.0 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:577c490dc3c3cb3335ca8fbf5842ac7e
|
3.0 MB | Preview Download |