Full text
INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 AUGUST 2025 AUTHOR(S): Julia Sophie Troppens EP-ADT-TR SUPERVISOR(S): Stefano Veneziano
CERN openlab Report PROJECT SPECIFICATION In 2025 the ATLAS trigger system will be running its first machine learning-based anomaly detection trigger, which includes both L1 and HLT algorithms. The objective of a 2025 summer project will be to study the first data selected by this trigger, and use these results to inform future searches and trigger improvements looking towards 2026 and the HL-LHC. INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 1
CERN openlab Report ABSTRACT In 2025, the first anomaly detection algorithm for the ATLAS Level-1 trigger in Run 3 was implemented using machine learning. This report investigates variations of the algorithm with the focus on input normalisation strategies aiming to identify improvements for the 2026 implementation. Since the algorithm must operate under strict latency and hardware constraints, the number of operations required for normalisation is highly restricted. This study evaluates four approaches: (i) power-of-two standard deviation scaling (current model), (ii) refined power-of-two standard deviation scaling, (iii) floating-point standard deviation scaling, and (iv) full power-of-two standardisation including mean subtraction. A particular focus is placed on evaluating the impact of normalisation on model performance for (ii) the refined power-of-two standard deviation scaling and (iv) the full power-of-two standardisation. Both methods are studied using a VAE-GAN trained on 2×106Enhanced Bias and Zero Bias events, allowing a comparison of how these normalisation strategies influence the anomaly detection score and reconstruction loss. The refined power-of-two standard deviation scaling successfully brings all input features to a comparable scale but still retains a strong dependence of the anomaly detection score on the leading jet pT. In contrast, the full power-of-two standardisation equalises the contribution of all jet and τlepton pTfeatures and incorporates muon information. However, this approach suffers from training instability as the learned rules strongly depend on the random seed. Nonetheless, the more balanced treatment of features appears promising for the development of a more robust algorithm. INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 2
CERN openlab Report TABLE OF CONTENTS 1 INTRODUCTION 4 1.1 CurrentStatus .................................... 4 2 Visual Analysis of Normalisation Effects on Input Distributions 5 2.1 Current Method: Power-of-Two Standard Deviation Scaling . . . . . . . . . . . 5 2.2 Refined Power-of-Two Standard Deviation Scaling . . . . . . . . . . . . . . . . . 5 2.3 Float-Precision Standard Deviation Scaling . . . . . . . . . . . . . . . . . . . . . 7 2.4 Power-of-Two Standardisation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3 Evaluating the Impact of Normalisation on Model Performance 7 3.1 Float-Precision vs Power-of-Two-Precision . . . . . . . . . . . . . . . . . . . . . 7 3.2 AD Model Alterations for Introducing Offset . . . . . . . . . . . . . . . . . . . . 8 3.3 Correlation of Standard Deviation Scaling and Standardisation . . . . . . . . . . 9 3.4 Performance of the Refined Power-of-Two Standard Deviation Scaling . . . . . . 9 3.4.1 Feature Distribution of Anomalous Events . . . . . . . . . . . . . . . . . 9 3.4.2 pTCorrelation with Anomaly Score . . . . . . . . . . . . . . . . . . . . . 11 3.4.3 Random Seed Dependence . . . . . . . . . . . . . . . . . . . . . . . . . . 11 3.5 Performance of the Power-of-Two Standardisation . . . . . . . . . . . . . . . . . 12 3.5.1 Feature Distribution of Anomalous Events . . . . . . . . . . . . . . . . . 12 3.5.2 pTCorrelation with Anomaly Score . . . . . . . . . . . . . . . . . . . . . 13 3.5.3 Random Seed Dependence . . . . . . . . . . . . . . . . . . . . . . . . . . 14 4 Conclusion 15 A Appendix 16 A.1 Code Correction for Rounding to next Power-of-Two . . . . . . . . . . . . . . . 16 A.2 Code Implementation of Mask with Offset . . . . . . . . . . . . . . . . . . . . . 16 A.3 Correlation of Standard Deviation Scaling and Standardisation . . . . . . . . . . 16 A.4 Additional Material for Analysis of the Refined Power-of-Two Standard DeviationScaling ...................................... 16 A.4.1 Anomalous Event Distributions . . . . . . . . . . . . . . . . . . . . . . . 17 A.4.2 Correlation Plots of pTand Latent Variables and Loss Scores . . . . . . . 18 A.4.3 Random Seed Dependence . . . . . . . . . . . . . . . . . . . . . . . . . . 20 A.5 Additional Material for Analysis of Power-of-Two Standardisation . . . . . . . . 21 A.5.1 Anomalous Event Distributions . . . . . . . . . . . . . . . . . . . . . . . 21 A.5.2 Correlation Plots of pTand Latent Variables and Loss Scores . . . . . . . 22 INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 3
CERN openlab Report 1 INTRODUCTION At LHC proton bunches collide every 25 ns, producing an enormous amount of collision data. Most of these events are well understood and not of interest for new physics analyses, and it is not feasible to record all events. Therefore, a trigger system is used to filter out interesting events that are stored, while the rest are discarded. The ATLAS trigger is a two-stage system: the L1 and the HLT. The L1 trigger is implemented in firmware and must operate under strict latency constraints. It reduces the event rate from the collision frequency of 40 MHz to about 100 kHz. This is followed by the software-based HLT, which has more relaxed latency requirements and further reduces the event rate to about 3 kHz for permanent storage. During Run 3, the L1 trigger incorporates the L1 Topological Processor (L1Topo), which receives input from the L1Calo and L1Muon subsystems. These subsystems identify high-energy objects in the calorimeters and muon detectors and provide information such as particle identification, momentum, and flight direction, as determined by the trigger algorithms. Based on this input, the L1Topo performs fast, topology-based trigger decisions (see [1,2] for detailed descriptions). In 2025, a new Anomaly Detection trigger (AD) was introduced into the L1Topo, incorporating machine learning methods. This system is designed to capture complex correlations of various input features that traditional approaches may miss. However, because of the strict timing limitations at L1, the ML models must remain extremely compact. 1.1 Current Status The current AD algorithm is based on a combination of a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN). It receives 44 inputs per event: •pT, η, φ from –6 leading jets –4 leading τleptons –4 leading muons •pTand φfrom missing ET The VAE learns to compress the data into a small latent representation of only six values and reconstruct it by reducing a custom MSE loss. The GAN learns to distinguish real from fake, thereby producing more realistic outputs. The model is trained on approximately 2×106Enhanced Bias and Zero Bias events. This allows it to learn how to encode and decode common events, while failing to accurately reconstruct rare ones. Anomalous events can therefore be distinguished from normal ones by applying a threshold to the reconstruction loss. Due to latency constraints, the Level-1 implementation uses only the encoder, mapping events into the latent space. The anomaly detection score is then calculated directly from the latent variables µ. AD score =µ2 1+µ2 2+µ2 3 Unpublished analysis results indicate that the model relies primarily on a limited subset of features, while others are ignored. This behaviour is undesirable, as our goal is to make a topological decision that accounts for the interplay among all features characterising an event. The AD algorithm is being updated for next year to address these issues. The new algorithm INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 4
CERN openlab Report must be completed by December 2025. Improving AD trigger performance is approached in two ways: by optimising the inputs and refining the model. This project focuses mainly on the former, investigating input normalisation and its impact on the model performance. 2 Visual Analysis of Normalisation Effects on Input Distributions Bringing input features to a comparable scale is essential for the training and performance of neural networks. When features vary widely in scale, those with larger ranges can dominate the loss function, biasing the model [3]. This leads to unequal feature contributions and suboptimal performance. Normalisation mitigates this by ensuring that all inputs contribute more evenly, enabling the model to capture complex correlations between features more effectively. 2.1 Current Method: Power-of-Two Standard Deviation Scaling The current input normalisation method is inspired by the standard score approach, which is commonly used in machine learning. z=x−µ σ However, the normalisation must be Level-1 compliant, meaning that computational steps must be minimised. Division by arbitrary numbers is not permitted, with the exception of powers of two, since these operations can be implemented as efficient bit shifts without additional computational cost [4]. Therefore, normalisation is reduced to a division by the closest power of two of the standard deviation σpow2. z=x σpow2 However, with this normalisation method, the scales of the input distributions for the 44 features still vary considerably — for example, the distribution of the first jet pTcompared to the third muon pT(see Figure 1). 2.2 Refined Power-of-Two Standard Deviation Scaling Among the input features are the leading objects of jets, taus, and muons. For objects that are not present in a given event, the corresponding inputs are set to 0. Since it is rare for all of these objects to be produced in a single event, a large fraction of the input features are often zero. The high fraction of zero entries leads to very small standard deviations for some features. This is leads to a wide distribution range of muon features and higher leading taus. To address this, a second method calculates the standard deviation σpow2using only the nonzero entries of each feature. Additionally, a minor inaccuracy in the rounding to the next power of two has been corrected, as described in more detail in the Appendix A.1. These adjustments result in a more consistent scale across all input features (see Figure 2). INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 5
CERN openlab Report Figure 1: Input distributions of the first and sixth jet pTand the third muon pTin Monte Carlo events containing two Zbosons. The inputs are normalised by the current power-of-two standard deviation scaling method (see Section 2.1). Figure 2: Input distributions of the first and sixth jet pTand the third muon pTof Zero Bias events. The inputs are normalised by the refined power-of-two standard deviation scaling method (see Section 2.2). INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 6
CERN openlab Report 2.3 Float-Precision Standard Deviation Scaling The third method uses floating-point precision of the standard deviation σof non-zero entries to investigate the effect of rounding and whether it has a significant impact. This method is intended for study purposes only, as division by arbitrary numbers is not feasible under the Level-1 constraints. 2.4 Power-of-Two Standardisation In all above normalisation methods the pToffset is noticeable. Exemplary, the comparison of the first and sixth jet pTis shown in Appendix 2. Its effects on the model are unclear. The fourth method includes mean µsubtraction as part of the standardisation, in addition to dividing by the standard deviation of non-zero entries σpow2. z=x σpow2 −µ σpow2 Dividing the input value xand the mean µindividually is important for the event masking described in Section 3.2, where bit precision plays a critical role. It is not yet clear whether the additional subtraction is feasible in firmware, but if it yields good results, this approach can be investigated further. 3 Evaluating the Impact of Normalisation on Model Performance In this section, the impact of the normalisation method on the performance of the anomaly detection (AD) algorithm is analysed. In unsupervised learning, model evaluation cannot rely on efficiency metrics, since no ground truth labels are available for the dataset. Instead, the performance of the algorithm is assessed indirectly by visualising and analysing its predictions. 3.1 Float-Precision vs Power-of-Two-Precision First, the correlation between the AD scores obtained with float precision and power-of-two precision was examined. The two plots on the left in Figure 3show a moderate correlation between the predictions of the two models trained with different precisions. For comparison, the effect of the random seed is shown on the right, indicating that the impact of precision choice is of a similar order of magnitude. Figure 3: Correlation of AD score between the same model: trained with float precision and with power-of-two precision (left) and trained with two different random seeds (right). INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 7
CERN openlab Report In the following, the analysis focuses only on the refined power-of-two standard deviation scaling (see Section 2.2) and the power-of-two standardisation (see Section 2.4), as these are the only candidates considered for practical implementation. 3.2 AD Model Alterations for Introducing Offset The input data contains zero entries for objects that do not exist in an event. Since the AD algorithm should focus on reconstructing existing objects, a mask is applied in the following steps: •Data and reconstructed data are masked for the calculation of the MSE loss during training and model prediction. •Data are masked for the autoencoder during encoder and decoder training. •Reconstructed data are masked when used as input for discriminator training. In the case of standard deviation scaling, the mask filters out entries that are exactly zero. For full standardisation, which also includes subtracting the feature mean, the mask must instead compare each feature value to its offset. Here, the influence of bit precision becomes relevant. A detailed description of both mask implementations, including their precision settings, is provided in the Appendix A.2. Figure 4: Correlation of AD score between models with 3 different masks: no mask, the old implementation and a new implementation. Figure 4illustrates the effects of the masks on training. The first row corresponds to standard deviation scaling: both the old mask and the new definition with offset = 0 yield INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 8
CERN openlab Report 4 Conclusion In this project, four different normalisation methods were introduced and their impact on model performance for anomaly detection was analysed. The comparison between standard deviation scaling using floating-point precision and using the next power-of-two was brief and not sufficient to demonstrate that the next power of two can reproduce results comparable to floating-point precision. Given the strict deadline for finalising the algorithm, and the fact that floating-point division is not feasible in practice, the focus was shifted towards studying and optimising the methods that are realistically implementable. The investigated metrics are not sufficient to fully understand the behaviour of the algorithm, but the power-of-two definition of the model was able to identify consistent rules across different random seeds. This is an indication that a global minimum was reached in other words consistent rules to identify anomalous events. However, the results also show that anomaly classification heavily prioritises the pTof the leading jet. As a consequence, the seed-independent rule learned by the model is based on a cut on a single feature, while muon features are hardly taken into account. Since L1Topo already provides trigger algorithms that make decisions based on cuts on individual features and particle multiplicities, this strong focus on the leading jet pTis redundant. The actual goal is to develop an algorithm capable of identifying rules in the data that define anomalous events beyond such simple cuts. How many events are classified as anomalous that would not otherwise have been triggered, and what their characteristics are, has not yet been investigated. Therefore, it is difficult to assess how well the algorithm actually performs. Nevertheless, it is necessary to establish a way to balance the contributions of all features more equally to optimise this trigger algorithm. The model using the power-of-two standardisation addresses the issue, treating all jet pTfeatures and τlepton pTfeatures equally. And also muons contribute to the anomaly classification. However, this normalisation introduces new complications. Training stability is low, meaning that the learned rules are strongly influenced by random seeding, and no clear rules corresponding to a global minimum are found. The training behaviour is strongly affected by the offset subtraction of the pTfeatures. The training configuration was optimised for the current model, while this model has to learn more complex rules. Therefore tuning the training duration and learning rate could potentially improve the performance. Another possible reason for the strong variation is that the chosen architecture is too small to compactly encode all relevant event information in the latent space. This limitation of model size could also affect the standard deviation scaling case, but it is less apparent there due to the strong emphasis on the pTcut of the first leading jet. To address this issue, the problem could be reduced to fewer input features, or the model could be supported by techniques such as knowledge distillation. Furthermore, the reconstruction score used for training and the AD score used for event classification lead to different outcomes. The AD score should be redefined based on the latent variables to provide a better approximation so that the algorithm learns and predicts using the same score so that it is able to improve its predictions during training. INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 15
CERN openlab Report A Appendix A.1 Code Correction for Rounding to next Power-of-Two The old implementation computed log2() and rounded to the nearest integer to get the exponent. This way, 2exponent does not always yield the next power of two. In the corrected version the boundary between rounding up and down is raised from 0.5 to 0.584962500721156. Listing 1: Fixed normalisation code. def round_to_pow2 (X): """ Round input to the nearest power of 2. """ exponent = np . floor (no . log2 (X )+0.4150374992788439) X_pow2 = 2 ** exponent return X_pow2 A.2 Code Implementation of Mask with Offset Listing 2: Mask implementention with (mask2) and without (mask1) offset subtraction. mask_1 = tf . keras . backend . cast ( tf. keras . backend . not_equal (data , 0) , keras . backend . floatx ()) mask_2 = tf . keras . backend . cast ( ~ tf . experimental . numpy . isclose ( data , offset , atol =1e -3 , rtol =0) , keras . backend . floatx ()) The above code shows the implementation of the masks for the standard deviation without offset subtraction and the standardisation with offset subtraction. Mask 1 returns true if the corresponding entry in the data is not equal to zero. Mask 2 returns true if |a−b| ≤ atol + rtol ·max(|a|,|b|). The parameters that were used are atol = 10−3and rtol = 0. A.3 Correlation of Standard Deviation Scaling and Standardisation Figure 11: Correlation of the reconstruction score between two models trained with data normalised with standard deviation scaling and standardisation for four different random seeds. A.4 Additional Material for Analysis of the Refined Power-of-Two Standard Deviation Scaling INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 16
CERN openlab Report A.4.1 Anomalous Event Distributions 2 4 6 8 10 12 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Distribution of jet 1 pt All Events 10.75 % non-zero Anomalous Events 99.69 % non-zero 2 4 6 8 10 12 14 0.00 0.05 0.10 0.15 0.20 Distribution of jet 2 pt All Events 4.34 % non-zero Anomalous Events 46.08 % non-zero Figure 12: pTdistribution of the second and third jet for anomalous events compared to all events. The left plot shows that 99%anomalous events contain a second jet, with pTvalues remaining high but with a broader peak compared to the first jet. The right plot shows that anomalous events contain a third jet ten times more frequently than all events, while the pT distribution matches the overall sample. This suggests that the presence of a third jet influences the anomalous classification, while its pTvalue is less important. 2 4 6 8 10 0.0 0.1 0.2 0.3 0.4 Distribution of tau 1 pt All Events 10.38 % non-zero Anomalous Events 70.06 % non-zero 5 10 15 20 25 0.00 0.05 0.10 0.15 0.20 0.25 Distribution of tau 2 pt All Events 3.51 % non-zero Anomalous Events 33.14 % non-zero Figure 13: pTdistribution of second and third τlepton. The trend is similar to the jet distributions: for higher leading τs the high pTpeak broadens and the distribution of anomalous events aligns more with the over all distribution. Hence, the existence of a particle influences the anomalous classification more than its pTvalue. 3210123 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 Distribution of jet 0 eta All Events 25.48 % non-zero Anomalous Events 100.00 % non-zero 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 Distribution of jet 0 phi All Events 25.48 % non-zero Anomalous Events 100.00 % non-zero Figure 14: ηand φdistributions of the first jet for anomalous events compared to all events. The similar shapes of the distributions indicate that no preferred flight direction is associated with anomalous classification. These ηand φdistributions are exemplary for all other particles. INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 17
CERN openlab Report A.4.2 Correlation Plots of pTand Latent Variables and Loss Scores INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 18
CERN openlab Report INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 19
CERN openlab Report A.4.3 Random Seed Dependence Figure 15: Correlation plots of the anomaly detection (AD) score and the reconstruction score of the same model trained with different random seeds. The inputs are normalised by standard deviation scaling. INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 20
CERN openlab Report A.5 Additional Material for Analysis of Power-of-Two Standardisation A.5.1 Anomalous Event Distributions 0246810 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Distribution of jet 1 pt All Events 10.75 % non-zero Anomalous Events 100.00 % non-zero 0246810 0.00 0.05 0.10 0.15 0.20 Distribution of jet 2 pt All Events 4.34 % non-zero Anomalous Events 94.98 % non-zero Figure 16: pTdistribution of the second and third jet for anomalous events compared to all events. The left plot shows that 99%anomalous events contain a second jet, with pTvalues remaining high but with a broader peak compared to the first jet. The right plot shows that anomalous events contain a third jet ten times more frequently than all events, while the pT distribution matches the overall sample. This suggests that the presence of a third jet influences the anomalous classification, while its pTvalue is less important. 0246810 0.0 0.1 0.2 0.3 0.4 Distribution of tau 1 pt All Events 10.38 % non-zero Anomalous Events 99.76 % non-zero 0 5 10 15 20 0.00 0.05 0.10 0.15 0.20 0.25 Distribution of tau 2 pt All Events 3.51 % non-zero Anomalous Events 98.43 % non-zero Figure 17: pTdistribution of second and third τlepton. The trend is similar to the jet distributions: for higher leading τs the high pTpeak broadens and the distribution of anomalous events aligns more with the over all distribution. Hence, the existence of a particle influences the anomalous classification more than its pTvalue. INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 21
CERN openlab Report A.5.2 Correlation Plots of pTand Latent Variables and Loss Scores INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 22
CERN openlab Report INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 23
CERN openlab Report REFERENCES [1] G. Aad et al. “The ATLAS experiment at the CERN Large Hadron Collider: a description of the detector configuration for Run 3”. In: (2024). [2] G. Aad et al. “The ATLAS Trigger System for LHC Run 3 and Trigger performance in 2022”. In: (2024). [3] Sergey Ioffe and Christian Szegedy. “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”. In: (2015). [4] Wikipedia. Arithmetic shift.url:https://en.wikipedia.org/wiki/Arithmetic_shift (visited on 08/27/2025). INPUT NORMALISATION FOR THE ATLAS ANOMALY DETECTION TRIGGER FOR RUN 3 24