scieee AI-readable full text Open interactive document viewer

Design and Hardware-Aware Optimization of a Unified Neural Network for Muon Momentum Reconstruction in the ATLAS HL-LHC Trigger System

Resende, Francisco; Carnesale, Maria; Rojas, Rimsky; Veneziano, Stefano

Abstract

Real-time muon momentum reconstruction for the ATLAS trigger system at the High-Luminosity LHC (HL-LHC) presents a significant challenge due to stringent latency and hardware resource constraints. This work details a complete workflow for developing a deployable, high-performance neural network to replace the traditional parametric algorithm. We introduce a unified model architecture that consolidates 128 distinct reconstruction scenarios into just four models, overcoming data scarcity issues and drastically simplifying hardware implementation. To meet the demands of FPGA-based deployment, we employ a software-hardware co-design strategy centered on High-Granularity Quantization (HGQ). This advanced framework optimizes the model by assigning heterogeneous, trainable bit-widths to individual parameters, guided by the Effective Bit Operations (EBOPs) metric, which strongly correlates with on-chip resource usage. Our final, optimized model demonstrates a significant improvement in momentum resolution over the baseline while adhering to the strict latency and resource budget of the trigger system. This study provides a comprehensive blueprint for integrating sophisticated, hardware-aware machine learning models into real-time high-energy physics applications.

Full text

Design and Hardware-Aware Optimization of a Unified Neural Network for Muon Momentum Reconstruction in the ATLAS HL-LHC Trigger System Francisco Resende Supervisors: Maria Carnesale Rimsky Rojas Stefano Veneziano September 14, 2025 Abstract Real-time muon momentum reconstruction for the ATLAS trigger system at the High-Luminosity LHC (HL-LHC) presents a significant challenge due to stringent latency and hardware resource constraints. This work details a complete workflow for developing a deployable, high-performance neural network to replace the traditional parametric algorithm. We introduce a unified model architecture that consolidates 128 distinct reconstruction scenarios into just four models, overcoming data scarcity issues and drastically simplifying hardware implementation. To meet the demands of FPGA-based deployment, we employ a software-hardware co-design strategy centered on High-Granularity Quantization (HGQ). This advanced framework optimizes the model by assigning heterogeneous, trainable bit-widths to individual parameters, guided by the Effective Bit Operations (EBOPs) metric, which strongly correlates with on-chip resource usage. Our final, optimized model demonstrates a significant improvement in momentum resolution over the baseline while adhering to the strict latency and resource budget of the trigger system. This study provides a comprehensive blueprint for integrating sophisticated, hardware-aware machine learning models into real-time high-energy physics applications . 1 Contents 1 Introduction and Physics Motivation 3 2 Summary of the ATLAS Experiment and Muon Spectrometer 4 2.1 The Muon Spectrometer (MS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.1.1 Structure ....................................... 4 2.1.2 DetectorTechnologies ................................ 5 2.2 The ATLAS Trigger and Data Acquisition (TDAQ) System . . . . . . . . . . . . . . . 5 2.2.1 Field-Programmable Gate Arrays (FPGAs) . . . . . . . . . . . . . . . . . . . . 7 3 Muon Transverse Momentum Reconstruction 8 3.1 The Baseline Parametric Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.1.1 Three-Segment Reconstruction: The Sagitta Method . . . . . . . . . . . . . . . 8 3.1.2 Two-Segment Reconstruction: The Angle Difference Method . . . . . . . . . . 9 3.1.3 Momentum Parameterization . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 4 Neural Network-Based Momentum Reconstruction 9 4.1 Motivation for Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 5 Methodology and Tools 10 5.1 Simulated Dataset and Feature Engineering . . . . . . . . . . . . . . . . . . . . . . . . 10 5.2 Hardware Target: FPGA Implementation . . . . . . . . . . . . . . . . . . . . . . . . . 11 5.3 Hardware-Aware Optimization Framework . . . . . . . . . . . . . . . . . . . . . . . . . 11 5.3.1 Gradient-Based Bit-Width Optimization with HGQ . . . . . . . . . . . . . . . 11 5.3.2 A Refined Resource Metric: Effective Bit Operations (EBOPs) . . . . . . . . . 12 5.3.3 Software-Hardware Co-Design Workflow . . . . . . . . . . . . . . . . . . . . . . 12 6 Model Design and Optimization Strategy 13 6.1 Baseline Implementation and Feasibility Study . . . . . . . . . . . . . . . . . . . . . . 13 6.2 Model Architecture Consolidation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 6.3 Hardware-Aware Co-Design with High-Granularity Quantization . . . . . . . . . . . . 16 6.3.1 Hardware Implementation Constraints and Methodological Refinements . . . . 18 7 Final Hardware Validation and Performance Measurement 20 8 Discussion 22 9 Conclusion and Outlook 23 A GitLab Repository and CI/CD Pipeline 24 B Automated CI/CD Workflow for Model Development and Validation 24 B.1 Environment and Data Preparation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 B.2 Modeling Strategies and Workflows . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 B.2.1 Baseline and Pilot Studies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 B.2.2 The8-ModelStrategy ................................ 24 B.2.3 The4-ModelStrategy ................................ 25 2 1 Introduction and Physics Motivation It is planned that the CERN Large Hadron Collider (LHC) will begin operations in 2030 with an instantaneous luminosity increased by approximately an order of magnitude. To sustain physics output under these conditions, the accelerator and its experiments are undergoing significant upgrades to handle substantially higher proton–proton collision rates. Since many processes of interest in searches for new physics occur with very small cross-sections, the experiments require a highly selective trigger system capable of identifying and recording relevant events in real time. Within this challenging environment, final states containing muons are of paramount importance for a wide range of key physics analyses. The significance of muon signatures is intrinsically linked to the exploration of the mechanism of electroweak symmetry breaking (EWSB). This phenomenon will be probed via precision measurements of the Higgs boson’s couplings, particularly in rare and clean decay channels such as H→ZZ∗→4µand H→µ+µ−. Muons are also central to studies of Vector Boson Scattering processes, for instance pp →jjW±W±, which are crucial for understanding the nature of EWSB. Furthermore, muons provide essential signatures in searches for physics beyond the Standard Model (BSM). In a wide array of theoretical frameworks, the ability to achieve excellent momentum resolution and maintain high selection purity for muons is critical for the precise measurement of rare decays, such as B0 s→µ+µ−, where the transverse momentum (pT) can be as low as 3 GeV. Conversely, it is equally vital for identifying high-pTmuons, with momenta on the order of the TeV scale, which are characteristic signatures in searches for new high-mass resonances [2]. To select these crucial muon signatures, the ATLAS muon trigger plays a central role by identifying tracks with transverse momenta above about 20 GeV. Achieving this requires precise momentum determination, which is based on the hit information from the precision Monitored Drift Tube (MDT) chambers. The current baseline for the High-Luminosity LHC relies on traditional reconstruction algorithms implemented on large Field-Programmable Gate Arrays (FPGAs). However, modern machine learning methods offer the potential to surpass these algorithms. This project investigates the implementation of neural networks for muon transverse momentum estimation in the ATLAS first-level trigger. The study is carried out in two complementary phases. In the first phase, the model’s architecture is developed in Keras and systematically optimized through an extensive hyperparameter search using the Optuna framework to assess the feasibility and performance potential of neural networks for this task in a floating-point, software-based setting. In the second phase, a new hyperparameter optimization campaign is performed within the High-Granularity Quantization (HGQ) framework, where quantization and hardware-related constraints are directly incorporated into the training and optimization process. This enables the discovery of models that are both accurate and efficiently implementable on FPGA hardware through the HLS4ML toolchain. The final quantized networks are deployed and validated on an FPGA-based accelerator card, with their performance—measured in terms of transverse momentum resolution—evaluated against the traditional baseline algorithm. The results demonstrate the advantages of a hardware–software co-design approach for real-time event selection in the challenging conditions of the High-Luminosity LHC. 3 2 Summary of the ATLAS Experiment and Muon Spectrometer The ATLAS experiment at the LHC is a general-purpose particle detector designed to reconstruct particles produced in high-energy collisions. It is composed of several sub-detector systems arranged in layers: •Inner Detector: Precisely tracks charged particles near the collision point [11]. •Calorimeters: Measure the energy of absorbed particles [10]. •Magnet System: Superconducting magnets generate a strong magnetic field to bend the paths of charged particles [12]. •Muon Spectrometer (MS): The outermost system, which identifies and measures the momentum of muons [1]. •TDAQ: A trigger and data acquisition system that selects and records interesting events [13]. 2.1 The Muon Spectrometer (MS) The primary role of the Muon Spectrometer is to identify muons, which are the only charged particles that typically penetrate all the inner layers of ATLAS, and to measure their momentum with high precision. It achieves this using a large superconducting toroidal magnet system that bends the muons’ trajectories. The degree of curvature of the particle’s trajectory is inversely proportional to its momentum. Figure 1: Schematic illustration of the Muon Spectrometer on the left, one specific phi sector on the right 2.1.1 Structure The MS is divided into a barrel region and a endcap redgion, each subdivided into two sides. Each of these is segmented azimuthally into 16 sectors (eight ”large” and eight ”small” sectors) that overlap to ensure complete coverage. 4 2.1.2 Detector Technologies The MS uses two main types of gaseous detectors for its function: 1. Fast-Response Trigger Chambers: These detectors have excellent time resolution (a few nanoseconds), allowing them to quickly identify high-momentum muons and assign them to a specific collision event. This includes Resistive Plate Chambers (RPCs) in the barrel and Thin Gap Chambers (TGCs) in the endcaps [14, 1]. 2. Precision Tracking Chambers: These provide highly accurate position measurements (to less than 0.1 mm) to precisely reconstruct the muon’s curved path. The primary technology for this is the Monitored Drift Tubes (MDTs). In the high-intensity regions of the upgraded detector, new technologies like MicroMegas and sTGCs are also used [14, 1]. 2.2 The ATLAS Trigger and Data Acquisition (TDAQ) System At the HL-LHC, proton bunches, each containing approximately 1011 protons, will collide at the center of the ATLAS detector with a frequency of 40 MHz. In each bunch crossing, the average number of simultaneous proton-proton interactions (pile-up, ⟨µ⟩) is projected to increase from approximately 50 to 200, generating a dense shower of secondary particles that must be registered by the detector [18, 6]. The upgraded ATLAS Trigger and Data Acquisition (TDAQ) system is designed to maintain low pT thresholds for a variety of physics signatures while managing the substantially higher event rates and harsh pile-up conditions. 5 Figure 2: Schematic illustration of the trigger and acquisition system The TDAQ architecture employs a two-level system. The first stage is a hardware-based Level0 (L0) trigger, which processes coarse-granularity data from the calorimeter and muon systems at the full 40 MHz bunch crossing rate. The L0 trigger issues a decision with a maximum latency of 10 µs, reducing the event rate to 1 MHz. Events accepted by the L0 trigger are then forwarded to a software-based Event Filter (EF). In this second stage, full event reconstruction is performed, including high-fidelity track reconstruction using the new all-silicon Inner Tracker (ITk). This enables sophisticated event selection algorithms to reduce the final output rate to 10 kHz for permanent storage and subsequent offline analysis [6]. The MDT Trigger Processor (MDTTP) is a central component of the Phase-II upgrade to the ATLAS muon trigger and readout system [7]. It is designed to incorporate precision coordinate data 6 from the Monitored Drift Tube (MDT) chambers into the hardware-based Level-0 (L0) muon trigger. The primary goal is to improve the transverse momentum (pT) resolution for muon candidates at the hardware level, thereby enabling a high trigger efficiency for signal events while maintaining a low L0 trigger rate under high pile-up conditions [9, 19]. Figure 3: Schematic illustration of the MDT Trigger Processor To achieve this, the MDT front-end electronics will be upgraded to facilitate a continuous, triggerindependent transmission of hits to off-detector systems [2, 9]. Trigger candidates from the Resistive Plate Chambers (RPCs) in the barrel, Thin Gap Chambers (TGCs) in the endcaps, and the New Small Wheels (NSWs) in the forward region will be processed by Sector Logic (SL) boards. These boards provide a Region-of-Interest (RoI) and a bunch-crossing identification (BCID), which serve as a seed for the MDTTP’s track reconstruction algorithm. The MDTTP identifies MDT hits compatible with the RoI and BCID, reconstructs track segments within each MDT chamber, and combines segment parameters to determine the muon’s pTwith high precision [9]. This upgrade is designed to substantially enhance the muon trigger performance, thereby expanding the physics discovery potential of the ATLAS experiment. The part of the MDT Trigger Processor (MDTTP) highlighted within the blue box in Figure 3 corresponds to the module responsible for estimating the muon transverse momentum (pT). The baseline algorithm currently used for this task relies on a parametric model derived from the geometry of reconstructed track segments within the MDT chambers, with the specific parameterization depending on whether two or three track segments are available for a given muon candidate. The neural networks developed in this project aim to surpass the performance of this baseline approach, as will be discussed in more detail in Section 3. A first hardware prototype of the MDTTP has been developed on the ”APOLLO” AdvancedTCA (ATCA) platform, featuring a Command Module equipped with a powerful AMD Xilinx Virtex Ultrascale+ FPGA and high-speed optical transceivers [16, 9, 15]. 2.2.1 Field-Programmable Gate Arrays (FPGAs) Field-Programmable Gate Arrays (FPGAs) are integral to the LHC trigger systems. These are semiconductor devices composed of a matrix of configurable logic blocks (CLBs) interconnected by a programmable routing fabric. Modern FPGAs contain a vast number of logic gates, on-chip Random Access Memory (RAM) blocks, high-speed serializers/deserializers (SerDes), and specialized resources such as digital signal processing (DSP) slices [21]. 7 A primary advantage of FPGAs is their post-fabrication reconfigurability. Unlike ApplicationSpecific Integrated Circuits (ASICs), which have a fixed, unchangeable function, an FPGA can be reprogrammed in the field to implement new logic or adapt to changing requirements. This architecture enables the implementation of complex, parallel digital circuits defined via a Hardware Description Language (HDL), thereby offering a powerful combination of software-like flexibility and the highthroughput performance characteristic of custom hardware. 3 Muon Transverse Momentum Reconstruction 3.1 The Baseline Parametric Algorithm The baseline algorithm for estimating muon transverse momentum (pT) relies on a parametric model derived from the geometry of reconstructed track segments within the MDT chambers. The specific parameterization depends on whether two or three track segments are available for a given muon candidate. 3.1.1 Three-Segment Reconstruction: The Sagitta Method For candidates with three reconstructed segments, the track curvature is quantified by the sagitta, s(Figure 4). In the barrel region, the sagitta is defined as the displacement of the segment in the middle chamber from the straight line connecting the segments in the inner and outer chambers, measured within the magnetic bending plane. In the endcap region, a similar quantity known as the pseudo-sagitta is calculated. It measures the deviation of a segment in the New Small Wheel (NSW) from the line extrapolating between segments in the middle and outer endcap MDT stations. Figure 4: Schematic view of the concept of the sagitta of three segments and of the polar angle difference ∆βbetween two segments [3] 8 further gains. Additionally, the per-sector datasets, while sufficient for minor adjustments, may lack the statistical power to drive more significant improvements without risking overfitting. Despite these possibilities, our findings strongly suggest that the eight-model strategy provides a robust and nearoptimal solution without the added complexity of per-sector fine-tuning. Figure 6: Resolution plots of neural network model (8 base models) scenario and the finetuned version on a specific phi sector. Finally, a unified strategy was developed to further simplify deployment. In this approach, the twostation and three-station cases were combined into a single model per region by introducing padding for the missing inputs in the two-station configuration. This reduced the total number of models to only four, and in practice allowed for one model per FPGA. Eliminating the need for multiplexers or model switching between the twoand three-station cases, this final design maximizes accuracy, ensures sufficient training statistics, and significantly streamlines the hardware implementation. The neural network approaches demonstrated a significant improvement in resolution over the baseline algorithm in most tested scenarios, as observed in figure 5. However, a notable exception was observed for small endcap sectors that utilize information from three segments, as can be observed in figure 7. In these specific cases, the NN models and the baseline exhibited comparable performance, a behavior that is still under investigation. 15 Figure 7: Resolution plots of baseline scenario and the two different neural network approaches - bad case. 6.3 Hardware-Aware Co-Design with High-Granularity Quantization With a consolidated architecture and the feasibility of the approach confirmed, the focus shifted to developing a hardware-aware model using the High-Granularity Quantization (HGQ) framework. Before launching the full-scale HPO, a preliminary study was conducted to empirically map the tradeoff between model accuracy and hardware cost. The HGQ framework controls this balance via the loss function, L=Lbase +β·EBOPs + . . . , where the regularization coefficient βdetermines the penalty for on-chip resource usage. Using a smaller-scale model for the barrel A 3-sector region, we systematically varied βacross a range of values. This process allowed us to directly observe how adjusting the strength of the EBOPs regularization term impacted the final validation loss and the resulting hardware cost. The analysis provided a crucial, data-driven understanding of the model’s sensitivity to quantization, which was essential for defining the strategy for the subsequent large-scale optimization. In figure 8 one can see the impact of this beta parameter, effectively reducing the the amount of EBOPs. One can also observe that the amount of epochs to achieve this decrease to be considerably large. 16 Figure 8: Validation loss and EBOPs against epochs while training. Figure 9 illustrates the estimated FPGA resource consumption for two models: a baseline trained with β= 0 and a hardware-aware model trained with β= 10−8. The comparison highlights a substantial reduction in the required Digital Signal Processors (DSPs) and Look-Up Tables (LUTs) for the β-regularized model. Figure 9: Comparison of estimated FPGA resource consumption. The hardware-aware model trained with β= 10−8(left) shows a significant reduction in DSP and LUT usage compared to the baseline model trained with β= 0 (right). To showcase the performance of the models with different values, the same model was tested on different βvalues ranging in different orders of magnitude from 10−12 to 10−4. One can see the resolution plots for some of these values on figure 10 against the keras architecture with floating point numbers. 17 Figure 10: Resolution plots of the same neural network architecture with different βvalues The Optuna HPO was re-configured to optimize the HGQ-based models. The search space included not only architectural parameters but also the trainable bit-widths for weights and activations, guided by the principles of gradient-based bit-width optimization. After finding the best model, this model was trained with a specific beta value enabling the optimizer to directly minimize the estimated hardware cost on the final candidate model. 6.3.1 Hardware Implementation Constraints and Methodological Refinements To ensure the successful synthesis of our models onto FPGA hardware, several key constraints and methodological refinements were adopted during the development process. A primary hardware constraint was imposed by the hls4ml toolchain, which does not support matrix multiplications with more than 4096 entries. To provide a robust margin for synthesis, all dense layers in our models were proactively constrained to a maximum of 3800 matrix entries. Furthermore, a limitation regarding activation functions was discovered during the initial development phase (the eight-model case). The hyperparameter search space originally included various activation functions, such as LeakyReLU. While these models trained successfully with HGQ’s quantized layers in Keras, a compatibility issue arose during the hardware conversion step: the hls4ml toolchain does not currently support the direct translation of quantized LeakyReLU layers. This finding suggests that while the training framework is capable of utilizing these activations, their hardware deployment is contingent on future hls4ml feature support. Consequently, for the subsequent four-model case, the HPO search space was refined to exclusively use the standard ReLU activation function, which is fully supported for conversion by hls4ml. Although this represents a minor modification to the search space, the comparative analysis of the results is presented with a consistent structure to maintain clarity and facilitate a direct parallel with the unquantized Keras model analysis. It is also worth noting that an alternative path to quantization exists. By replacing the HGQ QDense layers with standard Keras Dense layers, one could leverage hls4ml’s built-in Post-Training 18 Quantization (PTQ) capabilities, offering a different strategy for creating hardware-optimized models. The following resolution plots provide a comparative analysis of the 4-model and 8-model strategies against the baseline scenario. It is important to clarify the conditions of this comparison: the baseline resolution is calculated using floating-point precision, representing an idealized software benchmark rather than its final hardware performance. In contrast, the resolutions shown for the neural network models are derived from their fully quantized versions, offering a realistic projection of their on-chip performance on an FPGA. Furthermore, the models presented in this section were optimized solely with high-granularity quantization, without the inclusion of the EBOPs minimization term in the loss function. Figure 11 illustrates a scenario where both the 4-model and 8-model quantized networks achieve a momentum resolution comparable to the idealized floating-point baseline. In contrast, Figure 12 depicts a case where the models exhibit a slight degradation in resolution relative to the baseline. Figure 11: Resolution plot for barrel region, A side, phi sector 13, 3-station case 19 Figure 12: Resolution plot for barrel region, A side, phi sector 3, 2-station case 7 Final Hardware Validation and Performance Measurement To validate the theoretical resource estimates from EBOPs and measure real-world performance, a selection of the most promising models from the HGQ optimization was targeted for hardware deployment. The models were processed through the hls4ml toolchain to be synthesized into firmware for an AMD/Xilinx Alveo FPGA-based accelerator card. The validation pipeline was designed to ensure bit-level functional equivalence and to measure key hardware performance indicators. The automated workflow proceeded as follows: 1. Golden Reference Generation: First, the software-based HGQ model’s predictions on a test set were generated. This output served as the software golden reference, establishing the definitive baseline for bit-level comparison. 2. Bitstream Synthesis: The model was then converted into a final hardware bitstream (.xclbin) using hls4ml, targeting the Vitis accelerator platform. 3. Three-Level Inference Comparison: A bit-level accuracy check was performed by comparing the outputs from three distinct stages: the Keras software prediction, a bit-accurate HLS Csimulation, and on-card hardware inference using the synthesized bitstream on the FPGA. This empirical testing provides an essential bridge between software simulation and practical hardware implementation. By analyzing the results, we confirmed the efficacy of the HGQ optimization and measured the final, on-chip performance indicators, including the inference latency for a single muon prediction and the actual resource utilization (LUTs, FFs, DSPs, and BRAMs). This process validates the viability of the final models for deployment in a low-latency trigger environment. The hardware implementation of the 4-station model demonstrates exceptional efficiency, as quantified in 20 Figure 13. Even for the model optimized purely for accuracy (β=0), the on-chip resource utilization remains minimal, and the design achieves a processing latency of 80 ns. Figure 13: Resource Consumption Report of the Neural Network Model The resources shown on figure 13 show the total resource consumption for the muon transverse momentum estimation part of the MDTTP trigger. The aggregate on-chip consumption is minimal, utilizing less than 3% of the total available device resources. At the single-sector level, Look-Up Table (LUT) utilization is approximately 12%, with a substantially lower footprint for Digital Signal Processors (DSPs) and Flip-Flops (FFs). A crucial aspect of this workflow is the bit-level validation between the software model’s output and the actual on-card hardware inference. While the goal is bitaccurate equivalence, minor discrepancies can arise from the hardware implementation process. These mismatches typically fall into two categories. The first is a minimal, constant error corresponding to a single-bit difference in the least significant bit (LSB), which is an expected and acceptable artifact of fixed-point arithmetic. The second category consists of slightly larger, variable errors. These deviations often stem from subtle differences between the software simulation and the synthesized hardware, such as the precision of internal accumulators in DSP slices or hardware-friendly approximations of activation functions implemented by the hls4ml toolchain. 21 8 Discussion The results of this study successfully demonstrate that a neural network–based approach can not only match but significantly exceed the performance of the traditional parametric algorithm for muon momentum reconstruction. The key to this success lies in the synergy between strategic model design and advanced hardware-aware optimization. The decision to consolidate the model architecture from 128 highly specialized models to four unified ones was a critical turning point. This strategy directly addressed the problem of data scarcity in fine-grained detector sectors, mitigating underfitting and improving the model’s ability to generalize. Furthermore, the use of padding to handle both twoand three-station cases within a single model architecture proved to be a powerful simplification, enhancing training statistics while simultaneously streamlining the final hardware deployment by removing the need for model multiplexing on the FPGA. The application of the High-Granularity Quantization (HGQ) framework was instrumental in bridging the gap between a high-accuracy software model and a resource-efficient hardware implementation. Unlike conventional methods that apply a uniform bit-width, HGQ’s ability to assign precision on a per-parameter basis allowed the optimization to allocate resources where they were most needed, drastically reducing the model’s footprint. The strong correlation between the EBOPs metric and the final synthesized resource usage validated its role as a reliable proxy, enabling efficient hardwareaware hyperparameter optimization without costly and time-consuming synthesis-in-the-loop cycles. This co-design methodology proved essential for navigating the complex trade-off between physics performance and the stringent latency and resource constraints of the ATLAS trigger system. A notable constraint during this work was imposed by the hls4ml toolchain, which currently limits synthesizable dense layers to a maximum of 4096 matrix entries. This directly constrained the hyperparameter search space, preventing the exploration of wider architectures. However, our findings demonstrate that advanced optimization techniques like HGQ can produce models that are highly resource-efficient even as they approach this size limit. This suggests that significant performance gains may still be achievable. If future versions of hls4ml remove this constraint, or if methods are developed to circumvent it, it would open the door to deploying even more powerful models without necessarily exceeding the hardware budget. While this work was conducted on a highly realistic simulation, the ultimate test will be its performance on real collision data. Potential discrepancies between simulation and reality, such as unmodeled detector effects or variations in background conditions, may require further calibration or domain adaptation techniques to ensure robust performance. Nevertheless, the framework established here provides a solid foundation for such future refinements. 22 9 Conclusion and Outlook This work has presented a comprehensive workflow for the design, optimization, and deployment of a neural network for real-time muon momentum reconstruction in the ATLAS experiment at the HL-LHC. We have demonstrated that a data-driven approach, leveraging the full granularity of detector information, can achieve superior momentum resolution compared to the established baseline algorithm, as shown in Figure 10. The principal achievements of this study are threefold. First, we validated the feasibility of using a neural network to outperform the traditional parametric method. Second, we developed a unified and scalable model architecture that maximizes training statistics by consolidating 128 distinct reconstruction scenarios into just four regional models. Third, through an advanced software-hardware co-design framework based on High-Granularity Quantization, we produced highly compact and efficient models that meet the strict latency and resource constraints of the FPGA-based trigger system. The effectiveness of this hardware-aware optimization was demonstrated by a significant reduction in required DSPs and LUTs, as highlighted in Figure 9, without compromising physics performance. The final, optimized models represent a deployable solution that promises to enhance trigger efficiency and background rejection in the challenging high-pile-up environment of the HL-LHC. While this study provides a complete blueprint for integrating sophisticated machine learning models into real-time physics applications, it also opens several promising avenues for future development: •Validation with Real Data: The immediate next step is to deploy the developed models in a testing environment within the ATLAS TDAQ system. Evaluating performance on real collision data is crucial for validating the simulation-based training and identifying any necessary calibrations to account for discrepancies between simulation and reality. •Exploration of Alternative Model Architectures: While the feed-forward neural network proved effective, other architectures could offer advantages. Investigating Boosted Decision Trees (BDTs) would provide a valuable performance benchmark, while recurrent architectures like Long Short-Term Memory (LSTM) networks could better capture the sequential nature of the detector hit data. •Advanced Feature Engineering: Further research could focus on the input feature representation. This could involve using an autoencoder to learn a more compact and informative latent space representation of the input hits, potentially simplifying the learning task for the primary network. •Refinement of Hardware-Aware Training Strategies: More dynamic strategies for navigating the accuracy-resource trade-off could be explored. This includes implementing a beta annealing schedule, where the regularization strength βis gradually increased throughout training, potentially finding a better point on the accuracy-resource Pareto frontier. 23 A GitLab Repository and CI/CD Pipeline For the purposes of reproducibility and automated model training, a GitLab repository with a Continuous Integration/Continuous Deployment (CI/CD) pipeline was established. The repository contains all the source code for data processing, model training, and hardware validation used in this project. The full repository is available at: https://gitlab.cern.ch/atlas-nextgen/work-package-2.2/ngt2.2pt-estimation-mlbased B Automated CI/CD Workflow for Model Development and Validation To ensure reproducibility and enable a systematic exploration of different modeling strategies, the entire research workflow was automated using a comprehensive Continuous Integration/Continuous Deployment (CI/CD) pipeline on GitLab. The pipeline, defined in the .gitlab-ci.yml file, orchestrates a sequence of stages that handle everything from initial data preparation to the final hardware validation of the trained models. The key stages and workflows are detailed below. B.1 Environment and Data Preparation The pipeline begins with two foundational stages: •Environment Setup: A common job configuration (.base) defines the Docker container image, which includes all necessary software dependencies such as TensorFlow, Keras, hls4ml, and HGQ. It also handles authentication for accessing distributed file systems like EOS. •Data Conversion: The conversion-to-npz-files stage uses GitLab’s parallel:matrix feature to execute a massively parallel job. This stage takes the large, monolithic ROOT file as input and partitions it into 128 distinct .npz files, one for each azimuthal phi-sector of the detector. These smaller files serve as the primary data source for all subsequent training and validation tasks. B.2 Modeling Strategies and Workflows The pipeline is designed to execute and evaluate several distinct modeling strategies in a sequential and organized manner. B.2.1 Baseline and Pilot Studies Before training the main models, two preliminary stages establish context and verify the toolchain: •Baseline Performance: The baseline stage runs the traditional, non-machine learning reconstruction algorithm to establish the performance benchmark that the neural network models must surpass. •Pilot Study: The in the benninging stage serves as a self-contained proof-of-concept. It performs a full end-to-end workflow on a single data slice, including fixed-model training, an β-sweep, and final resolution analysis, validating the entire software stack. B.2.2 The 8-Model Strategy This workflow explores the approach of training eight large, regional models and then specializing them. 24